# Hardware and speed
## What you need
- An Apple Silicon Mac, with macOS 14 or later.
- About 110 GB of free SSD space for the model.
Choose Apple menu → About This Mac to check your chip and memory. The
installer has been tested on macOS 14 and 15; model runs have been tested
on macOS 26. Windows, Linux, and Intel Macs are not supported by this engine.
Slotstream is built for Macs with 16 to 64 GB of memory, where the model
cannot fit. It runs on 96 GB and larger Macs too, where the model fits in
memory, but it is not optimized for them; see
[Who it's for](../README.md#who-its-for).
**Compatibility-tier support is coming soon.** An 8 GB Mac can't run the
current model: even the smallest memory plan needs more
memory than it has, so Slotstream refuses to start instead of swapping. On
other Macs, close memory-heavy apps before running the model.
To check your own Mac without downloading or loading anything, run:
```sh
slotstream doctor
```
## Understanding speed
A **token** is a small piece of text, often part of a word. `tok/s` means
tokens per second. Historical warm-decode results describe replies after the
model's cache has warmed up. The development tests below have their own
request protocol. Prompt-processing measurements appear separately below.
The first reply also needs time to load the model and
process your question. Long conversations take longer to process.
Your chip, SSD, and other running apps affect speed. A memory size alone
isn't enough to predict it.
Sevra Desktop's current development build keeps Original 4-bit in Auto,
based on the comparisons so far. It selects a memory budget when the model
loads. Separate runtime management adjusts the cache as available memory
changes. Quantization, the memory ceiling and runtime adjustment have
independent controls. Quantization never changes during a response.
The results below are reference measurements; speed for your current settings
remains unverified unless a matching configuration has been qualified.
## Results
These reply-generation results were measured on real Macs, using different
releases and settings. The newest development tests and the older release
benchmarks are listed separately; they do not establish a release-to-release
speedup at identical settings.
| Mac | Memory | Reply speed |
|---|---|---|
| **MacBook Pro, M5 Pro (our development Mac), October 5 development tests at a 33 GB budget** | **48 GB** | **20.91–23.69 tok/s** |
| **MacBook Pro, M5 Pro (our development Mac), 0.2.19 at a 22 GB target** | **48 GB** | **15.86 tok/s** |
| Same M5 Pro, 0.2.16 configuration at a 20 GB target | 48 GB | 13.47 tok/s |
| Same M5 Pro, historical 0.2.3 result | 48 GB | ~12 tok/s |
| Mac mini, M2 (base storage) | 16 GB | 1.41 tok/s |
| MacBook Pro, M4 Pro | 24 GB | 5.41 tok/s |
| MacBook Air, M5 | 32 GB | 6.22 tok/s |
| MacBook Pro, M4 Max | 36 GB | 8.41 tok/s |
| MacBook Pro, M3 Max | 64 GB | 12.38 tok/s |
| MacBook Pro, M4 Max | 64 GB | 15.93 tok/s |
| Same M4 Max, model on a 10 Gb/s external SSD | 64 GB | 2.98 tok/s |
| MacBook Pro, M5 Max, 0.2.3, auto (34.6 GB target) | 128 GB | ~21–22 tok/s |
| Same M5 Max, 0.2.3, 48 GB target | 128 GB | ~26.9 tok/s |
| Same M5 Max, 0.2.3, 73 GB target | 128 GB | ~31.5 tok/s |
The historical M5 Pro release result is the 0.2.19 benchmark of the
shipping forecast against the 0.2.18 forecast (1.10x faster, 14.38 to 15.86 tok/s,
identical output). Both arms used smaller prompt passes and disabled prefix
caching, so more of the same budget held experts. These are controlled
benchmark settings, not today's automatic configuration; the 0.2.16 row is the pre-release benchmark of that release's
configuration. The M5 Pro results are from the author; the
others are community reports. The M5 Max rows are outside the target range:
the model fits in memory on that Mac, and engines that keep it resident
report faster replies there.
The 18 GB size still needs reports, and 8 GB Macs don't run the
model. Open the details below for versions, settings, and credits.
### October 5 development tests
On October 5, 2026, the original 4-bit model ran on the same 48 GB M5 Pro
with a 33 GB total memory ceiling, using the Desktop startup recipe. Source commit
`754426d7915d26f007fed2c5ec53d633d0664596` and its exact CI-built executable
are pinned in the [complete run record](../db/sources/runs/2026/10/2026-10-05-practical-desktop-ceiling.md).
The recipe retains its full context allowance, prefix cache, two-token
speculation and corrected expert forecast. Each workload ran three times:
| Workload | Median generation speed |
|---|---:|
| Short prompt | 20.91 tok/s |
| Longer context | 23.69 tok/s |
| Completed coding answer | 21.59 tok/s |
Rates count committed output tokens after the first, divided by the complete
generation timer. The first two workloads use fixed-length replies; the
coding answer runs to completion. All repetitions pass the process-memory
and timing checks with no system swap-ins or swap-outs. The filesystem cache
is uncontrolled. These short tests do not establish sustained-session speed,
a minimum across prompts, or performance on another Mac.
The larger budget holds more experts in RAM and reduces generation reads.
This result uses the original quantization; it does not qualify a smaller
pack or establish a speedup over the historical release benchmark. The
[measurement summary](../db/records/measurements/quantization-screen-2026-10-02.md#actual-desktop-ceiling-october-5)
also compares the research three-bit pack and preserves the lower-budget
results. Desktop's total ceiling is distinct from the independent CLI's
automatic model-ceiling policy; this row does not measure a default CLI run.
Full results and test conditions
| Mac | Memory | SSD | macOS | slotstream | Plan | Warm decode | Long prompt | Reported memory | Reported by |
|---|---|---|---|---|---|---|---|---|---|
| MacBook Pro, M5 Pro | 48 GB | internal, 2 TB | 26.6.2 | 0.2.19 | 22 GB target, two drafts, corrected decode forecast, ~100 experts/layer | 14.38 to 15.86 tok/s with the corrected forecast, arm medians over counted cells from 24 held-out pairs | not measured | not recorded | [@carloslfu](https://github.com/carloslfu), 2026-09-16 |
| MacBook Pro, M5 Pro | 48 GB | internal, 2 TB | 26.6.2 | 0.2.16 candidate | 20 GB target, two drafts, decode lookahead, ~88 experts/layer | 11.79 to 13.47 tok/s with the lookahead, arm medians over eligible runs from 34 held-out pairs | not measured | not recorded | [@carloslfu](https://github.com/carloslfu), 2026-09-13 |
| MacBook Pro, M5 Pro | 48 GB | internal, 2 TB | 26.6 | 0.2.3 | auto: 33 GB target, ~152 experts/layer | ~12 tok/s; 12.8 with `--mtp` at a 28 GB memory target | ~220 tok/s at a 4096-token pass (est.) | 32 GB (estimate) | [@carloslfu](https://github.com/carloslfu), 2026-09-02 |
| Mac mini, M2 | 16 GB | internal, 256 GB | 26.6.2 | 0.2.2 | auto: 10.2 GB target, ~21 experts/layer | **1.41 tok/s** | not measured; `context-check` postdates 0.2.2 | 6.1 GB | [@flol's report](https://github.com/carloslfu/slotstream/issues/5), 2026-09-02 |
| Same Mac mini, M2 | 16 GB | internal, 256 GB | 26.6.2 | 0.2.3 | auto: 10.7 GB target, ~25 experts/layer | **1.48 tok/s** | 11 tok/s for 8192 tokens (12.1 min) | 8.1 GB RSS on the long prompt | [@flol's re-run](https://github.com/carloslfu/slotstream/issues/5#issuecomment-5525389653), 2026-09-03 |
| MacBook Pro, M4 Pro | 24 GB | 512 GB; location not specified | 26.6.2 | 0.2.24 | auto: 15.9 GB target, ~53 experts/layer; no draft head or lookahead at this size before 0.2.25 | **3.57 tok/s**; 3.85 to 3.97 in a later round | 93 tok/s for 8192 tokens (18.0 GB target) | 16.6 GB process peak on the long prompt | [@davidcavazos's report](https://github.com/carloslfu/slotstream/issues/41), 2026-09-25 |
| Same M4 Pro | 24 GB | 512 GB; location not specified | 26.6.2 | 0.2.25 | auto: 17.4 GB target, ~58 experts/layer; draft head and decode lookahead on | **5.41 tok/s**; 5.58 and 4.92 in the same round | 86 tok/s for 8192 tokens (18.0 GB target) | 16.7 GB process peak on the long prompt | [@davidcavazos's re-run](https://github.com/carloslfu/slotstream/issues/41#issuecomment-5850068364), 2026-09-26 |
| MacBook Air, M5 | 32 GB | 1 TB; location not specified | 26.6.2 | 0.2.11 | 22 GB target, ~75 experts/layer planned | **6.22 tok/s** | 126.28 tok/s for 8192 tokens, 2048-token passes | 17.75 GB RSS on the long prompt | [@arczhi's report](https://github.com/carloslfu/slotstream/issues/12), 2026-09-07 |
| MacBook Pro, M4 Max | 36 GB | internal, 1 TB | 26.0.1 | 0.2.22 | auto: 27.1 GB target, ~90 experts/layer | **8.41 tok/s** | 166 tok/s for 8192 tokens (27.1 GB target) | 24.8 GB process peak on the long prompt | [@JohnClarkson's report](https://github.com/carloslfu/slotstream/issues/26), 2026-09-20 |
| MacBook Pro 14", M3 Max | 64 GB | 512 GB; location not specified | 27.0 | 0.2.18 | auto: 48.1 GB target, ~119 experts/layer | **12.38 tok/s** | 213 tok/s for 8192 tokens (34.6 GB target) | 30.1 GB process peak on the long prompt | [@merken's report](https://github.com/carloslfu/slotstream/issues/20), 2026-09-16 |
| MacBook Pro 16", M4 Max | 64 GB | internal, 1 TB | 27.0 | 0.2.22 | auto: 48.1 GB target, ~119 experts/layer | **15.93 tok/s** | 270 tok/s for 8192 tokens (34.6 GB target) | 30.1 GB process peak on the long prompt | [@YenHub's report](https://github.com/carloslfu/slotstream/issues/22), 2026-09-19 |
| Same M4 Max, external SSD | 64 GB | external, 1 TB, USB 3.2 Gen 2 (10 Gb/s) | 27.0 | 0.2.22 | auto: 48.1 GB target, ~119 experts/layer | **2.98 tok/s** | 51 tok/s for 8192 tokens (34.6 GB target) | 30.2 GB process peak on the long prompt | [@YenHub's report](https://github.com/carloslfu/slotstream/issues/23), 2026-09-19 |
| MacBook Pro 16", M5 Max | 128 GB | internal, 2 TB | 26.6.2 | 0.2.3 | auto: 34.6 GB target, ~152 experts/layer | ~21–22 tok/s with speculative decoding | not measured | not measured; server path only | [@waterliu1981's update](https://github.com/carloslfu/slotstream/issues/6#issuecomment-5520489176), 2026-09-03 |
| Same M5 Max | 128 GB | internal, 2 TB | 26.6.2 | 0.2.3 | manual: 48 GB target, ~253 experts/layer | ~26.9 tok/s with speculative decoding | not measured | not measured | same report |
| Same M5 Max | 128 GB | internal, 2 TB | 26.6.2 | 0.2.3 | manual: 73 GB target, ~401–441 experts/layer as reported | ~31.5 tok/s with speculative decoding | not measured | not measured | same report |
The historical memory values retain their original measurement limits. The
M5 Pro figure is a planner estimate, and older reported values do not establish
the kernel lifetime footprint peak added by the reporting correction. These
hardware configurations have not been requalified with the new counter.
The 16 GB M2, 24 GB M4 Pro and 32 GB M5 Air results are below the planner's
estimates, the M4 Pro's 5.41 tok/s on 0.2.25 against its ~8, and the 36 GB
M4 Max's 8.41 tok/s is just under its ~9. The 64 GB M3 Max and
the 64 GB M4 Max on its internal SSD are above their ~11 estimate, and the
128 GB M5 Max result is above its estimate. The planner uses the M5 Pro curve
and doesn't model these differences.
The 64 GB M4 Max separates disk speed from everything else: with the same
plan and release, it decoded 15.93 tok/s from its internal SSD and 2.98 tok/s
from a 10 Gb/s USB drive that read 0.9 GB/s. A 16 GB Mac with a fast SSD would
still help separate disk speed from memory capacity at the small end.
The three 0.2.22 reports ran without the decode-forecast file added in 0.2.19.
Their logs print `no correction at lookahead/tap-correction-attention-rank128-v1.safetensors`.
`slotstream pull` fetches the 37.5 MB file; through 0.2.24, the download
`slotstream run`, `serve` and `launch` offer on first use did not. The 24 GB
M4 Pro's 0.2.25 re-run had it.
The Air's long-prompt test explicitly used a 22 GB target with vision and
speculative decoding off. Its full warm-server command and system load were
not supplied. The M5 Max row uses the reporter's updated results after
moving from 0.2.1 to 0.2.3. Community results have not been independently
rerun by the author.
Full methods, raw reports, and limits are in [MEASUREMENTS.md](../MEASUREMENTS.md):
the M5 Pro throughout, the M2 in C1, the M5 Max in C2, the M5 Air in C3, the
M3 Max in C4, the 64 GB M4 Max in C5, the 36 GB M4 Max in C6 and the 24 GB
M4 Pro in C7.
### What the columns mean
- **Plan**: the target and cache size `slotstream doctor` prints with nothing
else running. Auto sizes down while other apps hold memory, so say what was
open.
- **Warm decode**: tokens per second on the third identical request to a
running server, once the expert cache has warmed up. The first generation
in a fresh process is colder and slower; report it too.
- **Long prompt**: prefill tokens per second from `context-check`, which
reads a synthetic prompt through the real engine with process-budget and
real-headroom safeguards. Keep paging observations with any timing result.
- **Peak**: the process-memory bound reported by `run` and `context-check`,
combining native lifetime physical-footprint and RSS peaks with current
usage; request samples are separate observations. This is measured separately
from the plan's estimate.
## Recent prompt-processing results
Changes shipped in 0.2.23 have separate prompt-processing measurements on the
48 GB M5 Pro at a 10 GB target. Each row has three clean matched pairs and
identical generated token IDs within every pair. Times are medians within
each arm; the percentage is the median of paired reductions, so it need not
equal the percentage calculated from the two displayed medians.
| Workload | Matched control | Median times, control → enabled | Median paired time reduction |
|---|---|---|---|
| 16K inventory prompt, MTP on | Larger-read workspace policy off | 155.22 s → 53.94 s prefill | 65.40% |
| 2K prose follow-up, MTP off | Prefix checkpoints disabled | 30.73 s → 4.42 s request | 85.62% |
The inventory fixture reads 16,387 synthetic tokens with two MTP drafts and
emits 11 tokens before its stop token. Both arms use fused attention; the
control disables fused-workspace accounting. It qualifies the automatic
read policy before release, not the complete change from the prior release.
See the [policy qualification](../db/records/measurements/mtp-prefill-policy-2026-09-21.md).
The prose fixture tests the installed release. It reads 2,090 tokens on its
first request, then 2,092 on the follow-up, reusing 2,048 and emitting the
same 16 capped output tokens in both arms. The table measures only the
follow-up. Only one complete two-request pair passes the timing gates, so the
percentage is not a repeated full-session result. Earlier releases already
had prefix caching; this is its benefit against disabled checkpoints, not an
incremental release gain.
The [published-release audit](../db/records/measurements/published-prompt-speed-audit-2026-09-22.md)
also covers other prompt types, lengths, budgets, MTP settings and disk reuse.
Short requests show no consistent speedup. Paging-affected long-request
comparisons stay excluded from qualified timing claims; the larger-memory
release comparison and kernel-only attribution have too few clean pairs for
a repeated claim. None of these percentages updates the warm reply-speed
ranges or establishes a speedup on another Mac.
### Fresh installed-release first reads and exact repeats
The same development Mac and memory target were measured with ordinary
caching and planner-owned settings. The table separates prompt processing
from the full repeated request, which includes a capped reply:
| Prompt | Eligible first reads / repeats | First-read prefill range | Median repeated request |
|---|---:|---:|---:|
| 2K code | 4 / 3 | 13.86–28.11 s | 2.79 s |
| 2K prose | 3 / 3 | 14.91–26.51 s | 3.22 s |
The prospective desktop load and process-page-in screen passed for the
included observations. Several runs had system swap-ins; the stricter global
no-swap subset is insufficient for a repeated first-read claim. Request
history changed read batching despite an unchanged memory plan. The linked
[measurement](../db/records/measurements/release-prefill-2k-2026-09-22.md)
keeps server-first prompts and later misses separate and preserves every
excluded request. These observations do not establish a new decode headline,
a general ETA correction, or a whole-release speedup.
## Does more memory help?
Within Slotstream, yes: a larger expert cache reduces SSD reads and improves
reply speed. From 96 GB the model fits in memory; Slotstream runs there and
benefits from a larger cache, but that is not the case it is optimized for,
and engines that keep the model resident report faster replies there. See
[Who it's for](../README.md#who-its-for).
The clearest community evidence is
[@waterliu1981's cache sweep](https://github.com/carloslfu/slotstream/issues/6#issuecomment-5520489176)
on the same M5 Max, using Slotstream 0.2.3 with speculative decoding enabled:
| Total-process memory target | Reported warm reply speed |
|---|---|
| 34.6 GB (auto) | ~21–22 tok/s |
| 48 GB (manual) | ~26.9 tok/s |
| 73 GB (manual) | ~31.5 tok/s |
All three runs used the same Mac with 128 GB installed memory. The targets
are decimal GB budgets, not measured process peaks or installed-memory
requirements. This comparison supports a gain from allocating more memory
on that machine; comparing its auto result with the M5 Pro alone would not
isolate the effect of memory.
The manual rows are the reporter's warm-speed summaries, without the repeated
per-run timings supplied for auto. They have not been independently rerun or
remeasured on 0.2.16. The report does not establish a universal scaling curve,
a best automatic target, or a larger qualified context window. See
[memory defaults and overrides](../README.md#why-doesnt-slotstream-use-all-of-my-ram)
to try a larger target while leaving room for macOS and other apps.
## Speed estimates
### Planning ranges
The README's estimates combine the real reports above with the development
Mac's measured configurations and planner curve. They are rough expectations
across hardware and settings, not a fitted scaling model or statistical
confidence intervals. Endpoints are rounded outward to whole tok/s.
| Installed RAM | Estimated warm reply speed | Basis and main inference |
|---|---|---|
| 16–<24 GB | ~1–6 tok/s | The M2 mini reported 1.41 tok/s; the M5 Pro-based 16/18 GB simulations estimate about 3.5 to 5 tok/s before the decode lookahead's gain. The upper end has not been measured on a real Mac in this band. |
| 24–<48 GB | ~5–16 tok/s | The 24 GB M4 Pro reported 5.41 tok/s on 0.2.25, rounded outward to 5; the M5 Pro measured 15.86 tok/s on 0.2.19 at a 22 GB process target, rounded outward to 16. That benchmark used different prompt-workspace and cache settings from today's automatic plan, and the upper end assumes a comparable chip and SSD. The 36 GB M4 Max reported 8.41 tok/s on 0.2.22 and the 32 GB M5 Air 6.22 tok/s on 0.2.11. The M4 Pro's first report, 3.57 tok/s on 0.2.24, ran before 0.2.25 turned on its draft head and decode lookahead; its SSD read cold experts through the engine at 3.7 GB/s, about a third of the development Mac's rate, and other apps held memory during both runs. |
| 48–<96 GB | ~15–27 tok/s | The lower reference rounds down from the 48 GB M5 Pro's 15.86 tok/s on 0.2.19 at a 22 GB target, below the CLI's 33.6 GB automatic target; the newer Desktop-recipe measurement above uses a distinct 33 GB total ceiling and does not re-anchor this estimated band; the 0.2.16 result at a 20 GB target was 13.47 tok/s, and the older ~12 tok/s result remains historical evidence. A 64 GB M4 Max reported 15.93 tok/s on 0.2.22. A 64 GB M3 Max reported 12.38 tok/s on 0.2.18, below this range; a rerun on the current release is pending. The upper end transfers the M5 Max's 26.9 tok/s at a 48 GB process target to a comparable Mac with enough available memory. That run used a 128 GB Mac; it was not a measurement of a 48 GB Mac. |
| 96 GB+ | ~20–32 tok/s | The 128 GB M5 Max reported about 21 to 22 tok/s in auto and 31.5 tok/s at a 73 GB process target. Applying this range to other Macs in the band is an estimate. This row is outside Slotstream's target range: the model fits in memory from 96 GB. |
The 96 GB+ row's lower endpoint allows for the same reporter's roughly 20 tok/s
warm auto runs on 0.2.1; the main results table uses the updated 0.2.3 report.
The upper ends of High and the 96 GB+ row assume an M5 Max-class chip, fast
internal SSD, speculative decoding and manual targets that leave room for macOS and
other apps. A 48 GB process target cannot consume all of a Mac's installed
48 GB; it needs a larger machine. These ranges mix releases, so they are not
predictions for a single current build. No release-speedup multiplier was
applied to community reports.
A slow SSD, older chip, different prompt, draft acceptance or memory pressure
can produce results outside the ranges. The 64 GB M4 Max decoded 2.98 tok/s
with the model on a 10 Gb/s external drive, far below its band. More RAM
helps only when the engine can use it to reduce a bottleneck; the band labels do not establish a causal
speed ranking. In particular, there is no measured performance boundary at
96 GB. The shared context recommendation reflects the current planning
guidance, independently of reply speed.
### Automatic memory plans
The columns were checked against a build of `main` after 0.2.24, which
streams the draft head's experts on smaller caches and runs the decode
lookahead in plain decode. They describe the plans in auto mode, which picks the
context window along with the target and speculative decoding. The draft file
is available and no other apps hold memory. Simulated RAM is in decimal GB; a
Mac's marketed memory capacity can produce a different decimal-GB device
reading and target. Auto picks 32,768 tokens through 32 GB of simulated RAM,
65,536 at 36 GB, 32,768 at 48 GB, 131,072 at 64 GB and 262,144 from 96 GB.
| Simulated RAM (decimal GB) | Automatic memory target | Speculative decoding | Decode lookahead | Automatic context window |
|---|---|---|---|---|
| 8 GB | No plan fits | Not applicable | Not applicable | Not applicable |
| 16 GB | 10 GB | Off | On | 32,768 |
| 18 GB | 11.5 GB | Off | On | 32,768 |
| 24 GB | 16 GB | On, experts streamed | On | 32,768 |
| 32 GB | 22 GB | On | On | 32,768 |
| 36 GB | 25 GB | On | On | 65,536 |
| 48 GB | 33.6 GB | On | On | 32,768 |
| 64 GB | 43.2 GB | On | On | 131,072 |
| 96 or 128 GB | 54.7 GB | On | On | 262,144 |
These are allocation plans, not measured performance tiers. The planner's
M5 Pro-based warm-decode estimates for the rows without speculative decoding
are ~3.5 tok/s at 16 GB and ~5 tok/s at 18 GB of simulated RAM. They leave
out the decode lookahead, which made plain decode 1.05x to 1.11x faster on
the development Mac. At 24 GB the draft head now runs with its experts
streamed. The historical 22 GB benchmark measured 15.86 tok/s on 0.2.19 with two
drafts at about 100 experts per layer (14.38 with the 0.2.18 forecast).
Although its total budget matches the 32 GB simulation, its smaller prompt
passes and disabled prefix cache leave a different expert pool. It does not
measure the current automatic plan. The earlier estimate of about 10 tok/s
came from a two-draft measurement at 76 experts per layer on 0.2.14; the real
M5 Air result above was slower. A shared memory budget does not establish
matching runtime settings or speed.
The 15.86 tok/s result with the corrected forecast at about 100 experts per
layer, and the 13.47 tok/s result of 0.2.16 at about 88, are measured references,
not predictions for larger caches. Larger caches have not been timed with
0.2.19 yet; the M5 Max sweep above
demonstrates gains beyond auto with an earlier release. Do not apply the
development Mac's release speedup to those community figures.
**Auto mode picks the memory target, cache size, speculative decoding and
context window.** It takes the largest window of 32,768, 65,536, 131,072 or
262,144 tokens that keeps speculative decoding, the draft head's resident
experts and the decode lookahead as the 32,768-token plan has them, keeps
one complete conversation of that length for follow-up turns, and adds at
most 10% to the planner's estimate for a typical request of 2,000 prompt
tokens and a 400-token reply. `slotstream doctor
--sim-ram ` shows every candidate and its reason. At 24 GB a 65,536-token
window would add 22%, and at 32 GB it would stream the draft head's experts,
whose reads the estimate does not price.
At 36 GB it adds 9%, as the cache drops from 96 to 75 experts per layer.
At 48 GB, auto keeps the original cache because its size exceeds the measured
decode range: the estimate cannot price the loss, even when it reports little
or no change. At 64 GB, 262,144 would add 18%.
From 64 GB the larger window's memory comes from room the 32,768-token plan
leaves unused, so the cache keeps its size and the target rises above that
plan's 34.6 GB, to 43.2 GB at 64 GB and 54.7 GB from 96 GB. Real available
memory and Metal limits can change these decisions; a busy start applies the
same cache and speed rules before choosing its window. A larger Mac can
therefore receive a smaller window when widening it would sacrifice cache
whose performance benefit is unmeasured. `--max-context N` fixes any
window up to 262,144; on a 32 GB Mac, `--max-context 65536` gives the larger
window with the draft head's experts streamed.
For prompts near 32,768 tokens, the planner estimates about 3 minutes of
prefill from 24 GB and 6.4 minutes at 16 GB; near 65,536 it estimates
about 9 minutes at 24 GB, where the draft head's experts stream and the
prefill pass is smaller, and about 8 minutes from 32 GB. These estimates use the M5 Pro's prefill curve,
not measurements on those memory sizes. The planner's historical prefill
curve has not been recalibrated for the new read policy; the bounded results
above cannot supply a multiplier for every pass size, prompt and context.
Windows above 128,256 tokens have no
calibrated estimate yet. On the development Mac, a full 131,072-token prompt
took 38 minutes to read at a 16 GB target, and its passes slowed as the prompt
grew: the planner's estimates, which ignore position, came within a few
percent of the measured time at 65,536 tokens and fell about a third short of
it past that. Leave room for the reply in the
configured window. Startup, queueing, images and reasoning before visible
answer text add to the user's wait.
The repeated target from 96 GB is the intentional conservative default: the
33 GB base ceiling plus the draft head and the full window's charge. Auto does
not increase its ceiling for the M5 Max's demonstrated larger-cache gains. These simulated plans
describe allocation policy; they do not measure speed. See
[memory defaults and overrides](../README.md#why-doesnt-slotstream-use-all-of-my-ram).
## How to measure
Context is a startup choice. Since 0.2.17 auto picks it for each Mac and
`--max-context` fixes it; 0.2.14 added the feasibility report and request-wait
controls described here.
Use `doctor --json` with the intended `--max-context` and memory policy to inspect the feasible
window before loading. A memory-feasible window does not promise a short wait:
the request-to-first-token budget defaults to 30 minutes, including preparation
and queueing. Setting `--max-prefill-wait 0` disables only that time policy.
Keep the configured window, prompt count and required reply count with each
result. Capacity evidence needs a complete prompt and reply and process memory
within budget, including sampled footprint and lifetime peaks. Record global
paging separately; it does not identify which application caused it. Report MTP and vision separately;
a text-only capacity result does not qualify those modes or answer quality.
Allow about ten minutes once the weights are downloaded. Close other
memory-heavy apps if you want clean speed measurements, and exclude timing
intervals affected by paging. Functional checks can run with other apps open
when the memory and pressure safeguards permit it. Run one model
process at a time.
To share your Mac's results, follow the [measurement steps](TESTING.md#measure-your-mac),
then open a [measurement report](https://github.com/carloslfu/slotstream/issues/new?template=measurement-report.yml).
Allow about ten minutes once the model is downloaded. Reports are credited
to their authors.
The full [engineering notes](ENGINEERING.md#speed) explain prompt-processing
time, memory use, and the methods behind the performance claims.