# Measured facts Every number here was produced by running something on this machine, not estimated. Estimates live in PLAN.md and are replaced from here as they are measured. Dev machine: MacBook Pro, Apple M5 Pro, 48 GB unified, 18 cores, 2 TB SSD (APPLE SSD AP2048Z), macOS 26 (Darwin 25.6.0), Swift 6.3.3, page size 16384 B. ## M0.1 — Metal / memory limits (mlx `device_info`, 2026-08-28) | Quantity | Value | |---|---| | `memory_size` | 51,539,607,552 B (48.0 GiB) | | `max_recommended_working_set_size` | 40,200,896,512 B (**37.4 GiB**) | | `max_buffer_length` | 30,150,672,384 B (**28.1 GiB**) | | architecture | `applegpu_g17s` | | `iogpu.wired_limit_mb` | 0 (default) | Consequences: - Total footprint budget on this Mac is **≤ 37.4 GiB**, not 48. The `pro48` preset (~32 GB) fits with ~5 GB headroom. - `max_buffer_length` caps a **single MLXArray at 28.1 GiB**. The slot pool is 9 separate tensors (gate/up/down × weight/scales/biases), the largest being `gate_proj.weight` at 8/27 of the pool, so a 27 GB pool → 8.0 GiB largest tensor. Not binding here, but a **single-tensor pool layout would have been**. Keep the 9-tensor layout. ## M0.2 — Model ground truth (byte-exact, from safetensors headers) Source: `pipenetwork/Qwen3.8-Flash-Next-MLX-4bit` @ `aa7c790e804b`, 4-bit MLX, group_size 64 for experts. Read via HTTP range requests over 11 shard headers (3,215 tensors) — no full download needed. Index `total_size` = 103.770 GB, matched. | Component | Measured | % | Plan estimate | Verdict | |---|---|---|---|---| | Routed experts | **67.948 GB** | 65.5% | 67.9 GB | ✅ exact | | N-gram / PLE store | **32.000 GB** | 30.8% | 28.8 GB | ⚠️ +11%, structure differed | | misc (norms + lm_head) | 1.255 GB | 1.2% | — | | | Gated DeltaNet | 1.189 GB | 1.1% | — | | | QSA attention | 0.376 GB | 0.4% | — | | | Hyper-connections | 0.367 GB | 0.4% | — | | | embed_tokens | 0.358 GB | 0.3% | 0.72 GB (w/ lm_head) | ✅ | | Shared experts | 0.133 GB | 0.1% | 133 MB | ✅ exact | | Routers | 0.126 GB | 0.1% | 126 MB | ✅ exact | | ple (non-store) | 0.019 GB | 0.0% | — | | | **TOTAL** | **103.770 GB** | | ~102 GB | ✅ within 2% | | **RESIDENT** (all − experts − ngram) | **3.822 GB** | | ~3.3 GB | ⚠️ +16% | Expert record geometry (drives `experts.bin`): | Tensor | Shape | dtype | Bytes/expert | |---|---|---|---| | `gate_proj.weight` | [512, 640, 320] | U32 | 819,200 | | `gate_proj.scales/biases` | [512, 640, 40] | BF16 | 51,200 each | | `up_proj.*` | same as gate | | 921,600 total | | `down_proj.weight` | [512, 2560, 80] | U32 | 819,200 | | `down_proj.scales/biases` | [512, 2560, 10] | BF16 | 51,200 each | | **record total** | | | **2,764,800 B** | ✅ Exactly the plan's number. Padded to 16 KiB → 2,768,896 B (169 pages). Per-layer experts: **1.4156 GB** (plan: 1.42 GB ✅). ### N-gram store — plan was structurally wrong | Property | Plan assumed | **Actual** | |---|---|---| | Rows | 20.0 M | **320,001,536** (128 shards × 2,500,012) | | Row dim | 2560 | **160** (= 2560 / 16 ngram heads) | | Row bytes @4-bit | 1,440 | **100** (80 weight + 10 scales + 10 biases) | | Quant group size | 64 | **32** (shape [rows, 5] over 160 dims) | | Total | 28.8 GB | **32.0 GB** | | Lookups/token | ~16 rows ≈ 23 KB | **16 rows = 1,600 B of data** | Structure (from `qwen4_exp.py`): `ngram_heads = (ngram_size−1) × heads_per_ngram = 2 × 8 = 16`; each head has its own ~20M-entry table sized to a distinct prime (`_nth_prime_after(19_999_999, g+1)`); tables are concatenated and split into 128 shards of 2,500,012 rows. Index = `splitmix64`-derived multipliers XOR-mixed over the token n-gram, mod the head's prime, plus the head's offset. PLE is injected at **layer index 1** (`ple_layer_ids: [2]`, 1-based). **Design implications:** the per-token n-gram payload is 20× smaller than planned (1.6 KB, not 23 KB), but it is **16 scattered ~100 B reads**, and each row's weight/scales/biases live in *three different tensors* — so a naive reader does **48 scattered reads per token**. The repack must interleave each row's three parts into one contiguous 100 B (pad 128 B) record, turning 48 reads into 16, and the row cache should be **page-granular (16 KiB holds ~128 rows)**, not row-granular. ## M0.3 — MLX slot-pool capability (the M3 entry gate) `Tools/slotbench3.py`. Batch = 48 experts (132.7 MB), the plausible per-token miss set. | Pool | Strategy | ms/batch | GB/s | Verdict | |---|---|---|---|---| | 5.66 GB (2,048 slots ≈ 43/layer) | contiguous slice | 1.24 | 107.31 | PASS | | | `slice_update` | 1.34 | 98.74 | PASS | | | **batch scatter** | **2.70** | **49.22** | PASS | | | per-slot assign | 9.19 | 14.44 | PASS | | **27.10 GB (9,800 slots ≈ 204/layer)** | contiguous slice | 1.17 | 113.80 | PASS | | | `slice_update` | 1.18 | 112.72 | PASS | | | **batch scatter** | **1.77** | **74.86** | PASS | | | per-slot assign | 7.79 | 17.03 | PASS | - **In-place confirmed**: peak memory 27.32 GB against a 27.10 GB pool + 0.13 GB staging. No full-pool copy. - Throughput *improves* with pool size (49 → 75 GB/s), confirming cost is proportional to the batch, not the pool. - **Gate passed with ~12× margin**: writing 48 experts costs 1.77 ms vs ~22 ms for the SSD to deliver those bytes at 6 GB/s. Slot writes are not the bottleneck. ⚠️ **Methodology note — a wrong first measurement.** `slotbench.py` (v1) called `mx.eval()` after *every single expert*, which measured per-write GPU sync, not the copy, and reported 1.54 GB/s at 27 GB — an apparent gate failure. Batching writes before one `eval` is what a real engine does; v1's number is an artifact. Retained in-repo as a caution. `gather_qmm` correctness: `mx.gather_qmm` over the pool vs `mx.quantized_matmul` on the same expert → **max abs diff 0.000e+00** (bit-identical). vs dequantize-then-dense → 6.6e-3 relative (expected 4-bit quantization error). ## M0.4 — Swift feasibility (mlx-swift 0.31.6, mlx-swift-lm main) Resolved and inspected. **Far more prior art than the plan assumed:** | Need | Status in Swift | |---|---| | Slot-pool gather | ✅ `MLX.gatherQuantizedMM(x, w, scales:, biases:, rhsIndices:, transpose:, groupSize:, bits:, mode:, sortedIndices:)` — `Ops.swift:1468`, wraps `mlx_gather_qmm` | | MoE block | ✅ `SwitchGLU` / `QuantizedSwitchLinear` in `MLXLMCommon/SwitchLayers.swift`, incl. a `gatherSort`/`scatterUnsort` fast path and a custom Metal unsort kernel | | **Gated DeltaNet** | ✅ `Qwen3NextGatedDeltaNet` + `gatedDeltaUpdate` in `MLXLLM/Models/Qwen3Next.swift`, with `conv1d`, `dt_bias`, `A_log`, and a `decodeConv` fast path | | Attention / norms | ✅ `scaledDotProductAttention`, `rmsNorm` in MLXFast | | Slot writes | ✅ `MLXArray` subscript assignment (`ArrayAt.swift`, `MLXArray+Indexing.swift`) | | **QSA indexer** | ❌ must implement (genuinely novel) | | **Hyper-connections** | ❌ must implement (small: 2 low-rank matmuls + sigmoid) | | **N-gram / PLE** | ❌ must implement (streaming-critical) | **This materially shrinks M3.** The plan's long pole assumed porting GDN from scratch; `Qwen3Next.swift` is a near-drop-in (qwen4_exp splits `in_proj` into qkv/z/b/a where qwen3_next fuses them — mechanical). Remaining novel Swift work is the QSA indexer, hyper-connections, and the PLE path. ### Swift probe — the mechanism actually runs (`swift-probe/`) A Swift executable that allocates a slot pool, preads expert records from disk, and gathers over the pool. Measured on this Mac: | Step | Result | |---|---| | Slot pool alloc (1.42 GB, 512 slots) | 0.52 s | | `gatherQuantizedMM` vs `quantizedMatmul` | **max abs diff 0.0 — bit-identical, PASS** | | Slot batch scatter (48 experts) | 2.67 ms = **49.77 GB/s**, in place | | pread 48 records QD1 → QD16 | 15.70 → 59.94 GB/s (cache-warm store; see M0.5 for cold) | The Swift slot-write figure (49.77 GB/s) independently reproduces the Python measurement (49.22 GB/s at the same 2,048-slot pool) — two languages, same kernel path, agreeing to 1%. **The SlotPool architecture is sound in the target language.** ### ⚠️ Build blocker: mlx-swift's *bundled* Metal shaders require Xcode **Scope note added 2026-08-30, because this section was later misread as banning custom kernels.** What follows is about building the shader library mlx-swift ships with. It says nothing about writing a *new* kernel: `MLXFast.metalKernel` JIT-compiles Metal source at runtime through the Metal framework, `GatedDelta.swift` already uses it as the shipped fast path, and a fresh kernel was verified compiling and running on this CLT-only machine. mlx-swift's own README states: *"SwiftPM (command line) cannot build the Metal shaders so the ultimate build has to be done via Xcode."* Confirmed here — a `swift build -c release` links fine but produces **no metallib**, and every MLX call dies with `Failed to load the default metallib`. This machine has **Command Line Tools only, no Xcode**, so `xcodebuild` is unavailable. **Workaround found and verified**: MLX's loader (`device.cpp:load_colocated_library`) searches, in order, `mlx.metallib` → `Resources/mlx.metallib` → SwiftPM-bundle `default.metallib` → `Resources/default.metallib` → `METAL_PATH`, all relative to the binary's directory. Copying the **prebuilt metallib that ships with the Python `mlx` wheel** (`.venv/lib/python3.12/site-packages/mlx/lib/mlx.metallib`, 182 MB) next to the executable **as `mlx.metallib`** makes everything work. Naming it `default.metallib` does *not* work at that path. Caveat: the borrowed metallib is from Python mlx 0.32.2 while mlx-swift vendors MLX 0.31.1. It worked for every kernel this probe exercised (quantize, gatherQMM, scatter, eval), but a version-skewed metallib is not a shipping strategy. **This is a real, unplanned constraint on M7 packaging** and PLAN.md §4.5's "`make install` from source" — the release build needs Xcode (~15 GB) on the build machine, or a vendored metallib built once and shipped as a package resource. Add it to the risk register and decide before M7. ## M0.5 — Disk (the number the whole IO model rests on) **Methodology matters here and two earlier attempts were wrong.** Final method (`Tools/coldread.c`): every offset read **at most once**, spread across 57 GB of real model shards (≫ the ~35 GB usable page cache), `F_NOCACHE` + `F_RDAHEAD 0`, so neither the page cache nor the SSD's own DRAM can serve a repeat. APPLE SSD AP2048Z (2 TB), never-repeat cold random `pread`: | Record | QD1 | QD4 | QD8 | QD16 | QD32 | |---|---|---|---|---|---| | **expert 2.7648 MB** | **9.46 GB/s** | 16.95 | **17.25** | 17.18 | **17.30 GB/s** | | 16 KiB | 0.27 | — | 2.04 | — | 4.54 | | 4 KiB | 0.08 | — | — | — | 1.11 | Latency floor: 292 µs per 2.76 MB read at QD1 (→160 µs at QD8+); **53.6 µs** for 4 KiB and 60.1 µs for 16 KiB at QD1 — a genuine NVMe latency, which is the independent evidence that these reads reach the device and are not cache hits. **This SSD is ~2.5–3× faster than the plan assumed (5–7 GB/s).** 17.3 GB/s exceeds PCIe 4.0 x4; Apple's controller is integrated into the SoC rather than behind a discrete PCIe link, so it is not PCIe-bound. Treat it as measured-on-this-machine, not as a number every Mac will hit — base-storage MacBook Airs will be far slower, and Stage C on real small Macs must re-measure. **Design consequences (large):** - Expert-sized records are the sweet spot: they reach 55% of peak at **QD1** and saturate by QD8. Small-page IO is 100× worse at QD1 — which is the quantitative case against the page-granular mmap approach and *for* the record layout. - Decode IO cost per token (480 expert-uses × 2.7648 MB × (1−h) ÷ 17.3 GB/s): | h | miss/token | IO ms/token | IO-bound ceiling | |---|---|---|---| | 0.98 | 26.5 MB | 1.5 | 650 tok/s | | 0.90 | 133 MB | 7.7 | 130 tok/s | | 0.50 | 663 MB | 38 | 26 tok/s | | **0.00** | 1,327 MB | **77** | **13 tok/s** | **Even a zero-hit cache sustains ~13 tok/s from IO alone.** The plan's small-Mac tiers (`lite16` 4–9, `edge8` 1–4 tok/s) were far too pessimistic *on the IO axis*; the binding constraint on small machines is memory and compute, not bandwidth. - Dense-sweep prefill: a full 68 GB sweep costs 3.9 s → at 8k-token chunks that is ~2,100 tok/s IO-bound, so prefill will be compute-bound at every useful chunk size. Superseded attempts, retained as method cautions: `Tools/diskbench.c` on a freshly-written 12 GB file reported 28.8 GB/s sequential / 72 GB/s random — impossible, because the file was wholly in the unified buffer cache and **`F_NOCACHE` does not evict pages that are already cached**. `purge` requires sudo (declined). A second run against one 10 GB shard gave 24.95 GB/s at QD32, still inflated by repeat reads within the SSD's own cache. Only the never-repeat numbers above are trustworthy. ## M0.6 — Compute | Measurement | Value | |---|---| | Unified-memory bandwidth (bf16 add, 537 MB arrays) | **235.1 GB/s** | | Single 4-bit `quantized_matmul`, batch 1, 7.37 MB weights | 156.2 µs = **47.2 GB/s** | The batch-1 quantized matmul reaches only 20% of memory bandwidth: at this size the kernel is launch/occupancy-bound, not bandwidth-bound. So the naive "3.375 GB of active weights ÷ 47.2 GB/s = 14 tok/s" extrapolation is **not** a valid compute ceiling — the real model issues many ops per layer with better parallelism (and the MoE path gathers 11 experts per layer at once). The honest compute number has to come from running the model, not from extrapolating one kernel. ## M0.7 — The naive path fails (why slotstream exists) Downloaded the full 4-bit conversion (97 GB on disk, 11 shards + tokenizer) and ran it through stock `mlx_lm.load()` on this 48 GB Mac. **Result: the machine went to 48.8 GB of swap and the process was killed before emitting a single token.** Root cause, found in `mlx_lm/utils.py:load_model`: ```python model.load_weights(list(weights.items()), strict=strict) if not lazy: mx.eval(model.parameters()) # <- materialises all 104 GB ``` `load()` defaults to `lazy=False`. So the out-of-the-box Python path is not merely slow on a 48 GB machine — it is fatal, and it takes the whole machine into heavy swap on the way down (the exact failure mode PLAN.md §3.4 predicted for mmap-and-pray, now observed rather than argued). ### `lazy=True` fixes loading — and then fails for a second, deeper reason With `load(..., lazy=True)`: | Step | Result | |---|---| | `load()` | **0.4 s, 0.00 GB active, 0.00 GB peak** | That is the plan's "first token in seconds, never a full-model load" claim, confirmed on the real 97 GB checkpoint. But the run then **died silently during prefill of a 63-token prompt**, with no traceback — killed while paging. **Why, and why it matters:** lazy mapping defers materialisation, it does not bound it. A 63-token prefill routes to ~all 24,576 expert records (coverage ≈ 1−e^(−10·63/512) ≈ 70% per layer, and near-total across 48 layers), and every touched expert becomes a **live MLX array with no way to release it**. Residency therefore climbs monotonically toward the full 68 GB and the process dies. Nothing in the stock path can evict a touched expert. **This is the precise gap slotstream fills.** Lazy mmap gives you deferred loading; it does not give you a *bounded working set*. The slot pool — fixed capacity plus an eviction policy — is what converts "loads lazily, then dies" into "runs in a chosen footprint forever". The observation also independently confirms §3.3's premise that on-demand caching is the wrong mode for prefill: the dense sweep exists exactly because prefill's expert coverage approaches 100%. Measured IO during that naive page-fault prefill: **21.8 KB mean transfer, ~271 tps, 6.1 MB/s sustained** (an earlier sample caught a burst at 9,012 tps / 148 MB/s). Against the 17.3 GB/s this same SSD delivers on record-sized reads, page-fault-driven streaming runs **two to three orders of magnitude below device capability**. ### M0.8 — The decisive finding: MLX cannot sparsely materialise a mmap'd tensor A 5-token raw prompt should touch only ~9% of experts (≈6.2 GB) and still died. So I measured the two gather paths directly, on lazily-loaded real tensors: | Operation | Data actually needed | **MLX materialised** | Amplification | |---|---|---|---| | `mx.take(ngram_shard, 16 rows)` | 1.3 KB | **200 MB** (the whole shard) | ~150,000× | | `mx.gather_qmm(x, experts, rhs_indices=top-10)` | 27.6 MB | **471.9 MB** (all 512 experts of the layer) | 17× | Scaled up: the expert path materialises **1.4 GB per layer → 68 GB** across 48 layers, and the n-gram path materialises **250 MB per touched shard → up to 32 GB** across 128 shards. Together ≈100 GB. That is the whole checkpoint, and it is why every stock run died regardless of prompt length. **This converts slotstream's central design choice from an optimisation into a requirement.** MLX offers no sparse-materialisation path out of a memory-mapped tensor: any gather or take over a lazily-loaded array evaluates the entire source tensor. Therefore the only way to run this model in bounded memory under MLX is exactly the plan's architecture: 1. `pread` precisely the records needed (2.7648 MB per expert; 16 KiB pages for n-gram rows), 2. place them into a **pre-allocated, bounded, fully-resident** pool, 3. gather over that pool — where every element is already resident, so no materialisation surprise exists. The measurements in §M0.3 confirm each step is fast: `gatherQuantizedMM` over a resident pool is bit-exact, slot fills run at 49–75 GB/s in place, and the SSD feeds records at 17.3 GB/s. **The design is not just viable — under MLX it is the only option, and every component of it has now been measured working in isolation.** Honest scope note: because of this, **no end-to-end generation of the full model was achieved on this 48 GB Mac in this session.** The stock path cannot do it, and the bounded path requires the slot pool that M3/M4 build. What has been proven is that every mechanism the bounded path depends on works, and that nothing simpler will substitute for it. Practical note for the runbook: `mlx_lm.load()` must never be called on this model without `lazy=True`, and `slotstream doctor` should refuse to start a configuration whose resident set exceeds the measured working-set limit rather than letting the OS swap. ## M3/M4 — The Swift engine exists and its correctness is measured (2026-08-28) The full engine was built (`Sources/`): qwen4_exp in Swift over mlx-swift — GDN (vendored `gatedDeltaUpdate`), QSA + indexer, MoE over the SlotPool (`gatherQuantizedMM`), hyper-connections, PLE/n-gram with CPU hashing + row dequant, tokenizer + Jinja chat template (swift-transformers), sampler, prefill/decode loop, Ollama-compatible server, CLI. ### Verification results (all against the Python reference implementation) | Test | Result | |---|---| | Chat template (system+user, non-thinking) | **token-for-token identical** to `transformers.apply_chat_template` | | N-gram row ids (6-token prompt × 16 heads) | **exact match** (splitmix64/prime/XOR/floormod port) | | N-gram row CPU dequant vs `mx.dequantize` | **exact** to printed precision (gs32 4-bit + bf16 rounding) | | Layer 0 (GDN + MoE-over-slot-pool + hyper-conn) | **bit-exact** (max abs 0.00000) | | Layer 1 (adds PLE injection, streamed rows) | **bit-exact** | | Layer 2 (GDN + MoE) | max abs 9.8e-4, RMS-rel 0.13% | | Layer 3 (QSA attention) | max abs 1.0e-2, RMS-rel 2.4% | ### The parity method finding: pin the MLX version or you measure the wrong thing First parity runs compared against Python **mlx 0.32.2** and showed 3–4% everywhere. Rerunning the identical computation under **mlx 0.31.1** (what mlx-swift 0.31.6 vendors): the same `quantized_matmul` differs between 0.31.1 and 0.32.2 by up to 0.5 absolute (0.19% of max) — **kernel changes between MLX versions dominate porting error**. Against version-matched goldens, my first divergence (hyper-connection `down` matmul) became **0.0 — bit-identical**. The residual layer-2/3 drift enters at a few bf16 ulps in one low-rank projection (mlpHC `down`, 558/15,360 elements at ≈3 ulps) with bit-exact inputs — consistent with mlx-swift's vendored MLX commit not being byte-identical to the 0.31.1 wheel for one kernel variant, then amplified by RMS-norm rescaling into the next layer. Since layers 0–1 prove every structural path (streaming MoE, GDN, PLE, hyper-connections, embeddings) bit-exact, this is numerics skew, not porting error. Parity gate adopted: layers 0–1 must be bit-exact; deeper layers RMS-rel ≤ 3e-2 (tracked, with the ulp origin documented). ## M4/M5/M6 — End-to-end results (2026-08-28, this Mac, zero tuning) ### The headline: the full 125B+51B model generates on this 48 GB machine `slotstream run --prompt "Why is the sky blue?" --greedy --experts-per-layer 181` (24 GB cache): | Metric | Cold (first run) | Warm (server, 2nd request) | |---|---|---| | Engine start | 2.3 s (page-cached residents: 1.1 s) | — | | Prefill (18–23 tok) | 1.9–4.6 tok/s | 13.2 tok/s | | **Decode** | **7.8–10.4 tok/s** | **20.0 tok/s** (has not reproduced on 0.1.6 — see the 2026-08-30 re-anchoring) | | Expert hit rate | 0.837 (cold cache) | higher (persistent pool) | | Peak Metal memory | 27.3 GB (181 experts/layer cached) | — | | Output | fully coherent, correct Rayleigh-scattering answer | deterministic across requests | **Golden equivalence (the design's core invariant) passed on the full model**: caching only **30 of 512 experts per layer** (1,446 global slots, 5.9% coverage, hit rate 0.556) produced **byte-identical greedy output** to the 181-per-layer cache — streaming placement provably does not touch the math. That starved run peaked at **7.3 GB** total at **5.6 tok/s**: the lite16 tier already works in emulation, at ~10× the plan's original 4–9 tok/s low-end estimate... and within its band despite the cold cache. ### Ollama API surface `Tools/api_test.sh` (raw-socket transport; this sandbox proxies curl/urllib — external clients on a normal machine are unaffected): `/api/version`, `/api/tags`, `/api/chat` non-streaming + NDJSON streaming, `/api/generate`, `/v1/chat/completions` non-streaming + SSE with `[DONE]`, `/api/embed` clean reject — **all pass**. Instruction following through the whole stack verified ("Reply with exactly: SLOTSTREAM OK" → `SLOTSTREAM OK`; "2+2" → `4`). Engine code: ~2,300 lines of Swift (`Sources/`), single binary + colocated metallib via `make build`. ### The memory planner and its promise (2026-08-28) The default UX is now zero-flag auto-tune (`SlotstreamCore/Plan.swift`), and the constants in it are derived from the measurements above, not chosen: - **Fixed (non-pool) footprint modeled at 3.9 GB** = resident weights 3.822 GB + n-gram row cache ≤0.13 GB. Measured whole-run peaks actually came in at pool + ~3.3 GB (27.3 @ 24.0-GB pool; 7.3 @ 4.0-GB pool) — the model errs ~0.6 GB high on purpose so the announce over-promises memory use, never under-promises. - **`--memory-gb G` derivation**: pool = G − 3.9 − 0.5 margin. Promise test, measured: `--memory-gb 8` → 27 experts/layer, **actual peak 7.0 GB** (predicted 7.5, target 8.0), 5.2 tok/s decode, byte-identical greedy output — now a standing gate in `Tools/verify.sh`. - **Auto policy**: target = min(70% RAM, working set − 2 GB). On this Mac: min(36.1, 38.2) = 36.1 GB → 239/layer (31.7 GB pool) → **actual peak 35.0 GB** (predicted 35.6) under the 40.2 GB working set. Announced at startup and in `/api/show` `details.memory_plan`. - **est. tok/s in the announce** was log-linear between the two measured anchors (30/layer = 5.6, 181/layer = 20.0) and flat above 181 (decode is kernel-launch-bound there, §M0.6) — labeled "est. from M5 Pro anchors" because other machines' SSD/GPU shift the curve. Spot check: the 8-GB-target run's 27/layer estimated ~5.2 and measured 5.22. **This curve was replaced on 2026-08-30** — it over-promised 25 to 45% through its own middle; see the re-anchoring section below for the measured replacement. - **Floor**: 640 global slots (~14/layer, §SlotPool) ⇒ minimum honest target 6.2 GB; below it `--memory-gb` refuses with the arithmetic spelled out. ### The availability clamp (2026-08-28): auto respects what other apps hold Static sizing alone had a first-impression failure mode: a 48 GB Mac with 30 GB already in use would still get a 36 GB target → swap storm. Auto now also reads **currently reclaimable memory** and clamps to it (minus max(1.5 GB, 5% RAM) slack) when that is the binding constraint. **Choosing the "available" definition** (probed live on this machine): | Candidate | Value at probe time | Verdict | |---|---|---| | `kern.memorystatus_level` (= `memory_pressure` "free %") | 88% = 45.4 GB | rejected — counts other apps' compressible/swappable memory as free; sizing a GPU pool against it *causes* the swap storm | | `host_statistics64`: free (raw counter incl. speculative) + purgeable + external file-backed | 33.8 GB | **adopted** — pages reclaimable without compressing or swapping anyone | **Live pressure test** (21.5 GB hog of distinct `bytearray`s over a 64 MB urandom block — page-level incompressible, allocation at memcpy speed): | Phase | Reclaimable | Auto target | Result | |---|---|---|---| | Before | ~29–33 GB | 26.7–34.4 GB (clamp gently binding — this session's apps) | note printed | | Hog up | 13.2 GB | **10.7 GB** (47/layer) | run completed: **9.4 GB actual peak**, 6.4 tok/s cold, coherent output, no thrash | | Hog killed | 36.9 GB | 34.4 GB | sprang back, no restart needed for the *next* process | Semantics: the clamp applies to **auto only** and only at startup; explicit knobs are honored unchanged with an informational "only X GB is reclaimable" note. On a pristine machine the clamp sits above the ceiling and never binds, preserving deterministic sizing. Cross-machine behavior is pinned by seven simulated-setup gates in `Tools/verify.sh` driven by `doctor --sim-ram/--sim-working-set/--sim-available` (48 pristine → 36.0 unclamped; 48 busy → 15.4 + note; 16 pristine → 9.8 unclamped; 16 busy → 6.2 floor + heavy-paging warning; 8 GB → 6.2 floor + too-small warning; 128 GB → fully resident; explicit 30 GB on busy 48 → honored + info note). ### The elastic pool (2026-08-28): serve resizes itself while running A startup-time size can't be right for a daemon's lifetime — the machine's state keeps changing. The governor (`Governor.swift`) resizes the pool between requests, under the engine's generation lock. Mechanics chosen for their memory transients: **grow** gathers contents into the larger tensors one piece at a time (transient ≤ one piece ≈ 1 GB; growth only happens when availability covers it) — cache stays warm; **shrink** frees the old tensors *before* allocating the small ones (transient = max(old,new), never the sum) and restarts cold — under pressure, holding two pools to preserve warmth would spike memory at exactly the wrong moment, and a cold cache refills from SSD in seconds. **Correctness across live resizes** — `slotstream elastic-check`, now a standing verify gate: four greedy generations in one process across 30 → 181 → 30 → 50 experts/layer (grow-with-copy, shrink-cold, regrow) are **byte-identical** (6.4 s / 9.6 s / 7.0 s / 5.4 s). **Live experiment 1 — growing 21.5 GB hog against a running server** (first policy iteration): the poll shed the pool in a cascade 220 → 183 → 145 → 108 → 75 → 37 → 14 experts/layer (29.2 → 1.8 GB) as the hog grew; a request under pressure and a request after recovery were both **byte-identical** to the pre-hog baseline; after the hog exited the pool grew back **with contents kept**, and `/api/show` tracked every step. **Policy learning (why the triggers are absolute GB, not relative):** with the feasibility replan crediting everything a restart would release (pool + fixed footprint — without the fixed credit the equilibrium double-reserves ~4 GB), the honest adjustment under contention is a few GB regardless of pool size — a "shrink at <85% of current" trigger can never fire on a 30 GB pool. Final policy: shrink when desired ≤ current − 1 GB (one-step convergence), grow when desired ≥ current + 2 GB after 60 s of calm and cooldown; pressure events shed absolute chunks (warning ≥2 GB/15%, critical ≥4 GB/50%), repeating until calm. Startup counts as the first resize — launch-time availability can undercount for a minute (page-reclaim lag from a predecessor process), and growing on that transient caused churn until the cooldown covered it. **Live experiment 2 — passive hog, final policy:** equilibrium *held* (no shed): macOS chose to swap the idle hog's pages rather than raise pressure, and keeping the hot pool while the OS pages out idle memory is the correct allocation — the earlier cascade penalized the active workload to protect idle bytes. **Live experiment 3 — ACTIVE 24 GB hog (every page touched continuously for 75 s):** still no OS pressure event on this machine — macOS absorbed the overcommit by compressing/swapping the *idle slots of our own pool* (Metal shared-storage buffers are ordinary pageable VM). Honest status: the pressure-event path is implemented and its arithmetic reviewed, but it has **never been observed firing live** here — the availability poll is the primary actor in practice, events are the backstop for machines/loads that do reach system pressure (`sudo memory_pressure -S` would test it directly but needs root). ### First-run closure (2026-08-28): pull, browser clients, small-Mac stress **`slotstream pull`** (weights acquisition — the missing half of install UX) is implemented and proven against the real network: - Manifest of all 24 files (103.8 GB) with upstream LFS sha256, embedded at build time from the pinned revision (`PinnedModel.swift`) — integrity never depends on a live API. Disk-space check before any bytes move. - `pull --verify` hashed the full local 103.8 GB against upstream in **14 s** (parallel SHA256, ~7.4 GB/s): **24/24 match** — the dev copy is byte-exact provenance-verified. Now a standing verify.sh gate. - Interrupted download **resumed from the exact byte offset** (356 MB into a 10 GB shard, HTTP Range), and survived a live **HTTP 429** rate-limit with backoff-retry from the same offset. - A shard deliberately truncated to 2.0 GB was completed over the network and **passed sha256** — range-stitching is byte-exact. - A part with one flipped byte at offset 1e9 was caught by the hash gate, **deleted, and refused** with a clear message; the rerun re-downloads fresh. - Model names now resolve (`--model` defaults to the pinned name → dev checkout or `~/.slotstream/models`), and missing weights say "run: slotstream pull" instead of a stack trace. **Browser clients (CORS)**: all responses carry `Access-Control-Allow-Origin` and OPTIONS preflight answers with methods/headers/private-network. Preflight verified on the wire; then a **real browser** ran a streaming `/api/chat` from page JS: 5 NDJSON chunks, exact expected content, CORS header present. The cross-origin probe from a public https page was blocked by the test browser's own request filter (`net::ERR_BLOCKED_BY_CLIENT` — an extension-level block, not a server refusal), so that path is unproven here; the case real web GUIs use (localhost page → localhost API) is the one proven. **Small-Mac stress (what is testable without the hardware)**: a 659-token prompt (three prefill chunks) at the absolute floor — 14 experts/layer, 1.8 GB pool — completed with no pin exhaustion and coherent output at **6.1 GB peak**, 3.9 tok/s decode. The same prompt at the pristine-16-GB-Mac auto size (`--memory-gb 9.8`) peaked at **9.6 GB — promise held**, with the long-prefill transient consuming 0.3 GB of the 0.5 GB planning margin (that is what the margin is for). Note: decode after a long prefill runs slower than the short-prompt anchors (3.5 vs ~7 tok/s at 41/layer) — KV/indexer overhead plus a prefill-polluted cache; the est. table is anchored on short prompts. **Soak (bounded)**: `serve --memory-gb 10` ran ~40 min with a request every 45 s — 28/28 requests succeeded, latency flat at 6–7 s, RSS flat across the whole window: **no leak, no drift, no crash**. (The request loop itself paused when the machine slept; the server rode through it.) **RSS finding → MLX cache limit**: that soak surfaced a real hidden footprint — **15.1 GB RSS for a 10 GB-target server**. MLX's allocator retains freed transients (per-request KV caches, activations) in an unbounded internal cache; the Metal "peak" metric doesn't show it, real process memory does. Fix: `GPU.set(cacheLimit: 2 GB)` at engine init. Measured after: **6.0 GB RSS, flat across requests, identical 6–7 s latency** — real process memory now tracks the announced plan instead of exceeding it by 50%. **verify.sh is now memory-adaptive**: the heavy equality gates size themselves to what is reclaimable (181/layer when ≥32 GB, 60/layer otherwise, printed when scaled) — the equality properties are size-independent, and the suite must be runnable on a 16 GB contributor machine, or a busy 48 GB one, without swamping it. Current full battery: **15/15 PASS**. **Operational lesson (learned the hard way):** stacking two slotstream instances plus a full test load (browser, builds, request loops) on one 48 GB machine overcommitted the host and crashed it. The governor protects a single auto-sized instance against the rest of the system; it cannot protect against deliberately stacked model processes. Rule, now in the README: **one instance per machine**; test instances get small explicit `--memory-gb` sizes. ### One-command install (2026-08-28): v0.1.0 release + installer, proven end to end Release v0.1.0 ships `slotstream-arm64.tar.gz` (50.8 MB compressed: the 36.9 MB binary + the 131 MB mlx-0.31.1 metallib, built from commit `6a038fe`) plus a `.sha256` asset; names are stable because the installer fetches `releases/latest/download/`. `install.sh` (repo root) gates on Darwin/arm64/macOS ≥ 14, sha256-verifies the tarball, installs to `~/.slotstream/bin`, wires PATH (a wrapper in `/usr/local/bin` when writable without sudo, else one grep-guarded profile line), and offers a handoff to `serve` when a real terminal exists. `serve`/`run` now offer the download themselves: pinned model missing + a terminal → a size/destination/free-disk block and one `[Y/n]`. Proven live with an APFS-cloned copy missing only `config.json`: consent pulled the missing file over the network, ran the full 24-file verify (PASS), and generated correctly at a 6.9 GB peak. Decline exits 1 with "when you are ready: slotstream pull"; no terminal (piped stdin, no usable `/dev/tty`) exits 1 with the pull hint. `serve` now also prints a copy-paste curl and the client hint at bind time; the exact printed payload returned HTTP 200 NDJSON with the CORS header on the wire. End-to-end as a stranger, from GitHub, under an isolated `$HOME`: the README one-liner installed 0.1.0, appended the PATH line exactly once, printed the no-terminal fallback, and exited 0; under a pty it prompted and honored "n". From the installed directory (outside any checkout), `doctor` initialized Metal off the colocated metallib and a real generation produced the exact requested string at a 6.9 GB peak. Two installer findings: `[ -r /dev/tty ]` passes even with no controlling terminal, so the guard must actually open it (`(exec < /dev/tty)`); and raw.githubusercontent caches for ~5 minutes and ignores query-string cache busting, so after editing `install.sh` wait out the cache before re-testing. Not proven here: a truly clean machine (this Mac's dev checkout resolves first via the embedded path), and the handoff "y" branch was not exec'd live (it composes two proven pieces). **Clean-machine simulation (2026-08-28, follow-up):** with the dev checkout's weights hidden and a fresh `$HOME`, the full stranger chain ran live: one-liner → installer → "y" handoff (exec'd this time) → serve → "y" → the real 104 GB pull streamed (killed deliberately at ~0.6 GB). It exposed two defects, both fixed and re-proven in v0.1.1: (1) the weights presence check was `config.json` alone, and small files download first, so an interrupted first download passed the check and died later in engine load — `serve`/`run` now size-check every manifest file (plus `.part` progress) and the prompt says `have: N GB already here — the download resumes`; proven by resuming the interrupted state at the exact byte offset through the prompt. (2) `FileManager.homeDirectoryForCurrentUser` ignores the `$HOME` environment variable, so redirecting the download (external drive, tests) silently used the passwd home — `ModelLocator` now honors `$HOME` when set. **Per-OS Metal library (2026-08-28, follow-up):** the release tarball's metallib is built for macOS 26, but mlx-metal publishes separate builds for macOS 14, 15, and 26 — shipping the 26 build to older systems is the forward-compatibility direction that can fail. The installer now fetches the build matching the host's macOS from the mlx-metal 0.31.1 wheel (URL and sha256 hardcoded per OS; PyPI files are immutable) on macOS 14 and 15, and keeps the tarball's copy on 26 and later. Tested via an override on this host: the macOS 15 wheel downloaded, hash-verified, and extracted (107 MB vs the 26 build's 131 MB), Metal initialized from it, and a real generation ran at a 7.0 GB peak. A physical macOS 14/15 machine still hasn't run it, but each OS now gets exactly the library a from-source build there would use. Also fixed: re-running the installer says "PATH already set up" and appends nothing (verified one profile line after two runs). **CI-built releases with signed provenance (2026-08-28):** v0.1.0 and v0.1.1 were built on the dev machine and traceable only to a commit hash and checksum in hand-written notes. From v0.1.2, pushing a tag runs `.github/workflows/release.yml` on a GitHub macos-26 runner: newest-Xcode selection (mlx-swift needs Swift 6.3), the pinned-wheel metallib (`SLOTSTREAM_METALLIB_MACOS=26`), a smoke gate that fails the build unless `--version` equals the tag, packaging with sha256, GitHub artifact attestation (verify: `gh attestation verify slotstream-arm64.tar.gz --repo carloslfu/slotstream`), and publish with the commit and build-log URL in the notes. First live run found the macos-15 image's Swift too old for mlx-swift 0.31.6 (tools version 6.3); the macos-26 image with newest Xcode selected is the working recipe. Local asset builds are retired to a documented emergency fallback. **Weights mirror + multi-source pull (2026-08-29):** `pull` previously had a single hard-coded download base, so slotstream's availability depended on a third-party HF repo staying up. It now takes an ordered source list (env override `SLOTSTREAM_WEIGHTS_SOURCES`), tries each in turn, and skips straight to the next source on a permanent HTTP refusal instead of burning retries. Proven: bogus-primary falls back and completes (401 → next source, full verify pass); all-sources-bogus fails closed with "failed from all N source(s)"; default path unchanged. A byte-identical mirror now ships as the primary source: `carloslfu/Qwen3.8-Flash-Next-MLX-4bit` @ `852ebf6f` (README added in a later commit, which is why the pin matters — the pinned revision's files are exactly the manifest's). Upload took seconds rather than hours because HF's content-addressed storage already held these chunks from the upstream repo — the practical argument for mirroring *while* the source is alive rather than after it disappears. Verified three ways: all 24 files present at exact sizes; all 12 LFS sha256s on the mirror equal the pinned upstream hashes (i.e. the whole 103.8 GB is byte-identical, proven without downloading it); and a live `pull` of a hashed file (`tokenizer.json`) from the pinned mirror URL passed the hash gate with a full 24/24 verify. Integrity semantics are unchanged: sources supply bytes, the compiled-in manifest supplies truth. **README correction (2026-08-31).** The README claimed "every file is checked against a hash compiled into the binary". Only 12 of the 24 are: the 11 shards plus `tokenizer.json`, which is 103.783 of the 103.794 GB. The other 12 — configs, `merges.txt`, `vocab.json`, the index, 10.5 MB total, 0.01% — are size-checked and then structurally parsed on load. The user-facing promise (a corrupted download cannot become garbage tokens) survives, but the README now states the split. A full `pull --verify` re-measured at **7.7 s** wall for 103.8 GB (13.4 GB/s, ~6 readers wide at 559% CPU — consistent with the 17.3 GB/s sequential SSD figure and unrelated to the 4.5 GB/s random-`pread` expert path). **Download time, stated honestly (2026-08-31).** The README quoted "35 to 45 minutes on a fast link", which gets the cause wrong: the link is not the constraint. Hugging Face plateaus past four connections, and the two sessions above disagree about where — **50 to 57 MB/s** in the parallel-download work, **36.5 MB/s** in the R2 comparison — so 103.8 GB is a 30-to-47-minute job depending on Hugging Face's day, on a link that does 134 MB/s to a plain host. Past roughly 400 Mbps more bandwidth buys nothing; below it the user's link binds and the wait is ordinary arithmetic (200 Mbps 1h09, 100 2h18, 50 4h36, 25 9h13). The README now carries that table, `doctor` reports whether the disk can hold the weights and the best-case time, and the first-run prompt quotes the same estimate before asking. Both new surfaces were tested against a real 1 GB volume for the refusal path. ### Parallel weight download (2026-08-29): 8 connections, exact resume The pull was one connection streaming one file at a time, which measured 28 to 40 MB/s and put the 103.8 GB download near an hour. Before changing it, the question was where the ceiling actually is. Same 600 MB moved every time, from distinct offsets, to `/dev/null`: | connections | Hugging Face | note | |---|---|---| | 1 | 28 to 40 MB/s | | | 4 | 54 to 57 MB/s | | | 8 | 50 to 55 MB/s | | | 16 | 53 MB/s | | | 32 | 53 MB/s | | The plateau starts at 4 and never moves. Three controls place it: `https://ash-speed.hetzner.com/1GB.bin` gave **27 MB/s on one connection and 144 MB/s on eight**, so the link is not the limit; `hf_xet` 1.29.0, Hugging Face's own fastest client speaking the native xet protocol, downloaded model-00011 (2.19 GB) in 39.4 s, **55.7 MB/s**, the same number; and the reconstruction endpoint shows why, since every xet chunk URL resolves to the same `us.aws.cdn.hf.co` host the plain `resolve/` redirect lands on. The cap is per-IP rather than per-repo: 4 connections to the mirror plus 4 to pipenetwork gave 53.5 MB/s, no better than 8 to either alone, so sharding across mirrors buys nothing. **About 55 MB/s is Hugging Face's number for this client, and parallelism is what reaches it.** The download is now 64 MB chunks from every file in one shared queue, drawn by 8 workers over one URLSession, so connections stay busy across file boundaries and to the last byte. Each incomplete file keeps a `.partmap` beside its `.part`: one byte per chunk, flushed every 2 s after an `fsync` of the data it claims, so a map never promises bytes that are not on disk. Files are hashed and renamed on a background queue the moment their last chunk lands, so verification overlaps the download still running. `--connections N` and `SLOTSTREAM_PULL_CONNECTIONS` override the default. **The bug this found.** The first parallel build ran at 21 MB/s, slower than some single-connection runs, and only 1 of the 6 small files ever renamed. `session` was a `lazy var`, and Swift's lazy initialization is not thread-safe: 8 workers entering at once built several URLSessions, and `taskIdentifier` is only unique within one session, so in-flight state collided and 5 of 8 workers blocked forever on semaphores nothing would signal. 21 MB/s is 3/8 of the plateau, which is exactly the three workers that survived. The session is now built in `init`, and requests are keyed by an id the job assigns rather than URLSession's numbering. **Live, from scratch, against the real mirror:** 12.2 GB in 250 s (**48.8 MB/s** including startup, steady state 50), then killed with `kill -9`. 7 files complete, `model-00001.safetensors` among them: 10.04 GB assembled from 150 chunks fetched out of order across 8 connections, and it renamed, which only happens after its sha256 equals the pinned upstream hash. Resuming reported **91.9 GB to go** and continued `model-00002` from chunk 28 of 150, refetching only the 8 chunks that were in flight when the process died. **Full 24-file proof without spending 104 GB of network:** a Range-capable local server over the existing copy, so the real pull path runs at SSD speed. Pass 1 moved 99 GB in 45 s (**2.47 GB/s**, so the client is nowhere near being the bottleneck at 55) and was killed mid-flight with 22 files complete and two chunk maps at 95/153 and 19/33. Pass 2 computed **4.8 GB to go**, finished, and the full re-hash returned **VERIFY PASS: all 24 files match the pinned revision (103.8 GB)**. Battery 15/15 after the change. Net effect: about **35 minutes instead of about 50** for a first install on this link. The remaining headroom is not reachable on Hugging Face at any connection count; it would take hosting the weights somewhere without that cap, which the link would serve at 144 MB/s. ### Prefill: a bigger pass really is faster, measured at a matched pool (2026-08-30) The pass-size ladder had one anchor and one guess. 2048 was solid — **112.9 tok/s** on an 8,016-token prompt at a 16 GB target, mean of three interleaved runs. 4096 was extrapolated to 125, which is the same mistake the decode curve was making one section down. It cannot be measured at its natural home: a machine that *chooses* 4096 is on a 36 GB target and needs ~33 GB reclaimable, which has not been available. So the two were compared at a matched pool of 60 experts/layer instead, on the same 8,016-token prompt, interleaved: | round | chunk 2048 | chunk 4096 | |---|---|---| | 1 | 96.6 | **108.8** | | 2 | 76.3 | **92.2** | | 3 | 91.4 | **103.9** | | mean | 88.1 | **101.6** | 4096 wins every paired round: **1.15x**. Applied to the 2048 anchor that implies ~130, so the shipped estimate of 125 sits under the evidence rather than over it, and 8192 is credited with no further gain because nothing has measured one. Note what these absolute numbers also show: **prefill depends on pool size, not just chunk size.** The same 2048 pass gives 88 tok/s at 60 experts/layer and 113 at 67, because a bigger cache means fewer expert misses per pass. The estimator is a function of chunk alone, so read it as typical-for-a-machine- that-would-pick-that-chunk, not as a law. ### Warm decode re-anchored, and the live governor finally observed (2026-08-30) Two things had been asserted in the docs for two releases without being re-measured on the shipped build. Both turned out to need correcting. **The decode estimate was 25 to 45% optimistic in the middle of its own range.** The old curve interpolated between 30/layer = 5.6 tok/s and 181/layer = 20.0. Re-measured on 0.1.6 with the pool warm: | experts/layer | measured | old estimate | |---|---|---| | 30 | 6.0 | 5.6 | | 60 | 8.2 | 9.2 | | 120 | **11.2** | **14.8** | | 150 | 11.6 | 17.3 | Three checks make this trustworthy. **It is not a regression from 0.1.6's changes**: 0.1.5 and 0.1.6 were A/B'd interleaved at an identical 60/layer and came out the same (7.4 / 8.0, then 8.1 / 8.2). **It is not under-warming**: fourteen consecutive generations at 120/layer plateau at the second one and hold 11.0 to 11.4, so three samples is enough. And the curve is **nearly flat from 120 to 150**, meaning the plateau starts far below the 181 the old curve assumed. The 20.0 figure at 181/layer could not be re-verified: that config peaks at 27.4 GB against 26.6 GB reclaimable. Forcing it anyway — which is what `--experts-per-layer` is for, and it is never resized by the clamp — drove the machine to 158 MB free and 13 GB of swap, and produced a wide 12.5 to 18.6 band that is not a clean measurement of anything. The estimator now interpolates the verified points and **holds flat above them** rather than extrapolating to a number nobody has reproduced. Under-promising is the right failure direction for a planner. **The elastic governor had never been watched doing its job.** Its *policy* was covered through 19 branches, and `elastic-check` proved the pool can be resized without changing the math, but nothing had exercised the loop that connects them: poll, decide, take the generation lock, resize, update the plan, log. A 5 GB memory hog did not trigger it — correctly, because the replan credits back the pool and fixed footprint, so it still wanted *more* than it had. `slotstream elastic-drill` closes it using the availability seam instead: ``` start: 1615 slots (~34/layer) -> Nile, Amazon, Yangtze elastic: availability dropped — cache ~34 → ~18 experts/layer (4.5 → 2.4 GB pool, cold) squeeze: 862 slots (~18/layer) -> Nile, Amazon, Yangtze cooldown: held at 862 slots, as designed elastic: memory freed — cache ~18 → ~69 experts/layer (2.4 → 9.2 GB pool, contents kept) recover: 3312 slots (~69/layer) -> Nile, Amazon, Yangtze ``` Shrink, cooldown, grow, and **byte-identical output at every size**. One hazard found while building it, now documented on the seam itself: `availabilityOverride` does not make the allocation imaginary. Simulating 60 GB free on a machine with 7 GB made the governor take a real 25.4 GB pool and drove swap from 13 to 39 GB. Anything using that seam must bound the simulated value by `deviceAvailableGB()`, which the drill now does — and it skips rather than fails when the machine is too busy to leave shrink headroom. ### Prefill, second pass (2026-08-30): the cost model was wrong, read-ahead does not work With the prefix cache making follow-up turns cheap, the remaining latency is the per-request floor, and the first job was to find out what it is rather than assume. Splitting the pass into its three parts settled it immediately — an 8,016-token prefill at auto sizing: | part | seconds | |---|---| | reading expert records | 33.91 | | scattering them into the pool | 10.33 | | compute | 50.27 | | **sum** | **94.51** | | **measured pass** | **94.51** | They add up exactly, so the three phases are **fully serialized** — which is what made read-ahead look like a free 36%. It isn't; see below. **Read throughput is 4.5 GB/s, not the 17.3 GB/s the SSD measured, and queue depth is not why.** An expert record is nine separate pieces (gate/up/down x weight/scales/biases), so the pass issues 9N preads of ~307 KB rather than N of 2.76 MB, and the M0 size curve is steep. Swept: QD 12 and 32 tie at ~4.0 to 4.5 GB/s, QD 64 and 128 are *worse*. Nothing here is queue-depth-limited. Making these reads contiguous means an on-disk repack — the M2 container that was skipped by measurement — and this is the first evidence that would un-skip it. **The prefill cost model overcharged by 2x, and it was costing real speed.** It charged `(chunk - 256) x 1.8 MB` because it folded two different things into one term: pass activations, which scale with the chunk, and KV plus indexer state, which scales with the context. Measured separately, with the pool pinned so peak minus the 14.1 GB base is the pass: | chunk | charged | actually measured | |---|---|---| | 1024 | 1.38 GB | 1.30 GB | | 2048 | 3.23 GB | 2.19 GB | | 4096 | 6.91 GB | 4.30 GB | And context state really is small and really is separate: going from a 4,016 to an 8,016-token prompt at fixed chunk moved peak by **0.1 GB**, exactly the ~27.6 KB per token the prefix cache already accounts for. Because a pass was priced at twice its cost, the planner kept choosing 1024 where 2048 is strictly better. Holding total memory fixed and trading pool for pass size on a 4,021-token prompt: | chunk | pool | prefill | decode | peak | |---|---|---|---|---| | 1024 | 77/layer | 65.2 s | 7.3 s | 15.4 GB | | 2048 | 67/layer | **47.9 s** | **6.6 s** | **14.9 GB** | | 4096 | 47/layer | 42.9 s | 9.0 s | 14.4 GB | 2048 dominates 1024 on every axis — faster prefill, faster decode, lower peak. 4096 buys a little more prefill and gives back more decode, so it should only be reached where the pool is already past the decode plateau, which a proportional cap does on its own. The cost model is now the measured 1.30 MB per chunk token and the cap is a quarter of the pool budget instead of a fifth. **Result, three interleaved runs on an 8,016-token prompt at a 16 GB target** (interleaved because single runs here vary by 15% or more, which is enough to invent a result that is not there): | plan | runs | mean | |---|---|---| | chunk 1024, 77/layer (before) | 80.3, 94.8, 82.9 s | **86.0 s — 93.7 tok/s** | | chunk 2048, 67/layer (after) | 65.5, 76.7, 71.8 s | **71.3 s — 112.9 tok/s** | Faster in every paired run, and peak went down, not up. The `--memory-gb` promises still hold with headroom: 8 -> 7.5 GB peak, 12 -> 11.5, 16 -> 15.1 on an 8k prompt. M5's ≥150 tok/s target is still not met; compute is now the majority of the pass and closing that needs a grouped-GEMM kernel, which cannot be built on this machine (mlx-swift Metal shaders need Xcode — see the risk register). **That was wrong, and corrected 2026-08-30**: the Xcode constraint covers building mlx-swift's *bundled* shader library, not writing a new kernel. `MLXFast.metalKernel` JIT-compiles Metal source at runtime, this repo already ships one for gated DeltaNet, and a fresh kernel was verified compiling and running on this CLT-only machine. The grouped-GEMM work is not blocked. **Cross-layer read-ahead was built, measured, and removed.** Since the phases are serialized and a pass above ~512 tokens routes to essentially every expert of every layer, reading layer L+1's whole set while layer L computes is exact rather than speculative, and looked like it should hide most of the 33.9 s. It does not: | | prefetch off | prefetch on | |---|---|---| | run 1 | 64.3 s | 67.7 s | | run 2 | 67.9 s | 78.3 s | Read time barely moved (19.7 -> 19.5 s, 19.4 -> 17.6 s) while compute rose. The reads already saturate with 12 concurrent lanes, so a background reader mostly steals CPU from the thread feeding the GPU, and on unified memory it competes for the same bandwidth. It also cost 1.4 GB of pool to hold a layer's expert set. The implementation is gone; what remains is the timing instrumentation that disproved it. **Do not rebuild this without first making the reads contiguous** — the premise that there is idle IO capacity to overlap into is what measurement refuted. ### Behavioural quality probe (2026-08-30), and what it is not `Tools/quality_probe.sh` runs 15 checkable prompts — factual recall, arithmetic, sorting, instruction obedience, translation, code, cloze — against a live server and requires all of them. It currently passes 15/15. **This is not the FP8 comparison the plan asks for.** That needs an inference credential for Qwen3.8-Flash-Next FP8 (Qwen's own DashScope, or an aggregator carrying it); none is provisioned, and it is paid, so it stays blocked pending a decision. What this probe does is catch *gross* quantization or architecture damage, and give any future re-quantization or kernel change a gate to fail. One item was deliberately removed after it failed: the bat-and-ball question, answered 0.10. That is the classic System-1 trap and a 6B-active model misses it on its own merits, so keeping it would have made the probe flaky in a way that says nothing about the conversion. A damage detector may only contain items the unquantized model reliably gets right. ### The conversation prefix cache (2026-08-29): flat time-to-first-token Every request called `model.makeState()`, so a chat re-prefilled its whole history every turn and paid again for tokens it had already processed. The fix is to keep the state that produced one reply and let the next request extend it when its prompt starts with exactly the ids that state consumed. **What it buys.** Eight turns over HTTP against `--memory-gb 16` (77 experts per layer), each turn adding ~155 tokens of context, temperature 0, measured as time to the first streamed token: | turn | prompt tokens | TTFT cached | TTFT uncached | |---|---|---|---| | 1 | 152 | 7.29 s | 5.68 s | | 2 | 307 | 7.15 s | 8.96 s | | 3 | 462 | 6.03 s | 10.16 s | | 4 | 617 | 6.30 s | 11.65 s | | 5 | 772 | 6.87 s | 11.05 s | | 6 | 927 | 6.50 s | 13.99 s | | 7 | 1082 | 6.04 s | 18.60 s | | 8 | 1237 | **6.01 s** | **25.81 s** | | whole conversation | | **52.2 s** | **105.9 s** | Cached TTFT is **flat** — it does not care how long the conversation is, because only the new tokens are prefilled. Uncached it grows linearly and is already 4.3x worse by turn 8 of a conversation that is only 1,237 tokens long. Turn 1 is slightly *slower* cached, which is the honest cost of storing the state. The ~6 s floor is not prefill. It is the fixed per-request cost of a cold expert cache at this pool size; the prefix cache removes the part that scales, not the part that does not. ### One slot was not enough: what a real client did to the cache (2026-08-30) The cache shipped holding **one** conversation and evicting it on any miss. That kept peak memory provably unchanged, and it passed every synthetic test — `prefix-check` drives a clean three-turn chat and saw reuse on every follow-up. Then it met Open WebUI, and scored **0 hits and 7 misses** across a two-turn chat. The reason is not subtle once you watch the traffic: Open WebUI fires a **title-generation request** immediately after each reply, with a completely different prompt, and follows it with tag and follow-up-suggestion calls. With one slot, the chat's state is evicted by the title request before the user's next turn ever arrives. Every client that decorates a conversation this way — titles, tags, suggestions, embeddings — defeats a single-slot cache the same way. **A one-slot prefix cache is a cache that only works in benchmarks.** The fix is to hold several conversations (four) against one shared token budget, evicting least-recently-used, and to pick the *longest* matching prefix so a follow-up resumes the deepest state available. Re-running the identical Open WebUI chat afterwards: **1 hit**, and the five remaining misses are the auxiliary requests, which are genuinely new prompts and now occupy their own slots instead of destroying the chat's. This costs what the first design avoided: several held states are additive, so the retention ceiling is now **charged against the memory budget** rather than being free. At a 16 GB target that is 0.9 GB, and the pool drops from 67 to 60 experts per layer. That is the honest price of the cache working at all outside a test harness. The lesson worth keeping: **the synthetic gate could not have found this.** `prefix-check` drives the conversation itself, so nothing ever interleaves. Testing against one real client changed the design. ### The equivalence question, and why the answer is not "byte-identical" The obvious gate for this feature is the one every other cache here gets: output must be byte-identical to a cold rebuild. **That gate is not achievable, and the reason is worth writing down, because the first version of this work was specified with it and it took a failing test to find out.** Reusing a state means the same tokens were pushed through the model in a different batching — a prefill of 22, then single-token decode steps, then a prefill of 20, rather than one pass of 69. MLX picks kernels and reduction orders by tensor shape, so a different batching sums the same values in a different order, and floating point is not associative. Swept over a 64-token sequence, **all 63 possible split points** produce different logits from a single pass. (A 28-token sequence happened to be exact for splits 12 through 16, which is coincidence, not alignment — chasing that pattern was a dead end.) So the honest question is not "is it identical" but "is it *more* perturbing than something already accepted as correct". The control is re-chunking a plain prefill, which this project already ships, already gates, and nobody disputes: | sequence | prefix reuse | prefill re-chunk control | |---|---|---| | 28 tokens | 1.99% | 3.42% | | 100 tokens | 4.37% | 4.48% | | 196 tokens | 3.63% | 5.90% | (max |delta| as a fraction of the logit spread, greedy, 30 experts per layer) **Prefix reuse moves the logits less than prefill re-chunking already does**, and the gap does not grow with depth — which is the line between harmless re-association and a state that is accumulating error. That is the gate `slotstream prefix-check` enforces: the delta must stay inside the control's band, it must stay flat with depth, the cached path must be deterministic run to run, and reuse must actually be happening (without that last one the rest is vacuous). A by-product worth recording: the existing "byte-identical output at every prefill chunk size" result is **luckier than it reads**. The underlying logit deltas between chunk sizes are several percent of spread; the text matches because top-1 usually survives that, not because the computation is the same. One of the three probes above already shows top-1 flipping. Anything that demands bit-exactness across re-batching on this model is demanding something the numerics do not offer. **Consequences that shipped with it.** `--no-prefix-cache` pins the old behaviour for reproducibility and is what parity work should use. Retention is capped at a tenth of the pool budget (and never above the context limit, past which reuse is impossible anyway, since a match needs a strictly longer prompt). A cache miss releases the old state *before* the replacement is allocated, so exactly one state is ever live and peak memory is unchanged — what rises is the idle floor, up to ~0.9 GB. The governor drops the retained state before it shrinks the pool further: it is the cheapest thing to give back, costing one re-prefill, where a starved pool costs every token after it. ### Weights hosting: Cloudflare R2 tested and rejected (2026-08-29) The section above ends on an open question — the link has headroom Hugging Face will not give, so *host the weights somewhere without that cap*. Cloudflare R2 is the obvious candidate: zero egress at any volume, ~$1.50/month for 104 GB. It was tested against a real bucket rather than argued about, and **it does not lift the cap**. Method: a throwaway public R2 bucket in the author's own Cloudflare account, a 256 MB random object plus four distinct 64 MB objects, `curl` to `/dev/null`, 8 connections, compared same-session against Hugging Face and against a raw datacenter host (Hetzner Ashburn) standing in for "what the link can actually do". | source | 1 connection | 8 connections | |---|---|---| | Hetzner US-East (raw datacenter, no CDN) | 21.9 MB/s | **133.8 MB/s** | | Hugging Face | 28 to 40 MB/s | 36.5 MB/s | | Cloudflare R2 via `r2.dev` | 28 cold / 36 warm | **50.9** (one object) / **42.4** (four objects) | Three things fall out of this. **R2 ties Hugging Face; it does not beat it.** Both sit in a 36 to 51 MB/s band while the same link sustains 133.8 to a plain datacenter host. Switching hosts buys nothing measurable. **The cap is not per-object.** Spreading 8 connections across four distinct objects — the shape of the real 24-file pull — measured *lower* (42.4) than hammering one (50.9), so no download-client change routes around it. **The one untested path is a custom domain.** Only `r2.dev` could be measured, and Cloudflare documents it as development-only and rate-limited; the unthrottled path is a custom domain on a Cloudflare zone. None of the author's domains are on Cloudflare DNS (Namecheap and Vercel), so testing it requires a zone migration first. That is the only remaining reason to think R2 might still win. Two by-products worth keeping: - **Upstream from this location is 5.0 MB/s** (measured pushing the 256 MB object). Populating a 104 GB bucket from here would take about 5.8 hours, so any future migration has to be driven from a machine with a real uplink, not this one. - **Both CDN paths cap in the same band while a raw host does 133.8.** That may be ISP shaping of CDN traffic, or edge-side per-client shaping, or Popayán peering. It is one link in Colombia, so it is *not* evidence about what users elsewhere see from either host — which is also why hosting should not be re-litigated on this measurement alone. **Decision: stay on Hugging Face.** It is free, already mirrored under the author's own account, already covered by the ordered fallback list and the compiled-in sha256 manifest, and measurably no slower. Revisit only on real reports of slow downloads from other geographies, and then by testing an R2 custom domain. ### Adversarial review of the serving layer (2026-08-29): 0.1.5 An end-to-end adversarial pass over the whole system. The numerical core came through clean; the serving layer did not. Everything below was reproduced against a running server before it was fixed, and each case is now a gate in `Tools/api_robustness.sh`. **Held up under attack.** The acceptance battery genuinely passes (re-run, not taken on trust). Weight provenance over all 103.8 GB, layer parity against the Python reference, cache-size and live-resize byte-equality, and the `--memory-gb` promise all hold. `MoELayer` matches the reference block line for line. The slot-pool floor survives full 256-token prefill chunks. The planner refuses or clamps every out-of-range knob and explains itself. Parallel download and resume are exact: after `kill -9` mid-transfer the chunk map claimed 1,073.7 MB against 1,369.1 MB actually on disk — under-promising by the in-flight partials, which is the safe direction — and the resume continued from those chunks rather than restarting. Multi-source fallback works (HF answers 401 for a missing repo, the file moves to the next base). Preallocation is genuinely sparse: 103.8 GB apparent, 3.3 MB on disk. **Three inputs killed the process.** No SIGPIPE disposition was ever set, so a client vanishing mid-stream terminated the server with signal 13 — which also meant the `alive` flag threaded through every streaming handler was dead code, since `write` could never return `-1`. `"seed": -1` trapped in `UInt64(v)`, and `"num_predict": -1` trapped forming `0 ..< -1`; both are Ollama's *documented* defaults for "random seed" and "generate until EOS". Fixed by ignoring SIGPIPE (plus `SO_NOSIGPIPE` per socket), and by clamping every sampling knob in one place, `SampleParams.sanitized()`. **Streaming silently corrupted some responses.** A reply beginning with a character whose UTF-8 spans several tokens lost its opening: asking for five emoji returned `🚀🔥⭐❤️🌳` unstreamed and `⭐❤` streamed, on all four streaming surfaces. The incremental detokenizer cleared its token list whenever nothing had been emitted yet — exactly the state a leading emoji is in while it waits for the token that completes it. The first fix exposed a second, subtler bug: diffing decoded text by `Character` drops a scalar that merges into the grapheme cluster already sent, so `❤️` streamed as `❤` (the U+FE0F variation selector vanished). The diff is now scalar-exact. **The sampler had an unguarded 0/0.** Any filter that empties the candidate set — `top_p` at or below 0, `min_p` above 1 — made `probs / probs.sum()` produce NaN, after which the server emitted token 0 forever (`!!!!!!`). The normalization is gone: the uniform draw is scaled by the unnormalized CDF total instead, which removes the division and, because `u < 1`, also removes the float-tail case where every CDF entry compared below `u` and the pick ran off the end onto a zero-probability token. **Two limits were missing.** There was no context cap at all, and the memory plan does not model KV growth: a 7,960-token prompt peaked at 8.3 GB against a 7.9 GB plan, and under `--memory-gb 8` it reached 7.9 GB, consuming almost the entire 0.5 GB planning margin. KV plus indexer state costs **27.0 KiB per token** (12 QSA layers x [2 x 2 heads x 256 dims x 2 B] + 12 x 128 x 2 B), so 32k tokens would add 0.91 GB and break the promise outright. Prompts are now capped at 32,768 tokens (`--max-context`) and refused with a 400. Separately, the connection handler had no read timeout and one thread per connection: 300 idle sockets produced 314 threads. Reads now time out, connections are capped at 64, and request bodies are bounded like headers already were. **QSA indexer, previously untested past its budget.** A 7,960-token prompt with the answer planted in the first sentence retrieved it correctly, so the sparse path is exercised end to end for the first time. It also priced the naive prefill honestly: **30 to 32 tok/s**, i.e. about four minutes to first token at 8k. That, not memory, was what made long contexts impractical — see the next section, which sized the prefill pass from the memory plan and took it to 92 tok/s (about 90 s at 8k). **The load path, found by asking "is anything missing?".** The first pass concentrated on the request path and waved the checkpoint reader through because the weights are hash-verified. That was the wrong test: `--model` accepts any directory. Pointing it at a directory with no safetensors, with a corrupt header, or with another model's tensors each trapped (exit 133) after printing the memory plan. `serve` on a port already in use — running it twice, the single most likely operator mistake — loaded the entire model and *then* hit `fatalError("bind failed")`; it now fails in 0.03 s with a sentence naming the fix, because the port is claimed before the model loads. `--max-context 0` was validated after the load too. `posix_memalign` failing in the expert read path was a force-unwrap. And `pull --verify` used `attributesOfItem`, which does not follow symlinks, so a model directory of symlinked weights reported all 24 files corrupt and sent the user into a 104 GB re-download. Each of these is now a gate in `Tools/planner_gates.sh`, which needs no weights. **Smaller things corrected.** `/api/version` reported 0.1.0 from a hard-coded string (now one constant, checked against the binary in CI); `/api/tags` reported a weights size 23.3 MB off the manifest (now read from it); the planner's floor was described as 14 experts/layer in every user-facing message while being 13.3 (640/48), with one `doctor` screen printing both; `Geometry` carried a comment claiming it was "validated against config.json at engine init" when no such check existed (now `Geometry.check` runs in `Qwen4ExpModel.validate`); the n-gram row-cache counters were never reset, so they accumulated across every request in a `serve` process; `IndexerCache` re-concatenated the whole cache each token where `KVCache` next to it grows in blocks; the safetensors header length was read with an alignment-requiring `load`; the GatedDelta kernel guarded `Dk % 32` but not the `Hv % Hk` ratio it also assumes; `SS_DEBUG_LAYER` was read from `ProcessInfo` 48 times per token; and in `pull`, the periodic map flush could `fsync` and rewrite the map of a file whose descriptor the finishing path had already closed. Also fixed for compatibility, each now a gate: `stop` sequences were accepted and ignored, OpenAI array-form message content was silently dropped, an empty prompt sampled a first token from an uninitialized tensor, malformed JSON returned 500 "chat template failed", `HEAD` returned a body, `/api/generate` omitted `prompt_eval_duration`, and `done_reason`/`finish_reason` was always "stop" even when the run hit the token limit. **Deliberately not changed.** The presence penalty is applied before the temperature division, which matches HuggingFace's processor order; the consequence is that API `temperature: 0` is argmax over *penalized* logits while CLI `--greedy` zeroes the penalty, so the two are not identical by construction. Attempts to make them diverge on real prompts did not succeed. Unknown model names are still accepted rather than 404'd: one model exists and its name is advertised in `/api/tags`, so leniency costs nothing. ### Closing the three deferred gaps (2026-08-29): 0.1.5 **Prefill: 40 -> 92 tok/s, byte-identical output.** Prefill is expert-stream- bound, not compute-bound: one pass activates nearly every expert of every layer, so the whole 68 GB expert set is re-read roughly once per pass and halving the number of passes halves the bytes moved. Measured on a 7,960-token prompt at 30 experts/layer cached: | tokens per pass | prefill | MLX peak | |---|---|---| | 256 (old default) | 40.0 tok/s | 8.3 GB | | 512 | 49.6 | 8.6 | | 1024 | 67.1 | 9.2 | | 2048 | 91.8 to 104.8 | 10.1 | Greedy output is byte-identical across all four, checked at 2,980 and 7,960 tokens with the sparse indexer active — the indexer's block boundaries are absolute positions, so chunking cannot move them. The pass size is therefore sized from the memory plan rather than fixed: at most a fifth of the pool budget, which gives 256 at the floor, 512 at an 8 GB target, 1024 at 24 GB and 2048 from 32 GB up. An 8k prompt on this machine went from 199 s to 87 s. Budgeting it exposed the KV gap concretely. Charging only the pass's own activations (1.1 MB/token, which is what they measure in isolation) left `--memory-gb 12` peaking at **exactly 12.0 GB** on an 8k prompt — the promise held with no headroom at all, because the long context that motivates a big pass also carries ~27 KiB/token of KV and indexer state the pool math never modelled. Charging the two together at 1.8 MB/token is what they actually cost: `--memory-gb 12` now peaks at 11.4 GB and `--memory-gb 8` at 7.7 GB on the same prompt, the latter *better* than the 7.9 GB it used to reach while also prefilling 50 tok/s instead of 30. The long-prompt case is now its own gate; the previous one used a 6-token prompt and could not see any of this. **Sampler: golden, and it matched first try.** `Tools/sampler_ref.py` reimplements the sampler in numpy float32. Both sides build their logits from the same splitmix64 stream using only exactly representable float operations, so the comparison is exact rather than approximate. 14 configurations agree token for token — greedy, pure sampling, top-k 1, tight nucleus, min-p, presence penalty with accumulation, the out-of-range values the sanitizer clamps, seed 0, and the real 248,320-entry vocabulary. This closes "sampler implemented but not golden-tested"; the sampler was extracted into its own `Sampler` struct so it runs on synthetic logits with no checkpoint loaded, which also makes it a CI gate. **Governor: policy tested without a memory hog.** The gap was that only the resize *mechanism* was proven (`elastic-check`, byte-identical output across grow/shrink); the *policy* had only the one-off 21.5 GB-hog observation. Rather than repeat that — unsafe on a shared machine and unrepeatable — the decision was extracted into `GovernorPolicy.decide`, a pure function of (current size, availability, recent history). `slotstream governor-check` drives all 19 branches deterministically with no model loaded: shrink, grow, both dead-bands, both cooldowns, warning and critical pressure, repeated pressure converging to the floor, and the cap. Writing it surfaced a genuine subtlety. A test that held availability fixed while shrinking the pool showed the governor ratcheting down step after step instead of converging. That state is unreachable — freeing pool memory raises what is reclaimable by exactly that amount — and the credit in `desiredSlots` (`available + pool + fixed footprint`) exists precisely so the answer does not depend on how much is held at the moment. Modelled correctly it converges in one step, and the invariant is now asserted directly: two states holding the same total memory must want the same size. ### Honest gaps (not yet done) Dense-sweep prefill and cross-token prefetch. Sizing the pass from the memory plan took prefill from 40 to 92 tok/s (8k prompt: 199 s to 87 s), but the sweep is still naive and prefill remains the slow axis for long prompts; decode after a long context also runs below the short-prompt anchors (5.0 tok/s at 30/layer after 7,960 tokens). Real ≤16 GB hardware validation (floor and 16-GB-target behavior emulated and stress-tested here, but never run on a physically smaller, slower-SSD Mac). A hosted web GUI (e.g. Open WebUI) driven end to end (browser streaming client proven; the full product not installed here). LaunchAgent `install` (foreground `serve` is the supported mode). The planner now charges KV and indexer growth through the prefill-pass budget rather than modelling context length directly, so a prompt far longer than the 8k the coefficient was fitted to is still bounded by `--max-context` rather than predicted. The governor's policy is tested as a pure function, which is stronger than the one-off hog measurement but is not the same as observing the daemon under live pressure. Closed 2026-08-29 (see the two sections above): the QSA indexer past its 2048-token budget, the sampler golden, and the governor policy. ## Reference implementation `Tools/reference/qwen4_exp.py` (vendored, from the pinned conversion) — the port oracle. Confirms: PLE at layer 1; QSA indexer returns `None` when `kv_len ≤ 2048` (so **dense attention is exact only up to the 2048-token budget** — the review-pass correction was right); GDN state fp32; router in full precision (`quant_predicate` excludes `mlp.gate`); MTP and vision tower dropped by `sanitize`. mlx-lm 0.31.3 already provides `gated_delta_update`, `SwitchGLU`, `ArraysCache`; `qwen4_exp` itself is **not** in mlx-lm 0.31.3 (confirms the open-PR status). --- ## Summary — what M0 settled **Verified correct in the plan** (no change needed): expert record geometry (2,764,800 B exactly), routed-expert total (67.948 GB), per-layer expert block (1.4156 GB), shared experts (133 MB), routers (126 MB), total checkpoint (~104 GB), 16 KiB record padding arithmetic, and the review-pass correction that QSA's dense path is exact only up to the 2048-token indexer budget. **Corrected by measurement**: n-gram store structure (320 M rows × 160 dims × 100 B = 32.0 GB, group size 32 — not 20 M × 2560 × 1440 B = 28.8 GB) and its per-token cost (1.6 KB, not 23 KB); resident floor (3.822 GB, not ~3.3); SSD throughput (17.3 GB/s, not 5–7); the memory ceiling that actually binds (Metal working set 37.4 GiB, not 48 GB of RAM); and the low-end decode estimates, which were pessimistic on the IO axis by roughly an order of magnitude. **Discovered, unplanned**: (1) MLX cannot sparsely materialise a memory-mapped tensor, which makes the bounded slot pool mandatory rather than optional and merges M3/M4 into one gating milestone; (2) mlx-swift's Metal shaders cannot be built by SwiftPM CLI — Xcode or a vendored metallib is required, which changes M7 packaging; (3) `mlx_lm.load()` defaults to `lazy=False` and will drive a 48 GB Mac to 48 GB of swap; (4) decode is kernel-launch-bound (batch-1 matmul reaches 20% of memory bandwidth), which displaces IO as the top performance risk. **De-risked**: the M3 entry gate (slot writes 49–75 GB/s, in place, ~12× faster than the SSD can feed them), `gatherQuantizedMM` bit-exactness in both Python and Swift, and the existence of Swift GDN/MoE prior art. **Not achieved**: no end-to-end generation of the full model, because the stock path cannot produce one on this machine and the bounded path is M3/M4 work. No expert locality curves from the real model (the trace collector and simulator are built and the simulator is validated on synthetic input, but collecting real traces requires the same bounded forward pass). M1's h-curves remain open — though they no longer gate viability, only tier sizing. ### One more toolchain constraint (found the hard way) `swift test` is impossible on this machine: neither XCTest nor swift-testing ships with Command Line Tools — both require Xcode. Acceptance testing therefore lives in `Tools/verify.sh`, which is strictly stronger anyway: it runs the n-gram-id golden, the chat-template golden, the bit-exact layer-parity gate, and the full-model golden-equivalence test against the real checkpoint. Current status: **4/4 PASS.** ### The auto memory target: 70% of RAM was the wrong shape (2026-08-31) Auto targeted `min(70% of RAM, workingSet - 2)`, then clamped to `available - max(1.5, 5% of RAM)`. Two things were wrong with that, and only the second one matters. The small one: of *reclaimable* memory it took up to 93%, leaving a fixed ~2.4 GB cushion on a 48 GB Mac however much was free. The real one: **the target scaled with RAM, and the speed did not.** Measured a GB at a time with `doctor --json`: | RAM | old target | est. decode | |---|---|---| | 48 GB | 33.6 GB | ~12 tok/s | | 64 GB | 44.8 GB | ~12 tok/s | | 128 GB | 89.6 GB (83 GB peak) | ~12 tok/s | A 128 GB Mac gave up 83 GB to run at the speed 33 reaches. The decode curve plateaus at 150 experts/layer (11.2 tok/s at 120, 11.6 at 150) and the plan lands 150 well before it runs out of machine. **33 GB is the knee**, and it is a two-sided one: it is the smallest target where the cache clears the decode plateau *and* the budget still affords the 4096-token prefill pass (125 tok/s against 113 at 2048). Sweeping 34 to 84 GB moves neither number. That the old 70% rule produced 33.6 GB on a 48 GB Mac is why this never showed up: it was accidentally right for the author's machine and wrong for every larger one. Auto now takes `min(33, ramPercent% of RAM, workingSet - 2)`, and the ceiling also makes the availability clamp stop binding on a quiet 48 GB Mac — 22.8 GB stays free instead of 14.4, with no loss of speed. `--max-ram-percent P` exposes the share. Deliberately no knob lifts the ceiling: `--memory-gb`, `--experts-per-layer` and `--pool-gb` already do that precisely, and full residency (512/layer, ~88 GB) stays reachable for anyone wanting to test the one unreproduced hint of a further step at 181/layer. **A bug this found: more memory could plan a slower machine.** `--memory-gb 26` planned a *smaller* cache than 25 (116 against 128 per layer) and a slower decode, because crossing a quarter of the pool budget doubled the prefill pass from 2.7 to 5.3 GB — more than the GB just added. Pass sizing now scores candidates by `prompt/prefill + reply/decode` seconds rather than "biggest that fits", which prices both sides in the same unit. Swept 7 to 90 GB, the estimated request time never gets worse as the target grows; `Tools/monotonic_plan.py` is that invariant as a gate. Two dead ends on the way there, both from optimising decode alone: requiring a pass step to cost *no* decode collapsed it to 256 tokens on small machines (94 to 40 tok/s prefill), and a 2% tolerance still let three steps through. The total-time objective subsumes both. **Method note.** The first version of the monotonicity gate read the banner and reported a regression at 31 GB that did not exist: the banner rounds tok/s to whole numbers, so 11.58 prints as 12 and 11.19 as 11. `doctor --json` now emits `est_warm_tok_s` and `est_prefill_tok_s` unrounded, and the gate reads those. Anything asserting on a plan should. ### Automatic memory default: evidence scope and retained policy (2026-09-09) This clarification keeps the current automatic memory policy and every historical run unchanged. It separates the evidence behind the policy from the predictions produced by that same policy. The preceding [[records/measurements/the-auto-memory-target-70-of-ram-was-the-wrong-shape-2026-08-31]] describes a sweep using `doctor --json`. Its flat larger-target speed values are planner predictions, whose decode curve already stops extrapolating at its upper verified anchor. They are not an independent benchmark of larger allocations. The historical wording that nothing larger decoded or prefilled faster must be read as a statement about those estimates, not a measured universal result. The real ladder in [[records/measurements/warm-decode-re-anchored-and-the-live-governor-finally-observed-2026-08]] records 11.2 tok/s at 120 experts/layer and 11.6 at 150 on the development Mac. It shows diminishing gains over that measured interval. The larger-cache observation at 181 was not reproducible without memory-pressure contamination and remains excluded from a stronger throughput claim. The 33 GB base target remains the best-supported operating choice so far for the implemented model and planning objective, accommodating expert cache and prefill workspace without consuming more memory solely because a machine has it. This is a policy judgment, not proof of an optimum on every hardware/workload pair. The 70% RAM share is an upper bound on auto, not a lower bound. The draft charge and the existing availability/Metal bounds remain separate. A better comparable hardware/workload result can justify changing the default. Until then, explicit sizing is the supported way to explore another tradeoff, with its documented fixed-cache behavior. A model-free planner gate proves target selection, arithmetic and diagnostics; it does not measure allocation, physical peaks or speed. No new benchmark or numerical default is introduced by this clarification. Controlling decision: [[records/decisions/auto-target-is-the-33-gb-knee-not-70-percent-of-ram]]. General engineering contract: [[records/design/measured-operating-policies]]. ## Community evidence incorporated on 2026-09-13 [[records/measurements/c2-macbook-pro-m5-max-128gb-community]] now explicitly surfaces the larger-target sweep already preserved in its original source. The same M5 Max reportedly ran faster as its manual memory target increased beyond auto. This is positive evidence that a larger allocation can help; the conservative development-Mac default is not established as the best tradeoff on that machine. The public tables and memory FAQ now make this scope explicit. Runtime defaults remain unchanged, pending qualification of a hardware-specific allocation policy. No new model run was performed. ### The --memory-gb promise did not hold on real prompts (2026-08-31; resolved below) `--memory-gb 10` **peaked at 12.4 GB** on a 7,960-token prompt. Characterised so the fix would not have to start from scratch; the cause and fix are in the Resolution below, and a clean build of the committed tree now peaks at 8.6 GB on the same prompt with the full battery green. All runs greedy, 8-16 tokens out, RSS peak as the binary reports it. **It is not the expert pool.** At the 1.8 GB floor cache the same prompt still peaks at 10.9 GB. Sweeping prompt length at that fixed floor pool: | prompt | peak RSS | |---|---| | ~15 tokens | 5.4 GB | | ~250 tokens | 10.2 GB | | ~1,000 tokens | 10.5 GB | | ~3,900 tokens | 10.9 GB | | ~7,960 tokens | 10.9 GB | So it is **a step, not per-token growth**: ~4.8 GB appears between a trivial prompt and a 250-token one, then roughly 90 KB/token after. Anything that budgets this by context length will mis-size it. **It is additive in the pool.** Pool 1.8 GB gives 10.5 GB peak and pool 8.0 GB gives 16.9 GB on the same prompt: +6.2 GB of pool costs +6.4 GB of RSS. The overhead above the pool is ~8.7 GB where the plan models ~5.8. **But it is not a clean constant either**, which is why this is left open rather than patched by inflating `fixedFootprintGB`: | target | peak | verdict | |---|---|---| | 10 GB | 12.4 | over by 2.4 | | 16 GB | 18.4 | over by 2.4 | | 20 GB | 19.6 | **under**, holds | Two candidates ruled out. The n-gram row cache is capped at 400,000 rows of 160 floats, about 0.3 GB with overhead — an order of magnitude too small to be the step. `MLX.Memory.cacheLimit` is already pinned at 2 GB in `Engine.swift`, so an unbounded allocator cache is not it either. Raising `fixedFootprintGB` to cover the gap would push `minMemoryGB` from 8.1 to over 10, which refuses the very target the gate tests, and would shrink every pool to pay for something whose location is still unknown. The next step is to find where the ~4.8 GB step is actually allocated on the first real prefill, not to widen the constant around it. ### Resolution: bound expert-load staging (2026-08-31) The discontinuity was the cold expert-fill shape. One 256-token layer can route all 512 experts, and `SlotPool.ensure` handed all misses to one `ExpertStore.readBatch`. At 2,764,800 bytes per record that is a **1.415 GB raw batch** across nine tensors; its MLX scatter materialization and the source buffers coexist at the high-water point. It looked unrelated to prompt length because the batch jumps to nearly every expert as soon as a prompt is large enough, then cannot grow beyond 512. `ensure` now reads and fully scatters at most **32 records at a time**. Each slice is evaluated before the next is read, so no later slice can overlap its staging lifetime. The internal reads remain queue-depth-parallel and the pool still pins the complete routed set before any MoE math, so this changes memory and scheduling, never weights, routing, or output. The exact 7,960-token `--memory-gb 10` acceptance prompt after the change: | metric | before | bounded staging | |---|---:|---:| | process RSS peak | 12.4 GB | **8.6 GB** | | target verdict | over by 2.4 GB | **under by 1.4 GB** | | prefill | — | 7,960 tokens / 192.74 s = **41.3 tok/s** | | prefill split | — | 99.13 s I/O + 44.92 s scatter + 48.69 s compute | That preserves the planner's conservative 40 tok/s estimate for a 256-token pass while removing 3.8 GB from the observed high-water mark. The answer remained `SEVENTEEN`, and the ordinary small-vs-large cache equivalence gate remains the correctness check. No fixed-footprint inflation is needed. ## M9 — MTP self-speculative decode: conversion, parity, accept curve, and where it pays (2026-09-01) The design note said the multiplier only exists in the launch-bound regime and that memory decides. Everything measured this session agrees with that shape — including the one result that looks negative and isn't. ### The head exists again (the pinned conversion had dropped it) The pinned community conversion strips all `mtp.*` tensors (`sanitize` drops them), so the draft head was rebuilt from the official release without downloading it: the 31 MTP tensors live in 28 of the official repo's 131 shards, and safetensors headers give exact byte ranges, so `Tools/mtp_convert.py` range-requests precisely those tensors — **4.9 GB in 92 s** instead of ~250 GB — then applies the same transforms the community conversion applied to the main model and quantizes to the same recipe (4-bit, group 64, affine; router, gates, `index_qk_proj`, and norms stay bf16). Output: `mtp.safetensors`, **1.471 GB** (the design note estimated 2.25), sha256 in `mtp.provenance.json`. Two conversion facts were verified rather than assumed: - **The +1 norm centering is real and uniform.** For four main-model norms the official raw tensors were fetched and compared against the pinned converted ones: mixer hc_norm +2.7497 → +3.7498, attn hc_norm −0.2638 → +0.7362, indexer q_layernorm −0.0372 → +0.9628, q_norm +0.2833 → +1.2833 — exactly +1.0 each. The MTP-only `pre_fc_norm_*` weights (no main-model analog, not in the reference's CENTERED list) follow the same convention: vLLM builds them as GemmaRMSNorm (the 1+w form), and the raw embedding norm sits in a tight band around −0.764 — sensible as 1+w ≈ +0.24, pathological as a bare negative scale. - **The forward semantics come from the only public implementation.** vLLM's `Qwen4ExpMultiTokenPredictor` ("scheme A"): fc_embedding on the normed token embedding, a SHARED fc_hidden on each of the four normed hyper-connection branches of the PRE-final-mixer multi stream, embedding added to every branch, one full-attention decoder layer, the head's own mixer for the lm_head path — and the pre-mixer stream, not the collapsed one, feeds the next chained draft step. ### Parity: bit-exact, after two false alarms worth recording `slotstream mtp-parity` compares the Swift head against the MLX Python reference (`Tools/reference/mtp_ref.py`) on a stored fixture. Final result: **max abs 0.00000 on all four outputs** — prefill sample/multi and cached decode sample/multi are bit-identical. Getting there surfaced two lessons: 1. **Random fixture inputs are adversarial for THIS layer.** The MTP block's norms run hot (raw q_norm mean 2.68 vs 0.28 on a main layer), and its attention logits reached **680** with top-2 gaps as small as **0.5** on random inputs. Sub-ulp cross-implementation noise flips near-tie argmax keys and reads as a 20% output error. Feeding Python's sdpa the Swift-dumped q/k/v byte-for-byte returned Swift's output exactly — the ops were never wrong, the near-tie lottery was. The fixture now uses REAL captured inputs (`slotstream mtp-fixture-inputs`). 2. **The reference must run on the mlx the Swift build pins.** mlx-swift is 0.31.x; a fixture generated under Python mlx 0.32.2 disagreed at 5–15% (kernel reduction orders moved between versions), regenerated under 0.31.1 it is bit-exact. Same lesson as the layer-parity work, now written down: `make_mtp_fixture.py` runs under `.venv31` and says so. ### The accept curve — measured, previously unpublished anywhere `slotstream mtp-accept` runs plain greedy decode and, at every position, chains the draft head then rolls it back, scoring drafts against the tokens the model actually produced. Four prompts (prose, code, list, arithmetic), 96 tokens each, 380 scored positions: | chain depth | prefix accept | E[tokens/round] | fetch-free ceiling | |---|---:|---:|---:| | 1 | **85.8%** | 1.86 | ×1.37 | | 2 | 71.0% | 2.57 | **×1.48** | | 3 | 53.8% | 3.11 | ×1.43 | | 4 | 41.3% | 3.52 | ×1.38 | The head predicts the model's next-next token at 85.8%. The last column is the round arithmetic with the pass costs `mtp-passcost` measured (below): a k-token verify costs about 1 + 0.16k single passes, a rebuild likewise, a draft step 0.05, so it is the speedup with every expert resident, and no real cache reaches it. The 0.2.0 version of this table assumed the verify pass was free and read ×1.52–1.96. ### Where it pays — measured at every cache size that fit (2026-09-01, redone 2026-09-02 on the release) In-process A/B (`slotstream mtp-bench`: one engine, one warm pool, the two decode paths alternated, greedy, one 300-word prompt, 192 tokens, medians of interleaved pairs) on the released 0.2.0 binary at its draft depth of 4, at every cache size the machine could hold with headroom that evening: | target | experts/layer | plain tok/s | speculative tok/s | ratio | round cost, in plain tokens | |---|---:|---:|---:|---:|---:| | 10 GB | 20 | 5.80 | 3.21 | **×0.55** | 5.2 | | 14 GB | 29 | 5.92 | 4.07 | **×0.69** | 4.2 | | 16 GB | 42 | 6.35 | 5.59 | **×0.88** | 3.3 | | 18 GB | 57 | 7.48 | 7.20 | **×0.96** | 3.0 | Every run drafts the same 268 tokens and keeps 124 (46.3%): 67 verify passes for 192 tokens, 2.87 tokens per round. The last column is 2.87 × plain/spec, what one round costs in units of the plain path's per-token time, and it is the number to watch: it falls from 5.2 to 3.0 as the cache grows, because the verify pass's five tokens fetch fewer experts, and the paths break even when it reaches 2.87. The 18 GB row is five pairs at 0.91–0.99 each; a first three-pair run at that size overlapped another session's build and threw a 4.5 tok/s pair on both arms, discarded per the discipline above. The dev-build figure from the day before (×0.96 at 54/layer, 6.77 → 6.52) reproduces. At a fixed target, which is what the flag does for a user, the loss is larger, because the head's 1.6 GB comes out of the cache. `run --memory-gb 14 --greedy` on a cold process, two rounds each: `--mtp off` plans 40 experts/layer and decodes at 6.97 / 7.18 tok/s (hit rate 0.606); `--mtp on` plans 29/layer and decodes at 4.69 / 4.55 (hit rate 0.391, 124/268 drafts accepted): **×0.65**. Auto keeps the head off until the cache still reaches 120/layer after the charge, which on this Mac is a **28 GB** target (27 GB plans 127/layer, 115 after the charge; the "~26 GB" the 0.2.0 docs said was never read off `doctor`, and is corrected). ### The plateau: its ceiling measured, then the A/B itself (2026-09-02) The A/B at ≥120/layer, where auto turns the head on, waited most of the day behind Carlos's apps: the plan's conservative 27 GB peak against 20 to 28 GB reclaimable. It ran late in the evening (below; the real footprint was 20 GB). First, the premise the ×1.5–1.9 arithmetic stood on: that verifying five tokens in one pass costs about one token's pass once nothing has to be fetched. `slotstream mtp-passcost` runs a pass twice from one checkpoint at eight real positions of a greedy continuation and times the second run, when the pool's miss counter reads zero (it did, at every position): | pass, every expert it needs resident (57/layer) | ms | × one token | |---|---:|---:| | verify 1 token | 48.3 | 1.00 | | verify 2 / 3 / 4 tokens | 56.7 / 64.2 / 72.1 | 1.17 / 1.33 / 1.49 | | verify 5 tokens (depth 4) | 79.8 | **1.65** | | rebuild 1 / 2 / 3 / 4 kept tokens | 47.1 / 54.8 / 62.7 / 69.1 | 0.97 / 1.13 / 1.30 / 1.43 | | one draft step (head + lm_head) | 2.3 | 0.05 | **The premise was false.** Each token in the batch adds ~8 ms, a sixth of a single pass, linearly: a five-token pass gathers up to five times the expert weights of a one-token pass, and that does not ride free on launch overhead the way a dense batch-1 matmul suggested. With the measured accept curve (85.8 / 71.0 / 53.8 / 41.3% for chains of 1 to 4, so 3.52 tokens per round at depth 4) a fetch-free round costs 4 × 0.05 + 1.65 + 0.71 (the expected rebuild) = 2.55 plain tokens, and the speedup with **every expert resident** is capped at **×1.38** at depth 4, ×1.43 at depth 3, **×1.48 at depth 2**, ×1.37 at depth 1: the linear cost favours short chains. These are ceilings. On the real plateau the plain path is far from fetch-free (11.6 tok/s at 150/layer against the 20.7 tok/s a 48 ms pass implies, so some 40% of every token is still fetch; the never-reproduced 20.0 at 181/layer looks like this fetch-free rate glimpsed once) and the verify pass fetches for several tokens. **The "×1.5–1.9" in the 0.2.0 docs was arithmetic on this premise and is withdrawn.** ### Depth, and the plateau A/B that moved the default from 4 to 1 (2026-09-02) Same A/B at 57/layer, five pairs per depth (`SLOTSTREAM_DRAFT_DEPTH`): | depth | tokens/round | drafts accepted | plain → spec (column medians) | ratio | pair ratios | |---|---:|---:|---|---:|---| | 4 (0.2.0) | 2.87 | 46.3% | 7.48 → 7.20 | **×0.96** | 0.91–0.99 | | 2 | 2.34 | 66.5% | 7.32 → 7.38 | ×1.01 (pair median **×1.12**) | 0.85–1.27, two pairs disturbed | | 1 | 1.81 | 80.2% | 7.36 → 8.34 | **×1.13** | 0.96–1.27 | Depth 4 loses consistently; depths 1 and 2 are indistinguishable at this size and both ahead of plain. Calibrated on these rounds, the verify pass's fetch grows to about 1.7× / 2.6× / 3.1× a single token's for 2 / 3 / 5 tokens (consecutive tokens share experts, so it is not 5×), and projecting those onto the plateau (38 ms of fetch in an 86 ms token) put every depth near ×1.2. Then the plateau itself, once the machine had the room: the same A/B at `--memory-gb 28` with `--mtp auto`, which is the configuration auto ships (122 experts/layer after the head's charge), five pairs per depth, nothing else running, 20 GB real footprint: | depth | plain → spec tok/s (column medians) | ratio | pair ratios | |---|---|---:|---| | 4 (0.2.0, 0.2.1) | 10.41 → 9.15 | **×0.88** | 0.82–1.02 | | 2 | 10.33 → 11.63 | **×1.13** | 1.10–1.17 | | 1 | 10.09 → 11.79 | **×1.17** | 1.04–1.20 | The projection held for depths 1 and 2 and was optimistic for depth 4, which pays more rebuild and wasted fetch than the calibration allowed: the depth 0.2.0 and 0.2.1 ship loses even where auto turns it on. The default is now **1**: the best measured at the size that matters, the least wasted work on a rejection, and the most robust to a prompt the head reads badly. Depth 2 would win only once the rebuild goes away (its fetch-free ceiling is ×1.48 against ×1.37). Plain decode here reads 10.1–10.4 tok/s against the 11.2 anchor at 120/layer: the bench's pool is warmed by one repeated generation, and the anchor was taken on a different prompt; the ratio, not the level, is the measurement. Nothing about the head's cache, the gates, or the parity fixture depends on the depth. ### The rebuild eliminated, and the numbers that ship (2026-09-02) The rebuild after a rejection was the largest avoidable cost left: a full pass over the kept tokens (0.97 to 1.43 pass-equivalents, above) on every round that rejected, because the GDN recurrent state could not be rewound. It can be recorded instead. While a verify pass runs, each linear-attention layer now steps the GDN recurrence one token at a time and keeps the state after every position, and slices the conv windows the same way; a rejection swaps in the recorded state at the last kept position, trims the attention caches, rebuilds the n-gram context from ids, and slices the pass's own multi stream for the draft head. No model compute. Two things had to be true for this to be safe, and both are gated in `mtp-check`. The stepped recurrence must match the fused kernel, and it does to the bit (0.0000% of logit spread; the state is fp32 between steps exactly as inside the kernel). And the state a rollback leaves must be the plain path's state for the same kept tokens, which it is up to re-association: the kept token's projections came out of a two-row batch, so its recurrent tensors differ from a one-row build by 6.5e-2 (GDN state) and 4.3e-2 (conv window) relative, against 1.1e-1 and 7.1e-2 for the plain path merely re-chunked, and one more step moves the logits 3.4% of their spread against a 3.3% re-chunk control. A wrong window would read order one. The old rebuild recomputed the kept tokens one-row and matched the plain path exactly; the price was most of a pass per rejection, for a difference the batched verify had already accepted in the logits it sampled from. The plain path is untouched: recording is on only inside a verify pass, and the layer-parity gate still reads bit-exact. What it buys, same A/B, five pairs each: | depth | 57/layer before → after | 122/layer before → after | |---|---|---| | 1 | ×1.12 → **×1.20** | ×1.17 → **×1.24** (10.30 → 12.77 tok/s) | | 2 | ×1.12 → **×1.20** | ×1.13 → **×1.27** | Coverage beyond one prompt and greedy decode, all at 122/layer, depth 1, with the rollback: a code prompt reads ×1.33, a list prompt ×1.19, and the server's default sampling (temperature 0.7, top-p 0.8, top-k 20, presence 1.5, seeded so both paths draw the same stream) ×1.18 at 72% accept against 80% greedy. Before the rollback, at 57/layer, sampling read ×1.10 against ×1.11 greedy in the same session, so sampling costs a few percent of the gain, not the gain. `mtp-bench --sample` is the switch. What this means for the policy: auto's 120/layer floor stands, now on measurement rather than argument. At that size the head returns +17% at depth 1 for 1.6 GB that would otherwise buy experts past the plateau (decode is flat from 120 to 150), so it is free; below the floor every point is a loss at depth 4, and the few percent depth 1 would gain there the displaced cache would eat. The head is still an opt-in artifact (`Tools/mtp_convert.py`), so the gain reaches nobody until it ships with `pull`; that is the next product decision, not an engineering one. ### Correctness story (a claim from the design note corrected) The note said greedy speculation is "byte-identical to plain greedy". With the prefix-cache result in hand that claim was corrected before shipping: every emitted token's logits still come from the main model, but the verify pass computes them in a k+1-token batch, and re-batching re-associates sums exactly the way prefill re-chunking does — near-tie argmax flips are possible and observed. The shipped gates (`mtp-check`) are: two speculative runs byte-identical (determinism), a follow-up turn continues a speculative conversation through the prefix cache, the accept rate is not degenerate, and plain-vs-spec divergence is REPORTED, not gated to zero. Sampling semantics are exact by construction: draws happen sequentially off the verified logits, only for tokens the plain loop would also have sampled, so the rng stream and presence-penalty evolution match the plain path token for token. One gate had to be rebuilt twice to be honest. The cross-request check first asserted "a continuation of a speculative conversation produces tokens" and failed — not because the state was wrong, but because a 48-token turn-1 reply is usually cut mid-think, and the model legitimately answers some continuations of that context with an immediate EOS. 0.2.0 shipped it with a plain-path control ("the speculative path must not be the one that goes silent"), which passed on its deciding run: reused state " Paris", control " Paris.". Moving the draft depth to 2 flipped it the other way — reused state EOS, control " Paris." — with turn-1 text that differed between the paths by a near-tie flip, so the control was comparing two different contexts and the assertion was a coin toss either way. The gate now measures what it means: the next turn's logits from the reused speculative state against a cold rebuild of the same ids, bounded by three times what re-chunking a plain prefill moves them (the `prefix-check` method). On the deciding run it read 7.36% of the logit spread against a 6.16% control at depth 2 (top-1 differs) and 5.77% against 6.16% at depth 4 (top-1 same); a misaligned state reads tens of percent. State rollback is O(1): the recurrent caches' arrays are REPLACED each step (the GDN kernel emits a fresh state_out), so a checkpoint holds references, and KV/indexer buffers roll back by offset. A rejection costs one re-run of the kept tokens (the GDN state cannot be rewound — same constraint the prefix cache lives with). RoPE positions: the head trains with entries at position i+1; the port keeps 0-based cache positions. All rotations shift by the same constant and RoPE attention depends only on relative positions, so scores are mathematically identical — noted in MTP.swift rather than adding a shift parameter. ## M10 — The pull ran on one TCP connection (2026-09-01): HTTP/2 coalescing, measured and fixed A user reported a 22 MB/s average for the whole 104 GB. The investigation overturned three numbers this document had published, all measured from the same home link in Popayán, and found that `pull` had never opened more than one TCP connection. ### The finding: eight requests, one connection `PullJob` made one `URLSession` with `httpMaximumConnectionsPerHost = 8` and ran eight workers against it. Apple's documentation for that property says: "HTTP/2 and later run multiple requests over a single connection and thus ignore this property," and "This limit is per session, so if you use multiple sessions, your app as a whole may exceed this limit." Hugging Face speaks HTTP/2, so every pull since 0.1.4 ran eight streams on one connection. Proof from inside the process: a reproduction of the exact configuration, instrumented with `URLSessionTaskMetrics`, showed all eight 206 bodies on the same local port, `isReusedConnection: true`; with one session per worker, eight distinct ports. `nettop` on the real `slotstream pull --connections 8` showed one data-carrying flow. Interleaved same-minute pairs on the home link: | round | one session, 8 tasks | 8 sessions, 8 tasks | |---|---|---| | 1 | 16.0 MB/s, 1 connection | 41.5 MB/s, 8 connections | | 2 | 28.5 MB/s, 1 connection | 56.8 MB/s, 8 connections | Nothing else serializes the workers (read end to end: eight threads, each `dataTask` + semaphore); the delegate's `pwrite` path is not the limit (the local-server pull below runs it at 3.2 GB/s); the same server IP served both modes; request counts are identical. ### Why one connection is slow: window over round trip macOS caps a single TCP receive window at 4 MiB (`net.inet.tcp.autorcvbufmax`). Per-connection throughput cannot exceed window / RTT. The TCP handshake to the host that serves the bytes is 99 ms from Popayán, giving a ceiling near 42 MB/s before any loss (measured 25 to 40); from Helsinki it is 35 ms, so the window does not bind there (measured 72, set by the far end). Where the bytes come from: `resolve/` answers 302 to `us.aws.cdn.hf.co/xet-bridge-us/...`, the LFS bridge Hugging Face documents as reconstructing a Xet-backed file for plain HTTP clients. That hostname resolves through GeoDNS to Amazon EC2 addresses (checked against `ip-ranges.json`): `us-east-1` for Colombia, `eu-west-3` (Paris) for Helsinki. The responses carry CloudFront headers, but the TCP connection terminates at EC2, one region per continent, so every user pays a regional round trip per connection. ### Hugging Face has no per-client cap; the home link has one Same probes, two vantage points, separate HTTP/1.1 connections, totals across all connections: | source | Popayán home, 1 conn | Popayán, 8 conn | Helsinki 1 Gbit/s, 1 conn | Helsinki, 8 conn | |---|---|---|---|---| | Hugging Face bridge | 31 | 52 to 55 | 71 to 72 | 103 to 106 (108 at 16, 109 at 32) | | Cloudflare R2, direct (Ollama's bucket) | 49 | 62 (62 at 16) | 68 | 108 (110 at 16) | | Cloudflare edge, generated bytes | 41 | 57 | 103 | 113 | | Cloudflare cache hit (nodejs.org) | 37 | 47 | 100 | 114 | | Hugging Face, 8 streams on one HTTP/2 connection | — | 35 | — | 73 | The Helsinki port is 1000 Mbit/s (`/sys/class/net/eno1/speed`); 108 to 114 MB/s is the port. So the 2026-08-29 conclusion "Hugging Face caps this client at about 55 MB/s, 4 through 32 connections all plateau there" measured the home link, which today caps *every* host in a 47 to 62 band, and was read through a client that was using one connection anyway. Chunk size is not a lever: at 8 connections from Helsinki, 64 MB chunks moved 98 MB/s and 256 MB 108.5, the difference being slow-start amortization on one-shot `curl` connections that the persistent sessions in `pull` do not pay. ### Corrections to earlier sections - **"Hugging Face caps at 36 to 57 MB/s however many connections you open"** (README, CLAUDE.md, the parallel-download section above): withdrawn. That was one connection on a capped link. - **"Cloudflare R2 tested and rejected"**: the comparison could not discriminate. It ran through `r2.dev`, which Cloudflare documents as rate-limited and bandwidth-throttled, from a link on which R2's real read path (62), Cloudflare's edge (57) and Hugging Face (55) land in one band. From Helsinki, R2 direct and Hugging Face both fill the port. The custom- domain path remains untested; it is not needed for speed, only for independence from Hugging Face. - **"The link does 144 MB/s against Hetzner"**: could not be repeated. The Ashburn speed host now limits clients to two connections; its per-connection rate today (18.7 MB/s) matches the 08-29 run, so the link claim rests on that one measurement. Nothing else exceeded 62 MB/s from Popayán today. - **The ETA hint** quoted 50 MB/s as the mirror's ceiling; it now quotes the fastest measured rate for this client, 100 MB/s on a gigabit link, and says so. ### The fix and its gates One `URLSession` per worker (`httpMaximumConnectionsPerHost = 1` each, shared delegate; the delegate already keyed tasks by its own `taskDescription`, so sessions cannot collide). After every session has completed a chunk, `pull` prints the distinct TCP connections it measured from task metrics — the first version counted after two chunks per worker and once reported "15 of 16" because one slower session had not finished its first chunk; it now waits for all of them. | gate | result | |---|---| | `Tools/static_gates.sh` | pass (planner 64/64, pull-check, runtime-check, installer, llms-full current) | | `pull --verify` on the installed weights | VERIFY PASS, 9.5 s | | full 24-file pull from a local Range server (SSD) | "8 connections in use", 3,183 MB/s, VERIFY PASS in 64 s | | live mirror, `--connections 8`, 50 s | 8 flows in `nettop`, "8 connections in use", 57 to 63 MB/s cumulative (38 to 41 same evening before the fix) | | live mirror, `--connections 16` | 16 flows, "16 connections in use", 63 MB/s (the link) | | live mirror, `--connections 1` | 15 MB/s | | kill at 3.0 GB, rerun | "101.0 GB to go", continues from the chunk map | | Linux, 1 Gbit/s Hetzner port, Docker, exact `Pull.swift` via `Tools/pull_bench_linux.sh`, 8 connections, whole 103.8 GB | VERIFY PASS, 24/24 files; 112 MB/s average across the whole 103.8 GB, about 15.5 min wall (started 03:07 UTC, verified by 03:23); the port is 1000 Mbit/s | What this does not settle: the reporting user's own ceiling (his link, not the client, once eight real connections exist); anything above 1 Gbit/s (public reports put Hugging Face at 500 MB/s to 1 GB/s with `hf_transfer` on 10 Gbit links, unmeasured here); whether Popayán's 62 MB/s band is the ISP or CDN peering; and the R2 custom-domain path. ### Post-release: 0.2.1 installed through `install.sh` (2026-09-01) | check | result | |---|---| | `install.sh` one-liner | installed 0.2.1 to `~/.slotstream/bin`; asset sha256 matches; `gh attestation verify` exits 0 | | installed `pull --verify` | VERIFY PASS, 24/24 | | installed `pull`, default 8 connections, 50 s | 8 flows in `nettop`, 44 to 47 MB/s cumulative on the home link | | `Tools/e2e_release.sh` against `serve --memory-gb 8.1` | 30 of 31; the one failure, "empty prompt refused", asserted the pre-0.2.1 400 for a chat with no messages, which 0.2.1 answers with `done_reason: "load"` as its changelog and `api_robustness.sh` say — the check is now aligned and passes | The 0.2.1 report line counted every distinct connection since start, so two early reconnects made it say "10 connections in use" for eight workers. It now keeps one entry per session, the connection that session most recently carried a body on, and reports the distinct count once every session has one: "8 connections in use" with 8 flows in `nettop`. ## M11 — Hosting is not the lever below ~3 Gbit/s (2026-09-02): the bound per link speed, the bridge versus Hugging Face's own edge, and what the default leaves on the table M10 fixed the client and left one question open: could any hosting (R2, Vercel Blob, a CDN at the edge) make the 103.8 GB pull faster than Hugging Face does? This section answers it from the transfer chain, vendor documentation, and live DNS/HTTP checks made on 2026-09-02. No new transfer was measured; every throughput number below is from M10 or is derived, and is labelled as such. ### The chain Wall time is bytes over the slowest of four terms: the client's link, the source's per-client rate, connections × TCP window over round trip, and disk plus hashing. The bytes are fixed (4-bit weights do not compress). Per term: - **Client link.** The binding term at every vantage M10 measured. From Popayán, Hugging Face, R2 direct, Cloudflare's edge and a Cloudflare cache hit all landed at 47 to 62 MB/s over 8 connections; from Helsinki all four landed at 106 to 114, the 1 Gbit/s port. - **Window over round trip.** macOS caps the receive window at 4 MiB (`net.inet.tcp.autorcvbufmax`), so one connection tops out near 42 MB/s at 99 ms (Popayán to the Virginia bridge), near 60 at 70 ms (San Francisco to Virginia), near 140 at 30 ms. Eight connections cover a gigabit link up to roughly 300 ms of round trip; past that the default falls short. - **Per-stream at the origin.** Object-store read paths cap near 70 MB/s per connection: from Helsinki one connection got 71 from Hugging Face and 68 from R2 direct, while one connection to Cloudflare's edge got 103 and to a cache hit 100, which is the port. An edge cache has no per-stream cap that a gigabit link can see; an origin read path does. - **Per node.** The `resolve` redirect lands on `us.aws.cdn.hf.co/xet-bridge-us/…` (checked 2026-09-02: HTTP/2, `206` with a correct `Content-Range` on a 10,039,592,993-byte shard, signed URL valid for about an hour). That hostname resolves to eleven A records, all EC2 and none inside CloudFront's published ranges (`ip-ranges.json`, service `CLOUDFRONT`). AWS documents 5 Gbps per single flow and 5 Gbps of internet egress per instance under 32 vCPUs (50% of the NIC above), so one bridge node is at most about 625 MB/s shared by every client on it, and the client's eight sessions almost certainly land on the same node. - **Disk and hash.** Chunks stream to disk with `pwrite` as they arrive and files hash on completion; neither binds below several GB/s. ### The bound per client link | Client link (round trip to the bridge) | Today's client, 8 connections | Best possible | Gap | |---|---|---|---| | 10 to 100 Mbit/s, any distance | link-bound | link-bound | none | | 100 Mbit/s to 1 Gbit/s, up to ~150 ms | 95 to 100% of link (Helsinki: 106 at 8, 110 at 16) | 100% | 0 to 5% | | 1 Gbit/s at 250 to 300 ms | ~80% of link, derived from the window term | 100% | ~20% | | 2.5 Gbit/s at 70 ms | link-bound: 8 × ~60 MB/s covers ~290 MB/s | link-bound | none | | 5 to 10 Gbit/s at 70 ms | ~480 MB/s at 8; ~625 at 32 if on one node | ~1.1 GB/s, the Mac's 10 GbE NIC | ~2×, unmeasured | In time: 1 Gbit/s is 15.5 min measured (M10); 2.5 Gbit/s is about 6 min and 10 Gbit/s about 1.6 min by the link alone, neither measured. ### Hosting, from first principles and the vendors' own documentation | Path | What the documentation and checks say | Verdict for speed | |---|---|---| | Hugging Face `resolve` bridge (today) | EC2 fleet, one region per continent seen (us-east-1, eu-west-3); anonymous `resolve` limited to 3,000 requests per 5 minutes per IP; signed bridge URL lasts about an hour | Equal to everything else at ≤ 1 Gbit/s; per-stream ~70 and per-node caps above that | | Hugging Face Xet edge | `transfer.xethub.hf.co`, `cas-bridge.xethub.hf.co` and `cdn-lfs*.hf.co` resolve inside CloudFront's published ranges; xorbs are ≤ 64 MiB objects; CloudFront caches responses up to 50 GB; the protocol is public (token from the Hub, `GET /v2/reconstructions/{file_id}`, signed multi-range xorb fetches, LZ4 and byte-grouping chunk compression) with Rust and TypeScript reference clients | The only edge path that exists today for these exact bytes, free; the plain `resolve` client cannot use it; unmeasured | | Cloudflare R2 + custom domain | No throughput limit documented on a custom domain (`r2.dev` is throttled); cacheable object limit 512 MB on Free/Pro/Business and 5 GB on Enterprise, `.safetensors` not a default-cached extension, so every 10 GB shard is served from R2 origin; zero egress, about US$1.60 a month of storage | Equal at gigabit (108 at 8 from Helsinki); an edge only if the shards are re-chunked below 512 MB; buys independence from Hugging Face, not speed | | Own mirror on CloudFront + S3 | Caches the 10 GB shards whole; city-level POPs; about US$9 of egress per install | Faster only above ~3 Gbit/s per client | | Vercel Blob | Amazon S3 underneath; served from 20 regional hubs on the network Vercel describes as cost-optimised "where ultra-low latency isn't essential"; cache cap 512 MB per blob, so every shard is origin on every request; US$0.05/GB transfer plus US$0.06/GB Fast Origin Transfer on each miss, about US$11.4 per install; simple-operation limits of 20/s (Hobby) and 120/s (Pro), and each 64 MB range on a >512 MB blob is one operation | No physics advantage over the bridge; a bill for the same speed | Verdict: for every client at or below about 3 Gbit/s, which is every home link and every 1 Gbit/s datacenter port, hosting cannot move the number, and the default already sits at the physical ceiling. Above that, an edge with the chunks hot a few milliseconds away is faster than the bridge, and the edge that already exists for these bytes is Hugging Face's own Xet path through CloudFront, which our HTTP `resolve` client bypasses. ### What the default leaves on the table, from the code `PullTuning` in `Pull.swift`: 8 connections by default, cap 32 (`SLOTSTREAM_PULL_CONNECTIONS`, `--connections`), 64 MB chunks, one chunk in flight per session, and every chunk is a fresh request to `huggingface.co/…/resolve/…` that follows the 302 to the bridge. Consequences: - More connections cost no memory: bodies stream to disk through `pwrite`. - Each chunk idles two round trips (resolve, then redirect) before its body starts: 0.14 to 0.2 s per 64 MB chunk that takes 1.1 to 1.5 s per connection at 60 to 42 MB/s, 10 to 15% per connection. Hidden while the link is the limit; paid in full when the connections are the limit (far gigabit, multi-gigabit). - About 1,620 `resolve` calls per install count against the 3,000 per 5 minutes anonymous limit; fine at gigabit (1.75/s), still under it at 10 Gbit because the whole install is 1,622 chunks, but with no margin for retries. Three client-only changes would make the default the physical best on every link up to what one bridge node can push, with no hosting change: 1. **Adaptive concurrency.** Start at 8, add connections while the aggregate rate still rises by a real margin, stop when it flattens, cap 64. A slow link stops at 8 on its own; a 10 Gbit/s link climbs. Hugging Face's own `xet-core` client does this. 2. **Two chunks in flight per connection**, so the redirect gap never empties the pipe (a second stream on the same HTTP/2 connection is exactly what the per-worker session allows). 3. **Resolve once per file** and reuse the signed bridge URL until it expires, about an hour later (fall back to `resolve` on 403). Removes a round trip per chunk and takes the pull off the anonymous rate limit entirely. Past that, at 5 to 10 Gbit/s, the per-node cap and the ~70 MB/s per-stream origin limit remain, and only the Xet edge path or a spread across bridge nodes (URLSession cannot pin a session to an IP) gets the rest. ### The fastest path is not the WAN at all For a second Mac or a team: a peer copy over 10 GbE takes about 1.5 minutes, over a Thunderbolt bridge about 1 minute, and an external SSD is faster than any link. A `pull --from ` mode would beat every hosting change for repeat installs; nothing here is implemented. ### Open, each needing an explicit go - **One hour on a rented 10 Gbit/s host** (billed): the current client at 8, 32 and 64 connections against Hugging Face, next to `hf_xet` on the same model. Decides between client tuning and implementing Xet. - **A Xet download path in Swift**, roughly a week: token endpoint, reconstruction call, signed multi-range fetches, LZ4 and byte-grouping decompression, reassembly, then the existing sha256 gate. - **Pull from a LAN peer.** ### What this does not settle Nothing above 1 Gbit/s has been measured by this project; CloudFront's per-client rate on a cold xorb set; whether a public repo issues an anonymous Xet read token (the bridge URL carries `user_id=public`, which suggests yes); whether Hugging Face's GeoDNS ever hands a West Coast client a West Coast bridge (only us-east-1 and eu-west-3 were seen); and the cause of Popayán's 62 MB/s band (M10). ## N2 — the prefill sweep: grouped GEMM over staging, contiguous reads, no pool writes (2026-09-02) Prefill was the last unmet target: the plan asked for ≥150 tok/s on an 8k prompt and 0.2.2 read 113 at best. Splitting the old pass had already said where the time went (io 33.9 / scatter 10.3 / compute 50.3 s on 8k) and what would close it: a grouped GEMM over the routed experts instead of one gather per token, and contiguous reads instead of nine ~307 KB pieces per record. Both landed on 2026-09-02, with the scan-resistant sweep §3.3 designed, and this section is the record: what the old pass was actually doing, what the sweep does instead, the gates that say it computes the same thing, and the numbers, every one an interleaved A/B on the dev Mac between the 0.2.2 code (`e09bcac`, the commit before this work) and the sweep, same prompt, same target, one model process at a time. **What the old pass was doing.** `MoELayer` already called MLX's `gatherQuantizedMM` over the pool, but with one row per (token, expert) and unsorted indices. That reaches MLX's per-row matvec kernel, which re-reads an expert's weights once per token that routes to it — about forty times per 2048-token pass. MLX has a grouped kernel (`gather_qmm_rhs`) that reads the weights once per tile of tokens, but it takes it only for sorted indices and only when a call has at least 16 rows and four rows per expert of the weight array it is handed. Handed the pool, `E` is the slot count: a 2048-token pass has 20,480 rows and would reach the kernel only below 5,120 slots (106 experts per layer), so the kernel — and the arithmetic — would have switched with the cache size, which is exactly what the golden- equivalence invariant forbids. Every miss also went through the pool: 32 records at a time, nine preads per record, a scatter into the slots, and a CLOCK state flushed by every long prompt. **What the sweep does.** A pass of 256 tokens or more (`SweepTuning.minTokens`) sorts its rows by expert (a counting sort on the CPU, 20 ms per prompt) and walks the layer's experts in groups of 32: experts already resident are copied out of the pool, the rest are read from the checkpoint with one `pread` per piece per run of consecutive ids, and each group is one grouped GEMM per projection over that group's rows, `sortedIndices: true`. A group short of MLX's rule is padded up to it with repeats of its last row, so the kernel a row meets depends on the routing alone. The GPU works on one group while the CPU reads the next; at most two groups of staging exist at once. The pool is never written by a sweep. On the final pass of a prompt, each layer's most-used experts — its fair share of the pool — are admitted, so decode starts on the prompt's hot set instead of cold. Passes shorter than 256 tokens, decode, and speculative verify passes gather over the pool as before, so nothing below the threshold changed. ### The numbers: 8k, prose, the floor, the ladder, and decode after a long prompt Every row is an interleaved A/B on the dev Mac between the commit before the sweep (`e09bcac`, the 0.2.2 code) and the sweep, same prompt, same flags, one model process at a time, real footprint sampled from `top` with a watchdog that kills a run under 2.5 GB reclaimable. Single runs here vary by 10 to 15%, so the headline is three rounds and the rest are one or two. **The 8k acceptance prompt** (the 7,960/8,073-token `verify.sh` prompt, three sentences repeated) at `--memory-gb 16`, which plans a 1024-token pass and 54 experts per layer: | round | 0.2.2 code | sweep | |---|---|---| | 1 | 100.5 tok/s | 191.8 | | 2 | 83.8 | 165.5 | | 3 | 89.7 | 195.0 | | **mean** | **91.3** | **184.1 (×2.02)** | | shipped build, two more rounds | 94.6 / 90.1 | 209.5 / 161.8 | | peak RSS | 13.0 GB | 13.2 | The split says where the ×2 came from: reads 30.4 → 25.7 s, the pool scatter 10.4 → 1.2 s (that is the copies of resident experts; the pool is never written), and compute 49.2 → 21.9 s, on the same 98,000 records read either way. **Ordinary prose** — a 34,000-character excerpt of PLAN.md, 10,490 tokens, at `--memory-gb 16`: | build | prefill | |---|---| | 0.2.2 code | 66.1 / 66.6 / 67.3 / 67.6 tok/s | | sweep, rows one at a time | 107.9 / 106.5 | | **shipped: sweep + parallel n-gram rows** | **140.1 / 139.2 (×2.1)** | Prose routes to nearly every expert of every layer (181,475 records against 98,872 for the same pass size on the acceptance prompt) and is the honest number for a pasted document. **Small targets**, the 7,960-token prompt, one round each: | target | 0.2.2 code | shipped build | |---|---|---| | 8.1 GB floor (13/layer, 256-token pass) | 50.8 tok/s, peak 7.7 GB | **93.1 tok/s, peak 6.2 GB** | | 10 GB (20/layer) | 41.3 tok/s, peak 8.6 GB (2026-08-31 record) | **87.7 tok/s, peak 7.2 GB** | At the floor the pass reads 826 GB of records in 31 passes; the reads went 83.4 → 56.3 s and the scatter 26.3 → 0.7, and the peak fell because the MLX buffer cache is capped while a small target reads a prompt (the knobs section). Both memory promises hold with more headroom than before. **`context-check --tokens 8192` at `--memory-gb 16`**, the dense synthetic prompt the tool reads: 0.2.2 code 64.2 tok/s (127.7 s, peak 13.1 GB); sweep 152.4 tok/s (53.7 s, peak 13.2), the shipped build 154.8, ×2.4. **The pass-size ladder** at a matched pool of 60 experts per layer (`--experts-per-layer 60`, the 8,073-token prompt), which is what the planner's estimator now carries: | pass | 0.2.2 code | sweep | sweep, peak | |---|---|---|---| | 256 | 45.6 tok/s | 87.5 | 14.1 GB | | 512 | — | 128.2 | 14.1 | | 1024 | 87.0 | 169.2 | 14.0 | | 2048 | — | 210.8 | 13.7 | | 4096 | 103.2 | 222.3 | 13.8 | And 4096 at the 16 GB target by `SLOTSTREAM_PREFILL_CHUNK` override: 246.7 tok/s at a 12.8 GB peak against 107.9 at 14.6. The estimator reads 85 / 125 / 165 / 205 / 220, rounded down; with it the planner now picks a 2048-token pass from a 20 GB target where it picked 1024, and the request-time scoring and `Tools/monotonic_plan.py` still hold across the 7 to 90 GB sweep. The full-context waits in `doctor` moved accordingly: 3.0 min for 32k on the 48 GB plan (was 5.5), 6.4 on the 16 GB tier (was 13.7). **Decode after a long prompt**, 48 tokens after the 8k prompt at 16 GB, three rounds: 5.62 / 6.00 / 5.52 tok/s with the final pass admitting the prompt's hot experts, 5.28 / 4.31 / 4.94 without. The N2 exit was ≥150 tok/s at 8k on the dev Mac. It is met at a 16 GB target, and by every pass size of 1024 tokens and up on a 60-per-layer pool; the auto plan itself is the one row still missing (next section). ### The gates: the sweep is the same computation The sweep computes the same thing the pool path computes, and the claim is gated the way prefix reuse is: not by byte identity, which re-batching the same tokens cannot give (MLX picks kernels and reduction orders by shape), but by the band that re-chunking a plain prefill already moves the logits by. `slotstream sweep-check` loads the model at the floor pool, reads a 549-token prose prompt, and requires five things. Its reading on the shipped code: | gate | reading | |---|---| | deterministic: the sweep twice on an empty pool | identical logits | | inside the band: sweep vs the pool path, one pass | **3.32%** of logit spread, against a prefill-rechunk control of 5.09% (pool path whole vs in 7-token passes) and a bound of 3× the control; top-1 same | | re-chunked: the sweep whole vs in 256-token passes | 3.15% | | blind to the pool: the sweep after the pool path loaded the prompt's experts (638 copied out of the pool instead of read) | **bit-identical** to the cold sweep | | admission leaves the pool consistent: after a generate whose last pass admitted the prompt's hot experts | pool path identical, sweep identical | The sweep moves the logits *less* than re-chunking the plain prefill does. The bit-identity on a warm pool is the invariant PLAN §6 asks for, restated for the sweep: whether an expert was copied out of the pool or read from the checkpoint, the same bytes reach the same kernel. That needed one deliberate piece of engineering. MLX takes its grouped kernel only when a call has at least 16 rows and four per expert of the weight array it is handed; with the pool as that array the rule would have flipped with the cache size, and with groups it would have flipped with how many of a group's experts were resident. So the sweep hands the kernel one group of at most 32 experts at a time and pads a short group up to the rule with repeats of its last row, dropped from the output. The kernel a row meets is then a function of the routing alone. What the band means for greedy text: on the 8,073-token acceptance prompt at a 1024-token pass the pool path answers `SEVENTEEN` and the sweep opens a reasoning chain — a near tie at the second token that the 3.3% moved. That is the same class of effect the prefill-rechunk control produces on its own: at a 256-token pass both paths answer with an empty line instead, on the same prompt, in the same session. `sweep-check` is in `Tools/verify.sh` next to `prefix-check`; the rest of the battery — the 0–1 layer parity gate, golden equivalence across cache sizes, `elastic-check`, `prefix-check`, `mtp-check`, the memory promises on the 7,960-token prompt, and `context-check` — runs unchanged, since decode, speculative verify passes, and any pass under 256 tokens still take the pool path exactly as before. ### What set the knobs: group size, lanes, recycling, the cache cap, admission, and the n-gram rows Every knob was set by an A/B on the 8,073-token acceptance prompt at a 16 GB target (a 1024-token pass, 54 experts per layer), one round per arm unless stated, interleaved, `SLOTSTREAM_SWEEP_TRACE=1` splitting the pass into reads, time spent waiting for the GPU, the CPU row sort, copies out of the pool, and the rest. **Where the time goes.** Reads 22.4 s, waiting for the GPU 1.6 s, sorting 0.02 s, copies out of the pool 1.0 s, everything else 20.1 s, for 43.5 s in all. The group loop is bound by the reads: the GPU finishes a group before the next one is in, so the sweep waits for it 4% of the time. The 20 s that are neither is serial by construction — the router, attention, and the layer's tail run while no read is outstanding, because the next layer's experts are not known until its router has run. Reads move at 11–13 GB/s against the 17.3 the SSD delivers on 2.7 MB records; runs are cut by the resident experts between them and by the six small scale/bias pieces per record. **Group size** (`SLOTSTREAM_EXPERT_LOAD_BATCH`, the experts per staging group): | group | prefill | peak RSS | |---|---|---| | 16 | 183.3 tok/s | 13.0 GB | | **32** | 169.9 | 13.2 | | 64 | 158.5 | 13.7 | | 128 | 160.6 | 14.1 | Reads are flat (22.9–24.1 s) at every size; the rest grows with the group, and so does the peak. 16 and 32 tie inside the run-to-run band (the 32 default read 174–195 in the paired rounds); the documented default stays at 32. **Read parallelism** (`SLOTSTREAM_IO_QUEUE_DEPTH`): 4 lanes read 39.1 s for 120.6 tok/s, 12 read 23.5 s for 179.3, 32 read 24.4 s for 162.5. The default of 12 stands, on contiguous runs as it did on the nine-piece reads. **Staging buffer recycling was built, measured, and dropped.** The hypothesis was that a fresh 88 MB set per group paid a page fault per 16 KiB and an unmap on release, some 270 GB of both per prompt. Recycling the sets through the arrays' finalizers changed the read time not at all (23.6 s against 22.7) and cost 6% of prefill (163.6 / 166.1 tok/s against 177.5 / 174.2 in two paired rounds) plus 0.3 GB of peak from the sets it held. Whatever the reads are waiting on, it is not page faults. **The MLX buffer cache is capped while a small target reads a prompt.** The sweep allocates arrays whose sizes vary from group to group, and MLX's cache keeps every freed size up to its 2 GB limit: the trace read the cache at 2.16 GB by the end of every long prompt, which is where the sweep's higher peak came from (13.2 against 13.0 GB at 16 GB; 7.9 against 7.7 at the floor). Capping it at 512 MB for the duration of the prompt costs 6% of prefill at 16 GB, where the memory does not matter, and nothing at the 8.1 GB floor, where the pass is read-bound — and there it took the 7,960-token prompt's peak from 7.9 GB to **6.2** on the shipped build (93.1 against 88.3 tok/s), and from 8.8 to 7.2 at a 10 GB target (87.7 tok/s). The engine therefore caps it only when the plan's expected peak is 12 GB or under; `SLOTSTREAM_PREFILL_CACHE_MB` forces a value at any target. **Admission warms decode, measured.** The last pass of a prompt writes each layer's most-used experts, its fair share of the pool, so decode starts on the prompt's hot set. Three interleaved rounds of 48 decode tokens after the 8k prompt: 5.62 / 6.00 / 5.52 tok/s with admission against 5.28 / 4.31 / 4.94 without — **5.7 against 4.8**, every round in the same direction, for about 0.2 s of copies on the final pass. It only ever touches the pool on that pass; every earlier pass of a long prompt leaves it alone, which is the scan resistance §3.3 asked for. **Prose was paying for its n-gram rows, one at a time.** The trace on a 10,490-token excerpt of PLAN.md read the same 54 s of "everything else" at a 1024-token pass and at a 4096-token pass, so it scaled with tokens, not passes; the acceptance prompt, three sentences repeated, read 20 s. A token needs sixteen ~100 B n-gram rows, three `pread`s each at the SSD's ~55 µs latency, and `NgramStore` fetched them one at a time on the calling thread; repeated text hid that behind the row cache and prose could not. Reading every missing row of a pass on 32 lanes before its embedding is assembled took the excerpt from 103.7 s to 80.1 s (101 → 131 tok/s); the rows are the same bytes in the same cache order, so nothing else moved. This cost was in the old path too — its 89 s of "compute" on the excerpt against 49 on the acceptance prompt was the same 35 s. ### What the sweep does not settle - **The auto plan on this Mac is not measured.** Every number above ran at a 16 GB target or below, because the 33 GB auto plan (a 4096-token pass, 152 experts per layer) needs ~32 GB reclaimable and the machine had 24–27 while this was measured. The closest runs are the 4096-token pass at a matched 60-per-layer pool (222 tok/s) and at the 16 GB target by override (247); the planner's 220 for 4096 is taken from the lower one. A quiet-machine `context-check --tokens 8192` at auto is the run that turns it into a measurement. - **Prose is the honest number and the planner's is the acceptance prompt's.** The estimator's ladder, like the one it replaced, is measured on three sentences repeated; ordinary prose routed to nearly every expert of every layer, read 4% more bytes at the same pass size, and came out ~30% slower even after the n-gram fix (131 against 184 tok/s at 16 GB). The tier rows and the full-context waits inherit that. - **The rest of the pass is serial, and the one lever left costs a layer of staging.** With the group loop read-bound and the GPU waiting on reads, the remaining time is the router, attention, and the layer tail, which run while nothing is being read because the next layer's experts are unknown until its router runs. A pass of 256 tokens or more routes to nearly all 512 experts of a layer, so reading layer L+1's whole set during layer L would be exact rather than speculative — the 2026-08-30 read-ahead, whose premise (spare IO capacity) the contiguous reads have now created. It would cost a layer of staging, 1.4 GB, which small targets do not have. Not built; the read-ahead decision stands until it is measured. - **Reads stop at 11–13 GB/s.** The SSD reads 17.3 on 2.7 MB records at queue depth 8 and up; the runs here average a few records between resident experts, and six of the nine pieces per record are 51 KB scale and bias rows. Recycling the staging buffers did nothing, so the gap is not page faults. An on-disk repack that put a record's nine pieces together (the skipped M2 container) is the other thing that would move it. - **Decode after a long prompt** is measured once, at 16 GB, on 48 tokens: 5.7 against 4.8 tok/s with and without admission. Whether the admitted set is the right one further into a reply, and what it does to a second long paste in the same conversation, is not measured. ## C1 — Mac mini M2, 16 GB, base storage (community, 2026-09-02) The first measurement from hardware that is not the dev Mac, and it contradicts the design consequence M0.5 drew from that machine's disk. Reported by `@flol` in [issue #5](https://github.com/carloslfu/slotstream/issues/5), raw output in [[sources/community/2026/09/2026-09-02-mac-mini-m2-16gb-flol]], machine in [[records/machines/mac-mini-m2-16gb]]. Mac mini M2, 16 GB, internal 256 GB, macOS 26.6.2, slotstream 0.2.2, auto plan (10.2 GB target, ~21 experts per layer planned, ~19 achieved; 11.7 of 17 GB was reclaimable, so auto sized down from its usual 10.7). | | measured | the plan said | |---|---|---| | Warm decode | **1.41 tok/s** (three identical runs) | ~4 tok/s | | Cold decode | 1.39 tok/s (128 tokens in 91.95 s) | — | | Prefill, 28-token prompt | 2.6 tok/s (io 8.66 s + scatter 0.80 s + compute 1.47 s) | ~40 tok/s at a 256-token pass | | Read throughput | **1.5 GB/s** (4,778 records, 13.2 GB) | 17.3 GB/s on the dev Mac | | Peak RSS | 6.1 GB | ~9.2 GB | **The estimate was not merely optimistic; it was below the disk's floor.** The pool's hit counters are zeroed after prefill (`Generate.swift`), so the reported 0.434 is the decode hit rate, not the run's. Decode routes 48 layers x 10 experts = 480 expert-uses per token, so 271.7 of them missed, at 2.7648 MB each: **751 MB read per token**. At 1.5 GB/s that is **501 ms of IO per token, a ceiling of 2.00 tok/s** before a single multiply. The planner quoted 3.8 tok/s at 19 experts per layer — 263 ms per token, half the time the reads alone take. No cache policy or kernel could have reached it. The rest of the step reconciles: 1.41 tok/s is 709 ms, so IO is 71% and the remaining 208 ms is compute and scatter — 2.4x the dev Mac's 86 ms plateau step, which is the right order for an M2 against an M5 Pro. **What this falsifies.** M0.5 concluded, from a 17.3 GB/s disk, that "even a zero-hit cache sustains ~13 tok/s from IO alone" and that "the binding constraint on small machines is memory and compute, not bandwidth". Re-run M0.5's own table at 1.5 GB/s: | h | miss/token | IO ms/token | IO-bound ceiling | |---|---|---|---| | 0.98 | 26.5 MB | 18 | 56 tok/s | | 0.90 | 133 MB | 89 | 11 tok/s | | 0.50 | 663 MB | 442 | 2.3 tok/s | | 0.00 | 1,327 MB | 885 | **1.1 tok/s** | On this machine bandwidth is the binding constraint, and the tier curve — a function of experts per layer alone, with no bandwidth term — cannot see it. M0.5's measurement of *its own* disk stands; only the generalization to small Macs is withdrawn. That the caveat was already written into M0.5 ("base-storage MacBook Airs will be far slower, and Stage C on real small Macs must re-measure") is why this is a confirmation of the caveat, not a surprise. **What this does not establish.** The drive's independent ceiling was not measured: 1.5 GB/s is what slotstream's `pread` path achieved, not a `Tools/coldread.c` figure. The dev Mac's pool path reaches only 4.5 of its 17.3 GB/s, so a path that inefficient here would imply a ~6 GB/s drive, which base storage is not — but that is inference from public specifications, and `coldread` on this machine is the measurement that settles whether the remaining loss is the disk or the nine 307 KB `pread`s per record. One cold run and three warm runs, one machine, one storage configuration. **The prefill number is not comparable to the plan's.** A 28-token prompt pays a nearly full cold expert fetch amortized over 28 tokens; the planner's ~40 tok/s describes a 256-token pass. Correcting for that still lands near 7 tok/s, not 40, but no 256-token pass was run here. The peak is likewise not planner slack: the 5.3 GB fixed footprint budgets a full 32k context that a 128-token run never allocates. **Still current after the prefill sweep.** The sweep takes passes of 256 tokens or more; this prompt was 28, and decode always takes the pool path, so both numbers describe `main` as well as 0.2.2. The one column the sweep would move is the long prompt — and that step failed, because `context-check` postdates the v0.2.2 tag while `docs/HARDWARE.md` already told reporters to run it. 0.2.3 shipped `context-check` on 2026-09-02, so that column is measurable on this machine now and is the one number worth re-running here: the sweep roughly doubled prefill on the dev Mac, but it did so by turning nine 307 KB `pread`s per record into contiguous runs, and a disk already near its sequential ceiling has far less of that to give. The prediction on record is that this machine gains from the grouped GEMM's share of the pass and little from the reads — well under the dev Mac's 2x. ## The 0.2.3 re-run, recorded 2026-09-24 `@flol` re-ran the whole procedure on 0.2.3 the day after the first report, built from source, and posted it as a [comment on issue #5](https://github.com/carloslfu/slotstream/issues/5#issuecomment-5525389653). It was not recorded at the time. The raw text is in [[sources/community/2026/09/2026-09-03-mac-mini-m2-16gb-flol-0-2-3-rerun]]. | | 0.2.2 report above | 0.2.3 re-run | |---|---|---| | Auto plan | 10.2 GB target, ~21 experts per layer | 10.7 GB target, ~25 experts per layer | | Warm decode, three identical requests | 1.41 tok/s | **1.48 tok/s** | | Cold decode, 128 tokens | 1.39 tok/s | 1.42 tok/s | | Cold reads, 28-token prefill | 13.2 GB at 1.5 GB/s | 13.2 GB at 1.5 GB/s | | Long prompt, 8,192 tokens | step failed | **12.1 min, 11 tok/s**; peak RSS 8.1 GB against a 9.7 GB plan | For the re-run the reporter closed every other app and heard no fans; 12.9 of 17 GB was reclaimable, against 11.7 GB for the first report, so auto planned a slightly larger cache. **The long prompt is the number this section asked for.** 8,192 tokens took 12.1 min, 7.6 times the ~1.6 min the planner printed before starting; its prefill estimate of ~85 tok/s has the same missing bandwidth term as the decode estimate. Once passes were running, the progress lines' remaining-time estimates (8.4, 5.7 and 2.9 min) followed the actual pace. **What it does not settle.** The prediction above, that this machine gains little from the sweep's reads, needs a long-prompt time on 0.2.2 to compare with, and 0.2.2 could not run `context-check`. Warm decode moved from 1.41 to 1.48 tok/s with a slightly larger cache, inside the run-to-run spread, so it is not a speedup claim. The hardware table keeps the 0.2.2 row and adds this run beside it. ## C2: MacBook Pro M5 Max, 128 GB (community, 2026-09-03) Reported by `@waterliu1981` in [issue #6](https://github.com/carloslfu/slotstream/issues/6). The original report and its follow-up are preserved in [[sources/community/2026/09/2026-09-03-macbook-pro-m5-max-128gb-waterliu1981]]. MacBook Pro 16-inch (Mac17,7), M5 Max, 128 GB, internal 2 TB SSD, macOS 26.6.2. The reporter described an idle machine. The auto plan was 34.6 GB with about 152 experts per layer and speculative decoding enabled. The first report used Slotstream 0.2.1 and summarized warm decode at about 19 to 21 tok/s. The same author then remeasured with Slotstream 0.2.3, reporting a checksum-verified binary replacement and unchanged weights. That [follow-up](https://github.com/carloslfu/slotstream/issues/6#issuecomment-5520489176) is the source of the current hardware row: | Repeated 256-token request | Reply speed | |---|---| | 1 | 21.02 tok/s | | 2 | 21.53 tok/s | | 3 | 22.11 tok/s | The report summarizes this as **21–22 tok/s** with the auto plan and speculative decoding. Two longer warm runs returned 22.83 and 22.10 tok/s. The current surface uses the reported range instead of a best run. The report's prefill and peak figures were planner estimates, not measured long-prompt speed or process RSS. Keep both columns unmeasured. The manual cache-size sweep is reported separately below so its gains are visible without attributing them to automatic sizing. This is one community report, not an independent rerun or a comparison made under the same conditions as the M5 Pro and M2 measurements. It supports a machine-specific row, not a promise for all Macs with that memory capacity. ## Larger-cache results surfaced on 2026-09-13 Rechecked issue #6 and its follow-up through the live GitHub API on 2026-09-13. The existing immutable source already preserves the complete sweep. The public tables had retained only auto, omitting evidence that more allocated memory helped on this same machine. | Total-process target | Experts per layer, as reported | Slotstream 0.2.3 warm decode | |---|---|---| | 34.6 GB (auto) | ~152 | ~21–22 tok/s | | 48 GB (manual) | ~253 | ~26.9 tok/s | | 73 GB (manual) | ~401–441 | ~31.5 tok/s | All rows are the same 128 GB M5 Max, internal 2 TB SSD, with speculative decoding enabled. Targets are decimal GB budgets, not observed peaks or requirements for installed memory. Preserve the reporter's approximate expert-count range rather than replace it with today's planner output. The original 0.2.1 sweep also reported gains at larger targets; the current public comparison uses only the follow-up's 0.2.3 values. This within-machine comparison is evidence of a benefit from increasing the memory target, beyond the different-chip comparison against the M5 Pro. The manual rows are approximate summaries without the individual repeated request timings supplied for auto. They are not an independently reproduced, interleaved benchmark or a universal throughput curve. No new process-memory, long-context, correctness, or 0.2.16 performance qualification is established. Do not multiply these figures by the development Mac's later release speedup. The README's earlier flat 13.5 tok/s values extrapolated the *scope* of a 20 GB-target development-Mac measurement while holding its numerical value constant. Replace them with named measured configurations and explicit gaps. Retaining the conservative automatic target does not negate this community result or establish an optimum on larger Macs. ## C3: MacBook Air M5, 32 GB (community, 2026-09-07) Reported by `@arczhi` in [issue #12](https://github.com/carloslfu/slotstream/issues/12), preserved in [[sources/community/2026/09/2026-09-07-macbook-air-m5-32gb-arczhi]]. MacBook Air M5 (2026), 32 GB, 1 TB SSD, macOS 26.6.2, reported Slotstream 0.2.11. The report does not specify whether the model was on internal or external storage. It uses a 22 GB plan, with about 75 experts per layer planned and a 2048-token prefill chunk. Three identical requests to one running server returned **6.29, 6.28, and 6.22 tok/s**. The public hardware row uses **6.22 tok/s**, the third request, matching the measurement procedure. The short cold run returned 6.60 tok/s and a 15.7 GB process RSS peak; it is not the warm result or long-prompt peak. The long-prompt command explicitly sets a 22 GB target, vision off, MTP off, 8192 tokens, and physical-footprint sampling. The reported JSON completed without aborting: 8192 prefill tokens in 64.8707 seconds at 126.28197 tok/s, process RSS peak 17.75475 GB, and sampled physical-footprint peak 20.58214 GB. The full hardware row rounds prefill to **126.28 tok/s** and process RSS to **17.75 GB**. RSS and physical footprint are different metrics, not interchangeable versions of the same peak. The exact command and output remain in the source. The reported maximum context and memory plan differ from the default setup; these results do not establish the cost of the default or every larger conversation. The warm requests' full launch command and system load were not supplied. This review verified that `--sample-footprint` exists in the published v0.2.11 source, but did not independently rerun the reporter's binary. This adds a real 32 GB Mac to the hardware reports. It does not turn the planner's roughly 9 tok/s estimate into a measurement, or isolate the effects of chip, cooling, storage, context, and settings from one another. ## C4: MacBook Pro M3 Max, 64 GB (community, 2026-09-16) Reported by `@merken` in [issue #20](https://github.com/carloslfu/slotstream/issues/20), preserved in [[sources/community/2026/09/2026-09-16-macbook-pro-m3-max-64gb-merken]]. MacBook Pro 14-inch (2023), M3 Max (`applegpu_g15s`), 64 GB, 512 GB SSD, macOS 27.0, Slotstream 0.2.18. The report does not say whether the SSD is internal or describe other load. Auto planned a 48.1 GB target with about 119 experts per layer, speculative decoding, the decode lookahead and a 262,144-token window. | | Reported | |---|---| | Warm decode, three identical requests | 11.68, 12.49 and **12.38 tok/s** | | Cold decode, 128 tokens | 11.47 tok/s; 74 of 106 drafts accepted | | Cold reads, 18-token prefill | 9.4 GB of experts at 6.1 GB/s | | Long prompt, 8,192 tokens at context-check's 34.6 GB target | 39 s, **213 tok/s**; process peak 30.1 GB against a 33.6 GB plan | The hardware row uses **12.38 tok/s**, the third request. The planner estimated about 11 tok/s for the served plan. The warm requests' prefill rates of about 34 million tok/s are an artifact of the measurement recipe, not a prefill result: the repeated prompt was reused whole, and `prompt_eval_count` counts reused tokens while `prompt_eval_duration` times only what was read. **Below the band's estimated floor.** This is the first report from a Mac with 48 to less than 96 GB other than the development Mac, and it sits below the ~15 tok/s floor the README gives that band. The 64 GB M4 Max in [[records/measurements/c5-macbook-pro-m4-max-64gb-community]] ran the same auto plan and, by its logs, the same pre-0.2.19 decode forecast, and decoded at 15.93 tok/s. The cold decode splits put the difference in both parts of the step: reads took 4.00 s here against 2.69 s there, over 5,351 records against 4,424, and the rest of the step took 7.1 s against 5.6 s. That points at the older chip and the smaller SSD more than the release, though the two runs also differ in release (0.2.18 and 0.2.22) and neither was rerun. The 1.10x that 0.2.19's corrected forecast measured on the development Mac has not run here. The floor stays until this Mac is rerun on the current release, with `slotstream pull` run first; the hardware guide names this report beside the range. One report, one run of each step, not rerun by the author. ## C5: MacBook Pro M4 Max, 64 GB, internal and external SSD (community, 2026-09-19) Reported by `@YenHub` in [issue #22](https://github.com/carloslfu/slotstream/issues/22) (internal SSD) and [issue #23](https://github.com/carloslfu/slotstream/issues/23) (external SSD), preserved in [[sources/community/2026/09/2026-09-19-macbook-pro-m4-max-64gb-internal-yenhub]] and [[sources/community/2026/09/2026-09-19-macbook-pro-m4-max-64gb-external-yenhub]]. The same reporter proposed the hardware rows in [pull request #25](https://github.com/carloslfu/slotstream/pull/25). MacBook Pro 16-inch (November 2024), M4 Max (`applegpu_g16s`), 64 GB, macOS 27.0, Slotstream 0.2.22, each run after a fresh reboot with only a terminal open. Both reports used the same auto plan: a 48.1 GB target with about 119 experts per layer, speculative decoding, the decode lookahead and a 262,144-token window. The first read the model from the internal 1 TB SSD, the second from a Crucial X10 Pro 1 TB external SSD over USB 3.2 Gen 2 (10 Gb/s). | | Internal SSD | External SSD | |---|---|---| | Warm decode, three identical requests | 14.95, 16.22 and **15.93 tok/s** | 2.81, 3.04 and **2.98 tok/s** | | Cold decode, 128 tokens | 15.30 tok/s | 2.89 tok/s | | Cold reads, 28-token prefill | 13.2 GB at 7.6 GB/s | 13.2 GB at 0.9 GB/s | | Long prompt, 8,192 tokens at context-check's 34.6 GB target | 30 s, **270 tok/s**; process peak 30.1 GB | 2.7 min, **51 tok/s**; process peak 30.2 GB | The rows use the third requests, 15.93 and 2.98 tok/s. The warm prefill rates in the millions of tok/s are the recipe artifact described in [[records/measurements/c4-macbook-pro-m3-max-64gb-community]]. **Same Mac, plan and release, 5.3 times slower from a 10 Gb/s drive.** This is the first pair in the store that isolates the disk. The external drive read 0.9 GB/s against the internal SSD's 7.6 GB/s, warm decode fell from 15.93 to 2.98 tok/s, and the long prompt fell by the same factor, from 270 to 51 tok/s. The planner assumes a disk like the development Mac's and printed about 11 tok/s for both. **Both runs used the pre-0.2.19 decode forecast.** Each context-check log prints `[expert-lookahead] boundary forecast: no correction at lookahead/tap-correction-attention-rank128-v1.safetensors`: the 37.5 MB correction file was absent, so the engine ran the earlier forecast. Through 0.2.24 only `slotstream pull` fetched that file; a model downloaded before 0.2.19, or through the download `slotstream run`, `serve` or `launch` offers on first use, lacked it. 0.2.19's 1.10x was measured on the development Mac with the file present and is not applied to these numbers. One reporter, one run of each step on each disk, not rerun by the author. ## C6: MacBook Pro M4 Max, 36 GB (community, 2026-09-20) Reported by `@JohnClarkson` in [issue #26](https://github.com/carloslfu/slotstream/issues/26), preserved in [[sources/community/2026/09/2026-09-20-macbook-pro-m4-max-36gb-johnclarkson]]. MacBook Pro (November 2024), M4 Max (`applegpu_g16s`), 36 GB, internal 1 TB SSD, macOS 26.0.1, Slotstream 0.2.22, run just after a reboot with a terminal and one screen-sharing window open. Auto planned a 27.1 GB target with about 90 experts per layer, speculative decoding, the decode lookahead and a 65,536-token window. | | Reported | |---|---| | Warm decode, three identical requests | 8.48, 8.61 and **8.41 tok/s** | | Cold decode, 128 tokens | 8.63 tok/s; 74 of 106 drafts accepted | | Cold reads, 28-token prefill | 13.2 GB of experts at 5.5 GB/s | | Long prompt, 8,192 tokens at context-check's 27.1 GB target (about 121 experts per layer) | 49 s, **166 tok/s**; process peak 24.8 GB against a 26.1 GB plan | The hardware row uses **8.41 tok/s**, the third request; the planner estimated about 9 tok/s. This is the first 36 GB report. Like [[records/measurements/c5-macbook-pro-m4-max-64gb-community]], its context-check log prints `no correction at lookahead/tap-correction-attention-rank128-v1.safetensors`, so it ran the pre-0.2.19 decode forecast. The warm prefill rates in the millions of tok/s are the recipe artifact described in [[records/measurements/c4-macbook-pro-m3-max-64gb-community]]. One report, one run of each step, not rerun by the author. ## C7: MacBook Pro M4 Pro, 24 GB (community, 2026-09-25) Reported by `@davidcavazos` in [issue #41](https://github.com/carloslfu/slotstream/issues/41), preserved in [[sources/community/2026/09/2026-09-25-macbook-pro-m4-pro-24gb-davidcavazos]]. MacBook Pro (2024), M4 Pro (`applegpu_g16s`), 24 GB, 512 GB SSD, macOS 26.6.2, Slotstream 0.2.24, run after a fresh boot with one or two terminals, Safari, Stats and Activity Monitor open, and no swap before or after. Auto planned a 15.9 GB target, sized down from the usual 18.0 GB because 17.4 GB was reclaimable, with about 53 experts per layer and a 32,768-token window. At that window the 0.2.24 plan ran without the draft head and without the decode lookahead. | | Reported | |---|---| | Warm decode, three identical requests | 3.61, 3.52 and **3.57 tok/s** | | A second warm round, posted later | 3.85, 3.97 and 3.95 tok/s | | Cold decode, 128 tokens | 3.17 tok/s | | Cold reads, 28-token prefill | 13.1 GB of experts at 3.7 GB/s | | Long prompt, 8,192 tokens at context-check's 18.0 GB target (about 78 experts per layer) | 1.5 min, **93 tok/s**; process peak 16.6 GB against a 17.0 GB plan | The hardware row uses **3.57 tok/s**, the third request of the first round; the planner estimated about 8 tok/s. This is the first 24 GB report, and it falls below the 24 to less than 48 GB planning range of ~6–16 tok/s. Two known differences from the development Mac may account for the gap. The planner assumes an SSD like the development Mac's 17.3 GB/s, while this 512 GB SSD read cold experts at 3.7 GB/s. And 0.2.25 enables the draft head, with streamed experts, and the decode lookahead at 24 GB, which 0.2.24 did not. A rerun on 0.2.25 would separate the two. While using the server from Pi, the reporter saw disk reads of about 2 GB/s; the server reported no tok/s there. One report: two warm rounds and one run of each other step, not rerun by the author. ## The 0.2.25 re-run, recorded 2026-09-27 `@davidcavazos` re-ran the procedure on 0.2.25 and posted it as a [comment on issue #41](https://github.com/carloslfu/slotstream/issues/41#issuecomment-5850068364), preserved in [[sources/community/2026/09/2026-09-26-macbook-pro-m4-pro-24gb-davidcavazos-0-2-25-rerun]]. It was not a fresh boot: other apps held about 5.5 GB, and 419 MB of swap stayed unchanged before and after the tests. | | 0.2.24 report above | 0.2.25 re-run | |---|---|---| | Auto plan (`doctor`) | 15.9 GB target, ~53 experts per layer; no draft head or decode lookahead | 17.4 GB target, ~58 experts per layer; draft head and decode lookahead on | | Warm decode, three identical requests | 3.61, 3.52 and **3.57 tok/s** | 5.58, 4.92 and **5.41 tok/s** | | Cold decode, 128 tokens | 3.17 tok/s | 5.57 tok/s; 74 of 106 drafts accepted | | Cold reads, 28-token prefill | 13.1 GB at 3.7 GB/s | 13.1 GB at 3.7 GB/s | | Long prompt, 8,192 tokens at context-check's 18.0 GB target | 1.5 min, 93 tok/s, ~78 experts per layer; peak 16.6 GB against 17.0 GB | 1.6 min, 86 tok/s, ~72 experts per layer; peak 16.7 GB against 17.0 GB | The hardware row now uses **5.41 tok/s**, the third request and the round's median. On 0.2.25 this Mac decoded about half again as fast: the draft head accepted 70% of its drafts, and the long-prompt log shows the corrected decode forecast loaded. It is still below the planner's ~8 tok/s and below the ~6 floor the 24 to less than 48 GB range had, which now rounds down to ~5 ([[records/measurements/hardware-planning-ranges-2026-09-13]]). The cold read rate did not move: 3.7 GB/s through the engine, and the reporter saw 2 to 3 GB/s of disk reads during warm decode. The first report set that against the development Mac's 17.3 GB/s, which is a raw SSD figure. Through the engine the development Mac read cold experts at 11.5 GB/s in [[sources/runs/2026/09/2026-09-05-optimization-cache-policy-screen]] and 12.6 GB/s in [[sources/runs/2026/09/2026-09-18-memory-budget-native-verification]], so this Mac reads about a third as fast, not a fifth. Both plans were also sized down because other apps held memory, to 15.9 and 17.4 GB from the usual 18.0 GB. One re-run cannot separate the SSD from the plan size. The plan column is `doctor`'s plan, as the hardware guide defines it. The warm server's own plan was not posted either time. The first report's cold run held about 38 experts per layer; the re-run's held about 58, matching its `doctor` plan. One re-run, one round of each step, not rerun by the author. ## Decode: where the time goes, and the two knobs that moved it (2026-09-03) Decode had no equivalent of the prefill split, so "decode is slow" could not be attributed without guessing. `run` now prints one, and it says decode at a small cache is not mostly compute. On 48 tokens at 30 experts per layer, hit rate 0.576: | part | seconds | share | |---|---|---| | reading expert records | 3.24 | 44% | | scattering them into the pool | 1.47 | 20% | | everything else (compute and dispatch) | 2.66 | 36% | The scatter is the surprise. It moves 203 records per token, 561 MB, in 30.6 ms — about 18 GB/s, against the 49 to 75 GB/s the slot-write microbenchmark measured in both Python and Swift (§M0.3). The gap is not the copy: it is one full GPU sync per layer, 48 per token, because `SlotPool.ensure` finished every batch with `eval(pools)`. The gather that follows in the same layer already depends on those arrays, so MLX orders it correctly without the sync; all the sync bought was releasing the staging buffers a layer earlier. **Two knobs, measured separately, then together.** Interleaved rounds, greedy, one prompt, output compared byte for byte. *Read lanes for the pool path.* The 12 lanes the sweep uses were tuned on its long contiguous runs, which are throughput-bound. A layer's handful of decode misses is nine ~307 KB pieces per record, which is latency-bound, so 12 lanes leave the queue empty between waves. Three rounds at 30 experts per layer: 12 lanes 6.82 tok/s, 32 lanes 7.14, 64 lanes 7.16. The two paths now carry separate numbers (`SLOTSTREAM_IO_QUEUE_DEPTH` stays 12 for the sweep, `SLOTSTREAM_POOL_QUEUE_DEPTH` is 32) rather than one compromise: raising the sweep's lanes to 32 measured *slower* on 2026-09-02 (179 → 162 tok/s). *Finishing the scatter.* Three modes, three rounds each: `sync` (the shipped behaviour) 6.17 tok/s at a 7.5 GB peak, `async` 7.11 at 8.9 GB, `none` 7.06 at 7.8 GB. `async` buys nothing over `none` and costs 1.1 GB more, so the default is `none`. *Together*, five interleaved rounds, 96 tokens, on a quiet machine: | | decode | peak RSS | |---|---|---| | sync scatter, 12 lanes | 6.93 tok/s | 7.5 GB | | lazy scatter, 32 lanes | **7.63 tok/s** | 7.8 GB | **×1.10, and the output is byte-identical across all ten runs.** The same pair measured ×1.14 earlier in the day while the machine was building; ×1.10 on a quiet machine is the number to quote. The 0.3 GB is real and is the honest price: staging buffers and their allocator-cache churn live until the layer's `eval(h)` instead of being released mid-layer. `--memory-gb 10` still peaks at 7.3 GB on the 7,960-token prompt, so the promise holds with the same headroom it had. One caveat the split itself carries: with the lazy default the `scatter` column no longer measures the scatter, only the cost of issuing it. The work reappears in `compute`, which is why that column rises from 5.87 s to 8.45 s while the total falls from 16.37 s to 14.65 s in the same pair. Not measured: whether the same two knobs help at the 120-to-150-experts-per-layer sizes auto reaches, where the miss path is a smaller share of the token. Those configurations need ~26 GB reclaimable and the memory-safety rules keep test runs between 8.1 and 10 GB, so this is a small-cache result. ## The pass peaks on a plateau, not on one transient (2026-09-03) Two changes that should each have taken a few hundred MB off a prefill pass moved the measured peak by **0.00 GB**. Both were built, measured, and kept only where they cost nothing. The reason they failed is the useful part, and it was found by measuring rather than by a third guess. **Why the attention transient looked like the target.** MLX 0.31.1 admits the fused prefill kernel only for head dims 64, 80 and 128 (`sdpa_full_supported_head_dim` in `scaled_dot_product_attention.cpp`). The QSA layers run at head dim 256, so every pass longer than 8 tokens takes the unfused path in `fast.cpp`, which materialises the whole `[24, pass, context]` score matrix. That transient grows with pass × context, which is exactly what `PrefillSchedule` halves the pass to stay ahead of. **Query blocking is exact, and faster in isolation.** The fallback aligns queries to the end of the keys (`arange(kL - qL, qL + (kL - qL)) >= arange(0, kL)`), so a block `[lo, hi)` of a pass starting at context `base` sees exactly keys `[0, base + hi)` and reproduces the same mask rows. `AttnProbe` on the real shapes, no weights: | pass over context | whole | blocked 512, eval per block | difference | |---|---|---|---| | 4096 over 8,016 | 55.4 ms, 1.83 GB | 42.4 ms, 0.40 GB | 0.00e+00 | | 4096 over 32,768 | not built (6.4 GB) | 226.7 ms, 1.16 GB | 0.00e+00 | Bit-identical at every block of 256 and up; a block of 128 measured 1.6e-3 of logit spread, which is why 256 is the floor. **The per-block `eval` is load-bearing**: without it MLX builds the whole graph before evaluating and holds every block's score matrix at once — 6.5 GB at 4096 over 32,768, exactly what not blocking costs. **End to end it bought nothing.** Interleaved rounds on the 7,960-token prompt at a pinned 20-experts-per-layer pool, peak RSS: | pass | whole | blocked | |---|---|---| | 512 | 7.35 GB | 7.40 GB | | 1024 | 7.70 GB | 7.75 GB | | 2048 | 8.50 GB | 8.50 GB | and a 16,384-token `context-check` read 8.58 GB whole against 8.64 GB blocked, 200.7 tok/s against 208.7. Output byte-identical throughout. **The phase trace says why: the pass peaks on a plateau, not on one transient.** `SLOTSTREAM_MEM_TRACE=1` evaluates each phase of each layer and records the high-water active memory. At a 2048-token pass, against a 7.60 GB layer-end baseline: | phase | high-water | with attention blocked | |---|---|---| | QSA attention | 8.57 GB | **7.99 GB** | | PLE (layer 1) | 8.57 GB | 8.57 GB | | MoE sweep | 8.25–8.32 GB | 8.25–8.30 GB | | hyper-connections | 7.82–7.93 GB | unchanged | Blocking removes 0.58 GB from attention exactly as the probe predicts, and the process peak does not move because the PLE layer holds the same 8.57 GB. Four transients sit within 0.6 GB of each other, so **no single one of them is the peak, and lowering one alone can never lower the process**. The same reasoning explains the second null: not gathering the sweep's replicated rows up front (rows × hidden, 105 MB at a 2048-token pass) measured 8.50 GB against 8.50 GB. **What shipped, and why.** Query blocking is kept but its threshold is set so it is a **no-op at every configuration the planner produces today**: it engages only above `PrefillSchedule.measuredQueryKeyProduct`, which is where the schedule already shrinks the pass. That keeps the measured envelope unchanged while capping a transient that would otherwise grow without limit as the context cap rises. The per-group row gather is kept because it holds less for the same result and is bit-identical (`sweep-check` reads 3.320% against the same 5.089% control, unchanged). Neither is an optimisation and neither should be quoted as one. **What this says about the next attempt.** Lowering the pass's peak means bounding attention, PLE and the MoE sweep together; any one of them alone is wasted work. The phase trace is the tool for checking that, and it should be run before, not after. The measured marginal cost of a pass is also lower than the planner charges — the peak rose 0.35 GB from a 512 to a 1024-token pass and 1.15 GB from 512 to 2048, about 0.78 MB per chunk token against the 1.30 MB the planner charges — but that is a target-driven number and re-anchoring it needs its own runs at 8.1, 10 and 16 GB before the constant moves. ## M1 closed: expert locality on a real trace, and the eviction policy is not the lever (2026-09-03) M1 has been open since 2026-08-28. The simulator (`Tools/cachesim.py`) and a collector for the Python reference were built then, but no trace was ever taken, because collecting one needs a bounded forward pass and at the time none existed. The engine has been able to produce one for days; the decode split (above) is what made it worth doing, since it puts 44% of decode time at a small cache in the expert reads, and which records the cache keeps is what sets that. **The trace.** `SLOTSTREAM_ROUTER_TRACE` records every routing decision. `slotstream run --experts-per-layer 30 --max-tokens 220 --greedy` on a mixed prose-and-code prompt: 10,608 MoE calls, 123,360 expert-uses, engine hit rate **0.557**. Only the 220 single-token decode steps are simulated — a prefill pass of 256 tokens or more takes the sweep, which never reads or writes the pool, so its routing is not cache traffic. **The shape of the workload.** 220 steps, 105,600 expert-uses, **9,956 distinct records touched of 24,576** (40.5%), and the top 10% of records serve **70.9%** of accesses. Two numbers follow. The working set for this one generation is 26.9 GB, which is why a 4 GB cache misses so much. And because 9,956 of the 105,600 uses are first touches, **the compulsory-miss ceiling is a hit rate of 0.906** — no policy and no size can beat it on this trace. **The policy question, answered.** At the engine's own sizes: | slots | experts/layer | CLOCK (engine, measured) | LRU | LFU-decay | hot+LRU (offline) | |---|---|---|---|---|---| | 960 | 20 | — | 0.473 | 0.337 | 0.515 | | **1440** | **30** | **0.557** | 0.568 | 0.480 | 0.603 | | 2880 | 60 | — | 0.707 | 0.708 | 0.752 | **CLOCK is already within about one point of LRU and well ahead of LFU-decay, so there is no eviction-policy win to collect here.** The `hot+LRU` column pins the globally hottest records chosen from the whole trace, so it is an offline upper bound rather than an implementable policy, and even it is only 4.6 points ahead at the size that matters. M1's remaining question — "cheapest adequate policy wins" — resolves to the one already shipped. What the numbers do point at is **capacity and warm start, not policy**: at 30 experts per layer the cache runs at 0.557 against a 0.906 ceiling, and LRU/LFU only reach that ceiling around 27 GB of pool. The concentration figure is the one lever the engine does not yet use — 10% of records serving 71% of accesses is what a persisted hot set, preloaded at startup, would exploit against the cold-start cost. The prefill sweep already admits a prompt's hot experts on its final pass; nothing carries a hot set across processes. Scope: one trace, one prompt, one machine, decode only. A second workload could move the concentration figure, and an agentic trace of many short turns over one prefix — the shape §8.1 cares about — is not covered. ## V1 — the vision tower: what it costs, what a picture costs, and whether it computes the right thing (2026-09-03) The checkpoint has always carried a vision tower — 333 `vision_tower.*` tensors that `Weights.swift` skipped by name — and the chat template has always rendered an image part to `<|vision_start|><|image_pad|><|vision_end|>`. What was missing was the tower itself, the splice, and the accounting. This is what it costs and how it was checked. ## What it costs | | | |---|---| | Tower weights | **0.898 GB**, 333 tensors, bf16 (unquantized) | | Load time | 3.4 s, from a cold process, reading only those tensors | | Charged to the memory plan | **no** — see below | | One 846x859 photograph | **702 tokens** (52x54 patches, 2x2 merged), plus the template's two sentinels | | One 1206x1570 photograph | 1,862 tokens | | Largest picture accepted | 2,304 tokens (the 1536² engine cap) | | `run --image` at `--memory-gb 10` | peak **8.5 GB** RSS, tower included (announced peak 9.0 GB) | The tower's bytes come from the checkpoint's own header (`VisionTower.residentBytes`), not from a measurement of the process, so the number cannot drift from the artifact. `Planner.visionResidentGB` rounds it up to 0.9. **It is a conditional charge, and that is a decision, not an oversight.** Every published memory number — the README's tier table, the 32 GB peak, the planner goldens — is measured against a plan with no tower in it, and most runs never send a picture. Folding 0.9 GB into the fixed footprint would move all of those for everyone, to buy a capability most requests do not use. So the plan states the cost instead of paying it, `serve` prints it as a conditional line, and `Engine.ensureVisionTower` checks the machine can afford it at the moment the first image arrives — refusing with a 400 rather than overcommitting. What is not allowed is the third option, which the first implementation took: allocate it silently and leave the printed plan wrong by a gigabyte. ## Whether it computes the right thing A vision tower fails silently. A transposed weight, a rotary embedding laid out in the wrong half, a merger norm applied after the 2x2 shuffle instead of before — each still produces embeddings of the right shape, and the model still writes fluent sentences about a picture it did not see. Nothing downstream notices. So `Tools/vision_ref.py` implements the tower a second time, from the transformers reference rather than from `Vision.swift`, and the two are compared on the same pixel tensor. | | cosine | worst token | |---|---|---| | slotstream (bf16) vs numpy float32 | 0.99870620 | 0.898964 | | slotstream (bf16) vs mlx float32 | 0.99870270 | 0.919064 | | slotstream (bf16) vs mlx bfloat16 | 0.99878263 | 0.950250 | | **mlx bfloat16 vs numpy float32** | **0.99840382** | 0.839186 | | mlx float32 vs numpy float32 | 0.99996241 | 0.997329 | Read the fourth row first. **The same reference implementation, at the two dtypes, disagrees with itself more than slotstream disagrees with either of them.** The tower is 27 residual blocks deep in bfloat16 — 8 mantissa bits — and the tokens the two dtypes differ on most are the low-norm ones, where a tiny absolute error is a large angle. The last row is what makes the rest trustworthy: two float32 implementations that share no kernels agree to 0.99996. So the gate is not an absolute tolerance. It requires the float32 implementations to agree, slotstream to sit inside the band the dtype itself spans, and slotstream to match the bfloat16 reference. That is the same shape of argument `prefix-check` makes for reuse against a cold rebuild, and for the same reason: an equality gate here would either prove nothing or never pass. **One thing that looked like a lever and is not.** The rotary angles are computed in Float and cast to bfloat16, which carries about two significant decimal digits of a cosine. Keeping them in float32 and rotating there was tried on the theory that this was the largest avoidable error in the tower. Measured, it made agreement slightly *worse* — 0.99847 against 0.99870 — because at this depth the residual stream's own rounding dominates and the two errors were partly cancelling. Reverted; do not re-derive it. ## Whether the rows land in the right places Parity proves the tower. It says nothing about whether its output reaches the language model at the positions the template reserved, which is the other half of the feature and the half a fluent answer hides. That is what the serving suite is for: the model has to name what is in the photograph. - **Both pictures, named.** A close-up of a dog: "Dog nose close-up". Green citrus on a tree: "Green fruit growing." - **Two pictures in one turn, in order.** Asked for the subject of the first and then the second, the answer is "Dog, pomelo" — the second image really is a pomelo tree. A swapped pair would have been equally fluent. - **The count is exact.** A request with the dog costs 704 more prompt tokens than the same request without it: 702 placeholders plus `<|vision_start|>` and `<|vision_end|>`. - **Speculative decode sees the picture too.** `mtp-check`'s vision leg produces output identical to plain decode over 24 tokens, at a 93.8% overall accept rate. A draft head fed the placeholder's own embedding rather than the tower's row would be self-consistent and blind, and would diverge. ## What broader testing found Two photographs are not a test. A set of images whose content is known exactly — four colour quadrants, five ordered bands, counted squares, rendered text, a 40x40 thumbnail, a 4:1 panorama, the same picture as PNG and as JPEG, a greyscale JPEG — was put through the same path, with one open "describe this" and one checkable question each. 16 of 16. The ones worth naming, because each would have been invisible to a suite that only asked for a description: - **Spatial layout is right.** Unprompted, the model reported "red (top-left), green (top-right), blue (bottom-left), and yellow (bottom-right)" — all four corners, correctly. A tower whose patches were ordered by row instead of by merge block, or whose rotary halves were swapped, describes four coloured squares just as fluently and puts them in the wrong corners. - **Vertical order is right.** Five bands read back "black, red, white, blue, green", top to bottom, exactly. - **Text is legible.** A hand-drawn bitmap "SLOT 42" came back as `SLOT 42`. - **Counting is not a lucky guess.** Three squares answered 3, five answered 5. - **The size extremes hold.** A 40x40 image, upscaled to the processor's minimum, still reads as magenta; a 1200x300 panorama places green on the left and orange on the right. **One real defect, which only an image with transparency could show.** The decoder drew into a premultiplied CoreGraphics context over fresh memory, so every transparent pixel arrived black. On a photograph this is invisible — photographs have no alpha. On the images people actually paste into a chat — a logo, a chart, a diagram, a screenshot exported with transparency — it is not: black text on a transparent background reached the model as black on black, and it answered *"the image is entirely black, with no discernible features or content."* The picture did not survive the decoder. The fix is two lines: fill the context white before drawing, which is what every viewer composites onto and therefore what the sender saw. The same image now answers `SLOT 42`. Opaque images are unaffected, byte for byte — the parity dump's pixel tensor hashes identically before and after — and the property is pinned weights-free in `vision-check`: a transparent pixel must normalize to +1, never -1. ### Formats, and two more defects A second pass fed the finished path the files a decoder gets wrong rather than the ones it gets right: nine containers (TIFF, BMP, JPEG 2000, HEIC, PSD, TGA, GIF, interlaced PNG, palette PNG), 8-bit greyscale, 16-bit RGB, an animated GIF, a 1x1 image, 199:1, a 2400x1800 downscale, and five files that should be refused — truncated, text with a `.png` name, zero bytes, a 40000x40000 header with one row of data, and 400:1. Every container came back with the same correct answer, and the refusals refused. Two things did not. **EXIF orientation was ignored.** A phone stores its sensor's pixels and a tag saying which way is up. Every viewer applies it; so does the reference (`transformers.image_utils.load_image` calls `ImageOps.exif_transpose`). slotstream did not, so a photograph tagged `Orientation=6` — which is most portrait photographs — reached the model rotated, and it named the stored corner rather than the displayed one. Four of the eight orientations also swap the axes, so the token count was wrong with it. All eight are now applied and pinned corner by corner against EXIF's own table in `vision-check`; the transverse case was wrong on the first attempt and that table is what caught it. **A truncated file was answered rather than refused.** ImageIO is lenient by design: half a PNG decodes to the rows it has plus blank space, and — measured, not assumed — `CGImageSourceGetStatusAtIndex` reports it *complete*, because all the data it was handed arrived at once. So an upload cut short by a dropped connection came back as a confident description of a mostly empty picture. The container's end marker is the signal ImageIO does not give: `IEND` for PNG, the end-of-image marker near a JPEG's tail (editors append after it), `;` for GIF, and no judgement at all for containers without an unambiguous terminator. ### What the request layer and the shared state survive Sixteen protocol cases: malformed `images` fields, an image part with no url, a `data:` URL that is not base64, `ftp://`, an unknown part type, an image in a system turn (a 400 carrying the template's own reason, not a 500), eight images in one turn all charged for, thirty images past the context ceiling (a 400 naming the cap, not an allocation), streaming, `think: true`, and the same request twice byte for byte. Six state cases, which is where a wrong cache key hides: five simultaneous first images against a server whose tower has never loaded (the lazy-load race), eight concurrent requests over five different pictures, five interleaved conversations each still answering about its own picture, byte-identical prompts with different pixels, an image arriving on the second turn, and a picture swapped underneath an identical history. All six. ### What vision costs a text-only user: nothing Built from the commit before any of this and compared directly. - **Three greedy prompts, byte-identical output.** The vision path is invisible to a request without a picture. - **`--vision off` reproduces the pre-vision plan exactly** — same target, same pool, same expected peak; the only difference in `doctor --json` is two new reporting keys. That is the conditional-charge decision holding: no published memory number moved. ## Reuse A follow-up turn on a conversation whose pictures have not changed re-uses the state built for them: the tower does not run again, and prefill reads only the new text. Measured on the same conversation at `--memory-gb 10`: **15.4 s for the first turn, 1.8 s for the follow-up.** Re-measured at the 8.1 GB floor, which is where `verify.sh` runs the suite: **13.7 s then 1.9 s**, and all 18 serving assertions pass at that size too — the smaller pool changes the speed, not the behaviour. That reuse is only safe because the cache does not key on token ids. Every image expands to a run of the *same* placeholder id, so two different pictures that resize to the same grid produce byte-identical prompts; an id-keyed cache would answer the second from the first's state. Each run therefore carries a SHA-256 of the bytes it came from, and a match requires the digests to agree in both directions. The end-to-end check for it is two requests with identical words and different pictures: the first says dog, the second does not. ### Optimization — retained state and terminal-forward confirmation The first qualified mechanisms are deliberately scoped. They do not establish the combined engine's performance or complete the unified program. Runs and exclusions: [[sources/runs/2026/09/2026-09-05-optimization-initial-implementation]], [[sources/runs/2026/09/2026-09-05-optimization-second-implementation]]. | Mechanism and workload | Result | Limits | |---|---|---| | Omit the ordinary decode forward after the last requested token; fixed short prompt, one output token, 640 slots, requested 8.1 GB | Five complete valid confirmation pairs: median paired request-time reduction **14.73%**, median paired saving **0.2056 s**; all five improved; identical output token IDs; decode expert records **472 → 0** | One of the originally scheduled five pairs had swap activity and was wholly excluded. One fixed-policy replacement pair supplied the fifth valid pair. The effect is an avoided terminal forward; it is not a 14.73% steady decode or whole-engine claim. | | Compact retained GDN/PLE windows; frozen 440-token prose, one output token, 640 slots | Three complete valid pairs: median paired allocator-active reduction **333,414,400 bytes**; median sampled physical high-water **6,446,354,128 → 6,118,690,488 bytes** | 20 ms samples are lower bounds, not continuous peaks. The maximum chunk was explicitly forced to 1024, with a separate transient allowance. Latency was variable; the predeclared memory-benefit plus median-latency-nonregression gate passed. No general prefill speed claim. | The final-forward confirmation protocol required at least 5% median paired request reduction, at least four of five positive valid pairs, exact IDs and serving acceptance. The measured confirmation clears those thresholds; the 74-case serving battery passed the earlier control combination. Native generation diagnostics checked pending-token ownership, continuation, EOS, stop callbacks and early cancellation. Later cache-reservation and cancellation/validity fixes require their own final integrated verification before release. Exact compact-MTP-row checks passed 366 assertions; compact n-gram storage passed 1,160 checks including eviction and tiny capacities; incremental indexer blocks passed 1,127 checks at 2,051 tokens, including speculative rollback and continuation. The broader MTP text/image/prefix suite also passed. These are correctness results, not throughput measurements. Build-provenance limits for intermediate checks are recorded with the raw runs. The value-only sampler threshold passed 16 NumPy reference gates. Its first three-pair request benchmark lost two pairs to swap activity; the remaining pair does not demonstrate a request-speed improvement. Keep that candidate experimental pending component and confirmation measurements. All these experiments used the same local model and M5 Pro machine. OS filesystem cache was uncontrolled and never globally purged. No result certifies another Mac or justifies raising the context or query-by-key safety envelope. ### Final integrated optimization results **Final status, September 10: the approved OPT00–OPT36 optimization program is complete and the exact qualified candidate is active locally.** Every required execution gate is closed. All thirty-seven outcomes are accounted for; selected changes are implemented and tested, rejected candidates stay disabled, and explicitly prerequisite-dependent research remains deferred. This is completion of the approved program, not a claim that every possible future optimization has been implemented. **Session record and later live-server benchmark.** The completed session's exact final closure is [[sources/runs/2026/09/2026-09-10-optimization-final-program-complete]]. The original work, all-item outcomes, accepted and rejected experiments, failed runs, resource/thermal scheduling corrections and activation history remain preserved below. The later three-prompt benchmark, every clean timing, explicit discarded attempts and observed-versus-estimated tokens/sec are now documented separately in [[records/measurements/user-server-throughput-2026-09-10]]. It does not change the original qualification or the conservative calibration decision. The actual working source matches all 150 qualified source inputs. The installed executable SHA-256 is `9268e4b2a3371e78a71d493d7788559a06a22498e8061a89278c4918c6764673`; the installed Metal library also matches the qualified asset. Installed defaults match the final qualification, and a bounded generation check returns exactly `OK`. The previous installation remains available for rollback. Resource admission, cleanup, free model lock and unchanged Git index checks pass, with no owned model/compiler left running. Nothing is staged, committed, pushed or released by this activation: [[sources/runs/2026/09/2026-09-10-optimization-final-local-activation-and-all37-outcomes]]. Final reconciliation checks the complete candidate matrix, source and artifact identities, all thirty-seven historical dispositions, empirical calibration, public claims and generated documentation. Database validation has zero errors and one unchanged historical `LOG_UNKNOWN_KIND` warning at `log.md:124`; it is not reported as warning-free: [[sources/runs/2026/09/2026-09-10-optimization-final-preactivation-closure]]. The installed activation adds a final delivery check to those already-passed gates. **Measured improvements and tradeoffs.** These are median paired changes within each named study. They cannot be added together or generalized to every workload. | Workload | Observed result | | --- | --- | | Matching prefix with a distinct new tail | Visible preview latency 40.61% lower; request duration 31.52% lower | | Completely repeated committed prompt | Visible preview latency 96.48% lower; request duration 70.96% lower | | Plain sustained generation, fixed 10 GB target | Active TPS +0.46%, effectively flat; sampled process peak 6.54% lower | | Fixed-MTP sustained generation, fixed 12 GB target | Active TPS 3.66% slower; sampled process peak 3.94% lower | | Short prose / sampled requests | Request duration 2.39% / 8.06% lower; sampled process peak 7.47% / 6.15% lower | The two sustained cohorts each complete thirty-two first and thirty-two measured responses of 512 outputs, with exact paired prompt/output IDs and text. They retain thirteen clean joint pairs with MTP off and twelve with MTP on. Fixed-MTP request duration is 3.56% longer, within the original 5% nonregression allowance. New-input sustained preview is effectively unchanged (0.48% later plain; 0.86% later fixed MTP); the large preview gains require actual committed-prompt reuse. Process-peak reductions are neither a percentage of total machine RAM nor permission to allocate permanent extra expert slots. Automatic-prefill prerequisite studies separately reduced preview latency by 17.82% to 25.65% across their three qualified profiles. Those use a prerequisite build and different memory targets; the smaller two profiles use more sampled peak memory. They are not additional final-composition percentage gains. The completed calibration retains the conservative seventeen-parameter family with no added residency credit or timing multiplier. **Preserved preactivation and earlier execution history follows. Pending, live and unqualified wording below belongs to the stated checkpoint and is superseded by the completed status above. Prior failed attempts remain failed.** **Current status, September 10: every required final candidate workload now qualifies; empirical calibration is closed. Exact local activation is next.** Both complete sustained cohorts pass the original acceptance rules, alongside all eight short paired studies, both sixty-request lifetimes, the full original verification battery, public consumer checks and isolated installed upgrade/rollback qualification. All 37 OPT00–OPT36 dispositions remain in scope, including rejected and prerequisite-dependent alternatives. Qualification does not mean every optimization improved every metric or that deferred hardware/aggregate-demand research has been implemented. The final MTP-on cohort completes all thirty-two first responses and thirty-two measured responses, each with 512 outputs, for 32,768 outputs. Every paired prompt ID, output ID and text comparison is exact, including excluded rows. Twelve clean joint pairs remain; rounds 1, 2, 3 and 9 retain their original thermal/VM exclusions. No partial result is reused, replacement response added or acceptance threshold changed. Maximum sampled request process peak is 9,958,183,832 bytes under the fixed 12 GB target and 1,362 slots in both arms. Source/artifact identity, deadline and owned-process cleanup all pass: [[sources/runs/2026/09/2026-09-10-optimization-final-long-on-complete-pass]]. The actual bounded launch evidence is [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-thermal-recovery-guarded-launch]]. **The final sustained result is a speed/memory tradeoff.** Plain decoding is effectively flat: median paired active TPS improves 0.4560390212% with 6.5436095540% lower sampled process peak. Fixed-MTP active TPS is 3.6550887835% slower and request duration 3.5592440567% longer, with 3.9434511045% lower sampled process peak. This passes the original 5% request nonregression allowance; it is not an MTP speedup. No universal performance, energy, unbounded lifetime or new hardware claim follows. The completed empirical decision retains all seventeen parameters of `m5-pro-reference-envelope-v1`, with zero new permanent expert-capacity credit and no prefill/decode speed multiplier. It reconciles both sustained cohorts, both lifetimes, the actual context allocation ledger, short resource envelopes and the separately qualified 192-response prefill prerequisites: [[sources/runs/2026/09/2026-09-10-optimization-final-empirical-calibration-decision]]. Source/application activation and the final installed smoke and record closure remain to be performed. **Fixed-MTP sustained comparison on the exact final candidate.** Values below are marginal arm medians across the twelve clean joint pairs; changes are medians of per-pair percentages, so they need not equal the ratio of those arm medians. | Metric | Reference median | Combined median | Median paired change | | --- | ---: | ---: | ---: | | Active generation, tok/s | 7.551143377 | 7.303870732 | 3.6550887835% slower | | Late generation, tok/s | 7.429163901 | 7.189488654 | 3.6161598506% slower | | Client request duration, seconds | 74.719622021 | 77.127360459 | 3.5592440567% longer | | Visible preview, seconds | 7.077233437 | 7.101940979 | 0.8592601150% later | | Sampled request process peak, bytes | 9,892,025,276 | 9,504,289,592 | 3.9434511045% lower | Two of twelve pairs improve active/late generation and client duration; five improve preview; all twelve reduce sampled process peak. Active TPS uses the 511 observed inter-output intervals. Late TPS uses the 384 intervals from output 128 through output 512. Neither includes prefill or work after the final output sample. The full run takes 11,793.320247542 seconds. The inner driver exits 1 because excluded rows remain present, while the original outer complete-cohort assessor qualifies the run. This is a bounded fixed-total-target comparison; it does not establish infinite equilibrium. Nominal thermal observations do not certify an unloaded host. The earlier failed cohorts remain failed and are preserved below as history. **Preserved earlier execution history follows. Any pending, live or unqualified status below describes its original checkpoint and is superseded by the current status above.** The exact final integration build has completed all eight original short-request paired studies, both repeated-request lifetimes, the full verification battery and installed upgrade/rollback qualification. **MTP-off sustained generation also passes. MTP-on sustained generation remains unqualified after a thermal stop; final empirical calibration and local activation remain pending.** These results qualify specific workloads, not the whole optimization program or an absolute fastest runtime. **Current execution, September 10: continue the complete MTP-on qualification with prospectively tested recovery from excluded completed fair responses.** The user instructed continuation based on actual capacity; the prior quiet-window hold is superseded. V667 changes the benchmark control flow only. It explicitly adds `recover_completed_fair_thermal_requests=true`: a complete nominal-to-fair response remains excluded from timing, then may continue only through the unchanged bounded nominal-readiness check. Full work, exact parity, original twelve-GB cap, thirty-two cells, sixty-four 512-output responses, eight-clean-joint-pair minimum and nonregression criteria remain. Invalid or missing observations, low power mode, serious/critical thermal state, footprint violations, swap-outs and bounded readiness failures still stop. All eighty-six model-free checks pass, including the full simulated driver, both recovery points, cleanup and the original joint assessor. A full cohort whose first responses are all thermally excluded still fails qualification. The initial test representation error and corrected exact expectations are retained. V668 performs actual bounded prelaunch readiness before the single frozen run; V669 and V670 preserve the original terminal reporting and calibration criteria with their new evidence locations rebound. No passed cohort is repeated or partial observation reused. No inference source or user installation has changed. Full MTP-on qualification, final empirical calibration, final all-item/source/artifact/claim/projection closure and exact local activation remain open: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-thermal-recovery-preparation]]. **Closed prior V653 attempt:** eight first responses and seven measured responses complete all 512 outputs and 511 intervals. The four available first-response pairs and three complete measured pairs retain exact prompt IDs, output IDs and text. The final reference first response completes in 111.1651487500003 seconds and reports nominal-to-fair thermal state; its measured response is never started. Maximum sampled request peak is 9,960,494,024 bytes, minimum sampled reclaimable memory is 14,449,229,824 bytes and no new swap-outs occur. Source/artifact proofs, reservation and cleanup pass with no remaining owned jobs. This is incomplete and unqualified, with no partial performance comparison: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-user-directed-thermal-stop]]. The prior driver treated this completed fair response as an immediate whole-study abort even though its existing idle readiness helper could already wait through fair. V667 corrects that control flow prospectively without changing the inference candidate or timing eligibility. Nominal observations still do not certify an unloaded machine, and host-load limits remain reportable. **Preserved execution history follows; its former live statements are superseded by the current status above.** **Current execution, September 10: the complete MTP-on repetition is running under the user's instruction to continue now.** The user-directed continuation supersedes the prior quiet-window availability hold. Actual V656 startup readiness observes 121.419873333 sampled nominal seconds across sixty observations, normal pressure and no competing model/compiler, then starts frozen V653 under deadline 2026-09-10T18:41:33.844074+00:00. The unchanged full workload, cooling schedule, 18 GB startup check, 12 GB sampled request cap, live resource/ownership guards, timing exclusions and original acceptance remain in force. No applications are closed and no prior partial results are reused: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-user-directed-guarded-launch]]. Only the closed readiness and launch evidence is captured at this point. The full cohort is live and has not qualified. Host load and power observations remain part of the existing serving evidence; nominal thermal state does not certify an unloaded machine. Final reporting must retain those environmental limits. V654 is the prepared terminal reporter and V655 the prepared empirical calibration closure for this exact cohort. Both remain unexecuted. Full sustained MTP-on qualification, final all-item/evidence closure and actual local source/binary activation remain required. **Earlier preparation before user-directed continuation:** V653 is a separate, unlaunched repetition of the complete MTP-on protocol after V637's thermal stop. Its protocol, driver, helper, inference candidate, cooling schedule, workload and acceptance rules remain byte-identical. The runner's sole change is its output directory; reversing that change reconstructs V637 exactly. All seventy model-free checks pass for the fresh identity, and the original native/eight-paired prerequisite evidence is revalidated during freezing. The failed partial cohort stays separate and contributes no replacement observations: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-quiet-interval-preparation]]. V654/V655 preserve the original terminal-report and calibration criteria with only their new executor/report locations rebound. V656 preserves bounded prelaunch readiness and checks the immediately preceding failed cohort's cleanup. All three reverse-source and syntax checks pass; none of these scripts has executed. A user-confirmed quiet interval and actual bounded memory/thermal readiness are required before a fresh 12,600-second work plus sixty-second cleanup reservation. No deadline is granted and no model is launched. Full MTP-on qualification, empirical calibration, final all-item/evidence closure and actual local activation remain open; the other passed gates and all thirty-seven dispositions stay intact. **Preserved prior execution, September 10: the distributed-cooling MTP-on cohort stopped and remains unqualified.** V637 terminates after 13 recorded cells: twelve measured responses and thirteen first responses, each with all 512 outputs. V649 independently verifies all 12,800 completed outputs, full prompt work, all 511 emission intervals per response, sampled request caps and exact prompt/output IDs and text across all six complete A/B pairs. Maximum sampled request peak is 9,910,768,536 bytes against the original 12 GB cap. Minimum sampled reclaimable memory is 15,214,362,624 bytes, maximum owned RSS is 4,945,199,104 bytes, and no new swap-outs occur. Exact source/artifact proofs, the reservation and owned-process cleanup pass, with no remaining owned jobs: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-distributed-thermal-stop]]. The stop occurs during round 7 reference's first response, which finishes its complete workload in 83.77691487499942 client seconds. The generator reports nominal before and fair after, with low power mode disabled and no generator swap-ins or swap-outs. The measured request in that cell is never started. The outer incomplete-cohort rejection preserves this actual resource failure; it does not qualify a partial comparison. Distributed cooling delivered twelve between-request pauses in the complete cells, but cannot guarantee a nominal thermal state throughout every continuous request. The reference-arm failure does not establish a candidate regression. The next required condition is a quiet external machine interval before any fresh complete comparison. No automatic retry, replacement observations, partial-data pooling, shortened workload or relaxed acceptance is authorized by this failed result. V642/V643 remain unused preparations bound to the failed cohort and cannot close empirical calibration or authorize activation. All previously passed exact-build gates and all thirty-seven accepted/rejected/conditional/deferred dispositions remain intact. Full MTP-on sustained qualification, final empirical calibration, final disposition/artifact/claim/projection closure and actual local source/binary activation remain open. No model is left running and no source activation, installation, stage, commit, push or release has occurred. **Earlier V637 preparation and launch, September 10:** V637 retains the same minimum 180 seconds idle per cell: sixty reserved seconds plus thirty sampled nominal seconds before the first request, then ninety sampled nominal seconds before the measured request. Both 512-output requests remain continuous. All original sixteen pairs, work and numerical checks, timing exclusions, hard resource/thermal stops, 12,600-second work allowance and sixty-second cleanup remain. This changes benchmark scheduling only; inference code and the candidate binary are unchanged. All seventy model-free checks pass. The between-request pause keeps the already ready child alive with its native model reservation, checks child and lock status before and after observations, preserves the live reclaimable floor and outer memory guard, and stops before another cell if readiness fails. Incomplete raw assessment now reports the original stop before reading nonexistent later responses. Explicit reverse transformations reconstruct the prior runner, serving driver and helper exactly. The failed V627 evidence stays failed and supplies no replacement rows. V638 observes 121.28519829199999 sampled nominal seconds across sixty actual observations with normal pressure and no competing model/compiler, then starts the frozen full cohort under deadline 2026-09-10T16:58:48.859112+00:00. That cohort subsequently stopped and did not qualify, as recorded above. Source and frozen preparation: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-distributed-cooling-preparation]]. **Completed delivery evidence reconciled:** V641 rechecks 194 required artifacts against their exact earlier captures, along with current bound drivers, all 150 external-consumer source inputs, candidate binary, Metal library and installer asset. The full build, metadata, static, consumer, resource, four real client, 25/0 verification and 31/0 installed upgrade/rollback results remain valid. The packaged candidate uses the exact activation key exercised by the isolated installer. V642/V643 preserve the original terminal-report and calibration criteria, with only their current MTP-on executor/report paths rebound; both remain unexecuted. These read-only checks do not close MTP-on qualification or activate the user installation. Exact evidence and the two corrected audit path-resolution attempts: [[sources/runs/2026/09/2026-09-10-optimization-final-distributed-launch-and-delivery-reconciliation]]. **Previous failed cohort, preserved:** V627 finishes two first and two measured responses, all with 512 outputs. The completed pair retains exact prompt IDs, output IDs and text in both phases. Maximum sampled request peak is 9,899,021,496 bytes, below the original 12 GB cap. The combined measured request changes from nominal to fair and records eight global swap-ins; the reference measured request records sixteen swap-ins. No new swap-outs occur. Source proofs, deadline and cleanup pass, with no remaining owned jobs. The original thermal guard stops the cohort; the outer missing next warmup is secondary. No partial performance comparison or replacement evidence is accepted: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-cooled-thermal-stop]]. The required 60-second cooldown and at least 120 sampled nominal seconds were delivered before each cell. The first and measured 512-output requests nevertheless run back to back, totaling 187.00727912500008 client seconds for the combined cell. Before-cell cooling was insufficient on this attempt. This does not isolate a code regression or prove a remedy. Next qualify bounded cooling between the two complete requests prospectively, preserving every output, original inference behavior, paired criterion and hard resource rejection. At that checkpoint no successor cohort had been frozen or launched. V630/V631 remain unexecuted report preparations attached to a failed cohort and cannot close calibration or authorize activation. Passed gates remain passed; final MTP-on sustained qualification, calibration, complete disposition reconciliation and actual local activation remain open. Binary SHA-256: `9268e4b2a3371e78a71d493d7788559a06a22498e8061a89278c4918c6764673`. The machine is the recorded M5 Pro Mac. For the eight short-request studies below, reference and combined arms use the same exact build, checkpoint and fixed expert pool within each study. This is a comparison of reference paths against the selected integration family, not an old-release comparison. Each study retains 32 measured and 32 first responses, interleaved A/B order, original resource eligibility and full output/work checks. No failed earlier run supplies replacement observations. Positive percentages below mean a reduction. Each percentage is the median of eligible paired reductions, not a ratio of marginal medians. Process-memory percentages describe sampled request peaks, not whole-machine RAM, permanent capacity or a continuous maximum. First-response results are assessed separately. | Workload | Target / fixed slots | Clean measured pairs | Request time reduction | Visible-preview reduction | Sampled process-peak reduction | |---|---:|---:|---:|---:|---:| | Unique prose, 440 prompt tokens, 16 outputs | 8.1 GB / 827 | 14 | 2.39% | 0.38% | 7.47% | | Sampled short prompt, 17 prompt tokens, 16 outputs | 8.1 GB / 827 | 15 | 8.06% | 0.24% | 6.15% | | Fixed MTP, 17 prompt tokens, 16 outputs | 10 GB / 838 | 16 | −2.00% | −0.35% | 4.82% | | Distinct tail, 445 prompt tokens, 256 reused by combined, 16 outputs | 8.1 GB / 640 | 16 | 31.52% | 40.61% | 1.59% | | Complete repeated prompt, 440 reused by combined, 16 outputs | 8.1 GB / 640 | 16 | 70.96% | 96.48% | 5.73% | | Unique prose with retention enabled, zero reuse, 16 outputs | 8.1 GB / 640 | 16 | 0.84% | −0.21% | 5.00% | | Actual default controls, 440 prompt tokens, one output | 8.1 GB / 640 | 16 | 2.98% | Not measurable: whitespace output | 7.09% | The eighth original short-one study, at 8.1 GB and 827 fixed slots, has 16 clean measured pairs and **11.28% lower paired request duration**. It measures terminal-work and compact-state effects in the selected family; one output cannot establish active decode TPS. The fixed-MTP short cohort passes its original memory/nonregression criteria, but its roughly 2% longer request duration is explicitly not a speed gain. Repeated-prompt preview benefits depend on actual prefix reuse and do not apply to unrelated prompts. Individual fixed-MTP, distinct-tail, complete-repeat and short-one evidence is preserved in [[sources/runs/2026/09/2026-09-09-optimization-final-mtp-resource-complete-pass]], [[sources/runs/2026/09/2026-09-09-optimization-final-distinct-tail-complete-pass]], [[sources/runs/2026/09/2026-09-09-optimization-final-complete-repeat-complete-pass]] and [[sources/runs/2026/09/2026-09-09-optimization-final-short-one-complete-pass]]. Full precision, first-response results, exclusions, exact work and raw measurements: [[sources/runs/2026/09/2026-09-09-optimization-final-prose-complete-pass]], [[sources/runs/2026/09/2026-09-09-optimization-final-sampled-complete-pass]], [[sources/runs/2026/09/2026-09-09-optimization-final-retention-complete-pass]] and [[sources/runs/2026/09/2026-09-09-optimization-final-actual-default-and-eight-paired-complete-pass]]. The actual command/target audit corrects two earlier prose/sampled plan labels that incorrectly said 10 GB. Their protocols, acceptance and percentage calculations were already correct: [[sources/runs/2026/09/2026-09-09-optimization-final-calibration-inputs-and-profile-label-correction]]. All 512 first/measured responses across the eight studies satisfy their actual sampled physical cap, including observations excluded from timing. The automatic-prefill mechanism also has three complete prerequisite studies on build `d3701afdb0540850f376ca9a696a2a0ffa31a67121362e341f87ebc9e27fd7dc`, before final default composition. Each has 32 first and 32 measured responses with 16 outputs, matching fixed pools within each profile, actual public planner metadata, and unchanged chronological compute geometry. All 192 responses and the original descriptive calculations have been independently rechecked. These results explain the prefill benefit; they are not an additional final-build benchmark and cannot be added to the combined percentages above. | Planner chunk / target / fixed slots | Prompt tokens | Clean measured pairs | Request reduction | Preview reduction | Prefill-time reduction | Sampled peak reduction | |---|---:|---:|---:|---:|---:|---:| | 256 / 10 GB / 961 | 1,027 | 14 | 21.89% | 25.65% | 25.66% | -16.03% | | 512 / 12 GB / 1,491 | 2,051 | 11 | 18.65% | 20.91% | 20.91% | -12.69% | | 1024 / 16 GB / 2,576 | 4,099 | 6 | 16.70% | 17.82% | 17.95% | 4.16% | Negative peak reductions mean more temporary process memory. All responses remain within their declared 10/12/16 GB caps. Whole-request, prefill and preview gains are specific to these eligible group sizes and inputs; sixteen-output timing does not establish sustained TPS. Original evidence: [[sources/runs/2026/09/2026-09-09-optimization-public-planner-256-complete-pass]], [[sources/runs/2026/09/2026-09-09-optimization-public-planner-512-complete-pass]], [[sources/runs/2026/09/2026-09-09-optimization-public-planner-1024-complete-pass]]. Exact reconciliation and unexecuted final-calibration preparation: [[sources/runs/2026/09/2026-09-10-optimization-final-prefill-reconciliation-and-calibration-preparation]]. The OS filesystem cache was uncontrolled and was not globally purged. Each process ran its original first request before its measured request. A fresh process is not evidence of a physically cold SSD; prefix reuse is reported separately above. Observed paired request ranges make the variability visible. Positive reductions mean faster; negative reductions mean slower. These are the extrema of these finite studies, not reliable population p95 estimates. | Workload | Observed paired reduction range | Faster pairs | |---|---:|---:| | Short one-output | 4.72% to 16.36% | 16/16 | | Unique prose | -3.59% to 9.90% | 13/14 | | Sampled short | 2.70% to 16.00% | 15/15 | | Fixed MTP | -5.27% to 3.13% | 1/16 | | Distinct prefix tail | 30.12% to 33.24% | 16/16 | | Complete repeated prompt | 66.23% to 72.89% | 16/16 | | Unique with retention | -4.36% to 8.42% | 10/16 | | Actual default one-output | -1.63% to 6.88% | 14/16 | The original three-pair serving A/A pilot, with the same arm configuration and unchanged client/generator swap counters, showed an apparent median request reduction of 0.86% and a maximum of 8.93%. It used an earlier build and an eight-output workload, so it illustrates variability rather than estimating the final build's noise distribution. No historical A/A percentage is subtracted from the current results, and no new threshold or significance claim is introduced. Near-zero timing changes remain descriptive; passing resource-benefit/nonregression criteria does not make them universal speed gains. Exact ratios, counts and unchanged exclusions: [[sources/runs/2026/09/2026-09-09-optimization-final-observed-spreads-and-original-aa-context]]. The complete MTP-off long-generation study uses a fixed total memory target of 10 GB, with 1,217 effective expert slots in both arms. All 32 first and 32 measured responses complete 512 output tokens, for 32,768 outputs in total. Every paired prompt, output-token sequence and output text matches, including excluded observations. Thirteen pairs qualify after rounds 4, 12 and 15 are excluded for global swap-in activity; there are no replacement rounds. The original cohort assessor and outer source/deadline/cleanup qualification pass. | MTP-off long-generation metric | Reference median | Combined median | Median paired improvement | |---|---:|---:|---:| | Active emission speed | 6.9522 tokens/s | 7.0013 tokens/s | 0.46% | | Emission speed after output 128 | 6.8849 tokens/s | 6.9567 tokens/s | 0.34% | | Whole request duration | 81.8221 s | 81.0804 s | 0.85% | | Visible-preview latency | 8.0920 s | 8.0241 s | -0.48% | | Sampled process peak | 7.9069 GB | 7.3922 GB | 6.54% | Higher throughput and lower duration/peak are improvements. The displayed arm medians provide context; the improvement column is calculated per pair, so it must not be reconstructed from those marginal medians. Active emission is 511 divided by the sum of all 511 inter-token intervals. The later window uses 384 intervals from output 128 through output 512. Both exclude prefill and work after the final sampled token. Nine of 13 pairs improve active and later throughput; all 13 reduce sampled process peak. The small timing shifts do not establish a material sustained-speed gain. This is bounded 512-output evidence, not indefinite equilibrium, another memory target or another chip. Across all first/measured responses, including timing exclusions, the maximum sampled peak is 7,938,182,984 bytes. Evidence and full precision: [[sources/runs/2026/09/2026-09-10-optimization-final-long-off-complete-pass]]. The first final MTP-on attempt refused its incompatible legacy large-memory profile before launching a model. V609 corrected delivery and completed both first responses with identical 512-output work, but eight global swap-in pages during combined startup triggered the legacy full-study abort. There were no swap-outs, at least 22.44 GB remained reclaimable, and the largest first-response sampled peak was below 9.88 GB. Both partial identities remain failed and supply no MTP-on performance comparison: [[sources/runs/2026/09/2026-09-10-optimization-final-long-on-delivery-refusal]], [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-long-startup-swapin-stop]]. The prospective V616 method preserves the exact original 12 GB, 512-output chat protocol and 18 GB startup requirement. Startup swap-ins invalidate the entire pair under the original joint assessor while the fixed cohort continues. Swap-outs, missing/invalid/reset counters, pressure and footprint violations retain stops. Full response/parity checks also apply to excluded pairs; no failed partial observations or replacement rounds enter the comparison. All 55 model-free checks pass. That V616 cohort later stops on a fair thermal observation and remains failed: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-long-startup-exclusion-preparation]]. V616 completes seven first and seven measured responses, preserving all 512 outputs in each. Round four combined begins in nominal thermal state and ends fair, triggering the original thermal stop. The missing later warmup reported by the outer assessor is a consequence, not the root cause. All three completed pairs preserve exact first/measured prompt and output IDs and text; all 14 responses stay below 12 GB, with a maximum sampled peak of 9,912,832,968 bytes. At least 20,621,852,672 bytes remain reclaimable, with no new swap-outs or swap changes in the failed request. This partial cohort remains failed and supplies no sustained comparison: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-long-thermal-stop]]. The subsequently failed V627 cohort increased cooling to a 60-second reserved cooldown plus 120 sampled nominal seconds per cell, within the unchanged 600-second readiness bound. All original requests, numerical gates, timing exclusions, memory/pressure/swap-out/thermal stops and full work allowance remain. All 60 model-free checks pass. The prelaunch review also corrects an overstrict sample-count assertion using an actual saved 120-second observation; V624's unused freeze is preserved. The real prelaunch readiness check passes before the fresh complete cohort starts; the cohort later stopped thermally and remains unqualified, as recorded above. Longer cooling is not a guarantee of constant clocks or equilibrium: [[sources/runs/2026/09/2026-09-10-optimization-final-mtp-cooled-cohort-preparation]]. Both original lifetime modes complete 60 requests: two warmup cycles plus eight measured cycles over six request positions. MTP off at 10 GB reaches a maximum sampled request peak of **9,037,173,152 bytes**. One prose observation has eight swap-ins and no swap-outs and is excluded from clean growth calculations without replacement; all requests still meet output/work and absolute-cap checks. MTP on at 12 GB reaches **10,828,060,144 bytes**, with all 60 observations resource-clean. Both modes pass exact per-position replay, the original 64 MiB active-growth and 256 MiB physical-end-growth bounds, bounded embedding storage and charged prefix capacity. This is bounded lifetime evidence, not a leak-free-forever claim: [[sources/runs/2026/09/2026-09-09-optimization-final-lifetime-off-pass]], [[sources/runs/2026/09/2026-09-09-optimization-final-lifetime-on-pass]]. The same exact build passes all 25 original top-level verification checks, including independent numerical parity, all 15 behavioral-quality items, all 74 API probes, full vision-serving inputs and resource checks. The public installer also passes all 31 installed checks and all five install/upgrade/E2E/rollback phases in an isolated root, including actual rollback generation. The user's installed binary remains unchanged at this checkpoint: [[sources/runs/2026/09/2026-09-09-optimization-final-complete-verification-pass]], [[sources/runs/2026/09/2026-09-09-optimization-final-installed-qualification-pass]]. The final planner constants retain the verified reference family. Exact-source review proves policy continuity after moving unchanged device-observation functions, and the actual final 2,048-token context rung fits its 8,997,561,600-byte ledger with a sampled peak of 8,536,952,072 bytes. No temporary vision saving is allocated permanently to expert residency, and no isolated prefill ratio multiplies the whole timing curve. The MTP-on sustained measurement and final empirical family decision remain open. Hardware claims stay on the measured machine; the fused rotation default is restricted to its qualified chip/model/OS and other platforms retain fallback dispatch: [[sources/runs/2026/09/2026-09-09-optimization-final-calibration-inputs-and-profile-label-correction]], [[sources/runs/2026/09/2026-09-09-optimization-final-control-and-fallback-review]]. The final terminal extractor and empirical-calibration report now follow the exact V627 cohort. Their original calculations remain, and the calibration report also checks the closed automatic-prefill prerequisites. Both reports remain unexecuted until the full MTP-on cohort qualifies; activation remains unapplied: [[sources/runs/2026/09/2026-09-10-optimization-final-cooled-launch-and-report-binding]]. ### Three-prompt serving benchmark and throughput expectations **What to expect from this observed setup.** On the recorded M5 Pro Mac, the existing auto-sized server produced **10.28 to 15.80 decode tokens per second** across the eight clean short-request observations. A rough expectation of **10 to 16 tokens/sec for similar short requests at these settings** is a description of this observed range, not a guaranteed floor, ceiling, sustained rate or estimate for another machine. Thinking was disabled, greedy sampling was used, and the MTP draft head was enabled. Model, output length, available memory, retained prompt state and background activity matter. The server's existing planner reported `est_warm_tok_s` of **10.63** at about **121 cached experts per layer**, and **11.47** after automatic residency grew to about **146 per layer**. These are planner estimates, distinct from the observed request rates. Its metadata retained a **27.9 GB advertised auto target** while the pool changed from **5,817 to 7,006 slots**. The target is not an observed peak-memory measurement. This small run does not recalibrate the planner or supply a new MTP multiplier. The original integrated sustained studies used **10 GB with MTP off** and **12 GB with fixed MTP on**, producing combined-arm median active rates of about **7.00** and **7.30 tokens/sec**, respectively. Those are different memory budgets, workloads and timing definitions from this larger auto-sized server's short requests. They remain in [[records/measurements/optimization-final-composition-2026-09-09]]; the present results neither replace them nor establish another optimization speedup. **Method and evidence.** Three synthetic prompts cover a sky explanation, a Python function and a longer fictional incident report. Each is sent as an initial request, an exact repeat, and a follow-up including the complete initial user/assistant exchange. Requests run one at a time against the user's already running server, with thinking disabled, temperature zero, seed seven and a maximum of 128 output tokens. Initial means first in this small study, not a cold process, SSD or expert pool. All nine completed responses stop naturally, and all three completed exact repeats match their initial output text. Exact prompts, raw streamed frames with arrival timestamps, outputs, server metadata, memory observations and driver bytes are preserved in [[sources/runs/2026/09/2026-09-10-user-server-three-prompts-small-benchmark]]. | Prompt | Case | Output tokens | First visible text | Complete response | Decode tokens/sec | | --- | --- | ---: | ---: | ---: | ---: | | Sky explanation | Initial | 76 | 3.94 s | 9.82 s | 12.92 | | Sky explanation | Exact repeat | 76 | 2.5 ms | 4.81 s | 15.80 | | Sky explanation | Follow-up | 47 | 1.51 s | 5.49 s | 11.81 | | Python function | Initial | 108 | 2.10 s | 10.52 s | 12.83 | | Python function | Exact repeat | 108 | 2.7 ms | 8.21 s | 13.16 | | Python function | Follow-up | 32 | 1.86 s | 4.98 s | 10.28 | | Long report summary | Initial | 103 | 5.56 s | 14.98 s | 10.93 | | Long report summary | Exact repeat (resumed) | 103 | Excluded: paging | Excluded | Excluded | | Long report summary | Follow-up (resumed) | 55 | 1.47 s | 5.51 s | 13.60 | **Interruptions and exclusions.** One additional long-summary repeat attempt is canceled on observed warning memory pressure and has no successful terminal frame: [[sources/runs/2026/09/2026-09-10-user-server-small-benchmark-pressure-attempt-discarded]]. An optional quiet-memory wait sends no requests and expires. The remaining cases then continue under the original live request guards. The completed resumed summary repeat overlaps **11,581 system swap-ins** and is excluded from the timing comparison: [[sources/runs/2026/09/2026-09-10-user-server-small-benchmark-paging-repeat-discarded]]. Its successful content check is not timing qualification. The eight reported timing rows have nominal sampled thermal state and no swap-counter change during their request. The aggregate source is mixed evidence; the two separate sources explicitly mark the discarded attempts/timings and the original source remains unchanged. Automatic residency changes across the interruption, so the resumed cases are not a matched fixed-pool comparison. Only one observation of each case is taken; no population percentile, significance test, universal rate or causal attribution to a single optimization is claimed. The two short immediate repeats show almost zero prefill time and first visible text in milliseconds; their remaining generation still takes seconds. The ordinary Ollama endpoint does not expose an exact cached-token count, so no unobserved hit count is claimed. The user's server is left running, with no other model/compiler launched and no applications closed. **Meaning of the timing fields.** First visible text is client time from sending the request to receiving a frame with non-whitespace answer text. Complete response is client time through the successful terminal frame. Decode throughput is the server's `eval_count / (eval_duration / 1e9)`, excluding prefill. It is not the independent inter-output-interval active TPS used by the sustained qualification studies. Follow-ups have different input and output lengths from their base prompt, so total durations are not like-for-like speed comparisons. **How to see your own numbers.** `slotstream run` prints separate prefill and decode rates after generation; its `--stats-json ` option saves the raw measurements. When using the existing server, `/api/chat` and `/api/generate` provide `eval_count` and nanosecond `eval_duration` in their successful final response; streaming clients receive them in the final frame. Calculate the ratio above when the duration is positive. The ordinary OpenAI-compatible response carries token usage without these Ollama duration fields. These are end-of-response observations, not a built-in continuously updating counter. See the CLI guide for commands. **Long-session context.** The original first-principles optimization program, all OPT00–OPT36 selected/rejected/deferred outcomes, superseded progress, failed attempts, resource and thermal-control changes, qualification evidence and actual local activation remain in [[records/plan/whole-engine-optimization-2026-09-04]]. Integrated preview, prefill, sustained throughput, memory and lifetime results remain in [[records/measurements/optimization-final-composition-2026-09-09]]. The final exact-source and installed-artifact closure is [[sources/runs/2026/09/2026-09-10-optimization-final-program-complete]]. This later user-server benchmark is observational follow-up evidence, not a replacement for that qualification. ### Hermes configuration correction and regression checks The earlier guide placed the main output limit in `model.max_tokens`, but Hermes CLI initialization did not forward that field to the agent. The earlier live integration gate passed a limit directly to `AIAgent`, so its successful requests did not establish that the copied guide applied the same limit. This was a documentation and test-coverage error on our side. It did not establish a Slotstream decoder failure. The revised guide defines a provider named `slotstream`, with its endpoint, placeholder credential, chat-completions mode and `extra_body.max_tokens` together. Main requests carry 4,096 output tokens; compression requests independently carry 4,096, and title requests carry 64. Both auxiliary tasks explicitly receive a 1,800-second timeout. The main stream watchdog remains separately configured. Reasoning stays an optional user preference. #### Configuration regression checks Unmodified Hermes commits `b1f003e18633298d549668b8e186af84cca45b76` and `4a39a3ff8bea45ab5a6b646ce26ced88a8fed079` each pass eleven cases through the real CLI configuration and request code, with synthetic HTTP and all real socket connections disabled. The guide's YAML is the input, not a separately maintained test configuration. The cases cover the clean guide, legacy and current providers named `custom` plus conflicting environment settings, optional reasoning, an edited main output cap, missing and disabled providers, local HTTP 503 and 401 responses, truncated and empty summaries, and the explicit default-profile diagnostic command with a conflicting sticky profile. Main requests preserve the configured endpoint and 65,536-token context. Failure cases make no cloud inference request, failed summaries preserve history, and titles successfully retry after unsupported constrained output is rejected. These are configuration and transport-construction tests, not real-model results. Raw configuration results: [[sources/runs/2026/09/2026-09-08-hermes-config-latest]] and [[sources/runs/2026/09/2026-09-08-hermes-config-reported-version]]. #### Real-model follow-through Released Slotstream 0.2.11, binary SHA-256 `7f540b73b5ff4cf48975ff122a3d17f57e53103e616ad84a76cfc71d551be5b8`, served unmodified latest Hermes using the revised guide through CLI initialization. The bounded server used `--max-context 65536 --mtp off --memory-gb 10`, with 30.4 GB reclaimable observed before launch and only one model process. This functional run is not a throughput or memory-capacity benchmark. With stale generic-provider settings present, the real terminal read executed once, the follow-up recalled the code without another read, and the tool-free length probe completed the full list after 1,095 generated tokens with `finish_reason: stop`. The request carried `max_tokens: 4096`. Every captured inference request stayed on loopback. A metadata connection was blocked by the test's network guard; this is not a claim that every Hermes feature is offline. The real compressor reduced the 39-message, 99,684-character fixture to 25 messages and 86,585 characters. Its 4,096-token request completed with `stop` after 724 generated tokens. The diagnostic code appeared only in the generated summary at index 4. An actual subsequent turn read 18,659 prompt tokens and recovered the exact code without executing another tool. This is forced compression and recall evidence, not an automatic-threshold stress test. The observed compression threshold was 55,705 tokens. The length request's cap, complete list and summary-only recall assertions were also checked against the preserved original captures after tightening those assertions in the gate. The multi-turn run used `127.0.0.1`; final CLI smoke runs preserve the guide's `localhost` spelling. Raw receipt: [[sources/runs/2026/09/2026-09-08-hermes-live-hardening]]. Both final CLI smoke runs returned exactly `OK`, exited zero and sent a 4,096-token main limit to `http://localhost:11434/v1/chat/completions`, with auxiliary requests independently capped and timed. The diagnostic command explicitly selects the default profile and classic CLI; the ordinary fresh guide also selects that frontend by default. This does not qualify every alternative Hermes frontend. Their wire captures are in [[sources/runs/2026/09/2026-09-08-hermes-cli-hardening]]. The unchanged released server passed 24 live OpenAI protocol checks, including streamed and ordinary tool-result loops, parallel call identity with reversed results, reasoning separation, invalid histories and unsupported format rejection. A deliberately truncated required tool produced an inference error; its streamed form published neither an executable partial call nor a successful completion terminator. Raw requests and responses are in [[sources/runs/2026/09/2026-09-08-hermes-protocol-regression]]. No new inference-engine behavior was needed for these configuration fixes. Both complete configuration suites passed again after final failure-assertion review, with the final gate and YAML hashes preserved. The owned test server was stopped and its port had no listener. Final verification: [[sources/runs/2026/09/2026-09-08-hermes-final-config-regression]]. #### Scope and remaining boundaries A named provider protects this setup from collisions with the generic `custom` path; editing that named provider or explicitly configuring a fallback can still change routing. A wrong active profile can still select different settings; the diagnostic command uses `--profile default` to select the configured root. Neither the screenshot nor these reproductions identify which setting was present on another person's laptop. The request override repairs the wire limit. It does not repair Hermes's separate internal reservation accounting. Keep compression enabled, allow room for input and follow the server's advertised maximum when adjusting limits. Tests of failed summary publication do not prove general memory quality or every long-conversation path. The server still rejects unsupported constrained JSON, and Hermes's successful plain-text title retry does not provide schema guarantees. The first long-output fixture left tools enabled and triggered calculator calls. The guard refused them and the run was interrupted; it is retained as discarded fixture evidence in [[sources/runs/2026/09/2026-09-08-hermes-long-output-fixture-discarded]]. The replacement disables tools only for its length probe. Earlier protocol, vision and memory-capacity qualifications remain separate; their direct-agent output-limit checks are not evidence for the old guide's CLI configuration path. #### Explicit proxy environments A final source and pure-function check on both Hermes versions confirmed that an explicit HTTP proxy is selected for the local endpoint unless its hostname is excluded. Setting both NO_PROXY and no_proxy to include localhost and 127.0.0.1 selects a direct connection for both spellings. The guide now documents preserving existing exclusions and checking the profile .env, which can override the shell. This is an additional environment condition; it does not explain a log whose resolved endpoint already names OpenRouter. No proxy server or inference request was used in this check. Raw result: [[sources/runs/2026/09/2026-09-08-hermes-proxy-exclusions]]. #### Automatic MTP guide correction The Hermes server command now leaves MTP on Slotstream’s normal automatic default. The public CLI already selected automatic MTP; the guide unnecessarily forced it off. The model-provider YAML and Hermes launch command are unchanged. The integration gate now captures the running server’s memory plan, checks the configured context, and accepts --expect-mtp on/off to require the selected state before any inference request. Planning-only inspection and synthetic metadata rejection checks passed; the evidence is [[sources/runs/2026/09/2026-09-09-hermes-automatic-mtp-preflight]]. No live Hermes inference was performed for this change because another benchmark held the model-process reservation. The earlier live Hermes runs remain explicitly MTP-off evidence. Automatic MTP integration, including the enabled path, is pending and must not be claimed from this planner output. ### Hermes integration: context qualification and OpenAI agent protocol The integration failure in [issue #11](https://github.com/carloslfu/slotstream/issues/11) has two independent causes: the released OpenAI endpoint rejects agent tool semantics before inference, and Hermes requires a context larger than the served default. The issue does not include the reporter's trace, versions, or configuration; these are independently reproduced failures, not a claim to have identified their exact first request. #### Protocol and client findings The OpenAI adapter now reuses the native tool schema, history, reasoning, image, and output parser. It validates call/result identity, restores assistant call order when results arrive out of order, and emits complete OpenAI calls with distinct IDs and stream indices. This reordering matters because the native template renders results positionally without call IDs. Required/named choices, incomplete calls, invalid history, and requested context inflation fail explicitly. Ordinary chat retains its prior template and sampler; tool turns use the existing agent sampler. Discovery now reports the actual runtime window and the available vision capability. An image-enabled server previously returned only the completion capability, causing unmodified Hermes to report that vision was unsupported despite successful image inference. Hermes's reasoning-off flags and bounded `options.num_ctx` are accepted and enforced. The unsupported constrained-output rejection now triggers Hermes's plain title retry. Constrained JSON generation and strict tool-schema enforcement remain unsupported; argument validation and execution belong to the caller. The local profile gives main responses and auxiliary summaries explicit output allowances (Hermes otherwise omits the auxiliary limit and inherits the small server default), keeps title and compression requests on the main local provider, and sets a bounded local stream watchdog. The qualified long prefill takes longer than Hermes's prior default watchdog, so transport keepalives alone do not justify leaving that client timeout unchanged. The guide uses speculative decoding off, matching the tested configuration. See `docs/HERMES.md` for the setup guide. #### Context qualification and its counterexamples | Run | Completed prompt + reply | Process peak | Expected peak | Result | |---|---:|---:|---:|---| | Released diagnostic, default-context plan | 65,520 + 1 | 9.35785704 GB RSS | 8.99949696 GB | Exceeds the estimate | | Requested-context plan, retained state charged | 65,520 + 1 | 10.056142544 GB sampled footprint | 9.260163328 GB | Exceeds estimate and 10 GB target | | State plus transient reservation | 65,520 + 1 | 9.735410056 GB sampled footprint | 9.742730496 GB | Fits estimate and 10 GB target | The first check enlarged its diagnostic engine limit after constructing the ordinary plan. The second correctly charged retained state but missed transient process memory; lifetime RSS alone missed that peak. The final plan charges both retained growth and a conservative reserve throughout the explicitly larger supported range before selecting the pool and prefill batch. The ordinary 32,768-token default and its allocation budget remain unchanged. The explicit ceiling is 65,536, with minimum-pool and availability checks. The successful run had no abort, inference error, pressure cancellation, or increase in system swap-outs. Its margin to the expected estimate was small; the requested total target still had headroom. This is one measured text-capacity and memory configuration, not a mathematical memory bound, a cross-hardware guarantee, long-context answer-quality evidence, or a throughput comparison. The diagnostic now also requires full prefill and a reply before reporting `fits`. Raw evidence: [[sources/runs/2026/09/2026-09-05-hermes-context-first-counterexample]], [[sources/runs/2026/09/2026-09-05-hermes-planned-context-counterexample]], and [[sources/runs/2026/09/2026-09-05-hermes-context-qualified]]. The failed runs remain intact. #### Final client acceptance Hermes release `v2026.8.31` at `29112bef099274229cadff79cdff7bf7b99c4b77` and main snapshot `9dd6634c5635321cf38840cc30e9b51226689128` both passed real terminal fixture execution, follow-up recall, and the title fallback on final binary `800502693480a7187f81fab231f8516dccc56a7926894e7c0e8c2ce2480ed3f8`. Both report package version 0.21.0; the commits distinguish them. The actual released CLI returned `OK` and exited zero. All captured main tool-enabled requests used the documented 4,096-token output allowance, `reasoning_effort: none`, `think: false`, and bounded `options.num_ctx: 65536`. Discovery selected 65,536 tokens and a 52,224-token compression trigger. The released agent also discovered vision and answered `Dog` for the fixture image. Its actual auxiliary compressor reduced the 39-message, 99,684-character fixture to 25 messages and 87,350 characters. The summary completed with `stop` after 999 output tokens, exceeding the ordinary 512-token server default and demonstrating why the separate auxiliary allowance matters. The diagnostic code existed only in the summarized middle; after compression it appeared only in the handoff at index 4. A subsequent actual agent turn read 18,367 prompt tokens and recovered the exact code. This is forced-compaction and recall evidence, not an automatic-threshold stress test or a general memory-quality score. The combined final configuration was `--memory-gb 11 --max-context 65536 --vision on --mtp off` with 30.710185984 GB reclaimable before launch. Its larger explicit target exercises context plus the vision tower, which the smaller text-only qualification did not combine. This functional test does not extend the separate full-context text memory measurement to a full-window multimodal bound. | Acceptance check | Exact candidate prefix | Result | |---|---|---| | Released Hermes agent, image, compaction/recall and actual CLI | `80050269` | Pass | | Current Hermes agent, terminal tool, follow-up and title | `80050269` | Pass | | OpenAI HTTP/SSE, parallel calls with reversed results, errors and image tools | `80050269` | 27 checks passed | | Existing image serving across all dialects and history reuse | `39d236e8` | 25 checks passed | | Existing ordinary API robustness suite | `666b59f8` | 74 checks passed | | Native gateway tool/result loop and actual Ollama text CLI | `39d236e8` | Pass | | Actual Ollama image CLI after capability correction | `80050269` | `Dog`, exit zero | | Full T0/T1 catalogue | `39d236e8` | 42 groups, 21,945 assertions passed | | Final T0 catalogue and static components | `80050269` | 31 groups, 20,087 assertions; static components pass | The final candidate differs from `39d236e8` only in Server.swift's vision capability advertisement. The prior candidate adds tool-result ordering and its assertions to the memory-reserved `666b59f8` source. Exact source identities bridge these results without pretending all gates ran on the same binary. The ordinary context default, plain-chat template/sampler, native gateway, image paths, and installer/planner behavior retain their passing regression coverage. Static validation reports zero errors and the pre-existing unknown log-kind warning; it does not mask that warning. The earlier real-client checks remain preserved, including the first compression fixture that grew instead of shrinking: [[sources/runs/2026/09/2026-09-05-hermes-openai-real-client-gates]]. Regression results and the vision-discovery counterexample are in [[sources/runs/2026/09/2026-09-05-hermes-adapter-regressions-and-vision-discovery]]. Final acceptance is in [[sources/runs/2026/09/2026-09-05-hermes-final-integration-acceptance]]. At qualification time the installed release was unchanged. The user subsequently requested installation; the normal local `slotstream` command now points to the exact verified Hermes binary and matching Metal library, with the previous install preserved for rollback. The installed runtime and context-planning checks passed; see [[sources/runs/2026/09/2026-09-05-hermes-local-install-verified]]. No issue reply or release has been published. A verified local binary, source archive, identity and complete client captures are retained under `.build/hermes-integration/`. This qualification applies to that frozen source closure; separate ongoing engine optimizations in the working tree are not silently included in these results. ### Configurable context component and interface contracts This unreleased candidate preserves the installed ordinary allocation projection and exposes one context/wait policy through CLI and library entry points. The request clock starts on acceptance, includes preparation and queueing, and stops at the first sampled token. Zero wait disables only time. Atomic Engine-owned leases account for simultaneous input, image and dispatch reservations; per-buffer guards price full replacement and provisional draft allocations before the work starts. Build 16 passes all ten bounded C07 shape/prefix cases on one identified binary, each with 1605 unchanged assertions, plus the 835-assertion HTTP matrix. Nine numerical cases have zero candidate drift; the short-tail64 case has one logits measure inside the unchanged band, with every greedy token agreeing. Normal Generator activation is implemented, but these short witnesses do not reach its real128K transition or qualify full-window capacity. Build 17 corrects the preserved build16 MTP exclusivity abort and passes actual MTP retention/cancellation/recovery (198), lifecycle (138), CLI (102) and planner (64). The adaptive-MTP mismatch exactly reproduces the already rejected V50 option, which remains disabled. Build 18 then passes fixed-MTP pressure cancellation/feasible recovery (80), HTTP (835) and T0 (33 groups / 22249). Four distinct conversations and their interleaved follow-ups complete, then a 4096+16-token main request completes; eight system swap-in pages exclude that run from capacity evidence. Actual Ollama and the published AI SDK gateway pass at32768. Hermes enforces its own64000-token minimum before inference at that ordinary window. At65536, actual Ollama, gateway SDK and pinned Hermes all pass; Hermes performs one real allowlisted tool read, follow-up, title and bounded request-context handling. Build20 also passes the actual Hermes CLI and forced compression extensions. Compression finishes normally, reduces the transcript, and preserves the exact diagnostic fact in a real follow-up answer. Image extensions still need a current-build run. This is not a test of the separate FX app or Sevra receiving-side action authority. Build 19 passes T0 (33 groups / 22251), the original public Swift API signatures from a fresh external SwiftPM consumer, and all281 complete-prompt MTP/vision cache assertions after arithmetic-epoch consistency was corrected. Its existing native regression also passes sampler/governor, n-gram/template/layer parity,8.1GB/10GB exact output, live cache resizing, prefix reuse, sweep and MTP-head parity. The last MTP check was intentionally stopped on a source counterexample: its legacy command loaded the head outside an MTP-off plan. Build20 compiles mandatory-head startup planning, bounded full governor/MTP test budgets and actual late schedule reporting. T0 passes33groups/22256 assertions, CLI116, sampler/governor17groups and the full static suite, including planner64, installer fixtures, nine qualification-driver groups and three owned-process cleanup groups. The full13GB governor and12GB combined MTP/image diagnostics remain unrun. The optional V193 read-scope checkpoint integration was imported only after build20's group ended and is unbuilt in this worktree; it does not inherit build20's acceptance. The shared V193 continuation passes T0 33/22263, CLI116, HTTP835 and image-reuse76, then stops on eight plain-governor failures. A successful startup floor advisory was incorrectly treated as live feasibility. The prospective governor now checks the complete ledger against credited physical headroom and the total target at every context. Its32new pure assertions and both native pressure variants remain unrun; C05 is open for this correction. All model runs above overlapped a coordinated download; their timing and capacity interpretations are excluded. No full-window capacity ladder, retained-state capacity result, new public limit, release or installed-release acceptance is claimed. The default remains32768 and the public implementation ceiling65536. The engineering plan owns remaining C01-C22 gates. Raw failed, interrupted and passing runs remain linked below; updated source does not inherit an earlier build's native acceptance. ### Governor fixture classification after physical-fit correction (2026-09-06) [[sources/runs/2026/09/2026-09-06-configurable-context-governor-advisory-fixture-counterexample]] records shared V198: compilation passes; T0 returns32of33groups and22295 assertions with20failed expectations in four old ordinary-context fixtures. All32new exhaustion/recovery assertions pass. Both reviewers independently verify through the frozen public doctor that the old10GB whole-availability rows require7.919289600/7.921999104GB while only7.425GB remains after safety slack. A successful legacy startup advisory cannot imply live feasibility. The diagnostic-only correction retains every original10/18/44GB input, checks refusal/floor/still-unavailable behavior for the exact four invalid advisories, and adds12GB feasible near-floor rows. Existing settlement/cooldown/recovery assertions remain for feasible plans. The production guard, deadbands, allocation goldens and numerical tolerances are unchanged. The revised fixture is unbuilt at this capture, and both native governors still must pass; C05 remains open. No CLI or model case followed the failed T0 run. The verified transport release and exact evidence packet are integrated while preserving context documentation, shipped Hermes history and restored original historical sources. This prepares a common successor source closure; it is not context capacity qualification. P5 and P6 remain open, with default32768 and public ceiling65536 unchanged. ### Shared regression and excluded full-resource interval (2026-09-06) [[sources/runs/2026/09/2026-09-06-optimization-merged-correctness-v202]] records the exact shared V202 candidate: all 19 correctness groups pass, including T0 33/22346, both native governor variants at 80 assertions, the 832-assertion combined scope lifecycle, and the 835-assertion HTTP suite. The production feasibility correction and revised advisory classification now pass their relevant pure and native checks. Experimental defaults remain off. This correctness result does not qualify a larger context window. [[sources/runs/2026/09/2026-09-06-configurable-context-full-resource-swap-exclusion]] preserves the full 13 GB elastic drill's incomplete first attempt: the first answer is exact and the sampled footprint remains below the ceiling, but eight global swap-ins stop the diagnostic before a second completed answer. Two controls without an owned model also observe swap-ins. No process is attributed from global counters, no resource criterion is relaxed, and the full MTP/image diagnostic does not start after the failure. [[sources/runs/2026/09/2026-09-06-configurable-context-vision-target-counterexample]] records independent tower/reference parity and the subsequent 10 GB image suite's four failures. Full-size fruit and two-image requests receive the correct typed memory refusal. The raw 21-pass report also exposes one fixture false positive: an empty error answer counted as different image content. Its prospective correction requires successful, nonempty responses. The full-image successor reserves its actual attention workspace at 14.5 GB with a 3072-token prefill reservation and a 20.5 GB real preflight; the ordinary quality probe remains at 10 GB. Its outcome is not yet recorded here. P5 has not run. Full resource, retained/minimum/pass-transition capacity and remaining client/release/installed acceptance remain open. The public ceiling is still 65536 and the default is 32768. Neither a plan nor an excluded interval is a 262144-token qualification. ### Full images and actual clients complete (2026-09-06) [[sources/runs/2026/09/2026-09-06-configurable-context-full-image-and-actual-client-pass]] records all 25 full-image assertions passing on V202, including successful nonempty responses for the different-picture predicate. The original images use an explicit 14.5 GB total target and a 3072-token prefill reservation, after a 20.5 GB real preflight. That override is local to the full-image server. The original 10 GB refusal remains preserved. A separate ordinary 10 GB server passes all 15 quality assertions. Actual pinned Hermes at 65K passes its tool/follow-up/title/image flow with exactly one fixture tool execution, and actual Ollama returns the dog image's correct subject. `Tools/verify.sh`, the helper and testing instructions now carry that exact full-image profile. The helper's semantic correction was exercised; its subsequent header edit is documentation only. No runtime source, model ceiling, default option, image fixture or memory criterion changed. Timing and capacity conclusions remain excluded because global swap-ins continue. Full resource, P5 and release/installed acceptance remain open. ### Structured overflow refusals and verification completeness (2026-09-07) [[sources/runs/2026/09/2026-09-07-configurable-context-ollama-overflow-wire-counterexample]] preserves the original HTTP suite's 73/74 result and four independently recorded live HTTP 400 responses. Length enforcement works, but the old shell assertion matches obsolete prose and both Ollama overflow branches omit the structured code/details used by other refusals. A focused source successor keeps the legacy string and status, adds the existing typed refusal wrapper, and exercises 21 new pre-header/code/no-success assertions across all seven dialect variants. The shell suite now checks status and code directly. Its shared build and affected acceptance are pending. The full verification script also now treats missing vision-reference dependencies and insufficient full-image headroom as failed acceptance, while preserving the no-launch behavior and rerun guidance. The reviewed provider-free shell/predicate checks live in the shared V209 evidence; no mandatory vision skip can silently yield a passing full battery. The successful V202 client source is marked discarded for resource/timing measurement because its recorded global swap-ins are nonzero. Its exact raw body and passing correctness assertions are preserved unchanged. The flag is measurement disposition, not a claim that those client assertions failed. The independent runtime correction leaves default 32768, public 65536, all optional defaults and the memory ledger unchanged. [[sources/runs/2026/09/2026-09-06-optimization-vision-acceptance-profile]] [[sources/runs/2026/09/2026-09-07-optimization-typed-context-wire-counterexample]] ### Additional OpenAI code field found by the unchanged matrix (2026-09-07) [[sources/runs/2026/09/2026-09-07-optimization-typed-context-wire-counterexample]] preserves the shared V211 successor. Build, T0 22346 and CLI116 pass; the extended context matrix passes 854 of 856 assertions and stops before the two actual HTTP suites. All four Ollama structured-code cases pass. Only the OpenAI JSON/SSE structured-code checks fail: their local helper puts the code in the existing type field but omits error.code. The one-line V214 correction adds error.code while retaining type, message and HTTP status. The same 856 assertions and actual suites are being rebuilt and rerun on V215. No fixture, criterion, default, context ceiling or memory constant changes. The V211 result remains preserved, not replaced. A byte-exact local rollback copy of installed public v0.2.11 is prepared with both binary and Metal hashes; its doctor accepts configured65536. This is preservation and a model-free configuration check only: rollback itself and installation of the new context work have not run. ### V215 current interface acceptance and unrun capacity matrix (2026-09-07) [[sources/runs/2026/09/2026-09-07-optimization-typed-context-correctness]] records the complete affected successor: T0 33/22346, CLI116, native context serving856, actual API robustness74 and all four actual overflow cases pass. The additive Ollama and OpenAI typed-code corrections close the preceding preserved wire failures. Source/driver/Metal identities and raw outputs remain exact. Native context serving records8 swap-ins and API robustness28, with zero swap-outs: their correctness predicates stand, while the source is discarded for resource/timing use. No inference arithmetic or memory credit changed from V202. The older V211 source receives the same metadata-only resource disposition correction with its failed raw/body bytes unchanged. The actual CLI also records exact262144 planning boundaries: cold accepts 20,583,511,296 bytes and retained accepts20,585,806,080 bytes; each rejects a one-byte smaller target. These are current discrete-policy planning results, not measured physical minima. The retained case explicitly carries four 1021-token warm inputs and their interleaved follow-ups; its5795-slot retained budget exceeds the4096 rounded slots that setup requires. The extra C19 protocols and exact raw neighbors are at [[sources/runs/2026/09/2026-09-07-configurable-context-v215-minimum-and-transition-protocols]], and the ordinary cold/retained protocols at [[sources/runs/2026/09/2026-09-07-configurable-context-v215-capacity-protocols]]. All eight profiles/sixteen main rungs are prepared and unrun, with full model payload verification still required. No inference from synthetic filler is an answer-quality or latency-calibration result. A further idle control records16 global swap-ins over83.21061301231384s, including a45s wait with no owned model. This is preserved and excluded at [[sources/runs/2026/09/2026-09-07-configurable-context-v215-idle-swap-control]]. Full resource, full-window capacity, release, installed acceptance and actual rollback remain open. Public65536 and default32768 are unchanged. ### Final harness review and resource hold (2026-09-07) The bounded release-shell review is complete. Its exact20case-tested correction now requires the expected doctor exit2 under pipefail as well as typed insufficient_memory, validates the nested discovery cap before any POST, and requires context_length_exceeded alongside HTTP400. This supersedes the preceding pending harness status; the installed model suite itself is unrun. The original failed proposal and passing successor are preserved at [[sources/runs/2026/09/2026-09-07-optimization-default-and-release-acceptance-review]]. All143runtime source files still match the exactV215archive in both worktrees, recorded at [[sources/runs/2026/09/2026-09-07-configurable-context-v215-prepared-state-audit]]. The final readiness snapshot has21,662,023,680reclaimable bytes, below the ordinary capacity profile's25GB prerequisite, with further background swap activity. There is no owned model/compiler and the model lock is free. The raw snapshot is at [[sources/runs/2026/09/2026-09-07-configurable-context-v215-capacity-preflight-hold]]. No capacity or full resource run was attempted. The eight frozen profiles remain unconsumed; P5/P6, remaining resource acceptance, installation and rollback remain open. Default32768 and public65536 remain in force. ### Selected-candidate acceptance correction (2026-09-07) [[sources/runs/2026/09/2026-09-07-configurable-context-candidate-selection-acceptance]] preserves the full five-file acceptance correction and its first counterexamples. Static, planner and installer gates consistently use the explicitly selected candidate, including paths with spaces/apostrophes. The installer packages that executable and its colocated Metal library, then verifies exact installed bytes through fresh/repeated activation, checksum refusal and legacy upgrade. The static entrypoint retains the mandatory qualification-driver checks. All64 planner assertions remain unchanged. Six static-selection fixtures, ten real-install.sh selection/fault fixtures, and the unchanged nine qualification-driver groups pass. The old installer fails seven of the same ten cases. A separate private fixture installation of V215 also passes exact binary/Metal identity; the real user installation and rollback remain unexercised. This does not close release/installed C22. On the loaded machine, the old quoted planner path produces18pass/46fail. The corrected quoted and normal paths both produce58pass/6fail, with the same six malformed-checkpoint diagnostics blocked by the real startup memory guard. A captured direct fixture shows9.2GB reclaimable and an8.1GB target, then insufficient allocation headroom before checkpoint parsing. These failures are preserved; no expected error or guard is weakened. A complete64/64 planner and full static pass remain required on a quiet machine. All143runtime source files still match frozenV215 in both worktrees; no frozen P5 driver changes. P5/P6 remain open. Eight prospective profiles/sixteen main rungs remain unrun; no model payload verification or capacity run starts during this work. Public 65536 and default32768 remain unchanged. Loaded-machine results carry no resource or timing qualification. ### Final acceptance-script integration (2026-09-07) [[sources/runs/2026/09/2026-09-07-optimization-verification-selected-paths]] records the full verification entrypoint preserving the selected executable through its evaluated call sites; its nine focused fixtures pass. Mandatory vision, context-qualification and installer checks remain enforced. [[sources/runs/2026/09/2026-09-07-configurable-context-sampler-exit-and-binary-selection]] preserves the five original sampler false passes and the corrected eight-case full-shell fixture, plus all17real V215 sampler/NumPy/governor checks passing. API version checks now require both the selected quoted executable and a successful process exit. No new full74-case API server run is claimed. [[sources/runs/2026/09/2026-09-07-optimization-mandatory-sampler-and-final-handoff]] records the final mandatory static wiring: sampler8/static9/syntax3pass, exact integrated files,95public-number checks and zero validation errors. All143runtime sources and the frozen capacity drivers remain unchanged. The shared context plan, source closure, public docs and client fixtures are now integrated. P5/P6 remain open: all eight capacity profiles/sixteen main rungs are unrun; the current six headroom-blocked planner checks, remaining full resource tests, actual final clients, release/installation and rollback still require their own passing evidence. Public65536/default32768 remain. ### Final release-response acceptance review (2026-09-07) [[sources/runs/2026/09/2026-09-07-optimization-serial-and-installed-gate-integration]] records16original installed-gate false passes and one quoted-binary-path failure, followed by19/19passing complete-shell fixtures and mandatory static wiring. The separate successor at [[sources/runs/2026/09/2026-09-07-configurable-context-openai-release-completion]] preserves those19cases and exposes nine OpenAI false passes. All28complete-shell fixtures pass after requiring successful curl, one successful text completion, nonempty assistant content and a positive integer completion-token count. These are local process/response fixtures; no real installed model, socket, release or rollback ran. The original failed scripts and raw results remain preserved. The existing mandatory static entrypoint runs the expanded suite. [[sources/runs/2026/09/2026-09-07-optimization-cached-planner-build-and-typed-parity]] separately records the successful temporary cached build and441exact typed planner comparisons, with shared V215runtime/release/build state restored. It grants no new context capacity or performance claim. All eight V215P5 profiles/sixteen main rungs remain unrun. Full resource acceptance, P5/P6, actual installed/release/rollback gates remain open; the original ordinary capacity preflight remains25GB, public context65536 and default32768. The source/protocol dependency split is preserved at [[sources/runs/2026/09/2026-09-07-optimization-planner-metadata-and-context-boundary]]. The full262Kexpansion campaign remains necessary for this context-expansion goal, while unchanged32768/65536optimization retains its own applicable gates. No context profile is waived. That source also records the corrected e2e_release.sh executable bit; verify.sh remains at its tracked0644mode. ### Portable software acceptance; native qualification deferred (2026-09-07) [[sources/runs/2026/09/2026-09-07-configurable-context-portable-software-acceptance]] records the exact source-only policy pass, nine passing Python suites, adversarial evidence fixes, and portable source/binding checks. See the current context plan addendum for Carlos's explicit native-testing deferral. These results qualify software predicates and geometry only; no new model capacity, numerical result, memory measurement, real client or release is claimed. The public/default/mode ceilings remain unchanged. ### Explicit per-window software checks (2026-09-07) [[sources/runs/2026/09/2026-09-07-configurable-context-window-cli-matrix]] preserves the selected V215 binary/archive hashes and exact outputs for the completed window matrix, including public-ceiling and model-ceiling refusals. The current testing guide documents the reusable command. The identified candidate passes the public planner and bounded scheduling checks; the source-only platform-method extraction is reported as a source difference. These are software contract results, with no fresh model load, tensor capacity, throughput, parity or release qualification. The initial test-only ledger assertion failure and its corrected validator are both retained. ### Decode wall-time attribution with natural GPU boundaries **The largest individual elapsed-time category is the expert file-read phase.** In the qualified 512-output observation it occupies 33.6654% of decode, staging preparation adds 1.6095%, GPU encoder execution occupies 31.9764%, and the remaining CPU work/coordination occupies 32.7487%. The completed smaller-cache control spends 42.8458% on reads plus staging, so a generic claim that expert I/O is always 40–45% is unsupported. These percentages describe specific workloads and memory settings, not every use of Slotstream. **Protocol and scope.** Apple M5 Pro, 48 GiB unified memory, internal 2 TB SSD, macOS 26.6.2. Native Slotstream base `449f3841d65cbca8346ab6ae8092eb0948d92dbd`, mlx-swift `0bb916c67f4b9e5c682cbe02a42c701c93ab5021`. The working HEAD later differs in the version string and a server helper's visibility, not the measured inference implementation. The frozen profiling binary SHA-256 is `3121d93bdd1040cdb0500b4444a043cc0e59fb51c6b3ad12696a2dad0250471c`. Every arm uses a fresh server, the same 521-token text prompt, one same-length warmup, greedy generation, MTP depth one when enabled, prefill chunk 256, no elastic resizing and no prefix retention. Profiled and disabled arms alternate order by round. The large profile has a 24 GB total target and 5,702 expert slots, 15.765 GB of quantized weights. The small control has a 10 GB target and 1,217 slots, 3.365 GB. This is not the auto-sized demo server's exact cache configuration. Decode excludes model load, prefill, request queuing, and trace serialization. It uses the generator's own decode interval, including callbacks and decode state work. On/off pairs preserve exact output-token IDs, read-byte counts and slot counts. The 512-output pair takes 43.943496584 seconds disabled and 42.447254083 seconds profiled. Only one of its three requested pairs passes the no-swap gate. Its measured native workload reads 129,674,649,600 bytes in 46,902 expert records, approximately 253.271 MB per emitted token. Cache hit rate is 81.6340%; 190 of 321 drafts are accepted. These are decimal bytes/MB/GB. The 128-output large-cache cohort retains three of four pairs, reads 27,048,038,400 bytes per measured response, and has an 84.3221% cache hit rate. **Partition of 100% of elapsed decode.** The long column is its sole qualified trace. The short column divides summed exclusive category seconds by summed decode seconds across all three clean traces, including the slow pair. It is not a best-of selection. Small display-rounding differences do not change the exact 100% partition. | Elapsed-time attribution | 512 outputs, one clean pair | 128 outputs, three clean pairs | | --- | ---: | ---: | | SSD file-read workers and waits | 33.6654% | 30.2487% | | Allocate and wrap RAM staging buffers | 1.6095% | 1.7536% | | GPU writes from staging RAM into the expert cache | 1.8817% | 1.6831% | | GPU expert calculations and output mixing | 8.1593% | 8.6271% | | GPU cache writes and expert calculations together | 4.6890% | 4.0888% | | GPU recurrent layers, with associated normalization and routing | 9.7901% | 10.2493% | | GPU attention layers, with associated normalization | 2.4684% | 2.5197% | | GPU remaining normalization, routing and residual operations | 2.1355% | 2.2359% | | GPU vocabulary, ngram embedding, MTP head, sampling and other passes | 2.8523% | 2.9522% | | CPU Metal command preparation and driver calls | 9.9556% | 11.2440% | | CPU graph evaluation, scheduling and synchronization | 10.4297% | 10.4863% | | CPU model graphs, cache bookkeeping, MTP state and output | 12.3635% | 13.9113% | | Total before display rounding | 100% | 100% | **What loading means.** On a miss, `ExpertStore.readBatchChecked` dispatches and joins worker lanes issuing `pread` calls into aligned temporary RAM buffers. Each complete quantized expert record is 2,764,800 bytes across nine pieces. Staging is wrapped as MLX arrays; later GPU scatter kernels copy the quantized bytes into persistent cache slots. The CPU and GPU share physical RAM, so this is SSD to staging RAM, then a copy within RAM, then GPU reads/calculations. There is no discrete-GPU PCIe upload stage in this Mac path. The file layer requests `F_NOCACHE` and disables read-ahead. The ordinary checkpoint path ignores those calls' return values, so this trace establishes file-read-path latency and requested bytes, not a hardware measurement of NAND service time or guaranteed physical SSD bytes. The read row includes worker scheduling, syscall/OS work and waits, with separately measured staging allocation/wrapping removed. **Cache behavior behind the timing.** These profiles use one global slot pool shared across layers. Each slot holds one quantized expert record identified by its layer and expert ID. A hit sets its reference bit and pins the slot for the active readers. Misses are deduplicated within the request, assigned unpinned victims by the CLOCK scan, read in bounded batches, and installed through dependent scatter graphs. The scan skips pinned slots and clears reference bits before reconsidering recently used entries. The optional layer-local floor branch was not exercised in these measurements. Layer completion releases the lifetime dependency that permits subsequent slot reuse. This policy preserves recently reused experts without keeping all 512 experts of every layer resident; the hit rate depends on the actual routing sequence. Under MTP, batched verification, duplicate expert requests and accepted output counts differ, so compute bytes per emitted token from the native read counter rather than assuming one ordinary forward per output. The ngram lookup path has separate storage/caching and timing; its measured decode work is included in the smaller CPU/GPU categories. **Why the GPU/CPU rows are honest but not individual-kernel timers.** CPU host scopes, MLX Metal command-preparation spans, and native GPU encoder timestamps share a checked clock. Analysis sweeps all endpoints and counts each interval once: observed GPU execution first; then host read/staging/ngram-fetch phases; then CPU Metal encoding; then the remaining host scope. Concurrent CPU work is represented under the GPU category during that overlap. This is a declared wall-time attribution, not a unique causal decomposition, the sum of CPU utilization and GPU utilization, or pure arithmetic time. A GPU encoder's envelope includes memory access, scheduling and barriers. CPU evaluation/synchronization means time in that host scope after observed GPU execution and Metal encoding have been removed; it is not all active CPU computation or all sleeping. Natural GPU passes can include several functional operations. Cache writes and expert calculations share a row when they occur in the same pass or overlap on the GPU; that time cannot be split without changing execution. GDN and attention rows likewise include associated normalization and routing in their natural passes. The remaining tensor row includes normalization, routing and residual operations. Untagged GPU passes occupy only 0.062293% of the long trace and stay explicitly visible in the full partition. MTP native draft/verification timers overlap these categories and must not be added again. GPU-pass union and host/encoder attribution account for every nanosecond of the measured interval, with zero uncovered time, invalid GPU samples or overlapping host labels; CPU/GPU clock scale is one. **Profiler qualification and limits.** Mode4 keeps natural encoder boundaries; counter-buffer allocation/rollover and trace bookkeeping still have a cost. Mode0 disables the active instrumentation in the same diagnostic build, so it retains the compiled tag/branch scaffolding and is not a separate pristine-binary A/B. The source stays isolated from production. The short large-cache paired timing changes range from -1.938% to +14.964%, with a median -0.558%. The long pair is -3.405%. Negative values are timing variability, not profiler speed gains; these observations do not prove zero overhead or a fixed overhead bound. The separate encoder-splitting profiler increases its short pair's duration by 10.73% and is excluded. | Protocol | Clean matched pairs | Disabled median, seconds | Profiled median, seconds | Median paired profiling change | | --- | ---: | ---: | ---: | ---: | | 10 GB, MTP off, 32 outputs | 1 | 4.228180 | 4.160559 | -1.599% | | 10 GB, MTP off, 128 outputs | 1 | 16.893156 | 16.837556 | -0.329% | | 24 GB, MTP on, 128 outputs | 3 | 10.129347 | 9.933006 | -0.558% | | 24 GB, MTP on, 512 outputs | 1 | 43.943497 | 42.447254 | -3.405% | Every accepted decode has unchanged global swap-in and swap-out counters, nominal thermal state at the generator observations, and low-power mode disabled. Counter readings and process contention checks are sampled, not continuous proof that the machine was otherwise idle. Earlier long attempts, early endings, headroom refusals, the counter allocation failure and every swap-contaminated pair remain in [[sources/runs/2026/09/2026-09-10-decode-attribution-discarded-attempts]]. No clean sibling of an invalid pair enters the timing comparison. The 128- and 512-output measurements do not establish a workload-independent distribution or long-context behavior. **First-principles implication.** Each layer's router must produce expert IDs before the host can resolve misses. The host then waits for the required bytes and submits dependent cache/expert work; layer evaluation also provides the lifetime boundary needed to reuse cache slots safely. Sparse activation reduces weights needed per step, but miss traffic remains hundreds of MB per output here, and repeated CPU/GPU handoffs remain costly even on cache hits. SSD reads are the largest single category; CPU graph/Metal coordination is comparable in aggregate. Removing the file-read phase alone would leave about two-thirds of the observed time, before accounting for changed overlap and cache-copy work. A faster SSD or larger cache alone is therefore not evidence of a proportional whole-decode speedup. **Production and reproducibility.** All instrumentation lives in the ignored detached worktree and isolated MLX dependency copy. No production source, installed executable or installed Metal library was replaced for this measurement. The original installed binary hash matches the restoration proof, and the normal demo server is again listening on port 11434. Raw accepted GPU/CPU traces are losslessly archived and round-trip hash checked. The exact native sources, later driver checks, offline analyzer, fixture, raw results and exclusions are bound by [[sources/runs/2026/09/2026-09-10-decode-wall-time-attribution]]. Analysis can be rerun by decompressing the retained JSON traces and running the archived `analyze.py`; `summarize_final.py` specifies the cohort and grouping arithmetic. Full exclusive partition of the qualified long trace follows; the CPU loop category includes checkpoints, allocation admission, state/setup and loop work outside the explicitly tagged scopes. | Exclusive interval category | Percent of 512-output decode | | --- | ---: | | Host: expert file read | 33.665360% | | Host: graph evaluation and synchronization | 10.429658% | | CPU: Metal command preparation and driver calls | 9.955589% | | GPU: GDN layer preparation (normalization, recurrence, routing) | 9.790136% | | GPU: expert calculations and output mixing | 8.159304% | | GPU: cache writes and expert calculations together | 4.689038% | | Host: residual/other | 4.653391% | | GPU: attention layer preparation (normalization, attention) | 2.468408% | | Host: state reconciliation | 2.421033% | | GPU: normalization, routing and residual operations | 2.135506% | | GPU: cache writes | 1.881667% | | Host: staging allocation and wrapping | 1.609544% | | GPU: MTP head and state updates | 1.501168% | | Host: hyperconnections and normalization | 1.086041% | | GPU: final mixing and vocabulary projection | 1.035660% | | Host: GDN recurrence | 1.012478% | | Host: routed experts | 0.584328% | | Host: cache slot write | 0.504666% | | Host: GDN projections and convolution | 0.466066% | | Host: attention | 0.442764% | | GPU: ngram embedding and preparation | 0.212956% | | Host: sampling | 0.201171% | | Host: cache lookup and eviction | 0.190812% | | Host: output callback | 0.170326% | | Host: draft head | 0.158004% | | Host: shared expert | 0.132315% | | Host: router | 0.130325% | | Host: attention indexer | 0.091195% | | Host: GDN output | 0.069897% | | GPU: other tensor passes (untagged) | 0.062293% | | GPU: sampling | 0.040244% | | Host: ngram PLE | 0.037321% | | Host: embedding | 0.009062% | | Host: ngram row prefetch | 0.001880% | | Host: output projection | 0.000394% | ## Lossless model download: complete package, integrity and CDN delivery This section preserves the original codec and v0.2.10 deployment measurements. Hosting has since moved to Hugging Face in v0.2.11; see [[records/measurements/hugging-face-lossless-download-2026-09-06]]. The codec and model bytes are unchanged. # Lossless model transport: complete package and delivery The complete model package uses **16.12% fewer bytes**: 105,264,463,248 original bytes become 88,295,438,048 package bytes, saving 16,969,025,200 bytes. The package includes its embedded manifest; normal clients transfer only the 88,294,086,225 object bytes. Original files, tensor bits, and installed size remain unchanged. This completes the earlier sample-based compression investigation. The original first-install bottleneck was the volume of network transfer, with client concurrency and redirect costs affecting how fully a route was used. Rehosting raw multi-gigabyte shards alone did not remove that byte cost. The new layout changes the representation and makes each immutable object small enough for ordinary edge caching. ## Evidence and qualification scope [[sources/runs/2026/09/2026-09-05-slotpack-package-and-regressions]] records complete, independent Mac and Linux builds with identical manifests and compressed objects. Both hash every original file, round-trip every object, and prove exact file coverage. The native codec, manifest validation, real compressed HTTP faults, and raw multi-chunk compatibility checks pass. Linux sanitizer-backed fuzzing completed 233,401 runs without a reported error. [[sources/runs/2026/09/2026-09-05-slotpack-memory-counterexample-and-repair]] preserves a failure of the new candidate, the incorrect first repair, a bounded regression counterexample, and the successful correction. Completed autoreleased operations retained payloads until a long-lived worker finished; an autorelease pool around the complete per-object iteration, including operation creation, fixes that ownership issue. The sustained 6.44 GB test completes with 352,403,456 bytes peak RSS, whereas the old candidate trips its safety guard. The full resumed CLI reconstruction independently verifies all original files and stays below its 4 GB test guard. [[sources/runs/2026/09/2026-09-05-slotpack-public-cdn-and-linux-qualification]] captures an uninterrupted fresh Linux public pull of every object, zero raw fallback chunks, and independent SHA-256 verification of all twenty-five original files. The native download completed in 840.6 seconds on the recorded server route, with one cache hit and 4,154 misses. This is a single diagnostic result, not a universal or paired hosting speedup. [[sources/runs/2026/09/2026-09-06-slotpack-mac-and-final-candidate-acceptance]] captures the complete Mac public installation: an empty destination, a deliberate interruption, resume, and independent original-file verification. No existing model chunks seeded the new copy. Peak sampled RSS across its two download attempts was 1,558,462,464 bytes. The resumed segment reported zero raw fallbacks; the interrupted segment did not emit a final fallback counter. Do not describe this as one uninterrupted fresh timing sample. The final candidate preserves existing Swift API function references and nonescaping log forwarding. Both native clients revalidate/reuse the complete downloaded model without transfer. The actual Mac CLI source/transport overrides and signal handling pass, as do all transport, static, sampler, external-consumer, and coverage gates. The final binary loads the downloaded model at a 10 GB target and answers the bounded greedy prompt correctly. These are correctness and bounded-memory results; shared-machine timings remain diagnostic. Public release and installed-binary acceptance are recorded below. [[sources/runs/2026/09/2026-09-06-slotpack-ci-interruption-fixture-correction]] preserves the failed v0.2.9 release gate and its test-only correction: interruption now follows durable progress, including intentionally delayed startup. The complete corrected local suite passes. No v0.2.9 asset was published; v0.2.10 is the corrected published release. [[sources/runs/2026/09/2026-09-06-slotpack-ci-coverage-regression-closure]] preserves the first v0.2.10 main-CI ratchet failure after every functional gate passed. Direct resumed-byte and CDN-status tests raise local downloader line coverage to 96.87% without changing production code or lowering any floor. The corrected main-CI result is recorded with publication acceptance. [[sources/runs/2026/09/2026-09-06-slotpack-v0210-publication-and-installed-acceptance]] closes publication: [v0.2.10](https://github.com/carloslfu/slotstream/releases/tag/v0.2.10) is the latest published release, both public CI workflows pass, and the ordinary installer activates the exact signed archive. The installed release passes default/source/cancellation checks, revalidates and reuses all original files without transfer, independently verifies their hashes again, and loads the reconstructed model to return READY at a bounded 10 GB target. The passing main-CI revision adds test coverage only; its production sources match the signed release. Fresh `slotstream pull` now uses the lossless CDN package by default. ## What compression and CDN caching change The exact compressed fraction is 0.838796260. When the network is limiting, the ideal transfer-time reduction is therefore the same 16.12037% on a slow or fast connection. Decoding, file writes, and final verification overlap downloads where possible. They impose a processing ceiling on faster links, so whole-install time cannot be inferred from the fraction alone. Transfer-only estimates are about 2 hours at 100 Mbps or 8 hours at 25 Mbps. These rounded estimates exclude protocol overhead, retries, changing throughput, and any unhidden processing. The exact idealized byte arithmetic is: | Nominal link | Original transfer | Compressed package transfer | Ideal saving | |---|---:|---:|---:| | 25 Mbps | 33684.6 s | 28254.5 s | 5430.1 s | | 100 Mbps | 8421.2 s | 7063.6 s | 1357.5 s | | 500 Mbps | 1684.2 s | 1412.7 s | 271.5 s | | 1000 Mbps | 842.1 s | 706.4 s | 135.8 s | | 5000 Mbps | 168.4 s | 141.3 s | 27.2 s | CDN hits affect latency, route throughput, and origin load. They do not remove the remaining client download bytes and must not be counted as another fixed percentage reduction. A new edge can miss the cache and fetch from R2; a repeated read of the same immutable object can hit. Real HIT responses and full-client results must be preserved as measurements rather than inferred from cache eligibility. ## Implementation and operational contract The qualified default is a stable Slotpack v1 representation: bounded independent objects containing packed four-bit weights plus losslessly predicted BF16 metadata. The model is never requantized. Object hashes, reconstructed-chunk hashes, and original whole-file hashes remain mandatory. The pinned manifest rejects gaps, overlap, unsafe paths, unknown formats, and length/range overflow. Each network worker owns a persistent URLSession. Concurrency starts at eight and tests larger counts only while throughput improves; explicit counts remain fixed. The decoder queue and per-iteration object lifetime bound memory. Disk admission accounts for reconstructed output; verified resume bits follow synced writes. Cancellation drains work. Optional-file cleanup cannot race in-flight writes, and missing optional files no longer cause readiness to request another download indefinitely. Cloudflare R2 stores immutable content-addressed objects under a manifest-specific prefix behind the public custom domain. The application wildcard Worker is excluded for that hostname. The original Hugging Face pins remain an independent raw fallback. Existing raw resumes and explicit source overrides retain their semantics. See `docs/DOWNLOAD-FORMAT.md` and the producer/qualification tools under `Tools/slotpack/`. ## Hugging Face transport: unchanged bytes and publisher cost The unchanged lossless Slotpack package is published in the separate public repository [carloslfu/Qwen3.8-Flash-Next-MLX-4bit-Slotpack](https://huggingface.co/carloslfu/Qwen3.8-Flash-Next-MLX-4bit-Slotpack), pinned at commit `13ec15dcebdddc817b57f0f9087c5ef82018f10e`. The original raw repository and its original pinned revision remain intact. Keeping the representations in separate repositories prevents ordinary Hugging Face clients from downloading both. The original compression and integrity measurements in [[records/measurements/lossless-model-download-2026-09-05]] remain valid: this move changes hosting, not the model, codec, object bytes, manifest or reconstructed file hashes. [[sources/runs/2026/09/2026-09-06-slotpack-hugging-face-publication]] captures committed-object identity verification and anonymous public reads. [[sources/runs/2026/09/2026-09-06-slotpack-hugging-face-linux-and-gates]] captures a complete fresh anonymous download using the exact production Swift/C engine and embedded Hugging Face default in the Linux bandwidth instrument. Every original file passes the client checks and independent GNU sha256sum verification; the client reports zero raw fallback chunks. Its download counter reports 799.2 seconds; subsequent explicit verification and independent hashing are outside that counter. This is a single route diagnostic, not a paired host comparison or universal installation-time claim. The final native Mac static and installer checks also pass. [[sources/runs/2026/09/2026-09-06-slotpack-hugging-face-full-mac-qualification]] additionally records a complete native Mac installation begun in an empty directory, interrupted by a hostname-resolution failure and resumed without external chunks: every original file independently matches. The completed resumed segment reports zero raw fallback chunks; the interrupted segments emitted no final fallback counter. Public main CI passes all functional, consumer, catalogue and coverage gates. [[sources/runs/2026/09/2026-09-06-slotpack-v0211-publication-and-r2-retirement]] closes publication and installed acceptance: the signed [v0.2.11 release](https://github.com/carloslfu/slotstream/releases/tag/v0.2.11) passes the ordinary installer, actual CLI selection/cancellation checks, complete model reuse and independent hashes, and a bounded loaded-model reply. Hugging Face resolver throttling can require a full reset-window wait. The shared HTTP transport now honors its RateLimit reset and longer Retry-After values, with a bounded, cancellable wait. Real HTTP fixtures verify retry and prompt cancellation. Original compressed-object, reconstructed-chunk and whole-file checks, raw compatibility and resume remain intact. The legacy weights.sevra.page hostname now redirects through Cloudflare's static asset service to the same exact Hugging Face commit. The deployed version serves assets directly, has no bindings, and does not invoke a Worker function. Its staged public object checks pass. [[sources/runs/2026/09/2026-09-06-slotpack-hugging-face-legacy-compatibility]] additionally verifies the unchanged released v0.2.10 client following the live redirect, reconstructing a missing model shard and metadata with exact original hashes and zero raw fallback. The complete native Mac installation also passes after preserving and resuming its network-interrupted progress. The redundant model-only R2 bucket has now been emptied and deleted, with absence and continued legacy redirect delivery verified in the publication receipt. Under its current [public storage policy](https://huggingface.co/docs/hub/en/storage-limits), Hugging Face provides best-effort public hosting without publisher charges per download. Limits still apply. No paid plan, metered proxy or billing fallback is enabled. [Static asset redirects are free](https://developers.cloudflare.com/workers/static-assets/billing-and-limitations/); the redundant model-only R2 bucket is deleted, so it no longer accumulates new storage/read usage. Deletion does not remove any charges that may already have accrued. Cloudflare volume/address analytics now cover only legacy redirect traffic. Hugging Face's default model counter follows selected query files, while the compressed client fetches hash-named objects and embeds its manifest. Neither counter measures completed compressed installations. GitHub release acquisition proxies remain available; no telemetry or counting-only request was added. ### v0.2.12 release qualification **Status: candidate closed without publication. See [[records/measurements/release-qualification-0-2-13]] for the corrected release.** The original unified optimization campaign is complete in its recorded scope. This release binds that implementation to the intervening public documentation, configuration and operating-policy changes. The inference algorithms and tuning constants are unchanged; planner/help presentation and the version are updated. The public context ceiling is unchanged. Larger context qualification remains a separate plan, not a claim made by this release. ## Clean-build identity correction The first CI run compiled successfully but failed to bind its artifact because SwiftPM replaced `.build/release` with an architecture-specific symlink, losing the pre-build receipt. The Makefile now asks SwiftPM for its real output path before and after the build. A compiler-free fixture reproduces the original failure, verifies a clean build in a path with spaces, proves non-build targets do not invoke SwiftPM, and still refuses source mutation during the build. All five identity tests pass. The public archive will include the CI identity and reconstructible source archive inside its attested payload. Independent external-consumer and coverage builds now run in separate CI jobs. Every original CI command is retained. This parallelism uses separate hosted runners; local model and compiler work remain serial. ## Coverage review The completed instrumented job at `4927998` passed 44 check groups and 27,342 assertions with no failed or skipped check, then failed the coverage ratchet. It reported 82 new files with no baseline and ten lower per-file percentages. This is T0/T1 coverage; it does not instrument the separate native-model suite. Two floors remain unchanged. The server comparison exposed removed stock-SDK no-op and image/text conversion checks, which are restored. The context-budget and governor checks are also restored with explicit refusal/hold/shrink cases. The prefix-cache comparison lost only an attributed closing brace; added tests exercise partial image checkpoints, processor identity mismatch, invalid and overflowing extents, state preservation on rejection, and reset/drop behavior. Fresh CI must confirm both retained floors. For the other eight files, exact unchanged-line mapping found no previously covered line losing its hit. Absolute hit counts stayed equal or increased; new native-only paths enlarged their denominators. Their T0/T1 baselines are updated individually from the preserved report, with no automatic global reset. New-file baselines use measured coverage, including explicit zeroes for native-only diagnostics. The gate still rejects missing or regressing floors. | File | Previous hit/found lines | Candidate hit/found lines | Previous floor | Measured floor | |---|---:|---:|---:|---:| | `Sources/Slotstream/Engine.swift` | 101/683 | 101/1087 | 14.79% | 9.29% | | `Sources/Slotstream/Generate.swift` | 81/429 | 114/1305 | 17.25% | 8.74% | | `Sources/Slotstream/Model.swift` | 23/278 | 34/973 | 8.27% | 3.49% | | `Sources/Slotstream/NgramStore.swift` | 6/294 | 6/447 | 2.04% | 1.34% | | `Sources/Slotstream/VisionPrompt.swift` | 33/54 | 33/118 | 58.93% | 27.97% | | `Sources/SlotstreamDiagnostics/Diagnostics+Governor.swift` | 146/175 | 161/198 | 83.43% | 81.31% | | `Sources/SlotstreamDiagnostics/Diagnostics+Runtime.swift` | 119/136 | 361/422 | 86.73% | 85.55% | | `Sources/SlotstreamDiagnostics/Diagnostics+Vision.swift` | 511/561 | 612/689 | 90.31% | 88.82% | No aggregate percentage is presented as coverage of the complete inference implementation. Final release testing will run the original full model battery and installed-release checks against the actual public download. Evidence: [[sources/runs/2026/09/2026-09-10-release-0-2-12-prepublication-corrections]]. Performance remains in [[records/measurements/optimization-final-composition-2026-09-09]] and practical serving observations in [[records/measurements/user-server-throughput-2026-09-10]]. ## Malformed checkpoint gates on small CI hosts The next static job reached the planner checks and exposed six tests that assumed enough real host memory to reach checkpoint parsing. The production startup guard correctly runs first and refused the small hosted runner. The fixtures now independently validate the exact production CheckpointIndex through the existing packed-artifact verifier. Every tiny invalid fixture fails metadata construction before payload reads, model allocation or writes. The normal run path is still required to exit cleanly with the same parser error or its earlier, explicitly identified memory refusal. Five regression tests exercise both permitted paths and reject wrong or absent parser errors, success exits, traps, unexpected startup errors and output creation. All six actual malformed fixtures also passed locally with their exact diagnoses. Production memory guards and inference code are unchanged. Final CI confirmation remains pending. Evidence: [[sources/runs/2026/09/2026-09-10-release-0-2-12-planner-fixture-correction]]. ## Restored coverage confirmation The next instrumented CI job passed all 44 groups and 27,381 assertions, with no failure or skip. PrefixCache coverage rose to 86.83% and Server to 11.64%, clearing both retained floors. Their floors now rise to these measurements, alongside Plan at 80.14% and Governor at 28.80%. One remaining percentage flag concerned Diagnostics+PrefixCapacity itself. Its added negative-input assertions enlarged LLVM's counted denominator from 137 to 187 lines while hit lines grew from 128 to 171. No covered unchanged line lost a hit. The assertion that would report unexpected success for an invalid image correctly remains unexecuted. This diagnostic floor is updated individually to 91.44%; the production floors are not reduced. Applying these floors to the preserved CI report passes the full ratchet. Evidence: [[sources/runs/2026/09/2026-09-10-release-0-2-12-restored-coverage]]. Final static CI and installed public-artifact acceptance remain pending. ## Candidate closed without publication This candidate did not publish a release archive. Its instrumented CI passed all 44 groups and 27,381 assertions and the coverage ratchet, but the isolated static-entry-point fixture omitted the newly added planner test dependency. The release workflow stopped before packaging. The v0.2.12 tag remains at 4ccb2cf and is not moved or deleted. The corrected release is tracked in [[records/measurements/release-qualification-0-2-13]]. Evidence: [[sources/runs/2026/09/2026-09-10-release-0-2-13-harness-correction]]. ### v0.2.13 release qualification **Status: candidate closed without publication. See [[records/measurements/release-qualification-0-2-14]].** This release carries the completed unified optimization campaign and the intervening public integration/documentation work. The preparation history, clean-build correction, restored serving/prefix checks and scoped coverage review are preserved in [[records/measurements/release-qualification-0-2-12]]. The v0.2.12 tag failed before publication and remains unchanged. The correction adds the planner-test dependency to the isolated static entry-point fixture and verifies that its failure stops the pipeline before native checks. All 22 entry-point tests and the planner helper tests pass locally. Both CI and release workflows now execute these checks before the native build, while retaining the complete later static gate. The complete local static invocation against the older installed 0.2.11 source-qualified executable passed its Python, runtime, pull and transport checks, including sustained memory. Its single planner failure concerns the new large-machine help wording absent from that old executable; it is not counted as a passing current-release gate. The separate installer suite passed. Current-source CI and full public-artifact acceptance remain required. The only compiled-source change after the passing instrumented candidate is the version string from 0.2.12 to 0.2.13. Inference algorithms, tuning constants, context limits and the original model acceptance workload remain unchanged. The public archive must pass checksum and signed provenance verification, match its source identity, and pass the original full local model battery and installed-release API checks. A prior installation is retained for rollback. Evidence: [[sources/runs/2026/09/2026-09-10-release-0-2-13-harness-correction]]. Performance remains in [[records/measurements/optimization-final-composition-2026-09-09]] and [[records/measurements/user-server-throughput-2026-09-10]]. These comparisons are between selected/reference paths within the same build, not a direct A/B against the previous public release. Sustained decode was flat or slightly slower in its measured profiles; this release makes no universal TPS claim. ## Candidate closed without publication The hosted static suite again failed six startup-path fixture checks after its preceding gates passed. The fixture did not recognize RequestFailure's human-readable model-allocation errors, accepting only coded planner errors. The old wrapper hid the precise command output; the revised gate preserves it. The failed v0.2.13 tag is unchanged. The corrected gate and build-once release flow are tracked in [[records/measurements/release-qualification-0-2-14]], with evidence in [[sources/runs/2026/09/2026-09-10-release-0-2-14-candidate-gates]]. ### v0.2.14 release qualification **Status, September 11: v0.2.14 is published, installed and serving locally. CI, all 25 original model gates and all 31 installed-release checks pass. The final unchanged governor and long-prompt memory tests completed with zero swap activity. Full local public-artifact acceptance is complete for this exact release.** Release: [v0.2.14](https://github.com/carloslfu/slotstream/releases/tag/v0.2.14), tagged at `ac7d7e53ca2397bb8840a97f028509fbb0a8f688` after the complete [main CI run](https://github.com/carloslfu/slotstream/actions/runs/34539578315) passed. The [publication run](https://github.com/carloslfu/slotstream/actions/runs/34541172114) published the exact successful CI archive without recompiling it. Public archive SHA-256: `2fd5bfab8073b29095148274b0e5857a43dc20bc530bcacb4c9180ccc5ac2b52`. The public download matched the CI artifact exactly, passed GitHub attestation verification, and reconstructed every source file with matching hashes. The ordinary public installer installed those exact binary, Metal, identity and source-archive bytes. The previous installation remains available for rollback. The normal demo server was restored after the September 11 checks, reports `0.2.14` and returned `OK` to the bounded smoke request. Chrome and Wispr Flow were reopened. That response is not a performance benchmark. ## Release and CI qualification This release carries the completed unified optimization campaign, intervening public integration/documentation work, and completed decode-profiling records. Profiling instrumentation remained isolated and is not in production sources. The unpublished preparation history remains in [[records/measurements/release-qualification-0-2-12]] and [[records/measurements/release-qualification-0-2-13]]. Those failed tags are preserved and have no binary release archives. The corrected planner fixture independently requires the exact malformed checkpoint error before checking startup. It accepts the exact early memory refusal descriptions and prints both command diagnostics on failure. Seven helper regressions, 22 static entry-point regressions and nine archive-verifier regressions passed. The successful hosted run then passed the full static, runtime, sampler/governor, check-catalogue, public-library and coverage jobs. The instrumented catalogue passed 44 groups and 27,381 assertions, with no failed or skipped group. Coverage floors remained enforced. Main CI archives the candidate, tests it and verifies that the tested bytes still match the archive. Publication requires successful main CI for the exact tag commit, verifies the checksum, identity, reconstructed source and version, then signs and publishes the same archive. Failed or unfinished CI is ineligible. The only compiled-source difference from the earlier passing instrumented candidate is the version string. A concurrent documentation-only child commit was preserved for the final documentation push; its compiled source closure is identical to the release commit. Evidence: [[sources/runs/2026/09/2026-09-10-release-0-2-14-candidate-gates]] and [[sources/runs/2026/09/2026-09-10-release-0-2-14-publication]]. ## Installed public-artifact acceptance The original full suite ran unchanged against the installed public binary. It finished **20 passed, five failed, with no required skip**. The failures came from four model runs: governor, combined MTP/vision, long-prompt memory, and context-check, whose fit and memory assertions both failed on swap activity. All observed peaks in those excluded intervals were below their explicit targets and there were no new swap-outs. The original strict zero-swap gates still rejected them; these are preserved failures, not passing measurements. | Area | Current public-artifact evidence | | --- | --- | | Original model battery | 25 of 25 unique gates qualified across the initial pass and unchanged targeted reruns | | Speculation and vision | Complete original 12 GB diagnostic passed determinism, exercised vision and validated its zero-swap memory interval | | Context diagnostic | Original 2,048-token, 10 GB check passed both fit and zero-swap memory assertions | | Quality and serving | 15 quality probes, 74 API robustness checks and 25 vision-serving checks passed in the original full suite | | Installed-release suite | All 31 checks passed, including concurrent clients, prefix reuse and disconnect recovery | | Full live-governor memory interval | Passed September 11 with zero swap activity; all four earlier invalid intervals remain preserved | | Long-prompt memory interval | Passed September 11 with zero swap activity and exact recall; all three earlier invalid intervals remain preserved | Reruns retained the exact installed identity, original driver hashes, targets, prompts, token limits, full cooldown and acceptance assertions. Each used a bounded readiness check requiring normal pressure, adequate reclaimable memory, no competing model/compiler and a stable zero-swap preflight. No production code or acceptance assertion was weakened. The first governor-only completion protocol was never executed; its superseding attempts and the independent remaining-case checks are preserved with their outcomes. The passed installed-release suite was reused by exact binary identity instead of being repeated. System-wide counters do not identify the process responsible for swap-ins. During the September 10 attempts, Chrome and Wispr Flow were running; neither was closed. Their presence is not proof that either caused a particular event. The normal demo was restored after all test processes stopped. A quiet preflight by itself does not prove that the subsequent interval stayed quiet, which is why the rejected attempts remain invalidated. Exact initial run and prospective correction: [[sources/runs/2026/09/2026-09-10-release-0-2-14-governor-invalid-interval]] and [[sources/runs/2026/09/2026-09-10-release-0-2-14-initial-local-acceptance]]. September 10 targeted attempts: [[sources/runs/2026/09/2026-09-10-release-0-2-14-local-requalification]]. Installed API and normal-service restoration: [[sources/runs/2026/09/2026-09-10-release-0-2-14-installed-api-and-serving]]. ## Final memory acceptance, September 11 Carlos authorized temporarily closing Chrome and Wispr Flow. Both apps and their helpers exited before testing. No Slotstream server was listening at startup. Each model ran alone after the original 30-second stable-swap, normal-pressure readiness check, requiring at least 16 GB reclaimable for the governor and 13 GB for the long prompt. Governor startup observed 30.06 GB reclaimable. The twelve original driver hashes match the frozen protocol and the release tag; only output and frozen-driver paths changed in the wrapper. No workload, cooldown, token limit, memory ceiling or assertion changed. | Final check | Result | Maximum observed process memory | Ceiling | | --- | --- | --- | --- | | Full governor drill | Shrink, full 60-second cooldown and regrowth passed; all three nonempty output-ID sequences identical; zero swap-ins and swap-outs | 10,984,885,632 bytes | 13,000,000,000 bytes | | Original long prompt | 7,972 prompt tokens; four output tokens; completed answer `SEVENTEEN`; both memory and recall gates passed; zero swap-ins and swap-outs | 8,094,469,432 bytes | 10,000,000,000 bytes | Both process-memory intervals contain complete 20 ms sampling. The maximum also considers the lifetime RSS high-water mark and final physical footprint. These are bounded-memory observations, not memory-saving percentages. The aggregate is the initial 20 passing gates, three previously qualified targeted gates, and these two new passes. All are tied to the same published binary and original acceptance drivers; the separate 31-check installed suite is reused by exact identity. Earlier failed attempts remain failed. Closing apps allowed this attempt to qualify, but system-wide counters cannot assign the earlier swap-ins to either app or establish a causal diagnosis. After both model processes exited, normal `slotstream serve` was restored and answered `OK`; both authorized apps were reopened. Exact raw results, frozen driver identity, unchanged assertions and restoration receipts are in [[sources/runs/2026/09/2026-09-11-release-0-2-14-final-memory-acceptance]]. ## Acceptance and performance scope No required local v0.2.14 release-acceptance gate remains open. This conclusion covers the published artifact at `ac7d7e53ca2397bb8840a97f028509fbb0a8f688`. The later adoption of two draft tokens and the prospective Expert Lookahead plan are separate source identities and work. This rerun neither installs nor qualifies those changes; the public v0.2.14 binary remains unchanged. The completed optimization campaign and its original qualified candidate are separate evidence from this new public-artifact acceptance. Established performance scope remains [[records/measurements/optimization-final-composition-2026-09-09]] and [[records/measurements/user-server-throughput-2026-09-10]]. Selected/reference comparisons within that build are not a direct A/B against the prior public release. Sustained decode was flat or slightly slower in the measured profiles; there is no universal TPS improvement claim. Functional rerun durations and the excluded swap intervals do not create new performance claims. ### MTP draft depth at equal total RAM: no new default qualified **No new draft-depth default qualified. Keep depth one as the existing conservative default; these runs do not prove it is optimal.** Depths zero (MTP off), one, two and three were exercised, but the repeated comparison could not obtain enough resource-clean pairs. No production setting was changed. **What was tested.** Installed Slotstream 0.2.14 on the 48 GiB M5 Pro with its internal 2 TB SSD and the local 4-bit Qwen3.8-Flash-Next checkpoint. Every measured response followed a same-prompt 128-token warmup in a fresh server. Each measured response emitted 128 tokens. Greedy sampling, seed 42, thinking disabled, prefill chunk 256, maximum context 32,768, prefix retention off and elastic resizing off were fixed. The four planned setting orders were 0/1/3/2, 1/2/0/3, 2/3/1/0 and 3/0/2/1. Prompt order rotated by round. These are short-context and short-output measurements; the code tutorial in the 24 GB cohort generated mostly introductory prose. The 20 GB cohort replaced it with a prompt that begins directly with Python code. Prose and arithmetic fixtures were retained. No cross-budget or cross-fixture pooling is valid. **Why equal total memory matters.** The plain arm does not load the MTP head and can spend that memory on cached experts. All speculative arms load the same head regardless of depth. Letting plain decoding keep an unnecessary loaded head would conceal this tradeoff. | Total process target | Plain expert slots | MTP expert slots, any tested depth | | --- | ---: | ---: | |24 GB|6,281|5,702| |20 GB|4,834|4,255| Each slot holds 2,764,800 bytes, so the 579-slot difference is 1.6008192 GB of quantized expert-cache capacity. The targets are budget inputs, not claims of measured allocation: accepted measured requests peaked below 22 GB in the 24 GB cohort and below 18 GB in the 20 GB cohort. No explicit pool override or simulated availability was used. **Qualification rule, frozen before timing.** A candidate needed at least three clean same-round pairs against depth one on each of three workloads, an equal-workload geometric mean of median paired throughput ratios of at least 1.03, and no workload median ratio below 0.95. A pair is excluded when either member fails. Throughput is the native emitted-token count divided by native decode seconds; model loading, prompt prefill and HTTP serialization are outside that decode interval. First-request and client timing remain in the raw evidence. Larger acceptance percentages do not substitute for this end-to-end decode criterion. **What completed and why it stopped.** | Cohort | Planned measured responses | Completed measured responses | Individually clean | Clean matched pairs against depth one | | --- | ---: | ---: | ---: | --- | |24 GB|48|11|7|One plain/one-draft pair on arithmetic; none for depths two or three| |20 GB|48|15|5|One three-draft/one-draft pair on prose; none for plain or depth two| The 24 GB preflight refused the next launch with 28.58 GB reclaimable against its 29 GB minimum. A separate 20 GB study was started at a 25 GB preflight after the optional user preference question received no reply; this was an explicitly stated working assumption, not user approval of a new default. The two early 20 GB code baselines both encountered system swap activity. With four fixed rounds, at most two eligible code pairs remained possible, fewer than the required three for any candidate. A later code request also reported non-nominal thermal state. The remaining work was stopped for resource-based inability to qualify. This was an unplanned early stop and the original full matrix was not completed. The prospective 512-output confirmation never ran. **The two surviving matched comparisons are observations, not repeatable rankings.** At 24 GB the arithmetic prompt produced 12.0561 tok/s with MTP off and 13.2185 tok/s at depth one, a 1.09642 ratio in its sole clean pair. At 20 GB prose produced 10.7757 tok/s at depth one and 11.2826 tok/s at depth three, a 1.04704 ratio in its sole clean pair. Neither pair meets the repetition or workload coverage requirement. Other individually clean results cannot be paired with a contaminated sibling to advertise a speedup. All individual rates and all excluded timings are retained in the linked analysis. **What the evidence does establish.** The stock runtime exercised the requested depths; MTP acceptance, verification counts and file-read counters were captured. All 26 completed warmup/measured pairs reproduced identical output-token sequences within their setting. Changing draft depth can change the generated token sequence because the target model evaluates different batch shapes; cross-depth identity is not promised and output differences remain visible. These runs therefore compare generated workloads, not forced identical token paths or model-quality scores. No complete-program, reasoning-quality, sampled-distribution or long-context qualification was performed. A larger draft can reduce target-model traversals per emitted token, but it also runs more draft-head steps, verifies more positions, performs rollback/reconciliation, and may read experts for rejected positions. The resident head also reduces expert-cache capacity at a fixed total budget. The balance depends on acceptance, verification batch cost, routing, cache size and prompt. This is why a depth recommendation from a fully resident GPU server cannot establish the optimum of this SSD-offloaded Mac configuration. The counters record requested expert-file bytes, not measured physical NAND traffic. **Resource and instrumentation limits.** Swap counters are global to the Mac; these data cannot attribute an increment to the model or another app. Four 24 GB and nine 20 GB measured responses had swap activity and were excluded. One additional 20 GB response failed the thermal gate. No new swap-out count was observed across either cohort; existing swap occupancy was not erased. Thermal and low-power observations are OS policy readings, not direct GPU clock or temperature measurements. The watcher sampled resource conditions every two seconds, and the stock generator sampled physical footprint every 20 ms. Clean readings do not prove complete host isolation or zero timing noise. The common sampler's overhead was not independently calibrated here. GPU/CPU profiling instrumentation was not enabled or added. **Next qualifying test.** Use a new frozen cohort after the Mac is quiet, with nominal thermal state, stable swap counters and the declared memory headroom. Repeat all four depths in balanced order, keep the actual-code fixture and same total RAM, then confirm any eligible improvement with the prospectively described longer output. Preserve these attempts and do not replace their failed rows. Until then, the correct statement is that one remains the default and the optimum is unproven. **Cleanup and evidence.** Every launched test process exited, the native model lock was free, no competing model/build job remained in the scoped inventory, and the installed executable and Metal hashes matched their pre-test identities. The demo server was not restarted. Raw evidence, reproducible analysis, protocol limits and the stop reason are in [[sources/runs/2026/09/2026-09-11-mtp-depth-fixed-total]]. The earlier depth study remains historical evidence in [[records/measurements/the-rebuild-eliminated-and-the-numbers-that-ship-2026-09-02]]; this incomplete study neither supersedes it nor changes any published speed claim. **Subsequent operating decision, September 11:** After this measurement was recorded, Carlos explicitly adopted two drafts as the default and requested implementation alignment. See [[records/decisions/draft-depth-defaults-to-two]]. Statements above about the unchanged one-draft default describe the state at measurement closure; they are not the current operating policy. The adoption does not change these results, exclusions or failed performance-qualification conditions. ### Complete MTP depth comparison with automatic 40% RAM **For mixed use on this Mac, two draft tokens are a reasonable practical setting at the tested automatic 40% RAM policy. Two and three were effectively tied in the clean repeated comparison: three did slightly better on code and worse on arithmetic. This is a scoped recommendation, not proof of a universal optimum or qualification to change the shipped one-draft default.** All 54 planned measured responses completed; the longer comparison did not obtain a clean matched pair. No production setting changed and no model was left running. **Complete study.** Installed Slotstream 0.2.14, the local 4-bit Qwen3.8-Flash-Next checkpoint, 48 GiB M5 Pro and internal 2 TB SSD. The user explicitly requested testing while continuing ordinary work. Their applications remained open. Each cell launched one fresh server, completed a same-prompt 128-token warmup, settled, generated its measured response and stopped the server. Forty-eight primary responses crossed MTP off and depths one, two and three with prose, actual Python generation and arithmetic explanation, repeated four times. Six further responses compared the top two settings at 512 output tokens, one per setting and workload. All 108 warmup/measured requests completed; all 48 short responses emitted 128 tokens and all six longer responses emitted 512. The depth orders were 0/1/3/2, 1/2/0/3, 2/3/1/0 and 3/0/2/1; workload order rotated. Temperature was zero, seed 42, thinking off, top-p 1, top-k 0, min-p 0 and presence penalty 0. These are short-context greedy performance fixtures, not reasoning-mode, sampled-distribution, completed-program correctness or model-quality qualification. The configured context ceiling was 32,768, but these prompts contained only 68, 119 and 130 tokens. That ceiling is not a tested context length. **Memory and serving configuration.** All arms used `--max-ram-percent 40`, automatic expert-cache and prefill sizing, elasticity on and prefix caching on. An explicit `--memory-gb` pins the cache and disables elasticity in the current implementation, so it was deliberately omitted for this multitasking profile. Forty percent of this machine's physical RAM is 20.616 GB in decimal units; that was also the study's measured-footprint ceiling, not a claim that every request allocates that amount. | Setting | Effective expert slots | Automatic prefill chunk | | --- | ---: | ---: | | MTP off | 3,764 | 2,048 | | One, two or three drafts | 3,667 | 1,024 | These values held in every measured response. The plain arm spends much of the freed draft-head budget on a larger prefill reservation. This is a comparison of the actual automatic serving policies, not an isolated fixed-pool plain-versus-MTP experiment. The speculative arms do have the same pool. Every measured request reused its entire prompt, with zero new prefill tokens. This captures warm repeated-prompt decode, not first-request loading or a long conversation. The 512-token responses extend beyond the 128-token warmup output. **What the clean short timings show.** Values below are median native decode tokens/sec, with the eligible count in parentheses. Native decode time excludes model loading and prompt prefill. Client timing and first-visible-text timing are preserved separately. Unequal eligible subsets mean dividing these column medians is not a paired speedup estimate. | Workload | Off | One draft | Two drafts | Three drafts | | --- | ---: | ---: | ---: | ---: | | Prose | 6.70 (4) | 7.88 (4) | 9.48 (2) | 9.39 (4) | | Python code | 7.43 (2) | 9.80 (3) | 11.40 (4) | 12.15 (4) | | Arithmetic explanation | 8.71 (2) | 9.80 (2) | 10.16 (2) | 9.58 (3) | The paired comparison is more useful for choosing a setting. Each ratio uses the same workload and round, and a pair is excluded if either member is ineligible. | Candidate relative to baseline | Prose | Code | Arithmetic | Equal-workload aggregate | | --- | ---: | ---: | ---: | ---: | | Two versus one | +20.4%, 2 pairs | +15.5%, 3 pairs | -3.0%, 1 pair | +10.5% | | Three versus two | +0.2%, 2 pairs | +3.9%, 4 pairs | -4.9%, 2 pairs | -0.3% | | Two versus off | +25.3%, 2 pairs | +56.3%, 2 pairs | +16.5%, 2 pairs | +31.7% | The aggregate is the geometric mean of the three workload median throughput ratios. These are descriptive estimates with small samples, not confidence bounds. Different comparisons use different eligible rows; do not subtract their aggregate percentages to manufacture another comparison. In particular, the direct two-versus-three evidence is essentially a tie, not a decisive overall win. **Finalist selection and longer responses.** The frozen exploratory rule normalized eligible rates within each round/workload block, took each workload's median normalized rate, then used their geometric mean. Settings needed observations in all three workloads. Its order was two, three, one, off, so two and three advanced without using the fallback. The normalized scores are for selection, not measured percentage improvements. They differ from direct paired estimates because the eligible comparison sets differ. | Longer workload | Two drafts, tok/s | Three drafts, tok/s | Usable matched pair? | | --- | ---: | ---: | --- | | Prose | 9.84, excluded for swap | 8.54, excluded for swap | No | | Code | 11.52, excluded for swap | 8.52, excluded for swap | No | | Arithmetic | 10.13, eligible | 8.08, excluded for thermal state | No | The excluded values remain visible as evidence of completed work. They cannot support a two-draft speedup claim. Longer verification is therefore inconclusive, despite all six responses completing. **Why the shipped default remains unchanged.** The prospective default-change rule required at least two clean same-round pairs against depth one for each workload, an equal-workload aggregate gain of at least 5%, no workload median slowdown exceeding 5%, and at least two clean longer pairs with median gain of at least 3%. Only one arithmetic pair against depth one survived, and there was no clean longer pair. Depth one also did not advance to the longer comparison. No default candidate qualified. The practical suggestion to use two for this user's mixed workload is separate from changing a default for all users, memory budgets, prompts and sampling modes. **Why more drafts are not automatically faster.** Each additional guess adds draft-head work and another position to the expensive target-model verification pass. Rejected guesses still consumed verification and expert-read work. In the repeated 128-token code fixture, depth two accepted 82 of 92 drafts and depth three 91 of 108; verification passes fell from 46 to 36. In arithmetic, depth two accepted 72 of 110 and depth three 76 of 156; verification passes fell only from 55 to 52. Code therefore gained more useful work from the deeper batch. Acceptance and verification counts were identical across all four repeats of each setting, as were the generated token sequences. Across depths, target batch shapes can change the output sequence; this was not a forced-identical-token comparison or quality test. File-read counters describe requested bytes, not measured physical NAND traffic. **Stability evidence and its limits.** All 54 measured responses and 54 warmups completed without a crash, watchdog safety stop or memory-pressure cancellation. Sampled OS memory pressure stayed normal. Maximum native sampled physical footprint across both phases was 16.64380632 GB, and minimum watcher-observed reclaimable memory was 7.972814848 GB. There were zero new global swap-outs, but 2,038 swap-in pages across the whole study. Existing swap occupancy was not cleared, and global counters cannot identify which process caused an increment. Twelve short timings and four longer timings were excluded for swap activity. One further longer timing failed the thermal rule. macOS reported both nominal and fair thermal states during the full run; automatic settling pauses allowed two later primary measurements to start nominally. All clean timing intervals passed the declared thermal/power rules. That does not prove fixed GPU clocks, no heating or complete host isolation. Ordinary application activity still produces timing noise. The two-second watcher also queried `/api/version`. There were no metadata request errors; p95 response time was about 1.04 ms in the primary phase and 1.01 ms in the longer phase. Watcher scheduling-lateness p95 was 5.16 ms and 10.08 ms, respectively. These are narrow server/scheduling responsiveness proxies, not a measurement of other applications' UI latency, fan noise or energy use. The stock 20 ms physical-footprint sampler's common overhead was not independently calibrated. **Practical use.** A scoped way to select the suggested configuration is: ```sh SLOTSTREAM_DRAFT_DEPTH=2 slotstream serve --max-ram-percent 40 --mtp on ``` Leave elasticity enabled and avoid an explicit memory/pool override for this profile. Three is a reasonable alternative for code-heavy use, but the clean repeated difference from two is small. The study does not justify larger depths, another machine, sampled decoding or a universal setting. It also does not guarantee stability under arbitrary additional memory/compute load. **Evidence and cleanup.** Raw data, all exclusions, frozen drivers, exact requests, reproducible analysis and cleanup proof are in [[sources/runs/2026/09/2026-09-11-mtp-multitask-auto40]]. All 54 owned model PIDs exited, no model remained, and the native model lock was free. The installed executable and Metal hashes were unchanged and all 150 compiled-identity files still matched. Diagnostics existed only in terminated test-child environments. The demo was not restarted. This complete automatic-memory study remains separate from the earlier fixed-memory attempts in [[records/measurements/mtp-depth-fixed-total-inconclusive-2026-09-11]] and does not replace their evidence or published historical speed claims. **Subsequent operating decision, September 11:** After this measurement was recorded, Carlos explicitly adopted two drafts as the default and requested implementation alignment. See [[records/decisions/draft-depth-defaults-to-two]]. Statements above about the unchanged one-draft default describe the state at measurement closure; they are not the current operating policy. The adoption does not change these results, exclusions or failed performance-qualification conditions. ### Two-draft default implementation and bounded checks **Two draft tokens are now the source default, explicitly adopted by Carlos after the completed depth study.** This change implements [[records/decisions/draft-depth-defaults-to-two]]; it does not change the performance study's original qualification outcome or establish a universal optimum. `Generator.defaultDraftDepth` owns the value shared by the Swift library and CLI help. Missing or invalid `SLOTSTREAM_DRAFT_DEPTH` values fall back to two; valid 1...16 overrides remain supported. Existing programmatic overrides, context bounds and the opt-in adaptive policy remain intact. The MTP activation floor, global automatic RAM share, weights, sampling and rollback algorithms did not change. Existing depth-one diagnostics and frozen historical studies retain their explicit settings. **Verification on the 48 GiB M5 Pro.** A guarded, two-job build passed. The complete T0 suite passed all 33 checks and 25,412 assertion items, including default/fallback and override coverage. All static gates passed, including planner, installer, downloader, generated-document and claim checks. Native CLI help prints two drafts from the shared constant. The bounded serving probe used a fresh process for each arm, a 10 GB target, MTP on, vision off and one 32-output-token greedy response without warmup. With no depth environment variable, the default arm recorded 28 drafted tokens over 14 verification passes, exactly two per pass. With the explicit depth-one override, the second arm recorded 16 drafts over 16 passes. Both completed without memory-pressure cancellation. These demonstrate configuration behavior; their timings are not a performance comparison or full correctness qualification. **The full MTP/vision diagnostic remains incomplete.** Both 12 GB attempts stopped on the native global swap guard. The first passed its initial 48-token greedy determinism and speculation-ran checks before stopping. Its observed swap-in delta was four pages; the retry's was eight. Neither observed a swap-out increase within its native measurement window, but both returned `memory_validated: false` and are discarded for qualification. No guard was relaxed and no failed result was replaced. All 150 compiled source inputs still matched the captured candidate identity at closure. The unpublished candidate SHA256 is `e32e9cd33b569984bc5c15f21b79ae2cb601f99392ffb6ea6e5aedf12f460be7`. The installed public 0.2.14 artifact remains unchanged and retains its former default unless overridden; this task did not publish or install a release. Test-only bench details were confined to child environments. Every owned model process stopped and the native lock was free. Help, README, CLI/engineering documentation, agent instructions, claims and current plan annotations now refer to two. Prospective Expert Lookahead accounting includes the pending token plus two drafts, three position biases and a conservative three-position storage bound, with a fresh P0 freeze required. Historical one-draft multipliers remain labeled as historical one-draft evidence. Raw commands, logs, build source, request/response wires, resource failures and cleanup evidence were captured first in [[sources/runs/2026/09/2026-09-11-draft-depth-two-default]]. The prior performance comparison remains [[records/measurements/mtp-depth-auto40-multitasking-2026-09-11]]. ### Lifetime footprint reporting and fixed-budget audit The current source still lost earlier GPU memory peaks. This audit reproduced the defect, replaced the counter with the native lifetime physical-footprint high-water, and checked fixed-budget planning and allocation. It found no silent 48 GB override failure in the exercised planner and runtime paths. The fix is local and unreleased. ## Confirmed defect and correction The old counter took the maximum of lifetime RSS and current physical footprint. RSS can omit GPU allocations, while current footprint falls after those allocations are freed. A peak could therefore disappear from the report. Against the unchanged baseline production counter, a bounded Metal test observed about 267 MB, freed the buffers, and then reported only about 65 MB as the peak. Seven assertions failed. The corrected counter retained the earlier peak after release and passed the same GPU lifecycle, CPU allocation, concurrent-read and independent-sampling checks. No model was loaded for that reproduction. The implementation reads `task_vm_info.ledger_phys_footprint_peak`, checks that the returned kernel structure covers the field, and rejects unavailable or invalid signed values. The compatibility peak combines the native lifetime peak, lifetime RSS and current footprint. Current usage remains distinct. The optional `lifetimePhysicalFootprintPeakBytes` statistics field preserves decoding of older saved observations. Common generation completion and all Engine early returns now record memory, including cancellation and refusal paths that could previously leave zero values. The CLI distinguishes lifetime and sampled peaks; context-check no longer mislabels the combined process peak as RSS. The memory acceptance gate includes the native lifetime counter when supplied, while retaining request sampling and the unchanged zero-swap requirement. Lifetime peaks include loading and earlier requests, so they cannot be attributed exclusively to the latest request. Historical values have not been retroactively remeasured. The historical automatic-plan 32 GB figure on the public surfaces is now correctly identified as an estimate, and its claim uses a specific needle rather than matching unrelated hardware tiers. ## Fixed-budget checks and limits The public doctor interface passed 304 cases across simulated 32, 48, 64, 96 and 128 GiB Macs, several explicit targets and context sizes, MTP off/on/auto, available-memory refusals, option precedence and malformed/nonfinite inputs. These cases load no model and do not qualify those devices. In the simulated 64 GiB, default-context, MTP-off case, an explicit 48 GB target enlarged the pool from 7,280 to 12,705 slots. It was not silently held at the automatic plan. An explicit expert or pool setting still takes documented precedence, and inadequate available memory causes a refusal rather than an undisclosed smaller fixed cache. Two real local servers at 8.1 and 10 GB confirmed that the runtime plan matched doctor and that the larger plan increased allocated pool bytes by 1,031,270,400 and startup physical footprint by 1,031,012,352. Six requests checked short/long/short lifetime accounting, stable fixed pool sizing and current `/api/ps` usage. Four CLI cases covered normal completion, empty input, exhausted context and preparation refusal. A doctor invocation does not reconfigure an already running server. A memory budget also need not be fully used while context and temporary reservations are idle. One request in the integration sequence had global swap-ins. It is retained as excluded from resource/performance qualification; the startup plan and allocation observations are functional evidence only. These scaled local runs do not establish behavior on an M2 Ultra 64 GiB machine or prove the cause of any uninstrumented customer observation. ## Validation and artifact scope The first frozen candidate passed the 44-group catalogue with 27,397 assertions, all 26 governor cases, the serving robustness suite with 74 checks, and vision serving with 25 checks. Real-model checks covered output equality across budgets, growth/shrink/regrowth, the full elastic governor drill, prefix reuse, sweep behavior, MTP, vision parity and long-context recall. The initial acceptance battery reported 20 passed and five failed; every failure was a strict resource exclusion for global swap-ins, with no target overrun or new swap-outs observed in those intervals. Isolated reruns qualified the full MTP/vision diagnostic, short memory gate and both context-check gates with zero swap. Together with the original valid gates, 24 of the 25 acceptance gates qualify on the same accounting and engine implementation. The 7,972-token recall repeatedly answered SEVENTEEN and stayed below the 10 GB target, but continued to observe four global swap-ins. Its zero-swap memory gate remains unqualified. This audit does not call the complete acceptance battery passed or attribute global paging to a particular process. A final rebuild changed only the human-readable context-check label. `final-source-comparison.json` proves that this is the only compiled-source difference from the candidate used for the broad model battery. The final build passed the full static gates and another isolated 2k context check. The native positive and negative controls, source archives, fixture failures and resource exclusions remain in the linked raw archive. The original counter's failing negative control is expected evidence, not a failed corrected implementation. The final source matches the reconstructed build identity, and the model lock is free. No installed binary, release, tag or remote branch was changed. No inference about an unobserved customer command, server configuration or measurement tool is warranted. ### v0.2.15 prepublication qualification **Testing follow-up, September 11:** [[records/measurements/release-0-2-15-open-apps-testing-2026-09-11]] records all 31 additional candidate API checks passing and four retained MTP paging exclusions with every user application left running. Chrome and Wispr Flow have been reopened; the earlier pause request is superseded by the instruction to leave other work alone. Publication and installation remain pending. The earlier checkpoint below is preserved as history. **v0.2.15 is prepared and pushed, but not published or installed. Complete main CI passed. Of the 25 original model gates, 24 now qualify; the combined MTP/vision diagnostic still lacks a zero-swap memory interval.** Release source: `ee4d1af5b3d63c2b5670c814b40b25432415eb46`. [Main CI 34592671081](https://github.com/carloslfu/slotstream/actions/runs/34592671081) passed every coverage, weights-free and public-library job. The exact downloaded CI archive passed source and identity verification. Archive SHA-256: `4f28e283daadcde7020789718e94f757190625c88297c78952d365e0c5454af0`; binary SHA-256: `31eefbbb4791beddb0f8674ab1c1875c2eb1c4a034f5cdd0fa1abcba31e373cf`. Its 150 compiled inputs match the tagged-version preparation source. There is no v0.2.15 tag yet. ## Changes and atomic commits The release candidate includes the adopted two-draft default, compatibility-preserving lifetime physical-footprint peak reports, native memory-lifecycle and fixed-budget regressions, corrected memory documentation, and the reviewed Expert Lookahead execution plan. Expert prediction and prefetching remain planned work. Existing depth overrides, MTP activation and automatic RAM policies remain intact; there is no new universal throughput claim. The default/alignment commits are `0b600d8` and `7ea0941`. Final plan review is `b7487fa` (`docs(plan): finalize expert lookahead execution gates`); memory implementation, tests and evidence are `546f292` (`fix(memory): retain lifetime GPU footprint peaks in reports and gates`); version/changelog preparation is `ee4d1af` (`chore(release): prepare v0.2.15`). The later qualification record is documentation only and does not change the candidate's compiled inputs. ## Actual acceptance state The complete original battery finished 23 passed and two failed. Long-context recall and memory, context diagnostics, 15 quality probes, all 74 API robustness checks, vision parity and all 25 vision-serving checks passed. The original 7,972-token prompt returned `SEVENTEEN`, completed with four output tokens and peaked at 8,102,153,528 bytes under its 10 GB limit, with zero swap activity. The earlier preversion long-prompt gap also closed in its own identity; the CI binary independently passed, so no identity substitution is needed. The governor and MTP/vision intervals initially failed strict system-swap guards. The first unchanged targeted governor retry also recorded four swap-ins. The second passed full shrink, the original 60-second cooldown and regrowth with all three nonempty output-ID arrays identical, complete memory observations and zero swap. Maximum sampled footprint was 10,386,689,384 bytes under the 13 GB ceiling. The first targeted MTP/vision rerun passed text/vision determinism, speculation-ran and rollback checks but recorded 12 swap-ins. The second stopped early after four swap-ins. Both observed peaks were below the 12 GB target and neither native interval recorded new swap-outs, but both correctly reported `memory_validated: false`. They remain failed resource intervals, not passing measurements. The original battery's MTP interval also recorded 32 swap-outs. All exclusions are preserved. No workload, target, cooldown or acceptance assertion was relaxed. Readiness checked real headroom, normal pressure, nominal thermal state and 120 stable seconds of swap counters; that preflight does not guarantee a quiet later interval. ## Closure and remaining work Every owned model process stopped and the native model lock is free. The installed public executable remains v0.2.14. No tag, publication, public-download attestation or installed-v0.2.15 API test has occurred. Test-only detail settings were confined to child environments. Carlos authorized temporarily closing Chrome and Wispr Flow and reopening them afterward. Both apps were already closed when inspected. At restoration time the computer-use tool reported that the Mac was locked and could not unlock automatically, so reopening remains pending a manual unlock. Separate local verification work was active in a VM; this establishes concurrent workload, not the cause of any particular swap event. Permission to pause that other task temporarily is pending and was not inferred from the Chrome/Wispr authorization. Next: obtain a clean original MTP/vision interval without disrupting another task, tag the exact already-passing CI commit, let publication reuse that archive, verify the public checksum/source/attestation, install those bytes, run the separate installed-release suite, stop the model and reopen the authorized apps. Do not call this candidate released or fully accepted until those steps have actually completed. Exact raw evidence and failed intervals: [[sources/runs/2026/09/2026-09-11-release-0-2-15-prepublication]]. Related implementation: [[records/measurements/lifetime-footprint-reporting-and-fixed-budgets-2026-09-11]] and [[records/measurements/draft-depth-two-default-2026-09-11]]. ### v0.2.15 testing with apps open: API passes, strict MTP resource gate remains unqualified **Testing with applications left running has finished its bounded runs: all 31 candidate API checks passed. The original model suite remains 24 of 25 gates qualified because all four new full MTP/vision attempts were rejected by the strict system-wide swap guard. Full model acceptance, publication and installation are not complete.** Carlos explicitly required that no applications or other work be closed. This phase obeyed that instruction. The earlier proposed pause of another task was never performed and is superseded by this constraint. Chrome and Wispr Flow were observed running at closure. Every owned model and test server is stopped; the native lock is free. The candidate is still source commit `ee4d1af5b3d63c2b5670c814b40b25432415eb46`, binary SHA-256 `31eefbbb4791beddb0f8674ab1c1875c2eb1c4a034f5cdd0fa1abcba31e373cf`. All 150 compiled inputs match; the successful main CI and earlier 24 model passes remain the evidence in [[records/measurements/release-0-2-15-prepublication-2026-09-11]]. No runtime source, model weight, inference setting or native acceptance assertion was changed to obtain these results. ## Additional API qualification The complete 31-check end-to-end script passed against the exact uninstalled CI candidate at a 10 GB target, with MTP on and vision off. This covers API/CLI compatibility, streaming, Unicode, long prompts, malformed/hostile input, seeded sampling, prefix reuse, four concurrent clients and disconnect recovery. Its default-depth probe recorded 30 drafts over 15 verification passes and populated the native lifetime footprint statistic. No depth environment override was present. This confirms configuration behavior and ordinary functional operation; it is not a speed comparison or a zero-swap measurement. It does not claim that the new binary was installed. ## Remaining strict memory gate | Attempt | Maximum observed process bytes | Native swap-ins | Native swap-outs | Outcome | | --- | ---: | ---: | ---: | --- | | 1 | 10,422,621,560 | 4 | 0 | Resource interval rejected | | 2 | 10,362,000,712 | 12 | 0 | Resource interval rejected | | 3 | 8,507,675,112 | 4 | 0 | Guard stopped early | | 4 | 10,372,159,392 | 4 | 0 | Resource interval rejected | All are below the original 12,000,000,000-byte ceiling, but every attempt returned exit 1 and `memory_validated: false`. The last reached passing text/vision determinism, speculation-ran and numerical recording/rollback checks; its global paging guard then prevented complete cross-request MTP state qualification. This is not all 25 gates passed, and partial functional successes do not substitute for the missing full gate. The error specifically includes global paging, and the counters prove swap-ins; they do not identify which process caused them or establish a runtime allocation defect. The optional two-minute stable-swap preflight repeatedly restarted even without a model running. The API suite used that idle time on one server, and the waiting wrapper detected it and did not launch a second model. After that server stopped, the task's idle preflight wrapper was deliberately interrupted and the remaining attempts used the original real-headroom preflight. This removed an extra waiting policy, not a native acceptance check. The workload, 48-token generations, 12 GB ceiling and strict native zero-swap assertions remained unchanged. All unsuccessful intervals and the wrapper lifecycle are preserved. ## Installed and release state The installed public executable remains v0.2.14. No v0.2.15 tag, publication, public-artifact installation or installed-v0.2.15 verification occurred. There is no background retry loop or newly running server. Bench details were confined to the test child's environment; production configuration was not changed. Further full acceptance requires a valid original MTP interval; this record does not authorize relaxing it or closing other work. Failed native attempts: [[sources/runs/2026/09/2026-09-11-release-0-2-15-open-apps-mtp-excluded]]. Passing API evidence: [[sources/runs/2026/09/2026-09-11-release-0-2-15-candidate-api-passed]]. ### Global paging policy: native checks with apps open **The full MTP/vision check now finishes and passes with host paging present.** Governor shrink/cooldown/regrowth and the context check also pass. The full static suite passes. This validates the corrected functional gate on a local source build; the updated CI artifact and release still require their own acceptance. Carlos explicitly removed the zero-global-swap acceptance rule. [[records/decisions/global-paging-is-diagnostic]] now separates process-budget and numerical acceptance from timing eligibility. Global counters remain visible; actual headroom, OS pressure cancellation, process ceilings, model exclusion and completed output remain enforced. Native MTP/governor receipts also enforce the kernel lifetime footprint peak, including freed GPU allocations. | Complete local diagnostic | Maximum observed process bytes | Ceiling | Native swap-ins / swap-outs | Result | |---|---:|---:|---:|---| | Original MTP and vision, including cross-request state reuse | 10,349,041,736 | 12 GB | 600 / 0 | Pass | | Full governor shrink, cooldown and regrowth | 11,000,302,952 | 13 GB | 32 / 0 | Pass | | Context check, 2,048 prompt tokens | 8,519,322,864 | 10 GB | 0 / 0 during generation | Pass | The governor preserved identical nonempty token IDs across all three generations. MTP kept its original determinism, vision execution, accept sanity, recording, rollback and reused-state logit criteria. The context request completed inside its plan. Regression fixtures additionally verify that increases in either global counter do not fail functional acceptance, while exceeded physical peaks, missing process-memory evidence, actual pressure cancellations and incomplete deliveries still fail. The MTP outer launch interval also recorded paging before the native receipt boundary: 604 swap-ins and 580 swap-outs in total. This does not attribute paging to any application and does not qualify clean timing. All applications and unrelated workloads stayed open; every native test process exited. Production generation settings and instrumentation defaults are unchanged. Source implementation commit: `0f7aae1`; documentation alignment: `48d11f2`. All 150 compiled inputs are bound in the local build identity and reconstructible archive. Historical excluded runs in [[records/measurements/release-0-2-15-open-apps-testing-2026-09-11]] retain their original results; they are not regraded. Full evidence: [[sources/runs/2026/09/2026-09-11-global-paging-policy-native-pass]]. ### v0.2.15 published, installed and accepted **v0.2.15 is published, installed and functionally accepted.** The exact CI artifact passed all twenty-five model gates, the published archive verified against that artifact with a valid attestation, the installer replaced 0.2.14 on this machine, and the installed binary passed all thirty-one end-to-end release checks. Every user application and unrelated workload stayed open throughout. Release: [v0.2.15](https://github.com/carloslfu/slotstream/releases/tag/v0.2.15), tagged on `48d11f288237e9b697264621297890eead7ffb0a`, published 2026-09-11T17:42:30Z. Archive SHA-256 `d9ea8246d7868620a0c9e0766f7ba637522739d66da8e5fa5d44b2190fc6d4fa`; binary SHA-256 `8abb02b639285335b4fc3113819ebc6fe084bbf017fab2b4a52de3e471a551e3`. The CI candidate, the re-downloaded public archive and the installed binary are byte-identical. ## What qualified | Phase | Result | |---|---| | Main CI 34626184507, commit `48d11f2` | Coverage, weights-free and public-library jobs all succeeded | | Model acceptance on the downloaded CI binary | 25 of 25 gates, 0 failures, 1,103.77 seconds | | Release workflow 34629147966 | Succeeded; published archive matches the CI artifact | | Provenance | `gh attestation verify` confirmed the sigstore bundle for `refs/tags/v0.2.15` | | Installation | 0.2.14 replaced by 0.2.15, exit code 0, digest re-checked | | Installed-release end-to-end | 31 of 31 checks, 0 failures | The twenty-five gates are the sixteen numbered checks plus weights provenance, planner gates, sampler and governor gates, the full elastic drill, speculative decode gates, the behavioural quality probe, the serving robustness suite, vision tower parity and the vision serving suite. The 150 compiled inputs and the eight frozen drivers were hashed before and after and were unchanged. ## Why this reaches twenty-five when the candidate stopped at twenty-four The combined MTP and vision gate had previously been cancelled four times by a rule that failed acceptance whenever any process on the whole Mac touched swap. That rule was removed by [[records/decisions/global-paging-is-diagnostic]]. This acceptance was therefore decided by the original numerical and work assertions, the actual process-footprint ceilings, real headroom and OS pressure handling, with global paging recorded separately as a diagnostic. Do not read the twenty-fifth gate as the earlier zero-swap condition being satisfied. It was not. The condition was retired as a test-design defect, and the preserved exclusions in [[records/measurements/release-0-2-15-open-apps-testing-2026-09-11]] stay as they were recorded. The acceptance interval observed 271 system-wide swap-ins and zero swap-outs, with reclaimable memory rising from 28,482,207,744 to 30,258,495,488 bytes. ## The shipped default, observed on the public binary The installed-release probe ran with no depth override and drafted 30 tokens over 15 verification passes, exactly two per pass, accepting 16. This is the two-draft default from [[records/decisions/draft-depth-defaults-to-two]] arriving through a published artifact rather than a local build. Its lifetime physical-footprint peak was 7,203,164,480 bytes against a 10 GB target, and that generator interval observed no paging in either direction. ## Limits No clean-timing or throughput qualification is claimed here. The acceptance interval deliberately shared the machine with ordinary work, which is what functional acceptance is now allowed to do and what a speed measurement still is not. The depth-two choice remains a mixed-workload preference, not a universal improvement. Every commit that has landed on `main` after the released build changes documentation, projections and store records only. None touches `Sources/`, `Package.swift`, `Package.resolved` or the `Makefile`, so the published artifact remains the exact build of `48d11f2`. Expert lookahead and expert prefetching are reviewed, planned and unimplemented. This release ships the plan, not that acceleration. At closure no Slotstream process remained, the native model lock was free, no application was closed or paused, and no persistent production instrumentation or environment setting changed. ### Expert Lookahead pilot: exact causal capture and bounded prefetch, but small predictors project no throughput gain **Outcome: a justified stop at the plan's offline continuation rule, with the whole pipeline built and proven exact.** The Expert Lookahead plan ([[records/plan/expert-lookahead-local-experiment-2026-09-10]]) was executed from P0 through P3a on the 48 GB M5 Pro, and the P2 prefetch runtime was built and checked for correctness. The pilot's calibrated replay shows that a perfect forecast under the bounded scheduler would project a 1.59x decode throughput ratio, but no trained or cheap predictor came close to the precision that requires: the best policy projects 1.0055x under the plan's 1.20x read-traffic bound and 1.064x with the bound removed at 4.9x file traffic. The rule therefore stops the experiment before native speed qualification (P3b, P4, P5). No throughput improvement is claimed; the measured improvement over the current deployment is none. **What was frozen.** Protocol `xla-pilot-20260911`: the pinned two-draft MTP text mode (`SLOTSTREAM_DRAFT_DEPTH=2`, adaptive speculation and draft-tail shortening off, prefill chunk 256, prefix retention off, elastic off), greedy seed 42, thinking off, maximum context 32,768, `--mtp on`. The 24 GB primary profile needed 29 GB reclaimable and the machine offered 28.9 GB at freeze, so the plan's separately frozen 20 GB profile with a 25 GB preflight ran instead: 4,255 global slots, about 89 experts per layer, which is the memory regime where demand reads take 37% of decode time. A public 307-request, 95-family corpus with a seed-1729 family split (train 242, validation 32, test 33 requests) was frozen with prompt hashes; test families were never read. **Capture (P1).** The collector writes one binary shard per request with lossless little-endian BF16 feature bits, exact ordered top-10 router IDs per layer and pass, three-position start features, full 48-layer `x2` for verification passes, demand events with hit, miss, victim and adoption lists and timings, sweep admissions, reconciliation labels and a post-prefill CLOCK residency snapshot. Collector on versus off produced identical output IDs, router digests and finish reasons on six correctness requests (1,218 outputs) at 8.8% wall overhead, of which the collector's own write and copy time was under one percent. The pilot captured 69 requests, 12,899 outputs, 5,205 verification passes and 4.40 GB of shards at a median 10.98 tok/s and 0.687 hit rate; the one request the engine's memory-pressure guard cancelled completed on the resumed run and its failed row is preserved. Data validation checked framing, geometry, finiteness, joins, pass completeness and recomputed every request's router digest from the shard. A CLOCK replay from the residency snapshot reproduced all 249,840 native demand events exactly, miss lists and victim slots included, so the offline scheduler evaluation stands on a calibrated cache model. **Training and offline evaluation (P3a).** On 12,540 training and 3,075 validation positions with the frozen recipe, the whole-pass models learned nothing beyond the per-layer frequency prior: G64 and G128 stopped below the prior's recall@16 of 0.121, and at the plan's 1e-3 nonconvergence rate G64, G128 and G256 all converged to the identical prior solution (recall@16 0.121, validation loss 0.6325). The completed-layer refiner L128 did learn: recall@16 0.193 at horizon one and 0.194 at horizon four after 20 epochs, rising from 0.162 on a nested 28-request prefix, so doubling the data bought about three points. In the replay with the native scheduler settings (32-record cap, 8 lanes, 4-layer window, service time calibrated from recorded demand events), timely miss-byte coverage under the 1.20x traffic bound was 1.39% for L128 (precision 7.8%), 1.48% for G128 with L128 refinement, 0.54% for G128, 0.46% for G64, and effectively zero for the recent-routes and frequency baselines; the corresponding projected ratios are 1.005x at best against the oracle's 1.587x. Precision is the binding constraint: with 7 cold misses per layer per pass out of about 22 routed experts, a top-4 candidate list that is right one time in thirteen cannot cover misses without multiplying read traffic. **Runtime (P2) and native correctness.** The raw-staging prefetch runtime exists and is exact: tickets with owned aligned buffers read through the store's checked seam, one shared lane budget with demand priority, a 32-record (88,473,600-byte) cap charged until the scatter graph releases the bytes, expiry at each completed layer, promotion of in-flight reads on demand, zero-copy adoption into the slot the ordinary victim scan chose, and a 128 MiB reserve the planner deducts before solving capacity. Weights-free checks cover the lane budget, ticket lifecycle with cancellation before, during and after a read, an injected fault, the EINTR seam, the byte cap, shadow mode, the recent-routes policy and adoption lifetime; the T0 catalogue passes 36 checks with 25,487 assertions. Natively, six correctness requests with prefetch on (G128 pack, top 32, 4,206 slots after the reserve) matched the baseline exactly while adopting 11,161 records (4,442 by promotion) out of 430,251 issued tickets with zero failures and the cap never exceeded; the exported G128 pack agreed with Python to a maximum score error of 0.00026 with zero top-16 set mismatches over 4,608 layers at 0.45 ms per row on the GPU; a request cancelled mid-decode with 7,947 tickets in flight was followed by an exact recovery request. Shadow forecasting cost 1.3 ms per verification pass, about half a percent of decode. **What the native runs also show.** With this predictor at top 32 the prefetch arm decoded at 5.2 to 6.8 tok/s against 10.1 to 13.0 tok/s without it: wrong forecasts consume read bandwidth and force promotion joins. Two single diagnostic pairs on one code prompt (7.41 against 8.40 tok/s with prefetch on, one control arm swap-contaminated; 12.95 against 12.53 tok/s in shadow mode) are mechanism diagnostics only, because swap-in counters never held still for the plan's 120-second readiness interval. No pair in this record is speed evidence. **Limits.** One seed, one memory profile, cold per-request caches, a 10k-output pilot; the sealed test families were never used; no native benchmark cohort ran. The oracle ceiling is an upper bound under an idealized service model and says nothing about how well any forecast can be learned. Every timing interval in this record is a diagnostic under [[records/decisions/global-paging-is-diagnostic]]. **Next decision.** Do not collect the 50k set or run P3b/P5 for this feature contract: the start-feature models are prior-equivalent and the completed-layer refiner's learning curve does not reach usable precision within the plan's budget. The evidence points at the feature contract rather than the mechanism: a same-layer pre-attention forecast (predicting layer L's experts from its own `x1`, which the router later reads through `x2`) would trade lead time for precision, and demand service was measured at 0.27 ms per record, so even a one-layer lead could be timely. That is a new protocol under the plan's extension table, not a continuation of this one. Nothing was installed, activated, published or committed; the installed 0.2.15 artifact is unchanged. Raw commands, identities, per-request counters, training logs and every excluded row are in [[sources/runs/2026/09/2026-09-11-expert-lookahead-pilot-offline-stop]]. ### Expert Lookahead probes: the model's own routers forecast routing one layer early at 60%, and optimal replacement would halve misses at the same memory **Outcome: the model's own routers, applied one layer early with no training, forecast routing four times better than the trained predictors of the pilot, and the slot cache has a 15-point hit-rate gap to Belady's optimum at the same memory.** These two offline facts, measured on the pilot's validation shards with the plan's exact twin, reopen the Expert Lookahead question that [[records/measurements/expert-lookahead-pilot-offline-stop-2026-09-11]] closed, and they are the evidence for the successor plan [[records/plan/expert-lookahead-2-replacement-router-reuse-memo-2026-09-11]]. No native run, no throughput claim and no code change to the engine are part of this record. **Router reuse (zero parameters).** Applying layer T's real router weights to the captured MoE input of layer T minus k, on 1,025 validation passes, reproduces T's exact top-10 experts in 59.5% of positions one layer early (recall@16 72.2%), 50.0% two layers early, 45.3% and 43.2% at three and four; a stride-zero self-check reproduces the real routing at 99.9996%, so the probe is faithful to the engine's router. Agreement is nearly uniform across depth (0.55 to 0.66 by eight-layer group) and lowest at layers 1 to 3 (0.43). For comparison the pilot's trained completed-layer refiner reached recall@16 0.193 and the frequency prior 0.121. Under the plan's bounded scheduler twin with exact replayed residency, a stride-1 top-10 forecast covers 48.8% of misses in time at 40.1% precision among issued non-resident candidates and 1.73x file traffic, which projects 1.23x by the previous plan's proportional formula and 1.17x by the read-cost model below; stride 2 top-10 covers 39.8% at 30.7% precision (1.90x traffic, 1.14x). Both are lower bounds on a native forecast, which would use the target layer's own hyper-connection mixing on the live streams and omit only that layer's attention sublayer, whereas the probe reuses the source layer's mixed input. **Replacement (Belady bound).** Replaying the same decode demand stream from each request's post-prefill CLOCK snapshot at the native 4,255 slots: CLOCK 69.3% hits, exact LRU 70.5%, segmented LRU 71.7%, decayed LFU 67.3%, and Belady's optimal policy 84.8%, which halves the misses (167,620 against 338,147). The optimal replay was verified against a brute-force reference on 200 random traces. Two thirds of misses (66.6%) are re-references of experts that were resident earlier in the same request. At 1.2x slots LRU removes 22% of misses and optimal 57%. This bound is dynamic and differs from the static hot-set bound of [[records/measurements/m1-expert-locality-on-a-real-trace-2026-09-03]] (4.6 points at 30 experts per layer); it does not by itself reverse [[records/decisions/clock-stays-the-eviction-policy]], whose condition is an implementable policy beating CLOCK by more than about one point on a second workload, and LRU is again within about one point of CLOCK here. **Structure.** A layer's demand read costs 0.61 ms for one missing record and about 0.21 ms per additional record (weighted fit 0.38 ms plus 0.21 ms per record); 97.3% of layer events have at least one miss; the lead from a completed layer to the next layer's demand is 1.16 ms (median), which is that layer's attention sublayer, and 5.5 ms two layers ahead. Layers 0 to 7 carry 22.2% of misses at 9.2 per event and have the weakest temporal locality (previous-position overlap 0.21 against 0.34 over all layers), but a repeated token routes to the same layer 0 to 3 experts 59.6% of the time, and 54% of positions repeat a token already seen in the request. Three positions per pass route to 22.4 unique experts per layer, of which 6.9 miss. **Limits.** One machine, one memory profile (20 GB, 4,255 slots), 13 validation requests, cold per-request caches, no native timing. The twin's projections ignore bandwidth contention from wasted reads and were contradicted natively once in the pilot at 2.6% precision; the ratios above are projections under an idealized service model. The twin counts a correct but late ticket as wasted and its record as a demand read, whereas the native runtime promotes and joins an in-flight ticket without a second read, so the traffic figures are upper bounds. The exact-residency accounting takes each pass's resident set once at pass start and ignores evictions by earlier layers within the pass, which slightly undercounts both coverage and waste. The recency policies (LRU, segmented LRU, decayed LFU) start each request from the snapshot's slot order, which carries no recency, while CLOCK replays its exact reference bits; the comparison is therefore mildly biased against the recency policies at the start of each request. The read-cost fit uses per-event medians, not a queueing model. The router probe's stride-2 forecast lacks all of the intermediate layer, not only the target's attention. **Next decision.** Proposed, not approved: the successor plan sequences an exact offline replacement-policy lab against the Belady bound, a native router-reuse forecast diagnostic and prefetch with admission tuned on the twin, and an early-layer token memo, each with its own gate and the previous plan's held-out B0 cohort as the final gate. The learned whole-pass predictor and the 50k capture are retired. Raw commands, hashes and per-request tables are in [[sources/runs/2026/09/2026-09-11-expert-lookahead-router-reuse-and-replacement-probes]]. ### Expert Lookahead 2: the model's own routers forecast 76% of the next layer's routing and the exact prefetch removes 63% of demand reads, yet every native screen loses throughput; replacement policy is not the lever **Outcome: no throughput gain; a complete, reproducible negative with the mechanism proven exact.** Expert Lookahead 2 ([[records/plan/expert-lookahead-2-replacement-router-reuse-memo-2026-09-11]]) was executed on the 48 GB M5 Pro from W0 through W4 with its follow-up screens. Its three ideas came out as follows. Replacement policy: in an exact replay of the pilot corpus, no implementable policy reaches the plan's port threshold (segmented LRU removes 7.6% of misses against Belady's 50.4%), so CLOCK stays. Router-reuse prefetch: the model's own routers applied one layer early reproduce 76.3% of the next layer's top-10 routing with no training, the native runtime prefetched on that forecast exactly and removed 63% of demand reads, and the calibrated twin projected 1.247x; yet every native screen lost throughput, 0.713x at first and 0.956x at best after a runtime revision and two lower-traffic settings, because the read-cost model credits the saved demand reads, which did materialize on the demand-IO clock exactly as modeled, and charges nothing for the runtime's own work on the critical path: the non-IO part of decode grew by 38% to 111% with prefetch on (joins on in-flight tickets, adoption work, the forecast and an unprofiled remainder on the GPU and memory side). Early-layer memo: its candidates cover 4.2% of misses at precision 0.57, but only 1.49% arrive in time in the twin (projected 1.005x), so it failed its offline gate and never ran natively. The measured improvement over the current deployment is none. **What was frozen.** Protocol `xla2-20260911`, the pilot's 20 GB profile (4,255 slots, 25 GB preflight) with the pilot's pinned two-draft MTP text mode, corpus, correctness, screen and B0 prompt sets, and all 69 pilot shards re-hashed byte for byte. The eligibility rule for timing is the one Carlos approved with the plan: readiness needs thermal nominal, low power off, reclaimable memory at or above the preflight and host swap-outs stable over 30 s; a pair is eligible when both arms are exact, no host swap-outs occur during either arm, the engine's own page-ins stay under 4,096 pages and thermal stays nominal; the old 120 s swap-stability verdict is recorded beside it. Three candidate binaries were frozen in sequence with their hashes and changes; each passed the T0 catalogue including the new deterministic scheduler check, which also caught a real issue-cursor defect before any timing screen. **Replacement policy (W1, offline).** Eight policies replayed with `ensureCore`'s batch-pin semantics on the 69-request demand stream at 4,255 slots, constants fit on 56 training requests and scored on 13 validation requests with the read-cost model fit on recorded demand events (0.38 ms per layer event with a miss plus 0.21 ms per record, within 0.2% of measured IO). Belady projects 1.188x (misses 0.496x native). Segmented LRU is the best implementable policy at 0.924x misses and a projected 1.023x; S3-FIFO 0.943x, ARC 0.948x, exact LRU 0.962x, W-TinyLFU 0.980x; this session's CLOCK-Pro 1.47x (an implementation limit, recorded as such). Layer-aware variants change the third decimal. A learned reuse-distance evictor trained on 6.3 million eviction candidates reaches 0.966x misses when choosing among 32 candidates at the hand. The port threshold was 0.85x, so replacement is not the lever at this size, and the CLOCK decision stands with a tightened reversal condition ([[records/decisions/clock-stays-the-eviction-policy]]). **Forecast (W2, native).** The runtime computes, at each layer boundary of a verification pass, the next layer's own router on the live streams (mix-only hyper-connection read after the previous layer's MoE add), keeps the top candidates per row with their margins, and evaluates them in the same sync as the layer. Its self-check reproduced the true routing in 189,936 of 189,936 rows and an offline recomputation from captured BF16 inputs matched 72,714 of 72,720 rows with six ties at the boundary. One layer early the forecast agrees with the real top-10 in 76.3% of positions (recall@16 0.882; the offline probe had 59.5%), two layers early 62.8%; the stride-1 lead is 1.04 ms at the median. Outputs with prefetch on, off and in shadow were identical on every request measured (six correctness requests at two profiles, 13 validation requests). The forecast's own cost measured 2.8% of decode in the paired overhead screen (12 of 12 pairs eligible, 0.977x), above the plan's 2% budget after its listed remedies. **Twin projection (W3).** With the native forecasts transplanted pass by pass onto the pilot's exact residency, 270 settings were swept; the selection rule chose strides [1], 10 candidates per row, issue cap 16: projected 1.247x, precision 0.601, timely coverage 0.652, traffic 1.43x, robust to 1.5x service time and to advanced deadlines (2.0x service gives 1.239x with 1.8% late tickets). The best admissible setting projected 1.257x. The plan's W3 gate (at least 1.10x) passed, so the native screens were justified by the protocol. **Native screens (W4).** Twelve-pair screens on the frozen code, prose and reasoning prompts, 128 outputs each after warmup, both arms charging the 128 MiB reserve, prefetch off against router prefetch on: | Screen | Binary | Setting | Eligible pairs | Ratio (aggregate; bootstrap 95%) | Demand records removed | Bytes read, on/off | | --- | --- | --- | --- | --- | --- | --- | | F2 | candidate 2 | stride 1, top 10, issue 16 | 12 of 12 | 0.713x (0.701 to 0.711) | 64.2% | 1.47x | | F2 revised | candidate 3 | same | 11 of 12 | 0.895x (0.845 to 0.894) | 63.4% | 1.48x | | Stride 1, threshold 0.062, issue 8 | candidate 3 | twin 1.175x at 1.17x traffic | 12 of 12 | 0.935x (0.919 to 0.934) | 51.6% | 1.17x | | Stride 2, threshold 0.062, issue 16 | candidate 3 | twin 1.153x at 1.40x traffic | 12 of 12 | 0.956x (0.942 to 0.955) | 48.1% | 1.39x | Candidate 3 added a batched adoption scatter (one pool write per piece per layer event) and lane priority for promoted tickets; it cut the join time per request from 2.5 s to 1.2 s and the adoption time from 2.5 s to 1.2 s on the same setting, which is the whole difference between 0.713x and 0.895x. No screen reached the 1.03x needed to proceed; the B0 cohort, the policy port (W5), the held-out qualification (W6) and the 24 GB cohort (W7) did not run. **Why the projection was wrong.** The twin and the native runtime agree on every count: issued tickets per pass within 2%, adoptions within 4%, misses within 1%. They also agree on the demand-IO clock: the read-cost model (0.38 ms per layer event with a miss plus 0.21 ms per record) reproduces the measured demand-IO time of every arm, control and prefetch alike, within 3 s (stride 2: 29.9 s measured against 32.4 s modeled, from 51.2 s in the control arm), so the SSD side did what the twin said and no traffic penalty is visible there. What the twin does not charge is the runtime's own work on the critical path, and that is where the whole loss sits: the non-IO part of decode grew from 73 s to 154 s on the first screen (+111%), from 66 s to 106 s after the revision (+61%), and from 73 s to 103 s and 101 s on the two thresholded settings (+41%, +38%). At stride 1 the host counters name most of it: 97.6% of adoptions were promotions of tickets still in flight when the demand batch arrived, the demand path waited 13.0 s on them and spent 13.5 s adopting, the forecast cost about 3 s, and 10 s remain unattributed. At stride 2 the joins and adoption enqueue vanished (0.15 s over twelve arms) and the non-IO part still grew 28 s, of which the forecast explains 3.6 s and 24 s are unattributed by any host counter. The candidates for that remainder are the GPU and memory work the adoption path adds per record and the demand path does not pay (a concatenation of each ticket's own buffers before the pool scatter, one array wrapper with a finalizer per piece per ticket, a fresh 2.76 MB allocation and release per ticket, 170,000 tickets per screen of which 45% expired unused), and they have not been profiled. The margin per record is thin: a prefetched hit saves about 0.26 ms of demand read, so any per-record cost above that on the critical path, or wasted reads at precision 0.55 to 0.60, turns the gain into a loss. The mechanism moved the cost from the SSD to the engine's own timeline rather than removing it. The lesson for any successor is that the SSD is not the constraint at this profile (it is idle for most of the decode and served the extra bytes as modeled); the constraint is what the engine does per prefetched record on the critical path, which must cost less than the 0.26 ms of demand read it replaces. **Limits.** One seed, one machine, one memory profile (the 24 GB profile never had its headroom); the forecast's coverage and precision were audited at a 12 GB diagnostic profile because another application's virtual machine blocked the 25 GB preflight for the shadow capture, while every timing arm ran at the frozen 20 GB profile; screens are 12-pair validation screens, not the held-out cohort; the eligibility threshold of 4,096 page-ins remained a placeholder that never decided a verdict (largest delta 6 pages); one pair was excluded for host swap-outs and is preserved; the forecast overhead exceeded its budget by 0.8 points; adoption-path revisions were made inside W4 on the plan's diagnosis order and are recorded as candidates 2 and 3, not as new protocols. No number here is a claim on any public surface. **Next decision.** Do not run the B0 cohort, port a policy or install anything from this line; the prefetch runtime, forecast event, twin and bench remain in the working tree as instruments. If the line is reopened, the first step is not a screen but a profile of one stride-2 arm attributing the 28 s of non-IO growth (GPU timeline, allocation, array churn), then an adoption path that costs what the demand path costs (speculative reads into per-layer contiguous staging so adoption is one scatter with no concatenation, buffers reused instead of allocated per ticket, expired tickets never allocated), then the stride-2 screen again, where joins are absent. The prize is bounded and worth stating: demand reads are 41.5% of decode at this profile, so a prefetch that hides all of them is at most about 1.7x, and the stride-2 setting with the runtime cost reduced to the forecast's 3.6 s would have measured about 1.17x on the same arms. Only after that do precision levers matter (a stride union that issues at stride 2 and cancels on the stride-1 refresh before the read starts) and forecast-protected eviction, which the captured native forecasts can evaluate offline in the twin before any native run. A trained predictor is not the next step: the zero-parameter forecast is already at 0.763 one layer ahead and the pilot's trained models never left the frequency prior. The predecessor's warm-start lever (10% of records serving 71% of accesses) is untouched by this result. Raw commands, identities, per-arm counters, parity outputs and every excluded row are in [[sources/runs/2026/09/2026-09-11-expert-lookahead-2-offline-lab]] and [[sources/runs/2026/09/2026-09-12-expert-lookahead-2-native-screens]]. ### Expert Lookahead 2, slot adoption: speculative reads straight into pool slots make the router-reuse prefetch a 1.14x validation-screen gain; held-out cohort pending a host restart **Outcome: the router-reuse prefetch is a measured 1.138x on the validation screen once adoption costs nothing, and the held-out cohort that would make it a headline is blocked by a host I/O stall that needs a restart.** After the staging screens of [[records/measurements/expert-lookahead-2-router-reuse-prefetch-native-screens-2026-09-12]] closed at a loss, the runtime was revised on that record's own diagnosis: the loss was the engine's per-record work on the critical path, not the SSD. **Slot adoption** reserves the CLOCK victim slot when a speculative ticket is issued and lets the worker write the record into that slot's pool memory, so adopting a prefetched expert is a map insert with no copy, no allocation, no array wrapper and no graph work; unused reservations recycle their slots before any live key is evicted. On the same twelve-pair validation screen, same prompts, same eligibility rule and same read-cost accounting, the setting the twin had projected at 1.153x (stride 2, top 10, issue cap 16, margin threshold 0.062) went from 0.956x with staging adoption to **1.138x with slot adoption** (11 of 12 pairs eligible, bootstrap 1.114 to 1.134, every eligible pair faster; code 1.089x, prose 1.170x, reasoning 1.157x), with outputs exact against the pilot and the shadow captures on the six correctness requests. **What the numbers say about the mechanism.** Over the eleven prefetch arms the demand-IO clock fell from 47.3 s to 28.5 s, within 1.4 s of the read-cost model's credit, and the non-IO part of decode grew only 6.4 s (from 66.2 s), about half of it the forecast itself; with staging adoption the same setting had grown non-IO time by 28 s. Demand records fell 47.3%; total bytes read rose 1.40x; precision was 0.545 (85,070 adoptions from 156,006 tickets, 70,855 expired, 3,532 by promotion). One-round pilots bracket the setting: 32 lanes 1.127x, a higher threshold (0.214) 1.122x, no threshold 1.085x, all with 16 lanes. So at this profile the prefetch is worth about 12% to 14% on validation prompts when it costs nothing per record, and the remaining cost is the forecast (about 3%) and the 45% of reads that expire unused. **What was tried and closed on the way.** Slot adoption alone (candidate 4) saved no time against staging on the six correctness requests because the join path forced the lazy demand scatter to evaluate early and wasted reservations evicted 177k live keys; both were fixed in candidate 5 before any screen. Forecast-protected eviction (CLOCK skipping the keys the next layer's forecast names) was evaluated offline on the exact replay with the native forecasts: 0.9985x misses at best, projected 1.0004x, closed. **The stall.** One pilot arm hung for 42 minutes inside a single `pread` of the prefill sweep's staging read after its warmup request had run slot-mode prefetch; a process sample showed every other thread idle. Killed, the process stayed in the kernel's exiting state holding the per-user model lock, and the next engine start blocked uninterruptibly at startup behind it. Two such processes now sit on this boot; no model process can start until the host restarts. The cause is unproven. Because the one kernel-level novelty of slot adoption was an uncached file read straight into GPU-shared pool memory, candidate 7 reads each piece into a host scratch buffer and copies it into the slot on the worker thread instead; it passes the T0 catalogue (138 checks) and awaits its parity run. The bench now kills an arm after 1,200 s and records the stall, and falls back to a separate lock path only when the default lock is held by a dead process and no live model process exists. **Limits.** A validation screen on the frozen validation prompts is not the held-out cohort, and the plan's B0 gate (aggregate at least 1.10, lower bootstrap bound above 1.00, no family median below 0.95, no family duration regression above 5%, two clean pairs per prompt over 36 pairs and 512 outputs) has not run. The screen binary (candidate 5) wrote speculative records straight into pool memory; the binary that will run the cohort (candidate 7) uses the scratch path and must be shown exact and re-piloted first. One pair was excluded for host swap-outs and is preserved. Both arms charge the 128 MiB reserve that slot adoption no longer needs (46 slots, about 1% of the pool), which slightly favors the control. No number here is a claim on any public surface. **Next decision.** After the restart: `w9/run-after-reboot.sh` runs the T0 checks, the six-request parity of candidate 7 at the screen setting, a one-round pilot, then the B0 cohort against the current deployment. If B0 passes its gate, the mechanism is a candidate for the production path with the engineering the plan lists as separate (acceptance gates, mode coverage, adoption of the controls as defaults); if the stall recurs under candidate 7, the mechanism stops until the kernel interaction is understood. The predecessor decisions stand: CLOCK stays, the memo stays out, no predictor is trained. Raw commands, candidate hashes, per-arm counters, parity outputs, the stall sample and every excluded row: [[sources/runs/2026/09/2026-09-12-expert-lookahead-2-slot-adoption-screen]]. ### Expert Lookahead 2, slot adoption: the held-out B0 cohort passes the plan gate at 1.105x **Outcome: the slot-adoption router-reuse prefetch passes the held-out B0 gate at an aggregate 1.105x.** The validation screen in [[records/measurements/expert-lookahead-2-slot-adoption-screen-2026-09-12]] measured 1.138x on frozen validation prompts; this record is the sealed held-out cohort the plan requires before the mechanism counts as proven. Candidate 7, whose file reads land in host scratch before a memory copy into the reserved slot, ran six families with two prompts each over three rounds at 512 measured outputs: 36 pairs, every one eligible under the `process-pageins-v1` rule, none excluded. **Correction (2026-09-13).** This record first reported an aggregate of 1.120x, a lowest family of 1.093 and medians of 12.24 to 13.59 tok/s. The cohort report took the upper of the two middle values whenever it formed a median over an even count, so each two-prompt family counted its faster prompt. Scored again from the same pairs with true medians, a family being the geometric mean of its prompt medians as the registered bootstrap already resampled, the aggregate is 1.105x. It still clears the 1.10 gate and the verdict is unchanged: [[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]. | gate | required | measured | | --- | --- | --- | | aggregate ratio (geometric mean of families, each the geometric mean of its prompt medians) | at least 1.10 | 1.105 | | lower bootstrap bound (10,000 draws, seed 1729) | above 1.00 | 1.090 | | lowest family | at least 0.95 | 1.065 (structured) | | family request duration | no regression above 5% | every family shorter, 0.885 to 0.956 | | clean pairs per prompt | at least 2 | 3 for all twelve prompts | Every family is faster: dialogue 1.142, prose 1.130, multilingual 1.126, reasoning 1.086, code 1.083, structured 1.065. Median decode throughput rose from 12.20 to 13.50 tok/s. **Mechanism.** Demand records fell 47.0% (median 59,390 to 31,462) with the expert hit rate unchanged (0.682 and 0.680), so the gain is reads moved off the critical path, not a larger cache. Precision was 0.557: 843,060 of 1,512,496 speculative records were adopted and 668,920 expired unused. Counting every issued record at full size, speculative reads add at most about 116 GB per request to 87.0 GB of demand reads, an upper bound near 1.24 times the control's 164.2 GB, for a 1.1 times throughput gain. Draft acceptance on the prefetch arm averaged 0.741. **What it settles.** The plan's initial proof passes for this setting: stride 2, top 10, issue cap 16, margin threshold 0.062, 16 lanes, slot cap 64. The host I/O stall that blocked the cohort before the restart did not recur on candidate 7 across the parity run, the pilot and all 72 cohort arms. Production adoption remains separate engineering: acceptance gates, mode coverage and making the controls a default. **Limits.** One machine at one profile and a single frozen setting. The earlier swap-stable rule would have admitted 28 of the 36 pairs; the headline uses the approved page-in rule as the protocol specifies. No number here is a public claim. Raw report fields, counters, commands and artifact hashes: [[sources/runs/2026/09/2026-09-12-expert-lookahead-2-b0-cohort]]. Rescoring: [[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]. ### Decode path serialization, round 1: deferring the per-layer GPU drain buys 3.8% to 5.1%; host time is waiting, not compute **Outcome: removing one of the two host-blocking synchronizations in every layer buys 3.8% to 5.1% with exact outputs, and the host's share of decode turns out to be waiting, which closes compilation and host point fixes.** Every MoE layer reads its routing back to the host and every layer ends in a full evaluation: 96 synchronizations per forward pass, so expert I/O, GPU work and host work run strictly in turn ([[records/measurements/decode-wall-time-attribution-2026-09-10]] partitions them 37.2, 30.1 and 32.7 percent, summing to 100). The end-of-layer drain exists so the next layer's ensure cannot scatter into a slot an unevaluated gather still reads. Pins already exclude a slot from every victim scan; they were retired one layer too early. With pins held for K + 1 generations and the drain taken every K layers, prefetch off, paired ratios against the original path are 1.025 at K=2, 1.038 at K=3, 1.040 at K=4, 1.047 at K=6 and 1.051 at K=8. Every paired comparison is above 1, demand records stay within 0.3% and tokens per verify pass are unchanged: the gain is overlap, not a different workload. Exactness held at K=4 on the six correctness requests and at every K in the sweep. **What closed.** - Whole-record speculative reads: 1.002 (0.993 to 1.013). Nine reads per record were real, but speculative reads run on worker threads off the critical path, and the demand path already read whole records. Four lanes cost 8.3% (0.917); 8, 16 and 24 are flat. - Nine implemented but disabled host accelerators, re-tested under slot adoption: paired mean 1.0012, standard deviation 0.0101. The earlier rejections stand. The router weight cache is the one with a consistent sign (1.010 on three pairs) and gets a powered re-test. - Residency speculation: only 2.2% to 2.5% of layer events need no demand read at the B0 setting, so skipping the routing readback on complete layers has almost nothing to act on ([[records/decisions/residency-speculation-waits-for-layer-completeness]]). - Graph compilation: the model thread's leaf samples are 69.7% GPU wait without prefetch and 73.5% with it, while MLX evaluation and Metal command encoding self time sits in the tens of samples at the top of the process-wide stack. The host share is waiting on the GPU and on reads, not graph construction ([[records/decisions/decode-host-time-is-waiting-not-graph-construction]]). **Also measured.** Prefetch halves the model thread's lock and condition-variable waiting (9.9% to 4.4%), consistent with its 47% cut in demand-read joins. The first reading of the flag sweep divided medians taken across prompts of different intrinsic speed and overstated every flag at 1.010 to 1.023; paired per-request ratios replaced it in both tools. **Limits and next.** Exploration on three prompts at 256 outputs, not a held-out cohort. The barrier was measured with prefetch off, because a pending forecast forces the drain in this build; deferred forecast consumption is written, and its composition with prefetch is the next measurement, followed by prefetch coverage depth, draft depth under per-depth protocols and the powered router-weight re-test. No public claim. Commands, tables, profiles, interruptions and hashes: [[sources/runs/2026/09/2026-09-12-decode-path-serialization-round-1]]. **Correction (2026-09-13).** The profile split above describes one of the model thread's two blocks in the `sample` output, the cooperative-queue block that runs the layer loop. The same thread's other block, where it reads records, is 91% file reads. Merged by thread id, the model thread spent 35.1% (prefetch off) and 41.3% (prefetch on) of its samples in GPU waits and 45.2% and 40.0% in file reads, and prefetch lowered locks and condition variables from 9.0% to 5.9% rather than from 9.9% to 4.4%. Host work outside waiting is 8% to 10% either way, so the compilation conclusion holds: [[records/measurements/decode-path-serialization-closing-profiles-2026-09-13]]. ### Decode path serialization, round 2: the B0 prefetch coverage is already the best point; closer or wider forecasts trade demand reads for waits **Outcome: the B0 prefetch setting is already the best coverage point in this sweep. Candidate lists deeper than ten are inert, and forecasts that are closer or wider trade demand reads for waits on reads still in flight.** Reference: the B0 setting (top 10, stride 2, margin threshold 0.062). Paired geometric means over six request-round pairs; outputs identical to the reference in every cell. | configuration | paired ratio | pairs above 1 | demand records | demand misses | promoted while in flight | join and adopt per run | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | top 32, stride 2 | 1.022 | 4 of 6 | 15,654 | 14,243 | 152 | 0.01 s | | top 24, stride 2 | 1.015 | 5 of 6 | 15,661 | 14,248 | 184 | 0.02 s | | top 16, stride 2 | 0.991 | 4 of 6 | 15,656 | 14,244 | 138 | 0.01 s | | top 10, stride 2 (B0) | reference | | 15,695 | 14,276 | 189 | 0.02 s | | top 16, strides 1 and 2 | 0.991 | 1 of 6 | 10,366 | 8,971 | 5,169 | 2.10 s | | top 24, stride 1 | 0.969 | 3 of 6 | 12,455 | 10,820 | 16,126 | 2.59 s | | top 24, no threshold | 0.555 | 0 of 5 | 7,098 | 5,752 | 14,032 | 7.27 s | **Deeper lists are inert, which makes them an A/A test.** Margins are measured against each row's tenth logit, so from the tenth rank down every margin is zero or negative and the threshold removes it. Top 16, 24 and 32 therefore issue the same reads as top 10: 28,598 to 28,602 candidates, with 0.8% more reads let through by the larger per-target issue cap. Their paired ratios, 0.991, 1.015 and 1.022, with single pairs from 0.848 to 1.085, measure the exploration sweep's own noise: a six-pair difference inside about 2.5% is not evidence here. Round 1's barrier results at K of 3 and above (1.038 to 1.051, every pair above 1) sit outside that band; K = 2 (1.025) sits at its edge. **Closer and wider forecasts arrive too late to pay.** Stride 1 forecasts one layer ahead. Its forecasts are more accurate (wasted bytes halve, 32.3 to 16.1 GB per run, and demand misses fall 24%), but 16,126 of its reads were still in flight when the layer needed the expert, and joining and adopting them took 2.59 s per run against 0.02 s at stride 2. Forecasting from both strides issues 35% more reads, cuts demand records 34% and spends 2.10 s joining and adopting, for 0.991 with one pair of six above 1. Removing the threshold issues 4.2 times the reads, queues a million deferred lane acquisitions and runs at 0.555. Fewer demand records is not the objective by itself: a speculative read pays only when it lands before its layer asks and does not queue behind other speculative reads. **Combination rule.** The pre-registered rule for step 6 (at least five pairs, paired mean at least 1.01, at least 80% of pairs above 1) selects top 24 at stride 2 for this lever, out of the A/A set. Its overrides are the top and the issue cap, so it carries no measurable change into the combined candidate, and no gain is attributed to coverage. The rule is permissive at six pairs, as this sweep shows directly; the held-out B1 cohort remains the test that counts. **Also measured.** Layer completeness stayed at 2.2% to 2.5% at every setting. One no-threshold cell was excluded for host swap-outs; system swap held 0.25 MB afterwards and every configuration peaked at 18.8 GB. **Limits and next.** Three exploration prompts at 256 outputs over two rounds, not a held-out cohort. The composition of the deferred barrier with prefetch is round 3. No public claim. Commands, tables, counters and hashes: [[sources/runs/2026/09/2026-09-12-decode-path-serialization-round-2]]. ### Decode path serialization, round 3: forecasts held to a deferred barrier lose 7% to 26% under prefetch **Outcome: holding router forecasts until a deferred barrier is exact but costs 7% to 26% under prefetch, with no pair above 1, so the barrier gain does not survive this design.** Reference k1-s2, the B0 prefetch setting with a barrier at every layer. Paired geometric means over six pairs (five for k3-s4); outputs identical to the reference in every cell; exact parity at K = 3, stride 3 on the six correctness requests. | configuration | paired ratio | pairs above 1 | range | demand records | adopted | promoted while in flight | deferred lane acquisitions | | --- | ---: | ---: | --- | ---: | ---: | ---: | ---: | | K = 1, stride 2 (reference) | | | | 15,685 | 13,706 | 111 | 38,096 | | K = 1, stride 3 (control) | 0.932 | 1 of 6 | 0.849 to 1.008 | 18,292 | 11,116 | 6 | 41,962 | | K = 3, stride 3 | 0.877 | 0 of 6 | 0.854 to 0.935 | 18,394 | 11,020 | 5,128 | 108,387 | | K = 3, stride 2 | 0.865 | 0 of 6 | 0.658 to 0.973 | 21,989 | 8,984 | 4,271 | 27,484 | | K = 4, stride 4 | 0.814 | 0 of 6 | 0.767 to 0.864 | 20,232 | 9,549 | 4,071 | 200,524 | | K = 3, stride 4 | 0.797 | 0 of 5 | 0.769 to 0.841 | 20,127 | 9,703 | 2,056 | 187,293 | | K = 8, stride 4 | 0.740 | 0 of 6 | 0.335 to 0.913 | 27,732 | 3,925 | 1,488 | 87,798 | **Why.** A forecast is worth its lead time, and holding it to the barrier spends that lead time three ways. At stride 2 a third of the forecasts target the barrier layer itself and reach the scheduler after it completed, so they are never issued: issued reads fall 35% and demand records rise 34%. A longer stride restores the distance but forecasts less accurately: stride 3 alone, at K = 1, adopts 19% fewer reads and runs at 0.932. And the reads that are issued leave in bursts at each barrier: at K = 3, stride 3, deferred lane acquisitions reach 2.6 times the stride-3 control, 5,128 reads were still in flight when their layer asked, and joining and adopting them took 1.69 s per run against 0.01 s. That configuration runs at 0.877 against the control's 0.932, a 5.9% loss at the same stride, where round 1 measured a 3.8% gain for the same period without prefetch. The loss grows with the period, to 0.740 at K = 8. **What follows.** Deferring the drain needs each forecast consumed before the next layer's routing, not at the next barrier. That routing readback is a synchronization the host takes anyway, so round 3b rebuilds with forecasts and completed-layer ticks riding it, which keeps their lead time within one attention block of K = 1 ([[records/plan/decode-path-serialization-2026-09-12]]). **Also measured.** Round 1's profiles rule out a third fold. Multi-token passes compact the linear-attention state windows with one evaluation each after the barrier, but that call site holds about 150 to 220 model-thread samples against about 3,700 to 4,000 at the barrier and 14,400 to 16,900 in the MoE call, so folding it could move at most about 1% of layer-loop time, inside the exploration noise band; it was not built. One k3-s4 cell was excluded for host swap-outs. **Limits.** Three exploration prompts at 256 outputs over two rounds. The design measured here is replaced in the source by round 3b's. No public claim. Commands, tables, counters and hashes: [[sources/runs/2026/09/2026-09-12-decode-path-serialization-round-3]]. ### Decode path serialization, round 4: draft depth 2 stays; wider verification passes cost expert reads **Outcome: draft depth 2 stays. Deeper drafts cut forward passes but widen every verification pass, and on an SSD-streamed MoE each verified position loads its own experts; depth 1 changes the greedy output.** Reference: depth 2 with the B0 prefetch setting, each depth under a protocol variant pinning it. Paired geometric means over six pairs. | draft depth | paired speed | pairs above 1 | forward passes | demand records per pass | seconds per pass | outputs identical to depth 2 | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | 1 | 0.957 | 1 of 6 | 1.310 | 0.744 | 0.800 | 0 of 6 | | 2 (reference) | | | 96 | 163.6 | | | | 3 | 1.022 | 4 of 6 | 0.829 | 1.254 | 1.181 | 6 of 6 | | 4 | 0.981 | 2 of 6 | 0.755 | 1.516 | 1.351 | 6 of 6 | | 6 | 0.877 | 0 of 6 | 0.687 | 2.069 | 1.661 | 6 of 6 | **Why.** Every draft position is verified in the same pass, and every verified position routes to its own experts, so demand records per pass grow with the positions: 25% more at depth 3 (a third more positions), 52% at depth 4 and 107% at depth 6. Acceptance falls at the same time, 10%, 22% and 41% below depth 2, so passes shrink less than their cost grows: the two nearly cancel at depth 3 (1.022, with four pairs of six above 1, inside the exploration noise band measured in [[records/measurements/decode-path-serialization-round-2-2026-09-12]]), and depths 4 and 6 lose. For an SSD-streamed MoE, verification width is paid in reads. **Depth 1 is not exact.** It diverged from depth 2 at the same output in both rounds of every prompt (outputs 96, 27 and 63). Under the interpretation registered before the round, a depth change alters the verification batch shape, which the engine already documents can flip greedy near-ties (`boundedDraftTail`), so depth 1 is recorded as not exact under a shape change and excluded. It was also slower, at 0.957. **Combination rule.** No depth qualifies. Depth 3's paired mean clears 1.01, but only four of six pairs are above 1, short of the 80% the rule requires. Step 6 keeps depth 2 and the protocol's pin. **Limits.** Three exploration prompts at 256 outputs over two rounds, on the round 3 binary at barrier period 1. No public claim. Commands, tables, counters and hashes: [[sources/runs/2026/09/2026-09-12-decode-path-serialization-round-4]]. ### Decode path serialization, round 5: the router weight cache holds at 1.017 over twelve pairs **Outcome: the router weight cache survives a powered re-test at 1.017 over twelve pairs, ten above 1, approximate 95% interval 1.003 to 1.031, with exact outputs; adding the specialized router top-k removes the gain.** Reference: the B0 prefetch setting. Four prompts (r0005, r0206, r0096 and r0074) over four rounds; eight of 48 cells were excluded for host swap-outs, leaving twelve clean pairs per configuration. | configuration | paired ratio | pairs above 1 | range | approximate 95% interval | router cache | | --- | ---: | ---: | --- | --- | ---: | | router weight cache | 1.017 | 10 of 12 | 0.982 to 1.056 | 1.003 to 1.031 | 256.9 MB | | cache plus router top-k | 1.002 | 7 of 12 | 0.945 to 1.058 | 0.984 to 1.020 | 256.9 MB | **What it does.** The router matmul promotes each BF16 router to FP32 on every call. The cache keeps a pre-materialized FP32 copy of every router, 256.9 MB, so the routing of each layer skips that conversion. Demand records are unchanged (17,873 against 17,877), so the gain is compute, not reads. Round 1 saw 1.010 on three pairs among nine flags whose paired mean was 1.0012 ([[records/measurements/decode-path-serialization-round-1-2026-09-12]]); at twelve pairs the effect holds, small but clear of zero. **Combination rule.** The cache qualifies (twelve pairs, paired mean at least 1.01, 83% of pairs above 1) and carries `SLOTSTREAM_OPT_ROUTER_WEIGHTS=1` into step 6. Router top-k does not. **Limits.** Exploration prompts at 256 outputs; the interval is a normal approximation on log ratios. Swap-outs grew during the round and removed a sixth of the cells. The cache adds 256.9 MB of resident memory. No public claim. Commands, tables and hashes: [[sources/runs/2026/09/2026-09-12-decode-path-serialization-round-5]]. ### Decode path serialization, round 3b: forecasts on the routing readback fix round 3 but add about 1% **Outcome: consuming router forecasts at the next routing readback removes round 3's loss and stays exact, but the gain it leaves is small: at barrier period 4, 1.014 over the four pairs not affected by two slow reference cells, all four above 1; the other periods sit between 0.96 and 1.00 on those pairs.** Reference: the B0 prefetch setting with a barrier at every layer. Outputs identical to the reference in every cell; exact parity at K = 8 on the six correctness requests. | barrier period | paired ratio, all clean pairs | pairs above 1 | without the r0206 pairs | pairs above 1 | median pair | | ---: | ---: | ---: | ---: | ---: | ---: | | 2 | 1.065 | 3 of 6 | 0.991 | 1 of 4 | 0.999 | | 3 | 1.084 | 3 of 5 | 0.992 | 1 of 3 | 1.007 | | 4 | 1.087 | 6 of 6 | 1.014 | 4 of 4 | 1.025 | | 8 | 1.125 | 3 of 4 | 0.998 | 1 of 2 | 1.103 | | 16 | 1.058 | 3 of 6 | 0.960 | 1 of 4 | 0.999 | **Two slow reference cells carry the headline ratios.** Both r0206 reference cells ran at 10.99 and 9.78 tok/s, slower than every other r0206 cell at barrier period 1 that evening (11.6 to 13.3 across rounds 2 to 5). Their demand records and decode I/O time match the other configurations' r0206 cells, and the excess, 1.2 to 2.1 s in round 0 and 3.8 to 4.1 s in round 1, is outside I/O. Neither tripped the eligibility rule, which sees swap-outs and page-ins but not processor or GPU contention from other processes. A slow reference inflates every ratio in its pair, so r0206 adds ratios of 1.12 to 1.37 to every period; without those pairs the periods read 0.96 to 1.01. **Round 3's loss is gone.** At K = 3, stride 2, round 3 ran at 0.865 with a third of its forecasts never issued ([[records/measurements/decode-path-serialization-round-3-2026-09-12]]). Here the same configuration issues and adopts as many reads as K = 1 (28,405 and 13,720 against 28,386 and 13,708) and demand misses are unchanged. Forecasts reach the scheduler one attention block later, so about 2,800 to 3,700 reads are still in flight when their layer asks, against 394 at K = 1, and joining and adopting them takes 0.14 to 0.18 s per run against 0.03 s. **Combination rule.** The pre-registered rule reads all clean pairs and selects K = 4 (six pairs, all above 1), carrying `SLOTSTREAM_DECODE_BARRIER_LAYERS=4` into step 6; K = 8 had only four clean pairs. The selection stands as registered: the setting is exact, and the screen against the shipped path and the held-out cohort measure what it adds. This round supports a gain of about 1% at K = 4, not 8.7%. **Limits.** Three exploration prompts at 256 outputs over two rounds; three cells excluded for host swap-outs. The reading without r0206 drops pairs after seeing them and is not the registered estimator. No public claim. Commands, per-pair ratios, counters and hashes: [[sources/runs/2026/09/2026-09-13-decode-path-serialization-round-3b]]. ### Decode path serialization, step 6: the combined candidate screens at 1.104 against shipped, 1.018 over B0 **Outcome: the combined candidate (B0 prefetch, barrier period 4 with forecasts on the routing readback, the router weight cache and an inert coverage override) is exact and screens at 1.104 against the shipped path on the exploration prompts, against 1.084 for B0 prefetch alone: 1.018 on top of B0, five of six pairs above 1.** Six clean pairs per configuration. | configuration | paired ratio against shipped | pairs above 1 | paired ratio against B0 prefetch | median decode seconds | median demand records per output | median decode I/O share | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | shipped | reference | | | 19.18 | 131.3 | 41.3% | | B0 prefetch | 1.084 | 6 of 6 | reference | 18.31 | 69.3 | 26.1% | | combined | 1.104 | 6 of 6 | 1.018, 5 of 6 above 1 | 17.70 | 68.9 | 29.6% | **Reading.** The combination adds 1.8% on top of B0 prefetch, in line with the parts measured alone: the router weight cache at 1.017 ([[records/measurements/decode-path-serialization-round-5-2026-09-12]]) and barrier period 4 at about 1.01 once two slow reference cells are set aside ([[records/measurements/decode-path-serialization-round-3b-2026-09-13]]); the coverage override changes nothing. The combined arms spend less time outside I/O than B0 prefetch (about 12.5 s against 13.5 s per run) and a little more waiting on reads (about 5.2 s against 4.8 s), consistent with forecasts reaching the scheduler one attention block later. B0 prefetch screens at 1.084 here against its held-out 1.105 (rescored from 1.120, [[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]) ([[records/measurements/expert-lookahead-2-b0-cohort-2026-09-12]]) because these are three prompts at 256 outputs, not the registered cohort. **Next.** Step 7 runs the combined candidate against the shipped path on the held-out B1 prompts under the B0 gate. **Limits.** Three exploration prompts at 256 outputs over two rounds; the medians are unpaired and the time split multiplies medians. No public claim. Commands, per-pair ratios and hashes: [[sources/runs/2026/09/2026-09-13-decode-path-serialization-combination]]. ### Decode path serialization, step 7: B1 cohort at 1.106 against shipped, evidence insufficient **Outcome: on the held-out B1 prompts the combined candidate decoded 10.6% faster than the shipped path (aggregate 1.106, bootstrap 1.073 to 1.128), with every family at 1.033 or above, every duration shorter and every output identical, but the registered verdict is not a pass: one prompt kept a single clean pair where the gate requires two, after host swap-outs from another application's virtual machine.** The plan allows no reruns inside a cohort, so this run records evidence insufficient and success false. **Correction (2026-09-13).** This record first reported 1.116 (bootstrap 1.073 to 1.129) with every family at 1.037 or above. The cohort report took the upper of the two middle values for medians over an even count; scored again from the same pairs with true medians, the aggregate is 1.106 and the verdict is unchanged: [[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]. | gate | required | result | | --- | --- | --- | | aggregate ratio | at least 1.10 | 1.106 | | bootstrap lower bound | above 1.00 | 1.073 | | family floor | at least 0.95 | 1.033 (reasoning) | | duration regression | at most 5% | every family shorter, 0.862 to 0.972 | | clean pairs per prompt | at least 2 | r0245 has 1 | | family | tok/s ratio | duration ratio | | --- | ---: | ---: | | dialogue | 1.175 | 0.862 | | prose | 1.167 | 0.928 | | multilingual | 1.124 | 0.918 | | structured | 1.077 | 0.949 | | code | 1.067 | 0.950 | | reasoning | 1.033 | 0.972 | **What it measures.** The cohort compares the whole candidate with the shipped path, so it prices B0 prefetch and the new levers together. It has no B0-only arm; B0 prefetch passed alone at 1.105 on the B0 prompts ([[records/measurements/expert-lookahead-2-b0-cohort-2026-09-12]], rescored the same way), where dialogue and prose also gained most. The increment of barrier period 4 and the router weight cache over B0 is measured only in exploration, at 1.018 ([[records/measurements/decode-path-serialization-combination-screen-2026-09-13]]). **Why the evidence fell short.** Three of 36 pairs were excluded for host swap-outs: r0244 round 0 and r0245 rounds 0 and 1. A virtualization process from another application held 8.41 GB by the end of the run, and the host swap-out counter rose by 13,960 during the cohort, although it had held still for five minutes before the cohort started. **Next.** A complete B1 rerun of the same candidate on a host without that memory pressure would test the same registered gates. It is a replication, not a top-up: its verdict would stand on its own, with this run reported alongside. No public claim. Commands, per-prompt ratios, exclusions and hashes: [[sources/runs/2026/09/2026-09-13-decode-path-serialization-b1-cohort]]. Rescoring: [[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]. ### Decode path serialization, step 7 replication: the combined candidate passes B1 at 1.114 against shipped **Outcome: the replication passes every registered gate. On the held-out B1 prompts the combined candidate decodes at an aggregate 1.114 against the shipped path (bootstrap 1.104 to 1.121), every family at 1.064 or above, every family's request duration shorter, 34 of 36 pairs eligible and every output identical.** Median decode throughput rose from 11.79 to 13.47 tok/s. **Correction (2026-09-13).** This record first reported an aggregate of 1.124 (bootstrap 1.105 to 1.122), every family at 1.072 or above and medians of 11.80 to 13.48 tok/s. The cohort report took the upper of the two middle values whenever it formed a median over an even count, so each two-prompt family counted its faster prompt, which is why 1.124 sat above its own bootstrap interval. Scored again from the same pairs with true medians, a family being the geometric mean of its prompt medians as the registered bootstrap already resampled, the aggregate is 1.114. Every gate still passes and the verdict is unchanged: [[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]. | gate | required | first run | replication | | --- | --- | ---: | ---: | | aggregate ratio | at least 1.10 | 1.106 | 1.114 | | bootstrap lower bound | above 1.00 | 1.073 | 1.104 | | family floor | at least 0.95 | 1.033 | 1.064 | | duration regression | at most 5% | none | none | | clean pairs per prompt | at least 2 | r0245 had 1 | met | | verdict | | not a pass | pass | | family | tok/s ratio | duration ratio | | --- | ---: | ---: | | prose | 1.176 | 0.923 | | dialogue | 1.169 | 0.865 | | multilingual | 1.136 | 0.914 | | structured | 1.075 | 0.953 | | code | 1.070 | 0.956 | | reasoning | 1.064 | 0.943 | **Standing of the two runs.** The replication was registered before it ran as a complete fresh run of the same candidate, binary, plan and gates; its verdict stands on its own and no pair from the first run entered it. The first run ([[records/measurements/decode-path-serialization-b1-cohort-2026-09-13]]) is reported alongside: it cleared every effect gate and failed only evidence sufficiency, after host swap-outs from another application's virtual machine. The two runs agree family by family within 0.012 except reasoning (1.033 and 1.064). **What it measures.** The cohort prices the whole candidate against the shipped path: B0 prefetch, barrier period 4 with forecasts on the routing readback, the router weight cache and an inert coverage override. B0 prefetch alone passed at 1.105 on the B0 prompts ([[records/measurements/expert-lookahead-2-b0-cohort-2026-09-12]], rescored the same way); the two cohorts use different prompts, so the increment over B0 is not read from their difference. It is measured in exploration at 1.018 ([[records/measurements/decode-path-serialization-combination-screen-2026-09-13]]) and, on one binary with sixteen pairs, by the attribution sweep registered with this replication. **Limits.** One machine at the 20 GB profile, text decode with MTP drafts at depth 2, the twelve B1 prompts. Defaults are unchanged and nothing is committed or installed; adopting the flags is separate engineering. No public claim yet. Commands, per-prompt ratios, exclusions and hashes: [[sources/runs/2026/09/2026-09-13-decode-path-serialization-b1-cohort-replication]]. Rescoring: [[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]. **Adoption (2026-09-13).** 0.2.16 turns this configuration on by default wherever the draft head runs ([[records/decisions/decode-lookahead-default-with-the-draft-head]]). Public claims citing this record: [[records/claims/decode-lookahead-1-11x-on-held-out-prompts]], [[records/claims/warm-decode-13-5-tok-s-with-the-decode-lookahead]] and [[records/claims/decode-lookahead-11-8-to-13-5-tok-s]]. ### Decode path serialization, attribution: each new part adds about 2% over B0 prefetch, 4.6% together **Outcome: on one binary, the two new parts each add about 2% on top of B0 prefetch and 4.6% together, every pair above 1 and outputs exact; B0 prefetch itself is 1.090 over the shipped path on these prompts.** Reference: B0 prefetch. Exploration prompts r0005, r0206, r0096 and r0074 over four rounds at 256 outputs; 78 of 80 cells clean. | configuration | against B0 prefetch | pairs above 1 | approximate 95% interval | | --- | ---: | ---: | --- | | shipped path | 0.918, so B0 prefetch is 1.090 over it | 1 of 15 | 0.894 to 0.943 | | B0 plus the router weight cache | 1.021 | 11 of 15 | 0.994 to 1.050 | | B0 plus barrier period 4, forecasts on the routing readback | 1.022 | 13 of 14 | 1.008 to 1.036 | | combined candidate | 1.046 | 15 of 15 | 1.024 to 1.069 | **The parts compose.** 1.021 × 1.022 = 1.044 against 1.046 measured together, so the router weight cache (compute saved in every router matmul) and the deferred barrier (synchronizations removed) do not overlap. Both agree with their earlier readings: the cache measured 1.017 in [[records/measurements/decode-path-serialization-round-5-2026-09-12]], and barrier period 4 about 1.01 on the pairs round 3b could trust ([[records/measurements/decode-path-serialization-round-3b-2026-09-13]]). The six-pair screen's 1.018 for the combination ([[records/measurements/decode-path-serialization-combination-screen-2026-09-13]]) sat inside its noise band; sixteen pairs place it at 1.046. **Apportioning the held-out result.** | step | gain | where measured | | --- | ---: | --- | | B0 prefetch over the shipped path | 1.090 here; 1.105 held out | this sweep; [[records/measurements/expert-lookahead-2-b0-cohort-2026-09-12]] | | router weight cache over B0 | 1.021 | this sweep | | barrier period 4 with forecasts on the routing readback, over B0 | 1.022 | this sweep | | both over B0 | 1.046 | this sweep | | combined candidate over the shipped path | about 1.14 here; 1.114 held out | this sweep; [[records/measurements/decode-path-serialization-b1-cohort-replication-2026-09-13]] | **Limits.** Exploration prompts at 256 outputs; intervals are normal approximations on log ratios. The parts were selected on these prompts, so the held-out 1.114 is the claim and these ratios apportion it. The held-out figures are the rescored ones ([[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]); the cohort report first gave 1.124 and 1.120 by taking the upper middle value of even-count medians. No public claim. Commands, per-prompt ratios and hashes: [[sources/runs/2026/09/2026-09-13-decode-path-serialization-attribution]]. **Adoption (2026-09-13).** The combined candidate is the 0.2.16 default ([[records/decisions/decode-lookahead-default-with-the-draft-head]]). docs/ENGINEERING.md quotes this sweep's split through [[records/claims/lookahead-parts-add-about-2-percent-each]]. ### Decode path serialization, closing profiles: decode waits split between file reads and the GPU **Outcome: on the whole model thread, decode waiting now splits between file reads and the GPU. With B0 prefetch the thread spends 36% of its samples in reads and 40% in GPU waits; the combined candidate cuts GPU-wait samples by 18% and adds 9% in reads. Round 1's recorded split covered only one of the thread's two profile blocks.** | bucket | round 1, prefetch off | round 1, B0 prefetch | B0 prefetch, round 3b binary | combined candidate | | --- | ---: | ---: | ---: | ---: | | file reads | 45.2% | 40.0% | 36.1% | 39.3% | | GPU wait | 35.1% | 41.3% | 39.8% | 32.9% | | other host work | 8.1% | 9.5% | 14.1% | 16.6% | | locks and condition variables | 9.0% | 5.9% | 6.1% | 6.3% | | allocation and memory copy | 2.6% | 3.3% | 3.8% | 4.8% | **What changed with the combination.** Fewer synchronizations and cheaper router math show up as 1,980 fewer samples in the IOKit trap (11,254 to 9,274). Reads rise by 885 samples (10,212 to 11,097), consistent with forecasts reaching the scheduler one attention block later, which the attribution counters also show as more reads still in flight when their layer asks. Host work outside waiting rises from 14.1% to 16.6% but stays a thin tail: no leaf above 433 samples, and `mlx::core::eval_impl` at 94 and 112 of about 28,000. **Correction to round 1.** `sample` lists the model thread once per dispatch queue. Round 1's parse kept the cooperative-queue block that runs the layer loop, whose shares reproduce the recorded 69.7% and 73.5% GPU wait, and missed the block where the same thread reads records, which is 91% file reads. Merged, round 1's model thread spent 35.1% (prefetch off) and 41.3% (B0 prefetch) in GPU waits and 45.2% and 40.0% in file reads, and prefetch lowered locks and condition variables from 9.0% to 5.9%, not from 9.9% to 4.4%. Host work outside waiting stays between 8% and 17% in every profile, so [[records/decisions/decode-host-time-is-waiting-not-graph-construction]] stands, with its evidence updated. **Where the next gains are.** Reads the model waits on are again the largest single cost, and round 2 showed that more speculative reads do not reduce them ([[records/measurements/decode-path-serialization-round-2-2026-09-12]]), so the lever is forecast accuracy at the same lead time. GPU waits are the other third. **Limits.** One prompt and one run per profile, in 45 s windows that can include warmup decode; shares are of model-thread samples, not wall time. No public claim. Commands, block tables and hashes: [[sources/runs/2026/09/2026-09-13-decode-path-serialization-closing-profiles]]. ### Decode lookahead default: checks and plans by Mac memory **Outcome: the 0.2.16 defaults are implemented and pass every weights-free check. Auto now runs the draft head and the decode lookahead from a 21 GB target, which puts 32 GB Macs and up on them at the default context; an 8 GB Mac is refused because even the smallest plan does not fit.** The decode lookahead is the exact configuration the held-out B1 replication measured at 1.114 ([[records/measurements/decode-path-serialization-b1-cohort-replication-2026-09-13]]); this record prices where the defaults turn on, not a new timing. **What changed.** The planner decides the lookahead with the head and charges its 373 MiB, the prefetch staging reserve and the FP32 router copies, before sizing the expert pool. The engine loads the qualified configuration when the plan chose it and no prefetch switch is set, then turns on the router cache and a four-layer barrier unless the environment names either. A deferred barrier falls back to draining every layer when a pass could not keep that many layers pinned. The governor re-plans with the engine's decision and charge. The head's automatic floor is 76 experts per layer after its charge ([[records/decisions/draft-head-auto-floor-76-per-layer]], [[records/decisions/decode-lookahead-default-with-the-draft-head]]). **Checks.** On the 0.2.16 build, 50 of 50 T0 and T1 checks pass (29,987 assertions), including a new 32-assertion check that parses the B1 candidate environment and requires it to equal the built-in configuration. The planner gates pass 73 of 73 with new checks at the floor, by Mac size, at the 65,536-token window and for the `SLOTSTREAM_OPT_EXPERT_PREFETCH=0` override. The static gates passed on the build before the version bump. Two checks failed on the first build and one on the second; each was a fixture or expectation error, recorded with its fix in the run. **Plans by Mac memory**, pristine what-if at the default window: | Mac RAM | automatic target | experts per layer | head and lookahead | plain-decode estimate | | --- | ---: | ---: | --- | ---: | | 8 GB | refused | | | | | 16 GB | 10 GB | 20 | off | 4.00 | | 18 GB | 11.5 GB | 28 | off | 5.54 | | 24 GB | 16 GB | 54 | off | 7.61 | | 32 GB | 22 GB | 74 | on | 8.68 | | 36 GB | 25 GB | 96 | on | 9.68 | | 48 GB | 33.6 GB | 161 | on | 11.60 | | 64 GB and up | 34.6 GB | 149 | on | 11.56 | At the 65,536-token window a 32 GB Mac's cache falls below the floor and runs without the head, while 36 GB keeps both; a 16 GB Mac's cache drops from 20 to 13 experts per layer, and a full window waits about 12.9 minutes against 6.4 at 32,768. From 24 GB a full 32,768-token prompt waits 3.0 to 3.3 minutes and a 65,536-token one 7.5 to 7.8 minutes. These plans back the public tier tables, their context recommendations and [[records/claims/recommended-context-65536-from-36-gb]]. **Limits.** Planner arithmetic on simulated machines; estimates use the M5 Pro curve and do not include speculative decoding. No model process ran: a functional run of the automatic path needs a target of at least 21 GB, which the host's reclaimable memory did not allow while another application's virtual machine held about 9.7 GB. Public speed figures come from the B1 measurement, the 0.2.14 depth study and the claims that cite them. Commands, build identities, per-size plans and hashes: [[sources/runs/2026/09/2026-09-13-decode-lookahead-default-tier-plans]]. ## Public documentation evidence and scope audit This is a documentation and evidence-scope audit, not a new model benchmark or a release qualification. It follows the README memory-tier correction in [[records/measurements/c2-macbook-pro-m5-max-128gb-community]]. ## Scope and evidence Read README.md, all Markdown guides under docs/, llms.txt, CONTRIBUTING.md, SECURITY.md and the current release-preparation notes. Checked the related claims, supersession notes, planner and context accounting, vision admission, CLI behavior, package/build instructions and published release metadata. Generated MEASUREMENTS.md and PLAN.md remain projections of the canonical records; historical source/run bytes and past release entries are preserved. [[sources/runs/2026/09/2026-09-13-public-docs-audit-plans]] captures the exact binary identity, eighteen default/larger-context simulations and the latest published release at audit time. They are planner observations only. The working tree concurrently contained unrelated source edits, so neither these observations nor the planner gates qualify those edits. ## Corrections - Kept the positive same-Mac larger-cache evidence visible. The old flat development-Mac number does not establish a larger-memory speed ceiling. - Labeled prompt-processing waits as M5 Pro-based estimates. They exclude startup, queues, image preparation and reasoning before visible answers. Near-full requests must leave reply room. - Distinguished decimal-GB simulations from marketed Mac memory and showed that the auto-target/MTP columns use the default window. Recommendations are separate simulations whose actual fit depends on available memory, Metal capacity, selected context and draft weights. - Limited the draft-enabled automatic ceiling to its default-context scope; larger windows add draft-context charges. The draft activation threshold is evaluated before the separate lookahead reservation. - Corrected image memory on the engineering, Hermes and AI-facing pages: auto/total-target plans reserve the tower inside the target; explicit pool-size settings preserve the pool and add resident bytes to the expected footprint. Both need real headroom and image workspace. - Clarified that main sequence-cache bytes per allocated token exclude recurrent, retained, draft and transient allocations. Manual targets keep request-memory safeguards despite disabling automatic resizing. - Scoped cache equality to fixed generation settings; total-target changes may also change prefill or MTP. Labeled the early engine-load figure historical and excluded current full-file verification from that timing. - Scoped queue-depth observations to their historical development-Mac run. Preserved the separate attribution study's supported component results after checking its newer evidence; they describe tuning prompts, not independent held-out improvements. - Distinguished candidate measurements and prepared version headings from a published installer release. Removed the stale unreleased label from features delivered earlier. - Corrected build-versus-T0 network requirements, removed a duplicated library paragraph, repaired the stale README memory reference, and updated stale contributor activation/plateau language. The claim records carry the refined scopes. The semantic review requirement is canonical in [[records/design/measured-operating-policies]] and projected into CONTRIBUTING.md. It supplements the existing text-match gate. ## Verification and limits All 73 existing planner gates passed against the available binary, and all 179 claim-to-surface checks passed. A local-link scan found all 160 checked links and fragments resolvable. Generated-doc parity, full-store validation, Markdown table-column checks and diff whitespace are checked after the final edits. The store has two unchanged historical-log warnings; no log history is rewritten to silence them. The direct sources support the corrections. This audit did not rerun real model performance, every client integration or other hardware, and does not certify a future client/library version. It does not assert that every possible documentation error has been ruled out. No runtime defaults changed in this task; no public release or push was performed. ## Follow-up: useful estimates, 2026-09-13 The user requested best-effort estimates after questioning High/Ultra again. README now presents broad ranges separately from measured configurations. [[records/measurements/hardware-planning-ranges-2026-09-13]] owns the endpoint construction and explicitly unmeasured hardware transfers. Installed RAM, process target and software version have separate measured-table columns. The hardware guide's allocation table no longer mixes allocation, measured speeds and larger-cache estimates in the same column. Old section anchors remain usable. Existing source bytes are unchanged. ## Hardware speed planning ranges and inference limits These are editorial planning estimates requested by the user, not a new benchmark. The following range rationale is projected manually into the hardware guide; the canonical evidence and revision scope live here. ## Estimate construction The README's estimates combine the real reports above with the development Mac's measured configurations and planner curve. They are rough expectations across hardware and settings, not a fitted scaling model or statistical confidence intervals. Endpoints are rounded outward to whole tok/s. | Installed RAM | Estimated warm reply speed | Basis and main inference | |---|---|---| | 16–<24 GB | ~1–6 tok/s | The M2 mini reported 1.41 tok/s; the M5 Pro-based 16/18 GB simulations estimate about 4 to 5.5 tok/s. The upper end has not been measured on a real Mac in this band. | | 24–<48 GB | ~6–14 tok/s | The 32 GB M5 Air reported 6.22 tok/s; the M5 Pro achieved 13.47 tok/s at a 20 GB process target. The upper end assumes a comparable chip and SSD with enough memory for that configuration; it has not been timed on a Mac in this band. | | 48–<96 GB | ~13–27 tok/s | The latest 48 GB M5 Pro result is 13.47 tok/s, rounded down for this estimate; its older ~12 tok/s result remains historical evidence. The upper end transfers the M5 Max's 26.9 tok/s at a 48 GB process target to a comparable Mac with enough available memory. That run used a 128 GB Mac; it was not a measurement of a 48 GB Mac. | | 96 GB+ | ~20–32 tok/s | The 128 GB M5 Max reported about 21 to 22 tok/s in auto and 31.5 tok/s at a 73 GB process target. Applying this range to other Macs in the band is an estimate. | The Ultra lower endpoint allows for the same reporter's roughly 20 tok/s warm auto runs on 0.2.1; the main results table uses the updated 0.2.3 report. The upper ends of High and Ultra assume an M5 Max-class chip, fast internal SSD, speculative decoding and manual targets that leave room for macOS and other apps. A 48 GB process target cannot consume all of a Mac's installed 48 GB; it needs a larger machine. These ranges mix releases, so they are not predictions for a single current build. No release-speedup multiplier was applied to community reports. A slow SSD, older chip, different prompt, draft acceptance or memory pressure can produce results outside the ranges. More RAM helps only when the engine can use it to reduce a bottleneck; the band labels do not establish a causal speed ranking. In particular, there is no measured performance boundary at 96 GB. The shared context recommendation reflects the current planning guidance, independently of reply speed. ## Supporting evidence - [[records/measurements/c1-mac-mini-m2-16gb-base-storage-community-2026-09-02]] - [[records/measurements/c3-macbook-air-m5-32gb-community]] - [[records/measurements/c2-macbook-pro-m5-max-128gb-community]], including the preserved original report and updated sweep - [[records/measurements/decode-path-serialization-b1-cohort-replication-2026-09-13]] - [[records/measurements/decode-lookahead-default-2026-09-13]], for the simulated small-memory curve, not timing on those Macs Low rounds outward from 1.41 and 5.54; Medium from 6.22 and 13.47; High from 13.47 and 26.9; Ultra from roughly 20 and 31.5 tok/s. The endpoints are approximate anchors, not calibrated minima, maxima or probabilities. Hardware transfers remain unmeasured. The two M5 Pro configurations also change software/settings with the target and cannot isolate the effect of memory. ## Revision and implementation scope README.md keeps a compact estimate table separately from actual reported configurations. docs/HARDWARE.md gives the full rationale and a separate automatic-allocation table. Revise the ranges as comparable hardware reports arrive; preserve the original evidence and keep inference explicit. Doctor can check a memory plan but cannot validate a speed or qualify hardware. No automatic target, runtime behavior, model benchmark, commit or push is part of this estimate update. ## Latest development-Mac reference, 2026-09-13 At the user's request, High now uses the latest development-Mac median of 13.47 tok/s as its lower reference, giving an estimated 13 to 27 tok/s range. The previous 12 to 27 range used the historical 0.2.3 result near 12 tok/s. That historical result remains public with its version. Medium still rounds its upper endpoint to 14. All transfer assumptions and uncertainty above remain in force; this is a revision to the editorial reference, not new performance evidence or a guaranteed minimum for all Macs in the band. The public current-result surfaces use the reported two-decimal medians 11.79 and 13.47 from [[sources/runs/2026/09/2026-09-13-cohort-rescoring-true-medians]]. No source bytes, model benchmark or runtime policy changed. The configuration was measured before release and adopted in 0.2.16; public wording distinguishes that benchmark from a rerun of the downloaded release binary. ## Re-anchored after 0.2.19, 2026-09-16 The 0.2.19 release benchmark measured the development Mac at a 22 GB process target, the automatic target of a 32 GB Mac: 15.86 tok/s with the corrected forecast and 14.38 with the 0.2.18 forecast, medians over the counted cells of each arm ([[records/measurements/corrected-forecast-release-benchmark-2026-09-16]]). That measurement replaces two inferences: - Medium (24 to less than 48 GB) now rounds outward from 6.22 (the 32 GB M5 Air on 0.2.11) and 15.86 (the M5 Pro at the 32 GB automatic target on 0.2.19): ~6–16 tok/s. The upper end still assumes a comparable chip and SSD; no Mac in the band has been timed on 0.2.19. - High (48 to less than 96 GB) now rounds down from 15.86 as its lower reference instead of 13.47: ~15–27 tok/s. The 48 GB Mac was measured at a 22 GB target, below its own 33.6 GB automatic target, whose larger cache has not been timed on 0.2.19. The upper end still transfers the M5 Max's 26.9 tok/s at a 48 GB process target. - The automatic-plan estimate of about 10 tok/s at 32 GB, taken from the two-draft measurement at 76 experts per layer on 0.2.14, is withdrawn in favour of the direct measurement. Low and 96 GB+ are unchanged. The ranges remain editorial estimates across releases, chips and SSDs, not calibrated intervals; no release-speedup multiplier was applied to community reports. Public surfaces: README.md and docs/HARDWARE.md. ## Benchmark-profile clarification, 2026-09-22 The historical throughput remains a valid controlled forecast comparison. Its protocol forces a 256-token prefill pass, disables prefix caching, uses two drafts, and disables adaptive speculation and the draft-tail experiment. That leaves about 100 experts per layer at the 22 GB target. The installed 0.2.23 normal-cache profile at that target instead plans 2048-token passes and about 74 experts per layer. Equal total memory therefore does not mean equal runtime settings. The historical figure does not measure the current automatic plan of a 32 GB Mac and cannot qualify a release-wide speed ratio. The new calibration attempt, retained raw observations, stricter prospective host-load screen and remaining gaps are recorded in [[records/measurements/release-speed-calibration-2026-09-22]]. ## Reports from 36 and 64 GB Macs, 2026-09-24 Three community reports added real Macs to the two middle bands ([[records/measurements/c4-macbook-pro-m3-max-64gb-community]], [[records/measurements/c5-macbook-pro-m4-max-64gb-community]] and [[records/measurements/c6-macbook-pro-m4-max-36gb-community]]): - 24 to less than 48 GB: a 36 GB M4 Max reported 8.41 tok/s on 0.2.22, inside ~6–16 tok/s. - 48 to less than 96 GB: a 64 GB M4 Max reported 15.93 tok/s on 0.2.22, and a 64 GB M3 Max 12.38 tok/s on 0.2.18, both from internal SSDs with the same auto plan. The M3 Max sits below the ~15 floor. Its report predates 0.2.19, so the floor stays ~15 until a rerun on the current release, with `slotstream pull` run first, says whether the gap is the release or the hardware. C4 records why the release is unlikely to close it: the M4 Max ran the same pre-0.2.19 forecast and still decoded faster. The public range names the M3 Max report beside it. - The same M4 Max decoded at 2.98 tok/s from a 10 Gb/s external drive. The ranges assume the model on a fast internal SSD; that result is the slow-SSD case described above, not a band endpoint. No release-speedup multiplier was applied to the community reports. ## A 24 GB report, 2026-09-26 - 24 to less than 48 GB: a 24 GB M4 Pro reported 3.57 tok/s on 0.2.24 ([[records/measurements/c7-macbook-pro-m4-pro-24gb-community]]), below the ~6 floor, and 3.85 to 3.97 tok/s in a later round. Two known differences may account for the gap. Its 512 GB SSD read cold experts at 3.7 GB/s, well below the development Mac's 17.3 GB/s, and the ranges assume a fast internal SSD. And on 0.2.24 a 24 GB plan ran without the draft head and decode lookahead, which 0.2.25 enables at that size. The floor stays ~6 until a rerun on 0.2.25 separates the release from the hardware; the public range names the report beside it. ## The 24 GB re-run, 2026-09-27 The 24 GB M4 Pro's 0.2.25 re-run decoded **5.41 tok/s** ([[records/measurements/c7-macbook-pro-m4-pro-24gb-community]]), up from 3.57 on 0.2.24, with the draft head, the decode lookahead and the corrected forecast on. That is the rerun the 2026-09-26 section waited for, and it stays below ~6, so the 24 to less than 48 GB range now rounds outward from 5.41 and 15.86: ~5–16 tok/s. The upper end is unchanged. That Mac read cold experts at 3.7 GB/s through the engine, about a third of the development Mac's engine rate, and other apps held memory during both of its runs; the range does not say which of these sets the gap. Corrections to the two sections above. The 64 GB M3 Max report gives a 512 GB SSD without saying whether it is internal, so only the M4 Max's reports are known to be from an internal SSD. C4's reason is weaker than "unlikely to close it": the M4 Max ran the same pre-0.2.19 forecast and still decoded faster, which points at the chip and SSD more than the release, but the two runs also differ in release, 0.2.18 against 0.2.22. And the 3.7 GB/s against 17.3 GB/s comparison set an engine read rate against a raw SSD figure; through the engine the development Mac read 11.5 to 12.6 GB/s (C7). ## Automatic context window: plans by Mac memory Weights-free checks and simulated `doctor` plans for the candidate that picks the context window for each Mac ([[records/plan/configurable-context-window-2026-09-06]]). Carlos asked on 2026-09-13 for auto to choose the best window for every memory tier and for `--max-context` to accept the model's 262,144 tokens, using best guesses from what the development Mac can measure. No model process ran for these plans, and nothing here is timed. **Rule.** Auto takes the largest of 32,768, 65,536, 131,072 and 262,144 tokens whose plan keeps speculative decoding and the decode lookahead as the 32,768-token plan has them, retains one complete conversation of that length, and adds at most 10% to the planner's estimated time for 2,000 prompt tokens and a 400-token reply. Tiers are judged on RAM and Metal working set. Live availability is applied at startup, which lowers the window rather than drop the draft head. Each candidate's automatic ceiling rises by that window's own charge. **Plans.** Candidate `3626ba67` built from `dfc9b36`, draft file available, availability equal to RAM, working set 75% of RAM: | Simulated RAM | Window | Target | Draft head | Experts per layer | Estimated request cost | Target and experts at 32,768 | Next window | |---|---|---|---|---|---|---|---| | 16 GB | 32,768 | 10.0 GB | off | 20 | +0.0% | same | 65,536: does not fit with one complete conversation retained | | 18 GB | 32,768 | 11.5 GB | off | 28 | +0.0% | same | 65,536: does not fit with one complete conversation retained | | 24 GB | 32,768 | 16.0 GB | off | 54 | +0.0% | same | 65,536: adds 17.7% to a typical request, above the 10% limit | | 32 GB | 32,768 | 22.0 GB | on | 74 | +0.0% | same | 65,536: turns speculative decoding off | | 36 GB | 65,536 | 25.0 GB | on | 75 | +8.9% | 25.0 GB, 96 | 131,072: turns speculative decoding off | | 48 GB | 65,536 | 33.6 GB | on | 140 | +2.3% | 33.6 GB, 161 | 131,072: adds 15.5% to a typical request, above the 10% limit | | 64 GB | 131,072 | 43.2 GB | on | 149 | +0.0% | 34.6 GB, 149 | 262,144: adds 17.8% to a typical request, above the 10% limit | | 96, 128 and 192 GB | 262,144 | 54.7 GB | on | 149 | +0.0% | 34.6 GB, 149 | model limit | The 8 GB simulation is refused, as before. Simulations at 40, 56, 72 and 80 GB, which match no current Mac, pick 65,536, 131,072, 262,144 and 262,144 tokens. **This Mac.** The development Mac reads 51.5 GB of RAM and a 40.2 GB working set. Auto picks 65,536 tokens at a 36.1 GB target with 158 experts per layer. Larger candidates: 65,536 chosen, adds 1.2% to a typical request; 131,072 rejected, adds 10.4% to a typical request, above the 10% limit; 262,144 rejected, turns speculative decoding off. **Checks.** On candidate `3626ba67`, T0 passed 38 of 38 checks and the planner gates passed 90 of 90, including the automatic section: each tier's window, quiet and busy starts, fixed caches and explicit windows. The development build of the same source bytes also passed the policy proxy with 964,209 Swift assertions, including case C23 for the automatic window, 130 of 130 context gates and all nine windows of the explicit-window matrix. The frozen default allocation stays byte-identical at an explicit 32,768 tokens on all four fixture tiers. **Limits.** These are allocation plans from M5 Pro-based estimates, not timed tiers. A Mac's marketed memory, Metal limit and live availability can differ from the decimal-GB inputs. The window charges are exact ledger arithmetic, but native capacity runs existed only through 65,536 tokens before this change. The representative request does not price long-prompt waits, which have no calibrated estimate above 128,256 tokens. ### Automatic context window: a 131,072-token read inside its plan The largest window the development Mac could hold on 2026-09-13 completed its whole prompt and reply inside the planner's memory ledger ([[sources/runs/2026/09/2026-09-13-context-capacity-131072-cold]]). It is native evidence for the 131,072-token automatic window of 64 GB Macs ([[records/measurements/automatic-context-window-plans-2026-09-13]]), taken at a much smaller target than those Macs use. **Capacity.** A 130,944-token synthetic prompt and a required 128-token reply filled the window at a 16 GB target, with the draft head, vision and prefix retention off. The sampled physical footprint peaked at 14.80 GB against the ledger's 15.00 GB expected peak. The run had no abort, pressure cancellation or runtime error. The frozen driver marked it failed because the Mac's global swap-in counter rose by 200 pages; under [[records/decisions/global-paging-is-diagnostic]] that is recorded, not a capacity failure. **Reading time.** The prompt took 38.1 minutes at 57.3 tok/s overall. Passes of 256 tokens slowed from 173 tok/s early in the prompt to 39 tok/s past 115,000 tokens, and the 128-token passes after 128,256 averaged 28.7 tok/s. The planner's anchors ignore position. For this schedule they estimate 6.4 minutes to 32,768 tokens, where 4.3 were measured, 12.9 to 65,536, where 12.5 were measured, and 25.1 for the 256-token part to 128,256, where 36.5 were measured. **Anchor decision.** A 128-token anchor was not added, although a preregistered rule allowed one. The anchor family is measured on an 8,016-token acceptance prompt at a matched cache and changes only with a complete envelope, and this synthetic filler read short-context 256-token passes at twice that family's rate. Waits above 128,256 tokens stay uncalibrated, and the public guides quote this run's 38 minutes as a measured reference instead. **Limits.** One cold text rung at one small target on one busy Mac, not the frozen campaign. Larger targets use larger early passes and a bigger cache, so this reading time does not predict a 64 GB Mac. No 262,144-token native run exists: every target from 18 to 24 GB was refused on this Mac at the time. The draft head at this window is recorded separately. ### Automatic context window: the draft head at 131,072 tokens With the draft head on, the 131,072-token window completed the same 130,944-token prompt and 128-token reply as the draft-off rung inside an 18 GB plan ([[sources/runs/2026/09/2026-09-13-context-draft-head-131072]]). This is the native evidence behind keeping the draft-head limit at the model limit for the automatic windows above 65,536 ([[records/measurements/automatic-context-window-plans-2026-09-13]]). **Parity.** The greedy reply matched the draft-off reply from [[records/measurements/automatic-context-window-131072-read-2026-09-13]] token for token. The head drafted 85 tokens and all 85 were accepted over 43 verify passes. The filler's continuation is highly predictable, so that acceptance rate does not describe real text. **Memory.** The sampled physical footprint peaked at 16.44 GB against the ledger's 17.00 GB expected peak and the 18 GB target, with no abort, pressure cancellation or runtime error. Global swap-ins rose by 229 pages, recorded as diagnostic under [[records/decisions/global-paging-is-diagnostic]]. **Limits.** One window, one small target and a head forced below its floor, so the decode lookahead was off. Reading the prompt took 40.3 minutes. No 262,144-token run exists with or without the draft head. ### v0.2.17 published, installed and accepted **v0.2.17 is published, installed and functionally accepted.** It ships the model's full 262,144-token limit and the per-Mac automatic window from [[records/decisions/automatic-context-window-per-machine]]. The exact CI artifact passed all twenty-five model gates, the published archive matched that artifact with a valid attestation, the installer replaced 0.2.16 on this machine, and the installed binary passed all thirty-one end-to-end release checks. Every user application stayed open throughout. Release: [v0.2.17](https://github.com/carloslfu/slotstream/releases/tag/v0.2.17), tagged on `d25dffff1f3f56e70ddb9f11b82e7dc1b843f041`, published 2026-09-13T19:49:26Z. Archive SHA-256 `66eb2ae95b325e75280d675fcf2afb3d61ef27db5500dea462faadef457b6042`; binary SHA-256 `1d761999c461c19237f4efa84947265c6253890ea57824ca62b7e2892e8f32aa`. The CI candidate, the re-downloaded public archive and the installed binary are byte-identical. ## What qualified | Phase | Result | |---|---| | Main CI 34775403549, commit `d25dfff` | Coverage, weights-free and public-library jobs all succeeded | | Candidate verification | Archive and binary digests recorded; 162 source files match the checkout | | Model acceptance on the downloaded CI binary | 25 of 25 gates, 0 failures, 1,065 seconds | | Release workflow 34778829372 | Succeeded; the public archive matches the CI artifact | | Provenance | `gh attestation verify` confirmed the archive for `carloslfu/slotstream` | | Installation | 0.2.16 replaced by 0.2.17, exit code 0, binary digest re-checked; 0.2.16 kept beside it | | Installed-release end-to-end | 31 of 31 checks, 0 failures, 72 seconds | ## The first acceptance run did not count The first run on the same artifact ended at 23 passed and 1 failed because the worktree it ran from had no mlx 0.31.1 reference environment, so vision parity could not run and the suite counted the skip as a failure. No product gate failed. The environment was linked in and the complete suite ran again on a fresh download of the same artifact, which passed all twenty-five gates. Both runs are in [[sources/runs/2026/09/2026-09-13-release-0-2-17-published-and-installed]]. ## The automatic window on the public binary At a 10 GB target with the draft head on, the installed server chose a 32,768-token window and printed the rule it applied: the largest of 32,768, 65,536, 131,072 and 262,144 tokens that keeps speculative decoding, retains one complete conversation and adds at most 10% to a typical request. This is the automatic window reaching a user through a published artifact rather than a local build. ## Limits No clean-timing or throughput qualification is claimed; the acceptance interval shared the machine with ordinary work. The installed-release checks ran at a 32,768-token window. The windows above that rest on the 131,072-token rungs ([[records/measurements/automatic-context-window-131072-read-2026-09-13]], [[records/measurements/automatic-context-window-draft-head-131072-2026-09-13]]) and the memory ledger; 262,144 tokens has not run natively. The open items in [[records/plan/configurable-context-window-2026-09-06]] are unchanged. ## Persistent prefix cache: exact restore, shared rows and restart resume at 10 GB **A persisted conversation state restores exactly, later turns write only their new rows, and a conversation resumes across real server restarts.** `serve --prefix-cache-dir` (off unless a directory is named) writes a text conversation's committed state after each reply of at least 2,048 tokens. A state is a head file (token ids, the fixed recurrent arrays, the draft head's last row) plus row segments; a state that descends from a persisted head references that head's segments and writes one new segment of only the rows added since. A save keeps the conversation's previous state, so the last reply can be regenerated after a restart, and removes older ones. A later prompt restores the longest persisted state it extends when memory retains nothing as long. Every payload and header carries CRC-32, and files belong to one identity: executable image, sampled weight content, cache geometry and optimization settings. The sources are `PersistentPrefixCache.swift`, `PersistentPrefixFormat.swift`, `PersistentPrefixPolicy.swift`, `PersistentPrefixSave.swift`, `PersistentPrefixRestore.swift` and `PersistentPrefixGenerator.swift`; [[records/design/measured-operating-policies]] classifies the tier's defaults. **Exact restore with shared rows.** In [[sources/runs/2026/09/2026-09-14-persistent-prefix-segments-exactness]], at 2,051 tokens and 640 pool slots, `optimization-state-check --variant persistent-prefix` passed 1,134 assertions and `persistent-prefix-mtp` passed 1,190, with none failed. The restored state matched the saved state in every representation field and in its next-token logits. A new cache over the same directory resumed the prompt plus a 259-token suffix with an empty memory tier, and its state and continued logits equaled the same continuation from an in-memory hit. That continuation was written as 122,365,681 and 123,001,483 new bytes while referencing 51,978,240 and 56,832,512 bytes of rows already on disk, against 167,765,980 and 172,642,364 bytes for the first save, and restoring it matched the memory continuation in every field and its logits. A regenerated reply resumed the kept parent exactly, and a request with `persistsPrefixState = false` restored but wrote nothing. Restores took 0.025 to 0.031 s. At this length most of a later write is the recurrent state every head carries; rows are the part that grows with the conversation. **Three turns and two restarts through serve.** In [[sources/runs/2026/09/2026-09-14-persistent-prefix-segments-e2e]], `serve --memory-gb 10` processes ran one after another; the plan retains up to 13,382 tokens in memory. Turn 1 read 3,849 prompt tokens in 42.2 s and wrote its 3,896-token state as 223.4 MB, a 115.7 MB head and a 107.7 MB segment. Turns 2 and 3 on the same server reused 3,896 and 3,968 tokens from memory and produced their first tokens at 1.69 and 1.79 s. Each wrote a new head and a segment of its new rows, 117.7 and 117.0 MB in 0.05 s, while referencing 107.7 and 109.7 MB of rows already on disk; turn 3 removed the turn-1 state and kept turn 2's, leaving two heads and 111.0 MB of rows in three segments. A new server over a copy of the directory taken after turn 2 restored the 3,968-token state (225.4 MB) from the segments turns 1 and 2 wrote in 0.04 s and produced its first token at 1.78 s, with turn-3 prompt ids (3,993, SHA-256 `e2165b2b927e0188…`) and output ids (22, `30eea6519ca6105e…`) identical to the first server's. Another server, over a copy taken after that restart, received turn 3 again as a regenerated reply: it restored the kept turn-2 state in 0.04 s, produced its first token at 1.68 s and returned the same output ids, and wrote nothing because the identical state was already on disk. `slotstream prefix-cache` listed that copy's two states and three segments (342.4 MB), and `--clear` removed all five files of another copy. A server with an empty directory read all 3,921 turn-2 tokens and produced its first token at 55.38 s. Peak process memory reported by the servers was 8.15, 6.59, 6.58 and 8.14 GB against the 10 GB target. The plan cached about 20 of 512 experts per layer, below the draft head's activation floor, so these requests did not speculate and their states carry no draft cache; `persistent-prefix-mtp` covers that path. **A conversation longer than memory retention.** In [[sources/runs/2026/09/2026-09-14-persistent-prefix-segments-long-conversation-e2e]], the same three turns ran over notes of 15,514 prompt tokens, longer than the 13,382 tokens the plan retains in memory, without a server that starts from an empty directory. Turn 1 read the prompt in 204.9 s and wrote its 15,561-token state as 546.0 MB in 0.18 s. Memory could not retain it, so turns 2 and 3 on the same server restored their previous states from disk in 0.08 and 0.07 s and produced their first tokens at 1.61 and 1.79 s. Each then wrote 117.7 MB in 0.03 s while referencing 430.2 and 432.2 MB of rows already on disk, where the earlier build rewrote the whole state, about 550 MB, every turn. After turn 3 the directory held two heads (231.5 MB) and three segments (434.2 MB). A restarted server over a copy taken after turn 2 restored the 15,632-token state in 0.09 s and produced its first token at 1.60 s, with turn-3 prompt ids (15,657, SHA-256 `0635963a10cadb70…`) and output ids (48, `9ce21908a4d4809e…`) identical to the first server's. Another server, over a copy taken after that restart, received turn 3 again as a regenerated reply: it restored the kept turn-2 state in 0.09 s, produced its first token at 1.63 s, returned the same output ids and wrote nothing. Peak process memory was 8.22, 6.68 and 6.69 GB. **An earlier build.** [[sources/runs/2026/09/2026-09-14-persistent-prefix-exactness]], [[sources/runs/2026/09/2026-09-14-persistent-prefix-restart-e2e]] and [[sources/runs/2026/09/2026-09-14-persistent-prefix-long-conversation-e2e]] measured the same tier before rows were shared, when every save rewrote the whole state and replaced the previous turn's file. They passed the same exactness and restart checks: 3,896- and 15,718-token states were written as 223.4 and 550.3 MB, restarted servers produced ids identical to the first servers', and first tokens came at 1.60 and 1.62 s against 50.62 and 210.53 s for servers with empty directories. **Limits.** These are single runs, neither paired nor interleaved, with every user application open. Reclaimable memory before each model process ranged from 24.9 to 29.6 GB. Timings describe this Mac's SSD and the fixture prompts. Every write still stores the fixed recurrent state, so a turn writes at least one head. The 20 GB quota, 2,048-token minimum and 30-day maximum age are provisional opt-in defaults, not measured optima. Removal order and maximum age are checked with synthetic states (`persistent-prefix-round-trip`), not under real traffic. Not covered here: requests with images (never written), many concurrent conversations, crash recovery, and other hardware. ## Shared prefixes: a system prompt reused across conversations, in memory and across restarts, at 10 GB **A system prompt processed once is reused by every later conversation that starts with it, in memory and across restarts, with identical outputs.** Issue #18 reported that `--prefix-cache-dir` states were per conversation: a second conversation with the same system prompt read it again. Both tiers were, because a state was written only at the end of a reply and reuse needs the state's ids to be a prefix of the new prompt. Now the prefill loop keeps a state at the last completed pass at or before two boundaries inside the prompt: where its system message ends (found from the template's `<|im_start|>system\n` and `<|im_end|>\n` ids, or named by `RequestController.sharedPrefixTokens`) and the longest head the prompt shares with a state a tier already holds, when either is 512 tokens or more beyond what the request reuses. The state is forked into the memory tier and, with a disk tier and `--prefix-cache-min-tokens` or more, written as a shared prefix. No pass is reshaped for it: the save point is the last existing pass end at or before the boundary, the 256-token grid by default. Shared prefixes are exempt from the ancestor removal a later save performs, evicted after conversations, listed by `slotstream prefix-cache`, and classed by lineages: a prefix two conversations start from is shared whatever wrote it. The sources are `Generate.swift` (the prefill loop), `PersistentPrefixGenerator.swift`, `PersistentPrefixPolicy.swift`, `PrefixCache.swift` and `Context.swift`; [[records/design/measured-operating-policies]] classifies the 512-token minimum and the pass-grid save point. **Exact reuse from memory and from disk.** In [[sources/runs/2026/09/2026-09-16-shared-prefix-exactness]], at 2,051 tokens and 640 pool slots, `optimization-state-check --variant shared-prefix` passed 745 assertions and `shared-prefix-mtp` passed 785, with 0 failed. A 2,051-id system message was kept at 2,048 ids during its own prefill, forked into memory and written to disk as a shared prefix in 0.196 s (172.3 MB; 0.085 s and 177.0 MB with the draft head), and the conversation's own state then referenced its rows. A second conversation with the same system message reused it from memory, and a fresh cache over the reopened directory restored it from disk in 0.036 s (0.029 s with the draft head); both continued with every state field and the next-token logits equal to a cold run of the same prompt. A sibling system message differing in its last 100 ids kept the 1,949-id head it shared at 1,792 ids, and a later prompt with the sibling resumed that head and matched the sibling's cold state. An explicit `sharedPrefixTokens = 700` on an unfinished system message kept 512 ids; a request with `persistsPrefixState = false` kept the shared prefix in memory but wrote nothing to disk. Two later turns of the first conversation replaced its first state while the three shared prefixes stayed, two classed shared and one parent. Resident memory at the end was 5.83 and 7.33 GB. **Through serve, across three restarts.** In [[sources/runs/2026/09/2026-09-16-shared-prefix-e2e]], `serve --memory-gb 10` processes ran one after another over one directory. Conversation A, a 3,663-token prompt whose system message ended 3,638 tokens in, kept it at 3,584 tokens during its prefill, the last 256-token pass end before the boundary, and wrote it as a shared prefix in 0.04 s (214.8 MB); its first token came at 37.33 s, and its own state then wrote 119.2 MB while referencing 99.1 MB of the shared prefix's rows. Conversation B, the same system prompt with another question, reused 3,584 tokens from memory and produced its first token at 3.05 s. A restarted server answered conversation C by restoring the 3,584-token shared prefix from disk in 0.03 s, first token at 4.39 s. Conversation D's system prompt S' shared its first 2,526 tokens with S: D kept that head at 2,304 tokens and its own system message at 3,584, writing both (330.5 MB in 0.14 s, the second referencing the first's rows) during a 51.85 s cold prefill. Another restarted server answered conversation E, with S', from the 3,584-token shared prefix in 0.04 s, first token at 2.59 s, and C's second and third turns restored C's states (1.10 and 1.19 s to the first token) and replaced C's first one while every shared prefix stayed: the final directory held 10 states, 3 of them shared prefixes (2,304, 3,584, 3,584 tokens), 11 segments, 1.37 GB. A server without a prefix cache read B's and C's prompts from scratch (21.23 and 45.47 s to the first token) and produced output ids identical to the reused runs (48 ids, SHA-256 `e256727d9d54b7e7…`, and 23, `59a097e2eff5eac5…`). 17 of 17 checks passed. Peak process memory reported by the servers was 8.11, 7.96, 7.17 and 8.63 GB against the 10 GB target. **Limits.** These are single runs, neither paired nor interleaved, with every user application open; reclaimable memory before each server ranged from 26.9 to 30.4 GB. Timings describe this Mac's SSD and the fixture prompts; the two cold prompts differed by more than twofold between servers, so first-token times are indicative, not a benchmark. The save point is a pass end, so a new conversation processes up to one pass of the system prompt again (54 tokens here). A shared prefix costs one head plus its rows on disk and one fork in memory; the 512-token minimum and the disk tier's 2,048-token minimum are provisional choices, not measured optima. Removal order with shared prefixes is checked with synthetic states (`persistent-prefix-policy`), not under real traffic. Not covered here: requests with images (never kept), prompts whose template renders the system header differently, many concurrent conversations, crash recovery, and other hardware. ### v0.2.18 published, installed and accepted **v0.2.18 is published, installed and functionally accepted.** It ships the persistent prefix cache: with `serve --prefix-cache-dir` a conversation's state is kept on disk, so a restarted server continues the conversation by restoring it instead of reading the whole prompt again ([[records/measurements/persistent-prefix-cache-2026-09-14]]). The exact CI artifact passed all twenty-five model gates on its second complete run, the published archive matched that artifact with a valid attestation, the installer replaced 0.2.17 on this machine, and the installed binary passed all thirty-one end-to-end release checks and the persistent prefix end-to-end check. Release: [v0.2.18](https://github.com/carloslfu/slotstream/releases/tag/v0.2.18), tagged on `829126e7b52c77981f5f02d6f7497d27262fe940`, published 2026-09-14T23:52:30Z. Archive SHA-256 `0e30342623f7140eba02046b6731699699d1f9c78f50b541dafe0f0e41113911`; binary SHA-256 `e8c77934be16007df99c8199163c0a3963f69b27dce01b70cf33793aeb04af4f`. The CI candidate, the re-downloaded public archive and the installed binary are byte-identical. ## What qualified | Phase | Result | |---|---| | Main CI 34894580252, commit `829126e` | Coverage, weights-free and public-library jobs all succeeded | | Candidate verification | Archive and binary digests recorded; 172 source files match the checkout | | Model acceptance on the downloaded CI binary | 25 of 25 gates, 0 failures, 1,131 seconds (second complete run) | | Release workflow 34910745261 | Succeeded; the public archive matches the CI artifact | | Provenance | `gh attestation verify` confirmed the archive, built by `release.yml` from `v0.2.18` at `829126e` | | Installation | 0.2.17 replaced by 0.2.18, exit code 0, binary digest re-checked; 0.2.17 kept beside it | | Installed-release end-to-end | 31 of 31 checks, 0 failures, 73 seconds | | Persistent prefix cache, installed binary | 11 of 11 checks; a restarted server restored 3,968 tokens from disk and matched the first server's prompt and output ids | ## The first acceptance run did not count The first complete run on the same artifact ended at 24 passed and 1 failed: the elastic drill read 13.9 GB reclaimable where it needs about 14 GB to demonstrate a shrink, so it skipped, and the suite counts a skipped required gate as a failure. No product gate failed. The complete suite ran again on the same artifact and passed all twenty-five gates. Both runs are in [[sources/runs/2026/09/2026-09-14-release-0-2-18-published-and-installed]]. ## Limits No clean-timing or throughput qualification is claimed; acceptance shared the machine with ordinary work and waited for memory headroom before it started. The installed-release checks ran at a 32,768-token window and a 10 GB target. The installed persistent prefix check used a 10 GB target and skipped its cold baseline, which the development measurement covers. ## Decode forecast taps: a more accurate forecast at the same lead time **Outcome: a more accurate expert forecast at the same lead time raises decode throughput by 11.1% over the configuration 0.2.16 to 0.2.18 shipped, with identical outputs, and becomes the 0.2.19 default.** The shipped decode lookahead forecast a layer's routing from the model's state two layers back. Reading the state after the previous layer's attention step instead, at the same point in time, raised top-10 forecast agreement from 0.62 to 0.73; a rank-128 linear correction of that forecast's router logits, fitted on the benchmark corpus's training requests, raised it to 0.80. On eight held-out prompts at a 20 GB target the corrected forecast decoded at 1.111 against the shipped configuration (bootstrap 1.102 to 1.123), all 23 counted pairs above 1, reading 21% fewer expert records during decode and wasting 73% fewer speculative bytes ([[records/measurements/decode-forecast-taps-readout-timing-confirmation-2026-09-15]]). The plan is [[records/plan/decode-forecast-taps-2026-09-14]]; the default change is [[records/decisions/corrected-decode-forecast-default-with-the-sidecar]]. Every step was registered before its data existed and read under a fixed gate. The sections below record them in dependency order: the offline gate and the native screen of the attention taps, the B1 confirmation that passed every effect condition but one hygiene condition, a co-routing prior that failed offline, the learned correction offline and natively, its screen and held-out confirmation, the issue-threshold screen that closed at 0.062, the attention readout that reached 0.86 agreement offline but lost natively, and the timing confirmation that produced the default's number. Four levers closed on their registered readings (the shared-expert variant by rule, the co-routing prior, the lower threshold and the readout); two passed (the attention tap and its learned correction) and ship together. Limits that apply to every section: one 48 GB M5 Pro, one checkpoint, the 20 GB profile with two drafts unless a section says otherwise, no multilingual prompt in the later held-out sets, and a correction trained on the pilot's 56 training requests. The prefetch twin projects read coverage and never counts in-flight joins or GPU cost, so its ratios are never claims; the numbers that stand are the native paired measurements. ### Decode forecast taps, steps 1 and 2: the attention tap passes its offline gate and the 18 GB screen **Outcome: forecasting from the streams after the previous layer's attention, at the same lead time as the shipped forecast, raises offline top-10 agreement from 0.6171 to 0.7292, and the native screen at 18 GB decodes 1.054 over the qualified configuration on the first twelve counted pairs with identical outputs, so the B1 confirmation runs.** Steps 1 and 2 of [[records/plan/decode-forecast-taps-2026-09-14]]. **Where the forecast reads.** The shipped forecast for target layer T applies T's own hyper-connection read and router to the streams after layer T-2's expert add, and reaches the scheduler on layer T-1's routing readback. At that readback the streams already hold T-1's attention output, so a forecast built on them (the `attention` tap) arrives just as early and misses only T-1's routed experts and T's attention; `attention-shared` adds T-1's resident shared expert. The capture command records any tap without a scheduler, and `SLOTSTREAM_EXPERT_PREFETCH_TAP` selects one for the router policy. **Offline gate** ([[sources/runs/2026/09/2026-09-15-forecast-taps-offline-stage]]): the 13 pilot validation requests captured at a 10 GB target with every tap as an observer, 141,450 rows over targets 2 to 47. | forecast | top-10 agreement | exact top-10 | recall at 16 | twin coverage at the shipped traffic | wasted reads | | --- | ---: | ---: | ---: | ---: | ---: | | boundary, stride 2 (shipped) | 0.6171 | 0.0223 | 0.7389 | 0.4389 | 136,416 | | attention tap | 0.7292 | 0.0530 | 0.8565 | 0.5393 | 102,443 | | attention tap with the shared expert | 0.7398 | 0.0603 | 0.8655 | 0.5492 | 99,104 | | boundary, stride 1 (one attention block late) | 0.7619 | 0.0885 | 0.8819 | 0.5785 | 89,203 | The gate asked for agreement at least 0.03 above stride 2 and matched-traffic coverage at least 0.03 above the shipped setting with no more wasted reads; the attention tap passes both by a wide margin. The shared-expert variant added 0.0099 of coverage against the registered 0.01 margin, so the plain attention tap went forward by rule. Adding the tap to the stride-2 forecast was below the tap alone at matched traffic. **Native screen at 18 GB** ([[sources/runs/2026/09/2026-09-15-forecast-taps-screen-18gb]]): the qualified configuration set explicitly (`combined`) against the same with each tap, on r0005, r0206, r0096 and r0074 at 256 outputs, interleaved. The 20 GB profile needed 25 GB reclaimable and 23.65 GB was available, so the screen ran at 18 GB as a mechanism screen. A first attempt stopped when `docs/LIBRARY.md` had been edited since the corpus froze; the corpus tool now reads every code and prose source from the git blob recorded at freeze. Another session compiled and ran its app on the same Mac, so a contention rule was registered before the screen resumed: a sampler records compiler, linker and engine processes every 5 s, a cell whose window holds one is excluded, and rounds are added until each arm has 12 counted pairs. Four rounds, 48 cells, every output identical to its reference. | configuration | counted cells | median tok/s | paired ratio (all pairs) | first 12 pairs | above 1 of 12 | | --- | ---: | ---: | ---: | ---: | ---: | | combined (reference) | 15 | 11.26 | | | | | attention tap | 15 | 11.97 | 1.047 (14) | 1.054 | 11 | | attention tap with the shared expert | 14 | 12.15 | 1.077 (14) | 1.078 | 11 | The reading needed a paired ratio of at least 1.01 with at least 8 of the first 12 pairs above 1 for the tap chosen offline; the attention tap passes. Mechanism, medians over pairs: records read in decode 0.869, reads issued 0.898, adopted 1.140, expired unused 0.563, wasted bytes 0.562. A screen decides only whether the confirmation runs; its ratios are not claims. ### Decode forecast taps, step 3: B1 at 1.058 against the qualified configuration, not passed as registered **Outcome: on the twelve held-out B1 prompts at 20 GB the attention tap decodes at 1.058 against the qualified configuration (bootstrap 1.031 to 1.088), every family at 1.019 or above, every output identical, but the run does not pass as registered: host swap-outs excluded six pairs and left r0033 with one counted pair, which fails the pairs condition.** Step 3 of [[records/plan/decode-forecast-taps-2026-09-14]]; run [[sources/runs/2026/09/2026-09-15-forecast-taps-b1-cohort-20gb]]. **Registration.** Written 36 seconds after the screen launched and before its first cell finished, the B1 gate for one change inside the qualified configuration: outputs identical in every pair; aggregate paired ratio (median of rounds per prompt, geometric mean per family, equal family weight) at least 1.02; bootstrap 2.5th percentile above 1.00; no family below 0.97; no family duration regression above 2%; at least two clean pairs per prompt. The hygiene rule added before the screen was analyzed counts a pair only when the cohort tool counts it under process-pageins-v1 and neither arm's window holds a compiler, linker or second engine. **Run.** Three rounds of twelve prompts at 512 outputs with a 128-token warmup, alternating arm order, 02:35 to 04:26 on 2026-09-15 with 26.2 GB reclaimable at launch. No arm window held a compiler, linker or second engine; the Sevra app peaked at 5.9% CPU, so the sensitivity reading equals the rule as written. | family | tok/s ratio | duration ratio | prompt medians (counted pairs) | | --- | ---: | ---: | --- | | code | 1.059 | 0.931 | r0062 1.049 (3), r0033 1.070 (1) | | dialogue | 1.028 | 0.976 | r0296 1.028 (3), r0295 1.028 (3) | | multilingual | 1.019 | 0.959 | r0244 1.066 (3), r0245 0.974 (3) | | prose | 1.105 | 1.005 | r0171 1.040 (3), r0173 1.173 (2) | | reasoning | 1.070 | 0.941 | r0124 1.088 (3), r0125 1.052 (2) | | structured | 1.068 | 0.955 | r0256 1.065 (2), r0257 1.072 (2) | | condition | required | result | | --- | --- | --- | | identical outputs | every pair | 36 of 36 | | aggregate ratio | at least 1.02 | 1.0578 | | bootstrap lower bound | above 1.00 | 1.0311 (upper 1.0880) | | family floor | at least 0.97 | 1.019 (multilingual) | | family duration | at most 1.02 | 1.005 (prose) | | counted pairs per prompt | at least 2 | r0033 has 1 | Median decode throughput over counted pairs was 12.17 tok/s for the qualified configuration and 12.98 with the tap. Paired medians, tap over qualified: records read in decode 0.878, reads issued 0.880, adopted 1.123, expired unused 0.569, wasted bytes 0.568, forecast evaluation seconds 0.545. **Standing.** Every effect condition passes and the mechanism matches the screen, but the registration makes a prompt with one counted pair a failed run, so this measurement changed no default and supports no public number on its own. The tap's default evidence came later, directly, from [[records/measurements/decode-forecast-taps-readout-timing-confirmation-2026-09-15]], which measured the corrected tap against the same qualified configuration on fresh prompts. ### Decode forecast taps, step 5: the co-routing prior fails its offline gate **Outcome: negative. A co-routing prior built from the previous layer's routed experts, which the host holds when the attention tap's forecast arrives, raises leave-one-request-out top-10 agreement by 0.0032 and matched-traffic twin coverage by 0.0094 against registered bars of 0.03 each, so no native diagnostic was built.** Step 5 of [[records/plan/decode-forecast-taps-2026-09-14]]; run [[sources/runs/2026/09/2026-09-15-forecast-taps-coroute-prior]]. CPU only, run between two rounds of the taps screen so no timed cell overlapped it. A per-layer pointwise mutual information table over pairs (expert routed at T-1, expert routed at T) was counted on the pilot's 56 training requests (4,180 verify passes, 12,540 rows per layer, alpha 20 smoothing as registered) and used to rescore the tap's 24 recorded candidates on the 13 validation requests as margin plus beta times the summed PMI over T-1's ten routed experts. | beta | top-10 agreement | exact top-10 | recall at 16 | | ---: | ---: | ---: | ---: | | 0 (the tap) | 0.7292 | 0.0530 | 0.8565 | | 0.02 | 0.7323 | 0.0465 | 0.8612 | | 0.05 | 0.7019 | 0.0272 | 0.8421 | | 0.1 | 0.6652 | 0.0139 | 0.8223 | | 1 | 0.5869 | 0.0023 | 0.7817 | Leave-one-request-out chose beta 0.02 in every fold (0.7323); alpha 5 gives 0.7251 and alpha 80 gives 0.7346. In the twin at the shipped setting's 284,812 issued reads, coverage is 0.5397 for the tap and 0.5490 with the prior, with 3,165 fewer wasted reads. Agreement peaks at the smallest nonzero beta and falls below the tap from 0.05 on, and exact rows fall at every nonzero beta: consistent with T-1's routes carrying little about T's choices beyond what the tap already reads from the same streams, though this probe does not test that explanation. The next accuracy step was the learned correction. ### Decode forecast taps, steps 6 and 7: the learned correction passes offline and natively **Outcome: a per-layer ridge correction of the attention tap's router logits, truncated to rank 128 and fitted on the pilot's training requests, raises validation top-10 agreement from 0.7292 to 0.7980 and twin coverage at the shipped traffic from 0.5397 to 0.6209 with 27,467 fewer wasted reads, in 35.2 MiB of FP16 weights; the engine's native form reproduces it (top-10 sets equal in 99.90% of rows, agreement 0.79799), so timing runs followed.** Steps 6 and 7 of [[records/plan/decode-forecast-taps-2026-09-14]]; runs [[sources/runs/2026/09/2026-09-15-forecast-taps-learned-correction-offline]] and [[sources/runs/2026/09/2026-09-15-forecast-taps-learned-correction-native-diagnostic]]. **Offline.** A 10 GB capture recorded the tap's router input and every layer's true router input for the pilot's 56 training and 13 validation requests (5,428 passes, 244,635 forecast records, 7.88 GB of shards). Per target layer, a ridge regression on 12,540 training rows corrects the tap's logits, with lambda chosen by five-fold cross-validation grouped by request (factor 1 for 29 layers, 0.1 for 17, 0.01 for 1); the rank-128 truncation keeps 87% to 98% of each correction's squared singular values. | forecast, validation rows | top-10 agreement | exact top-10 | recall at 16 | twin coverage | wasted reads | FP16 weights | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | attention tap | 0.7292 | 0.0530 | 0.8565 | 0.5397 | 102,326 | | | full correction | 0.8023 | 0.1062 | 0.9195 | 0.6269 | 72,821 | 117.5 MiB | | rank-128 correction | 0.7980 | 0.1012 | 0.9163 | 0.6209 | 74,859 | 35.2 MiB | The registered gate asked for agreement at least 0.05 above the tap, coverage at least 0.05 above the tap with no more wasted reads, and at most 64 MiB; the rank-128 form passes all three (+0.0688, +0.0812), the full form fails only the memory bound. Every target layer gains, from +0.030 (T=9) to +0.241 (T=47). The factors are `tap-correction-attention-rank128-v1.safetensors` (37,540,708 bytes, SHA-256 `37b00d3a32d1e1889a1794bbb8e97905a157a77c0508db620c1a11f2a895f7f5`), the file 0.2.19 ships as a sidecar. **Native.** `RouterTapCorrection` loads a `slotstream-tap-correction-v1` safetensors file (FP16 factors a and b, FP32 mu and delta, I32 targets), checks schema, tap, dtypes, shapes, target window and finite values, and identifies it by the file's SHA-256. The `attention-corrected` tap shares the attention tap's mixed input and router product and adds ((mixed - mu) a) b + delta per target, the wide product in FP16 and the narrow one in FP32; its factors stay resident and join the lookahead reserve in whole MiB. A capture of the 13 validation requests recorded the plain and corrected taps side by side: over all 141,450 rows the native corrected top-10 set equals the offline one in 99.90% (bar 99%), native corrected agreement is 0.79799 (0.7980 within 0.005) and the plain tap 0.72918 (0.7292 within 0.002). The forecast tap check grew to cover parsing, refusals, the reserve, the formula against a hand reference and loading against in-memory factors. ### Decode forecast taps, step 8: the corrected tap screens at 1.036 over the attention tap **Outcome: the corrected tap screens at 1.036 over the plain attention tap on the exploration prompts at 20 GB, 9 of the first 12 counted pairs above 1, every output identical, so the held-out confirmation runs.** Step 8 of [[records/plan/decode-forecast-taps-2026-09-14]]; run [[sources/runs/2026/09/2026-09-15-forecast-taps-learned-screen-20gb]]. A smoke of one 32-output request per arm at 10 GB first showed both arms starting with identical outputs and the corrected arm's identity naming the file (`correction=37b00d3a32d1e188`). The screen froze at 20 GB with 30.4 GB reclaimable: `attention` against `attention-corrected`, which differs only in the tap, the correction file and the reserve (128 to 164 MiB), on r0005, r0206, r0096 and r0074 at 256 outputs, interleaved, under the contention rule. Host swap-outs made three r0074 cells unclean, so the rule added a fourth round; 32 cells, 29 counted, 13 pairs. | round | r0005 | r0074 | r0096 | r0206 | | --- | ---: | ---: | ---: | ---: | | 0 | 1.047 | excluded | 1.043 | 1.017 | | 1 | 0.996 | excluded | 1.042 | 1.024 | | 2 | 0.963 | excluded | 1.086 | 1.290 | | 3 | 0.928 | 1.040 | 1.001 | 1.033 | The first 12 counted pairs in round order give 1.036 with 9 above 1 (bars 1.01 and 8); all 13 give 1.036 with 10 above 1. Median tok/s over counted cells: 14.17 for the attention tap, 14.30 with the correction. Median counters per counted cell, attention then corrected: records read in decode 13,760 and 12,370, reads issued 23,870 and 21,020, adopted 15,950 and 17,530, expired unused 5,650 and 2,420, wasted bytes 15.6 GB and 6.7 GB. The screen decides only whether the confirmation runs; its ratios are not claims. ### Decode forecast taps, step 9: the learned correction passes its held-out confirmation at 1.031 over the attention tap **Outcome: on ten held-out prompts from corpus families no earlier run had used, the learned correction decodes at 1.031 over the plain attention tap (bootstrap 1.013 to 1.041), every kind at 1.025 or above, three counted pairs per prompt, every output identical: the confirmation passes every registered condition.** Step 9 of [[records/plan/decode-forecast-taps-2026-09-14]]; run [[sources/runs/2026/09/2026-09-15-forecast-taps-learned-confirmation-20gb]]. **Prompts.** The B0 and B1 sets had informed decisions and the correction was trained on the pilot's training requests, so the prompts come from the 36 training-split families that no capture, fit, screen or cohort had used (code 16, reasoning 8, prose 6, dialogue 3, structured 3; no multilingual family remains): for each kind, its sorted family names shuffled with a fixed seed, the first two families, and from each the request with the smallest id: r0026, r0050 (code), r0304, r0292 (dialogue), r0186, r0195 (prose), r0156, r0128 (reasoning), r0268, r0271 (structured). A scan of 16,983 run files found requests from these families only in a copy of the corpus file itself. **Run.** Three rounds at 512 outputs with a 128-token warmup at 20 GB (30.6 GB reclaimable at launch), the attention tap first and arm order alternating, 05:18 to 06:41 on 2026-09-15. No arm had a host swap-out, no arm window held a compiler, linker or second engine, and the Sevra app peaked at 6.1% CPU, so all 30 pairs counted under both readings. | kind | tok/s ratio | duration ratio | prompt medians (counted pairs) | | --- | ---: | ---: | --- | | code | 1.027 | 0.988 | r0026 1.047 (3), r0050 1.008 (3) | | dialogue | 1.026 | 1.012 | r0304 1.023 (3), r0292 1.028 (3) | | prose | 1.025 | 0.991 | r0186 1.020 (3), r0195 1.029 (3) | | reasoning | 1.025 | 0.977 | r0156 1.017 (3), r0128 1.034 (3) | | structured | 1.055 | 0.971 | r0268 1.042 (3), r0271 1.068 (3) | | condition | required | result | | --- | --- | --- | | identical outputs | every pair | 30 of 30 | | aggregate ratio, equal kind weight | at least 1.02 | 1.0314 | | bootstrap 2.5th percentile | above 1.00 | 1.0132 (median 1.0312, 97.5th 1.0409) | | kind floor | at least 0.97 | 1.0245 (prose) | | kind duration ratio | at most 1.02 | 1.012 (dialogue) | | counted pairs per prompt | at least 2 | 3 for every prompt | 29 of 30 pairs were above 1 (r0304 in round 2 at 0.910). Median decode throughput over the pairs was 14.64 tok/s with the attention tap and 15.14 with the correction. Paired medians, corrected over attention: records read in decode 0.907, decode seconds 0.972, reads issued 0.922, adopted 1.084, expired unused 0.525, wasted bytes 0.523, demand misses 0.903, forecast evaluation seconds 0.996. Over these pairs demand reads take about 30% of decode time; the correction cut records read in decode by 10% and demand-read time by about 7%, which is the 3.1%. **Standing.** This is the correction's own confirmation, against the attention tap and not against the shipped configuration; the direct measurement against the shipped configuration is [[records/measurements/decode-forecast-taps-readout-timing-confirmation-2026-09-15]]. Limits: one machine, one checkpoint, the 20 GB profile, no multilingual prompt, and a correction trained on the pilot's training requests. ### Decode forecast taps, step 10: a lower issue threshold does not earn a confirmation **Outcome: with the correction on, a lower issue threshold does not earn a confirmation. At 0.031 the screen reads 1.015 on the first twelve counted pairs (10 above 1) and 1.013 over 16, below the registered 1.02 bar; at 0.0 it reads 0.994: every top-10 read issued arrives late and wastes 2.4 times the bytes. The threshold stays at 0.062.** Step 10 of [[records/plan/decode-forecast-taps-2026-09-14]]; run [[sources/runs/2026/09/2026-09-15-forecast-taps-threshold-screen-16gb]]. **Why it was tried.** With the correction, issued reads are adopted at a median 0.83 against the attention tap's 0.70, and the twin's threshold curve for the corrected forecast gives coverage 0.543 at 0.062, 0.582 at 0.031 and 0.670 at 0.0 with no traffic bound, while decode serialization round 2 had found natively that extra speculative reads did not reduce demand reads for the boundary forecast ([[records/measurements/decode-path-serialization-round-2-2026-09-12]]). **Run.** Another session's virtual machine held 9 GB at launch, leaving 22.35 GB reclaimable, so the screen ran at 16 GB as a mechanism screen (an expert cache of 2,746 slots against 4,206 at 20 GB, hit rates about 0.57 against 0.70). Arms identical except the threshold: t062 (reference), t031 and t000, on r0005, r0206, r0096 and r0074 at 256 outputs under the contention rule, rounds until 12 counted pairs per candidate. The Sevra app took the model lock at 13:50 and the sweep stopped at its 35th cell; it resumed at 14:28 and ran a fourth round, ending 14:47. 48 cells, 47 counted (one host swap-out exclusion), every output identical. | candidate | counted pairs | first-12 ratio | above 1 of 12 | all-pairs ratio | above 1 | reading | | --- | ---: | ---: | ---: | ---: | ---: | --- | | t031 (0.031) | 16 | 1.015 | 10 | 1.013 | 13 | passes the screen reading, below the 1.02 confirmation bar | | t000 (0.0) | 15 | 0.994 | 6 | 0.992 | 6 | fails | Median counters per counted cell, t062 then t031 then t000: records read in decode 17,780, 15,960, 10,670; reads issued 32,270, 36,510, 44,080; expired unused 3,832, 5,356, 9,033; wasted bytes 10.6 GB, 14.9 GB, 25.1 GB; reads still in flight when demanded 4,610, 5,978, 10,702; join seconds 0.074, 0.107, 0.365. Issuing every candidate cuts demand reads by 40% and still slows decode: the extra reads arrive late (twice the in-flight reads at demand time, five times the join time). The middle threshold buys a small, consistent gain that shrank over the rounds (1.026 in round 0, 1.005 in round 3) and would likely be smaller at 20 GB, where demand reads are rarer. The lever closes at this result. ### Decode forecast taps, steps 11 and 12: the attention readout is exact and reads 0.86, its correction 0.88 **Outcome: running the target layer's own attention sublayer early, on the forecast's approximate input against the resident caches, is exact (self-check 0.9997) and lifts offline top-10 agreement to 0.8614, against 0.7980 for the corrected tap; a correction fitted on top reaches 0.8790, short of the registered 0.90 bar, so the readout lever stopped at the offline result until it was reopened for a timing run.** Steps 11 and 12 of [[records/plan/decode-forecast-taps-2026-09-14]]; runs [[sources/runs/2026/09/2026-09-15-forecast-taps-readout-diagnostic]] and [[sources/runs/2026/09/2026-09-15-forecast-taps-readout-learned-correction]]. **Why.** The stride-1 boundary forecast, which knows the previous layer's routed experts exactly and misses only the target's attention, reads 0.7619, while the corrected tap reaches 0.7980 without knowing those experts; so most of the remaining error is the target's attention over the context, and that state is resident. A nonlinear probe of the tap's input was worse than the ridge at every layer tried (it overfits families), so the attention was computed rather than predicted. **Implementation.** `attention-readout` (tap 4), at layer T-1's routing readback, takes T's attention-side read of the streams holding T-1's attention output, runs T's attention sublayer on it against T's caches without writing them (recurrence state, convolution window and KV cache are read only), injects the output into the streams, then T's mixed read and router. `boundary-readout` (tap 5) is the same readout on T's exact input, an observer-only self-check. Two follow-up builds added an `after-demand` placement that submits the readout's GPU work asynchronously at the readback and consumes it after the source layer's demand reads, and `attention-readout-corrected` (tap 6); the correction file's header names the tap it was fitted on and a file fitted on the other tap is refused. All three builds passed both check tiers. **Diagnostic**, the 13 validation requests at 10 GB, 141,450 rows over targets 2 to 47: | forecast | top-10 agreement | exact top-10 | recall at 16 | recall at 24 | | --- | ---: | ---: | ---: | ---: | | boundary, stride 2 (shipped) | 0.6171 | 0.0223 | 0.7389 | 0.8125 | | attention tap | 0.7292 | 0.0530 | 0.8565 | 0.9162 | | corrected tap | 0.7980 | 0.1012 | 0.9163 | 0.9585 | | readout tap | 0.8614 | 0.2055 | 0.9656 | 0.9881 | | readout self-check | 0.9997 | 0.9970 | 1.0000 | 1.0000 | The self-check is 0.99998 over the eleven requests within the 2,048-token indexer budget and 0.9947 on r0178 (5,231 prompt tokens), where the forward attends sparsely and the readout densely. The readout gains on every request and every layer group. The twin projects timely coverage 0.687 at the shipped traffic against 0.620 for the corrected tap, before the readout's own GPU work, which the twin cannot price. **Correction on the readout.** A 3,360-second capture of the 69 requests with the readout tap and its inputs, then the same collect, fit and twin as the attention tap's correction: the rank-128 form lifts the readout from 0.8614 to 0.8790 (bar 0.90) and twin coverage from 0.688 to 0.716 (bar +0.03) with 9,266 fewer wasted reads. With the target's attention computed early the correction has less left to learn (+0.018 against +0.069 on the attention tap); the remaining 0.12 of agreement is the previous layer's routed experts, which no forecast can know before reading them. As registered the lever stopped here; the timing run that followed was a separate decision with its own registration ([[records/measurements/decode-forecast-taps-readout-timing-screen-2026-09-15]]). ### Decode forecast taps, step 13: computing the next layer's attention early loses 5% to 26% **Outcome: negative. Over four rounds and 77 counted cells with identical outputs, no readout arm had a single pair above 1 against the corrected attention tap: 0.952 in the readback placement (with or without its correction), 0.737 and 0.742 in the after-demand placement, and 0.85 to 0.89 or 0.70 on the long observation prompts. The readout reads fewer records but its GPU work costs more than the reads it saves, and the after-demand placement synchronizes the GPU at every layer. The lever closes.** Section 1 of the readout timing registration, step 13 of [[records/plan/decode-forecast-taps-2026-09-14]]; run [[sources/runs/2026/09/2026-09-15-forecast-taps-readout-timing-screen-20gb]]. **Registration and run.** Written after the readout correction's offline gate failed and before any configuration existed, stating the smoke's prior (7.71 against 6.18 tok/s on one cell) and the twin's projections (1.290 and 1.307 against 1.249) so neither could be read post hoc. Five arms at 20 GB, identical to the learned confirmation's corrected arm except the tap, the placement, the correction file and the reserve: `corrected` (reference), `readout-readback`, `readout-after`, `readout-corrected-readback`, `readout-corrected-after`. r0005, r0206, r0096 and r0074 at 256 outputs, rounds until 12 counted pairs per arm; the observation cells r0178 (5,231 prompt tokens) and r0222 (about 7,000) after each sweep. Three cells were unclean for host swap-outs; no arm window held a compiler, linker or second engine. | arm | counted pairs | first 12 | above 1 | all pairs | median tok/s | r0178 | r0222 | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | corrected (reference) | | | | | 14.56 | | | | readout-readback | 15 | 0.949 | 0 | 0.952 | 13.67 | 0.888 | 0.849 | | readout-corrected-readback | 15 | 0.952 | 0 | 0.952 | 13.48 | 0.886 | 0.846 | | readout-after | 16 | 0.738 | 0 | 0.737 | 10.71 | 0.712 | 0.699 | | readout-corrected-after | 15 | 0.745 | 0 | 0.742 | 10.36 | 0.702 | 0.706 | **Why.** The readout does what the offline audit said: 8% fewer records read in decode (13% with its correction), half the speculative reads expiring unused, half the wasted bytes. It loses because its GPU work is not free. In the readback placement the forecast evaluation grows by 0.56 to 0.70 s and the build by 0.19 to 0.22 s per 256-output cell, about 0.8 s on a 17.6 s decode, and the routing readback it rides returns later, so the source layer's own demand reads start later. Demand reads take about 30% of decode time at this profile, so an 8% cut in records is worth about 2.4% of decode, less than the readout costs; at long context the readout attends densely on the twelve full-attention layers past the indexer budget and the loss grows to 11% to 15%. The after-demand placement submits the readout with `asyncEval` and consumes it after the source layer's experts; consumption materializes the logits, which waits for every queued GPU operation, so each layer ends with a full synchronization that the deferred four-layer barrier exists to avoid (7.6 s in the selection and scheduling timers per cell against 0.14 s and 0.08 s). The engine keeps the readout taps as observer and opt-in tools; nothing recommends them. ### Decode forecast taps, step 13: the corrected tap decodes 1.111 over the shipped configuration on held-out prompts **Outcome: the corrected attention tap decodes 11.1% faster than the configuration 0.2.16 to 0.2.18 ship, on eight held-out prompts from families no earlier run had used: aggregate 1.111 (bootstrap 1.102 to 1.123), all 23 counted pairs above 1, every kind at 1.079 or above, every kind's request duration shorter, outputs identical in 60 of 60 cells. Median decode throughput over counted cells rose from 13.10 to 14.83 tok/s. This is the measurement behind the 0.2.19 default.** Section 2 of the readout timing registration, step 13 of [[records/plan/decode-forecast-taps-2026-09-14]]; run [[sources/runs/2026/09/2026-09-15-forecast-taps-readout-timing-confirmation-20gb]]. **Registration.** With no readout arm qualifying at the screen, the confirmation ran its registered default-evidence contrast: `corrected` (the learned confirmation's corrected arm: tap attention-corrected with the rank-128 file, reserve 164 MiB) against `qualified` (the B1 plan's combined arm, the shipped boundary forecast at stride 2), nothing else differing. Prompts by rule over the 26 training-split families no capture, fit, screen or cohort had used: r0058 and r0067 (code), r0289 (dialogue), r0189 and r0211 (prose), r0092 and r0132 (reasoning), r0265 (structured); dialogue and structured hold one prompt each, a stated limit. Reading: outputs identical in every cell; aggregate paired ratio (median of rounds per prompt, geometric mean per kind, equal kind weight) at least 1.02; bootstrap 2.5th percentile above 1.00; no kind below 0.97; no kind duration ratio above 1.02; at least two counted pairs per prompt; under the contention rule as written and the Sevra sensitivity. **Run.** A first launch stopped before any configuration on a launcher defect (the screen's "no candidate" passed as the word none) and its directory was set aside unused. The corrected launcher froze the protocol at 20 GB with 26.4 GB reclaimable (about 87 experts per layer, 4,193 slots) on the readout build. 48 cells at 512 outputs with 128 warmup tokens, arm order rotating every cell, 21:26 to 22:36 on 2026-09-15, then the observation cells r0178 and r0222 to 23:13. One cell was unclean for host swap-outs (r0289 round 2, corrected); no arm window held a compiler, linker or second engine. | kind | tok/s ratio | duration ratio | prompt medians (counted pairs) | | --- | ---: | ---: | --- | | code | 1.112 | 0.928 | r0058 1.115 (3), r0067 1.110 (3) | | dialogue | 1.079 | 0.937 | r0289 1.079 (2) | | prose | 1.096 | 0.956 | r0189 1.089 (3), r0211 1.104 (3) | | reasoning | 1.134 | 0.889 | r0092 1.137 (3), r0132 1.130 (3) | | structured | 1.136 | 0.903 | r0265 1.136 (3) | | condition | required | result | | --- | --- | --- | | identical outputs | every cell | 60 of 60 | | aggregate ratio, equal kind weight | at least 1.02 | 1.1114 | | bootstrap 2.5th percentile | above 1.00 | 1.1017 (median 1.1127, 97.5th 1.1232) | | kind floor | at least 0.97 | 1.079 (dialogue) | | kind duration ratio | at most 1.02 | 0.956 (prose) | | counted pairs per prompt | at least 2 | 2 for r0289, 3 for the rest | All 23 pairs were above 1, from 1.048 to 1.169, identically under the Sevra sensitivity. Paired medians, corrected over qualified: records read in decode 0.794, decode seconds 0.897, request seconds 0.918, prefill seconds 0.998, reads issued 0.796, adopted 1.283, expired unused 0.275, wasted bytes 0.274, demand misses 0.838, forecast evaluation seconds 0.523. Observation cells, outside the reading: r0178 ran 1.067, 0.994 and 1.088 over the three rounds and r0222 1.073 and 1.068 in rounds 0 and 1; the round-2 corrected cell on r0222 did not complete because the engine's memory-pressure guard stopped its prefill commit while the host swapped pages in, the guard doing its job under a host event, recorded and excluded. **Standing.** This is the direct measurement the default decision needed, consistent with and replacing the product of the two earlier links (1.058 for the tap on B1, 1.031 for the correction over the tap). It became the default in 0.2.19 ([[records/decisions/corrected-decode-forecast-default-with-the-sidecar]]); the shipping build's own default-path benchmark is recorded separately. Limits: one machine, one checkpoint, the 20 GB profile, no multilingual prompt, one dialogue and one structured family, observation prompts of at most about 7,000 tokens, a correction trained on the pilot's training requests. ### Corrected forecast default: checks, plans by target, the sidecar and a default-path smoke **Outcome: the 0.2.19 default is implemented, checked and smoked on the shipping build. The engine locates the checkpoint's correction file next to the weights, runs the corrected attention forecast and charges 409 MiB; without the file, or with `SLOTSTREAM_EXPERT_PREFETCH_TAP=boundary`, it runs the 0.2.16 forecast at 373 MiB. `pull` fetches the file from the public mirror and verifies it, and one 32-output request per arm at 22 GB produced identical outputs with the expected identities and charges.** The decision is [[records/decisions/corrected-decode-forecast-default-with-the-sidecar]]; run [[sources/runs/2026/09/2026-09-16-corrected-forecast-default-checks-sidecar-smoke]]. This record prices where the default engages and shows the wiring works; the throughput number is [[records/measurements/corrected-forecast-release-benchmark-2026-09-16]]. **What changed.** `RouterTapCorrection.shipped` locates `lookahead/tap-correction-attention-rank128-v1.safetensors` in the model directory and qualifies it only when its header serves the corrected attention tap and its whole-file SHA-256 equals the measured `37b00d3a32d1e1889a1794bbb8e97905a157a77c0508db620c1a11f2a895f7f5`; `ExpertPrefetchConfiguration.qualifiedDecode(correction:)` is the qualified default with the corrected tap, the file and a reserve of 128 MiB plus the file in whole MiB, and the previous default when nothing is located. The planner carries a matching `automaticCorrected(bytes:)` case: `DecodeLookahead.reserveBytes(correctionBytes:)` charges 373 MiB plus the file rounded up to 36 MiB, 409 MiB, before the expert pool is sized, and the engine's guard that the plan's reserve covers the configuration's holds. `TapCorrectionSidecar` pins the file's size, digest and mirror commit (`8c1f9c34e4567e83d46cebe1af432e8eba4f3ea8`); `pull` fetches it after the weights when it is absent or wrong, writes it atomically, and reports and continues on any failure, and `pull --verify` reports its status. The startup banner names the forecast in use and the reason. **Checks.** The shipping tree (git tree `511d6c6a0592d9518aa54b8f859bcc8968a46336` for the sources, built in an isolated export so no other session's uncommitted hunks were in it) passes 55 of 55 T0 and T1 checks with 30,400 assertions, including `expert-lookahead-forecast-tap` (65 assertions: the taps, the readout refusals, the correction file's location, a wrong digest, a readout-fitted file at the path, the resulting configuration and the boundary override) and `decode-lookahead-defaults` (39 assertions, with the 409 MiB charge at a 22 GB plan and the `automaticCorrected` ledger). The binary is `157eb4ab0366c6a7b54f9ffd47edd416d10017abf9ccc9a911614c2ec0266b42` and reports 0.2.19. **Where the default engages**, `doctor --mtp on` under the benchmark environment (prefix cache off, draft depth 2) on the development Mac: | target | experts per layer | lookahead | charge | | ---: | ---: | --- | ---: | | 10 GB | ~17 | off (below the head's floor) | | | 20 GB | ~74 | off (below the head's floor) | | | 22 GB | ~100 | on, corrected forecast | 409 MiB | | 22 GB with `SLOTSTREAM_EXPERT_PREFETCH_TAP=boundary` | ~101 | on, 0.2.16 forecast | 373 MiB | So the default engages where the 0.2.16 lookahead did, from the head's 76-per-layer floor ([[records/measurements/decode-lookahead-default-2026-09-13]]), 32 GB Macs and up at the default context; the file costs about three experts per layer of cache. **Sidecar, end to end.** The public mirror serves the file at the pinned commit with `x-linked-size` 37,540,708 and an ETag equal to the digest; a plain download hashed to the digest. With the shipping binary, `pull --verify` on the installed model passed all 25 files in 8.1 s and reported the sidecar present. With the file moved aside, `pull` verified the weights, fetched the 37,540,708 bytes from the mirror, verified the digest and wrote the file in 43.5 s; `pull --verify` then reported it present again. **Smoke.** A first smoke at 10 GB ran both arms to completion with identical outputs and no lookahead in either, because at that target the cache is below the head's floor; it is kept as `xla3-ship-smoke-10gb-no-lookahead` and is why the release benchmark's profile was amended to 22 GB before any timed cell. At 22 GB, one 32-output request (r0005, 16 warmup tokens) per arm: `previous` (the boundary override) ran with identity `router-reuse:strides=2`, the banner `373 MiB, charged above` and `boundary forecast selected by SLOTSTREAM_EXPERT_PREFETCH_TAP`; `default` (nothing set) with identity `router-reuse:tap=attention-corrected:correction=37b00d3a32d1e188`, the banner `409 MiB, charged above` and `corrected attention forecast: measured correction 37b00d3a32d1e188 at lookahead/tap-correction-attention-rank128-v1.safetensors`. Both exited 0 with the same 32 output tokens; no model process remained. One 32-token cell per arm is not a comparison. **Limits.** Weights-free checks and planner arithmetic for the charge; the smoke is functional only. The default's throughput on the shipping build is the release benchmark's, at 22 GB; the 20 GB confirmation used environment-configured arms. ### Corrected forecast release benchmark: the shipping build's default against the 0.2.18 forecast at 22 GB **Outcome: the shipping build's default decodes 1.10x faster than the 0.2.18 forecast on eight held-out prompts at a 22 GB target: aggregate 1.108 (bootstrap 1.092 to 1.147), 24 of 24 counted pairs above 1, every kind at 1.079 or above, outputs identical in every cell, median decode throughput 14.38 to 15.86 tok/s. The registered gate passes under the rule as written and the Sevra sensitivity.** Registered in `release-benchmark-preregistration.md` (sha256 `aeb81d63787595c365c2555de161ec9579216902f40de015afac46fc192dd657`) before any timed run of the shipping build, with a profile amendment from 20 to 22 GB written after the 10 GB smoke and before any timed cell; run [[sources/runs/2026/09/2026-09-16-corrected-forecast-release-benchmark-22gb]]. This is the number the README states for 0.2.19; the decision is [[records/decisions/corrected-decode-forecast-default-with-the-sidecar]]. **Arms.** Both on the shipping binary (`157eb4ab0366c6a7b54f9ffd47edd416d10017abf9ccc9a911614c2ec0266b42`, sources at git tree `511d6c6a0592d9518aa54b8f859bcc8968a46336`) with the protocol's pinned environment and nothing else: `default`, no prefetch variable, which located `lookahead/tap-correction-attention-rank128-v1.safetensors` next to the weights and ran the corrected attention forecast with 409 MiB charged (identity `router-reuse:tap=attention-corrected:correction=37b00d3a32d1e188`); `previous`, `SLOTSTREAM_EXPERT_PREFETCH_TAP=boundary`, the 0.2.18 forecast with 373 MiB charged (identity `router-reuse:strides=2`). The 22 GB target, a 32 GB Mac's automatic target, is the smallest whole-GB target at which the automatic plan runs the lookahead under the benchmark environment (prefix cache off, two drafts): at 20 GB the cache holds about 74 experts per layer, below the head's 76-per-layer floor, so the original 20 GB registration could not exercise the default path ([[records/measurements/corrected-forecast-default-2026-09-16]]). **Run.** The eight prompts the readout timing registration drew by rule (r0058, r0067, r0289, r0189, r0211, r0092, r0132, r0265), 3 rounds at 512 outputs with a 128-token warmup, arm order rotating every cell, protocol `xla3-release-bench-22gb` frozen at 22 GB with 32.5 GB reclaimable, process-pageins-v1 and the contention rule. Excluded cells: none. Observation cells r0178 and r0222 belong to a separate sweep outside the reading; they did not run: after the sweep ended at 09:39:43, reclaimable memory stayed near 22.6 GB, below the launcher's 27.5 GB bar, until the wait was stopped at 12:27 with no model process; they are outside the reading and remain to be run as a separate observation when memory allows. | kind | tok/s ratio | duration ratio | prompt medians (counted pairs) | | --- | ---: | ---: | --- | | code | 1.105 | 0.925 | r0058 1.098 (3), r0067 1.113 (3) | | dialogue | 1.079 | 0.933 | r0289 1.079 (3) | | prose | 1.096 | 0.948 | r0189 1.098 (3), r0211 1.095 (3) | | reasoning | 1.133 | 0.893 | r0092 1.128 (3), r0132 1.137 (3) | | structured | 1.126 | 0.915 | r0265 1.126 (3) | | condition | required | result | | --- | --- | --- | | identical outputs | every cell | yes | | aggregate ratio, equal kind weight | at least 1.02 | 1.1077 | | bootstrap 2.5th percentile | above 1.00 | 1.0919 (median 1.1137, 97.5th 1.1469) | | kind floor | at least 0.97 | 1.079 | | kind duration ratio | at most 1.02 | 0.948 | | counted pairs per prompt | at least 2 | r0058 3, r0067 3, r0092 3, r0132 3, r0189 3, r0211 3, r0265 3, r0289 3 | Counted pairs ranged from 1.040 to 1.297; the Sevra sensitivity reading gives aggregate 1.1077 over 24 pairs. Paired medians, default over previous: decode_read_bytes 0.805, decode_records 0.805, decode_seconds 0.903, hit_rate 0.999, prefetch.adopted 1.262, prefetch.demandMisses 0.851, prefetch.expired 0.262, prefetch.forecastBuildSeconds 1.353, prefetch.forecastEvalSeconds 0.534, prefetch.forecastSelectSeconds 1.740, prefetch.issued 0.765, prefetch.joinSeconds 0.916, prefetch.promoted 1.135, prefetch.wastedBytes 0.257, prefill_seconds 1.006, request_seconds 0.930. **Standing.** The public number for 0.2.19 rounds the aggregate down to 1.10x, reported with the medians 14.38 and 15.86 tok/s. The 20 GB confirmation with environment-configured arms on the readout build ([[records/measurements/decode-forecast-taps-readout-timing-confirmation-2026-09-15]]) read 1.111 and 13.10 to 14.83 tok/s; it is reported apart and is not comparable cache for cache. Limits: one 48 GB M5 Pro, one checkpoint, the 22 GB profile with two drafts, no multilingual prompt, one dialogue and one structured family, prompts used once before for the same contrast, and the shared machine's ordinary applications open (the contention rule excluded cells with a compiler, linker or second engine in their window). ### v0.2.19 published, installed and accepted **v0.2.19 is published, installed and functionally accepted.** It ships the corrected expert forecast as the decode lookahead's default, with the 37,540,708-byte correction sidecar that `pull` fetches next to the weights ([[records/decisions/corrected-decode-forecast-default-with-the-sidecar]]); its throughput number is the registered release benchmark ([[records/measurements/corrected-forecast-release-benchmark-2026-09-16]]). The exact CI artifact passed all 25 model gates, the published archive matched that artifact with a valid attestation, the installer replaced 0.2.18 on this machine, and the installed binary passed all 31 end-to-end release checks and the 11 persistent prefix checks; `doctor --memory-gb 22` on the installed binary plans the lookahead with the corrected forecast. Release: [v0.2.19](https://github.com/carloslfu/slotstream/releases/tag/v0.2.19), tagged on `95e21252192470857e7a89019e0987b6682453c4`, published 2026-09-16T19:39:38Z. Archive SHA-256 `910a0a82f0406e66aed131f1d9edc483a2a35f021cade86545b032a5eb10d05c`; binary SHA-256 `d26b529f5e6cb43288dc9d2f633ce79d7fc0c877daf0e01d7cc478670279bc96`. The CI candidate, the re-downloaded public archive and the installed binary are byte-identical. ## What qualified | Phase | Result | |---|---| | Main CI 35136690124, commit `95e2125` | Succeeded | | Candidate verification | Archive and binary digests recorded; 174 source files match the checkout | | Model acceptance on the downloaded CI binary | 25 of 25 gates, 0 failures, 1163 seconds | | Release workflow 35141835417 | Succeeded; the public archive matches the CI artifact | | Provenance | `gh attestation verify` confirmed the archive, built by `release.yml` from `v0.2.19` at `95e2125` | | Installation | 0.2.18 replaced by 0.2.19, exit code 0, binary digest re-checked | | Installed-release end-to-end | 31 of 31 checks, 0 failures, 85 seconds | | Persistent prefix cache, installed binary | 11 of 11 checks; a restarted server restored 3968 tokens from disk with identical ids | | Installed default at 22 GB | `doctor` plans `lookahead: on` with the 409 MiB charge and the corrected forecast; `pull --verify` reports the sidecar present and verified | ## Limits No clean-timing or throughput qualification is claimed here; acceptance shared the machine with ordinary work and waited for memory headroom before it started. The installed-release checks ran at a 32,768-token window and a 10 GB target, where the automatic plan runs without the lookahead. The installed persistent prefix check used a 10 GB target and skipped its cold baseline, which the development measurement covers. ### v0.2.20 published, installed and accepted **v0.2.20 is published, installed and functionally accepted, and Codex runs on the installed release.** The release serves the OpenAI Responses API at `POST /v1/responses`, the protocol Codex requires for a custom provider, and adds the additive embedding APIs the Sevra for Mac development app uses. The exact CI artifact passed all 25 model gates and a Codex and SDK check before the tag. The published archive matched that artifact with a valid attestation, and the installer replaced 0.2.19 on this machine. The installed binary passed all 31 end-to-end release checks, the 11 persistent prefix checks, and 18 of 18 Codex and SDK checks. Release: [v0.2.20](https://github.com/carloslfu/slotstream/releases/tag/v0.2.20), tagged on `f92021ac371c5b65a87a20541e03abe85b7159e1`, published 2026-09-17T03:02:26Z (September 16 in Bogota). Archive SHA-256 `54ba067bcfb2abec134a5329598451cd09155478057e13ac0ee809d46e50961b`; binary SHA-256 `08fd86d3ab2067ef041456f6b165765455eda21494aa616e8e0e634195bebda1`. The CI candidate, the re-downloaded public archive and the installed binary are byte-identical. Evidence: [[sources/runs/2026/09/2026-09-16-release-0-2-20-published-and-installed]] and [[sources/runs/2026/09/2026-09-16-codex-on-installed-0-2-20]]. ## What qualified | Phase | Result | |---|---| | Main CI 35153305706, commit `f92021a` | Succeeded, including the coverage ratchet; `context-proxies` and `docs` succeeded on the same commit | | Candidate verification | Digests recorded; 178 source files match a clean export of the commit | | Model acceptance on the downloaded CI binary | 25 of 25 gates, 0 failures | | Codex and SDK on the CI binary, before the tag | 9 of 9 checks; the SDK part passed 11 of 11 | | Release workflow 35176627953 | Succeeded; the public archive matches the CI artifact | | Provenance | `gh attestation verify` returned SLSA provenance v1 | | Installation | 0.2.19 replaced by 0.2.20, exit code 0, binary digest re-checked | | Installed-release end-to-end | 31 of 31 checks | | Persistent prefix cache, installed binary | 11 of 11 checks; a restarted server restored 3968 tokens from disk with identical ids | | Installed default at 22 GB | `doctor` plans the corrected forecast lookahead with its 409 MiB charge; `pull --verify` passes with the sidecar present | | Codex and SDK on the installed binary | 18 of 18 checks | ## Codex on the installed release Codex 0.148.0 followed `docs/CODEX.md`: the catalog script came from `main` with `curl`, and a provider entry pointed Codex at the server. | Job | Result | |---|---| | Create `hello.txt` and read it back | Completed in 114 s; an `apply_patch` add, then two commands; the file holds `SLOTSTREAM OK` | | Change OK to READY in that file | Completed in 144 s; Codex read the file, applied an update patch and showed the result; the file holds `SLOTSTREAM READY` | | Describe an attached picture | Completed in 104 s; the reply names a red rectangle on a white background | | Look at a picture with `view_image` | Completed in 117 s; the session file shows the call and the picture returned to the model; the reply names the red shape and its centered position | | OpenAI Python SDK 3.14.1 with strict validation | 11 of 11: a strictly valid response object, a streamed function call through the SDK's accumulator, a tool result round trip, and a reasoning turn with summary text and counted reasoning tokens | ## Documented constants The Codex guide and the API reference state four configuration facts. None is a speed measurement. - `stream_idle_timeout_ms = 1800000` is the Codex provider setting used in every run above. Codex's own default is 300,000 ms, and its idle timer counts events rather than bytes. - The server sends a `response.in_progress` event every 10 seconds while a prompt is read, from the constant in the Responses handler. `responses-events` checks that this keepalive is a real event, not a comment. - Without `max_output_tokens`, the reply budget is the gateway budget: a quarter of the served window, at least 256 and at most 8,192 tokens, and always below the window. `GatewayDialect.outputBudget` computes it, and `gateway-catalog` checks that it stays below the window. - The guide was verified with Codex 0.148.0, the version that ran the jobs above. ## Limits No clean-timing or throughput qualification is claimed. Acceptance shared the machine with ordinary work and with another session's model runs, and each phase waited until the model lock had been free for two minutes. The Codex job durations include prompt reads of about 10,000 tokens at a 12 GB target and describe this run only. The installed-release checks ran at a 32,768-token window and a 10 GB target, without the lookahead the automatic plan enables at larger targets. The Codex checks cover one Codex version, and Codex changes quickly. ## Speculative verify pass: split attention by default and an exact mode **Outcome: the speculative verify pass now splits its attention from 6,144 tokens of context. At the 22 GB default profile a verification round takes 12% less time at 16k, and on a quiet machine speculative decode ran at 11.82 against 11.67 tok/s (x1.013) with a 16,356-token prompt and x1.29 with a 32,740-token prompt.** The three-row pass of draft depth 2 had been falling off the backend's vector attention kernel onto the dense kernel, whose cost grows with the context. Run two rows at a time through the vector kernel, the fetch-free pass stays nearly flat: 32% cheaper at 32,740 tokens and 45% at 65,508. An opt-in exact mode goes further: a speculative run's output is identical to a plain run's in that mode (128 of 128 tokens at 16k), which no earlier pass guaranteed. Default: the split joins the deployment family (`SLOTSTREAM_OPT_VERIFY_SPLIT`, threshold `SLOTSTREAM_OPT_VERIFY_SPLIT_CONTEXT`); the exact mode stays off (`SLOTSTREAM_OPT_ROW_INVARIANT`) because it changes plain decode's rounding and costs speed. New gates: `mtp-rowcheck` in `Tools/verify.sh` and the weights-free `verify-pass-rows` catalogue check. Measured and not shipped: a gathered-key attention and a fused hyper-connection read. **Why the pass grew with the context.** The pinned backend admits its vector attention kernel only while query rows times the GQA factor is at most 32 (`scaled_dot_product_attention.cpp`); at this model's GQA of 12 that is two rows. A three-row pass takes the dense kernel, which reads every cached key and value of the twelve sparse-attention layers and builds the full score matrix. The audit of 2026-09-04 named this (§5). Synthetically, one layer's masked attention at three rows costs 0.32 ms at 4k keys, 1.16 at 16k, 2.31 at 32k, 5.29 at 65k and 15.78 at 131k. **The fix.** The split runs the rows through the vector kernel two at a time over the pass's keys and mask, and concatenates; below the indexer budget, where no selection mask exists, it builds the causal mask itself. It engages for passes of three to eight rows from the threshold on. One- and two-row passes and prefill chunks are untouched. **Fetch-free verify pass, warm, median of 4 positions, ms** (15 to 19 GB targets, 2026-09-16; warm misses 0 in every kept cell): | prompt tokens | target, window | k=1 | k=2 | k=3 dense | k=3 split | k=3 gathered keys | | ---: | --- | ---: | ---: | ---: | ---: | ---: | | 4,068 | 15 GB, 32,768 | 51.4 | 65.3 | 77.6 | 84.7 | 86.8 | | 16,356 | 15 GB, 32,768 | 54.5 | 69.0 | 96.8 | 81.5 | 83.2 | | 32,740 | 17 GB, 65,536 | 57.8 | 74.2 | 123.6 | 84.3 | 85.9 | | 65,508 | 19 GB, 70,000 | 64.2 | 81.2 | 180.6 | 99.2 | 91.5 | The third row's cost over the second is 12.3, 27.8, 49.4 and 99.4 ms at the four lengths with the dense kernel, against 16.2, 12.5, 10.1 and 18.0 ms split: the growth is the dense kernel. **Where the split starts to pay** (13 GB, each rung with the smallest window that holds it, 2026-09-17; split minus dense at k=3, same positions, same process): | prompt tokens | k=3 dense | k=3 split | split minus dense | | ---: | ---: | ---: | ---: | | 4,068 | 77.8 / 79.6 | 84.6 / 85.7 | +6.8 / +6.1 | | 6,116 | 93.9 | 92.2 | -1.7 | | 8,183 | 86.0 | 82.1 | -3.9 | | 12,279 | 97.1 | 89.3 | -7.8 | | 16,356 | 98.0 | 86.8 | -11.2 | The two 4,068-token values come from two rungs of the same prompt. The crossover lies near 5,700 tokens, so the split engages from 6,144. A second 4,068-token prompt with a longer reply measured the two within 0.3 ms, so the crossover moves with the content and the threshold errs toward the dense kernel. Rungs that ran beside another session's build, app check or large process were discarded; the run keeps them. **Decode on one warm engine, final build** (22 GB, draft head and decode lookahead on by the automatic plan, 128 greedy tokens, arms interleaved over two rounds after a warm-up; `spec` is the dense verify pass, selected with `SLOTSTREAM_OPT_VERIFY_SPLIT=0`; tok/s, the larger round is the median of two): | prompt, build | plain | dense | split | exact | plain-exact | | --- | ---: | ---: | ---: | ---: | ---: | | 16,356, first build | 9.19 / 9.23 | 10.93 / 10.92 | 11.75 / 11.80 | 12.21 / 12.21 | 8.68 / 8.74 | | 16,356, final build | 8.58 / 8.42 | 9.62 / 9.64 | 10.65 / 10.11 | 11.02 / 10.73 | 8.15 / 8.20 | | 32,740, final build | | 7.75 / 8.52 | 10.90 / 11.66 | | | | prompt, build | dense, ms per round | split, ms per round | exact, ms per round | plain, ms per token | plain-exact, ms per token | | --- | ---: | ---: | ---: | ---: | ---: | | 16,356, first build | 221 | 194 | 198 | 108.6 | 114.8 | | 16,356, final build | 251 | 220 | 222 | 117.6 | 122.3 | Rounds are 128 tokens over the verify passes (53 dense, 56 split, 53 exact), drafts and expert reads included, from the mean of the two rounds. The first build carried the same dense and split code with the split off by default, and its exact mode still attended two rows per call; the final build is the shipped code. Each arm repeats its own output exactly and decodes the same text in both runs. Greedy output diverges between arms at near ties (the dense pass from plain decode at token 76, the split at token 24), so the acceptance rates differ (69.8% dense, 63.4% split, 70.8% exact) and the throughput ratios include that; the per-round costs do not. The split saved 12% of a verification round in both runs. The whole machine ran about 10% slower during the final build's run (load 4.9 at its end against 2.7), so the absolute rates come from the first; paired within rounds, split over dense was 1.075 and 1.081 there and 1.107 and 1.049 in the final run. In both runs the exact speculative arm and the exact plain arm agreed on 128 of 128 tokens. At 32,740 tokens on the final build the split was 1.41 and 1.37 times the dense pass within rounds (medians 11.66 against 8.52 tok/s). That run's own rounds differ by 7 to 10% in absolute rate, on a machine still settling after a killed battery, so the ratio is the result and the rates are not. An earlier 32k comparison on the first build is discarded for the same reason, more severely (both arms 15% slower in its second round, 5-minute load 6.7); its paired ratios were 1.33 and 1.37. **Why a multi-row pass was not row-exact.** The pass is deterministic (a cold and a warm run agree at every k), and most large kernels compute each row alone (4-bit `qmv` below the backend's batch limit, gathered expert products, the fused RMS norm, RoPE, the pooled index keys). The difference comes from the model's dense matmuls, whose kernel changes with the row count: one row runs a GEMV and two or more a split-K GEMM, whose partial sums land in a different order. The dense matrices are the fp32 router (2560 to 512, 48 per round), the indexer projection (2560 to 640, 12), the GDN gate projections `in_proj_a` and `in_proj_b` (2560 to 48, 72), the shared-expert gate (2560 to 1, 48) and the hyper-connection inject weights (10240 to 4, 96); the hyper-connection down and up projections are 4-bit. The dispatch census counts 47.6 fp32 GEMVs and 167 bf16 `gemv_bm1_bn8` kernels per plain round (the narrow-output tiling: GDN gates plus inject) and 48.9 fp32 and 172 bf16 split-K GEMMs per verify pass. With seeded random inputs on the same backend, the stock multi-row matmul differs from one-row matmuls in 160 of 160 router cases, 82 of 160 indexer cases, 2 of 160 GDN gate cases and none for inject or the shared-expert gate. Amplified through 48 layers of routing, this puts row 0 of a two-row pass 1 to 3% of the logit spread away from the one-row pass, with a top-1 flip in one of twelve measured rows at 4k. The dense attention kernel adds its own deviation for the rows of a three-row pass. **A split row is not always plain decode's arithmetic either.** The vector kernel skips masked keys, so a split row reads the keys a one-row pass at its position reads, but the backend picks the kernel variant and its block layout from the call's key count, not the row's. On this GPU (`applegpu_g17s`) the two-pass variant starts at 1,024 keys and its block count changes above 1,024, 8,192, 32,768 and 65,536 keys; other GPUs switch at 4,096 keys, or at 16,384 and 65,536 (`scaled_dot_product_attention.cpp`). Where a row's own key count and the pass's straddle such a count, the row rounds differently: on seeded tensors, 11, 12 and 9 of 16 split rows differ from the one-row pass at 1,024, 1,025 and 8,193 keys, and none at 4,096, 16,384, 32,769, 65,536 or 65,537. Two-row passes, which the split leaves alone, carry the same one-key offset. This does not matter for the shipped split, whose rows were never exact, but it was a hole in the first exact mode. **The exact mode.** Three parts, together only when `SLOTSTREAM_OPT_ROW_INVARIANT=1` meets the split: - Dense matmuls on one kernel. `gather_mm` with one row per index runs `gather_mv`, which computes each row with parameters that depend only on the shapes: across rows 1, 2, 3, 5 and 8 on all five dense shapes, 1,000 of 1,000 seeded cases are bit-identical. It is not the one-row GEMV, which tiles differently when the input is at least 16 times wider than the output (against it the gathered rows differ in 4 and 8 of 160 GDN gate cases in two draws and 1 of 160 inject cases), so one-row passes take the gathered kernel too. The mode therefore changes plain decode's rounding, which is why it is a whole mode and off by default: the Python parity gates compare short prompts bit for bit against a reference that uses the stock kernels. - Attention one row at a time. From two rows up, each row attends in its own call over exactly the keys and mask a one-row pass at its position uses, so the backend picks the same variant and block layout on any GPU. - The index-block selection per row. Each row's blocks are scored and partitioned at a one-row pass's shapes (its query, the blocks complete at its position), because the scores go through a matmul whose kernel can depend on the pass's shape, and near-tied blocks could then rank differently. With 138 tied blocks at the budget, the whole-pass scores and selection happened to match the one-row ones, so this part is insurance. The pooled index keys are computed by kernels chosen by the block and row widths alone, so a row's pooled keys equal plain decode's. The equality is promised for passes of up to five rows, draft depth 4 or less: quantized matmuls compute each row alone only below the backend's batch limit, and the smallest limit in its table is 6 rows (M1 and M2 below Ultra, outputs wider than 4,096 such as the output head); on this GPU it is 10. Three checks hold the mode: - `mtp-rowcheck` (10 GB; a 973-token prompt extended to 1,019 tokens whose positions advance one token at a time so the rows' key counts run from 1,021 to 1,026, and a 2,826-token prompt above the indexer budget): every row of the two-row and three-row passes, and the state a three-row pass leaves, equal the one-row passes of the mode, 0 deviation and 0 top-1 flips at every compared position; the stock pass reads 1.3e-2 to 2.8e-2 of the logit spread. - `verify-pass-rows` (T1, no weights, 45 expectations): the row path at 2 to 8 rows on all five dense shapes; the split rows at a short context; the exact rows of 2-, 3- and 5-row passes at the eight key counts above, with and without a block selection; the per-row selection across a block boundary with tied blocks; and 4-bit products of 1 to 5 rows in three width classes. It measures what the stock and split paths change (the stock matmul in 24 of 24 router, 13 of 24 indexer and 2 of 24 GDN gate cases; the split rows at the counts above), so the check can fail on this backend. - The 16k decode comparison: the exact speculative run and the exact plain run agree on 128 of 128 tokens. **Cost.** At 16k a verification round in the exact mode took 198 and 222 ms against 194 and 220 split, 1 to 2% more, and plain decode in the exact mode 4 to 6% more per token (114.8 against 108.6 ms, 122.3 against 117.6): the gathered matmuls cost more than the stock kernels (the first build's exact mode differed from its split only there), and the final mode's extra attention call per round did not add a visible cost beyond that. Fetch-free at a 4,087-token prompt and a 13 GB target, medians of six positions: the one-row pass 51.5 ms stock against 54.7 exact, the two-row pass 65.0 against 73.4, and the three-row pass 77.4 stock, 77.9 split and 80.1 exact. Row 0 of an exact two-row or three-row pass was bit-identical there to the one-row pass at its position, while the stock and split rows sat 1.0% and 1.5% of the logit spread away from it. At a 16,375-token prompt and the same target the one-row pass read 97.9 ms stock and 104.3 exact, the two-row pass 121.0 stock, 115.2 split and 139.7 exact, and the three-row pass 167.3 stock, 148.8 split and 144.4 exact. The exact three-row cell sits below the split there, against the decode comparison's 1 to 2% above: the two measure different things, one verify pass at 13 GB against a whole verification round at 22 GB, and neither is a general claim about the exact mode's cost at three rows. The mode covers how decode and verify passes group tokens. Prefill chunks keep the stock kernels, so a conversation resumed from a cache can still differ from a fresh read where chunk boundaries differ; the Sevra for Mac records hold that open item and propose resuming from 256-token boundaries. **What this means for speculation.** Acceptance samples from the verify pass's own logits, so speculative decode is exact with respect to whatever that pass computes. The dense and split passes compute something 1 to 3% of the spread away from plain decode, so speculation on and off can give different text at near ties; `mtp-check` reports that divergence and does not gate it. The exact mode removes it. **The dispatch census and the fused read.** Per round, plain decode issues 6,341 compute dispatches in 746 command buffers (44.8 ms of GPU time) and a depth-2 round 8,171 in 737 (101.8 ms); matmul-class kernels are a tenth of the dispatches, and the median buffer holds 2 to 3 because every layer ends in a host sync for its router ids. About 1,300 to 1,500 dispatches per round are pointwise hyper-connection and gate arithmetic. A fused hyper-connection read, three compiled kernels for the gate, the stream mix and the inject around the unchanged projections, is bit-identical to the eager arithmetic in the engine but saved at most 0.7 ms on the one- and two-row passes of a quiet 4,068-token rung (50.7 against 50.9 ms at k=1), inside the noise; the repo's compiled normalization finish had already failed its serving gate for the same reason (a 0.05% regression). It is not shipped. The expert-cache scatter, the largest small-buffer cost (46 to 203 µs per dispatch), moves a slot at about 54 GB/s, the rate of a plain copy, so only fewer misses shorten it. **Measured and not shipped.** The gathered-key attention (each row's selected keys, 2,052 per row, attended directly) beats the split only at 65k (91.5 against 99.2 ms) and its bf16 score arithmetic cannot match the vector kernel, so it has no case inside the default windows. Padding verify rows to the tiled quantized kernel costs 2.8 to 3.8 times one row; a three-row quantized product costs 1.03 times one row for most shapes, so there is no per-row weight re-read to remove. Deduplicating expert reads costs 1.65 times the per-pair path. One gated-delta kernel over the verify window is not pursued: the fused recording measurement priced that chain at 0.04 to 0.19% of the round. A block-sparse prefill attention kernel is deferred: it serves prefill only, where the selected-attention path already applies. **Gates on the shipped build.** `Tools/verify.sh` passed on the frozen final build with the split on by default: 247 checks, no failure and no skip, including the Python-reference goldens, the streaming and elastic-pool equality, the governor drill, the prefix and sweep controls, the draft-head parity, the speculative gates with the new `mtp-rowcheck`, the memory-target promises, the long-prompt recall where speculative decode runs above the threshold, `context-check`, serving robustness, the symlinked model directory and the vision serving suite. The T0 and T1 check catalogue passes on the same build (61 checks, 30,634 assertions), and so do the static gates, whose claims gate holds 249 needles on the public surfaces. **Limits.** One machine; four to eight positions per rung; synthetic prose prompts; the crossover moves with content and was measured for draft depth 2 only (deeper drafts split into more calls and were not timed). The fetch-free table and the crossover rungs ran on development builds with the same split code; the decode comparisons ran on the final build. Each decode comparison is one run of two rounds, and its arms decode slightly different texts, so the ratios carry acceptance differences. A first 32k comparison on an earlier build ran on a loaded machine (both arms 15% slower in its second round) and is discarded. ## Quiet-machine recheck after the change was committed The two decode comparisons were re-run on a quiet machine after the commit ([[sources/runs/2026/09/2026-09-17-verify-pass-deployment-recheck]], binary `8ea3c360959fd20c`, same protocol, 1-minute load 1.89 to 2.29 and 35 GB reclaimable at each start). | prompt tokens | rounds | dense | split | ratio | | ---: | ---: | ---: | ---: | ---: | | 16,356 | 3 | 11.67 | 11.82 | x1.013 | | 16,356, second prompt | 3 | 12.82 | 13.90 | x1.085 | | 32,740 | 2 | 8.98 | 11.63 | x1.294 | The 16k ratio replaces the x1.079 measured earlier the same day, when both arms ran about 10% slower on a busy machine. The cause of the small 16k gain is visible in the arms' own acceptance: the split accepts 63.4% of drafts on that prompt against the dense arm's 69.8%, so it runs 56 verify passes to the dense arm's 53 and the cheaper pass mostly pays for the extra passes. That difference is a rounding artifact of the prompt and does not have a fixed direction: a second 16k prompt has the two arms accepting 70.8% and 73.1% and the split 1.085 times the dense pass, and the exact arm on the first prompt accepted more than the dense pass, 70.8% against 69.8%. At 16k the gain is therefore whatever acceptance the prompt draws, between the two measured ends. At 32k both arms run 57 passes at 61.4% and 62.3%, so the ratio there is the per-pass saving and nothing else, and it is the ratio that carries the change. A first 64-token attempt at 16k reported x0.936; 25 to 29 verify passes with a 10% swing between rounds is too short a sample, and the 128-token steps replace it. ## Coding agents: Claude Code, Codex, Pi, opencode and Hermes through slotstream launch at 12 GB **Claude Code, Codex, Pi, opencode and Hermes each ran real tasks on Slotstream through `slotstream launch`, and each agent's second session started from the instructions its first session read, also after a picture and after a server restart.** The work adds `slotstream launch `, which connects an agent to the running server for one run without changing the agent's own configuration; the Anthropic Messages API (`/v1/messages`), which Claude Code needs; `store: false` and the other no-op fields Pi and the OpenAI SDKs send on chat completions (issue #19); and the shared system prompt reuse of issue #18, with coding agents' requests keeping their shared prefix like a conversation. The sources are `CodingToolLaunch.swift`, `LaunchCommand.swift`, `AnthropicDialect.swift`, `Server.swift`, `PrefixCache.swift` and `Plan.swift`; `Tools/coding_agents_gate.sh` is the acceptance. **Setup.** One server at a time, `serve --memory-gb 12 --max-context 65536` on the development Mac, behind a proxy that records each request's prompt and reused tokens. Claude Code 2.1.270, Codex 0.148.0, Pi 0.85.1, opencode 1.18.31 and Hermes 0.21.1 each ran with a throwaway home in a fresh git project: a first session that writes `hello.txt` with the line `SLOTSTREAM OK` and reads it back, and a second session, a new conversation, that reads `note.txt`. Claude Code also ran a third session that describes a picture, and a session before and after a restart with `--prefix-cache-dir`. The Messages API phase checks token counts against usage, a shared system prompt reused despite a different attribution line, streamed thinking with its signature, a stop sequence, the overflow message, ignored fields, an unknown model and a token count during a generation, then the Anthropic Python SDK 1.6.0 with a streamed tool call and its result. **First pass: every task, no reuse across sessions.** In [[sources/runs/2026/09/2026-09-17-coding-agents-live-1]] all 13 Messages API checks, the SDK round trip and all 22 agent checks passed. Turns within a session reused the conversation, but every second session read its whole opening prompt again: Claude Code 15,487 tokens in 158.41 s to the reply, Codex 10,336 in 120.09 s, Pi 1,620 in 28.15 s, opencode 7,296 in 79.13 s, Hermes 12,267 in 102.98 s. The shared prefix saved during each first session was an optional snapshot, which may not displace a conversation, and with four conversations already held it was refused or evicted first. Requests that declare tools now keep their shared prefix like a conversation, and a shared prefix stays as recent as the conversation continuing from it (`SharedPrefixRetention.conversation`; `optimization-prefix-client-capacity` pins the new rule next to the old one). **Second pass: reuse across sessions, until the first picture.** In [[sources/runs/2026/09/2026-09-17-coding-agents-live-2]], 30 of 34 checks passed. Claude Code's second and third sessions reused 13,312 of 15,492 and 15,501 prompt tokens (48.89 and 50.60 s), and after a restart a new session restored the same 13,312 tokens from disk in 0.08 s and reused them (41.81 s). Then Claude Code's picture turn read all 16,346 tokens again (147.72 s), and the four failures followed from it: the first image makes the server reserve 0.9 GB for the vision tower inside the 12 GB target, and that re-plan dropped the retention allowance from 65,536 tokens to the budget share, 10,807, while the pool stayed at 640 slots and the plan's expected peak fell to 10.8 GB. From then on nothing above 10,807 tokens could be kept: Codex's turns read their whole 10,400-token prompt every time (81 to 104 s each), Hermes's second turn read 12,362 tokens (85.3 s), and no second session of Codex, Pi, opencode or Hermes reused anything. The re-plan now keeps, when the plan retained whole conversations, the largest retention for which the pool and prefill pass are at least what the fallback gives (`Planner.loadingVision`); at this plan that is 43,882 tokens with the pool unchanged and a 512-token prefill pass instead of 256, and `vision-check` pins the behavior at 11, 12, 16, 24 and 30 GB targets. **Final pass: everything.** In [[sources/runs/2026/09/2026-09-17-coding-agents-live-4]], with the retention change, the vision change and the launcher review fixes, all 15 Messages API checks, the SDK round trip and all 34 gate checks passed. The second sessions' first requests reused: | Agent | Reused of the prompt | Time to the reply | Second pass, same request | |---|---|---|---| | Claude Code | 12,800 of 15,293 | 43.72 s | 13,312 of 15,492, 48.89 s | | Codex | 9,728 of 10,336 | 17.79 s | 0 of 10,336, 85.53 s | | Pi | 1,536 of 1,620 | 6.86 s | 0 of 1,620, 27.45 s | | opencode | 7,168 of 7,296 | 26.37 s | 0 of 7,296, 80.97 s | | Hermes | 11,776 of 12,280 | 15.27 s | 0 of 12,267, 103.58 s | Claude Code's picture turn reused 15,587 of 16,332 tokens and answered in 11.02 s, and after it the server still held two conversations and the shared prefix (29,645 tokens) instead of nothing. Codex's turns reused their conversation again (10,416 of 10,526 in 9.74 s). After a restart, a new Claude Code session restored 12,800 tokens from disk in 0.07 s and answered in 38.78 s. Across the pass the agents sent 43 requests that read 75,035 prompt tokens and reused 264,943, against 175,889 read and 161,403 reused in 49 requests in the second pass. Claude Code's opening prompt was 199 tokens shorter than in the second pass (15,308 against 15,507) because the launch now denies its `WebSearch` tool, and the shared prefix ends on the 256-token pass grid, so it is 12,800 tokens instead of 13,312. **The launcher review.** An adversarial review of `slotstream launch` reported eleven defects, three of them ways a prompt reached a cloud provider without the user asking: Claude Code followed a Bedrock, Mantle or other cloud switch from a settings file and sent a settings file's API key to the port; Hermes sent side tasks, including the command-approval check, to OpenRouter or another keyed provider when the server was stopped or busy; and an opencode agent with its own model used that model's provider. The others: model names from the port written unchecked into configuration, Pi's `--provider` alone falling back to another default model, Pi's own commands becoming prompts, secrets in dry-run output, a stale Hermes window, Pi's models file losing its symlink and number formatting, the Codex instructions download ignoring the system proxy and prerelease versions, and loose messages. All are fixed and pinned by `launch-plans` (155 assertions). [[sources/runs/2026/09/2026-09-17-launch-review-reproductions]] reran the review's reproductions against stand-in servers in a sandbox without network access: every Claude Code case reached only the local stand-in and sent no user key, the opencode agent cases used the local model or failed, Pi reached the local stand-in (the launcher's earlier arguments reached the other provider), `codex cloud` and `--output-schema` were refused before any download, and the launched Claude Code listed no `WebSearch` tool. `Tools/hermes_config_gate.py` now fails the previous guide, whose approval check tried OpenRouter with the server stopped, and passes the current one, which pins all 18 side tasks Hermes 0.21.1 defines. **Limits.** Single runs on a machine with user applications open, one request at a time; the timings describe this Mac's SSD and these short tasks and are not a benchmark. Every agent ran with thinking off and short sessions: no conversation reached compaction, and no agent ran two sessions at once. Only the 12 GB target was run live; the vision retention change is checked on plans at other targets, and the elastic automatic plan, which re-plans its pool too, is unchanged. The review reproductions used stand-in servers, not the model. Agent versions change quickly, so each claim that names one is to be rerun after an update. ## The server slotstream launch starts: 54 live checks at 12 GB **`slotstream launch ` now starts the server itself when none is running, keeps it while agents use it, and stops it when they are done.** One command works on a Mac with nothing running: the launch starts `slotstream serve` in the background with the window the agent needs, prompt caches on disk and a log, shows the start until the model answers, registers the agent it opens, and leaves the server running for the next session. The server stops 30 minutes after the last agent exits, or at once with the new `slotstream stop`. The sources are `CodingToolLaunch.swift`, `LaunchCommand.swift`, `StopCommand.swift`, `ServerActivity.swift` and `Server.swift`; `Tools/launch_start_gate.sh` is the acceptance, and `launch-server` (94 assertions) pins the paths, messages, policy and status wire in T0. **Setup.** One model process at a time on the development Mac with user applications open, 35 to 36 GB reclaimable at each start, port 11531, `--memory-gb 12`, a throwaway HOME per command whose `.slotstream/models` links to the real weights, Claude Code 2.1.270, Pi 0.85.1 and Hermes 0.21.1. Full output: [[sources/runs/2026/09/2026-09-17-launch-background-server-live]]. **Live result: 54 of 54 checks.** | What was checked | Result | |---|---| | Nothing running, `--no-start` | Refused, named the `serve` command and `--no-start`, started nothing | | Nothing running, `--dry-run` | Printed the server it would start, wrote nothing, started nothing | | No agent named, a terminal | Listed the installed agents and planned the one chosen | | A Hermes folder without the `slotstream` provider | Refused before any server started | | First launch, nothing running | Claude Code answered in 424 s, including the server's start and its 15,284-token first prompt; the server kept running with 0 clients after it exited, in its own session, with the recorded process id, the disk cache, the memory target and the 30-minute stop | | Second launch | Answered in 162 s, reusing 41,038 prompt tokens from the running server, started nothing, named its log and the running server's 12 GB target against the 11 GB asked for | | Hermes against that 32,768-token server | Restarted it with `--max-context 65536` and answered in 118 s; the old process was gone and the new one recorded | | `slotstream stop` | Stopped it, removed the record, released the model lock, logged `stopping: asked to stop (SIGTERM)`; a second stop said nothing was running | | Two launches at once | One server, the other waited for it and both agents answered in 28 and 31 s | | `--idle-exit 0.25` with a registered process | Stayed while that process ran, past the idle time, and stopped 18 s after it exited, logging the reason | | A server started by hand | Hermes refused it and left it running; Claude Code used it and pointed at its window; `slotstream stop` stopped it | | Control-C while the server starts | Exit 130, the server stopped, its record removed | | `slotstream stop` while the server starts | Stopped it and said so; the launch reported `Slotstream stopped while starting` with the server's own SIGTERM line above it | **Why the automatic window, and not 65,536 for every agent.** The guides recommended `serve --max-context 65536` for coding agents. A launch that forced it on every Mac would cost more than it gives at small memory targets: `slotstream doctor` prices the same plan at 12, 10 and 9 GB in [[sources/runs/2026/09/2026-09-17-launch-window-doctor]]. The launch therefore keeps the automatic window, which is never below 32,768 tokens, and asks for 65,536 only for Hermes, whose own `MINIMUM_CONTEXT_LENGTH` refuses less. **Limits.** Single runs on a machine with user applications open, one request at a time; the timings describe this Mac's SSD and these short tasks and are not a benchmark. Only the 12 GB target ran live, with Claude Code, Pi and Hermes; Codex and opencode are covered by the plan checks and the older live pass, not by this gate. The idle stop was exercised at 0.25 minutes, not at its 30-minute default; what the default changes is when the memory comes back, not the mechanism. Three earlier passes of this gate, on earlier builds, are described in the run record: their failures were two wrong checks and one stale record a Control-C left behind, all fixed here. ### v0.2.21 published, installed and accepted **v0.2.21 is published, installed and functionally accepted.** The release starts a coding agent already connected to Slotstream with `slotstream launch`, and starts and manages the server it needs in the background: Claude Code, Codex, Pi, OpenCode and Hermes, with `slotstream stop`, an idle stop, a disk prompt cache, and the Anthropic Messages API behind Claude Code. The exact CI artifact passed every model gate before the tag. The published archive matched that artifact with a valid attestation, and the installer replaced 0.2.20 on this machine. The installed binary passed all 31 end-to-end release checks and 48 of the launch gate's 54, the six remaining being a phase that needs a coding agent this Mac does not have; the race it tests passed against the same binary with Claude Code in its place. Release: [v0.2.21](https://github.com/carloslfu/slotstream/releases/tag/v0.2.21), tagged on `be37bd707ae7517709cdfd13735bd8f471d86d79`, published 2026-09-18T17:19:10Z. Archive SHA-256 `80415bca1be76c8c223bdc1332eeb2774027b4bb6cd560213c30be5995af7c10`; binary SHA-256 `8a30e1e748c5e9294369260132ae3a876a9c65ca6dbd4aeed3081b788e0c167b`. The CI candidate, the re-downloaded public archive and the installed binary are byte-identical. ## What qualified | Phase | Result | |---|---| | Main CI 35365567576, commit `be37bd7` | Succeeded, including the coverage ratchet; `context-proxies` 35365567626 and `docs` 35365567589 succeeded on the same commit | | Candidate verification | Digests recorded; 188 source files match the commit's checkout | | Model acceptance on the downloaded CI binary | 24 of 25 gates in one run, 245 checks; the 25th, vision parity, ran beside it on the same binary and passed | | Release workflow 35373560427 | Succeeded; the public archive matches the CI artifact | | Provenance | `gh attestation verify` returned SLSA provenance v1 for the archive digest, ref `refs/tags/v0.2.21`, source `be37bd7` | | Installation | 0.2.20 replaced by 0.2.21, exit code 0, binary digest re-checked | | Installed-release end-to-end | 31 of 31 checks | | Launch gate on the installed binary | 48 of 54 checks; the six are the race phase, which needs Pi, not installed on this Mac | | The same race with Claude Code | 5 of 5: two launches at once, one server started, the other waited, both answered | ## Two defects qualification caught Neither was found by reading the diff; each was found by a gate. - **A real one, in the release's own new code.** Reading and discarding an oversized request body before its 413 had only the connection's thirty-second read deadline, so a client that declared a large body and then stopped sending waited thirty seconds for an answer the server had already decided on. The serving robustness suite read no response at all within its ten-second bound. `be37bd7` bounds the discarding at two seconds for each piece and ten in total: the stalled client now reads its 413 after 2.00 s, a client that writes all 40 MB still reads it 0.01 s after finishing, and the suite passes 74 of 74. `http-framing` holds the bound weights-free, so CI catches this class from now on. - **Two flaky gates, fixed rather than re-run.** A fifteen-second bound on each `doctor` in the memory-override gate, against a 5.3 s plan for an absurd target, and a three-second semaphore holding a prefetch lane in `expert-lookahead-forecast-merge`, which a descheduled checking thread outlived. Both failed main CI on a slow runner while the code under them was correct. ## Limits No clean-timing or throughput qualification is claimed. Acceptance shared the machine with another session's identical battery; two attempts ended before any gate ran, and the accepted run began only after 150 seconds with no model process anywhere. The installed-release checks ran at the automatic window and plan for this Mac. The one prefill rate recorded here, 3,819 tokens at 300 tok/s, describes that run only. ### v0.2.22 published, installed and accepted **v0.2.22 is published, installed, and functionally accepted.** The release fixes the memory-budget and automatic-context behavior: an explicit `--memory-gb` value is a process ceiling, while automatic context selection no longer spends measured decode cache on a larger window with worse expected request time. It also makes continued conversation reuse exact at chronological prefill boundaries. The README now explains the behavior in plain terms and points users to `slotstream doctor` and `--max-context`. Release: [v0.2.22](https://github.com/carloslfu/slotstream/releases/tag/v0.2.22), tagged on `cdda19bcaaf0b5593598bf47d0d9dbffdfdd4ab4`, published 2026-09-18T21:03:45Z. Archive SHA-256 `6301b7e02749f3f5a040136da2b9de15e7a980075a2356c90b60936b45adf1b4`; binary SHA-256 `ac6d194ee9fc986f775701f468cf1550575965e24f519a9d7daad3c4baa741a0`. The CI candidate, public archive, and installed binary are byte-identical. | Phase | Result | |---|---| | Main CI 35390709270 | Succeeded: public library, 46.31% coverage ratchet, full weights-free catalogue and candidate identity | | Context contracts 35390709264 | Succeeded on the release commit | | Candidate verification | Version 0.2.22; 190 source files match the checkout | | Release workflow 35394735158 | Succeeded with provenance; public archive matches CI byte for byte | | Installation | Public installer installed 0.2.22; installed digest matches CI | | Installed end to end | 31 of 31 after correcting the fixture to cross an exact checkpoint boundary | | Full local model battery | 27 top-level gates passed, including exact resume, quality, serving, and vision | | Full Mac app check | Passed release builds, 25 composer scenarios, real db.md, snapshots, and runtime safety | The first installed run's 30 of 31 result was a stale acceptance fixture, not a release defect. A tiny conversation cannot create an aligned checkpoint under the exact-resume policy. The corrected fixture crosses a real boundary and measured a cache hit increase from 8 to 9. The complete installed suite then passed. No clean-timing or throughput qualification is claimed. The installed acceptance ran on a shared machine with an availability-clamped automatic memory plan. ## Conversation resume: exact against a cold read, at one partial pass per turn **Outcome: a continued conversation now computes what a cold one computes, bit for bit, and a follow-up turn pays one partial prefill pass for it. Measured at 961 slots on a three-turn chat: follow-up prefill 2.47 s against 8.56 s, both well under the 26.3 s of reading the conversation cold.** Until this change a turn resumed whatever state the previous turn left behind, prompt read in passes and reply decoded a token at a time, and that state does not hold what reading the same ids holds. The difference measured 3.7% to 5.9% of the logit spread, inside the band re-chunking a plain prefill already moves them, and it crossed a token: on a 1,430-token agent turn a fresh read scored `>` at 0.9576 and `]` at 0.0421 for one position of tool-call syntax, the continued turn inverted them, and the model's first `file.edit` call arrived malformed. **What was actually different.** The arithmetic depends on how tokens were grouped into passes and on whether each one was read or generated: a 256-row pass sums a row in a different order from a one-row decode step, MLX picks reduction orders by shape, and the model's top-10 expert routing turns a small numerical difference into a different set of experts. Replaying the app's exact grouping with no cache at all reproduced the flipped token, so nothing was corrupt; the grouping alone was enough. **The rule.** A request now resumes only a state whose length is one of *its own* prefill pass boundaries and whose every token was read in those passes, under the same pass size, model and draft mode; everything after the boundary is re-read. Those boundaries are the same positions whatever the prompt's total length, which is why a snapshot one turn leaves is still a boundary of the next turn's longer prompt. The final partial pass is excluded, because a longer prompt reads past that position in one go, and so is the late-context regime, where a pass is measured against the origin of the read it belongs to rather than the position alone. **Exactness, cached against cold** (`prefix-exact-check`, raw next-token logits at the end of the prompt, as a fraction of their spread): | conversation | turn | resumed at | before | after | | --- | ---: | ---: | ---: | ---: | | 1,521-token history | 2 | 1,280 | 5.87% | 0.000000% | | 1,521-token history | 3 | 1,536 | 3.82% | 0.000000% | | 22-token history | 2 | nothing | 3.71% | 0.000000% | | 22-token history | 3 | nothing | 4.07% | 0.000000% | | identical prompt repeated | 1 | whole prompt | 0.000000% | 0.000000% | | second conversation, shared prefix | 1 | 1,280 | 0.000000% | 0.000000% | The two rows that were already exact are the two states the old engine retained that a fresh read also produces: a prompt repeated exactly, which reuses its own retained logits, and the 256-token common-prefix snapshot. **Cost** (`prefix-check`, 961 slots, the pool the Sevra Mac app plans at a 10 GB target, same three turns): | | rule off | rule on | | --- | ---: | ---: | | turn 1, cold | 12.46 s | 11.91 s | | turn 2 | 1.35 s, 25 tokens read | 3.70 s, 156 tokens read | | turn 3 | 1.12 s, 20 tokens read | 4.86 s, 177 tokens read | | follow-up prefill | 2.47 s | 8.56 s | | reading the conversation cold | 27.18 s | 26.28 s | A turn re-reads from the last boundary, so it pays for wherever its prompt ends relative to the grid, between nothing and one pass. At a 14 GB plan with speculative decoding and a 512-token pass, one turn landed 40 tokens past its boundary and read 1.85 s against 10.82 s cold, while another landed 524 past and read 4.33 s. **Second effects worth knowing.** Under the rule a conversation state can never be continued, so it is worth only its ids, which the next prompt's encoding splices in; it is now the first entry evicted and a boundary snapshot the last, and a deeper snapshot of the same conversation replaces the one it supersedes. Without that, a used snapshot pinned the resume point where it was and each turn re-read a little more of itself: the first build of this change resumed 1,280 at turn 2 and again at turn 3 instead of 1,536. **Not changed.** Different pass sizes still give different logits inside the same band, and the rule does not and cannot make them equal; `prefix-check` still measures that band (4.37% against a 5.90% control). A conversation shorter than one pass has no boundary and resumes nothing, which is why `prefix-check`'s own chat carries a longer history now. **The disk tier.** The optional persistent tier follows the same rule: it writes the boundary snapshot rather than the consumed conversation, and restores only a length that is a boundary of the incoming prompt. The snapshot is written before it is forked into the cache, so the state the next turn resumes carries its disk lineage and that turn's save references those rows instead of writing them again: 130 MB of new rows instead of a 236 MB full write. `Tools/persistent_prefix_e2e.py` passes its twelve checks, restart and regenerate included, with both turn-3 prompt and output ids equal to the first server's. Its conversation now carries a round of notes per turn, because a turn that adds only a short question stays inside the pass band its parent already wrote and correctly writes nothing. ## v0.2.22 release-candidate recheck The final v0.2.22 candidate repeats `prefix-exact-check` with 0.000000% prompt-logit deltas on every continued turn. Both follow-up turns resume their own prefill boundaries, edited history rebuilds, repeated prompts use their complete checkpoint and another conversation reuses its shared prefix. In this shared-machine run, the three-turn follow-up prefill took 7.87 seconds against 28.20 seconds cold. The complete native battery passes 27 top-level gates with no failures. Evidence: [[sources/runs/2026/09/2026-09-18-v0-2-22-release-candidate]]. ## Memory budget: preserve expert cache when context cost is unmeasured The automatic context selector used a speed estimate that is deliberately flat above the last measured cache anchor. Comparing two such plans treated removing useful expert slots as having no cost. A fixed process budget could therefore become reservations for a much longer context while a short request used far less memory. Automatic selection now declines any candidate that removes slots from a baseline cache above the measured decode range. It applies the same rule and request-time tolerance to live startup, including busy machines. A rejected candidate explains the cache tradeoff and omits a numeric relative request cost when that cost is unmeasured. Explicit context choices remain available. The memory flag remains a total process planning budget. The expert pool is allocated; runtime, conversation state and workspace allowances are not all materialized at startup. The banner and JSON distinguish those meanings. A lower measured footprint alone is not an error and is not a reason to fill RAM. This correction preserves the existing pool, conversation retention, allocation guards and allocator-cache limit. A reclaimable overflow cache is a separate optimization: copying evicted GPU slots would introduce synchronization on the miss path, while retaining incoming records duplicates the main cache. Neither is enabled without evidence that it improves end-to-end work within the same memory allowance. The correction does not claim that a static plan borrows unused context reservations. Evidence: [[sources/runs/2026/09/2026-09-18-memory-budget-regression]]. The new checks fail against the original policy and pass against the corrected production source. The CLI tests include an explicit memory budget with automatic context, which the older override matrix did not cover. ## Compiled software verification The first frozen allocation-fix CLI passes 319 memory-override cases, 90 planner checks and 130 context CLI assertions. The catalogue passes 63 groups, including 74 assertions for the new memory-budget regression. All completed static components pass after correcting the stale busy-start assertion. The production-source context policy passes; full evidence and the earlier fixture corrections are preserved in [[sources/runs/2026/09/2026-09-18-memory-budget-software-verification]]. The representative simulated 64 GB machine with a 48 GB budget and MTP enabled keeps approximately 33.1 GB of expert cache, compared with approximately 15.7 GB under the old automatic 262,144-token selection. It now chooses 32,768 tokens automatically; larger explicit windows remain available. This is allocation evidence, not a measured throughput improvement or native qualification of the full 48 GB target. ## Native verification The full native battery on the frozen allocation-fix build completed with 26 top-level gates passed and one failed. Passing gates cover pinned weight hashes, reference parity, identical outputs across memory targets, pool resizing, live governor shrink/regrowth, prefix reuse and exactness, the prefill sweep, speculative decoding and image memory, short- and long-request process memory, long-context recall, API behavior and image numerical parity. The image-serving suite passed 24 checks and failed one assertion about reusing image state on a follow-up. The same 24 passes and the same failure reproduce on the frozen pre-change checkout. Both snapshots include the separate conversation-resume work that was already present before this task. This correction does not change that code or relax the assertion. The failure remains open; the complete battery is not all green. The 10 GB short request peaked at 6.216 GB. The 7,972-token long request peaked at 8.094 GB and answered SEVENTEEN correctly. The separately authorized 12 GB image/MTP check peaked at 10.507 GB. These are process-memory and functional results on the development 48 GB Mac. Global paging is retained as a diagnostic, and these runs make no throughput claim. They do not qualify a native 48 GB allocation or a 64 GB Mac. Exact output and both frozen identities are in [[sources/runs/2026/09/2026-09-18-memory-budget-native-verification]]. This native run precedes the final headroom-report and help wording corrections. Those changes only expose the existing budget residual and clarify text; they do not alter allocation or generation. The final reporting checks are recorded separately below. ## Final reporting verification The final report shows planned expert cache at load, non-cache allowances and remaining budget separately. Near the minimum cache size, the planner can consume part of its nominal margin, so the report now calculates headroom from the existing ledger instead of displaying the full nominal margin. An unbudgeted raw pool does not invent total-process headroom. The reporting build passes 319 memory-override cases, 90 planner checks, 130 context CLI assertions and 964,237 production-source policy assertions. Its 48 T0 groups pass 28,623 assertions, including 85 in the memory-budget regression. The delivered build has the identical allocation and reporting sources; only two CLI help wording corrections differ. Exact archived-source comparison, a repeated T0 catalogue and context CLI suite, and direct human/JSON UX checks pass on that delivered binary. Evidence: [[sources/runs/2026/09/2026-09-18-memory-budget-final-reporting]]. These results supplement the earlier full native run; they do not reclassify its known image-reuse failure as a pass. The final source and binary remain local and are not a published release. ## v0.2.22 release-candidate qualification The v0.2.22 candidate closes the one known failure from the earlier native run. The old image-reuse fixture was shorter than the configured 3,072-token prefill pass, so it could not leave a stable boundary containing the complete image. Exact resume correctly rebuilt instead of reusing an inexact state. The corrected fixture crosses a real boundary without reshaping a pass and proves the first request stored it. The follow-up reused the prefix, skipped the vision tower and completed in 3.4 seconds. The complete vision suite passes 25 of 25. The full native battery now passes 27 top-level gates with no failures. It also repeats the pinned-weight hashes, parity, exact conversation resume, short- and long-request 10 GB bounds, speculative decoding, 74 serving checks and 15 behavioral probes. The repository gates pass 319 memory cases, 90 planner cases and 69 catalogue groups with 31,261 assertions. Evidence: [[sources/runs/2026/09/2026-09-18-v0-2-22-release-candidate]]. This qualifies the release candidate functionally on the shared 48 GB development Mac. It remains neither a native 48 GB allocation measurement nor 64 GB hardware qualification. ## Issue 21: confirmed serving bugs repaired, remaining crash and pressure reports unverified [Issue 21](https://github.com/carloslfu/slotstream/issues/21) combines confirmed server bugs, behavior already corrected after the reported release, deliberate compatibility limits, and failures that have not been reproduced. The confirmed streaming and conversation-splicing defects are repaired. This does not establish that every reported symptom is fixed. ## One-by-one disposition | Reported problem | Finding and resulting behavior | | --- | --- | | Long tool arguments arrive only at completion | Confirmed. The parser retained and repeatedly searched a growing parameter buffer; Chat Completions ignored provisional events. Declared string arguments now stream escaped JSON fragments during generation. The parser searches a bounded suffix and keeps completed executable calls separate from provisional fragments. | | Capped tool generation has no finish or usage | Confirmed. An incomplete call became an SSE error. Budget exhaustion now ends with `finish_reason: length`, requested usage and `[DONE]`, preserving partial arguments. Non-streaming follows the same contract. Partial JSON is not an executable completed call; undeclared tools and malformed completed output still fail. | | Follow-up prompts miss saved states | Confirmed splicing defect. A retained descendant could include several assistant turns, but matching compared it all with the first assistant message. Matching now slices and validates one assistant turn at a time, splicing only that turn's native IDs. This repairs that cause, not every nearly-identical prompt. | | Disk restore loses native history | Additional reproduced defect. Memory retained generated reasoning IDs; aligned disk checkpoints lacked the suffix after their numerical boundary. The intermediate restart reconstructed 1628 prompt tokens instead of 1654 and changed reasoning. Disk heads now optionally retain exact conversation IDs independently from numerical-state tokens. The repaired replay reconstructs 1654 prompt tokens and identical answer and reasoning, restoring 1536 numerical tokens. | | Sparse progress and stale tail ETA | Confirmed observability limitation. Progress appears after a completed pass once five seconds have elapsed, including a slow short suffix. ETA uses the recent interval and says “at this rate.” A separate request heartbeat reports the checked phase while a pass is still running. Initial estimates remain estimates. | | No reason given for cache misses | Added diagnostics for disabled/empty/evicted cache, prefix/model/image mismatch, incompatible pass boundaries, missing logits, fork failure and disk restore refusal. Hits report source and reused-token counts. | | Omitted `max_tokens` defaults to 512 | Already corrected after 0.2.20. Current policy uses one quarter of context, capped at 8192; explicit budgets remain available. Contract checks cover this bounded default. Unlimited generation is not promised. | | `store` and unknown optional fields rejected | `store: false` was already accepted after 0.2.20 and is now exercised by the live OpenAI gate. Persistence-requesting or unsupported semantics remain explicitly rejected. Arbitrary unknown fields are not silently ignored; the accepted surface is documented. | | Plain streamed cap has no terminal event | Not reproduced in official 0.2.20 or the repaired build. Both capped plain and unused-tools text streamed content, length termination, usage and `[DONE]`. Regression fixtures preserve that behavior. | | Ollama tools/tool-role rejected | Intentional adapter scope. Refusal now directs clients to `/v1/chat/completions`, with matching documentation. This work does not add Ollama tool calling. | | Silent whole-process exits | Not reproduced. Existing SIGPIPE protections remain. A live TCP reset during output was followed by a successful request in the same server. Writer failures now identify socket, drain, cancellation or queue causes. A numeric-conversion trap reproduced in 0.2.20 was already fixed after that release; no evidence links it to the reported exits. | | Busy pre-prefill stall after restore | Not reproduced with the available fixture. Added phase heartbeat and cancellation/deadline checks around encoding and between history-splicing steps. These improve diagnosis and cooperatively bound work; they do not prove the reported cause is fixed. | | Long-context pressure slowdown | Not reproduced or performance-qualified here. Pressure eviction and exact-prefix/boundary requirements remain intentional. No pressure workload was induced or reporter-hardware run performed. | ## Preserved invariants Conversation metadata never becomes a numerical resume boundary: only the original checkpoint tokens select restored state, and generated suffixes are reread under the existing fresh-equivalent pass rule. Metadata shares the head's disk quota, expiration, deletion and clear lifecycle. Legacy heads decode; invalid token metadata is refused. Atomic replacement preserves tensor bytes, using APFS cloning when available and a bounded copy otherwise. Requests prohibiting persistence write no conversation metadata. Chat Completions emits one call identity with ordered fragments. Incomplete JSON is exposed only with length termination marking the incomplete response. Completed calls keep validation and typed coercion. Other protocol adapters retain completed-call behavior, and nonblocking socket output remains bounded. ## Evidence and limits [[sources/runs/2026/09/2026-09-21-issue21-regressions]] preserves the official baseline, intermediate failed restart, final executable/source identities and raw evidence. The final build passed all 56 T0 checks (29468 assertions), a tensor-format round trip (94 assertions), all 24 OpenAI compatibility checks, live capped text/tool streaming, three-turn reasoning reuse, TCP-reset survival and identical disk-restart replay. A weights-free parser regression covers 16000 escaped Unicode fragments before closing the call; live fixtures use smaller explicit budgets to exercise termination. These are functional checks at an 8.1 GB target with MTP and vision off on the development Mac, not a clean benchmark, the full release battery or the reporter's long-context workload. Exact source is archived because other tasks edited the shared checkout. The exact client and complete logs were offered but not attached to the issue at capture time. Silent exits, the busy stall and pressure-related tail throughput remain open pending reproducible or diagnostic evidence. ## Issue 21: long conversations and full acceptance qualification The confirmed issue-21 streaming and conversation-cache defects now have passing long-context and full acceptance coverage. This extends [[records/measurements/issue21-serving-regressions-2026-09-21]], whose original findings and narrower evidence remain preserved. It does not establish a cause or fix for every symptom in the report. ## Results [[sources/runs/2026/09/2026-09-21-issue21-qualification]] preserves raw requests, streams, logs, failures, drivers and exact source/executable identities before this transcription. | Check | Result | | --- | --- | | Full acceptance battery | 29 passed, 0 failed, including MTP, vision, cache/logit equality, bounded pressure recovery, process-memory gates and API robustness | | Tensor/format tier | All 15 T1 groups passed | | Serving diagnostics | Context, output, pressure boundary and persistent-prefix checks passed; pressure/persistence variants also passed with MTP | | Final checkout T0/runtime | All 56 T0 groups passed, 29478 assertions, no skips; runtime check passed | | Final live API | Issue-21 streaming/termination/cache/disconnect suite, all 24 OpenAI checks and exact disk restart passed | | Final adaptive server | Kept the explicit 10 GB ceiling through startup cooldown and completed a request | | Long conversation | Prompt grew from 30288 to 51367 to 51413 tokens; reused numerical state advanced from 30208 to 51200 tokens, including after restart | The last follow-up reread only 213 tokens. Its reply, reasoning and usage matched after a real server restart. This passed at configured windows 65536 with a 10 GB target and 131072 with a 13.5 GB target, MTP and vision off. The larger-window follow-up and replay allowed 16000 output tokens and naturally produced 24. The window is a configured maximum; the largest actual prompt was 51413 tokens. The full battery used frozen executable `cc39e86eab88e3873d4c2fad47fb1e0505df63eb75d3d79da8da750c3e42fce0`. After another task added legacy Swift callable overloads and adaptive-policy guards, final checkout executable `7a4ae53bb0ce8e1fbdeb0df7f83685cea7c7b28e341a779814fb6e6fa48d0694` passed T0/runtime, live API/restart and bounded adaptive serving. The 29-gate battery was not repeated on that final executable. The run records each build separately and verifies final source hashes. ## Additional corrections made during qualification Broader tests exposed stale assumptions in diagnostics. Each was checked against actual engine behavior before changing the assertion: - Context-serving counted only retained conversation tokens even though ownership includes the numerical checkpoint too. It also expected all 515 prompt tokens to be reusable where the valid aligned boundary is 512. It now checks exact retained history and rereading of the suffix from the actual safe boundary. - Output-serving expected a generic inference error for low memory. It now requires the typed `insufficient_memory` response and HTTP 503. A second request remains refused while simulated headroom is zero, then succeeds only after a recovery bounded by real available memory and the normal governor policy. - The pressure-scope fixture injected pressure on a continuation poll that occurred after the read scope had committed. It now injects at an actual layer router callback inside the scope and explicitly requires an aborted read with zero committed tokens. For this focused test, it disables checkpoint splitting temporarily so the intended multi-pass scope exists, then restores the setting. Plain and MTP variants pass. These are test repairs, not a weakening of numerical resume boundaries or production memory protection. Earlier failed attempts remain in the evidence archives. ## What remains unproven The original intermittent whole-process exits and 20-minute busy stall did not reproduce. TCP-reset survival, long conversation completion and restart replay passed, but they cannot identify an unobserved failure's cause. A supplied 200-turn history also completed; it was not 200 real sequential generations. The exact client and complete logs were not attached to the issue at capture time. Live length-boundary tests use bounded generation; a 16000-token allowance is not evidence of 16000 generated tokens. The weights-free parser fixture separately covers 16000 escaped Unicode fragments. The configured 131K-window run is not full-window prompt qualification. No induced pressure workload or reporter-hardware test was run, and long-context tail timing on this active machine is not a performance benchmark. MTP/vision pass their acceptance fixtures, not every combination with this long-conversation workload. All owned servers were stopped and reaped. The repairs and evidence are local and have not been published as a release. Ollama tool calling remains outside the adapter's supported scope, with an actionable OpenAI-endpoint refusal; unsupported persistence semantics remain explicit refusals. ## Prompt speed: three qualified changes and one rejected attention upgrade Three improvements are implemented locally: larger guarded expert-read scopes, exact prefix-checkpoint/restart reuse, and one live thinking-to-answer session. The newer fused-attention candidate remains disabled. No release or cross-hardware speed guarantee is implied. ## Why these changes help A prompt has several costs: loading and encoding, fetching experts, dense/GDN work, attention, and optional cache I/O. Sharing an expert read across more existing compute passes reduces repeated SSD traffic. Once a scope already touches most experts, doubling its tokens need not double its expert bytes. Reusing an exact prefix removes work already performed. Continuing a live generation phase removes a second read of the thought without pretending generated state equals a fresh prefill on a later turn. Fused attention targets a different cost, avoiding intermediate score traffic, but its altered accumulation can change floating-point results and selected tokens. ## Results | Opportunity | Implemented behavior | Evidence and limit | | --- | --- | --- | | Larger expert-read scopes | Automatic scheduling may share up to 8192 tokens when the scheduled compute pass is 256 rows; other eligible pass sizes retain 4096. Every original compute pass, checkpoint boundary, memory check and smaller-scope fallback remains. | 817 numerical checks pass. Three eligible loaded-engine pairs give median throughput ratio 1.179187, about 15.2% less prefill time on the 8195-token fixture. Expert bytes fall from 104581324800 to 58511462400. The guarded automatic path also runs under 10 GB. | | Prefix/checkpoint and restart reuse | Read scopes end at the checkpoint actually selected, rather than an obsolete fixed checkpoint. Disk heads record producing pass size; mismatched or unknown arithmetic is refused. The app attaches a per-Home disposable cache for eligible ordinary conversations. | Core 2051-token follow-up reads 259 after reusing 1792, with bit-exact cold logits. Disk reopen, changed-pass refusal and matching-pass restore pass. App restart reuses 2048 and reads 168 instead of 2216, with the same answer. | | Thinking then answering | `Engine.generatePhased` owns both samplers under one gate and transfers live state. It consumes any pending token once and reads the transition suffix. Generated rows remain ineligible for cold-equivalent later-turn reuse. | Plain and MTP checks pass 19 assertions each. Plain transition logits match an independent reconstruction bit for bit. Real app forced closure, Answer now, natural completion, plain follow-up and metric equality pass. | | New fused attention | MLX 0.31.1 remains pinned. The isolated MLX 0.32.2 integration is not promoted. | 807 of 810 numerical checks pass, but three continuation/rollback final-token comparisons fail. This is failure of the project's equivalence gate, not proof of worse answer quality. Component speed or a non-interleaved full-model time cannot override it. | ## Privacy, memory and compatibility Ordinary app caches live in the owning Home's `.sevra/prefix-cache`, inherit the engine's existing quota/age policy, and are excluded from backups. Thinking, historical thinking conversations and incognito do not attach the disk tier. Home/privacy transitions clear held state before encoding; a direct thinking call detaches the tier too. Unavailable disk acceleration falls back to ordinary inference. The real app cache/phase sequence peaks at 8.081 GB under its 10 GB setting and preserves cache-file hashes through private requests. Legacy generation signatures remain available. The new phase API preserves request context/headroom policy, cancellation and independent output budgets. The legacy explicit cache mode retains its original shared-prefix behavior; aligned production reuse remains stricter. Full shared-prefix regression passes 745 checks, all 56 T0 groups and all 15 component groups pass, and the full native app and static scripts pass. Parity golden bytes remain unchanged. ## Timing boundaries and reproducibility The first fresh-process scope experiment has no eligible paired rounds and is discarded for timing. The separate loaded-engine experiment has three eligible pairs; rounds 0 and 4 are discarded. All raw results remain available, including intermediate diagnostic, compatibility and compilation failures. Checkpoint/app timings are single functional observations, not additional percentage-speed claims. The measured scope gain applies to this M5 Pro, fixture and bounded loaded configuration. Cold start, other Mac generations, larger compute-pass scopes and universal end-to-end latency remain unmeasured. Evidence: [[sources/runs/2026/09/2026-09-21-prompt-speed-fresh-scope-discarded]], [[sources/runs/2026/09/2026-09-21-prompt-speed-loaded-scope]], [[sources/runs/2026/09/2026-09-21-prompt-speed-engine-and-fused]], [[sources/runs/2026/09/2026-09-21-prompt-speed-app-qualification]]. Operating decision: [[records/decisions/prompt-speed-qualified-paths]]. ## Later reassessment on September 21 [[records/measurements/fused-attention-reassessment-2026-09-21]] and [[records/decisions/kernel-upgrade-fidelity-and-cache-equivalence]] correct the fused-attention interpretation above. The candidate still fails the original cross-kernel token-parity checks, but that does not establish inability to use it or lower answer quality. The ordinary rechunk control also changes a token; independent high-precision component comparisons favor the fused result; and the same fused backend passes the tested exact warm/cold continuation. The small semantic test contains the same arithmetic error under both backends. The original evidence remains preserved and the three other qualified paths remain adopted. ## Fused attention reassessment: usability, fidelity and qualification Fused head-dimension-256 attention can run on this M5 Pro with Slotstream's explicit sparse masks. The earlier rejection establishes a failure of cross-kernel token parity, not an inability to integrate the kernel and not demonstrated lower answer quality. This reassessment preserves the original run and its three candidate mismatches. ## What the rejection missed The existing diagnostic compares 256-token reference prefill, ordinary 512-token rechunking, and the fused 256-token candidate. It allows numerical drift relative to the rechunking control, but only asserted greedy-token equality for the candidate. The symmetric reassessment finds one changed token in the ordinary rechunking control and the same three changed choices in the candidate, across seven checkpoints. The synthetic prompt consists of arbitrary vocabulary IDs. There is no semantic oracle establishing that one of those predicted tokens is the correct answer. A changed arithmetic implementation and an inconsistent cache are separate questions. The standing warm/cold rule requires an identical computation when the same backend processes the same request, including its producing pass schedule. It explicitly does not require different pass sizes to agree. An upgrade can change floating-point rounding without breaking this rule. The new candidate reuses 2048 tokens and produces bit-identical cold logits on the tested continuation. ## Numerical fidelity and real tasks A float64 NumPy reference, computed from exactly the same BF16 inputs, is more independent than treating the old BF16 output as truth. Across 12 synthetic component cases, including sparse masks, odd lengths and keys up to 32768, fused attention has smaller relative-L2 error every time. The median old/new error ratio is 2.562232. Eight query rows across all 24 heads are checked in each case. This is component fidelity, not a full-model accuracy guarantee. The fallback materializes intermediate attention scores in the input dtype; the NAX implementation uses float accumulators and avoids the complete score tensor. They compute the same masked-attention formula with different rounding and memory traffic. The improved component fidelity is therefore consistent with the implementation, rather than a reason to expect bitwise agreement with the old path. Four real inventory tasks use 3935 to 4190 prompt tokens, 256-row compute, 4096-token read scopes and greedy generation without thinking. Both implementations produce identical output tokens in every task. Retrieval, JSON extraction and a typed tool call are correct. Both produce the same wrong arithmetic result, 966 instead of 714. Preserve that failure: the report passes 35 of 37 assertions, not every assertion. This limited comparison finds no candidate-only semantic regression; it does not certify general answer quality. Physical footprint remains below 10 GB. ## Integration and remaining scope The isolated Swift build uses MLX 0.32.2 and its matching metallib. The ordinary project still pins MLX 0.31.1. The tested newer dispatch defaults D256 to fused only for at least 1024 causal queries without an array mask. force_fused permits the supported kernel with the 256-query and explicit-mask geometry tested here. The later upstream array-mask dispatch change is distinct from, and unnecessary for, this forced path. This establishes feasibility on the M5 Pro. The production dependency is not upgraded by this research. Full backend/application qualification, MTP and persistent-reopen coverage under the new backend, larger context profiles, and other target Mac generations remain outside these new tests. Existing speed and memory gates continue to apply. The same-backend cache check remains strict; a cross-backend token mismatch is labeled as parity evidence, not silently rewritten as a quality failure or a passing parity result. ## Full-prompt performance The separate loaded-engine experiment processes the same 8195-token fixture, with the same model pool and compute/read schedule, under fused and unfused MLX 0.32.2. One warmup per arm precedes five alternating pairs. Two pairs are eligible; three are excluded for paging. Eligible unfused/fused times are 40.860659/38.723230 seconds and 41.259546/39.303033 seconds. The median paired throughput ratio is 1.052489, about 4.99% less prefill time. This is preliminary evidence from two clean pairs, not a release-qualified or universal speed claim. All prompt-completion and physical-memory checks pass. The whole-prompt gain is much smaller than the attention-operation gain because fusion does not remove expert SSD traffic or the rest of the model work. The measured candidate also reads about 0.4% fewer expert bytes because rounding can change routing, so this is an integrated-candidate comparison. Lifetime physical peaks do not establish a whole-process memory reduction. Longer contexts and larger compute passes may make attention's memory savings more valuable; that is an engineering hypothesis requiring a separate profile measurement. Correctness evidence: [[sources/runs/2026/09/2026-09-21-fused-review-correctness]]. Timing evidence: [[sources/runs/2026/09/2026-09-21-fused-review-performance]]. Operating interpretation: [[records/decisions/kernel-upgrade-fidelity-and-cache-equivalence]]. ## Diagnostic correction verified The main diagnostic now records every reference, control and candidate token choice while retaining its original parity assertions. The optimized main build, all 56 T0 groups and all 817 checks in the real 1024-token scope fixture pass. All 21 arm-token measurements and seven control-match observations are present. Production Slotstream sources match the previously qualified runtime byte for byte. Exact sources and final verification are retained in [[sources/runs/2026/09/2026-09-21-fused-review-verification]]. ## Qualified fused prefill integration and remaining bottlenecks The upstream fused D256 prefill path is integrated in the working tree, with matched CLI and Mac-app dependencies, shader packaging, cache identity and numerical/runtime qualification. This is local integration, not a published release. Automatic forcing is limited to the tested Mac17,9 / M5 Pro / macOS build 25G83 profile. Other devices retain MLX dispatch; explicit forcing still requires supported NAX hardware and BF16 prefill geometry. ## Measured gain and limits | Fixture | Comparison | Eligible pairs | Result | | --- | --- | ---: | --- | | inventory8195 | legacy to fused | 3 | 5.22% less prefill time, 1.0550x throughput | | inventory8195 | mlx32-default to fused | 2 | 1.80% less prefill time, preliminary only; fewer than three clean pairs | | inventory8195 | legacy to mlx32-default | 2 | 2.95% less prefill time, preliminary only; fewer than three clean pairs | All three integrated pairs improve prefill time, by 3.74% to 5.66%; median saved time is 1.55 seconds. The fused candidate stays below 8.23 GB physical peak in these three cells. The third normal-dispatch control has paging, so it is excluded from both component comparisons while the clean old/fused pair remains eligible. No 16K speed percentage is qualified. The main comparison uses this 48 GB Mac, a 10 GB engine target, identical raw inventory prompts, 256-row compute and fresh processes. It reports medians of eligible paired ratios. The original 8.1 GB pilot is retained separately and supplies no speed claim. These results do not imply the same gain at another memory budget, on another Mac, on warm-prefix hits or for every prompt. Read-scope schedules and routing can change the number of expert reads; the report preserves them for each pair. No across-study percentage is compounded with earlier prompt optimizations. The integrated old-backend comparison and the fusion-only comparison have separate eligibility counts. Fewer than three clean fusion-only pairs cannot establish a qualified fusion-only percentage. Adoption concerns the complete tested integration; it does not imply that all its gain came from attention. ## Integration and correctness mlx-swift is pinned at ab924c82ead3b970caaa1c0ac11171de23f0305a with MLX 0.32.2 and matching verified shaders. Upstream [D256 NAX attention](https://github.com/ml-explore/mlx/pull/3842) is by wyanzhao; [force_fused](https://github.com/ml-explore/mlx/pull/4185) is by hojin12312. The later [automatic array-mask dispatch](https://github.com/ml-explore/mlx/pull/4416) is by dwijenpatel and is unnecessary for this explicitly selected path. The integration and qualification here use those upstream mechanisms. Causal and explicit sparse masks pass an independent scalar Double oracle on actual BF16 inputs. Decode/short verify and unsupported dtypes preserve their paths. Both main-layer and draft-head Python comparisons pass under the new backend; the MTP reference is bit exact. Exact warm/cold cache logits, disk reopen, invalidation, pool changes, speculative rollback, vision, APIs and real app lifecycle checks pass. Backend/environment/GPU/OS identity prevents an old disk checkpoint from inheriting changed arithmetic. The old MLX 0.31 draft-head golden still differs and remains visible; no historical golden is regenerated and no tolerance is widened. The old arbitrary-rechunk heuristic is a reproducible optional diagnostic, while actual same-schedule cache equivalence stays mandatory. The catalogue passes 72 groups / 31,677 assertions; the final upgrade rerun passes seven acceptance groups. The preceding full battery's 26 passing live groups remain supported by identical production engine semantics, with comment-only engine changes and re-run CLI diagnostics documented by source hashes. All three real Mac checks, the scripted UI suite, bundle build, external Swift consumer and final static/installer gates pass. Numerical and behavioral evidence is bounded and is not a general claim of improved model quality. ## What limits further gains 1. The planner still prices the unfused query-by-context intermediates in `ContextWorkspace.prefillBytes`. Fusion removes the full per-head score/probability buffers, but indexer masks, dense/GDN activations, expert workspace and short-path fallbacks still need reservations. A backend-aware reservation and larger compute passes need independent physical-peak, cancellation, retained-cache and exact-resume qualification before adoption. 2. Memory-eligible expert-read scopes can dominate the outcome. At 8.1 GB the pilot reads more than 912 GB for 8195 tokens; at 10 GB the first baseline cell keeps an 8192-token scope and reads about 60 GB. At 16K, later scopes can contract as context and retained state grow. The opportunity is fewer rereads while retaining the process ceiling, not simply raising a memory limit or multiplying a kernel speedup into a full-request promise. 3. The pinned fused kernel loops over key tiles and computes QK before applying an arbitrary mask. Truly sparse tiled attention could avoid discarded-key work, but per-query selections make efficient matrix tiling and reuse difficult. Kernel-level measurements and actual selection locality are needed; the existing experimental scalar selected-attention kernel is not automatically promoted. 4. Expert matrix multiplication, GDN/dense layers, host synchronization and storage remain. Observed stage counters in the benchmark identify waits but do not fully attribute the remaining GPU work. A full trace should guide the next kernel change. The later upstream dispatch patch alone does not add another kernel gain to an already forced path. Raw evidence: [[sources/runs/2026/09/2026-09-21-fused-integration-correctness]], [[sources/runs/2026/09/2026-09-21-fused-integration-performance]], [[sources/runs/2026/09/2026-09-21-fused-integration-compatibility]], [[sources/runs/2026/09/2026-09-21-fused-integration-excluded-timings]]. Adoption: [[records/decisions/qualified-upstream-fused-prefill]]. ## Remaining long-prompt opportunities: tested gains and rejected alternatives Historical experiment and original verdict, preserved unchanged below. The later user-directed automatic policy is recorded in [[records/measurements/automatic-prefill-policy-2026-09-21]] and [[records/decisions/automatic-prefill-read-policy]]; its adoption does not upgrade this preliminary speed result. The remaining listed opportunities have executable tests and preserved negative results. The strongest candidate is the combination of larger expert-read groups and fused-aware workspace accounting. It remains opt-in: the bounded five-round study produced only one clean baseline/combined pair, short of the prespecified three. The existing upstream fused-attention default from [[records/measurements/fused-prefill-integration-2026-09-21]] remains unchanged. ## Results and decisions | Opportunity | Observed result | Decision | | --- | --- | --- | | Fused workspace accounting plus larger read groups | One eligible 16K pair: 137.63 to 80.27 s, 41.67% less prefill; all five combined observations improve, but four pairs are excluded | Keep bounded opt-in prototype; no qualified new default | | Larger groups alone | Two eligible pairs disagree; slow outlier retained | No independent promotion | | Accounting alone with original group cap | Read reduction in a screen; no clean paired speed qualification | Retain only as part of the opt-in combination | | Clear cached buffers before admission | Same expert-read count; no repeatable demonstrated gain | Remove prototype/control | | 512/1024-token compute | Fixed-pool screens are slower and read more; 1024 physically fits despite planner refusal | Retain existing compute defaults | | Scalar selected attention | About 0.28 times fused throughput on actual captured inputs | Reject | | GPU union/gather compaction, four query-group sizes | Every case slower; best about 33% slower | Reject | | GDN and expert-transfer attribution | Separate fixed-forward probes plus primary stage counters | Diagnostic evidence only; no unsupported kernel speed claim | The main study is this M5 Pro / 48 GB Mac, 10 GB engine target, 961 slots, inventory16387, compute256, MTP off and one greedy output. Filesystem cache is uncontrolled. Physical peak in combined primary cells stays below 8.54 GB. Requested expert bytes fall from about 1.239 TB to 356 GB while computation stays chronological. Requested bytes are not SSD device bytes. The result is first-prefill latency, not complete-answer latency. It must not be compounded with the earlier 8K 5.22% integration result. ## Qualification and implementation boundaries `SLOTSTREAM_OPT_AUTO_SCOPE_LIMIT=16384` permits larger automatic expert-read groups; `8192` or an absent control preserves the established maximum. It does not increase compute rows, bypass a checkpoint or grant memory. CPU retains the original maximum. Public preexisting scheduling signatures retain their behavior; optional Swift settings remain decodable when absent from older serialized objects. `SLOTSTREAM_OPT_FUSED_WORKSPACE=1` removes only full per-head score/probability reservations for the supported fused BF16 text path: 256 query rows, at most 16384 keys, MTP off and no selected-attention/terminal/small-query fallback. Images, CPU, other dtypes, larger contexts and fallback paths retain original accounting. The linear activation floor, indexer/mask allowance, copies, expert workspace, query-by-key envelope, actual process footprint, live headroom and reservation ownership remain. Both controls default off/absent. The experiment does not change the planner's general resident-pool or compute defaults. The actual 16K read envelope passes exact logits, all state bytes and continuation at a bounded diagnostic pool. The catalogue and corrected lifecycle, deployed cache/disk, MTP and synthetic-image rollback checks pass. Initial test-adaptation failures and their source-level explanations remain in the raw evidence. No golden or tolerance is changed. Timing uses frozen V2; the retained implementation narrows its eligibility and has separate correctness/source identities. ## What remains in the way 1. Performance qualification needs a stable host interval with at least three clean paired measurements. The five-round cap was reached; small swap-in counts are not waived and outliers are not removed. This study stops without promoting a new default. 2. Expert rereads are the dominant demonstrated opportunity. Even the combined primary arm still requests roughly 356 GB; its recorded I/O time is roughly 26 to 28 seconds. Actual memory admission can shrink later groups, especially with MTP. The conservative full-layer expert workspace and decode-pool residency still compete for space. Rebalancing that memory is a further experiment, not a proven gain here. 3. More compute rows are not automatically faster. In the matched floor-pool screens, larger passes lose read sharing and increase requested bytes. One larger-pass output changes; numerical/task qualification would be required even if a later configuration becomes faster. 4. These masks are sparse by query but dense across matrix tiles. Avoiding enough work without expensive gathers or losing matrix-unit throughput is the unsolved part. The tested scalar and compaction prototypes do not solve it. 5. Remaining dense/MoE/GDN work and synchronization require attribution under the actual long-prompt candidate before assigning another whole-request speed percentage. Fixed-forward profile numbers do not provide that attribution. Raw runs: [[sources/runs/2026/09/2026-09-21-prefill-opportunities-performance]], [[sources/runs/2026/09/2026-09-21-prefill-opportunities-excluded]], [[sources/runs/2026/09/2026-09-21-prefill-opportunities-components]], [[sources/runs/2026/09/2026-09-21-prefill-opportunities-correctness]]. Decision: [[records/decisions/prefill-opportunities-remain-experimental]]. ## Automatic prefill policy: validation and bounded adoption Historical initial automatic-policy validation, preserved below. MTP is now included through [[records/measurements/mtp-prefill-policy-2026-09-21]]; the earlier paging-excluded study is not relabeled successful. The engine now selects the combined fused-workspace and expert-read policy automatically on the existing qualified M5 Pro profile. Carlos explicitly requested automatic behavior instead of user opt-in after being told the speed result lacked three clean pairs. This is a bounded product adoption supported by exact-state, lifecycle, memory and integration checks. It is **not a claim that the frozen performance-adoption criterion passed**: the new study has zero eligible matched pairs, and the earlier 41.67% result remains one preliminary pair. ## Automatic behavior Deployment enables fused workspace accounting alongside the already qualified fused kernel. One internal PrefillReadPolicy checks the actual GPU/NAX capability, BF16 weights, attention geometry, MTP/image mode and attention fallbacks. Its default larger read envelope applies only to 256-query computation and ends within 16384 keys. At later positions, on other compute shapes, or when the attention path cannot use the reservation, the established automatic cap remains. Each candidate still passes the same physical-footprint, live-headroom and request-ownership admission. A larger maximum never grants memory, changes compute rows or bypasses a checkpoint. The Mac app, CLI and serving adapters inherit this through the shared engine. There is no new UI control or startup tuning benchmark. Explicit reference configurations and older serialized settings keep original reservations. SLOTSTREAM_OPT_FUSED_WORKSPACE=0 restores original accounting and automatic cap; disabling forced fusion also disables the combined automatic path. Independent read-cap overrides remain diagnostic tools, not a setup requirement. MTP, images, CPU and unqualified hardware retain their existing automatic policy. ## Evidence and limits | Round | Previous default prefill seconds | Automatic policy prefill seconds | Timing eligibility | | --- | ---: | ---: | --- | | 1 | 145.658 | 84.924 | Excluded: paging | | 2 | 142.374 | 76.459 | Excluded: paging | | 3 | 139.276 | 88.428 | Excluded: paging | These raw times are excluded from a qualified speed estimate. All primary default runs request 355.763 GB of expert data versus 1239.147 GB, with the same output IDs and compute schedule; requested bytes are not physical SSD bytes. All complete below 10 GB. The study stops after the three initial pairs because both allowed extensions could yield at most two eligible pairs. The original protocol, initial analyzer, explicit early-failure rationale, all cells and the controlled driver exit are preserved. No further extension or cross-study pooling earns a passed result. The full 16K state comparison, default catalogue, lifecycle, aligned checkpoint/disk, MTP/image, 8.1 GB boundary and fusion-disabled checks pass. Mac and static suites pass. See the raw correctness run for exact counts, peaks and the small-pass boundary correction discovered by the first catalogue attempt. Broader hardware, larger contexts and MTP workspace reductions remain unqualified; the guarded default falls back there. A reproducible performance regression should revise the policy, and a reliable speed percentage still requires a new independently frozen clean-host study. Decision: [[records/decisions/automatic-prefill-read-policy]]. Previous experiment and exclusions: [[records/measurements/prefill-opportunities-2026-09-21]]. Raw evidence: [[sources/runs/2026/09/2026-09-21-automatic-prefill-policy-timing]], [[sources/runs/2026/09/2026-09-21-automatic-prefill-policy-correctness]]. ## Automatic MTP prefill: phase accounting and bounded expert writes The supported fused text-prefill policy now works with MTP automatically. There is no new user opt-in. MTP changes retained tensors and peak memory, but it does not require disabling fused main attention. All three matched pairs are eligible. Median paired prefill-time reduction is 65.40%; median paired request-time reduction is 64.53%. ## Implementation Price the main and draft phases separately and reserve their maximum. The draft calculation includes the complete main multi-stream output while it is retained, the shifted first draft pass and cache positions, short tails and full draft-attention allowance. Main and draft replacement caches remain separately charged. All byte arithmetic saturates on invalid input or overflow. MTP resident weights can leave too little room for the larger main expert workspace. The engine can evaluate expert-buffer writes one piece at a time, which releases each old piece before replacing the next. Price the actual largest piece from pool shapes while retaining full original weights, staging, routed activations, retained frontiers and admission copies. This changes allocation lifetimes, not expert bytes or mathematical passes. Extra write barriers are a performance cost. In the qualified fused MTP path, choose this strategy only for groups beyond 8192 through 16384 rows that at least double the largest batched-write group fitting the current process budget. The existing physical-footprint, live-headroom, queued-request, cancellation and checkpoint guards still select or reject each candidate. Public optimization settings remain immutable; execution controls belong to the current group. This is an operating tradeoff, not a new hardware, memory or arithmetic ceiling. ## Qualification Final catalogue: 73 groups, 31841 assertions, all pass. Model assertion counts: mtp-equality16-final: 13, plain-equality16: 13, lifecycle: 1914, checkpoint: 25, mtp-vision: 874. A real server restart test forces MTP on in every server launch and checks the ordinary persistent-prefix path. Native Mac build/scripted regressions and the complete static suite pass. Functional and physical-memory checks retain global paging as diagnostics, independently of clean timing eligibility. The long MTP equality check requires actual grouping beyond 8192 and actual piecewise writes, while comparing raw prompt logits, all retained tensor bytes, teacher-forced continuation, speculative output IDs and chronological compute passes exactly. It exercises 16384 rows and 4878 piecewise writes. Complete process peak including both sequential arms, fingerprinting and continuation is 9.338720528 GB, below 10 GB. Candidate requested reads are 48580300800 bytes; the sequential control requests 1370118758400. Different allocation histories make this an equality/physical-memory check, not a speed comparison. The followup MTP/image lifecycle explicitly executes piecewise expert writes through cancellation during draft processing, checked read failure, rollback and exact retry. Automatic selection, checkpoints, disk restoration, head alignment, process/live-headroom fallback and image geometry remain covered. At 8.1 GB, forced MTP correctly refuses its additional resident head; normal automatic mode completes with MTP off, within the original target. The fusion-disabled MTP run also completes within its target. See functional-summary.json and each original receipt. No golden, tolerance or memory ceiling was relaxed. ## Paired latency The same frozen binary compares `SLOTSTREAM_OPT_FUSED_WORKSPACE=0` against the default automatic policy. Both retain fused attention. Configuration: inventory16387, 10 GB target, 640 slots, compute 256, MTP on with two drafts, greedy maximum 16 outputs, fresh processes, sampled physical footprint. Every primary cell actually emitted 11 tokens and finished at its stop token. The M5 Pro / 48 GB machine uses the pinned MLX 0.32.2 backend. Filesystem cache is uncontrolled; no purge. Source, binary, shader, model-header and exact token identities are preserved. Each cell waits for 120 consecutive nominal, normal-power seconds. Eligibility requires completion, no global swap activity, nominal power/thermal before and after, matching prompt/output IDs, pool, compute geometry and sampling, and physical peak <= 10 GB. Every primary pair has matching inputs, outputs and compute passes. All three matched pairs are eligible. Median paired prefill-time reduction is 65.40%; median paired request-time reduction is 64.53%. | Round | Control prefill seconds | Automatic prefill seconds | Pair eligibility | | --- | ---: | ---: | --- | | 1 | 153.279 | 56.106 | Eligible | | 2 | 155.217 | 52.859 | Eligible | | 3 | 155.869 | 53.936 | Eligible | The generated analysis.json preserves every exclusion, observed read group, requested read byte count, actual peak and drafted-token count. Requested expert bytes are engine requests, not physical SSD traffic. This is a synthetic long-prompt/MTP result, not a universal throughput or cross-hardware claim. The sequential equality diagnostic is separate and must not be substituted for fresh-process timing. ## Limits and remaining opportunities The result applies to this qualified M5 Pro profile, synthetic 16K prompt, small fixed target and MTP configuration. Do not compound it with the earlier standalone fused-kernel percentage. Other budgets and prompt shapes can choose different groups, and extra barriers do not help when the same group already fits. The first implementation's failed larger-scope assertion and the earlier fixed-group negative write experiments remain preserved. The primary cells report median prefill I/O time falling from 106.28 to 5.16 seconds, while median total prefill falls from 155.22 to 53.94 seconds. Scatter time increases from 1.32 to 2.85 seconds and reported GPU wait from 7.46 to 14.37 seconds. These counters may overlap and are not a complete additive profile. They support prioritizing remaining model computation and synchronization over expecting another comparable gain from read elimination at this profile. Raw medians and their interpretation are in timing-counters.json. The 16K key envelope, later-context reservations, full indexer/mask work, expert assembly and main-model compute remain limits on further gains. Broader key ranges and the write-strategy crossover across other prompt/budget profiles need separate measured qualification. Images retain their existing main-workspace policy. At qualification time these were local changes; this record is not a release receipt. Decision: [[records/decisions/automatic-mtp-prefill-read-policy]]. Raw evidence: [[sources/runs/2026/09/2026-09-21-mtp-prefill-policy-timing]], [[sources/runs/2026/09/2026-09-21-mtp-prefill-policy-correctness]], [[sources/runs/2026/09/2026-09-21-mtp-prefill-policy-v1]], [[sources/runs/2026/09/2026-09-21-mtp-prefill-policy-phase-screen]]. ## Remaining decode opportunities: current-backend tests do not qualify new defaults The two remaining decode proposals were implemented or exercised against the current release source and measured before deciding whether to keep them. Neither qualified a production default change. Production engine sources remain byte-identical to baseline; only hidden diagnostics and their correctness checks change. ## Scope and method Baseline is commit `14fb9aa3c253908cf7705b62780b28039ab42f92`, after release 0.2.23, on the M5 Pro 48 GiB development Mac, macOS build 25G83, with MLX 0.32.2 and the pinned Qwen Flash-Next checkpoint. Build identities, source archives, prospective protocols, raw results and every excluded or interrupted attempt are preserved in the linked runs. These experiments concern decode after prompt preparation. They do not establish another CPU/GPU/SSD percentage decomposition or a model-quality improvement. Both proposals target actual waits or a small amount of GPU pointwise work. The previous host elapsed-time categories are not a budget of removable graph-construction time; see [[records/decisions/decode-host-time-is-waiting-not-graph-construction]]. Deferring a barrier cannot eliminate the next router result dependency or an expert cache miss. Pointwise hyper-connection fusion leaves the weight projections, attention, expert reads and most model arithmetic unchanged. MLX 0.32.2 already compiles SiLU, further reducing the additional work available to fuse. ## Deferred layer barriers in plain decode The existing control was sufficient: compare `SLOTSTREAM_DECODE_BARRIER_LAYERS=1` with `4`, using the same frozen V1 executable. MTP and lookahead are off; memory target is 10 GB, actual pool is 1217 slots in every arm, elastic resizing and prefix caching are disabled, and the actual context cap is 32768. Every arm has a separate same-prompt warmup followed by a greedy measured request capped at 192 output tokens. Workloads are public code, reasoning and prose fixtures. Order alternates across rounds. Expert counts are application reads, not physical SSD-device bytes; filesystem cache is uncontrolled. The frozen adoption gate requires at least 3% median paired decode-time reduction across nine eligible pairs, no workload regressing by more than 1%, exact outputs, equal decode work, no more than 1% additional expert reads and physical peak within target. Global swap activity or nonnominal thermal observations exclude timing. An added readiness wait strengthens preparation without relaxing those gates. | Workload | Eligible pairs | Paired decode-time reduction | | --- | ---: | --- | | Code | 2 | +8.04% and -33.94% | | Reasoning | 3 | +1.79%, +1.66%, +0.44%; median +1.66% | | Prose | 1 | +2.87% | All six eligible pairs have exact output IDs/text and matched decode work. Their observed median reduction is 1.73%. A second code pair is paging-excluded. The slow third code pair passes the frozen interval gates and stays included. A macOS scanner was observed afterward during that cohort; there is no trace proving it caused the slowdown. This noisy outlier is neither silently discarded nor treated as a causal 34% regression estimate. The bounded study stopped for failure-only futility: even assigning 100% savings to all three remaining planned pairs would put the final median at only 2.8734%, below 3%. One in-progress prose process was stopped and its incomplete logs preserved. This stopping rule cannot declare success and does not manufacture a completed nine-pair study. The 8.1 GB follow-up was conditional on a promising 10 GB result, so it was not run. Larger budgets, other Macs and other workloads remain unqualified by this test. Decision: keep the existing one-layer default where lookahead is absent, the existing qualified four-layer lookahead behavior, and the explicit override. The small reasoning/prose signal remains a plausible workload-specific gain, but this test does not justify broadening the automatic default. ## Hyper-connection pointwise fusion The archived compiled gate/mix/injection prototype was restored behind a default-off control, with unchanged projections and an expanded exact BF16 self-check. Testing used actual production attention dispatch: stock at the shorter context, split above the existing 6144 threshold. This avoids treating a historical forced-split-at-4K result as production evidence. The component screen used one 13 GB process, a 16384-token context cap, actual contexts of roughly 4068 and 8183 tokens, rows 1 through 3, twelve paired positions after warmup and rotated mode order. Each arm starts from the same model-state checkpoint. The larger diagnostic target permits all experts in a verify pass to remain resident; it is not a small-memory end-to-end serving claim. Recorded physical peaks in the two eligible long-context runs are about 11.84 GB, below the explicit target. Initial screens revealed that one warmup pass does not always establish an all-hit CLOCK cache. Those timings were not promoted. The corrected diagnostic warms each actual mode separately, restores the same state, requires zero expert misses before and during timing, and records native VM/thermal observations around the timing loop. Startup paging remains recorded separately. A bounded optional settle after preparation requires ten nominal observations, with a 180-second abort limit; it is a diagnostic observation policy, not a guarantee that temperature cannot change between observations. | Actual ~8K context | Three-row paired median change | All timed expert misses | Logits | | --- | --- | ---: | --- | | Eligible process 1 | 0.640% slower | 0 | Exact | | Eligible process 2 | 0.727% slower | 0 | Exact | These two timing windows have unchanged native swap counters, nominal thermal observations and low-power mode off. The 24 paired three-row positions contain only one reduction of at least 3%. Even twelve remaining perfect observations could raise the planned final median only to 0.0692%; the third process was stopped during preparation under the failure-only rule. One-row medians improve by 0.49% and 0.93%; two-row medians regress by 1.23% and 1.04%. None supplies the required repeatable 3% component improvement. The 4K attempts were excluded for misses, paging or thermal conditions, so no clean 4K speed estimate is claimed. Decision: remove the production prototype and its control. There is no full-request fusion speed claim. Exact arithmetic alone does not justify carrying production complexity when the relevant measured component does not improve. Raw prototype source is retained for a future backend or kernel change. ## Retained implementation and verification The useful implementation is diagnostic-only: - A `decode-barrier` state-check variant compares the deployed arithmetic under one- and four-layer drains, including exact logits, main-model state tensors, ordered router traces, verify rollback at each kept row and continued logits. It also checks bounded pin generations and retained pins. The final check passes 1130 assertions. This main-model test does not by itself certify every MTP-head tensor. - `mtp-passcost` rebuilds long prefixes in bounded 256-token passes; its attention-mode comparisons prove each mode is all-hit before timing, fail on timed expert reads and emit raw per-position results plus timing-window conditions and physical peak. The optional thermal settle and additional probes exist only in the hidden diagnostic command. - The lifecycle check now covers both deployed aligned-prefix resume and legacy extend-only resume. Its old assertions incorrectly demanded reuse of tiny, unaligned decoded continuations under the new deployed policy. The original test failed identically at barriers one and four. The corrected test verifies that deployed mode rebuilds and creates a fresh aligned draft, while legacy reuse refuses stale draft verification. Cancellation at a committed 256-token prefill boundary remains covered. The final build passes 1130 barrier/state assertions and 154 lifecycle assertions at each barrier setting. The passcost smoke completes eighteen zero-miss samples, reports its timing conditions and peaks at 9036485168 bytes under the 13 GB target. This stock/split/exact smoke checks the diagnostic contract; it does not assert those distinct attention modes produce identical logits. The complete static suite passes, including 420 memory-override cases and installer checks. The prototype changes are not shipped, normal run/serve instrumentation is unchanged, and no inference service is left running. Raw successful, excluded, interrupted and initially failing checks are linked separately rather than overwritten. Runs: [[sources/runs/2026/09/2026-09-22-decode-opportunities]], [[sources/runs/2026/09/2026-09-22-decode-opportunities-excluded]]. Decision: [[records/decisions/decode-opportunities-stay-opt-in-2026-09-22]]. ### v0.2.23 published, installed and accepted **v0.2.23 is published, installed and accepted.** It ships the session's automatic long-prompt prefill policy, MTP phase accounting and bounded expert writes, pinned fused-attention backend integration, exact persistent conversation reuse and live thinking-to-answer continuation, together with the documented serving, memory and Mac app source changes. Release: [v0.2.23](https://github.com/carloslfu/slotstream/releases/tag/v0.2.23), published 2026-09-22T14:53:23Z from commit `14fb9aa3c253908cf7705b62780b28039ab42f92`. The CI candidate, public archive and installed executable match exactly. Archive SHA-256: `1b499652c33e2eb46af702c64b4ed26f62191b9538804567d057b8914f62ed22`. Executable SHA-256: `5cb612361887c2a317376721dbe5df94810da5230759c47169fc6fa2e9b30b89`. | Acceptance | Result | |---|---| | Exact-commit hosted CI | Main engine, external library consumer, coverage, Mac app build/checks, docs and context contracts passed | | Engine catalogue | 73 groups, 31,841 assertions; no failures or skips | | Full native model battery | 32 top-level gates passed; API robustness 74/74, quality 15/15 and vision serving 25/25 | | Automatic 16K MTP | 13/13 assertions; actual 16,384-token group and bounded writes; exact state/logits/output; 9.401 GB complete peak inside 10 GB | | Public distribution | Exact CI archive published, checksum and provenance verified; public installer upgraded the standard installation | | Installed serving | 31/31 with a 10 GB target and MTP on; test server cleaned up | The earlier shared-machine attempts and the first generation's unexplained exit are preserved in the raw run, including the successful isolated reproduction and the complete passing rerun. No assertion was removed or tolerance widened. The rerun retains stderr that the original harness discarded. Global paging observations remain diagnostics, and this acceptance adds no new performance percentage. The performance results and their limits remain in [[records/measurements/prompt-speed-qualification-2026-09-21]], [[records/measurements/fused-prefill-integration-2026-09-21]] and [[records/measurements/mtp-prefill-policy-2026-09-21]]. Their percentages describe different comparisons and must not be combined. Separate decode experiments made after this release was frozen are outside the tagged archive. ### Published v0.2.23 multi-prompt speed audit **The published improvements save real work, but they do not produce a universal speedup.** This audit adds 123 completed HTTP requests on the 48 GiB M5 Pro, comparing the installed v0.2.22 and v0.2.23 binaries and selected same-binary feature ablations. The strongest new repeated result is an 85.6% reduction in a cached prose follow-up's request time. That comparison disables checkpoints in the control; it is not an 85.6% release-to-release gain. Short requests have no consistent improvement. Many long-request timing pairs fail the declared no-paging gate, so they support mechanism and correctness observations but no new qualified latency percentage. The runtime under test is the already published [[records/measurements/release-0-2-23-published-2026-09-22]]. This audit changes only the benchmark validator and its regression tests. It does not change the inference engine, app defaults or release binary. #### Method and coverage The frozen protocols exercise prose, code and structured job records, short prompts through 24K tokens, first and warmed requests, MTP on/off, in-memory and disk reuse, and explicit 8.1, 10 and 24 GB memory targets. The model is `qwen38-flash-next-mlx-4bit`. Every run retains the model/build identities, effective plan, actual expert-read groups and compute passes, requested expert bytes, streamed output, generator/client timers, complete memory peak, thermal state and global paging observations. Stage timers overlap; requested expert bytes are application requests, not physical SSD traffic. No cold-device claim is made. Most first-request comparisons cap output at one token to isolate prompt processing. Prefix comparisons cap at 16 output tokens; they are bounded continuations, not complete-answer quality benchmarks. The initial separate uncached pilot includes longer decode workloads. The final disk diagnostic requires at least eight emitted tokens and produces 16 on every request. The normal-cache matrix and subsequent settled confirmation are separate prospective cohorts. Initial uncached pilots changed the expert-pool budget and therefore cannot represent normal cache-enabled defaults. Confirmation uses continuous nominal thermal settling, 120 seconds for long requests and 30 for short/cache requests. Arm order alternates. Three clean pairs within one protocol are required for a repeated timing claim; studies and previously published runs are never pooled to meet that threshold. Percent reductions are medians of paired reductions, not ratios of independently calculated medians. One- or two-pair results remain preliminary. Paging and thermal exclusions retain all slow outliers and are not evidence of an engine correctness failure. There are 115 completed matrix/pilot requests plus four initial and four corrected disk-diagnostic requests: 123 total. The matrix records 81 measured cells, of which 37 meet their original gates; cells are not paired comparisons. Reconciliation passes 1,035 emitted-metric checks on the 115 requests. The final disk diagnostic passes all 32 checks. These checks do not replace the release's independent numerical and quality acceptance. Two interrupted pilot requests and the initial disk fixture's empty-output failure remain separate evidence. #### New timing results | Comparison | Result | What can be concluded | |---|---|---| | 2K prose follow-up, v0.2.23 checkpoints disabled versus default, 10 GB, MTP off | Three clean pairs; median request 30.730 to 4.419 seconds; median paired reduction 85.62%; prefill 28.558 to 2.268 seconds | Exact 2,048-token reuse skips most repeated prefill. All 16 output IDs/text and decode work match. This is a reuse ablation, not a release delta. | | Complete two-request prose sequence, same study | Only one clean complete pair; summed request time 51.748 to 22.424 seconds | Preliminary only. The three clean follow-ups do not establish three clean complete sessions. | | 8K code, v0.2.22 versus v0.2.23, 24 GB, MTP on, planner-owned 1,024-row passes | One clean pair; request 47.556 to 44.496 seconds, 6.44% lower; prefill 6.48% lower | Preliminary larger-profile release benefit. Both plans use 3,698 slots and the same 4K + 1K read groups. Peak 18.812 to 18.496 GB, below 24 GB. | | Short warmed requests, v0.2.22 versus v0.2.23, 8.1 GB, cache disabled | Three clean exact-output pairs; paired request changes are 2.49% slower, 2.00% faster and 4.62% slower | No consistent short-request gain. This 827-slot setup study is separate from the normal-cache 640-slot profile. | | Short normal-cache first requests, 8.1 GB, MTP off | Two clean pairs; prefill 1.504 to 1.526 seconds, median paired 1.44% slower | Preliminary and small; no meaningful consistent improvement established. | | 2K prose first request, release comparison, 10 GB | One clean pair; prefill 17.057 to 14.197 seconds, 16.77% lower | Preliminary prefill observation. First output IDs differ between backends; not an exact-output full-request claim. The separate settled confirmation has no clean pairs. | | 8K code, same v0.2.23 backend, fused off versus fused on | One clean pair; prefill 31.112 to 30.665 seconds, 1.44% lower | Preliminary kernel attribution. Default workspace accounting gives 30.652 seconds, 1.48% lower than the same reference. All arms already use the same 8K read group. | The 24 GB screen does not force `SLOTSTREAM_PREFILL_CHUNK`. It requires 30 GB of real reclaimable memory before launch. It is not a measurement of the current larger automatic target near 32 GB. #### Long prompts: strong work reduction, excluded new timing percentages | Workload | Requested expert reads and actual grouping | Timing status | |---|---|---| | Normal-cache 8K code, old versus new, 10 GB, MTP off | About 1.036 TB to 70.630 GB; candidate uses 8,192 + 12 tokens | All three pairs have paging. Old client times 104.470/106.901/112.320 seconds; new 68.205/39.470/37.788. The slow first candidate stays in the record; its cause is unproven. | | 16K structured data, same new binary with workspace accounting off versus default, 10 GB, MTP on | 1,428.343 to 67.898 GB; reference begins at 6,656 then contracts to 256; default uses 16,384 + 13 | All three pairs excluded because reference requests page. Prefill reference 141.392/147.359/154.384 seconds; default 53.236/53.092/54.641. Exact first output and chronological compute match. | | 8K prose at the 8.1 GB floor, old versus new | 1,183.683 to 947.909 GB; new first group 2,048 then 256 | One paging-excluded pair. The tight budget limits the read-sharing benefit. | | 24K code, 10 GB, MTP on, old versus new | 3,647.565 to 1,294.300 GB; new first group 16,384 then 256 | One paging-excluded pair. The post-16K contraction remains visible. Peak stays below 9.326 GB. | An initial 16K prose release comparison also encountered thermal drift and paging and was stopped; it supplies no new clean speed percentage. All completed matrix peaks stay inside their requested target. Global paging counters do not identify the responsible process, and a 20-second idle control cannot establish the cause of paging during inference. #### Reuse, disk restore and the benchmark correction The MTP-on code follow-up study produces the same 16 tokens and decode work in all three pairs and reuses exactly 2,048 tokens. However, the frozen validator rejects candidate cells because warmup declines a second optional complete-prompt checkpoint after storing the useful boundary. The shared retention budget is 7,595 tokens; native counters show one warmup refusal, zero measured refusals and zero errors. Original invalid verdicts and timings remain unchanged. `Tools/serve_bench.py` now accepts prospectively declared exact refusal counts for every phase and arm, restricted to integer counts from zero through two. Existing protocols still require zero; storage errors, reuse, forks, stores, identity and resource gates remain strict. Regression coverage passes 51 tests, and offline replay of all six captured MTP cells passes the corrected functional validator. This is not retrospective timing qualification. The analysis itself passes 12 unit tests. The final 8K code diagnostic compares memory-only serve against an attached disk cache, using normal chat formatting and MTP off at 10 GB. Both requests emit the same 16 answer tokens in both arms. Disk restores 8,192 tokens; its follow-up takes 3.693 seconds versus 40.991 seconds for the memory-only miss. The first requests take 35.230 and 33.733 seconds respectively. Complete two-request sums are 38.924 versus 74.724 seconds. This is one functional pair with paging in different stages, so no qualified percentage is claimed. All memory targets and 32 functional checks pass. The preserved initial raw fixture restored 7,936 tokens but both follow-ups returned immediate EOS. Its same-output check was vacuous; independent reconciliation found the missing generated answers. The corrected fixture and minimum-output requirement were frozen before execution. The initial result is excluded from generated-answer equivalence evidence. #### Attribution and remaining opportunities 1. **Read sharing explains the largest mechanism gain.** Longer groups fetch an expert once for more chronological compute passes. Fused attention reduces intermediates and enables honest workspace accounting, but the whole policy gain is not a kernel-only gain. The upstream kernel is credited in [[records/measurements/fused-prefill-integration-2026-09-21]], not an invention of this audit. 2. **Larger planner passes and more than 16K keys still leave headroom.** `Sources/Slotstream/PrefillReadPolicy.swift` grants the expanded envelope only for 256-row queries and at most 16K keys. Read-only plans choose 1,024-row passes at 16/20/24 GB and 2,048 rows on the larger automatic profile. The 24 GB comparison retains the same read groups and nearly identical requested bytes. Qualifying workspace accounting at those actual pass sizes and beyond 16K is a concrete next experiment, with allocation, cancellation and exact-state gates before adoption. No unmeasured gain is assigned to it. 3. **Desktop still explicitly disables MTP.** `apps/macos/Runtime/Performance.swift` requests MTP off and a 32,768-token context. Engine/CLI automatic MTP gains therefore do not imply Desktop MTP gains. A change needs the app's feasible-budget guard reviewed too, since it uses a 33 GB base ceiling. This audit does not silently change that default. 4. **Long in-memory checkpoint retention is not always affordable.** `Generate.swift` selects the deepest boundary; `PrefixCache.swift` then charges stepped actual capacity plus the active reservation while preserving valuable other conversations. Both v0.2.22 and v0.2.23 refuse the 8K memory snapshots in the captured profile. This limitation predates the release. A shallower boundary might fit but also splits expert reads, so the deciding metric must include first-request cost and subsequent reuse. The disk diagnostic proves the existing persistent path can avoid this miss. Ordinary eligible Desktop Homes attach disk state; private/thinking workflows preserve their separate policy. 5. **Read admission depends on live allocation history.** Warm misses can contract groups. MLX's `set_cache_limit` changes a ceiling without immediately purging its existing pool; admission uses current physical footprint. The diagnostic establishes the symptom, not that clearing the allocator is the cure. Earlier cold-clear failures do not answer the warm-miss question. A bounded trim/admission experiment should measure the full sequence and cancellation cost before changing defaults. 6. **Current reuse is not all newly introduced.** v0.2.22 already has aligned and complete-prompt caching and the same deepest-boundary selection. v0.2.23 changes expert-read checkpoint boundaries, disk pass provenance, app attachment and live thinking-to-answer continuation. The new checkpoint ablation cannot be advertised as its incremental release speedup. CLI help also still describes an approximately 120-expert MTP floor while the planner constant is 76; this is a documentation discrepancy, not measured performance. Prior experiments on larger compute passes, selected scalar attention, GPU compaction, allocator clearing and fixed piece writes remain rejected or unqualified as recorded in [[records/measurements/prefill-opportunities-2026-09-21]]. Reviewed concurrent decode experiments also failed to justify defaults; their snapshot is in this audit's archive, separate from these new measurements. #### Earlier evidence retained, not counted as new runs [[records/measurements/prompt-speed-qualification-2026-09-21]] has three clean 8K pairs with 15.2% less prefill from guarded larger reads. [[records/measurements/fused-prefill-integration-2026-09-21]] has three clean pairs with 5.22% less prefill for the complete backend integration. [[records/measurements/mtp-prefill-policy-2026-09-21]] has three clean 16K/10 GB MTP pairs with 65.40% less prefill and 64.53% less request time from the same-backend policy comparison. These percentages describe different baselines and cannot be multiplied or pooled. The published release acceptance already covers independent numerical references, raw state/logits/continuation/output equality, 4,878 bounded writes, persistent restart, native app lifecycle and live thinking-to-answer handoff. Those checks were reviewed here, not rerun or counted in the 123 requests. No new repeated wall-time claim is made for thinking-to-answer UX. This audit does not qualify other hardware, the full approximately 32 GB automatic profile, unrestricted long sessions, whole-answer quality across all prompts or a complete GPU trace decomposition. Raw evidence: [[sources/runs/2026/09/2026-09-22-published-prompt-speed-audit]] and [[sources/runs/2026/09/2026-09-22-published-prompt-speed-audit-excluded]]. ### Public speed tables checked against the latest evidence The README's warm-reply ranges and hardware guide's measured decode table already contain the latest qualified public warm-decode reference: 15.86 tok/s on 0.2.19 at a 22 GB target. The recent installed-release audit does not replace that held-out sustained-decode benchmark. Short requests have no consistent gain and long first-request studies mostly cap output at one token. A recent prefill or cache percentage must not inflate the decode table or community reports. The public surfaces were missing the qualified prompt-policy and reuse results. README and HARDWARE now show the same scoped table: inventory/MTP prefill arm medians 155.22 to 53.94 seconds with 65.40% median paired reduction, and prose follow-up request medians 30.73 to 4.42 seconds with 85.62% median paired reduction. Each comparison has three clean pairs, matching generated IDs, one tested binary per comparison, and a 10 GB target on the 48 GB M5 Pro. The first is pre-release policy qualification; the second tests the installed release. Neither measures the whole release delta. The prose number covers the follow-up only and compares existing reuse against disabled checkpoints. Sources: [[records/measurements/mtp-prefill-policy-2026-09-21]] and [[records/measurements/published-prompt-speed-audit-2026-09-22]]. Their original eligibility and exclusions remain unchanged. The hardware guide links complete methods and explicitly excludes contaminated or insufficiently repeated timings from the public table. Read-only simulations of the installed 0.2.23 binary reproduce every current automatic memory/context row and the documented rounded full-window wait estimates. The planner still uses its historical reference curve. The narrow new experiments do not qualify a replacement curve across pass sizes, position, hardware, MTP and memory budgets, so no estimator or runtime setting changes. The guide now states that calibration limit explicitly. Raw plan verification: [[sources/runs/2026/09/2026-09-22-speed-tables-planner-review]]. ### Installed-release speed calibration and benchmark-profile correction No new idle-machine speed baseline or estimator calibration is qualified by this attempt. The installed release produced complete answers, but concurrent host activity and an incomplete repeat matrix prevent updating the README's throughput headline. The same evidence does identify a documentation error: a historical 22 GB controlled benchmark was described as though it measured today's automatic configuration. ## What ran The prospective protocol requested three rounds across eight existing public code, reasoning, prose, structured-output and dialogue fixtures. These are reused historical fixtures, not newly held-out prompts. Each prompt had a 128-token warmup followed by a natural answer with a 1024-token ceiling. The installed 0.2.23 binary used a fixed 22 GB total budget, normal prefix caching, automatic MTP/lookahead and planner-selected 2048-token passes. The effective expert pool was 3531 slots, or 73.5625 experts per layer. No inference implementation or optimization setting changed. The attempt completed 27 requests, including 13 naturally finished answers, before being interrupted during its second round. Raw response replay passes for all completed requests; the five applicable narrow arithmetic/JSON checks pass. These checks do not establish general answer quality or release parity. The maximum native lifetime process peak was 18.574897632 GB. The original functional pilot is separate and cannot become a timing anchor. One identical prose answer read 10.96 tok/s in the first round and 6.32 in the second, with the same output IDs, draft acceptance, forward-pass count and nearly identical expert reads. Active audio/video/browser work was subsequently observed, and the GPU remained busy after the owned model stopped. This supports refusing an idle-machine calibration; it does not prove exactly which app or mechanism caused every timing difference. All original automatic eligibility verdicts remain intact, while the separate population policy excludes the entire incomplete attempt from prospective idle calibration. No slow observation is silently removed to improve a median. Raw evidence: [[sources/runs/2026/09/2026-09-22-release-speed-calibration]]. ## Corrected benchmark interpretation The historical 15.86 tok/s result on 0.2.19 remains a valid paired forecast comparison. Its frozen protocol forces 256-token passes, prefix caching off, two drafts, adaptive speculation off and the draft-tail experiment off. That leaves about 100 experts per layer at the 22 GB budget. Normal-cache 0.2.23 instead plans 2048-token passes and about 74 experts per layer at the same budget. The memory budget alone does not identify an equivalent runtime configuration. README, HARDWARE and ENGINEERING now distinguish that controlled benchmark from current automatic behavior. The historical rough speed ranges retain their mixed-version, chip/SSD and configuration assumptions; they are not newly calibrated ranges. The earlier claim that the current 32 GB automatic plan had been measured directly is corrected. Historical records and their original results are preserved with a clarification, not silently rewritten. ## Prospective measurement controls The revised harness records anonymous background CPU totals and device GPU utilization. Before model launch it requires a continuous nominal, quiet interval; between requests it checks idle GPU activity as well. While the model runs, aggregate GPU use is diagnostic because it includes the model itself, and background CPU remains screened. Thresholds are benchmark screening choices, not physical limits or complete proof of isolation: at most 5% idle GPU utilization, 50% total background CPU and 25% for any one background process, with one core represented by 100%. Samples are taken roughly every two seconds. No process names, arguments or user activity content enter those captures. It also records the actual server memory plan before and after each request, rejects changes within a request, and does not pool different effective plans. Fixed profiles require measured reclaimable memory above the target plus 3 GB. Adaptive profiles require their expected physical peak plus 3 GB to fit both the independent VM reading and the planner's availability reading, with the declared ceiling retained. The production governor is not disabled to manufacture an automatic result. The original process-pageins-v1 timing screen and strict global no-swap sensitivity remain separate. Older studies retain their frozen verdicts. Analysis groups actual cache hits and misses separately, checks repeated generated IDs, and tests family-held-out estimate corrections without modifying the planner. The frozen prospective 31K study must check the read-policy boundary before any 16K result is generalized to a full context window. ## Still required Finish the quiet 22 GB repeat suite; measure the planned 10/16/24 GB and MTP-off profiles; measure new code/prose prompts at 2K, 8K, 16K and near 32K, including misses and repeats; and run the actual adaptive CLI and Desktop engine profiles when physical headroom permits. The observed actual-default preflights did not meet the measurement's physical headroom requirement, and no automatic model process was launched. Desktop's engine policy through HTTP remains distinct from Desktop UI latency. Other Apple Silicon hardware still needs actual access. The registered Linux server and Windows machine do not qualify this Apple engine. No community report or simulated memory plan is relabeled as a new measurement. The production estimator remains unchanged: its constants also influence automatic context/workspace decisions, so changing a display number alone would silently change policy without a complete measured envelope. ## Validation and final attempt status Both prospective idle pilot attempts ended without loading a model. The first exhausted its 900-second readiness window; its full observations are preserved in `idle-smoke-v2/`. A second attempt was stopped after continued background CPU work was independently identified as OS media-analysis activity. It left no model process. These are measurement-environment refusals, not inference failures. Validation: 13 analysis tests, 14 host-load parsing/gate tests, and independent replay of all 27 completed response streams pass. The five applicable narrow arithmetic/JSON output checks pass and are never used to select timing observations. The revised live-plan capture still needs its real-model functional pilot before a v2 timing campaign can qualify. Larger actual-default preflights failed the prescribed headroom test; no adaptive server launched. ## Windowed readiness correction The earlier pointwise host-load rule rejected ordinary interactive desktop bursts and prevented useful measurement. A separate v3 protocol now evaluates sampled load over the readiness or request window. This is a prospective timing screen for an interactive Mac, not proof of complete host isolation. Historical v1/v2 protocols, raw observations and verdicts remain unchanged. The v3 window permits mean total background CPU of at most 100% of one core and mean largest-process CPU of at most 50%. Before requests, mean device GPU utilization must be at most 5%. No more than 20% of samples may exceed the burst thresholds of 200% total CPU, 100% largest-process CPU or 20% idle GPU. Aggregate GPU during model work remains diagnostic. CPU values from ps are decaying estimates, not exact interval accounting. These are explicitly chosen screening limits, informed by the earlier false readiness refusals, not measured performance boundaries. They were frozen before any v3 model request. The independent memory, native process footprint, thermal, power, model-lock, known competing-job and paging checks remain in force. Readiness still requires two minutes of continuous memory and thermal eligibility before loading a model. Between-request readiness remains 15 seconds. Each readiness attempt is now bounded to five minutes. Windowed CPU/GPU screening does not reset the whole readiness interval for one brief desktop spike; sustained competing work still fails it. Raw snapshots retain the stricter v2 pointwise flags for sensitivity analysis. The host-load suite passes 23 tests, including burst tolerance, sustained-load rejection, unavailable telemetry and unchanged thermal/paging exclusions. The existing 13 analysis tests also pass. A simulated loopback metadata endpoint confirms the nested runtime-plan extraction, but does not replace the required real-model pilot. Evidence and the completed attempt status: [[sources/runs/2026/09/2026-09-22-release-calibration-load-screen-v3]]. The runtime, released binary, estimator and public speed tables are unchanged. No new speed gain is established by a benchmark-harness correction. ## Native pilot after the readiness correction After the 10 GB pilot refused insufficient memory, a separately frozen 8.1 GB functional pilot passed both requests on the installed 0.2.23 release. This qualifies the v3 harness live-plan capture on that profile. Both requests retained 640 slots, 256-token compute passes, a 32,768-token window and MTP off; before/after runtime plans agreed with the native effective pool. Exact raw response replay and token/timing accounting pass. Neither request observed global swap activity, and the native lifetime peak was 6.066900808 GB under the 8.1 GB ceiling. The model exited cleanly and was reaped. The responses were deliberately capped at 16 and 32 output tokens. They validate request/capture wiring, not complete-answer quality or a speed baseline. The pilot remains excluded from calibration, so its token rate cannot update README headline throughput or estimator anchors. A prospective analysis extension now evaluates full-prefill misses using leave-one-prompt-family-out corrections within the same realized plan. Cached tokens, pilots, unregistered populations, changing plans and unequal outputs cannot train the correction. Nineteen analysis tests pass. This is an exploratory diagnostic, not a fitted production policy or evidence of transfer to other prompt lengths. The subsequent 10 GB timing attempt and its exact disposition are preserved with the pilot at [[sources/runs/2026/09/2026-09-22-release-calibration-native-pilot-v3]]. Earlier memory refusals and frozen v1/v2/v3 attempts remain unchanged. No runtime optimization or estimator change is established by the pilot. ## Completed 2K installed-release measurements The subsequent v3 timing phase completed three fresh-server rounds and one prospectively declared code supplemental repetition. Thirteen of fourteen requests qualify under the primary screen; all response captures replay exactly. The measured ranges and cache reuse results are now published separately at [[records/measurements/release-prefill-2k-2026-09-22]], with immutable evidence at [[sources/runs/2026/09/2026-09-22-release-calibration-2k-v3]]. Server history changes the read batches even at an unchanged memory plan, and the strict global no-swap subset is insufficient for a repeated first-read claim. The broader profile matrix, history-independent ETA calibration and full-answer decode baseline remain unfinished; the production estimator is unchanged. ## Longer-prompt continuation and remaining limits A separate prospective 8K/16K code/prose phase completed eight requests in its first round. All raw captures replay exactly and the maximum native lifetime peak was 8.250117504 GB under the 10 GB target. Two requests failed the original thermal screen. One further repeat is excluded in derived analysis because a documentation checkout by this task overlapped it; the exclusion was registered before inspecting its result and original verdicts remain unchanged. Five observations remain eligible, with no fixture reaching three repetitions. Every 8K/16K exact repeat in this round reused zero prompt tokens. The normal runtime cache was enabled, but its 13,382-token-unit retention allowance could not keep the requested checkpoint beside the active future sequence reservation. The generator selected the last eligible pass boundary and checkpoint admission refused it. These are real limits in this measured configuration, not evidence that the cache is disabled or that larger-budget profiles have the same result. Initial read batches also differed from later reads despite an unchanged nominal plan. The second-round readiness attempt timed out before model launch: final-window background GPU averaged 15.90% against the frozen 5% screen. Thermal state had returned to nominal and memory headroom passed. All owned processes were reaped. Evidence: [[sources/runs/2026/09/2026-09-22-release-calibration-long-v3]]. Two follow-up hypotheses deserve matched experiments: trim genuinely unused allocator buffers before choosing an optional read scope when doing so could buy a larger scope, and retain a smaller existing pass-boundary checkpoint when the deepest one cannot fit. MLX cache-limit assignment does not itself immediately trim the cache; the allocator enforces the cap during allocation. Neither idea has been implemented or shown faster in this attempt. Both must preserve numerical pass boundaries, physical and reservation limits, active state, MTP and persistent-cache lineage. Do not turn these hypotheses into a speed claim. ### Installed-release 2K prompt timing and ordinary cache reuse The installed v0.2.23 release now has a repeated 2K prompt measurement on the 48 GiB M5 Pro with a 10 GB process target. This is a small-budget configuration on that Mac, not a measurement of a different Mac or a new automatic-default decode baseline. Normal prefix caching and planner-owned settings selected 961 expert slots, 256-token compute passes, a 32,768-token context and MTP off. Three fresh-server rounds rotated code and prose order. One prospectively declared supplemental code repetition followed a background-CPU exclusion. All fourteen requests completed and their raw responses replay exactly; thirteen qualify under the primary screen. Every first read and repeat of a fixture emitted identical sixteen-token output IDs. These capped replies do not establish complete-answer quality. The maximum native lifetime footprint was below the 10 GB ceiling; consult the raw per-request peaks rather than treating the configured target as measured usage. | Prompt | Eligible first reads / repeats | First-read prefill range | Median repeated request | |---|---:|---:|---:| | 2K code | 4 / 3 | 13.86–28.11 s | 2.79 s | | 2K prose | 3 / 3 | 14.91–26.51 s | 3.22 s | All formatted prompts contain 2,061 tokens. Mixed-order median first-read prefill is 14.8590 seconds for code and 26.1293 for prose; median full first-request time is 17.1836 and 28.4049 seconds, respectively. Three eligible first/repeat pairs per fixture give median paired end-to-end reductions of 85.4666% and 88.3625%. Those are the benefit of actual reuse in this release, not a matched feature-off experiment, a whole-release speedup or a new reply-generation rate. First-repeat comparisons also include warm process/OS caches. History matters. Server-first code reads used a 2,048-token read group followed by the thirteen-token tail; the later code miss used 256-token read groups. Prose shows the same order effect. Compute passes stayed at 256 tokens and output IDs were unchanged. Immediate repeats reused either the complete prompt or its 2,048-token boundary. The artifact reports server-first and later-miss observations separately; the mixed-order medians are descriptive summaries of this prescribed sequence, not history-independent ETA anchors. The primary screen is prospective v3 desktop load screening plus nominal power/thermal state, no host swap-outs and bounded process page-ins. Global swap-ins occurred on several runs: only one code first read and no prose first read passes the stricter global no-swap sensitivity. These results must not be described as a fully isolated or entirely swap-free calibration. The one automatically excluded repeat is retained, including its fast observed latency. No observation was selected by speed. A simple family-held-out multiplicative correction transfers poorly between these prompt/order mixtures. It is an exploratory diagnostic and confounds family with the prescribed server history. The production estimator remains unchanged. Larger prompt lengths, more memory profiles, real adaptive defaults and full-answer decode remain separate qualification work. Evidence: [[sources/runs/2026/09/2026-09-22-release-calibration-2k-v3]]. ## Issue 21 current-code reassessment: shipped repairs and remaining edge cases The earlier confirmed issue-21 repairs are present in published v0.2.23 and current production code. Publication and acceptance of the exact final release bytes are complete. This supersedes those two outstanding-work statements in the historical [[records/measurements/issue21-long-context-qualification-2026-09-21]], without changing its frozen evidence. A fresh review found two additional actionable edge cases, reproduced but not repaired in this audit. Evidence: [[sources/runs/2026/09/2026-09-22-issue21-current-review]]. Reviewed HEAD is `13934f50486c6211a4d70687c79ab9960436d8fa`; production `Sources/Slotstream` is identical to v0.2.23. Existing uncommitted diagnostic/CLI experiments were preserved. This audit made no production changes. ## New findings **P2: nullable string type arrays lose streaming and type fidelity.** `ToolDefinition.schema` in `Sources/Slotstream/ToolCallSplitter.swift:560` recognizes scalar `type` and the special nullable `anyOf` form, but treats `{"type":["string","null"]}` as unknown. Type arrays are a [valid JSON Schema representation](https://json-schema.org/understanding-json-schema/reference/type). The actual parser probe emitted 200 incremental argument deltas for the scalar and `anyOf` forms and none before close for the type-array form. It also converted the declared string `00123` to integer 123. Normalize the equivalent single-non-null type forms, preserve declared string bytes and add incremental, numeric-string, nullable-value and truncation cases. Genuine unions need an explicit conservative policy. These results establish buffering and wrong coercion; this audit did not test capped wire output for this schema shape. **P2: longest cache branch can hide a compatible retained branch.** `Engine.swift:565` asks `PrefixCache.peek(extending:)` for the longest descendant, then abandons splicing at line 571 if that descendant's assistant reply does not match the submitted history. Neither it nor the persistent selector tries a shorter compatible branch. A metadata-only probe retained both branches and reproduced selection/rejection of the wrong one through real library helpers, including the persistent policy. For omitted reasoning, fallback rendering loses the exact saved IDs and can cause avoidable rereading. This is a cache-reuse defect, not acceptance of an incompatible numerical state, and not a proven cause of the reporter's process exits. Select compatible assistant-turn candidates before committing to a descendant, across RAM and disk, and add branched-history regression coverage. Preserve the existing public longest-entry lookup contract if adding a separate candidate API. ## Original report, one item at a time | Reported behavior | Current disposition | | --- | --- | | Whole-context reread on follow-up | Linear omitted-reasoning conversations and persisted exact IDs now work, including the fresh 51K/restart run. The branch-selection edge case above remains; legitimate memory/provenance misses must still reread. | | Long tool argument delivered only at completion | Incremental string parsing and bounded scanning shipped and pass. Nullable type arrays remain a confirmed hole. | | Tools present but unused | Fresh bounded cap and natural completion pass, including a 16000-token allowance. No actual 16000-token generation was run. | | Streaming cap lacks terminal result | Fresh plain/tool cases return length, requested usage and DONE. The original generic missing-terminal failure did not reproduce independently. | | Omitted max_tokens becomes 512 | Current Chat Completions uses the advertised output budget, capped by remaining request context, rather than the historical 512-token cap. This was already repaired after v0.2.20. | | store rejected | store:false and declared compatibility fields are accepted. store:true and unsupported semantics retain explicit refusals. Arbitrary unknown fields are not silently accepted. | | Ollama tools rejected | Intentionally unsupported; refusal now directs clients to the OpenAI endpoint. | | Sparse progress and misleading ETA | Time-based prefill progress, recent-rate estimates, request-phase heartbeats and cache diagnostics are present. | | Intermittent daemon exits and prolonged busy stall | Still unreproduced. The exact reporter client and complete logs remain unavailable in the issue. | | Pressure-related long-context tail collapse | Not qualified as fixed. This review ran no induced-pressure workload or clean timing comparison. | ## What newer changes accomplished [[records/measurements/release-0-2-23-published-2026-09-22]] establishes public distribution, identical CI/public/installed bytes, 32 native acceptance gates, a 13-assertion 16K MTP gate and 31 installed-serving checks. This audit reopened the raw capture, verified all 134 members and inspected its results; it did not rerun that entire battery. The first release attempt's unexplained CLI exit code 2 remains preserved, followed by successful isolated and complete reruns. It cannot establish the cause of the reported daemon failures. The release adds automatic expert read grouping, the pinned fused MLX backend, MTP phase accounting/bounded writes, cache arithmetic provenance and backend/environment identity checks. These improve qualified prompt workloads and prevent unsafe reuse. [[records/measurements/published-prompt-speed-audit-2026-09-22]] separately demonstrates reuse savings and records long-prompt timing exclusions. The expanded read-sharing and reduced workspace accounting in `PrefillReadPolicy` and `ContextMemory` are deliberately confined to the measured 256-query/16K-key envelope and qualified platform. Outside it the established conservative accounting remains, even if a kernel itself still fuses. Removing those limits without allocation, parity, cancellation and timing qualification would not be a justified fix for the 51K tail. ## Fresh verification The current build passed all 57 T0 groups, 29659 assertions and runtime checks. The live issue-21 suite passed caps, incremental truncated string arguments, omitted-reasoning reuse, disconnect survival and restart. All 24 OpenAI compatibility checks passed. On the new backend, the long conversation grew from 30288 to 51364 to 51407 prompt tokens. Reuse advanced from 30208 to 51200 tokens. The last request reread only 207 tokens and returned identical answer, reasoning and usage after restart. The configured window was 131072, target 13.5 GB, MTP and vision off. Final replies allowed 16000 output tokens but naturally produced 21. These are functional results, not full-window, long-generation or pressure/timing qualification. All owned servers were reaped and the model lock was free. Required follow-up is to repair the two reproduced edge cases with focused regressions, then repeat the affected serving/cache checks. The issue-specific live scripts are documented but are not directly called from `verify.sh`, `e2e_release.sh` or the workflows at this checkout; include durable coverage of these cases in the accepted regression path. Obtain an exact recurrence/client trace before claiming a cause for the original exits or stall, and qualify any broader long-context performance change separately. Passing suites establish the checked behaviors, not a general bug-free guarantee. ## Issue 21 nullable tool strings and branched conversation reuse repaired Both additional defects from [[records/measurements/issue21-current-review-2026-09-22]] are repaired and covered by durable regressions. Changes are local and unreleased. Raw evidence was captured first in [[sources/runs/2026/09/2026-09-22-issue21-edge-fixes]]. ## Repairs `ToolDefinition.schema` recognizes a single non-null type in a JSON Schema type array, including both orders of string and null. Nullable strings now use the same incremental parser path as scalar strings and the existing nullable anyOf form. Numeric-looking strings retain their leading zeros and string type. Genuine unions, malformed members and unresolved declarations retain conservative handling. The parser's existing interpretation of literal string contents, including the word null, is preserved. `PrefixCache` now offers an internal matching lookup that considers transcript metadata from both memory and disk before choosing the longest accepted candidate. Engine validates the current assistant turn through that lookup, so an incompatible longer branch cannot hide a shorter matching branch. Candidate validation occurs outside cache locks and propagates request cancellation. The public longest-entry lookup remains available with its prior behavior. No numerical checkpoint admission, arithmetic provenance, token-equality requirement, expiry/identity filter or memory budget was relaxed. The issue-specific live suite is now part of Tools/verify.sh through Tools/issue21_e2e.py. It owns bounded servers, preserves wire evidence and checks exact restart replay. Its branch fixture uses deterministic supplied histories; the parser/output regressions also run in the ordinary catalogue and CI. ## Fresh results on the repaired build Executable SHA-256: `3e3de25a36f9265e5b11472452f311650143dd36f4e71f938c3794c6d0fc8a06`. Every recorded source hash still matched the checkout at functional capture. | Check | Result | | --- | --- | | Catalogue | 57 T0 and 16 T1 groups passed, 31907 assertions, no failures or skips | | Runtime diagnostics | Passed | | Nullable-string streaming | Character-split round trips and 16000-fragment output fixtures passed; real HTTP truncation streamed incrementally and ended with length, requested usage and DONE | | Branch regression | Same requests fail on the pre-fix executable and pass on the repaired build; seed answers and reasoning match exactly across builds | | OpenAI compatibility | All 24 checks passed | | Process and persistence | TCP-reset survival and identical answer/reasoning/usage after restart passed | | Numerical reuse | Exact-prefix check passed: generated tokens and prompt logits match cold reads bit for bit; edited history rebuilds | | Long conversation | Prompt grew from 30288 to 51364 to 51407 tokens, advancing reuse from 30208 to 51200; final reread was 207 tokens and restart was identical | The branch reproduction is stronger than the original metadata-only probe. A longer retained branch previously made Engine discard a compatible branch's saved reasoning, yielding a 1541-token follow-up prompt and 1024 cached tokens. The repaired request preserves the 1597-token prompt and reuses 1280 tokens. This demonstrates the repaired selection failure; it is not a universal performance benchmark or a claim that every cache miss is erroneous. ## Scope and remaining uncertainty Ordinary HTTP checks used an 8.1 GB target and context 32768; numerical equality used 10 GB. The long test used the established 13.5 GB exception at configured context 131072 with a real target-plus-3 GB preflight. MTP and vision were off for these serving workloads. All owned model servers were stopped and reaped. The largest actual prompt was 51407 tokens. A 16000-token allowance produced a short natural answer; it was not a 16000-token live generation. The full release battery previously passed on published v0.2.23, but was not rerun for this patch. Fresh coverage is the complete T0/T1 catalogue, runtime, targeted native equality, live OpenAI/issue-21 tests and long-context restart checks. The original intermittent daemon exits, prolonged CPU-busy stall and pressure-related tail slowdown remain unreproduced. These repairs cannot certify a cause or cure for those observations. No memory-hog experiment, reporter-hardware reproduction or clean throughput comparison was run. Existing diagnostic/CLI experiments were preserved; publication, version bump and installation were not part of this repair. ## Final validation The full static suite passed on the repaired executable, including syntax, harness, catalogue-support, download/installer, planner, memory-override and documentation checks. The source audit still matched the frozen build. Final evidence and process cleanup are linked from [[sources/runs/2026/09/2026-09-22-issue21-edge-fixes]]. The store retains its two historical log warnings; this repair does not rewrite historical logs. ## Publication follow-up The original repair capture above remains unchanged. These fixes were subsequently published as v0.2.24 on September 23. Full release qualification and public installation passed: [[records/measurements/release-0-2-24-published-2026-09-23]]. Post-release comparison and timing limits: [[records/measurements/release-0-2-24-performance-2026-09-23]]. ### v0.2.24 published, installed and accepted **v0.2.24 is public, installed and accepted.** It ships nullable-string tool streaming/type preservation, compatible conversation-branch selection and their regression coverage. Both repairs operate automatically. This patch adds no inference kernel or new tuning switch. The separate local decode experiments are outside the release. Release: [v0.2.24](https://github.com/carloslfu/slotstream/releases/tag/v0.2.24), published 2026-09-23T05:24:36Z from `814da126894959dfdf3e0a01babc43468ccee5c2`. The CI candidate, public archive and installed executable match exactly. Archive SHA-256: `2c942b6706febccd4e3fdf1930ba57347a86c21ae47ecc1e6b821323e8c36bde`. Executable SHA-256: `bbdfcaffa8959ac1ca3e39d1f804cc4491aaa98dd8f649d5f239713a6ba10499`. | Acceptance | Result | | --- | --- | | Exact-commit hosted CI | Engine, instrumented coverage, external library consumer, Mac app, docs and context contracts passed | | Engine catalogue | 73 groups, 31,907 assertions, no failures or skips, in release and instrumented builds | | Full native battery | 33 top-level gates passed; includes new issue-21 streaming/branch/restart, OpenAI compatibility 24/24, quality 15/15, robustness 74/74 and vision serving 25/25 | | Public distribution | Preserved CI archive published, public checksum/provenance verified, public installer upgraded the standard installation | | Installed serving | 31/31 with a 10 GB target and MTP on; owned server reaped | [[sources/runs/2026/09/2026-09-23-release-0-2-24-published-and-installed]] retains the commands, source/build identity, native logs and references, workflow output, installer and process-cleanup receipts. The historical backend reference remains visible alongside the passing current-backend reference; no tolerance was widened. These are functional acceptance results, not clean-host speed claims. Post-release comparison results are in [[records/measurements/release-0-2-24-performance-2026-09-23]]. The original intermittent daemon exits, CPU-busy stall and pressure-related tail slowdown from the issue review remain unreproduced; these two verified repairs do not establish their cause. ### v0.2.24 post-release responsiveness and reuse comparison **The released fixes improve delivered tool arguments and compatible branch reuse. No general latency or tokens-per-second gain is qualified.** The raw comparison in [[sources/runs/2026/09/2026-09-23-release-0-2-24-performance]] contains 72 requests across six sessions and three interleaved old/new pairs, using the installed v0.2.23 and v0.2.24 binaries. ## Repeated functional findings | Case | v0.2.23 | v0.2.24 | Interpretation | | --- | --- | --- | --- | | Nullable tool argument, capped at 256 generated tokens | 0 argument chunks; no argument value delivered | 239 argument chunks during generation in every round | Incremental client delivery now works; truncated output remains incomplete and must not be executed | | Nullable numeric-looking string | `{"content":123}` | `{"content":"00123"}` in every round | Preserves the declared string and its leading zeros | | Branch alpha follow-up | 1541 prompt tokens, 1024 cached, 517 uncached | 1597 prompt tokens, 1280 cached, 317 uncached | 200 fewer uncached tokens, about 38.7%, while retaining previously omitted reasoning | | Code/prose controls | Frozen input and output IDs | Exact match in all 12 comparisons | No token drift observed in these controls | The scalar-string control streams 239 chunks in both releases. The nullable change brings the equivalent schema onto that existing path. The branch result is a reduction in reread input, **not 38.7% faster inference**: the repaired request contains more correct reasoning, and output length changes from 30 to 24 tokens. It is not an equal-input model-throughput comparison. Branch seeds match across releases. Restart preserves the repaired answer, reasoning, prompt count and output count, with cached input advancing to 1536 tokens. The previous release's restart also reaches 1536 cached tokens but still omits the reasoning, leaving only 5 uncached tokens instead of the repaired 61. Faster completion of that incorrect shorter prompt would not demonstrate better reuse correctness. ## Method and timing limits The M5 Pro 48 GiB Mac ran a fixed 10 GB target, context 32768, MTP off and vision off. Two different 2048-token raw prompts, code and prose, were tested as first reads and repeats with 128-token output windows. The branch fixtures, tool schemas, ordering, readiness policy, process/memory sampling and analysis were frozen before the first request. OS file cache was uncontrolled. Installed functional acceptance separately tested MTP on; this comparison establishes no MTP performance gain. The shared-desktop screen accepted 68/72 requests. The stricter no-global-paging screen accepted 40/72. Required clean matched pairs per family were three; observed eligible counts were: | Family | Eligible matched pairs | | --- | --- | | Code first read / repeat | 0 / 1 | | Prose first read / repeat | 0 / 1 | | Scalar / nullable truncated argument | 0 / 1 | | Nullable numeric string | 2 | | Alpha / beta seed | 1 / 1 | | Branch follow-up / restart | 2 / 2 | No measured family meets the timing claim threshold. The one-token warmup mechanically meets the analyzer's pair count, but is explicitly excluded from all speed claims. All loaded/paging captures remain preserved rather than cherry-picked. The final audit replayed all 72 captures, checked the frozen inputs, and verified all 12 server processes exited zero and were reaped. The README throughput anchors and hardware estimates remain unchanged. Establishing a new general speed percentage requires three clean matched pairs with equivalent work; the corrected branch alone cannot provide that comparison. This patch contains no new inference computation optimization. The earlier long-prompt and fused-attention gains remain scoped to their own studies and must not be combined with these counts. ### Sevra desktop speed: short-turn rereads and disabled MTP **The actual development app wrote substantial replies at 10.5 to 12.6 tok/s, but spent 4.7 to 8.1 seconds before its first token.** Prompt processing dominates the delay in the short final reply. These are observed live-session timings, not new clean-host throughput anchors. Raw receipts: [[sources/runs/2026/09/2026-09-23-sevra-app-speed]]. | App request | Output tokens | Writing tok/s | First token, excluding load | Prompt tokens read / reused | | --- | --- | --- | --- | --- | | Bicycle explanation, first request | 258 | 12.53 | 4.68 s | 143 / 0 | | Rain explanation, warm | 304 | 10.46 | 6.94 s | 451 / 0 | | RAM/SSD explanation, warm | 271 | 12.58 | 5.65 s | 808 / 0 | | Brief greeting, warm | 2 | Not a steady-speed sample | 8.13 s | 1101 / 0 | The first request also loaded the model for 9.05 seconds. The final greeting spent 8.10 seconds reading context and 0.39 seconds generating its two tokens. The app's displayed rate correctly excludes prompt reading and model load; the tiny final denominator should not be read as steady throughput. All runs used a 33 GB automatic budget, context 32768 and thinking off. Expert hit rates for the substantial replies were 92.4%, 94.0% and 93.5%. No thermal or low-power restriction was observed. Some global swap-ins occurred, so these results are diagnostic. ## Concrete integration gaps 1. The launched app is an older development bundle, built September 21. Its embedded engine version declaration is 0.2.22, with additional then-uncommitted work. The public CLI update to 0.2.24 does not rebuild or replace the app's statically linked engine. Its manifest differs from current engine inputs. This establishes a stale integration, not a measured causal slowdown from every changed file. 2. Desktop's `PerformancePolicy.plan` explicitly requests `mtp: .off`. That source matches the running app's recorded input. The app therefore does not get speculative decoding; this is independent of the Think longer control. 3. Numerical-safe reuse admits only compatible complete compute-pass boundaries. A short prompt that ends inside a large pass supplies no eligible continuation boundary. All four observed app turns reused zero tokens. The present equivalent 33 GB MTP-off engine plan chooses 4096-token passes, illustrating the mismatch between long-prompt efficiency and short-conversation reuse. The old app's exact pass size was not directly instrumented, so that current plan is explanatory evidence, not a reconstructed live plan. 4. Automatic readiness releases the model after about ten idle minutes in this configuration. That saves memory but introduces another load on the next message. The observed first load was 9.05 seconds. ## MTP opportunity and limits A separate installed-v0.2.24 comparison, with one warmed 33 GB engine and fixed expert pool, measured median plain 13.50 tok/s versus speculative 17.45 tok/s over three rotating pairs. Each arm repeated its own output, but plain and speculative text differed from token 18. The observed ratio is about 1.29. Global paging occurred during the experiment, so this is **not a qualified general speedup or an app improvement already delivered**. It shows that enabling and qualifying the app's automatic MTP path is worth testing. This loop-only comparison retains draft-head memory in both arms and does not reproduce the app's independent MTP-off allocation. The next implementation work should rebuild and verify the app against current engine inputs, qualify automatic MTP in its ordinary and phased-thinking paths, and measure a short-conversation checkpoint/pass policy without weakening numerical provenance. Improve readiness based on cold/warm latency and memory measurements. Do not silently trade cache correctness for a smaller first-token number or infer a universal optimum from these four requests. No app code or saved performance preference changed in this investigation, and the app was reopened after the diagnostic. ### Sevra desktop defaults: MTP, useful checkpoints and verified reloads The development app now uses the engine's qualified automatic MTP policy, a short-chat checkpoint schedule, foreground-aware readiness and session-scoped model verification. This addresses the integration gaps in [[records/measurements/sevra-app-speed-2026-09-23]]. Code and failed trials are preserved in [[sources/runs/2026/09/2026-09-23-sevra-app-optimizations]]. The policy and revision criteria are in [[records/decisions/sevra-app-speed-defaults-2026-09-23]]. ## Final matched-input comparison Three alternating pairs on the M5 Pro/48 GiB Mac using the pinned Qwen3.8-Flash-Next 4-bit model, MLX 0.32.2, and the 0.2.24 engine plus this development patch: 26 GB total budgets, 32,768 context, fixed 96-token replies and identical input histories: | Turn | Plain writing tok/s | New policy writing tok/s | Plain reading | New policy reading | New policy reused tokens | | --- | ---: | ---: | ---: | ---: | ---: | | 1 | 11.51 | 16.03 | 3.88 s | 3.91 s | 0 | | 2 | 11.90 | 14.59 | 4.32 s | 5.94 s | 0 | | 3 | 11.67 | 19.03 | 4.67 s | 3.40 s | 512 | The median sum of prompt-reading and generation time across the three turns fell from **37.80 to 30.76 seconds**, an observed **18.6% reduction**. This includes the second turn's checkpoint-creation cost. It excludes model construction, UI overhead and the extra cold parity replay. Each within-policy cached result reproduced the cold output IDs exactly. MTP/plain text can differ, so this is matched token work and inputs, not a claim of identical answers across decoding methods. These are live-machine diagnostic timings: global swap-ins occurred and thermal state varied between nominal and fair. No new README throughput anchor or hardware estimate is justified. The earlier three-arm screen found automatic MTP already supplied most of the generation gain. Fixed smaller batches were slower for longer inputs; the hybrid keeps larger passes there. The 9,295-token inventory took roughly 31 seconds to read across the screening arms, with correct answers. That screening omitted a small optional correction charge from planning; the final comparison includes it. ## What changed and what it costs - MTP activates automatically only when its optional head is present and the fully charged plan qualifies. Both the head and configured lookahead correction fit inside the app's displayed total ceiling. The automatic Desktop ceiling stays 33 GB; custom ceilings remain adaptive. Smaller budgets retain ordinary decoding. - Below 1,536 prompt tokens, compute passes are at most 512 tokens or the smaller live plan. This makes an eligible checkpoint available earlier. Longer prompts keep the planned schedule. The larger workspace reservation remains available; it is not silently spent on more experts. Creating a checkpoint can slow an earlier turn, and crossing schedules requires a fresh read when arithmetic is incompatible. Selection happens under the generation lock after a governor resize and is stable across a thinking continuation. - Automatic readiness retains an already-loaded model while the app is foreground and has a visible non-minimized window. Background inactivity starts a fresh interval. Pressure, power saving and sleep retain their release behavior. This does not preload the model merely because the app opened. - The first load still hashes the pinned files. Within one inference owner's lifetime, an unchanged APFS file-identity/size/mtime/ctime signature can reuse that successful proof. Changed or replaced files, optional-file changes, a new process and other filesystems require fresh hashing. No proof or private inference state is persisted by this mechanism. - Immediate reloads wait only the remaining 1.05 seconds after model release before reading real availability. XNU's one-second shared statistics cache otherwise undercounted newly freed memory and disabled MTP in two reproduced reload checks. No synthetic availability credit or safety bypass is used. [Apple kernel source](https://github.com/apple-oss-distributions/xnu/blob/main/osfmk/kern/host.c) documents the caching window. ## Correctness and integration All six real-model profiles passed: precise phase metrics with actual MTP drafts, MTP disk restore/private isolation/cancellation/tools/crossover, thinking with Answer now, complete cited Cedar document creation after exact synthetic-fixture review, 10-to-9 GB lifecycle behavior, and ordinary 10 GB disk reuse. The MTP reload restored 2,048 prompt tokens at the full 26 GB budget; its immediate load was 2.085 seconds. The small-budget warm disk hit read only 168 tokens. Session-proof tests reject same-size corruption even after restoring mtime, optional arrival/removal, file replacement, symlink retargeting and mutation during verification, and honor cancellation. Native UI/runtime regressions, the final policy check, release builds, source-manifest verification, static gates and 73 engine catalogue checks (31,907 assertions) passed. The broader pure-policy sweep accepted 85 plans and safely refused 215. Functional acceptance is independent of global paging. The source changes preserve the independent CLI defaults and do not add another public release tag. The later [[sources/runs/2026/09/2026-09-23-sevra-verification-exact-mtime]] check strengthens the corruption test: it preserves inode, size and the exact nanosecond modification timestamp while corrupting the contents. Rejection and all scoped lifecycle checks passed. This test-only follow-up leaves the measured production implementation unchanged. ## Rebuilt app observation The final native replay is complete after unlock. The first replay also exposed an idle cache-growth memory overshoot, which was fixed and retested. Exact captures, failed observations, source changes and test results are preserved in [[sources/runs/2026/09/2026-09-23-sevra-native-replay-growth]]. Earlier source captures remain unchanged. The final rebuilt development app completed the same synthetic bicycle/rain/RAM/greeting sequence. Its native UI displayed the persisted metrics. Automatic resolved to 33.000 GB, with thinking off: | Request | Answer tokens | Writing tok/s | First token, excluding load | Read / reused tokens | | --- | ---: | ---: | ---: | ---: | | cold-bicycle | 225 | 16.19 | 4.86 s | 143 / 0 | | warm-rain | 266 | 16.54 | 4.78 s | 418 / 0 | | warm-caching | 243 | 16.31 | 7.98 s | 737 / 0 | | warm-short | 2 | Tiny reply | 4.13 s | 490 / 512 | Initial model preparation took 9.08 seconds. After Release memory now in the same process, preparation took 1.28 seconds; prompt processing still took 9.76 seconds to first token. This short conversation did not meet the existing disk-save minimum. The full final sampled interval, including idle and reload, peaked at 29.65 GB within the displayed 33 GB ceiling. The initial replay peaked at 38.49 GB between requests and failed process-memory acceptance. Warm growth now preserves slot indices while appending capacity piece by piece, removes the full occupied-row gather, and checks the temporary allocation against both physical footprint and real availability. When it cannot fit, the existing warm cache stays usable and growth retries later. Exact-byte, generated-output parity, live governor recovery and the full static gates passed; the engine catalogue now has 31,919 assertions. This is an engine correction shared by Desktop and the CLI, with no extra switch. These native rates are single live-host observations with different answer lengths and availability. They do not replace the matched-input comparison above or justify a new public throughput anchor. Raw paging and thermal observations remain in the capture. ## Remaining costs First launch still pays pinned weight verification. Uncached or schedule-incompatible context still requires real prompt computation and expert reads. MTP adds head memory and speculative work, so smaller plans must keep the existing activation guard. Short-batch checkpoint creation has an up-front cost. No universal maximum, cross-hardware speedup, or free continuation after every short message has been established. Revisit these operating choices with repeated complete-workflow measurements and the same numerical, privacy, process-budget and lifecycle gates. ## A shared prefix at a prompt's own resume boundary: the head is upgraded, and the gate that fails without it **Outcome: a disk head that already holds the ids of a shared prefix is upgraded with the flag instead of being answered as present, and a shared save that is skipped or fails now says why.** Until this change, whenever a prompt's own resume boundary fell on the same pass end as the shared prefix of its own system prompt, the head was written first as the conversation's checkpoint, without the flag, and the shared save that followed found those ids and treated the file as already correct. In the live session that exposed it — an installed v0.2.23 server with the disk tier on, one agent session of 172 log lines, 16:34 to 16:53 — **not one line contains `shared`**, and the directory ended with four states, every one `"shared": false`. The 24,576-token head that conversation wrote at 16:39:34 (`saved 24576 tokens (795.2 MB written)`) was removed 11 minutes later as a redundant ancestor (`saved 26624 tokens (144.1 MB written, 707.8 MB of rows reused) in 0.16 s, removed 1 older file`). The tier reused that one conversation well — eleven restores of 795.2 to 851.8 MB in 0.07 to 0.26 s — and shared nothing between conversations. **Why both writes land on the same position.** A conversation's checkpoint is written where the request may resume, its last prefill pass boundary. A shared prefix is saved at the last pass end at or before the prompt's system-message boundary, the save grid the 2026-09-16 record measured. Whenever the system message ends inside the same pass cell as that last resume boundary the two positions are the same one, and the checkpoint write comes first. The existing-head check compared the ids, the draft cache and the pass layout but not the flag, so the shared save answered `.present`; the head stayed classed as the conversation's own and the ancestor removal a later save performs took it. **The change, in two parts.** `PersistentPrefixSave` answers `.present` only when the head on disk already carries this save's shared flag (`existing.shared || !shared`); otherwise it writes, so the flag is in the file and survives a reopen. Rows are still referenced through `State.persistedLineage`, so the upgrade rewrites one head and no rows: the live directory's heads are 115,908,592 and 115,897,787 bytes, against 795.2 MB for that state's first write. The second part is visibility: a shared save that is skipped or fails reports `kept no shared N-token prefix: ` or `failed to write the shared N-token prefix: ` on the tier's event stream, which `serve` prints as `prefix cache disk: …`; a successful shared save keeps printing `saved shared N-token prefix (…)`, the line whose absence was the first symptom. **Evidence.** `persistent-prefix-round-trip` (T1, in the CI catalogue) now runs the sequence on synthetic states: a checkpoint written at the boundary without the flag; the shared save that must upgrade it and reference its rows rather than write them again; a reopened cache that must read the flag from the head; two later turns that must not remove it as an ancestor; its rows restoring exactly at every step; and a shared save below the write minimum reporting its reason. It passes **110 assertions**, and the catalogue passes **73 groups with 31,919 assertions and 0 failed**. With only the one-line condition reverted and everything else kept, **seven of the same 110 assertions fail**: `the shared save that follows upgrades the head: got present, want saved`, `rewriting the head references its rows instead of writing them again`, `the head is shared: got 0, want 1`, `two later turns keep it instead of removing it as an ancestor`, `its rows still restore exactly: got nil`, `a reopened directory reads the flag from the head, not from memory`, `and restores the upgraded head: got nil`. The field session's arithmetic — a 25,282-token prompt whose system and tools block ended 24,993 tokens in, inside the 24,576-to-25,600 pass cell, so its last resume boundary was 24,576 — is the diagnosis this gate encodes; the prompt itself was not preserved, so the gate, not the log, is the reproduction. Raw output and the log lines: [[sources/runs/2026/09/2026-09-23-shared-prefix-boundary-upgrade]]. **Limits.** Functional acceptance on one machine shared with other sessions, and no timing is claimed. The real-weight `optimization-state-check --variant shared-prefix --tokens 2051` passes on the fixed build, but it runs with `alignedPrefixResume` off, so it has no checkpoint at the shared boundary and never reaches the collision: it is a no-regression check on the disk write path, not evidence of the repair. Its draft-head variant was killed by the system before printing one check (signal 9, memory pressure) and is not evidence either way. No live server has run the fixed build yet; what to look for is `prefix cache disk: saved shared N-token prefix` in the session log, `"shared": true` in `slotstream prefix-cache`, and the reuse a second conversation then gets. The 2026-09-16 shared-prefix record's numbers stand and it is not superseded: its shared prefixes were saved at boundaries that did not coincide with a checkpoint, which is the case it did not cover. The price of the upgrade is one extra head write on every colliding boundary, 115.9 MB here, once per boundary and only where the two positions coincide; the skipped or failed outcome is reported rather than counted, so a loss shows up in the log and not only in a request's statistics. ## The fixed build in a live server: a colliding shared prefix is upgraded, reused and kept **Outcome: in a live server running the fixed build, a shared prefix whose boundary collides with a conversation's own checkpoint is upgraded with the flag, another conversation starts from it, a deeper state does not remove it, and it survives a restart on disk alone.** This closes, for one machine and one plan, the limit [[records/measurements/shared-prefix-boundary-upgrade-2026-09-23]] stated — *No live server has run the fixed build yet* — and produces the three observables it named: `saved shared N-token prefix` in the log, `"shared": true` in `slotstream prefix-cache`, and the reuse a second conversation then gets. **The collision through the engine's own save path.** A 3,639-token request whose system block and whose own last resume boundary fall in the same 256-token cell writes its checkpoint at 3,584 without the flag and then the shared save at the same 3,584 boundary, which upgrades the head instead of answering it present: ``` [20:06:19] prefix cache disk: saved 3584 tokens (214.8 MB written) in 0.11 s [20:06:19] prefix cache disk: saved shared 3584-token prefix (115.7 MB written, 99.1 MB of rows reused) in 0.03 s ``` The rows come through `State.persistedLineage`, so the upgrade writes 115.7 MB against the checkpoint's 214.8 MB and no rows are rewritten. The listing then reports `{'tokens': 3584, 'shared': True}`. The field session that exposed the bug logged no line containing `shared` in 172 lines and ended with four states all `"shared": false`. **A second conversation starts from it.** A different conversation with the same system block reuses the head, `prefix cache: reusing 3584/3643 tokens from memory`. That is the payoff the 2026-09-23 record could only predict: the head is no longer classed as the first conversation's own, so another conversation is allowed to start from it. **A deeper save keeps it.** Extending the first conversation to 4,850 prompt tokens writes a 4,608-token state, and the save line carries no `removed` clause — where the field session's equivalent save said `removed 1 older file` and took the 24,576-token head. The directory afterwards holds both: `4608 (shared false, continued)` and `3584 (shared prefix, continued)`. **It restores across processes.** With the server stopped and restarted on the same directory, a fresh process reads the shared head from disk alone: `restored 3584 tokens (214.7 MB) in 0.04 s`, then `prefix cache: reusing 3584/3643 tokens from disk`. **Evidence.** The regression gate on the same binaries passes 110 assertions, and the full T0+T1 catalogue passes 73 groups with 31,931 assertions and 0 failed. The live transcript, the plan the server printed, the state listings and the build identity are in [[sources/runs/2026/09/2026-09-24-shared-prefix-live-acceptance]]. **Limits.** Functional acceptance on one 32 GB MacBook Air shared with other sessions, one plan (`--memory-gb 10`, `--max-context 8192`, a 256-token prefill pass) and one prompt shape. The collision was constructed rather than met in the field: the system block and the prompt's last resume boundary were made to fall in the same 256-token cell, and the log confirms both writes at 3,584; the field workload — a 24,576-to-26,624-token agent session — is not re-run here. Elapsed times are incidental observations on a machine in ordinary use and are not a timing claim. The 2026-09-23 record's numbers stand and it is not superseded. The observed behaviour is one machine and one plan; the cross-conversation payoff in a real agent workload remains unobserved. ## Prefix-cache floor at 2048 and 1024 tokens across a restart (community, 2026-09-16) Reported by `@jasen215` in [issue #17](https://github.com/carloslfu/slotstream/issues/17), preserved in [[sources/community/2026/09/2026-09-16-prefix-cache-min-tokens-jasen215]]. Slotstream 0.2.18 (`main` at `ad89ecc`) on a 32 GiB Apple Silicon Mac, `serve --memory-gb 10 --max-context 65536`, one model process at a time. The prompt was a system prompt of about 1,900 tokens and one question. After the first turn the server stopped and a new one started over a copy of the state directory, so any reuse came from disk. The numbers are the server's own statistics. | Phase | Floor 2048 (the default) | Floor 1024 | |---|---|---| | Write after turn 1, 1,919-token prompt | nothing written | 170 MB | | Restart, 1,991-token prompt: tokens reused | **0** | **1,966** | | Restart: prefill | **45.5 s** | **6.0 s** | | Control: restart above both floors, 2,685-token prompt | 2,662 reused, 6.6 s | 2,662 reused, 6.5 s | | Disk after the last phase | 307 MB | 479 MB | The controls agree, so the difference comes from the floor alone. The write the floor avoids is small next to that re-read: on the 48 GB development Mac at 10 GB, later turns wrote about 117 MB each, a new head and their new rows, in 0.05 s, and a 225 MB state restored in 0.04 s ([[records/measurements/persistent-prefix-cache-2026-09-14]]). Pi's opening prompt, about 1,600 tokens ([[records/design/measured-operating-policies]]), is also below 2048, so the servers `slotstream launch` starts kept it in memory but never wrote it to disk. On this evidence the default fell to 1024 tokens on main on 2026-09-25, after 0.2.25. The cost is one head plus the new rows on each turn of a conversation between 1,024 and 2,048 tokens, within the same disk quota. Nothing below 1,024 was measured. **Rewinding after a restart.** A separate probe in the same report branched a three-turn conversation back to turn 1 after a restart and reused 0 tokens at either floor. The disk tier keeps a conversation's latest state and its parent, so the last reply can be regenerated, and removes older ones by design. Since 0.2.21 a prompt's system message and the longest head it shares with a kept state are also saved as shared prefixes when they reach the floor, so such a branch reuses its system prompt, and a later branch from the same point resumes from the head the first one saved. That behavior was not measured again here. One run per configuration on one machine, with single timings, as the report states. ## A continued conversation writes one state per turn: the end of a reply is not a shared prefix **Outcome: a continued conversation now writes one state per turn; the released 0.2.25 also rewrote the turn's own state as a shared prefix.** In [[sources/runs/2026/09/2026-09-26-prefix-turn-writes-e2e]], one `serve --memory-gb 10` process per build over a fresh cache directory ran a system prompt of 1,615 tokens and three turns with 320-token replies. On 0.2.25 the second turn wrote its 1,792-token state (122.8 MB) and then the same state again as a shared prefix (115.7 MB), and the directory ended with two shared prefixes: the system prompt at 1,536 and that turn's own state. On the fix the second turn wrote only its state, and the system prompt stayed the only shared prefix. Prompt and output ids were identical on every turn, and the three turns wrote 526.4 MB against 642.1 MB. **Why it happened.** Under aligned resume, the default since 0.2.22, the memory tier keeps each finished turn as a conversation entry holding its prompt and reply, but never continues it; the next turn resumes from the boundary checkpoint. The shared-prefix rule took the longest prefix the prompt had in common with any held state as a target, so the end of the previous reply became one. Once the reply crossed a pass boundary beyond what the turn reused, its save point was a pass end past that checkpoint: a second state when the new message crossed another boundary, otherwise the turn's own checkpoint, which 0.2.25's colliding-boundary upgrade ([[records/measurements/shared-prefix-live-acceptance-2026-09-24]]) rewrote with the shared flag. A shared prefix is never removed as a redundant ancestor, and a childless one is classed with states nobody continued, so each of these stayed until the quota or the age limit removed it, and a turn's latest state, once flagged, was the first candidate for eviction. In this run the second prompt matched the first prompt and all 320 reply tokens; the held entry ends one token earlier, because the last sampled token is never consumed. **The rule now.** A target counts only where the prompt parts from what another held state read as input. A state it extends outright, or parts from inside the reply that state generated, is its own conversation's earlier turn. The memory tier records where a conversation entry's prompt ended; disk heads under aligned resume hold prompt tokens only. The system prompt boundary is unchanged. **The existing gate did not reach the case.** `Tools/shared_prefix_e2e.py` passed on both builds with 320-token replies: its later turns ran after a restart, or after another conversation had evicted the entry. `Tools/prefix_turn_writes_e2e.py` keeps three turns in one process and exits 2 unless a later prompt reaches the old save point; on 0.2.25 it fails two of its four checks. **Limits.** One 48 GB M5 Pro in ordinary use, one plan (`--memory-gb 10`, 256-token passes), one prompt shape and the Ollama chat endpoint. Timings are incidental and not a claim. A client that sends a reply back changed was not exercised live; the rule for it is covered by `optimization-prefix-client-capacity`. ## Decode speed: GPU keepalive, direct demand reads, a streamed draft head and plain-decode lookahead **Outcome: four decode changes won on the development Mac with unchanged output, and the other ideas tried did not.** A GPU keepalive and direct demand reads together made decode 1.28x faster at a 10 GB target without the draft head and 1.22x at 22 GB with the head and lookahead, over the pairs with no swap activity. Streaming the draft head's routed experts through a 64-expert cache freed 1.2 GB for the main cache, which made the head worth running at 12 GB: 1.23x over plain decode with the lookahead at 28.4 experts per layer, over three swap-free pairs (1.21x over all eight). Letting the decode lookahead run in plain decode added 1.11x at 10 GB. Together, at a 16 GB target, the four decoded 1.79x faster than the shipped default of that commit. The keepalive raised energy per generated token by 7%. The measured configurations were environment-guarded prototypes on an export of commit 37fcb8e; the landed implementations have their own confirmation below. **Method.** One model process at a time on the 48 GB M5 Pro, a live desktop and 4.6 to 6.7 GB of swap in use. Every arm ran every prompt each round in rotated order, after a wait for reclaimable memory above the target plus 6 GB. Four public corpus prompts (r0005, r0206, r0096, r0074, `Tools/expert_lookahead_corpus.py`), greedy, 192 output tokens, two rounds: 8 pairs per comparison. The primary screen keeps completed runs with no global swap-out and no thermal warning; every run had none of the latter. The strict screen also requires no global swap-in, which removes most pairs on this machine; a timing claim needs three strict pairs. Ratios are geometric means of paired decode tok/s. Arms: `kadd` is keepalive plus direct reads, `la` the qualified lookahead configuration (`la_env.json`, the 22 GB plan's), `hs` the head's experts streamed through 64 slots. **Why decode waits.** Streamed decode is stop-and-go. At each layer the host reads the routing back and reads the experts the cache lacks before it submits the next burst. With MLX timing instrumentation, a one-token pass with every expert resident was 337 command buffers with the GPU busy 58% of 55.2 ms, 68.3 µs idle per buffer. An idle Apple GPU clocks down and starts the next buffer late. A one-thread kernel spinning on its own command queue kept it busy: 65% of 47.4 ms, 49.5 µs idle per buffer (single instrumented runs, `gpu_windows.py`). Five paired rounds of the fetch-free pass cost, whose runs recorded swap-outs only, put the one-token pass at 55.8 against 48.5 ms (medians). A Metal microbenchmark re-run on a quiet GPU shows why: a small kernel that runs in 49 µs back to back took 230 µs after any idle gap of 200 µs or more, and the next buffer started 0.12 ms after commit after a 200 µs gap and 0.62 ms after a 5 ms gap; with the spin kernel resident on a second queue, 52 µs and 0.07 ms. Demand misses were read into staging arrays and scattered on the GPU, one more dispatch and synchronization per layer; the direct path reads into host scratch and copies each record into its slot. **Decode, paired ratios.** All pairs, then pairs with no swap activity (n): | comparison | target | all pairs | swap-free | |---|---|---|---| | keepalive + direct reads vs base, no head | 10 GB | 1.298, 1.257 in a rerun | 1.274 (5), 1.279 (6) | | direct reads alone | 10 GB | 1.141 | 1.164 (4) | | keepalive alone | 10 GB | 1.069 | 1.067 (4) | | keepalive + direct, head and lookahead | 22 GB | 1.191 | 1.215 (5) | | keepalive alone, head and lookahead | 22 GB | 1.127 | 1.110 (5) | | direct reads over keepalive | 22 GB | 1.057 | 1.051 (6) | | plain lookahead over keepalive + direct | 10 GB | 1.123 | 1.112 (4) | | plain lookahead over keepalive + direct | 16 GB | 1.071 | 1.054 (4) | | streamed head + la over plain la, all kadd | 12 GB | 1.207 | 1.230 (3) | | streamed head + la over plain la, all kadd | 10 GB | 0.988 | 1.051 (3) | | resident head over plain la, all kadd | 12 GB | 1.010 | 1.018 (2) | | resident head + la over plain la, all kadd | 16 GB | 1.323 | 1.323 (5) | | streamed over resident head, kadd | 22 GB | 1.008 | 0.990 (4) | | streamed over resident head, kadd | 12 GB | 1.136 | 1.134 (2) | | streamed over resident head, kadd | 10 GB | 1.005 | 1.005 (3) | | streamed head + la + kadd vs shipped | 16 GB | 1.785 | 1.794 (4) | | streamed head + kadd vs shipped | 22 GB | 1.166 | 1.178 (3) | Every same-mode comparison kept identical output ids. Streamed and resident heads differed on 2 of 8 at 12 GB because the streamed plan's larger budget chose a longer prefill pass, which moves prompt logits within the known re-chunking envelope; with the pass pinned, the streamed head's ids equal the resident head's. Head against plain comparisons differ at near ties, as speculative and plain decode always have. The draft head's cache hit rate was 0.47 at depth 2, and its reads took 0.24 to 0.34 s per 192-token request. **Caches measured, experts per layer.** 10 GB: 20.0 plain, 17.1 with the lookahead, 17.4 with the streamed head, 13.3 with the resident head (the floor pool). 12 GB: 31.1, 28.2, 28.4, 22.7. 16 GB: 53.7 plain, 47.6 with the streamed head and lookahead, 39.4 with the resident head and lookahead. 22 GB: 73.6 with the resident head and lookahead, 82.7 streamed. Peak memory with the streamed head stayed at or below the resident head's at every target. **Energy.** IOReport SoC energy over the whole request at 16 GB, streamed head and lookahead on both arms: the keepalive raised energy per generated token 7.1% over four swap-free pairs (6.7% over eight), average power 26.7 to 32.5 W, for 1.17x decode (1.19x over eight). Against the shipped default the full configuration used 11.0% less energy per token (12.7% over eight), because it finished sooner. **Did not help.** A duty-cycled keepalive with 200 µs or 1 ms gaps (1.033 and 1.019, wide spread: the clock falls in the gaps). MLX's host spin-wait instead of blocking waits (0.711). Committing every 10 operations, with or without a buffer size limit (0.968, 0.995). A compiled one-row GDN step (0.978). User-interactive QoS for the generating thread (1.007). Layer-local eviction at the floor pool (0.906 plain). Overlapping the shared expert (0.986 at 16 GB, 0.967 at 22 GB). Draft depth 1 (0.979 at 12 GB) and 3 (0.701), and an adaptive depth up to 3 (0.801). The plain lookahead in place of the head at 22 GB (0.744). **Landed code.** The keepalive (`--gpu-keepalive`, `auto` on AC power) and direct demand reads (`SLOTSTREAM_OPT_DIRECT_DEMAND`) were confirmed on the landed build under heavier paging: every run saw swap-ins, so no pair is swap-free and these confirm direction, not a timing claim. Same binary, both switches on against both off, 8 pairs each: 1.228 at 10 GB (8 of 8 above 1, 1.111 to 1.307) and 1.116 at 22 GB (8 of 8 above 1, 1.006 to 1.223), identical ids. Against the installed 0.2.24 release: 1.150 at 10 GB (6 pairs without swap-outs) and 1.092 at 22 GB (8 pairs), identical ids and draft acceptance; the release differs by other commits and its build too. `decode-overlap-check` passed 1,672 assertions on the pre-release build. **Landed streamed head and plain-decode lookahead.** The streamed head, its floor of 28 and the plain lookahead were confirmed on their own landed build ([[sources/runs/2026/09/2026-09-24-draft-stream-landed]]), again with swap-ins in every pair, so these confirm direction, not a timing claim. At 12 GB, automatic mode, the streamed head with the lookahead at 25.5 experts per layer, against `--mtp off`, plain decode with the lookahead at 28.2: 1.205 over eight pairs, eight of eight above 1 (1.093 to 1.329). Plain decode with the lookahead against without it: 1.045 at 10 GB over seven pairs without swap-outs (six above 1, 0.965 to 1.136) and 1.050 at 22 GB with `--mtp off`, 75 against 78 experts per layer at a 65,536-token automatic window, over eight (seven above 1). Because the 10 GB ratio fell short of the prototype's, a later session interleaved the landed build and the prototype, each with and without its lookahead, at 10 GB: 1.082 landed and 1.064 prototype, eight of eight above 1 each, and the two builds within 1.3% of each other with the lookahead and 0.4% without, so the shortfall came from the machine's state that hour. Plain-decode and lookahead pairs kept identical ids; head against plain pairs differ at near ties as always. `draft-stream-check` passed 23 assertions, and `decode-overlap-check` passed again on the combined build. **Limits.** One Mac, one SSD and four prompts of 192 tokens; larger caches than 22 GB were not timed. Most comparisons keep fewer than three swap-free pairs; the claims use only those that keep three or more. The fetch-free pass-cost rounds recorded swap-outs but not swap-ins. The GPU span figures are single instrumented runs. ### v0.2.25 published, installed and accepted **v0.2.25 is public, installed and accepted.** It ships the GPU keepalive and direct demand reads, the draft head streaming its experts with an automatic floor of 28 experts per layer, and the decode lookahead in plain decode. It also carries the shared-prefix boundary fix from [#27](https://github.com/carloslfu/slotstream/pull/27), the decode-forecast download from [#28](https://github.com/carloslfu/slotstream/pull/28) and the documented Mac app changes. Release: [v0.2.25](https://github.com/carloslfu/slotstream/releases/tag/v0.2.25), published 2026-09-24T22:49:34Z from `a0cca848722bf27e7b896a948bff560936a14dbb`. The CI candidate, public archive and installed executable match exactly. Archive SHA-256: `24774d0755b16840f43b51e2a782ff0f68564ec5f9d75a8e4e17258515b3fb13`. Executable SHA-256: `d960143c783dda94bd4d43d304e5f9f7c0e195f7d00feacc2934885dd3d3a951`. | Acceptance | Result | | --- | --- | | Exact-commit hosted CI | Engine, instrumented coverage, external library consumer, Mac app and context contracts passed | | Engine catalogue | 75 groups, 32,004 assertions, no failures or skips, in release and instrumented builds | | Full native battery | 35 top-level gates passed, including both elastic governor drills, `draft-stream-check`, `decode-overlap-check`, quality 15/15, robustness 74/74 and vision serving 25/25 | | Public distribution | Preserved CI archive published, public checksum/provenance verified, public installer upgraded the standard installation from 0.2.24 | | Installed serving | 31/31 with a 10 GB target and MTP on, the draft head streaming its experts; owned server reaped | The first candidate, `5a54b68`, passed CI but failed both elastic drills and was not tagged. The drill predicted the governor without the decode lookahead's reserve, which plain decode now charges; the governor was right. The fix, `a0cca84`, changes only the drill, and both drills then passed on the released bytes. [[sources/runs/2026/09/2026-09-24-release-0-2-25-published-and-installed]] retains the commands, both candidates' native logs, the fix confirmation, source/build identity, workflow output, installer and process-cleanup receipts. The historical backend reference remains visible beside the passing current-backend reference; no tolerance was widened. These are functional acceptance results, not speed claims. The decode gains this release ships, and their limits, are in [[records/measurements/decode-perf-2026-09-24]]. ### v0.2.26 published, installed and accepted **v0.2.26 is public, installed and accepted.** It ships the prefix cache's 1,024-token disk floor from [#30](https://github.com/carloslfu/slotstream/pull/30), one prefix cache write per continued turn with the other review fixes from [#46](https://github.com/carloslfu/slotstream/pull/46), and the contributed fixes from [#31](https://github.com/carloslfu/slotstream/pull/31) to [#43](https://github.com/carloslfu/slotstream/pull/43) listed in the changelog. Release: [v0.2.26](https://github.com/carloslfu/slotstream/releases/tag/v0.2.26), published 2026-09-27T14:12:39Z from `a8a5294a8a826ee9356c900df76614646e756bba`. The CI candidate, public archive and installed executable match exactly. Archive SHA-256: `eb4f52d7655b7d1978c2ef28a19597c687d9e513d8ca0625eaefbc3a72547b3b`. Executable SHA-256: `37243fe333423b238e087448823b93a19712d9473937353fce9e145dfe10e645`. | Acceptance | Result | | --- | --- | | Exact-commit hosted CI | Engine, instrumented coverage, external library consumer and context contracts passed; the Mac app built, and its scripted checks failed only the intermittent scroll-position check that also fails on main | | Engine catalogue | 85 groups, 32,636 assertions, no failures or skips, locally before the candidate existed; CI ran the same catalogue on the candidate | | Full native battery | 35 top-level gates passed on the exact CI bytes, including both elastic governor drills, `draft-stream-check`, `decode-overlap-check`, quality 15/15, API robustness 74/74 and vision serving 25/25. Vision parity's two gates ran separately on the same bytes after the battery skipped them for a missing reference environment | | Prefix cache live gates | One write per continued turn 4/4 (turn 2 wrote 122.8 MB, not 0.2.25's 238.4 MB); disk tier restart 12/12 after its checks were corrected for the system prompt's shared prefix | | Public distribution | Preserved CI archive published, public checksum/provenance verified, public installer upgraded the standard installation from 0.2.25 | | Installed serving | 31/31 with a 10 GB target and MTP on; owned server reaped | The disk tier restart gate's five file-count checks failed identically on 0.2.25 and on 0.2.26, with byte-identical writes: they predated the shared prefix turn 1 keeps for its system prompt, while its restart, regeneration and output-id checks passed. Commit `6dcba83` corrects them. [[sources/runs/2026/09/2026-09-27-release-0-2-26-published-and-installed]] retains the commands, the candidate's native logs, the vision parity run, both live gates, the local pre-candidate checks, source/build identity, workflow output, installer and process-cleanup receipts. These are functional acceptance results, not speed claims. The per-turn write measurement is in [[records/measurements/prefix-turn-writes-2026-09-26]]. ### v0.2.27 published, installed and accepted **v0.2.27 is public, installed and accepted.** It ships the contributed fixes from [#48](https://github.com/carloslfu/slotstream/pull/48) (raw downloads that stop early when a file fails, `launch` reply deadlines, stricter `/v1/responses` and AI SDK gateway requests, prefix cache file checks and disk accounting, and checkpoints with empty tensors), [#52](https://github.com/carloslfu/slotstream/pull/52) (unambiguous HTTP body framing) and [#53](https://github.com/carloslfu/slotstream/pull/53) (sparse pin bookkeeping), and the corrected checks from the September 22 decode study in [#49](https://github.com/carloslfu/slotstream/pull/49), all listed in the changelog. Release: [v0.2.27](https://github.com/carloslfu/slotstream/releases/tag/v0.2.27), published 2026-09-30T19:16:02Z from `22ac07e7fe04bd45d844b39ebd8a86c1d3c6d4cf`. The CI candidate, public archive and installed executable match exactly. Archive SHA-256: `b5af97f06a7fcad105e6027aa00e6ec6d58e5c8d13234c5f95852c66f242494c`. Executable SHA-256: `fad3a30a1bcabec74d64690ddc538c8553026168d71094a97d31346bedf2eae4`. | Acceptance | Result | | --- | --- | | Exact-commit hosted CI | Engine, instrumented coverage, external library consumer and context contracts passed; the Mac app built and passed its scripted checks, including the scroll-position check fixed in [#50](https://github.com/carloslfu/slotstream/pull/50) | | Engine catalogue | 89 groups, 33,136 assertions, no failures or skips, locally before the candidate existed; CI ran the same catalogue on the candidate | | Full native battery | 35 top-level gates passed on the exact CI bytes with no failures or skips, including vision parity, quality 15/15, API robustness 74/74 and vision serving 25/25 | | Prefix cache live gates | One write per continued turn 4/4 (turn 2 wrote 122.8 MB, as in 0.2.26); disk tier restart 12/12 | | Public distribution | Preserved CI archive published, public checksum/provenance verified, public installer upgraded the standard installation from 0.2.26 | | Installed serving | 31/31 with a 10 GB target and MTP on; owned server reaped | A local build of the release commit also passed the same 35 gates before the candidate existed. [[sources/runs/2026/09/2026-09-30-release-0-2-27-published-and-installed]] retains the commands, the candidate's native logs, both live gates, the local pre-candidate checks, source/build identity, workflow output, installer and process-cleanup receipts. These are functional acceptance results, not speed claims. ## Initial quantization screen and bounded native decoding ### Corrected-prefetch lower-budget result, October 6 [[sources/runs/2026/10/2026-10-06-candidate-correction-lowbudget-result]] closes one clean pair at a ten-GB process ceiling on the owned Mac. Preparation is in [[sources/runs/2026/10/2026-10-06-candidate-correction-lowbudget-preparation]]. Both arms retain the same native three-bit weights, full context allowance and two streamed drafts; the corrected arm charges its additional forecast bytes and uses fewer cache slots. | Workload | Uncorrected committed tok/s | Corrected committed tok/s | Ratio | | --- | ---: | ---: | ---: | | Short | 8.353488 | 8.603241 | 1.029898 | | Context | 9.841657 | 10.336926 | 1.050324 | | Completed coding answer | 9.400903 | 10.251414 | 1.090471 | All committed token IDs and text match. The coding request-wall ratio is 0.934698. Peak physical footprints are 7,274,517,256 and 7,269,077,840 bytes, below the ten-GB ceiling; plans use 731 and 715 slots. All timing exclusions are empty. These are one pair of observations with uncontrolled OS file cache, not medians, a sustained guarantee, a contemporaneous original-pack comparison or proof for another budget/Mac. The frozen useful-gain criterion requires at least 1.05 on both fixed workloads and no material coding latency loss. Short fails, so the transfer recipe does not earn confirmation or promotion. Prefetch counters show fewer wasted requested bytes and demand misses, while forecast evaluation takes slightly longer. They explain the direction of the small observed gain but overlap and do not provide an additive wall-time decomposition. This does not provide a route from these rates to twenty. The frozen collector retains an older resource-identity mapping and stops after both model runs complete. The existing current validator, committed before the experiment, differs by exactly the branch recognizing the already implemented corrected-native profile. A separately pinned mechanical replay validates the immutable receipts against that exact identity with every threshold unchanged. The failed coordinator is retained; no model cell is replaced. The fourteen-GB pair remains unmeasured. Stop this practical transfer screen; further work needs a materially stronger evidence-backed lead, not another budget sweep. ### Corrected-prefetch continuation remains unmeasured, October 6 [[sources/runs/2026/10/2026-10-06-focused-continuation-host-refusal]] preserves the single bounded host admission wait. Neither arm launched because real reclaimable memory stayed below the prospective requirement. No timing result, candidate rejection or qualification follows. The earlier exact-output and physical-bound functional pass in [[sources/runs/2026/10/2026-10-06-candidate-correction-functional-only]] remains valid. A quiet, safely admitted pair is still needed to answer the speed question. ### Matched runtime control, October 6 [[sources/runs/2026/10/2026-10-06-matched-runtime-control-preparation]] freezes the diagnostic; [[sources/runs/2026/10/2026-10-06-matched-runtime-control-results]] preserves the completed clean pass and earlier preflight-refused attempt. One pass on the owned 48 GB Mac at a 14 GB ceiling gives: | Configuration | Short committed tok/s | Context committed tok/s | Coding committed tok/s | | --- | ---: | ---: | ---: | | Original default settings | 12.0401 | 14.8364 | 14.2470 | | Original with candidate verification, prefetch and allocation policy | 11.1519 | 13.3109 | 13.3980 | | Native three-bit with that same policy | 12.6323 | 15.5806 | 14.2376 | These are single screening observations, not new medians or hardware-tier promises. The matched original is 6.0–10.3% slower; three-bit beats that control by 6.3–17.1%. Against ordinary original settings, three-bit gains about 5% on both fixed workloads and ties the coding answer. All three produce the identical correct coding text. Peak physical footprints are 11.274, 10.879 and 11.139 GB. Plans keep 32K context and 1024-row maximum prefill, shortened to 512 for short prompts. Original/matched/candidate cache counts are 1635/1509/1941. The frozen runtime-overhead lead passes. Verification takes 19.837/16.114/5.344 seconds in the original versus 21.501/18.003/5.706 in the matched control. That phase includes compute, expert loading and synchronization; its increase is not proof that row-invariant projection alone causes the loss. The matched uncorrected forecast also wastes more speculative-read bytes than the original correction. Both matched arms preserve their own loaders and quantized routing/output, so the isolated comparison is the runtime recipe versus complete quantization/deployment benefit, not bit width alone. One tightly scoped next test will transfer the pinned forecast correction to the native candidate, charge its memory and require exact output plus a material complete-response gain. All product defaults and prior failed quality decisions remain unchanged. No observed configuration here reaches twenty tokens/s. ### Distinct mixed Q3_K component cost rejection, October 6 [[sources/runs/2026/10/2026-10-06-mixed-q3k-component-cost-rejection]] captures the prospective format audit, bounded source ranges, scripts, exact receipts and clean repeat. The independent GSQ/RCO release is pinned to revision df4f5bbd0a93e5f6a37a377d5d0cf67d89ff0f6e, whose API last-modified date is September 23. Its header contains 72 Q3_K, ten Q2_K and 62 Q2_0 expert tensors, with 43,332,403,200 packed expert bytes. That is 9.375% below the GSQ224 native expert payload, not a measured native-memory saving. Hierarchical scale layouts require different execution. Only a 12-MiB header prefix and 18,688,000 bytes of layer-zero expert samples are acquired; the complete shard hash and full useful quality remain unverified. The cheap component tests actual layer-zero experts zero through nine, using the dominant Q3_K gate/up and Q2_0 down recipe. Direct Q3_K Metal reductions compare against the independent pinned gguf decoder and float32 matmul. All four numerical cases meet the 0.02 maximum-relative bound; the largest observed error is 0.000874126. Batched versus individual custom reductions are exact. The Q2_0 down path reuses the existing affine representation with explicit BF16 rounding. These results do not transfer the publisher's quality from another runtime/embedding precision. Five alternating rounds of forty dependent synchronous iterations use separate rotating banks of at least 256 MiB per recipe. Frozen success requires both row counts to cost at most 0.90 times original and 1.05 times GSQ224. The first run's timing is excluded for CPU contention in [[sources/runs/2026/10/2026-10-06-mixed-q3k-component-timing-excluded]]. One separately frozen unchanged repeat passes timing eligibility and fails cost at both row counts: | Token rows per expert | Original median ms | GSQ224 median ms | Mixed Q3_K median ms | Mixed/original | Mixed/GSQ224 | | --- | --- | --- | --- | --- | --- | | 1 | 0.342198 | 0.285511 | 0.502665 | 1.468929 | 1.760576 | | 3 | 0.418998 | 0.417626 | 0.990269 | 2.363422 | 2.371185 | The clean repeat takes 133.682419 seconds including the retained 120-second quiet preflight. Physical lifetime peak is 1,715,259,096 bytes under the four-GB limit; nominal thermal state, real headroom and unchanged swap counters are retained. There is no full-model throughput claim. Close this prototype without Q2_K implementation, full loader/download, quality campaign or promotion. The negative component result is not universal format infeasibility. The same record preserves all four successful CI workflows for unchanged runtime source 0d85d9d50567e936d69e10b2abf8290ade92be1f. A current upstream check finds an idle GPU-residency fix and an M1-specific dispatch change, neither demonstrated to improve this sustained M5 workload; the prior fused-QMM source is unchanged. No product code, pack registry or Auto policy changes in this continuation. Native Python MLX hashes were captured after the runs and retain that limitation. ### GSQ224 image instruction loss and one-row batch rejection, October 6 [[sources/runs/2026/10/2026-10-06-gsq224-image-and-qmv-rejections]] completes both frozen screens. The image comparison reuses eight existing cases and twenty-four turns per arm, with no new prompts or relaxed graders. Original passes 8/8 cases and 24/24 turns; GSQ224 passes 4/8 and 20/24. The four new failures are bare values where an answer-key JSON object was required. All values are semantically correct: three bar counts of 3 and the color green. This is a structured instruction-following regression, not visual misrecognition. Both complete under 14.5 GB physical: original peaks at 12,246,179,544 bytes and candidate at 9,104,103,904 bytes. No clean timing claim follows. Two preserved preparation failures launch no model: a missing helper API and historical double-counting of verified hardlinks. The successful preparation uses the existing current physical allocation method and unchanged 430-GB staging cap. Reject the exact GSQ224 full-product recommendation under its frozen quality rule. Retain its text-speed and functional results within scope; cancel the unlaunched higher-budget pilot and remaining confirmation cell. Five completed speed cells are preserved without a final median. The independent one-row batch component preserves finite exact values in all eighteen shape/row cells and passes timing eligibility. However, lm_head costs 1.225 times the row-wise implementation at three rows and 1.301 times at five. Smaller-shape gains do not satisfy the frozen primary-shape/no-regression gate. The 21.302-second run earns no native implementation or speed claim. Ordinary batching's prior invariant failure remains separate and unchanged. These are specific closed routes, not universal impossibility evidence. ### GSQ224 owned vision and full-context functionality, October 6 [[sources/runs/2026/10/2026-10-06-gsq224-vision-and-context]] records the exact source and complete raw outputs. The existing catalogue passes all 104 groups. Explicit GSQ224 vision passes all 49 ownership, corruption, cancellation, bounded allocation, component equality, drafted/plain output, prefix and HTTP assertions. Its physical peak is 6,785,160,232 bytes. The same original vision tower and preprocessing bytes are retained. The existing full-context diagnostic consumes 32,768 tokens and passes all 1,445 assertions, with a 7,671,123,808-byte physical peak. Every recorded prefix, state restoration, continuation and over-cap refusal passes. This direct-model fixture uses 640 slots and a 128-MB allocator cache, with a ten-GB watchdog and real headroom. It is not a complete Engine speed measurement, an image-answer accuracy score or evidence that the fixed full-context allowance can simply be cut. GSQ image/lookahead combination and public Auto remain unadmitted. [[sources/runs/2026/10/2026-10-06-gsq224-confirmation-host-stops]] preserves five clean representative comparison cells and two pre-allocation quiet-start refusals. The remaining candidate cell has no result and the six-cell comparison is incomplete. No partial median is promoted. All four CI workflows for the measured predecessor a53ca0c39ee9288d67de273f5303dfbbb5916903 pass; final source acceptance remains separate. ### GSQ224 explicit four-draft pilot, October 6 [[sources/runs/2026/10/2026-10-06-gsq224-four-draft-cost]] changes only explicit depth to four, reusing the exact grouped depth-four state/recovery proof, full attention lifecycle, binary, weights, reserve and honest ten-GB planner. The same 81-token coding answer remains exact and passes all unchanged functional tests. Committed generation reaches 16.989207 tokens/s, 11.336% above the two-draft attention pilot. First text is 2.308700 seconds and total request 7.017247 seconds. Verification passes fall from 28 to 18; 63 of 72 proposals are accepted. Lifetime physical peak is 6,150,443,664 bytes, and timing exclusions are empty. The quiet-host preflight restarted once for observed unrelated CPU activity before a clean measured interval. It is not a failed timing retry. The frozen useful-gain and latency criteria pass. This is a single completed coding pilot and remains below twenty tokens/s. It earns the small interleaved representative comparison in the active plan, not production promotion or a global depth-four default. No new raw numerical corpus, model weights or build is added. ### GSQ224 attention lifecycle and coding improvement, October 6 [[sources/runs/2026/10/2026-10-06-gsq224-attention-cost]] preserves the initial empty-synthetic-fixture failure and the prospectively corrected continuation. The correction reuses the existing framed lifecycle fixture at identical lengths and its generated continuation token. All original equality, complete-work, memory and timing conditions remain. No token golden is regenerated. All 104 full Engine assertions pass, including the independent GSQ token golden, demand-versus-attention equality at 44/260/2054 tokens, persistent prefix restore, cancellation, live shrink/recovery under the saved ceiling, HTTP and exact byte ownership. The lifecycle peak is 7,213,700,592 bytes within its ten-GB physical watchdog. Its small fixed arenas and conservative planning allowance remain distinct from actual deployment planning. The subsequent honest ten-GB cost configuration reaches 15.259394 committed tokens/s, against the reused boundary pilot's 14.237665, a 7.176% gain. Its unchanged 81-token answer passes the same three pure-function/input-preservation tests. First text is 2.266354 seconds; complete request is 7.508721 seconds. Lifetime physical peak is 5,922,214,400 bytes. Timing exclusions are empty. It accepts 53 of 56 draft tokens in 28 verification passes. Forecast evaluation falls from 1.924225 to 1.262845 seconds and decode I/O from 2.068297 to 1.595937 seconds; overlapping counters do not provide additive causal attribution. This passes the frozen pilot gate and retains attention as the leading explicit GSQ research recipe. It is not paired confirmation, quality equivalence, another-Mac validation, production Auto admission or achievement of twenty tokens/s. Exact grouped depth-four state/recovery is already covered by the earlier speculation proof, so a separate prospective one-response depth study can test the high-acceptance workload without repeating numerical export or changing the adopted default. ### Original corrected-lookahead configuration control, October 6 [[sources/runs/2026/10/2026-10-06-original-corrected-lookahead-control]] records a prospective original configuration with two streamed drafts and explicit corrected lookahead. Two preparation failures precede the completed run: one protocol test protected historical admission, then a requested uncorrected mode conflicted with the installed correction identity. Both stopped before model allocation and are not memory-feasibility results. The corrected plan prices 428,867,584 lookahead bytes, keeps 694 slots and fits the ten-GB ceiling. It completes the unchanged coding task at 10.004366 committed tokens/s, first text 2.658422 seconds and whole request 10.654401 seconds. All three pure-function/input-preservation checks pass, with 81 tokens and no timing exclusions. It accepts 53 of 56 drafts, the same counts as the GSQ224 boundary pilot. GSQ224 remains 1.423145 times as fast in these single pilots. This configuration control supports further mixed-recipe optimization, not a causal quantization-only claim or paired qualification. The original automatic defaults remain unchanged. The explicit enabled original mode authenticates the installed correction, retains its complete rounded charge and is limited to V2 pilots. Affected protocol and applied-feature tests pass; unknown corrections, off/disabled execution under an explicit-on request, historical protocol widening and held-out widening remain refused. ### GSQ224 complete Engine cost and original drafting control, October 6 [[sources/runs/2026/10/2026-10-06-gsq224-engine-cost-pilot]] records the exact grouped Engine memory contract, preserved intermediate refusals, successful grouped component/admission checks and streamed speculative state/recovery. [[sources/runs/2026/10/2026-10-06-original-streamed-draft-control]] adds one cheap original control rather than attributing every configuration difference to quantization. These are single complete coding pilots on the owned 48 GB Mac at the same ten-GB process ceiling and same prompt. Every answer passes the unchanged three pure-function and input-preservation tests. GSQ224's answer also exactly matches the original Auto answer's 81 tokens. All timing eligibility checks pass. The older original Auto observation is reused; these are not interleaved paired confirmations. | Configuration | Committed generation tokens/s | First text seconds | Complete request seconds | | --- | --- | --- | --- | | Original Auto, no draft, automatic lookahead | 9.035368 | 2.507262 | 11.361321 | | Original forced two streamed drafts, automatic lookahead off | 9.519358 | 2.950460 | 11.353834 | | GSQ224, two streamed drafts, explicit boundary lookahead | 14.237665 | 2.232687 | 7.851238 | GSQ224 remains about 49.6% faster than the explicit original draft configuration, but lookahead settings still differ. It uses 640 slots versus original forced-draft 834. It accepts 53 of 56 draft tokens and takes 28 target passes instead of original Auto's 81. Target expert read counters remain about 32.94 versus 32.91 GB because cache/routing behavior differs; smaller records alone do not explain the result. Forecast evaluation records 1.924225 seconds within 5.618899 decode seconds, a concrete further optimization candidate. Those counters overlap and cannot be added into a wall-time decomposition. The mixed Engine lifetime physical peak is 5,924,508,184 bytes, within its full ten-GB planned ceiling. The preceding grouped component and streamed speculative checks pass without numerical tolerance changes. All 104 T0/T1 catalogue checks and 25 affected benchmark tests pass on the final source before full static/CI acceptance. This earns cheap explicit-lookahead controls and limited representative confirmation. It remains below twenty tokens/s, outside production Auto and unmeasured on other Macs. No new raw logits, weights or paid compute were used for these cost checks. ### GSQ224 native parity and completed tasks, October 6 [[sources/runs/2026/10/2026-10-06-gsq224-native-quality]] preserves the mechanical preparation refusal, native exact-artifact admission failure and corrected continuation. Reusing the completed traversal and full-layer reference, the candidate passes native layer comparisons, sixteen self-fed target steps across cache histories, and speculative verification/state/recovery. Tolerances, prompts and graders are unchanged. Original and GSQ224 both pass fifteen of sixteen complete tasks, with the identical pass/fail set. All four coding, four tool, multilingual and retrieval cases pass. Both fail sort-records by returning records instead of names. The frozen advance condition passes with no newly lost original success. This is a small useful-quality screen, not a general equivalence claim. The candidate task path is the existing fixed 640-slot reference probe, not a full Engine performance configuration. The complete corrected pipeline takes 588.534375 seconds; speculation peaks at 8,169,084,592 physical bytes. Functional durations do not establish serving throughput. The retained new numerical fixtures add 28,682,240 bytes within the declared 2.13-GB total raw allowance. Full Engine accounting, grouped execution parity, complete-response speed and product qualification remain distinct work. The original remains preferred until an eligible useful improvement is measured. ### GSQ224 authenticated source and positive proxy, October 6 [[sources/runs/2026/10/2026-10-06-gsq224-source-and-quality]] completes the previously pending source admission and bounded quality experiment. The exact first GGUF shard passes its 37,623,740,192-byte full SHA256 after the preserved first-transfer failure. The original source, complete GSQ file, header prefix and reference fixtures are independently reauthenticated for the quality run. The three-bit coding control exactly reproduces its converted tensor bytes and previous full-vocabulary hash. The following fixed-context scores compare sixteen selected positions per context against the existing VQ4.4 full-vocabulary reference, using original tokenizer IDs. Top-choice agreement is a distribution diagnostic, not completed-task accuracy or BF16 ground truth. | Context | Original four-bit agreement / KL | Prior333 agreement / KL | GSQ224 agreement / KL | | --- | --- | --- | --- | | Coding | 0.7500 / 0.366843 | 0.8750 / 0.475493 | 0.7500 / 0.416326 | | Tool result | 0.7500 / 0.764051 | 0.6250 / 0.918598 | 0.8125 / 0.442250 | | Multilingual | 0.8750 / 0.343711 | 0.6875 / 0.379610 | 0.8750 / 0.385272 | | Macro | 0.791667 / 0.491535 | 0.729167 / 0.591234 | 0.812500 / 0.414616 | The candidate passes the unchanged per-case and macro proxy criteria. All four forwards complete in 159.032416 seconds with a 2,186,544,568-byte lifetime physical peak, under the ten-GB bound. Functional duration is not serving throughput. No new raw logits or model payload are saved by this screen. It earns completed-task evaluation; it does not establish similar useful quality across tasks or the twenty-token target. [[sources/runs/2026/10/2026-10-06-gsq224-changed-projection-export]] then exports only the changed gate/up projections. Forty-eight files total 25,165,864,240 bytes, with exact manifest `dda8570d568459acb44dfbcf356dea437f12630fa461d9b0eef8c3185b523011`. Original down and nonexpert values remain in their authenticated parent. Every tensor is read back and every output file hashed before publication. The producer finishes in 55.238872 seconds with a 242,844,440-byte physical peak; these are conversion resource results, not inference results. The runtime record still includes original down and costs 1,945,600 bytes. Native parity, completed-task quality and useful complete-configuration speed remain required before product admission. ### GSQ224 source admission and frozen quality follow-up, October 6 [[sources/runs/2026/10/2026-10-06-gsq224-admission-preparation]] preserves the closed retirement scope, exact reconstruction sources, seven passing retirement checks, complete deletion receipt, failed first transfer, successful range connectivity recheck, repaired acquisition protocol, base-lineage recheck and unexecuted quality protocols. All forty-eight minmax223 files were authenticated before any deletion; 42,781,961,312 file bytes were retired with a 40,157,184-byte observed physical peak. Allocated staging immediately afterward was 366,837,399,552 bytes. The supported original model and the prior negative quality evidence remain unchanged. The first full-shard Hub attempt failed after a read timeout and DNS error, before authentication. It is an acquisition failure, not a GSQ numerical result. Bounded range acquisition is in progress at this checkpoint; no complete-source authentication or full-model quality is claimed. The next single screen uses the existing fixed contexts, control and proxy thresholds, with the same ten-GB physical cap and no exported weights or new raw logits. Base repository history still has one published non-README payload map, supporting a declared-lineage inference rather than an independently reproduced calibration/conversion. The active plan names the exact advance and stop conditions. ### Calibrated GSQ scalar compatibility and mixed component lead, October 6 [[sources/runs/2026/10/2026-10-06-gsq-native-format-and-mixed-screen]] preserves thirty raw metadata, protocol, code, refusal and numerical artifacts. Immutable GSQ model revision `ed59f92082b1e93c0e96d60a8b11aab089b52f09` declares Qwen Flash Next as its base. The actual bounded GGUF header has 144 routed projections, all Q2_0 with group size 64 and 18 bytes per block. Their source payload totals 33,973,862,400 bytes. The complete first shard is declared as 37,623,740,192 bytes with SHA-256 `69820c02ec7d0b45ef2ebb19d6620299db749fe2aded7f39f93c6b88b199b720`; that whole-file digest had **not** been verified at the format-screen checkpoint; the later full authentication is recorded above. Only immutable-revision HTTP ranges, extents and local range digests are checked. This is not installation authentication or proof of the exact original BF16 producer revision. Three declared layers (0, 23, 47), their first ten experts and all three projections supply 41,472,000 sample bytes. Independent scalar decoding and native FP16 affine reconstruction agree numerically for every sampled weight. The scalar grid is `(code - 1) * scale`; native repacking adds a bias equal to minus the scale. Two-bit weights with FP16 scales and BF16 inputs promote the native gathered operation to FP32, as confirmed by the returned dtype. Explicit FP16 execution avoids promotion, but its three-row combined gate/up/down cost is 1.2350/1.2093/1.1866 times original four-bit. One-row ratios are near one. It fails the frozen all-cell 1.05 cost ceiling. The process peaks at 802,440,272 bytes. Two early consecutive metadata-service CPU observations and one later isolated observation remain visible; none meets the frozen three-consecutive-sample exclusion. Thermal/power remain nominal, with no whole-host paging. The pre-allocation first invocation refuses stale helper hashes before output creation, MLX import or numerical execution. A separately frozen mechanical correction binds the already accepted current helpers, retaining identical samples, hypothesis and limits. The refusal is not a timing result or a scored retry. Existing minmax component data isolates the two-bit down projection as the three-row regression. A separately frozen follow-up therefore retains the original four-bit down matrix and uses calibrated GSQ two-bit gate/up with explicitly BF16-rounded scales and biases. No custom kernel is introduced. Five alternating rounds of forty iterations per cell give the following combined expert-operation costs; activations are deterministic synthetic probes, not a quality corpus. | Layer | One-row original / mixed, ms | Mixed/original | Three-row original / mixed, ms | Mixed/original | | --- | ---: | ---: | ---: | ---: | | 0 | 0.267236 / 0.260401 | 0.974422 | 0.379119 / 0.386411 | 1.019236 | | 23 | 0.262146 / 0.257251 | 0.981328 | 0.374168 / 0.382136 | 1.021297 | | 47 | 0.261270 / 0.260629 | 0.997548 | 0.371209 / 0.374944 | 1.010060 | Every mixed output is finite BF16. Maximum retained scale-rounding relative weight MSE is 0.00000294297024, below the frozen 0.00001 bound. About seven eighths of FP16 scales change under BF16 rounding, so this mixture is explicitly a new arithmetic/quantization recipe, not lossless FP16 scalar conversion. Its expert record is 1,945,600 bytes against original 2,764,800, a ratio of 0.703704. The run peaks at 801,637,336 physical bytes, with nominal thermal/power, no paging and no competing-CPU observations. All cells pass the frozen at-most-five-percent component-cost regression gate. This is a positive storage/kernel lead for one calibrated mixture only. It does not establish whole-model throughput, acceptable task quality or a lower-memory Mac recommendation. It cannot inherit the published full GSQ pack's quality: down matrices and non-expert components come from the original, while the retained GSQ scales are rounded. Source authentication/provenance and the existing bounded reference/task path are next. No weights are activated, no production code is changed, and the original remains preferred. ### Current two-draft VQ and larger-cache rejection, October 6 [[sources/runs/2026/10/2026-10-06-vq32-current-two-draft-cost]] preserves the current two-draft VQ3.2/original-dense cost screen. [[sources/runs/2026/10/2026-10-06-vq32-two-draft-cache-screen]] preserves its prospectively frozen cache follow-up, complete source patch/build identity, catalogue, physical supervision and restoration. Both use the exact existing 125-token coding prompt, 81 output tokens, 32,768-token configured context, original draft head and ten-GB physical ceiling. Both accept 52 of 58 drafts and produce the same complete answer as the preserved original coding case. No new payload or format is created. | Current VQ configuration | Conservative committed tokens/s | Full request, s | First emission, s | Physical peak, bytes | Cache hits | Cache loads | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | 608 records | 8.535129 | 21.125272 | 11.752244 | 7,484,364,464 | 0 | 42,128 | | 1,824 records | 8.734737 | 20.975328 | 11.816495 | 9,861,766,488 | 11,636 | 30,492 | The conservative rate counts output tokens minus the first, divided by complete request time after the first emission; it includes the short return/teardown tail after the last emission. Both runs start fresh and pass nominal thermal/power, no sustained competing CPU, zero global paging, exact output and physical-bound checks. They are one observation each with uncontrolled OS file caching, not repeated paired qualification or another Mac's performance. The reused original ten-GB coding reference is 9.035368 tokens/s and 11.361321 seconds for its full request, but it follows earlier requests in its process and uses its native complete decode timer. Preserve that instrument/history difference; the VQ request-time ratio does not establish a causal prefill slowdown. The first screen fails its original-referenced five-percent generation and ten-percent request gates. The larger bank is a distinct cache-working-set hypothesis from the older non-speculative trial. It reserves 3,583,180,800 bytes instead of 1,194,393,600, restores reuse, exercises slots 1,535 and 287 in the two classes, and returns with zero pins. All 104 native catalogue groups pass. Its generation ratio against the fresh-process small bank is only 1.023387 and request ratio 0.992902; its generation ratio against original is 0.966727. It misses the prospective twenty-percent mechanism gain and original-referenced usefulness gates. Reject this exact expansion and restore the prior accepted source/executable. No promotion, repeated confirmation or broader quality work follows. The code's short-prompt path yields no parallel-prefill batches on this prompt; the existing segmented path starts above 409 rows. Cache statistics and this dispatch fact are observations, not a causal wall-time breakdown. They do not support another implementation without a materially new measured mechanism. The earlier sixteen-task VQ3.2/original-dense quality result stays intact, but neither current cache configuration earns the complete speed or product recommendation gates. ### Fixed-row native complete-response screen, October 6 [[sources/runs/2026/10/2026-10-06-fixed-row-native-and-performance-screen]] closes the component lead. All 104 catalogue groups and all 86 existing native Engine lifecycle assertions pass, including exact plain/speculative outputs, cancellation, memory/prefix recovery and bounded resizing. Nine added assertions check the fixed-row implementation's one-through-eight-row equality, padding cropping, wide-prefill fallback and disabled-mode fallback. The lifecycle process peaks at 7,971,526,944 bytes inside its ten-GB watchdog. The separately versioned recipe retains the old exact attention, plain matmul and draft fusion. The subsequent single fourteen-GB pilot passes every frozen physical/timing condition and completes all three unchanged workloads. It compares against the completed three-round original medians; no original run is repeated. The natural coding text is identical. This is a directional engineering screen with uncontrolled filesystem caching on the owned 48-GB M5 Pro, not a statistical comparison or another-Mac estimate. | Work | Original median tok/s | Fixed-row pilot tok/s | Pilot/original generation | Pilot/original request time | | --- | ---: | ---: | ---: | ---: | | Short | 12.752474 | 12.885401 | 1.010424 | 0.996379 | | Context | 15.375956 | 15.496890 | 1.007865 | 0.944842 | | Coding | 14.968441 | 14.562017 | 0.972848 | 1.047671 | The pilot misses both the all-workload five-percent advancement criterion and the twenty-token alternative. Its successful component arithmetic and correctness checks do not overcome this complete-response result. Reject the exact recipe, with no paired repetition, larger quality campaign, Auto admission or supported-default change. The source is reverted and the prior fully accepted binary/source identity is restored; the complete experimental patch and receipts remain reproducible. The original remains preferred, and the full low-budget speed goal stays unmet. ### Native dispatch rejection and fixed-row follow-up, October 6 [[sources/runs/2026/10/2026-10-06-native-batched-invariance-failure]] records the full release build and all 104 catalogue groups passing, followed by a real failure of the unchanged native plain/speculative target-ID assertion. The first mismatch is the eighth generated token. This session stops the process and preserves its incomplete later checks; physical peak is 7,942,314,296 bytes. No speed pilot follows, no tolerance changes, and ordinary batched arithmetic is reverted. [[sources/runs/2026/10/2026-10-06-fixed-row-projection-screen]] is an incomplete follow-up with two isolated competing-CPU observations and an omitted direct candidate self-invariance comparison. Its timing is excluded. The separately frozen single retry in [[sources/runs/2026/10/2026-10-06-fixed-row-projection-clean-screen]] retains the exact shapes, rounds and advancement criteria, adds the omitted observation and avoids concurrent database/output work. It passes every timing/physical condition with zero paging or competing CPU and nominal temperature. All six candidate three-row outputs exactly match concatenated separately padded single-row projections. | Shape | Fixed/rowwise cost at three rows | Fixed/rowwise cost for two singles plus one triple | | --- | ---: | ---: | | QKV | 0.771 | 0.922 | | Output projection | 0.790 | 0.937 | | Hyper-connection up | 0.911 | 0.983 | | Hyper-connection down | 0.877 | 1.020 | | Shared-expert gate | 0.989 | 1.011 | | Vocabulary head | 0.467 | 0.732 | The three primary shapes all pass the predeclared gain criteria; none of the six blend costs regresses more than ten percent. Physical peak is 1,905,526,800 bytes. The synchronous component timings include common host/gather/tanh overhead and a rotating weight bank. The projection blend is not a measured draft/target wall-time model. This is a lead for a narrow native implementation retaining exact attention, plain matmul and draft fusion, with fixed groups of three quantized projection rows and padding cropped before return. Native lifecycle, full-response speed and focused quality are still required; no product/default promotion follows. ### Three-row resident projection screens, October 6 continuation [[sources/runs/2026/10/2026-10-05-fused-qmm-component-screen]] captures the pinned upstream fused four-bit QMM implementation at commit `187370a64fa11264880024b1dc227a7a767b5ba3` and its negative local screen. Six authenticated original dense shapes at one, three and four rows use rotating independent weight banks, five alternating rounds and forty dependent iterations per round. At the deployed three-row verification size, fused/ordinary cost ratios range from 0.9813 to 1.0247; none reaches the frozen ten-percent gain. All timing/physical gates pass. The process peaks at 1,938,769,936 bytes. No custom kernel is integrated. [[sources/runs/2026/10/2026-10-05-native-rowwise-projection-screen]] separately compares ordinary batched MLX with the actual reference-mode construction of three one-row MLX operations concatenated before evaluation. All six shapes improve in this component screen; the three predeclared large shapes exceed the ten-percent lead criterion. Five alternating rounds of forty dependent iterations use authenticated resident weights, rotating banks and equal host/gather/tanh work. All timing gates pass with zero paging/contention and nominal temperature; physical peak is 1,900,480,480 bytes. | Resident projection, three rows | Batched ms | Rowwise ms | Batched/rowwise cost | | --- | ---: | ---: | ---: | | Layer-zero QKV | 0.242285 | 0.321564 | 0.7535 | | Layer-zero output | 0.221211 | 0.275754 | 0.8022 | | Hyper-connection up | 0.186599 | 0.237910 | 0.7843 | | Hyper-connection down | 0.209330 | 0.238134 | 0.8790 | | Shared-expert gate | 0.174259 | 0.194274 | 0.8970 | | Vocabulary head | 1.481147 | 3.227148 | 0.4590 | These are synchronous dependent-chain component costs, including common non-QMM work, not isolated GPU timestamps or full-model throughput. Relative RMS output differences of about 0.0123–0.0227 make the proposed native recipe arithmetically distinct. The positive screen authorizes only the separately versioned native recipe's unchanged lifecycle/equality checks followed, if successful, by one full-response pilot. Reference arithmetic remains untouched. Neither source qualifies a pack or changes the supported original default. ### Actual Desktop ceiling, October 5 [[sources/runs/2026/10/2026-10-05-practical-desktop-ceiling]] captures three complete paired rounds at the original Desktop's existing 33-decimal-GB automatic ceiling. Both packs receive the same full ceiling on the 48-GB M5 Pro with at least 36 GB actual reclaimable memory at admission. All six physical, natural-completion and frozen timing checks pass, and every whole-host swap-in/out delta is zero. Exact-source engine, Mac and context CI pass. Temporary host controls are restored. | Work | Original tok/s | Three-bit attention tok/s | Paired ratio | Original request, s | Three-bit request, s | | --- | ---: | ---: | ---: | ---: | ---: | | short-256 | 20.91 | 19.95 | 0.956 | 14.102 | 14.553 | | context-256 | 23.69 | 21.11 | 0.889 | 17.705 | 17.123 | | coding-answer | 21.59 | 21.93 | 1.014 | 6.197 | 5.605 | The original's three workload medians and each of its individual repetitions exceed twenty committed tokens per second here. The candidate's short-workload median is 19.954682 and all three repetitions stay below twenty; do not round that into a passed target. Rates subtract the first output token from the numerator and use the complete decode timer. These controlled interactive bursts, uncontrolled filesystem cache and three-round pilot provide no conservative confidence bound, sustained-session guarantee or other-Mac result. The original remains the supported default. The candidate's first-text medians are 1.775, 5.043 and 1.961 seconds versus 1.906, 6.980 and 2.502 seconds for the original. It improves longer-prompt and coding latency, but loads in 10.854 versus 8.944 seconds. Every natural coding output has identical text and all 81 token IDs. Maximum supervised footprints are 28,943,666,616 bytes for the original and 26,128,172,288 for the candidate, both within the same 33-GB watchdog. The allocation plans expose a concrete tradeoff: the original reserves a 2,048-row prefill chunk and 7,509 expert slots; the candidate reserves 4,096 rows and 8,800 slots. Its larger prefill allowance costs 5,324,800,000 bytes, versus 2,662,400,000 for the original, although these prompts fit within 2,048 rows. A separate prospective check can use the existing explicit prefill allocation policy to trade workspace for cache, retaining the full process budget and acknowledging longer-prompt costs. This is a lead, not a measured improvement or permission to erase an allowance. ### Complete repeated attention comparison, October 5 [[sources/runs/2026/10/2026-10-05-practical-attention-comparison]] preserves the complete twelve-process comparison, the separate stopped prefix and restored host controls. Three paired rounds at each real ceiling pass physical budgets, natural completion and every frozen timing rule. Two whole-host swap-in deltas are 262,144 and 524,288 bytes, with no swap-outs in any process. The original supported pack remains unchanged. The smaller pack uses plain attention; the original retains its corrected predictor. The candidate now approaches original generation speed at 22 GB and reduces the longer-prompt request latency at both ceilings. It still decodes more slowly in all 14-GB cases, loads more slowly at both ceilings, and neither pack reaches twenty committed tokens per second. These controlled interactive-burst medians do not establish sustained-session, other-Mac or statistical qualification. Filesystem caching is uncontrolled. Rates use committed tokens minus one divided by the complete decode timer; ratios pair the original and candidate within each round. | Ceiling and work | Original tok/s | Three-bit tok/s | Paired ratio | Original request, s | Three-bit request, s | | --- | ---: | ---: | ---: | ---: | ---: | | 14 GB, short-256 | 12.75 | 12.55 | 0.984 | 21.647 | 22.033 | | 14 GB, context-256 | 15.38 | 15.11 | 0.982 | 25.831 | 24.955 | | 14 GB, coding-answer | 14.97 | 14.21 | 0.942 | 7.769 | 8.230 | | 22 GB, short-256 | 16.17 | 16.22 | 1.002 | 17.522 | 17.551 | | 22 GB, context-256 | 18.10 | 17.84 | 0.985 | 21.145 | 19.342 | | 22 GB, coding-answer | 18.23 | 18.62 | 1.021 | 6.880 | 6.732 | The natural coding answer has exactly equal text and all 81 token IDs across every process. The candidate's longer-prompt first text arrives at medians of 7.288 versus 9.206 seconds at 14 GB and 5.046 versus 7.056 seconds at 22 GB. Load medians are 10.967 versus 8.532 seconds and 11.050 versus 8.777 seconds. Peak physical footprints stay below 11.34 GB at the 14-GB ceiling and 17.93 GB at the 22-GB ceiling. The current evidence supports a limited prompt-latency benefit, not automatic promotion or a universal speed promise. Exact-head engine, Mac and context CI for the attention implementation pass. The existing original Desktop automatic ceiling is 33 decimal GB on this Mac; the reduced-budget measurements do not answer that default case, which is separately prepared with actual headroom admission. ### Native attention predictor probe, October 5 [[sources/runs/2026/10/2026-10-05-native-affine-attention-probe]] captures one prospectively frozen boundary/attention pair at a real 14-GB ceiling. The same three-bit bytes, two streamed drafts, 1,941 expert slots and complete workloads run through the common native Engine. All three outputs match exactly, both physical watchdogs pass and both cells are timing-eligible under the existing pilot policy. | Work | Boundary tok/s | Plain attention tok/s | Paired ratio | | --- | ---: | ---: | ---: | | short-256 | 11.84 | 12.63 | 1.067 | | context-256 | 14.71 | 15.61 | 1.061 | | coding-answer | 13.16 | 14.28 | 1.085 | The attention tap uses the already implemented one-layer forecast window without learned correction bytes. It reduces wasted reads and forecast evaluation time in the preserved counters. Those counters overlap and are not independent cost components. One directional pair and uncontrolled filesystem caching do not establish a general gain; the original remains the only supported default. Startup is not the mechanism claim because the diagnostic entry points differ in repeated metadata validation. Twenty-two runner tests and ten native protocol/lookahead/pack groups pass. The next unchanged three-round comparison against the original uses both real 14/22-GB ceilings; no new study framework is added. ### Complete practical comparison, October 5 [[sources/runs/2026/10/2026-10-05-practical-complete-comparison]] preserves all twelve fresh processes, the frozen complete analysis, the earlier incomplete attempt and host restoration. Three interleaved rounds at each real 14/22-GB ceiling pass physical budgets, natural completion and every prospective pilot timing rule. One process has 851,968 bytes of whole-host swap-ins and none has swap-outs; the declared one-MiB allowance is not a zero-paging claim. Two-minute quiet/nominal admission and a temporary, automatically restored Photos analysis pause make this a controlled interactive-burst comparison on one 48-GB M5 Pro, not sustained-session or other-Mac performance. Filesystem caching remains uncontrolled. The smaller pack generally decodes slower. It improves the longer prompt's first-text and full-request latency, but neither pack reaches the twenty-token target on these workloads. No automatic promotion follows. Rates below are three-run medians of committed output tokens minus one divided by the complete decode timer; paired ratios are medians of candidate/original rates within each round. | Actual ceiling and work | Original tok/s | Three-bit tok/s | Paired ratio | Original first text, s | Three-bit first text, s | | --- | ---: | ---: | ---: | ---: | ---: | | 14 GB, short-256 | 12.60 | 10.99 | 0.872 | 1.661 | 1.801 | | 14 GB, context-256 | 15.02 | 14.29 | 0.947 | 9.497 | 7.571 | | 14 GB, coding-answer | 14.53 | 12.96 | 0.892 | 2.479 | 2.584 | | 22 GB, short-256 | 15.80 | 15.40 | 0.976 | 1.778 | 1.829 | | 22 GB, context-256 | 17.71 | 16.85 | 0.954 | 7.802 | 5.488 | | 22 GB, coding-answer | 17.61 | 17.61 | 0.985 | 2.512 | 2.457 | The natural coding task emits the same complete 81-token answer in both arms. This single answer does not supersede the separate focused quality limitations. The analysis correctly leaves statistical qualification false; three pilot rounds do not provide its registered lower bound. Startup is also slower for the candidate. The prior excluded attempts remain excluded and do not contribute to these medians. A concrete optimization lead appears in the preserved prefetch diagnostics. The original uses the corrected attention forecast; the candidate still uses the older boundary tap and wastes more reads. These overlapping counters cannot isolate causality. Reuse the already implemented plain attention tap in an explicit bounded candidate probe before considering retraining or new kernels. The original correction's measured benefit does not transfer automatically to altered expert weights. ### Tiny swap-in diagnostics and preserved strict attempt, October 5 [[sources/runs/2026/10/2026-10-05-bounded-swapin-pilot-policy]] records one completed cooled run excluded for four 16-KiB whole-host swap-ins and zero swap-outs. Its second process is interrupted and drained; the attempt remains incomplete and timing-discarded. A separate prospective pilot permits at most 1 MiB of aggregate swap-ins per process, records all bytes, and still excludes any swap-out or larger delta. Twenty-two tests pass, including strict historical replay and byte boundaries. The same two-minute cooling and all physical/pressure/thermal/CPU guards remain. Fresh measurement is prepared and unexecuted; no zero-paging, speed or promotion claim follows. ### Complete retry and prospective cooling interval, October 5 [[sources/runs/2026/10/2026-10-05-practical-cooling-retry]] preserves all twelve repeated processes and their complete frozen analysis. Physical ceilings and natural-answer completion pass. Cells 6 and 9 are timing-excluded for thermal state and for sustained competing UI activity with paging, respectively. No clean medians or throughput qualification follow. All twenty tests pass for extending the existing admission interval from 20 to 120 continuously quiet nominal seconds while retaining the five-minute deadline and every other exclusion. The prospective fresh comparison is pending; it targets controlled interactive bursts and carries no sustained-session speed claim. ### Native arithmetic and focused quality, October 5 [[sources/runs/2026/10/2026-10-05-native-affine-arithmetic-trials-excluded]] preserves four complete executions of the same standalone bytes through the deployed affine kernels. The ordinary allocator cache and existing small-row verification remove two inherited research settings. The final tested recipe completes at both 14 GB and 22 GB within their process budgets. Competing CPU activity excludes every timing from clean comparative claims; the observed response-latency improvement is a lead for a later eligible comparison. No 20-token-per-second or general speed claim follows. [[sources/runs/2026/10/2026-10-05-native-affine-loader-and-focused-quality-launch]] captures the internal shared-loader descriptor, independently priced startup recipe, five passing native catalogue groups, loader plan/refusal checks and the focused quality launch. Both arms freshly run the same previously selected 25 pilot tasks, with existing bounded graders and no statistical noninferiority claim. [[sources/runs/2026/10/2026-10-05-native-affine-focused-quality]] records all fifty outcomes: original 17/25, native candidate 16/25. In family order (facts, multilingual, coding, instruction, tools), the original passes 4, 5, 3, 4, 1 and the candidate 4, 4, 3, 4, 1. The additional multilingual failure reaches the fixed output limit without a completed answer. One tool case improves and another regresses; both packs struggle on this small tool set. All ten sessions complete inside the physical and wall limits. These descriptive results do not establish broad quality equivalence. Native arithmetic does not inherit the reference path's earlier parity or quality qualification, and the original remains the only supported product pack. [[sources/runs/2026/10/2026-10-05-native-affine-lifecycle-and-draft-recovery]] records all 86 applicable native lifecycle assertions passing with a supervised peak of 7,948,523,784 bytes. It covers within-native plain/draft and lookahead equality, nonempty memory/disk/cold continuations, cancellation, pressure recovery and HTTP identity. The original draft-stream gate passes with a peak of 7,878,203,464 bytes. Two failed native fixture attempts remain preserved: immediate EOS prevented the original unframed prompt and arbitrary continuation token from exercising decode. The same lengths and assertions now use framed native inputs and the assistant's actual first token; reference fixtures remain unchanged. These are functional and process-memory observations, not qualified timing. [[sources/runs/2026/10/2026-10-05-native-practical-comparison-prepared]] prepares a prospective repeated comparison using the same existing runner and practical workloads. Native arithmetic has a distinct validated resource identity, and all sixteen runner checks pass. Four allocation proposals run without loading a model. The clean-start check refuses persistent unrelated CPU activity, so this preparation adds no speed observation or promotion evidence. ### Actual app activation and memory controls, October 5 [[sources/runs/2026/10/2026-10-05-real-app-activation-and-memory-controls]] preserves the first failed startup-observation assertion and both subsequent passing actual app checks. The original load and answer worked, but the verifier included separately owned router-cache bytes in its expected prefetch scheduler reserve. Correcting that comparison changes no model allocations or arithmetic. The unchanged activation check passes with a supervised peak of 5,816,603,376 bytes, including bounded health, durable commit, failed replacement, sequential rollback, retry, restart, commit cancellation and preserved-record repair. The original-pack memory-control check passes with a supervised peak of 6,624,908,304 bytes. A setting change during generation preserves the active response configuration, then unloads and reloads under the new nine-GB ceiling with fixed live management. Failed settings remain saved, automatic idle release still works, and drafts survive. These checks establish functional behavior on the available Mac; competing CPU excludes their timings from performance claims. They do not qualify the alternative pack or public distribution. ### Practical configuration pilot, October 5 [[sources/runs/2026/10/2026-10-05-practical-configuration-timings-excluded]] preserves the refused clean comparison, four completed cells of the separately frozen diagnostic pilot, its thermal stop before the fifth launch, and three completed streamed-draft recipe probes. Persistent Bluetooth-service CPU activity excludes every timing from clean performance claims. This is not a completed repeated comparison or a speed qualification. The explicit candidate recipes reduce observed memory use but show mixed response latency; the smaller bundle has not earned promotion. The first candidate recipe used generic Auto, which correctly left an unqualified draft threshold disabled. Explicit streamed-draft probes remove that mismatch. The candidate's reference arithmetic also imposes a 512-row prefill cap and separate precise operations across the trunk. Reusing the existing deployed affine kernels is the next bounded engineering check; reference parity and quality receipts cannot automatically qualify that different arithmetic. ### Standalone Engine and request cleanup, October 5 [[sources/runs/2026/10/2026-10-05-standalone-engine-request-cleanup]] records all 105 native checks passing for the complete standalone candidate with streamed original drafts and explicit uncorrected lookahead. Its physical process peak is 7,428,495,096 bytes, inside the ten-GB fixture watchdog. Short and sparse-context lookahead outputs match demand-only execution; memory/disk/cold prefix recovery, cancellation, pressure shrink/regrow and HTTP checks pass. The original attempt preserved output correctness but left returned speculative slots queued until the next demand. The shared Generator now drains those returns after joining readers at request completion; the exact unchanged assertions pass. Five expert-check catalogue groups and standalone metadata checks also pass. This closes standalone correctness for the tested recipe; clean practical performance and candidate product promotion remain pending. ### Standalone candidate assembled, October 5 [[sources/runs/2026/10/2026-10-05-standalone-affine-export]] preserves the complete frozen export and bounded payload retirement. The standalone bundle owns 78 files totaling 90,232,537,746 bytes before its manifest. Its independently audited complete manifest is `8f8c9a58558828a76d8eb6d40299ac472380adb456790f4f82330bdf5dc352d5`. Retirement removed only 96 authenticated reproducible files from the losing refit and redundant contiguous VQ copy; all outputs and numerical evidence remain. Staging after export is 412,690,948,096 bytes within the existing 430 GB ceiling. The exact bundle is admitted for research execution only. Native standalone parity and practical performance are next; product registration and release remain pending. The first implementation screen for [[records/plan/same-model-quantization-and-automatic-memory-2026-10-02]] establishes metadata geometry, bounded native row decoding and an existing-pack baseline. It does not qualify another pack, establish similar task quality, or achieve the proposed hardware-wide speed target. ### Evidence and method - [[sources/runs/2026/10/2026-10-02-quantization-baseline-v1]] preserves three fresh-process runs of the installed v0.2.27 binary at a fixed 10 GB target, with a frozen short prompt, 128 greedy output tokens, 32,768-token context, automatic draft policy and vision disabled. The small plan did not enable MTP. The primary rate is `(N - 1) / sum(interTokenSeconds)`, preserving draft/verification work between committed emissions. Legacy `N / decodeSeconds` remains separately named. - [[sources/references/2026/10/2026-10-02-vq-implementation-inventories]] preserves each candidate's own pinned config, source-runtime digest and complete shard-header inventory. Full weight payloads have not been downloaded or verified. - [[sources/runs/2026/10/2026-10-02-quantization-native-screen]] preserves the bounded native checks, exact fetched row ranges, fixture hashes, final build identity and every kernel sample. VQ row bits match the separate scalar oracle. This does not establish fused-dot or full-model parity. ### Observations The installed baseline's three primary committed rates were 8.50285, 8.45660 and 8.62913 tokens/s. Median: 8.50285. These are pilot observations for one short fixed-budget workload on the M5 Pro development Mac, not a replacement for the existing Auto-profile measurements. Process peak stayed below the declared target and generation endpoints were nominal. The original runner did not scrub or record ambient developer overrides; later validation confirms the preserved footprint and generator observations, but does not recover missing environment evidence. Host contention checks were endpoint snapshots. Do not promote these runs into a release qualification result. The kernel screen executed both projection shapes, ten routed experts and one/four/thirty-two input-token rows. The affine arms used the same synthetic dense source; VQ arms used independent spread synthetic codes. All 48 cost cells completed within a 0.579 GB process-lifetime peak. There was no observed global paging or non-nominal sampled thermal/power state. Synchronization, host dispatch and evaluation are included; this is not isolated GPU kernel time. At one input token, affine 4/3/2-bit medians were approximately 0.198/0.194/0.187 ms for the 640-by-2560 projection and 0.229/0.228/0.262 ms for the 2560-by-640 projection. Lower storage bits do not guarantee a faster operation. These small synthetic differences do not establish full-model gains or an optimal width. The materialized VQ paths took approximately 0.790 to 0.989 ms in those one-token cells. They first expand selected weights, round to half, convert to BF16 and run a gathered matrix multiplication. That is a substantially more expensive bounded fallback than the affine operations in this screen. It is also a different arithmetic contract from the pinned upstream fused kernels. Do not adopt this materialized path for production decode or use its result to reject optimized fused VQ as a whole. Prefill needs a separate representative routing study. ### Implementation implications and remaining gates `AffineQuantization` and `VQLayout` validate row sizes, packing, codebooks and checked byte arithmetic. The existing adapter now rejects inconsistent per-module descriptors before allocation; the production loader still admits only its existing affine layout. Mixed VQ record sizes and codebooks are recorded independently for each pinned pack, not inferred from advertised bits per weight. The full-vocabulary pilot scorer verifies manifests, file hashes, context identity and complete vocabulary coverage, then reports KL(reference || other), top-1 agreement and paired case deltas. It has only deterministic instrument tests so far. No real candidate logits, task-quality scores, confidence intervals or quality qualification exist yet. The next candidate integration must preserve the pinned fused-dot arithmetic, support mixed byte classes and bounded PLE decoding, and establish a memory-bounded reference execution path. Full model/logit parity and task evaluation precede any alternative pack selection, download activation or default promotion. Actual low-memory and 64 GB hardware still need their own runs. Reduced budgets on this Mac do not qualify other chips. ### Existing-path verification [[sources/runs/2026/10/2026-10-02-quantization-foundation-verification]] preserves the full T0/T1 catalogue (92 passed, no failures or skips), historical first-two-layer parity, static suite and complete Mac app checks. Native memory UI checks passed in Light, Dark and System appearances. This is scoped acceptance of the implementation above, not the full release battery. A separate real-model Mac app check passed lazy load, warm follow-up, a deferred 10-to-9 GB ceiling change, drained reload, an invalid saved 7 GB setting, automatic idle release despite that failed setting, and preservation of the draft and requested preference. The sampled process peak was 6.621959136 GB; released footprint was 0.636340312 GB; the slowest sampled metadata call took 0.0022507083194795996 seconds. Global swap-in/out counters stayed zero. These functional observations do not establish throughput eligibility. No new candidate was loaded. ### Full existing-engine acceptance and fused follow-up [[sources/runs/2026/10/2026-10-02-full-engine-quantization-foundation]] records the complete existing-engine battery on the foundation build: 35 passed, zero failed. It includes current layer and draft references, byte equality through resizing and prefix/sweep paths, both governor drills, process-memory limits, context recall, streaming/tool/restart behavior and full vision serving. This supersedes the earlier statement that this battery had not yet run, only for that recorded foundation binary. The later experimental fused path has its own component checks and is not a qualified full-model path. [[sources/runs/2026/10/2026-10-02-fused-vq-component-pilots]] records the native fused projection, the exact reviewed upstream Metal strings, source-bound builds and two cost pilots. All 78 selected-row cases match the Python MLX 0.32.2 binding bit for bit, across broadcast and per-expert inputs and both sides of the routed-pair dispatch boundary. They share the reviewed Metal implementation, so this checks bindings, packing, casts and dispatch rather than independently proving arithmetic. Constant-dot and invalid-index controls also pass. The final native catalogue reports 92 passed, zero failures or skips. The first fused pilot included a redundant GPU bounds reduction in each call. The second prepares validated CPU routes and private index arrays before timing, which matches inference's requirement to know expert IDs before SSD reads. It still times casts, dispatch, evaluation and synchronization. The frozen pilots remain separate; this is not an interleaved before/after speedup study. | One input token, ten routed experts | Affine 4-bit | Fused D8/K16384 | Fused D4/K2048 | Fused D4/K256 | Fused D2/K1024 | Fused D2/K256 | | --- | --- | --- | --- | --- | --- | --- | | Output 640, input 2560, milliseconds | 0.1995 | 0.2354 | 0.2271 | 0.2143 | 0.2602 | 0.2350 | | Output 2560, input 640, milliseconds | 0.2078 | 0.2419 | 0.2115 | 0.2007 | 0.2198 | 0.2290 | These second-pilot medians are much closer to the affine controls than the original materialization instrument. The process peak was 552,436,816 bytes and the internal thermal/power/paging eligibility checks passed. This keeps fused VQ worth testing with complete artifacts; it does not establish model quality, SSD savings, end-to-end throughput or a 20-token profile. No candidate weights were activated, and no full candidate artifact has been downloaded. The pinned D8 runtime changes reduction at its routed-pair boundary, so full draft verification must not assume one-row arithmetic is preserved. Native cache residency must also not alter dispatch. The next substantive dependency is a bounded full-model reference, including quantized PLE rows: the inspected upstream `ple_stream` helper only streams F16/BF16 `.weight` shards and explicitly leaves VQ `.codes` shards resident. Full reference parity, mixed allocation ownership, held-out quality and actual hardware evidence remain open. ### Bounded full reference and native PLE follow-up [[sources/runs/2026/10/2026-10-02-bounded-vq-reference-and-ple]] supersedes the absence statements above for the VQ 3.2 payload and bounded reference execution. All 139 tensor files, 76,976,259,433 bytes, were staged and checked against the pinned complete-file map. Every reference run rehashed them independently. This is an experimental local artifact, not a supported installation. The corrected first-four-layer traversal proof compares ordinary chunk-major execution and one-layer-at-a-time execution on 513 token IDs, including EOS boundaries and both recurrent and attention layers. Mixer output bits match exactly. The proof explicitly fixes the reference decoded-expert chunk to 32 and binds the architecture, runtime, local instrument and installed library bytes. Earlier failed or incompletely pinned attempts remain in the source. Two complete 48-layer VQ 3.2 forwards then succeeded. The six-token input produced six full-vocabulary rows with a 3,033,779,056-byte process lifetime peak. The 513-token input produced the final sixteen full-vocabulary rows with a 3,231,369,376-byte process lifetime peak. These are instrument feasibility and memory results, not native full-model parity, complete-task quality or generation speed. The reference materializes one layer at a time and streams quantized PLE rows; it is not the product expert cache. The Python PLE storage check passed all twelve real-fixture cases bit for bit against both the resident upstream PLE class and the independent scalar product oracle, with a 27,616,240-byte MLX peak. The native CPU PLE reader separately passes 49 storage/failure assertions and twelve real row-request comparisons across the three inspected layouts. It preserves duplicate order, checks requests before reading, discards incomplete results, and reproduces F16 multiplication followed by BF16 conversion. Its codebook representation and bounded call workspace still need to enter the eventual pack resource ledger; it is not wired into the product NgramStore. The native baseline logit exporter preserves full head batches before selecting rows. The six-token and 513-token functional runs completed under external memory/pressure supervision, with native lifetime peaks of 4,643,262,304 and 6,183,357,608 bytes respectively. These outputs are numerical fixtures, not held-out quality examples. The baseline exporter and reference produce raw receipts; the full-vocabulary scorer still requires explicit compatible artifact and preprocessing identities. [[sources/references/2026/10/2026-10-02-vq-tokenizer-compatibility]] records a material integration constraint: the pinned VQ bundles have a different tokenizer pre-tokenizer/decoder configuration. Their vocabulary, added-token mapping and normalized ordered merges match the original, but a direct Hindi example yields different token IDs. The original tokenizer and chat template match the deployed baseline. Controlled comparisons therefore explicitly freeze the original tokenizer and identical token contexts across arms. Do not replace candidate files silently or describe their supplied tokenizers as identical. The original model's complete visible history has one unchanged non-README payload map across three revisions. Both derivative cards declare that original model. Using its immutable current revision as the comparison's original-checkpoint identity is a documented lineage inference, not an independently replayed quantization conversion. Product qualification must name the exact chosen tokenizer/template and retain multilingual checks. The six-case owned raw-continuation protocol is frozen in `bench/quantization/logit-pilot-v1.json`. It includes prose, code, tool-result context, multilingual text, continuation and record retrieval; it is a distribution pilot, not a held-out agent benchmark. Native and VQ producers run sequentially with bounded memory and complete logit rows. Current performance claims, supported weight selection and defaults are unchanged. Full native VQ model parity, mixed allocation ownership, draft integration, held-out tasks, paired speed measurements and actual hardware qualification remain open. The bounded reference/PLE implementation additionally passed the complete static gates and all 93 native t0/t1 catalogue checks, with zero failures or skips. The first static attempt was invalidated by an in-flight edit to its test driver; the unchanged-driver rerun passed. Both attempts and the source-bound catalogue summary are retained in [[sources/runs/2026/10/2026-10-02-bounded-reference-verification]]. ### Reference normalization correction [[sources/runs/2026/10/2026-10-02-vq-pilot-normalization-diagnosis]] records the completed first matched pilot and its invalidation. Both VQ arms omitted the raw-norm +1 conversion required by the pinned architecture. Strict loading, finite outputs and direct-versus-streamed equality did not detect this shared semantic error. The old VQ feasibility runs and traversal proofs above establish resource feasibility only; they do not establish correctly normalized model outputs. The native affine outputs and isolated PLE/fused kernel checks are unaffected. The complete 4.4 artifact is now verified: 139 tensor files, 103,689,541,903 bytes. An independent bounded read of all 148 affected norm tensors in each VQ pack found that BF16 rounding of 1 + raw value reproduces every corresponding baseline tensor exactly. Gated delta-net normalization stays unchanged. The reference loader now makes this pinned artifact conversion explicit and records its normalization identity; the comparison adapter rejects old producer receipts. Fresh traversal proofs and VQ pilot outputs are required. None of the first pilot's apparent KL advantage is valid quality evidence. ### Corrected matched distribution pilot [[sources/runs/2026/10/2026-10-02-corrected-vq-distribution-pilot]] records fresh exact traversal proofs and both repeated VQ arms. All six contexts completed in each arm with the explicit raw-norm adapter; the unchanged native baseline outputs were reused. The normalization repair changes reference meaning, so these results replace the invalid comparison rather than combine with it. | Owned pilot context | Native affine KL to VQ 4.4, nats | VQ 3.2 KL to VQ 4.4, nats | Native top-1 agreement | VQ 3.2 top-1 agreement | | --- | --- | --- | --- | --- | | memory-prose | 0.224397 | 0.047676 | 0.8125 | 0.8125 | | python-interval | 0.366843 | 0.172623 | 0.7500 | 0.8125 | | tool-result | 0.764051 | 0.440440 | 0.7500 | 0.8750 | | multilingual | 0.343711 | 0.171254 | 0.8750 | 0.8125 | | continuation | 0.715288 | 0.036532 | 0.8125 | 1.0000 | | record-retrieval | 0.249774 | 0.113896 | 0.7500 | 0.7500 | Equal-case mean full-vocabulary KL is 0.44401073962586735 for the deployed native baseline and 0.16373688672074593 for VQ 3.2. Mean top-1 agreement is 0.7916666666666666 and 0.84375 respectively. The candidate has lower KL in every context, better top-1 agreement in three, equal agreement in two and worse agreement in the multilingual context. This supports continuing candidate engineering, not declaring similar task quality. Each context contributes its last sixteen teacher-forced positions; these are correlated owned pilot examples. There is no held-out confidence interval, real tool execution, long-context qualification, vision or draft evaluation. The VQ 4.4 arm is a quantized proxy, and the baseline and VQ arms also differ in their complete native/reference implementations, so this does not isolate weight quantization alone. Native full-model parity and complete-task evaluation remain mandatory. No speed was measured or pack promoted. ### Complete native expert records and ownership [[sources/runs/2026/10/2026-10-02-native-vq-complete-records]] extends selected-row kernel checks to complete real gate/up/down matrices and the pinned compiled SwiGLU composition. Both packs pass exact output-bit comparisons for layers 0 and 2, expert IDs 0, 1, 7 and 511, and one, two and three token rows with ten routes each. Duplicate routes preserve their order. These real layouts do not contain D8; its dispatch boundary remains covered by the separate fused fixtures. The fixed routes do not test the model router, shared expert or full-model behavior. The immutable staging batch admits only complete unique expert records. Its allocation-class key includes all three projection layouts; equal byte counts alone cannot make two banks interchangeable. Shared codebooks are counted separately. Review found that retaining mutable MLXArray objects did not protect admitted weights from caller-side context replacement. The v2 implementation keeps private array contexts, and independent mutation of each caller-owned codes, books and scales group preserves the expected outputs. This establishes value retention, not leases or pins for externally reused mutable cache memory. Both final native record checks pass 34 assertions, alongside the geometry, metadata and PLE storage checks. The final source-bound catalogue passes 93 checks with zero failures or skips, and the static suite passes. The recorded binaries and full outputs remain in the raw source. No full native candidate model, mixed mutable cache, draft integration, task-quality result or throughput profile is qualified by this checkpoint. ### Native dense-block candidate profile [[sources/runs/2026/10/2026-10-02-native-vq-trunk-profile]] establishes exact first-block parity under an explicit candidate arithmetic profile. The reference exports the real first linear-attention block and hyper-connection, with corrected folded norms, from both completely verified artifacts. Their complete fixtures are byte-identical. Native checks cover one, three and seventeen BF16 token rows, a subsequent continuation, retained convolution windows, FP32 recurrent state and non-mutating readout. All 75 trunk assertions pass. This does not cover the full candidate layer stack, QSA/PLE integration or speculative state recording. The first native attempt matched the hyper-connection outputs but failed recurrence-state bits and longer outputs. The final candidate profile matches the pinned reference's ordinary reduction, compiled decay and BF16 beta; the deployed kernel retains its compensated reduction and existing arithmetic. Grouped RMS and recurrent query/key normalization also follow the explicit candidate profile, and the candidate hyper-connection supports its quantized injection projection. The Swift binding's documented mlxNone context maps to the exact absent RMS weight in the C API. The failed attempts remain in the source, including a pre-allocation concurrency refusal and a repaired reference tuple-unpacking error. The final native binary is `4d07c1c6c8e629365b7415225dc45dc6ec83f7c1ac1e216fe177be90b5f8cd59`. A complete 48-layer forward on the unchanged memory-prose pilot IDs reproduces the previous deployed native logit SHA exactly: `7c2b5b7b78e507d28d3ca85b2e10a32519f0ebd7d67adb20b489bf6479e92f32`. This checks compatibility on that one case; it does not replace the complete existing-engine release battery or establish candidate quality. The catalogue passes 93 checks, with zero failures or skips, and the static suite passes. Candidate full-model checkpoint loading, mixed mutable caching, PLE/draft integration and all release qualification gates remain open. ### Native owned tensor-file primitive [[sources/runs/2026/10/2026-10-02-native-vq-owned-file-reader]] records the bounded native file reader needed by direct candidate loading. Its constructor verifies the complete payload through the descriptor it retains for later reads; its caller must separately authenticate the supplied file/header identities against a pinned pack manifest. Reads check tensor extents, cancellation and unchanged file metadata, with bounded allocations and syscalls. The tests preserve the distinction between owning a verified descriptor and following a replaceable filesystem path. All 34 storage assertions pass, covering cancellation, complete-file corruption, truncation, path reuse, lifetime retention, malformed extents, non-regular files and tensor coverage. The final catalogue passes 94 checks with zero failures or skips. This is a synthetic storage gate, not actual VQ artifact integration, model parity, asynchronous pool ownership or performance. The preceding trunk-profile static result is not presented as a fresh static run for this storage addition. ### Authenticated artifact reads and short complete-stack parity [[sources/runs/2026/10/2026-10-02-native-vq-authenticated-checkpoint]] records direct native loading from both pinned artifact inventories. Config/index bytes and the complete-file digest map are authenticated before data access. Demanded files are independently verified through retained descriptors; bounded selected expert and PLE reads match their existing fixtures. Counters describe demanded files, not a native read of every unused payload. Draft metadata remains independent and its tensors are excluded from this main-model probe. [[sources/runs/2026/10/2026-10-02-native-vq-complete-stack-smoke]] records a fixed three-token pass containing EOS and a subsequent one-token continuation. Both VQ 3.2 and 4.4 match all 320 observed boundaries exactly: hidden states through 48 layers, convolution/recurrent state, PLE convolution, QSA keys/values and raw indexer state, final mixer and complete vocabulary logits. Sampled native physical peaks were 1,926,531,184 and 2,003,257,576 bytes respectively; the functional harness enforced its separate process and actual-headroom bounds. Wall times include authentication and reads and are not committed-token throughput evidence. Global paging counters are retained as diagnostics. The first three native attempts failed at the first QSA layer. [[sources/runs/2026/10/2026-10-02-vq-bf16-sigmoid-arithmetic]] shows identical preceding hyper-connection intermediates and a BF16 sigmoid difference at input -6.84375. Explicit precise exponential and BF16 intermediate rounding match all 65,280 finite BF16 input patterns against the pinned Python GPU output. This is an observed arithmetic contract; an exact compiler-level root cause is not established. Only the candidate profile uses this operation. FP32 sigmoid and deployed model arithmetic remain unchanged. The native exhaustive diagnostic initially failed because a scalar test literal inferred FP64, which the GPU does not support. The first crash's accessor hypothesis was incorrect; subsequent explicit evaluation exposed the actual error. Those failures stay attached and separate from the passing v4 complete-stack engine. The final test uses an explicit Float literal. This advances step 4 without completing it. Ordinary prefill batch sizes, sparse indexer activation, native generated sequences, mutable cache pins and resize/late-I/O ownership, draft/rollback, vision, held-out task quality and speed remain unqualified. The observed four tokens cannot justify a production pack or a 20-token claim. [[sources/runs/2026/10/2026-10-02-native-vq-complete-stack-validation]] records the repaired exhaustive sigmoid gate, all 95 native T0/T1 checks passing with no skips, unchanged deployed full-model logit bits on the fixed memory-prose case, fresh baseline payload verification and six metadata rejection cases. The v7 inference sources are identical to the independently passing v4 complete-stack sources; only the diagnostic input/readback file differs. These checks do not replace the complete heavyweight app, governor and hardware qualification gates. The frozen v7 complete static suite also passes in [[sources/runs/2026/10/2026-10-02-native-vq-complete-stack-static]], including planner, memory-override, transport and installer gates. The first attempt stopped on a source-wrapper token list interpreted as a wiki-link; the authored wrapper was corrected without changing its raw transcripts. This result applies to the authenticated short-stack checkpoint, before subsequent batched-route changes. ### Partitioned fused-route checkpoint [[sources/runs/2026/10/2026-10-02-native-vq-partitioned-route-parity]] records bounded complete-record partitions with explicit original-batch arithmetic dispatch. Small local partitions cannot accidentally change the D8 reduction method. The component gates compare exact projection bits across the dispatch boundary, then exact complete SwiGLU at capacities of one, two, three and thirty-two experts, including duplicate routes and restored pair order. Both pinned packs retain exact original short-stack outputs and pass a new eight-token pass plus three-token continuation through all 320 observed hidden/state/logit boundaries. Each larger run exercises two staging partitions and at most 32 live expert records per batch. All 95 native catalogue checks pass without skips. The source, artifact and reference identities are frozen. Compilation and these numerical/ownership checks cover this increment; the preceding full static suite is not claimed as freshly rerun. This is synchronous immutable staging. It does not supply persistent residency, mutable cache pins, asynchronous generation fences, resize accounting or a production memory governor. The larger-batch upstream prefill path beyond 4,096 routed pairs remains a separate implementation/parity gate. Eleven functional tokens do not establish task quality, latency or committed generation speed. ### Segmented prefill component checkpoint [[sources/runs/2026/10/2026-10-02-native-vq-segmented-prefill]] records the pinned large-prefill dispatch audit and exact native expert composition for both packs at 410 and 512 prompt rows. The actual default is fused segmented GEMM, with recorded kernel flags and executed variant names. The reference's fixed decoded-expert chunk of 32 affects only its fallback; it is not evidence that the admitted large-prefill path materializes decoded matrices. No prior measured outputs are retracted by this dispatch clarification. The native wrapper uses the exact reviewed segmented kernel and its preprocessor specialization, preserves complete expert token segments across storage partitions, and restores routing order. Sixty-four real experts with skewed routes exercise two staging batches, partial tiles and both codebook placements. Complete SwiGLU outputs match exactly in both inspected layer families and both packs. All 95 native catalogue checks and the reference/source unit checks pass. These are component results with bounded process-memory supervision, not full-model prefill or timing qualification. Full-model attention, PLE and recurrent state at these batch sizes are the next parity gate; mutable residency, native generation, quality and speed remain open. ### Complete native prefill and rotary follow-up [[sources/runs/2026/10/2026-10-02-native-vq-complete-prefill]] records both packs passing the fixed prefill512-decode1-v1 profile. Each performs a 512-token pass and one continuation using retained state. All 320 boundaries match exact complete logical bytes, including full-vocabulary logits computed at the original pass shape. Every layer executes segmented prefill. Both runs reach eleven record-staging batches within a layer while retaining at most 32 expert records per batch. Native process peaks are 3,069,757,840 bytes for VQ 3.2 and 3,082,111,328 bytes for VQ 4.4, within the frozen 4 GB diagnostic bound. Global swap counters remain diagnostic; this is not a clean-timing result. The native readers independently verify 137 demanded files, totaling 73,780,799,521 and 100,494,081,991 bytes respectively. This is not a claim that unused optional files were opened. The reference producer independently verifies the complete artifact map. The low diagnostic footprint comes from releasing dense blocks and expert staging; it is not a production memory minimum. All 95 native T0/T1 checks pass with no skips. [[sources/runs/2026/10/2026-10-02-vq-prefill-rotary-arithmetic]] retains the failed native attempts. The first mismatch occurs at QSA layer three after three exact layers. Its preceding projection and normalization intermediates agree. The native inverse-frequency vector differs at 23 of 32 FP32 entries and matches a fast Metal power probe exactly. Precise power matches the pinned Python vector, and the repaired candidate's frequency vector plus all first-512-position FP32 sine/cosine bits pass their dedicated checks. A preliminary microscope script failed on an unsupported array.repeat method before its GPU experiment; its corrected successor and the failed transcript are retained. The deployed rotary path remains unchanged. This completes the bounded ordinary-prefill numerical checkpoint. It does not establish generated sequences, longer-context sparse selection, mutable expert residency, live governance, draft or vision behavior, held-out quality or committed generation speed. Those gates remain open. [[sources/runs/2026/10/2026-10-03-native-vq-prefill-validation]] records the complete static suite passing on this frozen corrected prefill binary, including 420 memory-override cases, planner, transport and installer checks. Eight malformed full-prefill manifests are rejected before execution. Fresh original-pack verification succeeds, and its frozen complete-vocabulary memory-prose logit bits remain unchanged. This is scoped regression evidence, not candidate app/governor, quality or speed qualification. ### Native autoregressive numerical follow-up [[sources/runs/2026/10/2026-10-03-native-vq-greedy-parity]] records the frozen greedy16-memory-explanation-v1 profile for both pinned packs. A 44-token owned literal continuation is encoded with the original tokenizer. Each reference and native implementation independently selects and feeds back sixteen argmax tokens, with the final sampled token left unconsumed. The native sequences and all 2,560 complete tensor boundaries per pack match exactly, including retained state and complete-vocabulary logits. Both finish at the length cap with 59 consumed tokens. Real sampled EOS termination is not exercised. Native process peaks are 1,881,000,024 bytes for VQ 3.2 and 1,849,722,944 bytes for VQ 4.4. Each uses up to seven staging batches in a layer and at most 32 live immutable expert records per batch. These are streamed diagnostic peaks, not production resident-memory floors or speed measurements. The corresponding Python process peaks are 2,783,791,960 and 2,943,159,272 bytes. Every run remains within its declared 4 GB bound. Twelve metadata/chain/stop/CLI rejection cases and all 95 native T0/T1 checks pass without skips. The build comparison binds unchanged engine inference sources to the prior fully validated prefill build; only the diagnostic and CLI dispatch change. Compilation, exact generation, refusal checks and the native catalogue cover this increment. No new full static, app, governor, vision or hardware qualification is claimed. Persistent caches, sparse long-context selection, draft, held-out quality and complete-configuration performance remain required. ### Sparse-selection checkpoint, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-sparse-selection]] records both packs passing the fixed 2053-token prefill plus one-token continuation. All 984 complete output/state boundaries agree with the independent reference, including the 24 actual Boolean sparse masks. Four large passes exercise segmented expert staging, and the final short pass and decode cross the selection threshold and partial block. This does not qualify the entire supported context range. The native peaks are 3,136,752,112 bytes for VQ 3.2 and 3,130,771,904 bytes for VQ 4.4, within the 4 GB diagnostic process bound. All eight malformed-fixture cases and 95 native checks pass, alongside the five reference unit tests. Global swap counters are retained as diagnostics, with no clean timing claim. These bounds describe the one-layer-at-a-time probe, not a production resident-memory floor. Mutable caching, production generation, draft, vision, held-out task quality and complete-configuration speed remain required. ### Resident-bank component checkpoint, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-resident-bank-component]] records synchronous mutable bank ownership for one complete-record allocation class. Separate contiguous backing allocations are validated before direct writes. Complete demanded sets are pinned before CLOCK replacement, publication follows all six checked pieces, and GPU completion precedes pin release. The inspected layers match the independent expert-output fixtures across cold/hot access and eviction. Coverage correction: the old 4.4 fixtures selected two layers from the same allocation class; see the audit below. Partial reads, cancellation, malformed pieces and reentrant clearing are refused, with retries and retained-output independence checked. Both packs pass all 244 complete-record checks, including the earlier composition/partition cases. Both segmented-prefill component fixtures and both full-model 16-step greedy regressions pass after sharing the projection composition. Those model runs still use immutable staging: this is not model-wide cache qualification. The bounded bank admits one to thirty-two records and has no asynchronous prefetch, resize or governor integration. Twelve malformed greedy manifests and all 95 native catalogue checks pass. Larger mixed-class residency, byte budgeting, complete-model cache wiring, draft, vision, quality and speed remain open. ### Rotary coefficient portability correction, October 3 [[sources/runs/2026/10/2026-10-03-vq-rotary-coefficients-and-ci-portability]] records main CI run 37098784524 failing the rotary digest checks in ordinary and instrumented catalogues. The M5 Pro measurements remain valid on their recorded machine; Metal precise power did not reproduce their inverse-frequency bits on the CI runner. The other 94 catalogue checks and public-library job passed. This is a reproduced portability failure, not a waived check. The candidate now stores the 32 FP32 inverse-frequency words extracted byte-for-byte from the verified independent Python fixture. The table is restricted to the pinned dimension 64 and base 10,000,000. The deployed public rotary constructor is unchanged. All expected hashes remain unchanged: this removes GPU power from coefficient construction without relaxing comparison or claiming correctly rounded mathematical power. On the corrected local binary, the frequency/angle digests, both 984-boundary sparse model profiles, all 95 native catalogue checks and the complete static suite pass. Static validation includes 420 memory-override cases, 97 planner checks, transport and installer gates. Bank/projection sources are unchanged from the preceding component and greedy validation binary. Remote requalification is pending at this capture; local success does not establish universal GPU arithmetic or declare the original CI failure remotely resolved. ### Component coverage correction, October 3 [[sources/runs/2026/10/2026-10-03-vq-allocation-class-coverage-correction]] withdraws the claim that the earlier real expert components exercised both allocation classes in each pack. Layers 0 and 2 span both VQ 3.2 classes, but share one VQ 4.4 class. VQ 4.4 requires layer 3 as the second representative. The earlier complete-record, segmented-prefill and resident-bank component results remain valid for their actual recorded layers. Full-model runs independently covered all 48 layers and remain valid. The artifact-bound geometry audit records the layer assignments and complete record sizes. The new allocation-classes-v1 fixture mode selects one real representative of each descriptor triple and requires the native checker to confirm two distinct classes. New fixture outcomes are separate evidence; the audit itself is not new numerical or performance qualification. ### Fixed mixed-class cache checkpoint, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-fixed-record-cache]] records both packs passing complete-model numerical checks with synchronous resident expert banks. Each validated allocation class has 96 physical rows and layer codebooks are retained once. Both 16-step greedy profiles match all 2,560 boundaries; both sparse profiles match all 984 boundaries. Hits and CLOCK evictions occur, and every pin is released. All runs remain within the 4 GB diagnostic process bound. The dense trunk remains streamed one layer at a time and large prefill uses separate immutable sweep staging, so these peaks are not production resident floors. New artifact-bound component fixtures cover both actual descriptor classes in each pack, including the corrected VQ 4.4 representative layer 3. Real-output checks exercise every physical bank row, with hot boundary-slot reads forbidden. Nine malformed coverage/identity/layer or cache-mode cases are refused, and all 95 local native catalogue checks pass. This does not resolve the separately observed remote rotary arithmetic issue. Larger effective caches, asynchronous ownership, resident trunk, draft, vision, held-out quality and complete-configuration speed remain open. No candidate is admitted to production or Auto. ### Exact finite rotary table checkpoint, October 3 [[sources/runs/2026/10/2026-10-03-vq-finite-rotary-table-portability]] records a second remote arithmetic failure: CI run 37101921994 agrees on the pinned inverse frequencies but differs on sine/cosine FP32 values. Host-double trigonometry was investigated and rejected because one unique admitted sine coefficient crosses a BF16 rounding boundary. Neither numerical tolerance nor existing golden changed. The candidate now embeds the complete 525,824-byte FP32 angle table from the pinned independent reference for positions 0 through 2053. Duplicate halves are reconstructed exactly. Both the generator and native reader bind its frozen digest, and requests outside that bounded research horizon are refused. This is a finite coefficient artifact, not a production context policy or universal transcendental function. The public rotary path is unchanged. All original inverse-frequency and 512-row FP32 digests pass, along with new whole-horizon and final-stride checks. Both packs preserve all 16 greedy tokens and 2,560 boundaries and all 984 sparse boundaries with resident expert caching. All 95 local native catalogue checks pass. The complete static suite passes after correcting its synthetic tool registry to include the new table-source suite; the initial harness failure is preserved. Remote requalification remains pending at this capture. The table does not establish full-model portability on other GPUs or admit a candidate to production, Auto, quality or speed qualification. ### Resident text checkpoint, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-resident-text]] records both packs retaining all 50 text-weight families, a 5,318,309,400-byte payload, alongside the fixed expert banks. Loading separately prices the largest 635,699,200-byte copy, checks real incremental headroom and refuses an undersized budget before payload reads. No partial owner is published. Both packs match all 16 greedy tokens and 2,560 full boundaries, and all 984 sparse boundaries. Physical peaks are 6,922,458,728 and 8,193,513,112 bytes for the VQ 3.2 greedy and sparse profiles, and 7,055,021,720 and 8,347,375,328 bytes for VQ 4.4. All fit the explicit 10 GB diagnostic process bound. These are measured configurations with small fixed banks, not an automatic production ceiling or speed qualification. The default streamed-text regression remains exact within its original bound, and all 95 native catalogue checks pass. The new eight-case CLI test supplies every required argument and matches the actual mode-validation error. It corrects a coverage gap in the earlier cache test: three cases omitted required arguments and therefore did not reach the intended mode guard. The old six coverage/identity/layer metadata refusals remain valid. Twelve malformed greedy manifests also pass their refusal checks. The preceding full static run is preserved as prior evidence, not claimed as newly rerun for this increment. Useful larger expert-cache capacity, asynchronous ownership, PLE caching/read parallelism, production generation and governor integration, draft, vision, held-out quality and speed remain open. No candidate is activated or admitted to Auto. ### Larger fixed-cache checkpoint, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-wide-record-cache]] records a second fixed research layout: 512 rows for the class used by forty-one/forty-two layers and 96 for the smaller class. The whole cache/book reservation is bounded by 1.8 GB before allocation. Only inspected U8/D2 and packed D4/K2048 families admit the enlarged banks; other families retain their earlier limit. This is not an automatic allocation policy. Real components match the reference through every admitted physical row and hot boundaries up to slot 511, within the existing 2 GB component bound. Both fully resident text configurations preserve all 16 generated tokens and 2,560 boundaries and all 984 sparse boundaries. Their highest process footprint is 9,433,913,080 bytes, within the existing 10 GB research bound. All pins are released, three fully specified wide-mode conflicts are refused, and all 95 native catalogue checks pass. On the same greedy fixture, VQ 3.2 cache hits rise from 311 to 2,379 and loads fall from 13,952 to 11,884. These aggregate counters include prefill and continuation; they are not a decode-only speed measurement. Hash observers and setup work remain inside the diagnostic process. Production serving, asynchronous read ownership, runtime resizing, task quality and complete-configuration throughput remain separate gates. No candidate is admitted to Auto. ### First complete-model generation-cost pilot, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-generation-cost-pilot]] binds a new frozen 44-token prompt / 128-token generation pilot, the exact binary and both artifacts. Separate validation first proves all sixteen complete vocabulary arrays and sampled tokens against the independent greedy reference while intermediate state observation is disabled. Measurement also omits final-logit hashing. Full finite-value and resource checks remain active, and all 138 main payloads are authenticated before the request interval. The process begins with empty expert banks but an uncontrolled OS page cache after authentication; this is not cold-SSD latency. | Pack | Three committed generation rates, tokens/s | Median rate | Median TTFT, seconds | Median authentication/resident setup, seconds | | --- | --- | --- | --- | --- | | VQ 3.2 | 3.07638 / 2.98809 / 3.04660 | 3.04660 | 9.26446 | 29.18210 | | VQ 4.4 | 2.54737 / 2.52407 / 2.52353 | 2.52407 | 10.51741 | 39.10702 | All six measurements complete 128 non-EOS tokens and meet the predeclared observed timing criteria. Three alternating paired rounds are preserved; there are no replacement runs. The rate is 127 divided by elapsed time from first to last committed emission. Full request time, every interval, startup observations, per-token thermal/power observations, global paging and separate prefill/decode cache counts remain in the receipts. Decode hit fractions are approximately 30.82% and 30.76% in this fixed small-cache experiment. The final capped output is a reasoning continuation, not a completed-task quality evaluation. EOS termination remains unexercised. The fastest representation here remains far below the proposed target. The older installed-pack pilot uses a different instrument and context policy and lacks complete ambient-override evidence, so no baseline speedup/slowdown is inferred from their numerical difference. This result does not qualify optimized VQ, larger contexts, MTP, vision, Auto or another Mac. Sixteen actual CLI/input refusal cases, the ordinary 3.2 full-state generation regression and all 95 native catalogue checks pass on this pilot binary. The complete e56f8fc remote CI separately confirms the prior finite-rotary checkpoint; the full static suite has not been claimed as freshly rerun here. ### CPU attribution follow-up, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-generation-cpu-profile]] preserves one separately budgeted sampled run. Its timing is discarded because sampling perturbs execution. Of 2,083 main-thread samples, 893 sit in the expert-piece pread subtree; GPU waits are also substantial. Process-wide leaf totals include other threads and cannot use the main-thread denominator. The profile supports testing bounded parallel demanded reads first; it does not establish a precise wall-time breakdown or a measured optimization gain. ### Bounded parallel demanded reads, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-bounded-parallel-record-reads]] records immutable authenticated read plans and bounded CPU workers. Each batch reserves all retained records plus per-lane syscall storage before it starts. Workers receive no MLX objects, mutable checkpoint map or bank addresses. All workers join before ordered owner-thread publication, including failure and cancellation. The entire demanded set stays pinned through evaluation and GPU completion. The initial twelve-lane choice is a measured comparison point, not a universal optimum. Both packs retain exact greedy and sparse reference results, within the unchanged 10 GB process bound. Maximum whole-model footprint is 9,435,043,552 bytes; reserved staging reaches 95,558,400 bytes for VQ 3.2 and 115,219,200 for VQ 4.4, under its 128,000,000-byte admission limit. Actual parallel overlap, malformed requests, cancellation with active siblings, drain-before-return, failed publication, retries, destructive reuse and pin release are exercised. The first build owns these full-model results. The second build changes only one Sendable annotation, refusal messages and diagnostic assertions; both pass all 96 native checks. The second also passes six actual mode conflicts and the complete static suite, including 420 memory-override cases and 97 planner checks. Source identities bind the distinction. [[sources/runs/2026/10/2026-10-03-native-vq-parallel-read-paired-pilot]] compares serial and parallel demanded reads for VQ 3.2 using that same final binary, prompt, resident text, 512/96 banks and 128-token workload. Both modes pass separate complete-logit prefix validation. Every measured 128-token sequence is identical. All three paired rounds meet the frozen observed timing conditions; no replacement runs occur. | Mode | Three committed generation rates, tokens/s | Median rate | Median TTFT, seconds | Median full request, seconds | | --- | --- | --- | --- | --- | | Serial | 2.99876 / 3.06689 / 3.08361 | 3.06689 | 9.25776 | 50.66785 | | Parallel | 4.27871 / 4.27802 / 4.28506 | 4.27871 | 3.41699 | 33.10364 | The median of paired throughput ratios is 1.3949069410128767, approximately 39.49% improvement. This remains below the target. Setup authenticates all main payloads before timing, so no cold-SSD claim is made. This short, capped reasoning continuation is not held-out completed-task quality. Sixteen input/validation refusals and both serial/parallel receipt-cross-binding refusals pass before output allocation. No default changes or hardware promotion follow from this pilot. Demanded-read workers now have bounded ownership. Predictive prefetch, service cancellation integration, dynamic byte budgeting/resizing, persistent PLE rows, production generation, longer contexts, draft, vision, held-out task quality and qualified Auto selection remain separate work. The experimental default remains serial and public model loading still rejects VQ. ### Parallel-path attribution, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-parallel-generation-cpu-profile]] preserves a separately budgeted five-second CPU sample whose timing is discarded. The main-thread tree has 2,088 samples: 471 in demanded-read work/join, 453 in bank output evaluation and 240 across the three repeated expert constructors. PLE row work appears in a smaller 28-sample block. The next test is a bounded code-only cache of identical kernel closures, preserving fresh array contexts and every bank lifetime rule, followed by exact reference checks and a new matched comparison. Attribution is not a measured gain. The same source separately records complete CI success for the earlier 9d5639a checkpoint. ### Kernel-object reuse checkpoint, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-kernel-object-reuse]] records a bounded cache of five exact Metal source/header specializations. It retains code only. Model tensors, array contexts, routing metadata, templates and bank leases remain independently owned. All five families execute at both input widths; both packs preserve every greedy and sparse reference boundary and pass the real bank ownership components. All 96 native checks pass, with maximum whole-model footprint 9,448,642,320 bytes inside the 10 GB bound. The frozen two-binary pilot permits exactly the kernel implementation, its extracted cache and diagnostic changes, with identical Metal-library bytes and one native profile. Both executables pass separate complete-logit validation. All six observed timing runs are eligible and every 128-token sequence is identical. Median committed generation rises from 4.27597 to 5.10865 tokens/s. The median paired ratio is 1.1913453519006945, about 19.13% improvement; median TTFT falls from 3.56642 to 3.20329 seconds. The predeclared engineering adoption rule passes, so code reuse is retained. This does not qualify task quality, the speed target, another hardware class or an alternative production pack. The build's intended preflight failed because its wrapper imported the helper from the wrong module, and the shell incorrectly continued to make. The source preserves the correction and the immediate during-build observation of over 22 GB reclaimable with normal pressure. That observation is not called a pre-build pass. Subsequent launches use a failing-command guard and the correct helper. All model runs separately pass their headroom checks. This increment repeats native numerical and ownership gates, not the complete static or app suites. Large transcripts in this new source are losslessly compressed as base64 text, with original/normalized byte counts and hashes and a round-trip decoder check. Older source captures remain byte-for-byte unchanged. ### Smaller-pack preparation, October 3 [[sources/runs/2026/10/2026-10-03-vq-2.1-artifact-preparation]] binds the VQ 2.1 research download to its exact revision, config and complete file map. The 139 tensor files total 51,426,593,465 bytes. Payload verification is in progress at this checkpoint. Staging admission is separate from numerical and product admission; reference and native allowlists still refuse this pack. Its three complete-record classes use 1,280,000, 1,382,400 and 2,611,200 bytes, with representative layers 2, 27 and 0 respectively. Compact 96-row banks total 506,265,600 bytes; shared books add 19,535,872 bytes. These are inspected storage geometries, not measured process floors or speed estimates. The bundled runtime differs from the reviewed newer runtime. An AST-only audit preserves the changed decoding and segmented-prefill expressions without executing the downloaded code. Qualification must explicitly bind a reviewed execution source and independently establish normalization, traversal and numerical agreement. As additional regression coverage of the committed kernel-object cache, all 78 real selected-row fused fixtures pass, including D8 and D4/K256 families absent from the complete larger-pack runs. The separate research download was active, so this is functional evidence only. Complete 2.1 parity, quality, memory ranges and throughput remain open. No alternative pack is offered or selected automatically. ### Complete smaller-pack parity, October 3 [[sources/runs/2026/10/2026-10-03-vq-2.1-payload-and-normalization]] supersedes the preceding in-progress staging state: all 139 tensor files passed full hashes. An independent CPU tensor audit found exact agreement for 148 RMS norms after one BF16 +1 fold and 36 unchanged gated norms. This authenticates storage and normalization; draft and vision execution remain unqualified. [[sources/runs/2026/10/2026-10-03-native-vq-2.1-three-class-parity]] captures the source-bound native implementation and all twelve sequential functional runs. All three expert record classes and physical research bank rows pass independent parity. Complete batched continuation checks 320 boundaries; compact and wide greedy each check 2560 boundaries over sixteen self-fed steps; sparse attention through 2054 consumed tokens checks 984 boundaries. Every check matches exactly. Compact greedy peaks at 7,018,174,080 process bytes, wide greedy at 7,694,063,712 and wide sparse at 8,814,728,856. These are bounded probe observations on the development Mac, not production memory floors or speed results. The wide banks hold 512 records in the dominant class and 96 in each other class, costing 1,038,745,600 bytes plus 19,535,872 shared-book bytes. Explicit profile metadata binds all three classes and their layer counts; no heuristic infers the dominant class. Reference execution explicitly binds the reviewed newer runtime and separately identifies the distinct older bundle. Missing or inconsistent identities and missing class coverage fail before model allocation. Both larger packs retain exact real-record, greedy and sparse parity. All 96 native catalogue checks, 22 native metadata refusals and six Python producer source refusals pass. Full static checks passed for the frozen native binary, including 97 planner and 420 memory-control checks. Final Python traversal/fetch admission changes followed that run's Python portion; nineteen focused tests cover those final sources. No timing is inferred from these functional runs. Public VQ loading, complete-task quality, feature qualification, Auto integration and the 20-token target remain open. ### Smaller-pack quality screen, October 3 [[sources/runs/2026/10/2026-10-03-vq-2.1-distribution-screen]] records all six new VQ 2.1 forwards against the identical original-tokenizer contexts and hash-verified corrected controls. Against the VQ 4.4 quantized proxy, mean case KL is 0.46529280439934456 versus 0.44401073962586735 for the installed baseline. Mean top-1 agreement is 0.6770833333333334 versus 0.7916666666666666. Coding and tool-result KL worsen; top-1 agreement worsens in five contexts and ties in one. The smaller representation therefore has an observed quality risk, despite exact native implementation parity and lower storage cost. VQ 2.1 stays out of promotion and the next speed-optimization work remains on VQ 3.2. The single-context VQ 2.1 throughput pilot is not started: screening a faster version of a candidate with this unresolved quality loss would not earn product integration. The payload, implementation and all negative evidence are preserved. Reconsider with a separately frozen held-out noninferiority result or a new, independently qualified same-checkpoint representation; do not tune these six cases into a final test. No pack is qualified by this pilot, and VQ 3.2 still requires complete-task and full-configuration gates. See [[records/decisions/vq-2.1-held-after-distribution-screen]]. ### Bounded allocator reuse, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-bounded-allocator-reuse]] records one native source change: resident research paths preserve an already bounded allocator cache across layers. The bound remains 128,000,000 bytes. Streamed text or a larger cache still clears every layer. Every arithmetic operation, finite check, real-headroom check, GPU drain and expert-bank lease remains intact. All three packs preserve every greedy and sparse reference boundary. All 96 native catalogue checks pass. The largest whole-model process peak is 9,499,711,224 bytes, below the unchanged 10 GB bound. The full static suite passes. Both exact performance binaries independently match the sixteen-step complete-logit reference before measurement. All three alternating pairs are eligible and preserve the same 128 generated token IDs. Median committed generation rises from 5.073901973129132 to 5.516094047564276 tokens/s; the median paired ratio is 1.0922703436950998. The median paired TTFT ratio is 0.9437188142389586. The frozen engineering adoption rule passes, so bounded reuse is retained. Separate experiments' gains must not be multiplied into a new measured result. This short-context, non-speculative, page-cache-warmed pilot remains below target and establishes neither completed-task quality nor another hardware profile. ### Independent draft metadata, October 3 [[sources/runs/2026/10/2026-10-03-vq-independent-draft-metadata]] authenticates the shared draft sidecar independently of all three main packs. Its own recipe is affine six-bit with group 32, covering 71 tensors and 20 quantized modules. Payload storage is 2,297,552,576 bytes, including 2,202,009,600 routed-expert bytes. This is storage geometry, not a process-memory floor or measured runtime reservation. The new CPU inventory tool binds the exact header and complete tensor extents, rejects inherited main-pack recipes and unknown overrides, and distinguishes header authentication from optional full-payload verification. Four focused tests and the static harness checks pass. The read-only source review identifies fused projection, per-stream hidden normalization and zero-centered norm conventions that differ from the existing native MTP adapter. No downloaded source was executed. A compatible independently validated draft adapter is still required; no draft acceptance, speed or product feature is qualified. ### Next candidate hypothesis, October 3 [[sources/runs/2026/10/2026-10-03-vq-allocator-profile-and-dense-overlay-feasibility]] preserves a separately sampled generation run. Its timings are excluded. Demanded reads, expert GPU evaluation and route evaluation remain substantial. A header-only audit identifies 498 compatible dense quantized modules whose existing same-checkpoint four-bit arrays could recover 2,424,832,000 stored bytes while retaining the VQ experts, n-gram tables, norms and unmatched tensors. The first audit's missing namespace prefix and corrected mapping are both preserved. This motivates an independently identified composite candidate, not a memory or speed claim. Authenticate every source payload used, bind the exact replacement map, prove direct-versus-streamed traversal and rerun the six-context screen before native integration. The original VQ pack's quality or parity cannot be inherited. No artifact is activated by this preparation. ### Dense four-bit composite screen, October 3 [[sources/runs/2026/10/2026-10-03-vq-dense-four-bit-overlay-screen]] identifies a separate same-checkpoint composite. It replaces 498 dense tensor triples with the installed affine four-bit arrays while preserving VQ 3.2 experts, PLE, norms and unmatched tensors. Both parent payload sets are freshly authenticated on every run. The complete replacement map and geometry have their own immutable digest; ordinary VQ parity and quality do not transfer. All seven frozen runs complete without retries. The four-layer traversal proof matches exactly across the 512-token boundary with 513 tokens, at a 5,927,277,560-byte process peak. The six complete forwards peak at 2,628,520,312 bytes. These are streamed reference observations, not native resident floors. Header-derived dense payload falls by 2,424,832,000 bytes; no speed or cache benefit has yet been measured. Against the VQ 4.4 proxy, mean case KL is 0.3936587962920319 versus 0.44401073962586735 for the installed baseline. Mean top-1 agreement is 0.8229166666666666 versus 0.7916666666666666. Top-1 improves in three cases and ties in three. KL improves in three and worsens in three, particularly retrieval; continuation dominates the favorable mean. Full VQ 3.2 retains the stronger prior distribution result. Keep this as a low-memory hypothesis for bounded native feasibility, not an established quality improvement or product candidate ready for promotion. Completed-task and held-out gates remain open. The comparator reuses control logits only after binding their manifests to original receipts and pre-candidate hashes. Changed artifact and unfrozen-source refusals pass. Copy admission keeps total raw logits below the frozen 2 GB cap. Four CPU tests, thirty static-harness tests and the full static suite pass on the unchanged allocator binary. No native composite execution, MTP, vision, dynamic memory range, Auto behavior or hardware speed profile is qualified. ### Draft normalization storage resolved, October 3 [[sources/runs/2026/10/2026-10-03-vq-draft-raw-normalization-audit]] authenticates both complete draft files and compares all nine normalization tensors using bounded CPU reads. All raw sidecar values, including those stored as F32, are exactly BF16-representable. None matches the installed folded tensors directly; adding one and rounding to BF16 matches every tensor exactly. This establishes the storage convention. Multiplication dtype, per-stream versus full-width normalization, fused projections and speculative acceptance still need explicit numerical qualification. Existing draft execution remains unchanged. ### Native dense-four-bit composite, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-dense-four-bit-greedy-parity]] preserves the independent reference, source-bound executable, compiler repairs, six refusal cases and all native outputs. Both the streamed and fixed-resident composite match all 2,560 complete logical boundaries across sixteen autoregressive steps. The fixed banks, codebooks, parallel demanded reads and allocator policy are unchanged from the parent profile. Streamed process peak is 1,181,058,704 bytes; resident peak is 5,136,863,552. The resident ledger observes all fifty families and exactly 2,893,477,400 payload bytes with a 317,849,600-byte largest load copy. These are this fixture's observations, not general admission floors. The same binary passes original 3.2, 4.4 and 2.1 greedy and sparse-continuation goldens. Their resident greedy/sparse process peaks respectively are 7,901,403,424 / 9,054,377,744; 8,410,127,008 / 9,534,019,368; and 7,768,873,128 / 8,884,131,600 bytes. All are below the declared 10 GB research envelope. The native T0/T1 catalogue also passes. No original golden or arithmetic tolerance changed. The map is a separate artifact, not a relabeling of full VQ 3.2. Its nine source shards are authenticated as whole pinned files, accounting for 83,770,065,126 bytes read for authentication rather than resident memory. Cross-artifact, partial-option, changed-map, wrong-parent and linked-manifest calls fail before result publication. Full affine acceptance, composite sparse continuation and timing are still pending at this checkpoint; no alternate pack is installed or offered by Auto. ### Dense composite prefill and fixed-cache cost, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-dense-composite-prefill-and-cost]] records new independent composite prefill and sparse references. The native ordinary-prefill check matches 320 boundaries with a 2,866,039,472-byte process peak. Both sparse paths match all 984 boundaries through 2054 consumed tokens; streamed peak is 3,123,775,912 bytes and fixed-resident peak is 6,973,444,416. This is finite-horizon correctness and measured process ownership, not a general minimum or longer-context qualification. After separate full-logit validation, the same executable completes three alternating full-VQ/composite timing pairs. All six requests satisfy the frozen thermal, power and paging eligibility checks, and each artifact reproduces its own full 128-token sequence. Cross-artifact output equality is not required. Median committed rates are 5.520602046506764 and 5.947090433995678 tokens/s; the median paired ratio is 1.0800631461693295. Median TTFT values are 3.1590660829970147 and 2.5441643340163864 seconds, with paired ratio 0.81918281617562. Median request durations are 26.25571970801684 and 23.914144957991084 seconds, with paired ratio 0.9132518003180788. These engineering observations do not qualify the 20-token target or provide a confidence bound. Both arms own the same 1,194,393,600-byte, 608-record bank allocation, 2,082,816 shared-book bytes and 95,558,400 maximum read staging bytes. Full VQ records 49,233 loads and 18,790 hits; the composite's different generated route sequence records 49,865 loads and 18,105 hits. The recovered dense memory was not used for larger banks. Every request's global paging counters stay at 20 swap-ins and 2908 swap-outs. Complete raw observations, including earlier paging, are retained. Authenticated load medians are 29.497434666001936 and 59.31279829199775 seconds. The research overlay additionally authenticates nine full baseline shards, totaling 83,770,065,126 bytes; this is startup I/O, not resident memory. A published composite pack and its load path have not been built. OS page cache is uncontrolled, and these are neither cold-SSD results nor complete-task latency or quality evidence. ### Shared dense dispatch affine acceptance, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-dense-composite-affine-compatibility]] preserves the full static suite and existing affine battery on frozen-dense-overlay-v2. The full battery reports 31 passes and four reference-related failures caused by the temporary driver resolving the venv interpreter symlink. The targeted repair uses the unchanged executable and the intact MLX 0.32.2 virtual environment: both independent reference producers and both native comparisons pass. This closes the existing affine compatibility requirement without changing the reference bytes, arithmetic or tolerances. The behavioral probe passes all 15 items and vision serving passes all 25 checks. These existing-pack results do not qualify composite task quality, features or throughput. ### Dense savings reinvested into fixed expert banks, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-dense-savings-reinvested]] compares the same composite and frozen binary with 512/96 versus 1536/288 record banks. The additional 2,388,787,200 bank bytes fit within the established dense payload saving; the process bound stays 10 GB. Both profiles match every greedy and sparse reference boundary. The larger greedy process peaks at 7,521,441,744 bytes and sparse at 9,328,595,360 bytes. Physical slots 1535 and 287 are actually executed in greedy generation. Sparse occupies only 1370 records and has no eviction, so its proof is kept distinct from the full-range greedy check. All three alternating timing pairs qualify, with exactly the same complete 128-token output across both arms. Control/enlarged arm medians are 6.0164231060/5.6931276441 committed tokens/s, 2.5375323750/2.8505367500 seconds TTFT and 23.6464410420/25.2083016250 seconds total request. Median paired ratios are 0.9462645070, 1.1265224763 and 1.0661911433 respectively. Each control request loads 49,865 records and hits 18,105; each enlarged request loads 34,782 and hits 33,188. Reduced reads did not create a speed win. This is one fixed-work non-speculative context with uncontrolled OS file cache, not a conclusion about cold SSD, larger budgets or the complete production stack. No automatic cache growth or candidate promotion is justified by it. ### Larger-cache CPU diagnosis, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-larger-cache-cpu-diagnosis]] records a separate ten-second CPU sample whose timings are discarded. Demanded reads and MLX evaluation dominate the observed main-thread stacks; the victim search barely appears. This is diagnostic evidence, not an isolated GPU cost model. The next bounded hypothesis is to compare the VQ reader's current buffered mode with the existing affine reader's uncached/random policy, retaining exact bytes, descriptor ownership and the same process envelope. It is not yet implemented or measured. ### Fixed-cache acceptance checkpoint [[sources/runs/2026/10/2026-10-03-native-vq-reinvestment-static-acceptance]] preserves the first static harness failure and its repaired full pass against the unchanged frozen reinvestment binary. The new test script required registration in the mocked suite inventory. No native numerical gate, reference or tolerance was relaxed. All four CI workflows for the preceding dense-composite commit passed. The larger cache remains an explicit research option and a losing cost result on the fixed workload; product Auto and the 20-token target remain open. ### Expert-containing shard read policy, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-expert-shard-read-policy-parity]] records an explicit research comparison of buffered descriptors versus checked uncached random-read hints on the nine authenticated shards containing experts. The same fixed dense composite and 1536/288 banks pass all 2560 greedy and 984 sparse boundaries under each policy, including exact generated tokens and cache state. The larger sparse peak is 9,388,167,512 bytes, within the unchanged ten-GB process bound. Ordinary readers remain buffered, and the policy covers entire shards, including their dense members. Integrity, descriptor ownership, cancellation and range guards are retained. [[sources/runs/2026/10/2026-10-03-native-vq-shard-policy-timing-admission-stopped]] preserves the separate failed timing campaign. Both lean validation runs pass, but their thermal state is fair and their timings are discarded. The first measured process fails its native initial headroom observation before model allocation; external snapshots do not identify the exact cause. There is no paired speed result or automatic retry. A later comparison needs a separately frozen stable-admission protocol without relaxing memory or timing requirements. Full production generation, larger contexts, task-quality, MTP/vision, dynamic allocation, Auto and real-hardware qualification remain open. ### Stable-admission shard-policy result, October 3 [[sources/runs/2026/10/2026-10-03-native-vq-shard-policy-stable-paired-cost]] records the separately frozen follow-up after the failed campaign. Every cell first passes thirty seconds of sampled nominal thermal state, low-power mode off, and the unchanged thirteen-GB floor in native and external VM observations. The bounded waiting step neither retries a failed cell nor relaxes native guards. The observer uses the timed binary's exact memory-observation source. Both full-logit validations and all six timing cells pass, with the same complete 128-token sequence. Buffered/uncached median committed decode is 5.8081/5.3094 tokens/s and request time is 24.9272/26.7625 seconds. The median paired ratios are 0.91614 for decode, 0.95239 for first-token latency and 1.07729 for request duration. Keep buffered reads as the default for this layout. The result does not establish a universal policy, a qualified new pack or the twenty-token target. The next declared storage hypothesis packs unchanged expert bytes into contiguous aligned records; no conversion or native packed-reader result is implied here. ### Lossless contiguous expert export [[sources/runs/2026/10/2026-10-03-vq-contiguous-record-lossless-export]] records one bounded CPU-only export from the fully authenticated VQ 3.2 parent. All 288 reconstructed source-tensor hashes match, including the six codes/scales pieces in every layer. Per-record padding is zero. Each payload begins at byte 16,384; aligned strides are 1,851,392 or 2,621,440 bytes. The 48 files occupy 47,866,183,680 bytes, within the frozen 48-GB output and 350-GB research-staging bounds. Staging before export was 254,809,699,818 bytes. Final observations are 246,923,960 current process bytes, 258,998,968 lifetime physical-peak bytes and 272,646,144 lifetime RSS-peak bytes, all within the one-GB conversion cap. The manifest SHA-256 is `230c53d8bea76e245c8c47863db3fc0f93c549865e7393b5a493c07b0f52b834`. Source/output stamps, independent tensor reconstruction, complete output file hashes and synced atomic manifest publication bind the artifact. Synthetic source-order, padding, corruption, truncation, extent and incomplete-write checks pass. This does not establish native reader parity, a speed gain, quality qualification or deployment readiness. ### Native contiguous record parity [[sources/runs/2026/10/2026-10-03-native-vq-contiguous-record-parity]] binds native producer `0c435ae491263b4bae35aaec29845d0e7f6507170ce411f3c1a811b3c1f52fff` to the frozen aligned export. All 96 catalogue groups pass. Both storage paths match every one of 2560 greedy and 984 sparse-context state/output boundaries without tolerance changes. Sparse coverage reaches 2054 consumed tokens. The contiguous reader verifies all 48 derived payloads totaling 47,866,183,680 bytes. Greedy split/contiguous lifetime peaks are 7,518,705,640 and 7,535,450,112 bytes. Sparse peaks are 9,373,421,864 and 9,383,448,896 bytes. The process ceiling remains 10,000,000,000 bytes. Maximum staging is 95,558,400 for split reads and 115,015,680 bytes for complete aligned reads, both below 128,000,000. The 1824-record greedy cache reaches physical addresses 1535 and 287; both arms report 10,308 loads, 3902 hits and 8484 evictions. Sparse does not claim full physical-slot occupancy. Three CLI refusal cases leave no output directory or model allocation. Buffered read policy, model values and bank capacities stay fixed. These functional runs are not throughput measurements. Authentication reads the parent, dense overlay and then candidate-only derived files; OS file-cache history remains uncontrolled and must qualify any later interpretation. Full original affine acceptance from the earlier composite checkpoint is not reported as rerun by these research-only changes. ### Paired contiguous-record generation cost [[sources/runs/2026/10/2026-10-03-native-vq-contiguous-record-paired-cost]] preserves both independent lean-path full-logit validations and all six interleaved measurements. All six timing cells pass the frozen observed eligibility rules and emit the same complete 128-token sequence. Both arms use the same composite, buffered policy, 1536/288 banks, original prefill sweep and ten-GB process ceiling. No run is retried or replaced. | Metric | Split ranges median | Contiguous records median | Median paired contiguous/split ratio | | --- | ---: | ---: | ---: | | Committed decode tokens/s | 5.776163947978307 | 6.872380896608456 | 1.1897828660167846 | | TTFT seconds | 3.0060759580228478 | 3.007833749987185 | 0.98991835190062 | | Request seconds | 24.993009916011943 | 21.581774708989542 | 0.8576591458049302 | Loading medians are 59.3513325420 and 77.1266143330 seconds. Maximum timed process peaks are 7,750,785,096 and 7,782,275,240 bytes. Every measured request has 34,782 record loads and 33,188 hits in both arms, so the improvement does not come from a larger cache or different generated route sequence. Median paired ratios and ratios of arm medians are different statistics. The complete research path improves generation while increasing authentication/load time. Candidate-only record files are read after both parent filesets, changing uncontrolled OS cache conditioning; this limitation was recorded before timing. The result cannot isolate positional-call count as its cause, claim cold-SSD speed, project standalone-pack startup, qualify another budget/Mac, or meet the twenty-token target. Main-model values are unchanged, but the underlying composite's held-out task qualification remains open. [[sources/runs/2026/10/2026-10-03-native-vq-contiguous-record-static-acceptance]] records complete static acceptance on this same frozen producer. No public loader, Auto choice, installed pack or download activation changes. The earlier complete affine model-loaded battery is preserved as earlier evidence and is not represented as rerun here. ### Separate original draft on real composite inputs [[sources/runs/2026/10/2026-10-03-vq-composite-draft-initial-parity-failure]] binds the CPU inventory, independent reference and preserved failed native comparison. The original head contains 68 tensors and 18 affine modules, with a 1,470,955,171-byte file, 1,470,946,816-byte payload and 419,430,400-byte largest tensor. Its own recipe is four-bit group-64. The complete-file and header hashes plus original-checkpoint conversion provenance are captured separately from the VQ trunk. The reference main prefill matches all 160 frozen boundaries and next token 760. The process releases the main stack before head computation, peaks at 2,424,162,872 bytes, and writes a 2,253,521-byte BF16 fixture from 43 composite-input head entries plus one cached step. The first native orchestration stops before launch because it repeats a lock preflight inside its own reservation. A separately recorded native-only correction executes under a parent-held lock and fails unchanged prefill parity at relative maxima 0.02927 and 0.02970; its cached outputs are exact. This also reproduces the standalone head loader's absent exclusion check without running a second model. [[sources/runs/2026/10/2026-10-03-native-vq-owned-draft-component]] adds owned full-file authentication, separate metadata and explicit retained-payload/load-copy checks. Its research-only arithmetic profile uses the reference's fast grouped normalization, the existing qualified BF16 sigmoid and finite rotary coefficients. All four unchanged fixture outputs then match exactly; offsets are 43 and 44, all outputs are finite, native peak is 2,064,894,112 bytes and bounded BF16 traces occupy 5,811,200 bytes. The public arithmetic remains deployed. This is a combined-profile comparison, not attribution to any single arithmetic operation. Twelve actual CLI refusal checks pass for process-lock contention, ambient overrides and corrupt, symlinked, truncated or FIFO inputs at the relevant fixture/configuration/sidecar boundaries. The process-lock fix applies to ordinary standalone MTP loading as well. This component result provides no speculative acceptance, committed-generation speed, long-context, vision or held-out quality evidence. No candidate is installed, served, selected by Auto or promoted. ### Final draft build admission status, October 3 [[sources/runs/2026/10/2026-10-03-vq-draft-final-build-admission-pending]] records a successful final native build and the broader local acceptance remaining unlaunched. The head and owned-loader source hashes are unchanged from the exact component producer; the final diagnostic adds an existing-output refusal and metadata/budget checks. One admission attempt refuses insufficient real memory. A separate bounded stable-admission campaign also launches no model and is interrupted after a separate llama-server workload is identified. The unrelated process is untouched. The quiet preflight now rejects known llama.cpp inference entrypoints before launch as well as Slotstream/build contention. Twelve context-qualification tests and thirty-two static-entrypoint tests pass; a read-only invocation reproduces the external-model refusal. Final local static/catalogue, baseline draft, streamed draft, image interaction and verification-row checks remain pending until exclusive model execution and their original real-memory bounds are available. Exact component parity is not substituted for those regressions. New independent reference campaigns must bind the changed safety-helper identity and satisfy their source-bound traversal requirements again. The full implementation remains in progress. No alternate pack has earned Auto integration or promotion, no new supported hardware profile reaches the target, and no release is implied by this checkpoint. ### Baseline product controls and completed local regressions [[sources/runs/2026/10/2026-10-03-baseline-auto-selection-and-live-memory]] records the baseline-only product selector, immutable manifest identity, independent quantization/ceiling/live-memory controls, configuration generations, deferred ownership and fixed-capacity pressure behavior. The native catalogue, full static suite, complete Mac runtime suite and all native settings appearances pass. Eight bounded model cells pass across the two preserved campaigns: ordinary/draft fixed-capacity pressure recovery, actual app deferred reload, research/public draft references, streamed draft equivalence, draft/vision interaction and draft row equivalence. Exact process peaks, real headroom, paging and elapsed times remain in the raw receipts; these functional observations are not throughput qualification. Two instrument failures are preserved. The first compiler observer missed independently grouped descendants and was stopped; the replacement enumerates parent links and completes under an explicit compiler-only envelope. The research draft initially rejects an inherited production diagnostic flag before loading; a fresh clean-environment continuation passes without tolerance or fixture changes. Neither failure is presented as a model numerical regression. The original deployment remains the only selectable pack and the registry's new hardware qualification list stays empty. The native VQ research path still lacks production service/speculation and dynamic memory integration. No alternative-quality, multi-pack activation or twenty-token claim follows from these checks. ### Candidate recording, exact verification and interrupted-state recovery [[sources/runs/2026/10/2026-10-03-candidate-state-recording-and-recovery]] preserves three source-bound builds and their bounded sequential campaigns. The first whole-model test fails 422 of 3,580 assertions, isolated to kept prefixes of one or two tokens and their continuations. Full-pass recording, the separate recurrence kernel and checkpoint restoration pass. The corrected explicit row-invariant verification mode passes all 3,580 short-context assertions on full VQ 3.2 and the original-dense composite. The expanded sparse-boundary campaign passes 4,952 assertions on each layout, including actual cancellation and recovery at layers 0, 17, 47 and the final output. Verification starts after 2,048 prompt tokens and stays within the existing finite rotary bound. Sampled process peaks are 9,170,212,240 and 6,384,979,400 bytes respectively. Every run retains the ten-GB process envelope, thirteen-GB real-memory preflight and three-GB headroom checks. Global paging is retained separately; elapsed times are functional diagnostics, not qualified throughput. The final executable also passes the native catalogue and the unchanged full/composite greedy and sparse references. These run in ordinary reference arithmetic, separate from the explicit verification profile. Each greedy check passes 7,717 assertions and each sparse check 2,974. Frozen tensors, tolerances and output IDs remain unchanged. The raw source preserves producer manifests, binary/Metal hashes, failed and successful reports and an archived final compiler input set. No target tokens are re-evaluated to repair a recorded speculative rejection. Draft acceptance, sustained speed, larger contexts, vision, production memory integration and pack qualification remain open. ### Target-verified composite speculation and shared-path regressions [[sources/runs/2026/10/2026-10-03-candidate-target-verified-speculation]] preserves the initial diagnostic build failure, the complete failing generation run and its corrected successor. The first runnable campaign fails 18 of 1,361 assertions, all in the draft key, value and indexer cache. Emitted tokens, target tensors, consumed boundaries and sampler state already match. The failure starts when token/hidden-branch flattening changes the draft fusion projection dispatch; it is not a target acceptance or sampler error. After candidate-only token-wise fusion, all 1,361 assertions pass, exercising 48 accepted drafts across the bounded cases. Greedy and sampled generation, EOS and callback stops preserve exact complete target/head logical state and RNG state. The sixteen-token greedy output equals the unchanged independent composite reference. Candidate process peak is 7,610,947,248 bytes within its ten-GB envelope. Sparse-boundary recovery without a draft passes all 4,952 checks on the same binary. The public overload now delegates embedding lookup into the same alignment/vision logic. Its existing full draft/image diagnostic passes with a 9,820,739,344-byte peak, and streamed-head checks pass with a 7,876,302,872-byte peak, within their unchanged twelve-GB target and fifteen-GB preflight. Every campaign cell preserves real headroom, pressure observation and the single-process rule. Runtime durations are functional observations only. No sustained speed, task-quality, arbitrary-context, candidate-image, production serving or Auto-promotion claim follows from these results. ### Authenticated extended rotary and native contexts [[sources/runs/2026/10/2026-10-03-candidate-extended-rotary-and-context]] preserves the independent coefficient producer, exact build identity, bounded native campaign and all raw results. The 67,108,864-byte F32 component covers 262,144 positions and matches the complete embedded 2,054-position prefix. All 21 native component assertions pass, including complete duplicated sine/cosine hashes, noncontiguous batch indexing, wrong kinds/extents/digests and early/mid-read cancellation. The native coefficient diagnostic peaks at 405,865,432 bytes; the separately supervised Python producer stays within its two-GB bound. Coefficient coverage is not an admitted native context. | Explicit composite context with original draft | Assertions | Failures | Observed process peak, bytes | | --- | --- | --- | --- | | 4,096 | 493 | 0 | 8,437,568,408 | | 8,192 | 509 | 0 | 8,498,680,704 | | 32,768 | 605 | 0 | 9,389,527,888 | Every model case uses 512-token prefill passes, the same small mixed-class banks and exact-verification arithmetic. Tests compare long-context draft discard, reproduced proposals, target logits, recorded target/head rollback and subsequent continuation, then consume the complete configured window and refuse the next token without state mutation. No target prefix is recomputed to accept a speculative result. All expert pins are released, and every case leaves a committed target/head boundary. The native catalogue also passes. These runs retain the ten-GB process ceiling, thirteen-GB real reclaimable preflight, three-GB remaining headroom and one-process rule. Global paging is preserved as diagnostic data. Elapsed functional-run times are not clean throughput evidence. The embedded finite path remains the default, and this experiment does not qualify full model-range execution, image input, held-out tasks, a public candidate Engine, Auto promotion or twenty tokens/s. The same source records successful remote core, Mac, context and docs CI for the earlier state-recovery commit. ### Frozen completed-task calibration [[sources/runs/2026/10/2026-10-03-candidate-completed-task-calibration]] binds a sixteen-case owned protocol covering instruction following, tool execution, tested coding fixes, multilingual structured answers and retrieval. The native producer authenticates the original tokenizer/template and freezes every input token once, avoiding tool-schema key-order differences across processes. Every arm uses that identical prepared protocol. Sampling is greedy, the admitted context is 8192 tokens, output is capped at 512 tokens and the process bound remains 10 GB with a 13 GB preflight and 3 GB actual headroom. | Arm | Completed tasks | Graded outcome | Observed process peak, bytes | | --- | --- | --- | --- | | Original affine four-bit, no draft | 16 of 16 | 15 of 16 pass | 7,994,971,224 | | Full VQ3.2, no draft | 14 of 16 | Ineligible incomplete arm; 11 observed cases pass | 10,020,674,016 at supervisor termination | | Original dense four-bit with VQ3.2 experts/PLE and two original-head drafts | 16 of 16 | 15 of 16 pass | 9,201,686,360 | The full VQ arm exceeds its process envelope during the first retrieval case and is terminated by its supervisor. Its last native receipt has only the earlier, lower peak; it does not contradict the external lifetime-peak observation. The composite runs once afterward as the originally frozen third arm. No preceding case, failed process or prompt is retried or replaced. Both complete arms return correctly sorted objects instead of the requested array of names in the same instruction case. All their tool calls execute successfully against deterministic local fixtures. Every coding fix passes its frozen pure-function tests without input mutation. The grader rejects incomplete answers, malformed or mismatched tool calls and unsupported coding authority; coding workers have bounded source, operations, memory observations and CPU/wall deadlines. The source includes the exact grader and eight passing instrument tests, twenty-one native refusal cases, the unchanged 1361-assertion generation control, the native catalogue and the full static pass. These are disjoint calibration examples, not held-out task estimates. The full VQ partial successes cannot be treated as a passing overall result. Functional request times are preserved but have no clean paired timing eligibility or confidence interval. No alternative pack, hardware speed profile, image support or production serving is qualified. The recorded next hypothesis is to omit unused vocabulary readouts on intermediate prompt passes, with exact state and continuation checks before any newly versioned measurement. ### Intermediate prefill readouts and task-memory repair [[sources/runs/2026/10/2026-10-03-candidate-prefill-readout-memory-repair]] preserves the next source-bound binary and prospective resource budget. The candidate's intermediate 512-token prompt chunks no longer compute full-vocabulary readouts that generation discards. They still consume every target and original-draft position. The final prompt chunk and recorded verification keep the existing full readout; ordinary forward observers and frozen numerical fixtures remain available unchanged. The native control passes all 2071 assertions, including every preceding speculative-generation check and new exact target/head state and continuation-logit comparisons after 17 and 512 tokens. Its observed physical-process peak is 8,488,327,024 bytes. Early cancellation and attempted readout omission during a recorded pass are refused without losing the preceding checkpoint. Both candidate arms are then repeated once under the unchanged sixteen-case protocol and ten-GB limit. Full VQ3.2 completes with a 9,365,920,248-byte peak and thirteen passing tasks. The original-dense composite with two drafts completes with an 8,492,242,872-byte peak and fifteen passing tasks. All fourteen previously completed full-VQ cases and all sixteen prior composite cases retain identical output tokens, finish reasons, parsed answers and speculative details. The original baseline is reused and remains fifteen of sixteen; it is not presented as rerun on this binary. The full VQ arm's prior 10,020,674,016-byte failure is retained. The change resolves that observed workload's memory failure without enlarging its envelope or changing its inputs. It does not qualify every possible 8192-token prompt, larger full-VQ contexts, held-out task quality or the twenty-token target. These sequential functional durations are not paired performance evidence. The native catalogue passes; the preceding complete static pass is scoped to its earlier binary, and the attached previous-commit CI snapshot still has engine and Mac jobs running. ### Parallel prefill records and completed tasks, October 3 [[sources/runs/2026/10/2026-10-03-candidate-parallel-prefill-read-validation]] preserves the complete staging ledger, source-bound build, native controls, task protocol, raw results, exclusions and acceptance. Twelve joined read lanes fill private complete records; immutable codebooks and plans are shared with decode. The reservation charges read results, per-lane scratch, retained MLX arrays and the largest join buffer, capped at 300,000,000 bytes. Every prefill scan preserves decode-bank hits, loads, evictions, occupancy, pins and capacity. Independent 410/512-row component fixtures pass for all allocation classes in VQ 3.2, 4.4 and 2.1. Sparse whole-model comparisons preserve every frozen boundary for those packs and the dense composite. Their observed process peaks are 9,183,156,008, 9,721,304,848, 9,088,718,704 and 7,105,663,392 bytes, respectively. The unchanged generation control passes with an 8,555,501,352-byte peak. No golden or tolerance changes. One prelaunch driver used a wrong fixture path; the next misread a successful JSON component report as a missing text label. Both failures and the successful native output are retained. The continuation runs only previously unexecuted checks. The complete sixteen-case composite calibration preserves all output tokens, parsed answers, finish reasons and speculative details, with an 8,510,592,976-byte process peak. The following three alternating serial/parallel pairs use identical prepared tokens, one fixed warm-up case, a short tool case and two retrieval cases, preserving all outputs and staying inside ten GB. Four timing cells fail the prospectively fixed competing-CPU condition. Only the final pair is eligible, so the complete study emits no clean median or promotion verdict. Raw timings remain available as excluded observations. This is not a sustained speed or twenty-token result. All eleven new actual CLI refusals and the complete static acceptance suite pass. All four remote workflows for both preceding source commits are successful. Parallel prefill remains an explicit research option; the installed model, public loading and Auto registry are unchanged. ### Controlled three-bit experts and independent calibration, October 3 [[sources/runs/2026/10/2026-10-03-affine-three-bit-expert-control]] captures a complete expert-only affine-three-bit transcode of the original four-bit checkpoint. Dense, PLE, draft and vision values remain original. The result contains 48 files and 52,848,290,992 bytes including headers, with 2,150,400 bytes per complete expert record. It still requires its original parent and is not an installable or reduced-download-size claim. The phase prospectively raises total research staging from 350 GB to 365 GB to price these additional outputs; historical protocols remain unchanged. The source-bound quantizer verifies every original source and complete output through owned descriptors. Numerical checks cover both expert shapes, all packed codes, byte-identical batch-eight versus individual conversion, tensor serialization and restored reconstruction. The component's internal physical lifetime peak is 577,864,664 bytes; conversion peaks at 257,082,112 bytes. The separate strict reference verifies every original PLE table against mapped gathers and preserves exact direct/streamed output over four layers and 513 tokens, peaking at 5,559,915,152 bytes. Each subsequent full-model pilot stays below ten GB; the largest observed peak is 2,291,976,352 bytes. All six contexts reuse the frozen original token sequences and corrected VQ4.4 reference. Candidate full-vocabulary logits are scored in memory and hashed, adding zero stored raw-logit bytes. The comparison uses the previously captured native original baseline and this explicitly identified Python control. It is not a causal proof that bit width alone explains every difference. | Context | Original baseline KL | Control KL | Original top-choice agreement | Control top-choice agreement | | --- | --- | --- | --- | --- | | Prose | 0.224397 | 0.284638 | 13/16 | 13/16 | | Coding | 0.366843 | 0.475493 | 12/16 | 14/16 | | Tool result | 0.764051 | 0.918598 | 12/16 | 10/16 | | Multilingual | 0.343711 | 0.379610 | 14/16 | 11/16 | | Continuation | 0.715288 | 0.192943 | 13/16 | 15/16 | | Retrieval | 0.249774 | 0.344421 | 12/16 | 14/16 | Macro KL is 0.444011 for the baseline and 0.432617 for the control; top-choice agreement is 76/96 and 77/96. Five contexts worsen KL, while the continuation improvement reverses the aggregate. This small calibration supports further native measurement but cannot qualify similar task quality, speed or product use. No output was discarded or sampled again to improve the outcome. The initial protocol setup used system Python without MLX metadata and stopped before model launch. The first static attempt then found that the new suites were absent from the static-entry-point fixture registry. Both failures are preserved. The corrected fixture registration and complete static suite pass; no model measurements were rerun for that repair. Public Engine loading, installed artifacts, the supported registry and Auto remain unchanged. ### Native affine reference, generation and completed-task calibration, October 4 [[sources/runs/2026/10/2026-10-04-native-affine-reference-generation-and-context]] captures the authenticated native adaptation of the already pinned affine-three-bit expert control. The expert recipe is three-bit/group-64, while dense, PLE and the separately loaded draft remain original. Pool/staging shapes use verified descriptors. Resident arrays and streamed rows retain the authenticated file owners; the public original loader and planner geometry keep their previous defaults. The research resident loader eagerly materializes the embedding tensor too, so these memory observations do not claim all potential row-cache savings. The six-token independent PR1788 reference and the sixteen-step self-fed reference both retain the predeclared 0.02 affine maximum-relative bound. Final native execution matches every observed hidden and vocabulary byte exactly at 640 cold, 640 reused, 800 grown and 640 shrunk slots; the deployed Generator emits the same sequence. The six-token process peak is 5,983,194,976 bytes and the longer-generation peak is 6,712,610,320 bytes. The first retained-pin diagnostic, deployed-arithmetic mismatch and longer-prefill attention mismatch remain in the source. The sorted expert branch leaves the latter failed logits unchanged; disabling forced fused prefill only in the explicit reference profile resolves it. A build containing a floating row-index division was stopped before model execution and corrected to integer division. The charged reference-storage ledger includes the previously omitted 525,824-byte short rotary table. Its corrected prior total is 1,970,759,168 bytes. The new hidden/readout fixture adds 12,789,760 bytes, and sixteen vocabulary rows add 15,892,480 bytes, yielding 1,999,441,408 charged bytes. The full census also lists separately budgeted original-draft and vision compatibility artifacts outside this reference-logit allowance. Everything remains inside the prospectively bounded research-staging account. Native checks save hashes and receipts rather than new raw logits. The original independent four-bit draft runs through the existing Generator, with explicit row-invariant verification and token-wise fusion. The initial diagnostic trapped because its independent canonical state had no draft cache; the source preserves that failure and faulting frames. After initializing that state as the Generator does, all 2,540 assertions pass, including actual accepted drafts, greedy depths one/two/four, sampled depths two/four, callbacks, EOS, verification cancellation and exact subsequent recovery. The process peak is 8,606,963,448 bytes. The later context-admission correction repeats those checks and adds requests at the finite window and beyond the reserved output limit; its final report contains 2550 passing assertions with a 8,473,138,480-byte peak. | Admitted native window | Passing assertions | Physical process peak, bytes | | --- | --- | --- | | 4,096 | 1,330 | 8,929,629,248 | | 8,192 | 1,346 | 8,878,134,312 | | 32,768 | 1,442 | 9,598,293,104 | Each context keeps the original draft live, consumes the entire window, checks every recorded-prefix target/head state and continuation exactly, then refuses an extra token without mutation. These deterministic contexts establish bounded execution and recovery, not long-context task quality or throughput. The final Generator also checks its actual admitted model window before state reservation and bounds provisional verification inside it; the baseline retains its full model limit. The same sixteen prepared calibration tasks and unchanged grader produce fifteen passes for the freshly repeated original baseline, draft-free affine control and two-draft affine control. All fail the same instruction requiring a names-only array, returning sorted records instead. All tool fixtures and pure-function coding tests pass. The repeated baseline has exactly its prior token streams and parsed answers; the two affine configurations also have identical streams across all sixteen cases. Their observed process peaks are 8,117,130,496, 6,765,843,064 and 9,120,127,000 bytes, respectively. Different cache capacities and arithmetic identities are explicit. The raw functional durations are not a clean paired timing comparison, and this small calibration cannot establish noninferiority. Existing historical/current layer parity, the cache-sweep gate and weights-free catalogue pass on the recorded frozen binary. The final static suite passes on its later source-bound binary. Builds retain exact before/after input manifests and original failed compiler diagnostics. No installed artifact, supported-pack registry, user preference, quality qualification or twenty-token speed profile is promoted. Candidate production/vision ownership, allocation/governance, transactional distribution, held-out quality and paired complete-configuration performance remain required. ### Pack-specific planning, Engine serving and observed-memory recovery, October 4 [[sources/runs/2026/10/2026-10-04-affine-engine-memory-and-governor]] preserves the new candidate integration and all predecessor failures. The immutable resource contract charges the admitted expert record, eager embedding and authenticated rotary table independently. Its workspace bound includes a complete reference layer, bounded assembly/staging and replacement pool backing during prefill admission. Capacity solving charges that extra scratch along with each additional slot; it cannot reinvest every byte saved by expert quantization while omitting workspace growth. Original planning keeps its existing costs and speed evidence; candidate estimates serialize as unknown and keep the wall-clock prefill guard active. The governor's live restart credit is capped by actual observed process ownership. The existing pure fixture seam retains its historical behavior when no observation is supplied, but every live path supplies one. A failed physical observation cannot create credit. Earlier full/small live drills shrink, honor cooldown and regrow with identical output. Subsequent actual pressure-with-draft, runtime-budget/first-image recovery and output-serving checks pass. Their predecessor failures are kept: bounded credit exposed a test recovery point with no allocator-release margin and a real first-image planner that reused stale zero availability after recovery. The fixture now reserves a small observed-credit margin, and the product planner takes a fresh bounded ownership snapshot. The final candidate Engine fixtures use fixed small arenas and a ten-GB physical watchdog. Their deliberately larger conservative planning envelopes are recorded separately; this is not evidence that a user can run those declared plans at a ten-GB ceiling. Complete-prompt retention is explicitly priced for both active and retained stepped capacity. A smaller fixture correctly refused retention, and that failed expectation remains evidence. All numerical expectations stay unchanged. | Final candidate mode | Passing assertions | Physical process peak, bytes | | --- | --- | --- | | Without draft | 57 | 6,873,813,888 | | Independent original draft | 57 | 8,375,162,784 | These checks execute the actual Engine and loopback HTTP handler. They verify owned Unicode tokenizer parsing after removal of the fixture's source copies, exact greedy output, full-prompt and aligned prefix reuse, disk identity/restoration, cancellation with no silent replay, pressure floor refusal, recovery, warm resize, and HTTP artifact identity. The original-loader/candidate-plan mismatch, oversized forward, unsupported image toggle, legacy draft loader and explicit different-quantization request all refuse before the incompatible operation. The final catalogue, historical/current original layer checks, sweep equality and static suite pass on the preserved source-bound binaries. Prior committed-checkpoint CI is also captured as completed successfully. This adapter is experimental and package-only. It still depends on the original parent plus controlled expert overlay and finite rotary artifact. It does not enable candidate images or streamed drafts, publish a standalone pack, qualify task noninferiority or establish a speed profile. The task evaluator's original-baseline branch now admits explicit drafting so the next comparison need not disable an existing baseline optimization. Those new comparative results are not part of this source. ### Complete draft pilot stopped by frozen timing conditions, October 4 [[sources/runs/2026/10/2026-10-04-complete-pair-exclusions-and-context-ci]] contains all raw task outputs, hashes, timing exclusions and host observations, plus the failed context CI and its local repair. Four completed runs used the original Engine with streamed original draft experts or the authenticated affine-three-bit control with a resident independent original head. Both requested two drafts. The equal physical watchdog did not make their allocation recipes or saved planning ceilings equivalent; this is a calibration comparison, not integrated Auto qualification. Each completed run passed 15 of the 16 frozen calibration tasks and retained the same sorting-format failure. Repeated runs of each artifact have identical prompt/output token IDs and terminal outcomes. This is repeatability on known tasks, not held-out noninferiority. The first original run met its timing conditions. Both candidate runs exceeded the frozen competing-CPU criterion. The second original run has a thermal/power exclusion, and the next preflight refused launch. No replacement rounds were run. Preserve the measurements as excluded timing evidence; there is no paired speed aggregate or evidence for the twenty-token target. All completed cases remained under the existing physical-process watchdog. Separately, context-proxies CI failed to compile because its standalone Swift source set omitted PackMemoryProfile. Adding the actual file to compilation and provenance fixes the observed error, and the isolated source proxy passes locally without an Engine or model. This correction does not widen a memory bound, edit a numeric golden or relax an experimental protocol. ### Durable original-pack activation and recovery, October 4 [[sources/runs/2026/10/2026-10-04-durable-model-activation-and-recovery]] preserves source-bound development builds, all scripted attempts, the full Mac suite, bounded resource protocols and actual model logs. The journal has a bounded state file, exclusive owner lease and synced atomic transitions. It stores fixed diagnostic codes and pack/settings receipts, never prompts, arbitrary error text or model bytes. Existing independent embedding APIs can leave the journal disabled; the Mac app enables it explicitly. The initial real campaign passes actual activation, partial-load failure, sequential rollback, explicit retry and restart at a sampled peak of 5,774,824,080 bytes. The unchanged legacy memory-control workflow passes at 6,621,500,336 bytes. Both use separate sequential processes under the existing ten-GB physical ceiling. A subsequent review fixed initial journal-open/write failures so they stop queue admission, and a cancellation check now follows durable commit but precedes publication. The new async test initially failed to compile; that diagnostic is preserved rather than presented as a product failure. The corrected full Mac suite passes, with an observed process-tree peak of 1,664,470,496 bytes. It includes all existing scripted suites and UI checks, plus an isolated settings executable launched without an adjacent Metal library from an empty directory. The final real activation campaign passes seven sequential load attempts in one process, including cancellation during commit and a later successful request; peak is 5,775,741,584 bytes. These are functional/resource observations, not clean speed measurements. The prior real legacy run is not relabeled as a run on the final binary. Failed requested settings remain visible. The previous runtime can be restored under its own applied ceiling, while saved requested preferences remain unchanged and other queued work waits. A retry revalidates files and current headroom; unreadable journal bytes are preserved. The scripted Light, Dark and System failure screens pass, and the final Light screen was visually inspected. Optional persistent caches attach only after health and commit. A cancelled activation does not load another model for rollback. The static monitor's inherited engine-binary field is not the Mac executable identity; the frozen build records bind the actual check binaries and sources. No installed executable, user Home or quantization default changed. No alternate pack, held-out noninferiority verdict, sustained speed profile or public release is qualified by these tests. ### Independent streamed draft and 32K recovery, October 4 [[sources/runs/2026/10/2026-10-04-authenticated-streamed-original-draft]] records the original four-bit head's separately authenticated streamed placement. Its non-expert tensors remain resident, and its bounded cache and scratch use the original head's record size. Full-file authentication, tensor range validation and descriptor ownership precede allocation. Parallel reads publish only after the entire batch and integrity checks succeed. The target's three-bit descriptor never configures the draft. The frozen streamed candidate matches all 22 previous resident-head observations exactly, including emitted tokens, consumed length and committed target/head state digests. The speculation report passes 2,689 assertions, including joined-read failure without partial cache publication, retry and mutation detection on an owned disposable sidecar copy. The real Engine report passes 58 checks covering memory/disk reuse, live governance, cancellation and HTTP. Their process peaks are 7,546,573,384 and 7,678,072,112 bytes. The original public streamed/resident draft and plain-lookahead regression passes at 7,875,680,232 bytes under its unchanged separate resource allowance. The candidate's 32,768-token context and exact recovery report passes 1,443 assertions at 8,097,697,704 bytes, inside the ordinary ten-GB process bound. These fixture peaks are neither the complete planner envelope nor speed evidence. All final static gates pass, with matching before/after source inputs. The final executable SHA-256 is `fc6faecc34177cada3820c23117e48ba20153ff41d377c6c73b4ebd88216491f`. An intermediate diagnostic compile failure remains preserved. Complete CI for `de75f797eec33b0429c92c9081146ad865ee6890` passes. Candidate planning now supports explicit streamed placement without inheriting the original pack's measured automatic placement threshold. No alternative becomes supported, installed, Auto-selected or quality/speed qualified. Candidate vision, held-out quality, complete-configuration performance and standalone distribution remain required. ### Equal-ceiling actual Engine calibration with timing exclusions, October 4 [[sources/runs/2026/10/2026-10-04-actual-planner-calibration-with-timing-exclusions]] preserves both build attempts, bounded input refusals and Engine regression, frozen task protocol, all six runs, original grader and complete host observations. Both arms request a 14 GB saved ceiling and physical watchdog, with a 17 GB real-memory preflight and 3 GB minimum headroom. They use two drafts and explicit original streamed experts, with vision, prefix retention and decode lookahead disabled. This larger allocation is the measurement itself; ordinary correctness fixtures remain under their prior bounds. The actual original planner chooses 2,162 slots and a 1,024-token prefill chunk; the conservative affine plan chooses 1,175 slots and a 512-token chunk. The smaller expert records alone therefore do not imply a larger cache: full replacement backing is still charged during admission. Every run passes 15 of the 16 unchanged calibration tasks, failing the same names-only sorting instruction. Repeated runs preserve prompt/output token IDs and terminal outcomes. Original physical peaks are 12,593,258,512, 12,537,176,056 and 12,528,508,824 bytes. Candidate peaks are 8,027,034,952, 8,034,866,504 and 8,023,676,208 bytes. These physical observations do not authorize lowering the planner's reserve without an ownership change. The first original run has CPU and process/request-paging exclusions; the first candidate run has a CPU exclusion. The later four runs meet their recorded timing conditions. The protocol requires all six eligible runs for its comparison, so no timing aggregate is published and no replacement runs are taken. The source is discarded for timing while retaining completed-task and memory evidence. This known calibration is not held-out noninferiority, sustained-rate certification, Auto qualification or a public release. ### Sequential expert replacement lifetime and unchanged state, October 4 [[sources/runs/2026/10/2026-10-04-sequential-affine-cache-allocation]] preserves the source-bound build, fourteen input refusals/accepted-header checks, complete native catalogue, real-model fixtures and final static suite. The copy contract is explicit and independent of the original profile. Each destination tensor finishes before the next replacement is issued. CLOCK decisions, staging bytes, canonical expert IDs, matrix geometry and reader lifetimes remain unchanged. For this authenticated record, the largest packed piece is 614,400 bytes per expert. All three packed pieces and all six BF16 scale/bias pieces remain resident where required. The conservative floor workspace at 640 slots and 512 query rows remains 2,517,897,216 bytes; the distinct sequential profile reserves 1,566,031,872 bytes. These are derived allocation bounds, not measured process savings. Larger caches still pay for their own largest destination replacement and gathered admissions. Invalid replacement bounds and mismatched resource identities refuse. | Check | Passing assertions | Physical process peak, bytes | | --- | --- | --- | | Speculation | 2,692 | 7,035,982,456 | | Engine | 78 | 7,360,418,904 | | Context | 1,444 | 8,083,443,528 | All 22 frozen prior resident-head observations match exactly for emitted IDs, consumed boundaries and committed target/head state digests. The additional speculation assertions prove the sequential path ran and its actual intrinsic reservation agrees with the profile. The Engine checks also compare complete tensor bytes and CLOCK keys/reference bits/hand across ordinary and sequential admissions, repeated resident hits and warm resize. The staged context consumes the full admitted window and preserves exact rollback and refusal behavior. Each model process retains the ordinary ten-GB watchdog and three-GB real headroom; the Engine integration fixture's separate conservative planning envelope is recorded in its receipt. The final static suite passes with identical before/after source manifests. Its sampled process-tree peak is 1,087,801,456 bytes. The frozen executable SHA-256 is `f8db9414ca20dc2f7bd706b0f9f63763396f514d414ae9c21f1605d0be6aa363`. All CI for the earlier streamed-draft commit `3608a4e10a94f240c89f4d59df5a224195a55f51` is successful. These checks do not qualify speed, held-out quality, vision, standalone distribution or another supported/automatic pack. A later performance experiment is a separate changed-candidate campaign. ### Sequential-copy actual-plan calibration with one timing exclusion, October 4 [[sources/runs/2026/10/2026-10-04-sequential-affine-calibration-with-timing-exclusion]] retains the complete prospective protocol, unchanged frozen tasks, source-bound executable, six run receipts, exact grader and host observations. The explicit allocation change is a new hypothesis; this campaign is not replacement timing for the earlier conservative profile. Both arms use the same 14 GB saved/physical ceiling, two streamed drafts and 8K requested context, with vision, prefixes and decode lookahead off. The preflight and headroom are unchanged at 17 GB and 3 GB. The actual original plan uses 2,162 slots and a 1,024-token prefill chunk. The sequential candidate uses 1,817 slots and a 512-token chunk, while retaining its separately charged replacement workspace. Each of the six runs passes 15 of 16 tasks with the same names-only sorting failure. Within-artifact outputs are exactly repeatable and match all prior allocation outcomes. Original process peaks are 12,529,786,776, 12,615,819,112 and 12,551,331,664 bytes; candidate peaks are 9,346,701,528, 9,338,886,432 and 9,361,578,272 bytes. Only the first candidate arm is excluded by sampled competing CPU load. The remaining five satisfy the recorded conditions, but the frozen all-six rule prohibits an aggregate. No replacement round or selected-round comparison is used. The source is discarded for timing while preserving functional and memory acceptance. Calibration, smaller observed memory and an unchanged task score do not establish held-out quality, sustained 20-token generation or Auto qualification. ### Bounded grouped affine expert checkpoint, October 4 The explicit grouped operator now computes the same affine target with a fixed thirty-two-expert RHS and bounded route tiles. It preserves the full-domain reference dispatch family, restores original router order and copies hot records into independent allocations before admitting them once in the existing global hot order. It changes neither original arithmetic nor default selection. Its resource profile deliberately retains the earlier conservative sequential-copy reservation pending separate allocation qualification. The frozen integration binary passes 264 exact component comparisons, sixteen admission trajectories, 2,693 speculative assertions, the prior twenty-three emitted-token/committed-state observations, Engine/HTTP and memory/disk/cold continuation and recovery comparisons, and all sixty-four staged context observations through 32,768 tokens. Component, speculation, Engine and context process peaks remain inside the prospective ten-GB ceilings. The complete source-bound runs, two corrected compile failures and prior-main CI identities are captured in [[sources/runs/2026/10/2026-10-04-bounded-grouped-affine-experts]]. This is correctness and bounded-memory evidence, not speed or held-out quality qualification. Next derive the grouped allocation's phase-by-phase bound, validate it independently and compare complete Engine plans under a new prospective protocol. Default Auto, supported packs and installed artifacts remain unchanged. Vision, held-out noninferiority, standalone distribution and full product qualification remain required. ### Explicit damaged-setup repair, October 4 A corrupt activation record now has an explicit recovery action. Under the existing owner lease, repair preserves an owned regular single-link record at an exclusive archive name, keeps the saved preferences and requires a new complete verification and health check. It refuses valid, replaced, symlinked or shared records and cannot run against an active or loaded owner. Ordinary retry still preserves the record in place. Queue resumption uses normal request admission and does not replay failed work or completed tools. The complete Mac suite passes the scripted failure/recovery cases and native Light, Dark and System screens. The frozen real-model sequence also passes corrupt-record refusal, explicit archive, reauthentication, new healthy generation and a completed response while retaining the exact damaged bytes. Its process peak is 5,776,773,872 bytes under the prospective ten-GB watchdog; the development suite peaks at 1,843,318,624 bytes under its six-GB tree limit. Source identities, raw outputs and screenshot digests are in [[sources/runs/2026/10/2026-10-04-explicit-model-setup-repair]]. This closes damaged-history recovery for the original supported pack, without qualifying an alternate pack or a public release. ### Grouped allocation phase bound, October 4 [[sources/runs/2026/10/2026-10-04-phase-bounded-affine-memory]] captures the independent grouped allocation hypothesis and corrected functional qualification. At the declared maximum 512-token pass the workspace formula charges 623,597,568 bytes at 640 slots, 874,887,168 at 1,000 and 1,579,603,968 at 2,000. These are conservative policy bounds, not measured allocation peaks. They include the complete largest destination-piece copy and hot-record coexistence. The same formula can exceed the previous reservation at full model residency; a test incorrectly asserting universal reduction was caught before model work and repaired. The corrected catalogue, nineteen input/protocol cases, speculation, Engine and staged 32K checks all pass. Catalogue, speculation, Engine and context peak at 1,439,450,384, 6,217,847,392, 6,617,650,168 and 7,833,505,584 physical bytes respectively. All twenty-three prior speculative output/state observations remain exact, the Engine retains the prior continuation/recovery observations, and context completes all sixty-four checkpoints. The unchanged 264-component/sixteen-admission gate remains the earlier run, not new evidence here. No performance or held-out quality inference follows. Original accounting and installed artifacts are unchanged. ### Grouped static acceptance checkpoint, October 4 The complete static suite passes against the same frozen grouped-memory binary, with matching before/after build inputs. This includes existing transport, installer, planner and public memory-override contracts. The guarded tree peaks at 1,089,783,992 physical bytes inside its six-GB ceiling, without loading a model. [[sources/runs/2026/10/2026-10-04-grouped-affine-static-acceptance]] preserves the complete raw gate and resource receipt. The new actual-plan timing pilot and remaining quality, vision, standalone distribution and integrated release gates are separate work. ### Grouped complete-plan calibration, October 4 [[sources/runs/2026/10/2026-10-04-grouped-affine-calibration-with-timing-exclusion]] records all six completed actual-plan runs. Both artifacts retain every prior same-artifact output and score fifteen of sixteen calibration tasks. At the same fourteen-GB ceiling the grouped candidate receives 2,116 slots, versus the preceding candidate's 1,817; the original receives 2,162 records of a different byte size. Candidate process peaks remain between 9,087,883,984 and 9,089,080,064 bytes, while original peaks range from 12,551,036,800 to 12,590,948,296. These are physical observations of the declared complete configurations, not permission to remove ledger reservations. The first five timing cells are eligible. Background indexing and media/photo analysis disqualify the final candidate cell under the frozen competing-CPU rule. Preserve the campaign without a comparative aggregate or replacement rounds. No local speed or twenty-token qualification follows. Continue independent owned-vision, held-out quality and standalone artifact work; leave background services and the supported original pack unchanged. ### Owned candidate vision and prospective outcome analysis, October 4 [[sources/runs/2026/10/2026-10-04-owned-affine-vision-integration]] records an explicit image-capable grouped affine path. It captures the exact original configuration and preprocessing values, retains authenticated tower tensor owners, and propagates cancellation through bounded tensor reads. The separate image resource identity retains the tested text profile and its arithmetic, with vision residency and workspace charged independently. Metadata geometry does not reopen the original paths after capture. The repaired frozen binary passes the native catalogue, all forty-nine image integration assertions and the unchanged original MTP/image diagnostic. Owned image features match the original component exactly. Two real tiny images exercise main-only versus drafted outputs, actual prefix reuse, HTTP equivalence, malformed-input refusal and text recovery, within the prospective physical ceilings. The first compiler failure is preserved; its missing diagnostic initializer argument was repaired before any model ran. This gate proves ownership and functional image integration, not image-answer quality or all advertised image capacities. No candidate enters Auto or distribution. [[sources/runs/2026/10/2026-10-04-paired-outcome-score-instrument]] records a tested prospective paired binary-outcome analysis helper. Its constrained-discordance score bounds follow Tango's method, with simultaneous family bounds for a fixed-stratum design. Published witnesses and independent likelihood optimization agree. It does not pool heterogeneous strata, does not claim exact finite-sample coverage and never emits model qualification. Corpus provenance, independent task units, complete grading, sample counts, margins, family weights, multiplicity, safety and the actual held-out run remain separate gates. No final task has been evaluated and no noninferiority result is claimed. The full static suite subsequently passes against that same frozen owned-vision binary with unchanged before/after native source inputs. [[sources/runs/2026/10/2026-10-04-owned-vision-static-acceptance]] retains the raw suite, planner/override/transport/installer checks and resource observations. This is implementation acceptance with no model loaded; it does not close the remaining held-out quality, image-capacity, standalone delivery, measured profile or promotion gates. ### Held-out grader and source preparation, October 4 [[sources/runs/2026/10/2026-10-04-heldout-grader-and-source-preparation]] records a Mac sandbox wrapper for the existing bounded coding executor. It preserves tuple/set fixtures, verifies copied source bytes, denies unrelated filesystem access, writes, network and forks, and treats failed worker startup as evaluation failure rather than model failure. The initial dyld/framework-launcher and test-registration failures are retained. Corrected helper tests pass on both local Python runtimes; the static-suite registration checks pass. These are changed-path checks after the separately recorded complete vision/static suite. The pinned MBPP test population yields 249 verified restricted repair fixtures after prospective grammar, literal, reference-success and input-preservation checks. Every original passes and a deterministic broken variant fails; every excluded source case remains recorded. The draft and corrected actual-source graders agree on all 320 syntax-eligible cases. This defines a restricted coding population before model answers; it does not establish arbitrary-program quality or a final sample. Exact public IFEval, MBPP, MGSM, BFCL and MMLU source versions and licenses are recorded. An isolated grader runtime is verified after a transient installer-observer failure, preserving the original failure. Correctly discovered, unchanged upstream IFEval tests all pass with fixed seeds and pinned tokenizer data. No final model answer, quality score or statistical qualification is produced. Multi-step executed tool evaluation, final independent task units and samples, margins/weights, runtime budgets and the held-out comparison remain open. Model, pack, public support and installed artifacts are unchanged by this instrumentation. ### Native conversation and offline tool-grader acceptance [[sources/runs/2026/10/2026-10-04-executed-tool-conversation-instruments]] binds final native binary `e08f824fb793477dce1ed60263950efb3dec791844601ce04c04a00b216c3d88` to its build sources and unchanged resource protocol. The model-free framing catalogue and protocol negatives pass. Both original and grouped candidate execute three real HTTP requests and two resets, including cached tool-result continuation and exact cold recovery. Original peak physical footprint is 7,587,501,360 bytes; candidate is 6,011,752,888 bytes, inside the prospective ten-GB watchdog. These tiny instrument fixtures do not establish comparative timing or task quality. Original/candidate actual plans retain their distinct byte ledgers and slot counts. The pinned BFCL base population has two hundred cases. Every unchanged gold trace passes the retained upstream state/response checker through the safe direct-call adapter, every empty trace fails, and a deliberately wrong filesystem fixture state fails. No reference call returns an execution or fixture error. The final adapter's largest observed worker footprint is 40,255,992 bytes under its prospective 256,000,000-byte ceiling. Six helper test groups pass on both Python runtimes, including real native denials and source-copy isolation; thirty-two static registration checks pass. These contain no candidate outputs and do not increase the number of independently evaluated model tasks. The original prototype and final guard/observation revision are both preserved. The same frozen native binary subsequently passes the complete model-free catalogue and static suite, with matching before/after build inputs. Physical tree peaks are 223,249,296 bytes for the catalogue and 1,115,949,264 bytes for static acceptance, inside the six-GB watchdog. [[sources/runs/2026/10/2026-10-04-executed-tool-static-acceptance]] preserves the full outputs and resource receipt. No model quality or timing inference follows. ### Disjoint quality protocol pilot [[sources/runs/2026/10/2026-10-04-disjoint-completed-task-quality-pilot]] preserves the fixed tasks, protocols, source/runtime identities, complete native transcripts, exact offline tool traces and both driver attempts. The first attempt fails before any model launch on a helper-directory hash. The corrected attempt finishes ten sessions in 2,729.316262 seconds. Original outcomes are facts 4/5, multilingual 5/5, coding repair 3/5, strict instruction following 4/5 and tools 1/5. The grouped affine control records 4/5, 4/5, 3/5, 3/5 and 1/5 respectively. There are three discordant pairs against the candidate and one in its favor, for totals of 17/25 and 15/25. All responses come from actual fourteen-GB, 32K native plans with streamed original drafts and prefix retention. The process watchdog, actual preflight/headroom, request/session/campaign limits and before/after grader identities pass. These are sizing and instrument observations, not eligible speed measurements or held-out noninferiority. The smaller pilot reply/step limits, output truncation, step exhaustion and strict BFCL reference-response requirement remain explicit. No failed case is replaced. All underlying task IDs remain excluded from every final sample. ### Stored affine-three-bit reconstruction refit [[sources/runs/2026/10/2026-10-04-affine-three-bit-refit-component]] records a prospectively bounded component screen of least-squares scale/bias fitting against the original four-bit parent. It uses the same three-bit/group-64 representation and actual BF16 dequantization error, with the unchanged control retained per group unless a trial improves it. The implementation is distinct from HQQ's robust proximal objective and uses no held-out activations or task answers. Twenty component cases pass, including two synthetic shape/batching cases and eighteen real expert projections containing 29,491,200 values. Independent packed-code decoding, stored serialization and group-error reductions agree. The real component squared-error reductions range from 49.1930% to 49.4405%, with zero group-level regressions; maximum individual absolute error grows in twelve of the eighteen. The process peaks at 434,455,344 physical bytes and finishes in 14.261859 seconds inside its frozen four-GB/13-GB-preflight/three-GB-headroom and twenty-minute envelope. No complete model runs or new model weights are produced. This supports pricing a separate full conversion, without establishing quality, performance, original BF16 equivalence or promotion eligibility. ### Complete affine refit and outcome-instrument acceptance [[sources/runs/2026/10/2026-10-04-affine-refit-full-screen]] records forty-eight refitted expert files totaling 52,848,290,992 bytes, produced in 3,355.410512125003 seconds with 966,214,376 peak physical bytes. Conversion stays within its separately frozen four-GB process and 430-GB staging reservations. Stored group squared error decreases by 49.28334082871262 percent. The unchanged six-context model screen reverses that local weight-error result. Refit's macro KL is 0.5175330957912493 and top-one agreement is 0.7708333333333334. Minmax remains at 0.4326172687996428 and 0.8020833333333334; the original is 0.44401073962586735 and 0.7916666666666666. Matching reference/baseline hashes and unchanged baseline scores are checked explicitly. [[records/decisions/hold-unconstrained-affine-refit]] therefore holds this recipe. The screen adds no raw logits and supports no held-out quality, speed or BF16-equivalence claim. [[sources/runs/2026/10/2026-10-04-complete-outcome-and-session-acceptance]] preserves the two hundred offline BFCL reference conversations through the completed outcome wrapper, full long-answer grader fixtures and seven prospective MOVER numerical diagnostics. These are zero-model instrument checks. The paired analysis remains asymptotic and uses independent task units, not tokens; the final method and sample must be selected before final answers. The same source then compiles and passes native V2 sessions for the original and admitted minmax artifact. Each refuses the explicit oversized reservation, performs actual calculator use and consumes its result, reuses a prefix and reproduces the same answer cold. Physical peaks are 7,540,446,440 and 5,952,065,976 bytes, inside the ten-GB envelope. All negative protocols, the complete native catalogue and full static suite pass with before/after build identity equality. These checks qualify instrumentation and identity ownership only. Held-out task outcomes, image quality, complete performance, standalone delivery and promotion remain open. ### Serial campaign acceptance and prospective final selection, October 4 [[sources/runs/2026/10/2026-10-04-serial-outcome-campaign-acceptance]] retains the exact code, failed and corrected unit checks, real instrument sessions, recorded-response replay, full static receipts and successful prior-main CI. On the previously excluded pilot tool case, the original completes with seventeen executed calls and thirteen model responses; the candidate completes with sixteen calls and nineteen responses. Both final outcomes pass. The paired instrument finishes in 384.7173854589928 seconds. This larger prospective request envelope does not replace the earlier smaller-envelope pilot outcome. After the input-envelope correction, a zero-model replay verifies all request histories and preserves those executed calls and outcomes. It also checks the historical native V2 success and context-refusal frames. Replay completes in 3.5612789160222746 seconds at a parent peak of 23,527,904 physical bytes. The complete subsequent static suite succeeds with a maximum sampled process-tree footprint of 1,076,316,296 bytes under its six-GB ceiling. These are instrument and process-bound observations, not speed or quality qualification. [[sources/runs/2026/10/2026-10-04-prospective-heldout-outcome-protocol]] captures the final protocol and exact task data before generation. There are 2,219 independent exact prompt groups across five families and 138 paired jobs. The coding deduplication removes one additional eligible duplicate, leaving 243 repairs. MMLU has 105 exact duplicate pairs before pilot-group exclusion and fixed hash selection. Multilingual translations share one underlying unit. Public-source training overlap and semantic dependencies beyond exact grouping are not ruled out. [[records/decisions/final-paired-task-quality-protocol]] states the prospective method, margins and inconclusive-result policy. No final model score, twenty-token result or promoted alternate exists at this checkpoint. ### Three-bit transport fallback checks, October 4 [[sources/runs/2026/10/2026-10-04-three-bit-lossless-transport-planning]] preserves the previous builder, its complete original plan, the changed source, production codec identity and both raw test logs. All four synthetic test groups pass on both Python runtimes, with byte-exact codec reconstruction and whole-file hashes. The original twenty-five-file, 4,155-object plan retains digest b2119ac3fb9a3a0eb534d87075ceb9b893fc3bcaad5aeee3622337d88be9dac0. Header-only planning covers the control's forty-eight files and 52,848,290,992 bytes in 1,776 independent objects; no model or full payload read occurs. No compression-size, complete transport, public-pull, performance or model-quality claim follows. ### Outcome cleanup failure and correction, October 4 Main CI for e1777ff10373d86d4fc49f2ed6c2f6c1cf899b07 fails in the campaign's deadline test with a process-group PermissionError and unclosed stream warnings, despite the earlier completed local static suites. [[sources/runs/2026/10/2026-10-04-outcome-watchdog-cleanup-ownership]] retains that log. A deterministic barrier fixture using the frozen runner observes two termination owners, then drains its own tiny child once. The corrected runner's ten unit groups pass under both local Python runtimes with ResourceWarning elevated to an error. This is child-process cleanup evidence with no model loaded, not a repeated task evaluation, final quality result or speed qualification. Full corrected CI and all remaining product gates stay open. ### Complete transport CI and optimized authentication fixture lifetime [[sources/runs/2026/10/2026-10-04-transport-ci-and-authentication-fixture-lifetime]] preserves exact-commit CI logs and statuses. The lossless fallback commit passes the complete main pipeline. The authentication follow-up passes the instrumented native checks, public-library import and Mac runtime/Xcode checks, while the optimized catalogue passes ninety-seven groups and fails the serial-fixture descriptor count. The instrument was observing borrowed descriptors after its last direct access to the owning array. The correction explicitly extends those owners through the cancellation, failure and mutation scans; no production reader change follows from this fixture repair. Corrected optimized CI and actual-weight loading measurements remain required. The four-lane authentication helper bounds concurrent CPU readers, verifies complete pinned files, retains original order and joins workers before publication or failure. Its component evidence does not determine an optimal lane count or a startup speedup. The running final quality campaign retains its previously frozen binary and helper identities. No alternate artifact, memory credit, Auto qualification or release is inferred from these CI checks. ### Retained tensor export and complete header geometry [[sources/runs/2026/10/2026-10-04-retained-tensor-subset-primitive]] records eleven synthetic copy/failure groups on each local Python runtime and thirty-two static entry-point checks. Real bounded chunk tails and independently re-read tensor digests agree; corruption, mutation, cancellation, disk-full failures and conflicting output paths are refused. These checks load no model and copy no real weights. The complete header plan retains 2,783 tensors with 35,822,021,112 payload bytes, excludes the 432 original expert tensors with 67,947,724,800 bytes, and prices canonical retained-file headers at a total of 35,822,408,896 bytes. Its digest is b6db15ccdc084438771ec9af99567d3020f8e50fb5647c518d54eff8d1d62ae6. The known-file subtotal of 90,231,754,946 bytes includes converted experts, original companions and the finite rotary artifact, but is deliberately not a complete deployment size or staging authorization. Final configuration/index/provenance and completion metadata remain to be fixed. Source payloads were not authenticated during this header-only run; eventual export must authenticate them through the copier and complete pack owner. The final publication review in the same source adds file-sync, directory-sync, link and competing-destination fault injection. All twelve groups pass on both Python runtimes. A late sync failure may retain a verified individual tensor file, but returns no successful receipt and cannot authorize a complete pack manifest. Existing competing output is preserved. The earlier complete header plan is unchanged and retains its original instrument identity. ### Instruction-worker ABI failure and prospective recovery, October 5 [[sources/runs/2026/10/2026-10-05-heldout-instruction-worker-recovery]] preserves the stopped coordinator and source audit. Forty-six paired jobs are complete. The first instruction job contains one complete original-model answer, no grade and a cleanly terminated native session after the worker import error. The preserved lineage totals 55,866.17972529016 active job seconds, 137 attempted model sessions and 25,698,304 allocated receipt bytes. Those costs remain charged; they are not task-quality results. A synthetic import reproduces the Python 3.9/cp312 regex mismatch. The exact same frozen worker passes synthetic success and refusal inputs under the already installed Python 3.12 image, within the unchanged bounds. Every BFCL class additionally imports inside its native sandbox with an empty synthetic call list. The preparation checks unchanged grading definitions, all source evidence and the exclusive saved-answer grading path. Forty-six unit checks pass under each local Python runtime. No held-out response was scored or inspected during preparation. The separately frozen recovery has yet to execute at this capture; complete final outcomes, product gates and performance remain pending. ### Host recovery and complete practical comparison, October 5 [[sources/runs/2026/10/2026-10-05-practical-benchmark-host-recovery]] preserves the authorized service recovery, both incomplete excluded attempts, the new prospective pilot policy, all twenty passing runner checks and every receipt from the complete follow-up. The native executable and both pack recipes are unchanged. The new policy retains short CPU spikes as diagnostics, excludes at least five seconds of consecutive observed competing CPU or any known heavy job, and waits for twenty quiet nominal seconds before launch, bounded to five minutes. These are operating heuristics, not proof that shorter activity has no effect. Old policies and results remain unchanged; no thermal, paging or allocation rule is relaxed. All twelve processes complete their fixed workloads and natural-answer completion flags inside their physical ceilings. The following values summarize all three runs of each complete configuration. They prove the observed process envelope and cache allocation, not a throughput advantage or task-quality equivalence. GB values are decimal; physical peaks are rounded. | Saved ceiling | Pack | Actual planned target (GB) | Expert slots | Largest observed physical footprint (GB) | | --- | --- | --- | --- | --- | | 14 GB | Native smaller pack | 14 | 1941 | 11.202 | | 14 GB | Original | 14 | 1635 | 11.370 | | 22 GB | Native smaller pack | 22 | 4923 | 17.847 | | 22 GB | Original | 22 | 3531 | 18.040 | Five timing cells are excluded for sustained background CPU or global paging. The complete frozen analysis has `all_timings_eligible: false` and `qualification: false`, and emits no clean medians. Individually eligible runs are not selected into a replacement comparison. This timing attempt is discarded. The original remains the only supported/public pack; a quiet-host repeat, alternative product activation/distribution and redistribution clearance remain open. ### Focused real-expert two-bit and mixed screen, October 5 [[sources/runs/2026/10/2026-10-05-focused-lowbit-screen]] preserves the frozen component protocol, complete results and both incomplete prefill/cache admission attempts. Ten declared experts from each of three original layers are fully authenticated before transient conversion. Five alternating-order rounds compare all-four-bit, all-three-bit, all-two-bit, two-bit gate/up with three-bit down, and three-bit gate/up with two-bit down. Activations are deterministic synthetic probes; the gathered RHS bank contains ten experts. These are component costs and reconstruction errors, not whole-model quality or throughput. The complete screen finishes in 29.286 seconds at a 618,316,592-byte lifetime physical peak, with nominal thermal observations, zero swap deltas and one isolated renderer CPU spike that does not meet the frozen sustained-contention exclusion. Three-row median gate/up/down kernel times for the promising mixed recipe are 0.379/0.366/0.370 ms across the declared layers, against 0.396/0.394/0.392 ms for all-three-bit. All-two-bit and the two-bit-down mix are slower on those rows. The mixed recipe reduces a complete expert record from 2,150,400 to 1,740,800 bytes relative to three-bit. That storage reduction does not itself predict generation speed. Two-bit relative weight MSE is about three and a half times three-bit; the nonlinear output probes also worsen. Advance only the two-bit gate/up, three-bit down lead to a small full-model distribution screen before any export. Preserve the original and current candidate. Do not advance all-two-bit or the two-bit-down mixture merely because they use fewer bytes. Different group sizes, calibration or VQ remain separate untested hypotheses. The first prefill/cache attempt launched no model; the retry completed only its automatic arm, then failed the unchanged quiet-start admission for the capped arm. Both restore their temporary host controls. There is no complete paired timing result and no selected partial benefit. The native diagnostic changes pass 24 Python runner tests and 88 native protocol assertions. Further quiet-host retries are deferred while the cheaper quality lead is tested. ### Focused lower-bit mixture quality screens, October 5 [[sources/runs/2026/10/2026-10-05-focused-lowbit-quality]] preserves two complete prospective screens. Each reproduces an existing coding control's complete vocabulary bytes before evaluating the fixed coding, tool and multilingual contexts. No new model payload or raw logits are saved. The reference remains quantized, and these are calibration positions rather than completed tasks. | Mixture | Comparator on the same three contexts | Comparator/candidate mean KL | Comparator/candidate top-choice agreement | Physical lifetime peak | Screen outcome | | --- | --- | --- | --- | --- | --- | | Affine two-bit gate/up, three-bit down; original non-experts | Prior all-three-bit control | 0.591234 / 0.666165 | 0.729167 / 0.687500 | 2,167,408,008 bytes | Passes the frozen loose screening limits; coding agreement loses two positions | | VQ2.1 experts/PLE/norms with original four-bit dense projections | Original pack | 0.491535 / 0.531509 | 0.791667 / 0.750000 | 2,387,675,560 bytes | Passes its frozen screen; coding ties, tool and multilingual each lose one position | The affine control checks every converted three-bit tensor against the earlier authenticated export and reproduces the coding vocabulary hash exactly. Its four forwards finish in 251.402 seconds. The VQ control reproduces the prior VQ3.2/dense4 coding vocabulary hash; the new target's own file/runtime identities and all destination shapes are explicit. Its four forwards finish in 105.222 seconds. Both stay below ten GB and retain actual headroom checks. These functional durations are not performance comparisons. Advance VQ2.1/dense4 to native parity, representative completed tasks and bounded cost measurement first because its existing artifact and kernels avoid another large export. The affine mixture remains a separate positive lead needing completed tasks. Neither is qualified for Auto, public installation, similar task quality or the twenty-token claim. Full VQ2.1's earlier adverse screen and all original evidence remain unchanged. The new VQ mixture has its own complete pinned metadata identity; native implementation and acceptance are subsequent work. ### Lower-bit VQ dense mixture native parity, October 5 [[sources/runs/2026/10/2026-10-05-vq21-dense-native-screen]] captures the source-bound research implementation, independent reference and native prefill/greedy checks. The exact mixture retains its own authenticated inventory, manifest and arithmetic identity. Native prefill and sixteen self-fed steps pass without tolerance changes. Physical supervision stays within ten GB. Unsupported flags stopped the first prefill invocation before allocation; a receipt-nesting error stopped the next driver after its successful prefill. The final continuation performs only the remaining greedy operations. These failures do not erase successful checks, and the results establish neither similar completed-task quality nor throughput. Existing original-pack behavior and public choices are unchanged. ### Lower-bit VQ speculative reduction correction, October 5 [[sources/runs/2026/10/2026-10-05-vq21-draft-reduction-repair]] records the initial 682 failed assertions among 2,071, then all 2,071 passing after the focused D8 correction. Depth one matched ordinary greedy generation while depth two and four changed answers: the D8 kernel switched numerical reduction above twenty routed pairs. Row-invariant verification now retains the single-token SIMD reduction through fifty pairs, with unchanged ordinary reference dispatch. Nonconstant component tests prove batch/one-row equality at one, two, three and five rows. The corrected run peaks at 8,376,964,832 physical bytes, below the existing ten-GB bound, and exercises 44 accepted drafts. Failed and corrected observations are both preserved. This is correctness evidence, not a throughput or completed-task-quality result. ### Lower-bit VQ mixed completed-task rejection, October 5 [[sources/runs/2026/10/2026-10-05-vq21-mixed-completed-tasks]] preserves every answer, grade and the unchanged frozen task protocol. The original scores fifteen of sixteen and VQ2.1/dense4 fourteen. Both fail sort-records. The mixture additionally fails filter-unique, returning `[2, -3, 0, -8, 4, -1]` where the original correctly returns `[-3, -8, -1]`. The tested coding, tool, multilingual and retrieval outcomes remain passing. The process peaks at 8,379,012,952 physical bytes within ten GB. Functional task time is not an eligible throughput comparison. The mixture misses the prospectively fixed fifteen-pass minimum. Reject advancement to a speed pilot or product integration without loosening the gate. This focused adverse result does not estimate broad quality loss or prove every lower-bit mixture unsuitable. The separate affine two-bit gate/up, three-bit down lead remains open. The complete existing static suite, all 104 native T0/T1 catalogue groups and Python overlay checks pass on the source-bound corrected binary. Their success preserves the numerical repair, not a quality qualification for this mixture. ### Mixed affine two/three-bit export and task rejection, October 5 [[sources/runs/2026/10/2026-10-05-affine223-completed-screen]] preserves the frozen duplicate-coalescing attempts, export, exact implementation identity, complete numerical checks and all sixteen answers. The first coalescing attempt stopped before changing any path because of Python-version compatibility; the corrected frozen attempt reverified all payloads and retained both names as hard links. The complete original installation and every unique weight remain intact. Allocated staging after coalescing is 364,793,344,000 bytes. The expert-only two-bit gate/up, three-bit down export writes 42,781,961,312 bytes in 141.133 seconds, peaking at 259,719,984 physical bytes. Its complete manifest is `1edbd2d7b01a1f15a185b3c4107feefb195c4e027c654c7f00dbdc981b4ad1c3`. All seven sequential reference/native operations complete within the frozen bounds. Traversal order is exact for four layers and 513 tokens. All forty-nine native prefill points have zero relative error through cold, reuse, grow and shrink cache histories. Sixteen self-fed greedy steps match across four cache histories; speculative target/state/recovery checks pass with seventy-six accepted drafts. The bounded numerical-reference addition is 28,682,240 bytes, with a prospectively declared 2.1-GB total raw-reference ceiling. Task process peak is 8,897,173,288 physical bytes within ten GB. These are functional checks, not eligible throughput measurements. Original scores fifteen of sixteen tasks; mixed affine223 scores twelve. Both fail sort-records. Only the mixture fails contradictory-source and spanish-structured by fencing JSON despite bare-JSON instructions, and hindi-structured by returning total sixteen instead of seventeen. Tested coding and tool cases still pass. The mixture misses the fixed fifteen-pass minimum; hold it without a speed pilot or product promotion. The earlier proxy screen is not a substitute for completed-task quality. Preserve the original and continue the separately frozen low-budget allocation probe. The exact mixed identity remains available only to the bounded research adapter, outside Engine memory recipes and Auto. ### Original ten-GB baseline and fourteen-GB allocation screen, October 5 [[sources/runs/2026/10/2026-10-05-original-lowbudget-allocation-screen]] captures the complete single retry following the preserved no-model admission failure. All three native processes pass their physical ceilings, natural completion and frozen timing eligibility; no failed sample is replaced or incorporated. These are single-process directional observations on the 48-GB M5 Pro, with uncontrolled OS file cache. The smaller process ceiling does not simulate another Mac's total RAM, SSD, chip or thermal design. | Actual recipe | Short / context / coding generation tokens/s | Short / context / coding request seconds | Long request seconds | Lifetime physical peak bytes | | --- | --- | --- | ---: | ---: | | Original, ten-GB ceiling, Auto ordinary decode | 9.343 / 10.156 / 9.035 | 28.994 / 39.011 / 11.361 | Not included | 7,750,505,344 | | Original, fourteen-GB ceiling, automatic 1,024-row prefill | 12.656 / 15.043 / 14.573 | 21.778 / 26.901 / 7.883 | 45.692 | 11,558,608,000 | | Same fourteen-GB ceiling, 512-row prefill cap | 13.022 / 14.747 / 15.037 | 21.266 / 31.463 / 7.770 | 55.007 | 12,283,747,840 | The fourteen-GB pool increases from 1,635 to 1,875 slots. Short/coding generation gains roughly three percent; context generation declines about two percent. Context and long-request wall time increase about seventeen and twenty percent. Both 8,356-token long lookups answer exactly `code-0853`, and short/coding IDs match the earlier original result and each other. The long answer has too few generated tokens to stand in for full-request utility. This fails the prospectively requested roughly five-percent generation gain across standard workloads without material request regression. Retain existing allocation, with no paired confirmation or new adaptive controller justified by this result. The full low-budget target remains unmet. ### Native-three-bit ten-GB draft recipe misses its gain gate, October 5 [[sources/runs/2026/10/2026-10-05-native-threebit-lowbudget-screen]] records one complete timing-eligible candidate process, compared with the separately frozen original ten-GB baseline. Candidate short/context/coding generation is 8.199/9.760/9.235 tokens/s against original 9.343/10.156/9.035. Candidate request wall time is 32.815/38.798/11.311 seconds against 28.994/39.011/11.361. Its coding answer's committed IDs match exactly. Candidate lifetime physical peak is 7,188,959,960 bytes and observed loading is 11.168 seconds; these are process observations with uncontrolled filesystem cache, not cold-SSD or other-Mac qualification. The candidate misses its declared useful-gain gate, so its unchanged draft-enabled recipe receives no repeated comparison or product promotion. In the short case, the 731-slot candidate has zero pool hit rate while original's 821 slots retain a 0.39098039215686275 hit rate. Candidate accepts 140 of 230 drafts and reduces target forwards from 255 to 115, but demanded expert reads rise from 69,310,771,200 to 77,947,699,200 bytes. Issued forecast bytes rise from 154,707,148,800 to 207,782,400,000. Decode read timers are 9.065 and 14.304 seconds. These overlapping counters are not additive causal wall time or physical SSD traffic. They motivate one different low-budget plain-decode/caching experiment, not a blind draft-depth grid or a claim that every three-bit recipe loses. The same source also preserves the isolated-policy CI compiler failure and its exact helper-relocation fix. The existing local model-free proxy passes all 965,028 assertions. Model arithmetic is unchanged; full earlier static/native results remain scoped to their original frozen build, and new-head CI is tracked separately. ### Native-three-bit ten-GB plain-decode follow-up also misses, October 5 [[sources/runs/2026/10/2026-10-05-native-threebit-plain-planning]] preserves the prospectively frozen input sequence: an off/depth-two inconsistency is refused before allocation, the corrected model-free Auto plan increases prefill, and the scored input retains the previous ten-GB arms' 256-row maximum so freed draft memory reaches the cache. No scored output informs these setup refinements. [[sources/runs/2026/10/2026-10-05-native-threebit-plain-screen]] captures the single complete run, with all physical and timing conditions passing and the original coding-answer IDs preserved. The 893-slot plain recipe reaches 9.826/9.371/8.750 tokens/s in short/context/coding order. Ratios against original are 1.0516046/0.9226965/0.9684619, and request-wall ratios are 0.9575561/1.0160672/1.0375347. Short-case cache hit rate rises from the draft recipe's zero to 0.4081209150326797 and demanded bytes fall to 57,275,904,000. Lifetime physical peak is 7,341,396,816 bytes. These are one-pilot process observations, with overlapping read counters and uncontrolled OS file cache. The mechanism restores cache reuse, but the complete recipe loses the fixed useful-gain gate because context and coding generation slow. Do not repeat or promote it. Combined with the separately captured allocation and lower-bit quality failures, the current evidence does not support the full low-memory twenty-token target. [[records/decisions/same-model-lowbudget-target-remains-unmet]] records the scope, stopping decision and explicit options without asserting universal impossibility or completion. ### Final pure-policy factory acceptance, October 5 [[sources/runs/2026/10/2026-10-05-affine-research-policy-factory-repair]] preserves the isolated-source compiler failure, the subsequent private-setter failure in full-module CI, and the final typed pure-factory repair. Private setters remain private; only artifact-independent descriptors cross into ModelConfig construction. The full release build, all 104 native catalogue groups and all 965,028 isolated-policy assertions pass on the corrected source. Earlier full static acceptance is reused in its unchanged scope. New-head CI is tracked separately, and the helper separation is not reported as a new model-performance measurement. [[sources/runs/2026/10/2026-10-05-final-quantization-source-ci]] closes all remote source workflows for `20cdda3`, including engine static/runtime safety, native catalogue, goldens, public-library, coverage, Mac scripted/snapshot/Xcode and context checks. Experiment-documentation commit `9be7f49` separately passes its documentation gate. No new model timing or quality result follows from CI acceptance. ### Corrected forecast transfer, October 6 [[sources/runs/2026/10/2026-10-06-candidate-correction-functional-only]] preserves the next narrow test. The speed pair is discarded: the first arm has system-wide paging and the second is refused before launch after the quiet-admission deadline. It supplies no speed verdict. A separate functional-only run passes exact token/text equality on all three workloads and the same completed coding answer. Physical peak is 11,211,529,128 bytes inside fourteen GB. The correction has its own resource identity, with an authenticated-header memory charge carried through replanning; the old native profile still refuses that changed charge. All native T0/T1 checks and the existing Python performance tests pass after a complete project-module rebuild resolved a stale incremental consumer crash. Keep the explicit research path available without changing original defaults, supported packs, Auto or user overrides. The only remaining question in this follow-up is a clean paired speed measurement. Neither the functional result nor the discarded timings establishes twenty tokens/s or a useful prefetch gain. ### v0.2.28 published and accepted **[v0.2.28](https://github.com/carloslfu/slotstream/releases/tag/v0.2.28) is published and accepted.** It ships maintained-pack selection in the CLI, observed-footprint memory planning, complete-allocation checks for fixed caches, and request-prefetch cleanup. Original four-bit remains the only supported pack. Experimental quantizations remain research-only; this engine release does not distribute a Desktop installer. Published 2026-10-07T19:24:58Z from source commit `71ebedd173aab089260f9085adc581cd9c20f1d4`. The complete engine source closure in the accepted CI archive matches current main; intervening commits change only documentation and research records. Release documentation is in `2e6b543b6d6cf81f24f023e39aea2c352ceeb48a`. Archive SHA-256: `1947c1b325f80097f0eac3f86da2ef7b091ced32d50e4487703f683f28b0980b`. Executable SHA-256: `fee022e64b37abd25260d2c5552fda4d9a66383c86e68449d814dd284c94d9e0`. | Acceptance | Result | | --- | --- | | Exact-source hosted CI | Engine, instrumented coverage, external library consumer, Mac app and context contracts passed | | Complete native battery on the CI executable | 35 gates passed, zero failures or skips; includes quality 15/15, API robustness 74/74 and vision serving 25/25 | | Public distribution | Published archive matches the preserved CI archive; checksum, GitHub provenance and build identity verified | | Public installer | Unchanged public installer fetched the latest release into an isolated installation; installed executable matches CI bytes; user installation and shell profiles unchanged | | Installed serving | 31/31 checks passed with Auto quantization, a 10 GB adaptive ceiling, 32,768-token context and MTP on; owned server reaped and no model processes remained | The linked run preserves raw native, installer and serving logs, commands, identities, workflow receipts and cleanup evidence. Global paging remains diagnostic: the native run observed swap-ins without increased swap-outs, while process-memory and functional gates passed. These checks establish functional and distribution acceptance, not a new throughput result or qualification of other Macs. The broader same-checkpoint speed target remains unmet, as recorded in the closed [[records/plan/same-model-quantization-and-automatic-memory-2026-10-02]] plan.