# Tesla V100 port The `sm_70` build is a peer hardware implementation of the same NInfer product surface as the default `sm_120a` build. It accepts all five registered artifacts and retains Text, Vision, MTP, prefix reuse, CLI, OpenAI Chat Completions/Responses, Anthropic Messages, and the 35B-A3B text-only DFlash route. It does not add a second artifact format or model runtime. ## Build CUDA 12.8 is required: CUDA 13 removed offline compilation for Volta. Configure the architecture and compiler explicitly: ```bash cmake -S . -B build-v100 -G Ninja -DCMAKE_BUILD_TYPE=Release \ -DCMAKE_CUDA_COMPILER=/usr/local/cuda-12.8/bin/nvcc \ -DCMAKE_CUDA_ARCHITECTURES=70 cmake --build build-v100 -j ``` The Volta implementation uses native FP16 tensor cores through CUTLASS where appropriate, exact Volta kernels for the registered attention/state contracts, and a vendored llama.cpp MMA flash-attention kernel for long causal prefill. INT8 group-64 KV applies the registered normalized D256 Hadamard transform to both persistent K and transient Q before the flash calculation. NVFP4 dense MLP gate/up and down weights are uploaded in checkpoint-native layout and then permuted in place into the Volta QPN fragment order during model load. The transform covers both the code and scale planes. Row-scaled FP8 weights receive a corresponding code-plane permutation. Both transforms preserve allocation sizes, record `VoltaQpnPrepacked` in the runtime weight view, and release temporary device storage before inference. The artifact on disk is not rewritten. Decode consumes the QPN layouts directly; wide prefill recognizes them and reconstructs FP16 for its tensor-core GEMM. FP8 projection groups stage a shared FP16 activation once and reuse it across their QPN launches. Learned MTP accepts one through seven draft positions. The native INT8 attention kernel handles verification widths through seven; the width-eight target round uses the qualified 6+2 attention composition. ## Qualification The default Volta run passes all 90 executed CTest cases; six real-artifact cases remain opt-in. The complete attention contract suite passes for causal BF16/INT8/FP8 paths, packed Vision attention, and DFlash context attention. The Qwen3.8-27B NVFP4 source-pressure case separately passes its real-artifact state-materialization path, and the performance sweep loads and exercises all five artifacts on 32-GiB V100 hardware. NVFP4 and K8V4 KV storage are not available on Volta. ## Single-request generation sweep The decode headline uses the short-context production target-round benchmark with two warmups, ten measured rounds, and CUDA Graphs. The sm_70 width-6+ target-verify regression is fixed (see below), so `--spec mtp` is no longer capped below upstream's [1,7] window; the headline below is the actual peak of a full K sweep on this corpus, at the optimized proposal head: | Qwen3.8-27B NVFP4 | K | Decode tok/s | Draft acceptance | |---|---:|---:|---:| | Current port | 1 | **219.0** | 99.2% | The full sweep behind that headline, same corpus and methodology: | K | Prefill tok/s | Decode tok/s | Draft acceptance | |---:|---:|---:|---:| | 1 | 1,102.5 | **218.98** | 99.2% | | 2 | 1,097.4 | 213.99 | 98.3% | | 3 | 1,094.9 | 209.24 | 97.5% | | 4 | 1,096.6 | 204.02 | 97.9% | | 5 | 1,100.3 | 199.58 | 97.1% | | 6 | 1,089.6 | 178.47 | 92.5% | | 7 | 1,087.2 | 180.74 | 93.3% | Narrow windows win outright on this corpus: acceptance is near-ceiling at every K, so round-verify cost dominates once there's little more accepted length left to buy. Treat this as a synthetic-corpus ceiling, not a general-generation rate -- see below for the practical, real-text sweet spot. The artifact sweep used the public `ninfer_bench` Engine route on a Tesla V100-PCIe-32GB with CUDA 12.8 and INT8 group-64 KV. Prefill is an isolated `pp2048` run. Decode is `pp2048+tg256` with CUDA Graphs and the optimized proposal head at the selected draft window. Each result has one discarded warmup and three measured repetitions; decode values are committed output tokens per second, not engine steps per second. Against the previous SXM2 DFlash sweep (same tool and methodology) the current PCIe round latency is a uniform ~6-7% higher across every draft window, so the preferred SXM2 recovers roughly that much; this decode workload is HBM-bound and the host and PCIe bus barely enter into it. | Model profile | K | Prefill tok/s | Decode tok/s | Draft acceptance | |---|---:|---:|---:|---:| | Qwen3.6-27B `groupwise-int` | 4 | 1,085.0 | 54.54 | 66.5% | | Qwen3.6-27B `nvfp4` | 5 | 223.8 | 55.22 | 54.4% | | Qwen3.8-27B `groupwise-int` MTP | 5 | 1,083.9 | 130.96 | 97.1% | | Qwen3.8-27B `nvfp4` MTP | 5 | 1,100.3 | 199.58 | 97.1% | | Qwen3.8-27B `groupwise-int` DFlash2 | 7 | 1,044.2 | 77.84 | 100% | | Qwen3.8-27B `nvfp4` DFlash2 | 7 | 1,059.0 | 126.32 | 100% | | Qwen3.6-35B-A3B `groupwise-int` DFlash | 4 | 686.2 | 139.58 | 90.9% | Decode throughput is acceptance-sensitive; see the MTP sweep above for how much. Real, less predictable text favors a wider window than this corpus does -- the practical production sweet spot is K=3. The DFlash2 rows use `--spec dflash2` at K=7, the measured peak of a 3-10 draft-window sweep on this same pp2048+tg256 shape; MTP stays ahead of DFlash2 on this corpus-continuation shape at every K tried. The Qwen3.6-35B-A3B `DFlash` row now reads 139.58 tok/s against a previously recorded 245.05 -- not a regression: the detailed 15-point DFlash K-sweep below was rewritten by the same commit that wrote 245.05 and puts K=4 at 120.97 tok/s, directly contradicting it. Today's 139.58 sits right on that sweep's own trend (close to its K=3 peak of 125.85); 245.05 was simply wrong from the moment it was typed. ### Varied-context DFlash2 sweep The Qwen3.8-27B NVFP4 varied-context sweep uses INT8 group-64 KV, CUDA Graphs, the optimized proposal head, and 128 generated tokens. The 2K through 32K rows use a verified non-repetitive 37,758-token corpus (only four duplicate 15-token windows) with one discarded warmup and three measured repetitions. Each backend column reports the best tested static K at that depth. | Context | No spec tok/s | Best DFlash2 | DFlash2 tok/s | Acceptance | Best MTP | MTP tok/s | Acceptance | |---:|---:|---:|---:|---:|---:|---:|---:| | 2,048 | 29.28 | K=7 | **68.33** | 35.3% | K=3 | 67.66 | 59.9% | | 8,192 | 28.02 | K=3 | 46.92 | 50.7% | K=3 | **92.74** | 78.6% | | 16,384 | 26.70 | K=3 | 39.95 | 40.5% | K=3 | **49.80** | 49.0% | | 32,768 | 22.13 | K=7 | **55.43** | 38.6% | K=4 | 55.04 | 57.8% | | 150,000 | 15.38 | K=7 | 41.09 | 53.7% | K=2 | **43.63** | 79.2% | The 150K row is a single measured repetition over a deterministic 151,032-token profiling corpus, constructed from shifted variants of the novel corpus. It is suitable for long-context timing but not a natural-language acceptance claim. DFlash2 K=7 is the best static DFlash2 choice at both ends of the tested range. The intermediate K=3 wins show that any adaptive policy should respond to observed acceptance, not context depth alone. MTP remains the general-purpose default. ### Context-lookup MTP MTP has a lossless fast path for output that reproduces the request's own context. The host finds the most recent earlier occurrence of the exact 16-token suffix and proposes up to fifteen tokens that followed it. The longer verification topology is selected only when the learned MTP drafts agree with the lookup prefix and every active batch row qualifies. Verification remains the authority for every emitted token; a miss or disagreement uses the ordinary configured MTP window. While lookup remains active, the graph produces one learned MTP token for the next qualification check rather than rebuilding the full configured proposal window. The target still verifies all fifteen lookup tokens, so this removes redundant serial draft work without weakening verification or changing ordinary-generation proposals. The copy qualification used the V100-PCIe-32GB, CUDA Graphs, INT8 group-64 KV, the optimized proposal head at the Volta draft-window cap of four, greedy sampling, and a 172-token prompt that both profiles reproduced verbatim. | Qwen3.8-27B profile | Context lookup tok/s | Mean output/round | |---|---:|---:| | `groupwise-int` | **130.9** | 12.91 | | `nvfp4` | **201.0** | 12.91 | These are context-reproduction results, not general-generation headlines. On ordinary prose the lookup eligibility test did not fire, output remained identical to the control, and decode stayed within measurement noise. ### DFlash The 35B-A3B production steady-round benchmark used an optimized proposal head, CUDA Graphs, batch one, and a 2,048-token established context. Each K used two warmups and ten measured rounds with real greedy acceptance, on the V100-PCIe-32GB. | K | Round ms | Acceptance | Mean output/round | Published tok/s | |---:|---:|---:|---:|---:| | 1 | 23.80 | 90.0% | 1.9 | 79.84 | | 2 | 27.14 | 90.0% | 2.8 | 103.16 | | 3 | 30.20 | 93.3% | 3.8 | **125.85** | | 4 | 33.89 | 77.5% | 4.1 | 120.97 | | 5 | 35.71 | 64.0% | 4.2 | 117.62 | | 6 | 40.36 | 51.7% | 4.1 | 101.57 | | 7 | 43.46 | 54.3% | 4.8 | 110.46 | | 8 | 50.01 | 47.5% | 4.8 | 95.98 | | 9 | 53.40 | 43.3% | 4.9 | 91.76 | | 10 | 55.77 | 38.0% | 4.8 | 86.07 | | 11 | 57.43 | 32.7% | 4.6 | 80.10 | | 12 | 63.15 | 30.0% | 4.6 | 72.84 | | 13 | 67.48 | 28.5% | 4.7 | 69.65 | | 14 | 70.43 | 25.7% | 4.6 | 65.31 | | 15 | 71.86 | 24.7% | 4.7 | 65.40 | K=3 is the measured DFlash default for this workload. Longer windows increase licensed length only slightly while proposal cost grows and acceptance falls. ## Preferred SXM2 device On a dual-V100 host, set `NINFER_GPU_UUID` to the UUID of the intended SXM2 and invoke `tools/v100/ninfer-sxm2.sh`. The launcher validates that the selected device is a `Tesla V100-SXM2-32GB`, sets `CUDA_VISIBLE_DEVICES`, and runs `build-v100/apps/ninfer`: ```bash NINFER_GPU_UUID= tools/v100/ninfer-sxm2.sh --help ``` Set `NINFER_EXECUTABLE` when the binary lives elsewhere. UUID selection prevents PCI enumeration from silently moving benchmarks to a PCIe-board GPU. The executable still receives `--device 0`, because CUDA remaps the selected UUID to logical device zero. For Qwen3.8-27B serving, pass `--prefill-chunk 2048`. On the preferred SXM2 with the NVFP4 artifact and INT8 KV, an 8,192-token prefill measured 1,137.0 tok/s versus 1,078.6 tok/s with the generic 1,024-token default, a 5.4% improvement.