V.* — the V-store¶
Token-ID-indexed storage for K and V tensors, one session per <ns>_pk /
<ns>_pv pair created by KV.PREFIX.REGISTER. The wire forms of
V.STOREBATCH and V.FETCH … RANGE / BATCH are in the
M14 externalized-attention table
and the KV.PREFIX.* page; this page covers the storage
formats.
Quantization tiers¶
Measured on a 20-question / 50-token-greedy BLEU eval against the standalone reference:
| Format | Storage vs FP16 | Mean BLEU | First-token | Notes |
|---|---|---|---|---|
| fp16 | 1.00× | 0.969 | 100% (20/20) | Bit-identical on 17/20, brief late drift on 3/20. Production default. |
| int8 | 2.00× | partial | 95%+ | Single-step argmax safe. Multi-token decode drifts; not benchmarked end-to-end here. |
| turbo4 | 3.51× | 0.538 | 90% (18/20) | Argmax preserved on first token, but compounds catastrophically over greedy decode. Single-step / classification only. |
Recommendation: ship vquant=fp16. Document int8 and turbo4 as opt-ins for greedy-tolerant single-step workloads (function-calling tool selection, classification, single-token routing) where the storage win matters more than multi-token fidelity.
The mlx4g32 tier¶
KV.PREFIX.REGISTER <ns> <kv_dim> mlx4g32 (and V.CREATE ... VQUANT mlx4g32)
stores K/V in mlx's QuantizedKVCache layout: int4, group 32, affine.
per token, per layer, D = kv_dim
[packed uint32 x D/8] 4-bit codes, element j of a group at bits 4*(j%8)
[scales fp16 x D/32]
[biases fp16 x D/32]
640 B/token at D=1024 against fp16's 2048 — 3.2x — which is what puts an 86,580-token cartridge for a 4B model inside a 16 GB machine (fp16: 12.5 GB).
Both parameters are measured, not conventional:
- group 32, not 64. g64 corrupted recalled facts at digit grain in real
generations (
"2430-04-22"for"2030-04-22"). Knowledge held as KV is bits-fragile the same way weight-held knowledge is. - affine, not symmetric. Storing a per-group scale and bias beat symmetric
by 2.92 pp on K, and K's error dominates the end-to-end result — measured
here as 0.038 mean abs error against
turbo4's 0.047 on the same input.
The scale is round-tripped through fp16 before the codes are chosen, so the quantizer targets the scale the reader will actually see.
Heterogeneous K/V (per-layer formats)¶
PionPromptCache(boundary_protect=N) switches the per-prefix register from the legacy uniform KV.PREFIX.REGISTER to two raw V.CREATE … SCHEMA calls (one per K/V side). K stays fp16 across all layers — K drives softmax routing and tolerates quant noise poorly. V uses fp16 for the first N + last N layers ("boundary protection") and the user's vquant for the middle layers.
# Boundary-protect K-V split: middle layers in fp8, boundary layers in fp16
pc = PionPromptCache(model, vquant="fp8", boundary_protect=2)
Measured on Llama-3.2-1B-Instruct-4bit, 3 prompts × 5 queries (Apple Silicon, single worker):
| Config | TTFT speedup | Throughput speedup | First-token agreement |
|---|---|---|---|
| Uniform fp16 | 2.75× | 2.74× | 14/15 |
vquant=fp8, boundary_protect=2 |
2.88× | 2.87× | 15/15 |
Faster AND more accurate than uniform fp16 on this workload — fp8 V on middle layers carries less wire data per layer; the fp16 boundary layers protect routing.
Trade-off: the SCHEMA path skips KV.PREFIX.REGISTER's cross-worker directory publish (single-worker visibility only). A future KV.PREFIX.REGISTER.SCHEMA server command would lift this.
boundary_protect requires kv_dim divisible by 32 when vquant ∈ {fp8, turbo4, turbo3, turbo2}. Llama-3.2-1B (kv_dim=512) and Llama-3.2-3B (kv_dim=1024) both qualify.