Skip to content

V.* — the V-store

Token-ID-indexed storage for K and V tensors, one session per <ns>_pk / <ns>_pv pair created by KV.PREFIX.REGISTER. The wire forms of V.STOREBATCH and V.FETCH … RANGE / BATCH are in the M14 externalized-attention table and the KV.PREFIX.* page; this page covers the storage formats.

Quantization tiers

Measured on a 20-question / 50-token-greedy BLEU eval against the standalone reference:

Format Storage vs FP16 Mean BLEU First-token Notes
fp16 1.00× 0.969 100% (20/20) Bit-identical on 17/20, brief late drift on 3/20. Production default.
int8 2.00× partial 95%+ Single-step argmax safe. Multi-token decode drifts; not benchmarked end-to-end here.
turbo4 3.51× 0.538 90% (18/20) Argmax preserved on first token, but compounds catastrophically over greedy decode. Single-step / classification only.

Recommendation: ship vquant=fp16. Document int8 and turbo4 as opt-ins for greedy-tolerant single-step workloads (function-calling tool selection, classification, single-token routing) where the storage win matters more than multi-token fidelity.


The mlx4g32 tier

KV.PREFIX.REGISTER <ns> <kv_dim> mlx4g32 (and V.CREATE ... VQUANT mlx4g32) stores K/V in mlx's QuantizedKVCache layout: int4, group 32, affine.

per token, per layer, D = kv_dim
  [packed uint32 x D/8]   4-bit codes, element j of a group at bits 4*(j%8)
  [scales  fp16  x D/32]
  [biases  fp16  x D/32]

640 B/token at D=1024 against fp16's 2048 — 3.2x — which is what puts an 86,580-token cartridge for a 4B model inside a 16 GB machine (fp16: 12.5 GB).

Both parameters are measured, not conventional:

  • group 32, not 64. g64 corrupted recalled facts at digit grain in real generations ("2430-04-22" for "2030-04-22"). Knowledge held as KV is bits-fragile the same way weight-held knowledge is.
  • affine, not symmetric. Storing a per-group scale and bias beat symmetric by 2.92 pp on K, and K's error dominates the end-to-end result — measured here as 0.038 mean abs error against turbo4's 0.047 on the same input.

The scale is round-tripped through fp16 before the codes are chosen, so the quantizer targets the scale the reader will actually see.

Heterogeneous K/V (per-layer formats)

PionPromptCache(boundary_protect=N) switches the per-prefix register from the legacy uniform KV.PREFIX.REGISTER to two raw V.CREATE … SCHEMA calls (one per K/V side). K stays fp16 across all layers — K drives softmax routing and tolerates quant noise poorly. V uses fp16 for the first N + last N layers ("boundary protection") and the user's vquant for the middle layers.

# Boundary-protect K-V split: middle layers in fp8, boundary layers in fp16
pc = PionPromptCache(model, vquant="fp8", boundary_protect=2)

Measured on Llama-3.2-1B-Instruct-4bit, 3 prompts × 5 queries (Apple Silicon, single worker):

Config TTFT speedup Throughput speedup First-token agreement
Uniform fp16 2.75× 2.74× 14/15
vquant=fp8, boundary_protect=2 2.88× 2.87× 15/15

Faster AND more accurate than uniform fp16 on this workload — fp8 V on middle layers carries less wire data per layer; the fp16 boundary layers protect routing.

Trade-off: the SCHEMA path skips KV.PREFIX.REGISTER's cross-worker directory publish (single-worker visibility only). A future KV.PREFIX.REGISTER.SCHEMA server command would lift this.

boundary_protect requires kv_dim divisible by 32 when vquant ∈ {fp8, turbo4, turbo3, turbo2}. Llama-3.2-1B (kv_dim=512) and Llama-3.2-3B (kv_dim=1024) both qualify.