Shared KV cache¶
Every request that starts with the same tokens — a system prompt, a few-shot block, a retrieved passage — makes the model compute the same K/V tensors for them again. Pion keeps those tensors after the first time, keyed by a namespace that names the exact prefix, so the next request fetches them instead of recomputing: from the same process, from another process, or after a restart.
This page is the concept: what a namespace must contain, how the cold and warm
paths work, and when Pion should compute the attention itself. The commands
are in the reference — KV.PREFIX.*,
V.*,
ATTEND.* — and the Python API is
pion-vllm-mlx.
Quick start¶
Start the server:
Then, from Python (install from a checkout with pip install -e 'pion-vllm-mlx/[mlx]' —
not on PyPI yet):
from mlx_lm import load, generate
from pion_vllm_mlx import PionPromptCache
model, tok = load("mlx-community/Llama-3.2-1B-Instruct-4bit")
pc = PionPromptCache(model, vquant="fp16")
system = "You are a support agent for Acme. Answer in one sentence." # shared by every request
prefix_ids = tok.encode(system)
ns = PionPromptCache.make_namespace("llama-3.2-1b-4bit", "fp16", system)
# The first call anywhere prefills locally, registers with Pion and stores the K/V.
# Every later call — this process, another one, or after a restart — fetches it.
cache = pc.get_or_prefill(prefix_ids, namespace=ns)
print(generate(model, tok, prompt=tok.encode(" How do I reset my password?", add_special_tokens=False),
prompt_cache=cache, max_tokens=40))
pc.stats() reports hits, misses, hit_rate, fetch_ms_total, store_ms_total.
Heterogeneous KV cache¶
PionPromptCache(boundary_protect=N) switches the per-prefix register from the legacy uniform KV.PREFIX.REGISTER to two raw V.CREATE … SCHEMA calls (one per K/V side). K stays fp16 across all layers — K drives softmax routing and tolerates quant noise poorly. V uses fp16 for the first N + last N layers ("boundary protection") and the user's vquant for the middle layers.
# Boundary-protect K-V split: middle layers in fp8, boundary layers in fp16
pc = PionPromptCache(model, vquant="fp8", boundary_protect=2)
Measured on Llama-3.2-1B-Instruct-4bit, 3 prompts × 5 queries (Apple Silicon, single worker):
| Config | TTFT speedup | Throughput speedup | First-token agreement |
|---|---|---|---|
| Uniform fp16 | 2.75× | 2.74× | 14/15 |
vquant=fp8, boundary_protect=2 |
2.88× | 2.87× | 15/15 |
Faster AND more accurate than uniform fp16 on this workload — fp8 V on middle layers carries less wire data per layer; the fp16 boundary layers protect routing.
Trade-off: the SCHEMA path skips KV.PREFIX.REGISTER's cross-worker directory publish (single-worker visibility only). A future KV.PREFIX.REGISTER.SCHEMA server command would lift this.
boundary_protect requires kv_dim divisible by 32 when vquant ∈ {fp8, turbo4, turbo3, turbo2}. Llama-3.2-1B (kv_dim=512) and Llama-3.2-3B (kv_dim=1024) both qualify.
The namespace is the contract¶
The namespace key is the only contract — it must encode every load-bearing piece of execution context. Mismatch is silent corruption, not a runtime error.
ns = PionPromptCache.make_namespace(
model_id, # e.g. "mlx-community/Llama-3.2-1B-Instruct-4bit"
tokenizer_hash, # changes invalidate the cache
rope_theta, # rope scaling settings
quant_format, # "fp16" / "int8" / "turbo4"
adapter_id, # LoRA / adapter identity if any
prompt_text, # the actual prompt
)
make_namespace returns sha256("|".join(parts))[:32].
The namespace is not a secret. Anyone who can reach the server and knows (or guesses) a key can read that prefix, so gate the server with --requirepass, and give each tenant its own credentials with --tenant (see the security model) rather than relying on unguessable keys.
How it works: the cold and warm paths¶
Cold path (first request per prompt):
client → mlx_lm.forward(prefix + suffix) → output
│
└── cache populated locally
client → KV.PREFIX.REGISTER → V.STOREBATCH per layer (K and V) → Pion
Warm path (every subsequent request on the same namespace):
client → KV.PREFIX.LOOKUP → +HIT
client → V.FETCH RANGE per layer → fp32 K and V tensors
client → MLX KVCache.update_and_fetch(K, V) per layer
client → mlx_lm.forward(suffix only, cache=rebuilt) → output
Pion's V-store stores per-token, per-layer values indexed by token ID. K is treated as just another value array — the wire format is the same. Boundary-layer FP16 protection is exposed via PionPromptCache(..., boundary_protect=N) — first/last N layers stay FP16 while middle layers go to the chosen vquant. Reduces drift on int8 by ~33%, on turbo4 by ~42%.
Stage 1 or Stage 2: who computes the attention¶
- Stage 1 (
KV.PREFIX.*+V.FETCH RANGE+PionPromptCache) — the client's inference engine runs attention locally on fetched K/V. Right when the client is on the same host as Pion (Apple Silicon unified memory) or when the inference engine doesn't expose a hook for offloading attention. - Stage 2 (
ATTEND.PREFIX.STORE/QUERY) — Pion's MLX sidecar runs attention on K/V that never leaves Pion's process memory after the initial push. Right when the inference engine accepts an external attention output (e.g., custom vLLM CacheEngine, exo'sgpu_attentionmode), or when many queries share the same K/V across sessions and the wire cost of Q+K+V re-marshaling dominates.
The two are complementary; the same prefix can live in both V-store (for clients that fetch K/V) and the MLX sidecar (for clients that offload attention).
Retrieved chunks, not just system prompts¶
HybridRetrievalCache extends the prompt-prefix cache pattern from "the
system prompt that's identical across requests" to "any retrieved chunk
that's been ingested before." The retrieval is still done by whatever
embedding model the consumer already uses (BGE, MiniLM, OpenAI,
text-embedding-3-small, anything stable). Pion's role is to skip the
prefill of the retrieved chunk by holding its K/V tensors keyed by
chunk_id.
from pion_vllm_mlx import HybridRetrievalCache
from mlx_lm import load
model, tok = load("mlx-community/Llama-3.2-1B-Instruct-4bit")
hr = HybridRetrievalCache(model) # inproc backend (default)
hr.ingest("eiffel_passage", tok.encode("The Eiffel Tower is..."))
cache, suffix = hr.prepare("eiffel_passage", tok.encode("How tall?\nAnswer:"))
# pass `cache` to mlx-lm generate — chunk K/V is already loaded
Backends¶
| Backend | K/V live | Precision | Server | Best for |
|---|---|---|---|---|
inproc (default) |
MLX arrays in a process-local dict | bit-perfect (state-setter pickling) | not required | single-process RAG |
pion |
KV.PREFIX.REGISTER + V.STOREBATCH/V.FETCH BATCH |
fp16 (BLEU ~0.97 inherited from Stage 1) | --kvcache --metal-attention -w 1 |
cross-process / cross-host |
Measured (Llama-3.2-1B-Instruct-4bit, 3 RAG cases)¶
| Backend | Quality | TTFT savings vs text-RAG (mean / range) |
|---|---|---|
| inproc | 100% token agreement | 55% / 37–73% |
| pion | functional answer-match parity | 46% / 29–61% |
Test: pion-vllm-mlx/tests/test_hybrid_retrieval.py.
First experiment: benchmarks/reproducers/stage0_hybrid_kv_injection.py.
Storage cost¶
INT4 K/V per token at single layer:
| Model | K-vec dim per layer | Per-token bytes | Per 256-token chunk |
|---|---|---|---|
| Llama-3.2-1B (16 layers, 8 KV heads × 64) | 512 | ~8 KB | ~2 MB |
| Llama-3-8B-class (32 layers, 8 KV heads × 128) | 1024 | ~32 KB | ~8 MB |
| Llama-3-70B (80 layers, 8 KV heads × 128) | 1024 | ~80 KB | ~20 MB |
Multiplies by N if multiple layers are cached. The hybrid pattern is worth it when the same chunks are retrieved repeatedly (FAQ, knowledge bases, doc search); the storage blowup over a single 768-dim embedding is amortized by the prefill saved per hit.
Why this is a separate API rather than a flag on PionPromptCache¶
Prefix caching's namespace contract (make_namespace(model, tokenizer,
rope_theta, quant, adapter, prompt)) bakes in everything that affects
the prefilled K/V. For RAG, the namespace contract is simpler:
hash(chunk_id) — the chunk text is the only thing that varies, the
model+tokenizer are implicit and stable. Mixing the two surfaces would
either pollute the prefix-cache namespace or hide the chunk semantics.
Separate API keeps each surface narrow and the contracts clear.
Multi-chunk¶
Top-K retrieval returns several chunks. set_shared_stub() registers a shared
prefix stub once, ingest_pack() stores each chunk pack, and prepare_multi()
composes up to max_packs (default 8) packs with exact positional re-rotation.
Encode each chunk and the suffix separately and concatenate tokens, not
strings. Keep the composition coarse: many small packs let distractor chunks
collide, and quality drops well before 20 packs.
Where it fits¶
Concrete fits for this build (single-instance, MLX, Apple Silicon, ≥95% hit rate, prefix-dominated):
- Multi-tenant SaaS with a fixed system prompt. The 9.3× warm-TTFT measurement (5 prompts × 30 queries, 96.7% hit rate) is exactly this shape.
- Local LLM apps on Apple Silicon (Mac/iOS). An embedded engine; chat with a reused system prompt.
- Mac cluster inference (exo, vllm-mlx). Cross-instance verified — multiple Macs share one Pion via TCP.
- RAG with a fixed document corpus, contiguous order. Cache
[system + chunks_in_canonical_order]. Arbitrary chunk recomposition is not safe (causal attention — chunk B's K is rotated for positions it will not occupy). - Few-shot prompts, function-calling, agent loops. Fixed system + tools + examples; user message varies. First-token agreement is what matters most.
- Code assistants with repo context. 6K-token repo prefix + short user query. Win grows with prefix length.
- Prompt-engineering iteration. 50+ test queries against one prompt; warm cycle dominates.
- A/B testing / replay harnesses. Namespace key ensures fresh cache when any input changes.
Where to go next¶
- Where Pion does not help — short prefixes, low hit rates, and the cases where recomputing wins.
- Persistence and durability — why a stored prefix survives a restart.
- Shared KV cache — full reference — every section of the design document, including measurements, status and the test inventory.