Skip to content

Shared KV cache

Every request that starts with the same tokens — a system prompt, a few-shot block, a retrieved passage — makes the model compute the same K/V tensors for them again. Pion keeps those tensors after the first time, keyed by a namespace that names the exact prefix, so the next request fetches them instead of recomputing: from the same process, from another process, or after a restart.

This page is the concept: what a namespace must contain, how the cold and warm paths work, and when Pion should compute the attention itself. The commands are in the reference — KV.PREFIX.*, V.*, ATTEND.* — and the Python API is pion-vllm-mlx.

Quick start

Start the server:

./pion-server --kvcache -w 1

Then, from Python (install from a checkout with pip install -e 'pion-vllm-mlx/[mlx]' — not on PyPI yet):

from mlx_lm import load, generate
from pion_vllm_mlx import PionPromptCache

model, tok = load("mlx-community/Llama-3.2-1B-Instruct-4bit")
pc = PionPromptCache(model, vquant="fp16")

system = "You are a support agent for Acme. Answer in one sentence."   # shared by every request
prefix_ids = tok.encode(system)
ns = PionPromptCache.make_namespace("llama-3.2-1b-4bit", "fp16", system)

# The first call anywhere prefills locally, registers with Pion and stores the K/V.
# Every later call — this process, another one, or after a restart — fetches it.
cache = pc.get_or_prefill(prefix_ids, namespace=ns)
print(generate(model, tok, prompt=tok.encode(" How do I reset my password?", add_special_tokens=False),
               prompt_cache=cache, max_tokens=40))

pc.stats() reports hits, misses, hit_rate, fetch_ms_total, store_ms_total.

Heterogeneous KV cache

PionPromptCache(boundary_protect=N) switches the per-prefix register from the legacy uniform KV.PREFIX.REGISTER to two raw V.CREATE … SCHEMA calls (one per K/V side). K stays fp16 across all layers — K drives softmax routing and tolerates quant noise poorly. V uses fp16 for the first N + last N layers ("boundary protection") and the user's vquant for the middle layers.

# Boundary-protect K-V split: middle layers in fp8, boundary layers in fp16
pc = PionPromptCache(model, vquant="fp8", boundary_protect=2)

Measured on Llama-3.2-1B-Instruct-4bit, 3 prompts × 5 queries (Apple Silicon, single worker):

Config TTFT speedup Throughput speedup First-token agreement
Uniform fp16 2.75× 2.74× 14/15
vquant=fp8, boundary_protect=2 2.88× 2.87× 15/15

Faster AND more accurate than uniform fp16 on this workload — fp8 V on middle layers carries less wire data per layer; the fp16 boundary layers protect routing.

Trade-off: the SCHEMA path skips KV.PREFIX.REGISTER's cross-worker directory publish (single-worker visibility only). A future KV.PREFIX.REGISTER.SCHEMA server command would lift this.

boundary_protect requires kv_dim divisible by 32 when vquant ∈ {fp8, turbo4, turbo3, turbo2}. Llama-3.2-1B (kv_dim=512) and Llama-3.2-3B (kv_dim=1024) both qualify.


The namespace is the contract

The namespace key is the only contract — it must encode every load-bearing piece of execution context. Mismatch is silent corruption, not a runtime error.

ns = PionPromptCache.make_namespace(
    model_id,           # e.g. "mlx-community/Llama-3.2-1B-Instruct-4bit"
    tokenizer_hash,     # changes invalidate the cache
    rope_theta,         # rope scaling settings
    quant_format,       # "fp16" / "int8" / "turbo4"
    adapter_id,         # LoRA / adapter identity if any
    prompt_text,        # the actual prompt
)

make_namespace returns sha256("|".join(parts))[:32].

The namespace is not a secret. Anyone who can reach the server and knows (or guesses) a key can read that prefix, so gate the server with --requirepass, and give each tenant its own credentials with --tenant (see the security model) rather than relying on unguessable keys.


How it works: the cold and warm paths

Cold path (first request per prompt):
  client → mlx_lm.forward(prefix + suffix) → output
                            │
                            └── cache populated locally
  client → KV.PREFIX.REGISTER → V.STOREBATCH per layer (K and V) → Pion

Warm path (every subsequent request on the same namespace):
  client → KV.PREFIX.LOOKUP → +HIT
  client → V.FETCH RANGE per layer → fp32 K and V tensors
  client → MLX KVCache.update_and_fetch(K, V) per layer
  client → mlx_lm.forward(suffix only, cache=rebuilt) → output

Pion's V-store stores per-token, per-layer values indexed by token ID. K is treated as just another value array — the wire format is the same. Boundary-layer FP16 protection is exposed via PionPromptCache(..., boundary_protect=N) — first/last N layers stay FP16 while middle layers go to the chosen vquant. Reduces drift on int8 by ~33%, on turbo4 by ~42%.


Stage 1 or Stage 2: who computes the attention

  • Stage 1 (KV.PREFIX.* + V.FETCH RANGE + PionPromptCache) — the client's inference engine runs attention locally on fetched K/V. Right when the client is on the same host as Pion (Apple Silicon unified memory) or when the inference engine doesn't expose a hook for offloading attention.
  • Stage 2 (ATTEND.PREFIX.STORE/QUERY) — Pion's MLX sidecar runs attention on K/V that never leaves Pion's process memory after the initial push. Right when the inference engine accepts an external attention output (e.g., custom vLLM CacheEngine, exo's gpu_attention mode), or when many queries share the same K/V across sessions and the wire cost of Q+K+V re-marshaling dominates.

The two are complementary; the same prefix can live in both V-store (for clients that fetch K/V) and the MLX sidecar (for clients that offload attention).

Retrieved chunks, not just system prompts

HybridRetrievalCache extends the prompt-prefix cache pattern from "the system prompt that's identical across requests" to "any retrieved chunk that's been ingested before." The retrieval is still done by whatever embedding model the consumer already uses (BGE, MiniLM, OpenAI, text-embedding-3-small, anything stable). Pion's role is to skip the prefill of the retrieved chunk by holding its K/V tensors keyed by chunk_id.

from pion_vllm_mlx import HybridRetrievalCache
from mlx_lm import load

model, tok = load("mlx-community/Llama-3.2-1B-Instruct-4bit")
hr = HybridRetrievalCache(model)        # inproc backend (default)
hr.ingest("eiffel_passage", tok.encode("The Eiffel Tower is..."))
cache, suffix = hr.prepare("eiffel_passage", tok.encode("How tall?\nAnswer:"))
# pass `cache` to mlx-lm generate — chunk K/V is already loaded

Backends

Backend K/V live Precision Server Best for
inproc (default) MLX arrays in a process-local dict bit-perfect (state-setter pickling) not required single-process RAG
pion KV.PREFIX.REGISTER + V.STOREBATCH/V.FETCH BATCH fp16 (BLEU ~0.97 inherited from Stage 1) --kvcache --metal-attention -w 1 cross-process / cross-host

Measured (Llama-3.2-1B-Instruct-4bit, 3 RAG cases)

Backend Quality TTFT savings vs text-RAG (mean / range)
inproc 100% token agreement 55% / 37–73%
pion functional answer-match parity 46% / 29–61%

Test: pion-vllm-mlx/tests/test_hybrid_retrieval.py. First experiment: benchmarks/reproducers/stage0_hybrid_kv_injection.py.

Storage cost

INT4 K/V per token at single layer:

Model K-vec dim per layer Per-token bytes Per 256-token chunk
Llama-3.2-1B (16 layers, 8 KV heads × 64) 512 ~8 KB ~2 MB
Llama-3-8B-class (32 layers, 8 KV heads × 128) 1024 ~32 KB ~8 MB
Llama-3-70B (80 layers, 8 KV heads × 128) 1024 ~80 KB ~20 MB

Multiplies by N if multiple layers are cached. The hybrid pattern is worth it when the same chunks are retrieved repeatedly (FAQ, knowledge bases, doc search); the storage blowup over a single 768-dim embedding is amortized by the prefill saved per hit.

Why this is a separate API rather than a flag on PionPromptCache

Prefix caching's namespace contract (make_namespace(model, tokenizer, rope_theta, quant, adapter, prompt)) bakes in everything that affects the prefilled K/V. For RAG, the namespace contract is simpler: hash(chunk_id) — the chunk text is the only thing that varies, the model+tokenizer are implicit and stable. Mixing the two surfaces would either pollute the prefix-cache namespace or hide the chunk semantics. Separate API keeps each surface narrow and the contracts clear.

Multi-chunk

Top-K retrieval returns several chunks. set_shared_stub() registers a shared prefix stub once, ingest_pack() stores each chunk pack, and prepare_multi() composes up to max_packs (default 8) packs with exact positional re-rotation. Encode each chunk and the suffix separately and concatenate tokens, not strings. Keep the composition coarse: many small packs let distractor chunks collide, and quality drops well before 20 packs.


Where it fits

Concrete fits for this build (single-instance, MLX, Apple Silicon, ≥95% hit rate, prefix-dominated):

  1. Multi-tenant SaaS with a fixed system prompt. The 9.3× warm-TTFT measurement (5 prompts × 30 queries, 96.7% hit rate) is exactly this shape.
  2. Local LLM apps on Apple Silicon (Mac/iOS). An embedded engine; chat with a reused system prompt.
  3. Mac cluster inference (exo, vllm-mlx). Cross-instance verified — multiple Macs share one Pion via TCP.
  4. RAG with a fixed document corpus, contiguous order. Cache [system + chunks_in_canonical_order]. Arbitrary chunk recomposition is not safe (causal attention — chunk B's K is rotated for positions it will not occupy).
  5. Few-shot prompts, function-calling, agent loops. Fixed system + tools + examples; user message varies. First-token agreement is what matters most.
  6. Code assistants with repo context. 6K-token repo prefix + short user query. Win grows with prefix length.
  7. Prompt-engineering iteration. 50+ test queries against one prompt; warm cycle dominates.
  8. A/B testing / replay harnesses. Namespace key ensures fresh cache when any input changes.

Where to go next