Skip to content

KV.PREFIX.* — the prefix registry

A namespace key names one exact token prefix for one model and quantization; the convention is app|version|model|vquant|prompt-id. Register it once, look it up from any process, and fetch the K/V through the V-store. All of these need pion-server --kvcache. From Python, PionPromptCache issues every command on this page for you.

Three production wrapper commands, plus the underlying V-store path (V.STOREBATCH / V.FETCH ... RANGE).

KV.PREFIX.REGISTER <ns_key> <kv_dim> <vquant> [BLOCKS <block_size> <hash_count> <hash_blob>]

Creates two V-store sessions, <ns_key>_pk (keys) and <ns_key>_pv (values), with the given quantization format. After REGISTER, standard V.STOREBATCH and V.FETCH ... RANGE work against the derived sids.

vquant ∈ {int8, turbo4, turbo3, turbo2, fp16, fp8, mlx4g32} (mlx4g32 = int4 group-32 affine). fp16 is the production default (BLEU 1.0 cross-instance, 0.969 vs standalone).

> KV.PREFIX.REGISTER my_app|v1|llama|fp16|prompt_a 512 fp16
+OK

Optional BLOCKS clause — carry a within-prefix block hash table for cache-aware routers. <block_size> is the token count per block (typically 16 or 64), <hash_count> is the number of u64 hashes, <hash_blob> is hash_count × 8 bytes of little-endian u64 hashes packed in token order. Hashes are stored on the K-side session and survive WAL replay + snapshot reload. Use the new KV.PREFIX.BLOCKS / KV.PREFIX.MEMBERSHIP commands below to query them.

import struct
hashes = [hash_block(...) for block in blocks]
blob = struct.pack("<" + "Q" * len(hashes), *hashes)
resp.call("KV.PREFIX.REGISTER", ns_key, str(kv_dim), "fp16",
          "BLOCKS", "64", str(len(hashes)), blob)

KV.PREFIX.LOOKUP <ns_key> [TOKENS <n>] [PREFILL_MS <ms>] → +HIT or +MISS

Tells the client whether both K and V sessions exist in V-store for this namespace.

> KV.PREFIX.LOOKUP my_app|v1|llama|fp16|prompt_a
+HIT

The options only change what the value receipt (PION.STATS) records for a hit; the answer is the same. TOKENS <n> credits the n tokens the caller actually restored instead of the namespace's own token count. That is for a client whose reuse spans several namespaces and reads rows with V.FETCH: pion-vllm-mlx serve stores a conversation as a chain of segments and names the leaf plus the total, once per restore. PREFILL_MS <ms> is that restore's measured cold prefill, when the caller has one. An unknown option or a missing value is refused with -ERR, and nothing is recorded.

KV.PREFIX.DROP <ns_key> [<ns_key> ...] → :<dropped>

Frees the K and V sessions of each prefix and writes the drop to the V-store WAL, so a restart does not replay the rows back. A namespace that is not held counts 0. This is for clients that run their own eviction policy. pion-vllm-mlx serve --pion-budget-gb keeps its lineages under a byte budget and evicts the least-recently-used leaf segment. The V-store's own LRU (it evicts when its 256 session slots are full) goes by last access alone. For a lineage, that is the root: it is written once and never touched while the leaf grows.

Important — V-store state ≠ Metal session-cache state. KV.PREFIX.LOOKUP reports only V-store registration (Stage-1 path). For Stage-2 wire-mode consumers (sparse-mask / fused-sparse / mlx-lm patch) that read K/V from the Metal SDPA session cache populated by ATTEND.PREFIX.STORE, also probe ATTEND.PREFIX.LOOKUP <sid> <layer_id>. A stale V-store HIT while the Metal cache is cold causes Stage-2 push-cold to skip — checking both is required. PionPromptCache.lookup(namespace) does this automatically when stage2=True.

KV.PREFIX.BLOCKS <ns_key> → bulk string [block_size 4B LE][block_count 4B LE][hashes] or +UNKNOWN

Returns the within-prefix block hash table registered via KV.PREFIX.REGISTER ... BLOCKS. The bulk-string body is 8 + block_count * 8 bytes: a block_size (u32 LE), a block_count (u32 LE), then block_count u64 LE hashes in token order. Used by cache-aware routers (and audit tooling) to reproduce the residency picture client-side, or to feed a KV.PREFIX.MEMBERSHIP probe with the canonical hash set.

Returns the simple string +UNKNOWN\r\n when the namespace exists but no block table is registered, OR when the namespace was never registered on any worker. When the namespace lives on a different worker (cross-worker directory hit), returns -ERR KV.PREFIX.BLOCKS session lives on worker N so the caller can pin its connection — same pattern as KV.PREFIX.WARM.

> KV.PREFIX.BLOCKS my_app|v1|llama|fp16|prompt_a
$1256
<binary: block_size=64, block_count=156, 156 × u64 hashes>

KV.PREFIX.MEMBERSHIP <ns_key> <hash_count> <hash_blob> → bulk-string bitmap or +UNKNOWN

The router supplies the block hashes it's looking for (in any order); the server returns a ceil(hash_count / 8)-byte bitmap, with bit i set iff probe hash i is in the namespace's registered table. Single round-trip, bandwidth-efficient: 100K tokens at block_size=64 → 1,562 blocks → 196-byte response.

Returns +UNKNOWN\r\n when no block table is registered or the namespace doesn't exist. Cross-worker rebound matches KV.PREFIX.BLOCKS.

Server-side compute is O(K log N) via binary search over a sorted parallel copy of the registered hashes (built once at REGISTER time). Measured at 125 µs p50 / 270 µs p99 e2e over loopback for K=N=1,562 on Apple Silicon (server-side compute well under its 100 µs target — most of the latency is loopback RTT for the 12.5 KB request).

probe_blob = struct.pack("<" + "Q" * len(probe_hashes), *probe_hashes)
reply = resp.call("KV.PREFIX.MEMBERSHIP", ns_key, str(len(probe_hashes)), probe_blob)
# Parse bulk-string body as a bitmap; bit i set ↔ probe_hashes[i] is cached.

KV.PREFIX.INFO → bulk string

Global stats: registered prefix count, total tokens, total fetches.

> KV.PREFIX.INFO
$104
registered_prefixes:5
total_prefix_tokens:5800
vstore_sessions:10
vstore_total_fetches:145

Also reported: vstore_bytes, the K/V bytes held across every session in its stored format. It's what a byte budget is measured against and what a snapshot writes; resident memory can be up to about 2× that, because buffers grow by doubling. wal_bytes is the on-disk size of the V-store WAL, which only grows until KV.PREFIX.SAVE snapshots and truncates it.

KV.PREFIX.SAVE [path] → +OK

Snapshots every V-store session to pion.vstore.<worker> (or a relative path), then truncates the V-store WAL. The snapshot is written to <path>.tmp, checked write by write, flushed to stable storage (F_FULLFSYNC on macOS) and renamed into place. The WAL is truncated only after all of that succeeds. A crash, a full disk or a power cut mid-save leaves the previous snapshot and the WAL as they were.

V.FETCH <session_id> <layer_id> RANGE <start_id> <end_id>

Single round-trip per layer-side, regardless of prefix length. Sidesteps the 64-token RESP frame limit that the legacy id-list form hits at ~60-token prefixes. Returns concatenated dequantized FP32 values.

KV.PREFIX.WARM <ns_key> <H> <D> [<attend_sid>] → +<N>

Server-side rehydrate of the Metal SDPA session cache from V-store. Walks both <ns>_pk and <ns>_pv sessions, dequantizes per layer to fp32, transposes from V-store layout [N, H*D] into ATTEND.PREFIX.STORE layout [H, N, D], and pushes each layer back via pion_metal_sdpa_store_kv. Returns the number of layers rehydrated as +N\r\n (or +N skipped_dim=K\r\n when some layers had a kv_dim ≠ H*D and were skipped).

The optional 4th arg overrides the ATTEND-side session id; default = <ns_key>. The production consumer (PionPromptCache._attend_session) stores ATTEND state under <namespace>_attn and passes that here.

Call this after ATTEND.PREFIX.QUERY returns -COLDMISS ..., or proactively before a hot batch of queries against a namespace that may have been evicted. Cost ≈ 50–200 ms depending on prefix length (one V-store dequant + one CPU transpose + one Metal copy per layer).

> KV.PREFIX.WARM my_app|v1|llama|fp16|prompt_a 8 64 my_app|v1|llama|fp16|prompt_a_attn
+16

KV.PREFIX.INFO reports four cold-tier telemetry fields:

sdpa_warm:256        ← live slots in the Metal SDPA session cache
sdpa_cold:1          ← demoted entries in the cold registry
sdpa_demotions:1     ← lifetime WARM→COLD transitions on this worker
sdpa_rehydrates:1    ← lifetime COLD→WARM transitions on this worker

Cold-tier state machine

Three states per (session_id, layer_id):

State Where it lives Wire signal Cost to query
WARM Metal SDPA slot (live K_buf/V_buf) +HIT from ATTEND.PREFIX.LOOKUP sub-ms (M=1 ~0.4 ms)
COLD Per-worker cold registry (metadata only); K/V on disk in V-store +COLD from ATTEND.PREFIX.LOOKUP; -COLDMISS … from QUERY* 50–200 ms after WARM, then sub-ms
MISSING Nowhere +MISS from LOOKUP; -ERR session not found from QUERY* full cold prefill needed

Eviction triggers when the per-worker WARM slot table (256 slots) fills. The LRU slot is picked by last_access_ns (mach_absolute_time), its Metal K_buf/V_buf are released, and (key, H, N, D, last_access_ns) move to the cold registry (SDPA_COLD_SLOTS=1024 per worker). A subsequent ATTEND.PREFIX.QUERY short-circuits via session_state(...)==2 and returns -COLDMISS …. The consumer issues KV.PREFIX.WARM and retries; the warm path stamps last_access_ns so the rehydrated session is now the most-recent, not the next eviction victim.

The PionPromptCache client handles this transparently — attend_query / attend_query_fused / attend_query_sparse_auto* detect -COLDMISS (RESP lane) or STATUS_COLDMISS=0x03 (binary fast lane), call _warm_namespace(...), and retry once. pcache.cold_rehydrates_observed exposes the count for telemetry.


The value receipt — PION.STATS

The server keeps a per-worker ledger of what the prefix cache actually did for you:

PION.STATS            → map of 16 fields (RESP3 %-map; RESP2 flat array)
PION.STATS RESET      → +OK, counters zeroed (uptime is not a counter)
INFO                  → the same numbers under a `# Pion` section
Field Meaning
kvprefix_hits / kvprefix_misses KV.PREFIX.LOOKUP answers
kvprefix_tokens_served prefix tokens whose prefill was skipped (K-side layer-0 token count at hit time, or the TOKENS the client named on KV.PREFIX.LOOKUP)
kvprefix_bytes_served V.FETCH payload bytes delivered
prefill_seconds_avoided cumulative prefill time skipped, measured + estimated
prefill_seconds_avoided_measured the part backed by client-reported PREFILL_MS only
kvprefix_hits_measured hits credited with a reported time
semantic_hits / semantic_misses AI.SEMANTIC_CACHE GET and every other cache_get caller
moe_hits / moe_misses MOE.EXPERT.FETCH tier hits
vector_queries FT.SEARCH + FT.HYBRID answered

Two kinds of number, kept apart on purpose. PionPromptCache times its local cold prefill and sends it as KV.PREFIX.REGISTER ... PREFILL_MS <ms>; every later hit on that prefix is credited exactly that time — a receipt, not an estimate. A prefix registered without PREFILL_MS (a hand-rolled client, or a server restart, since the reported time is deliberately not persisted) is credited tokens × 555 µs, the per-token cost measured on Llama-3.2-1B-Instruct-4bit on an M-series Mac (2,022 tokens: 1,218 ms cold, 95 ms warm). Larger models cost more per token, so the estimate is conservative for them, and it is labelled an estimate wherever it appears.

The counters are per worker — they live in the worker that answered — so a connection pool spanning -w N workers reads each worker's own receipt. Formatting happens on demand in the reply; recording is a counter increment on an already-slow-path command, so nothing on the KV fast path changed.

  • Shared KV cache — the concept, Stage 1 vs Stage 2, namespaces, durability.
  • V.* — V.STOREBATCH and V.FETCH … BATCH, the commands that move the bytes.
  • ATTEND.* — when Pion runs the attention itself.