ATTEND.* and ATTEND.PREFIX.*¶
Stage 2 of the shared KV cache: instead of shipping K/V back to the model,
Pion computes the attention over the cached prefix itself — natively on Metal
with --metal-attention, or through the MLX bridge — and the consumer
attaches only the suffix. This is the path behind the same-process TTFT row,
the sparse long-context selector and the pion-exo hook.
Wire forms¶
ATTEND.PREFIX.STORE <sid> <layer> <H> <N> <D> <K_blob> <V_blob>
ATTEND.PREFIX.LOOKUP <sid> <layer> → +HIT / +MISS
ATTEND.PREFIX.QUERY <sid> <layer> <H> <D> <top_k> <Q_blob> [<fa_window>] → bulk H*D fp32
ATTEND.PREFIX.QUERY_FUSED <sid> <layer> <H_q> <D> <S_suf> <H_kv>
<Q> <K_suf> <V_suf> <head_map> [<fa_window>] → bulk H_q*M*D fp32 (suffix+merge fused)
ATTEND.PREFIX.QUERY_SPARSE <sid> <layer> <H> <D> <K_sparse_max>
<Q> <indices> <counts> [<fa_window>] → bulk H*D fp32 (caller-supplied indices)
ATTEND.PREFIX.QUERY_SPARSE_AUTO <sid> <layer> <H_q> <D> <B> <K_top>
<H_kv> <Q> <head_map> [<fa_window>] → bulk H_q*D fp32 (server picks indices via block-mean top-K)
ATTEND.PREFIX.QUERY_SPARSE_AUTO_FUSED
<sid> <layer> <H_q> <D> <B> <K_top> <H_kv> <S_suf>
<Q> <K_suf> <V_suf> <head_map> [<fa_window>] → bulk H_q*D fp32 (sparse-prefix + dense-suffix + merge, single dispatch)
Native Metal SDPA via src/ffi/metal_compute.metal + src/ffi/metal_wrap.m, selected by --metal-attention (or --metal-attention-fp16 for vanilla mlx-lm precision parity). No Python, no Unix socket. D ∈ {32, 64, 96, 128, 160, 192, 256, 512} (D=512 is for Gemma 4 full-attention layers; dynamic threadgroup memory scales s_o per-PSO). Six kernels per D-PSO: sdpa_q1_fp32/fp16, sdpa_batched_q_fp32/fp16, sdpa_batched_q_fused_fp32/fp16, sdpa_q1_sparse_fp32/fp16, sdpa_q1_sparse_fused_fp32/fp16. Multi-worker (per-worker session caches with linear probing + tombstones). End-to-end M=1 ATTEND.PREFIX.QUERY median 0.441 ms at H=8/N=2048/d=128, faster than MLX raw compute (0.489 ms). Bit-equivalent to vanilla mlx-lm on tests/test_mlx_lm_patch.py (20/20 token agreement).
Sparse-mask path: server picks block-mean top-K from resident K/V (block size B, K_top blocks), runs sparse SDPA over the picked indices. Optional fused variant also merges a caller-supplied dense suffix in the same dispatch — the "wire-mode sparse" consumer in pion-vllm-mlx/pion_vllm_mlx/mlx_lm_patch.py uses it. 100% NIAH at 64K on Gemma-4-E2B-4bit (in-proc lane, sparse on full layers, K_block=64 K_blocks=8, 326× warm TTFT vs vanilla). Wire-lane consumer end-to-end on Llama-3.2-1B (GQA): same magic-number answers as vanilla. Validation gates: tests/test_attend_sparse_kernel.py, tests/test_attend_d512.py, tests/test_attend_sparse_auto.py, tests/test_attend_sparse_auto_fused.py, tests/test_wire_sparse_consumer.py.
The batched-Q form of ATTEND.PREFIX.QUERY (Q shape (H, M, D), M derived
from blob length) returns [output: H*M*D float32 | LSE: H*M float32] —
the LSE trailer is required by online-softmax merge. M=1 callers see no
wire-format change (no LSE trailer). See tests/test_attend_prefix_lse.py
for the end-to-end verification.
For decode (M=1), the per-layer round-trip dominates. For TTFT (M=large), batched-Q is 22.9 ms wire vs 754 ms unbatched at Llama-3.2-3B shape — 32.9× faster. The monkey-patch uses M=1 in the simple case shipped here; TTFT-batched M>1 is the optimization that closes the small-N regime.
When to use Stage 1 vs Stage 2¶
- Stage 1 (
KV.PREFIX.*+V.FETCH RANGE+PionPromptCache) — the client's inference engine runs attention locally on fetched K/V. Right when the client is on the same host as Pion (Apple Silicon unified memory) or when the inference engine doesn't expose a hook for offloading attention. - Stage 2 (
ATTEND.PREFIX.STORE/QUERY) — Pion's MLX sidecar runs attention on K/V that never leaves Pion's process memory after the initial push. Right when the inference engine accepts an external attention output (e.g., custom vLLM CacheEngine, exo'sgpu_attentionmode), or when many queries share the same K/V across sessions and the wire cost of Q+K+V re-marshaling dominates.
The two are complementary; the same prefix can live in both V-store (for clients that fetch K/V) and the MLX sidecar (for clients that offload attention).
Externalized attention commands¶
RESP commands on port 1974:
| Command | Pion path | GLIDE | Description |
|---|---|---|---|
| KV.STORE ⚠️ | SLOW | ❌ | Store KV cache tensor with HNSW-indexed embedding key. Experimental — TTL/MODEL stubs, capacity-capped 1000 entries, no LRU. See the external-parametric-memory Stage-2 reframe results §5. |
| KV.FETCH ⚠️ | SLOW | ❌ | Fetch nearest cached tensor by cosine similarity. Experimental — MODEL arg silently ignored on FETCH. |
| KV.INFO | SLOW | ❌ | KV cache store statistics (entries, blob bytes, hits, misses, capacity) |
| ~~KV.EVICT~~ | none | ❌ | Not implemented — documented in source comments only, no handler. |
| KV.PREFIX.REGISTER | SLOW | ❌ | Register namespace for shared prefix KV cache; creates <ns>_pk and <ns>_pv V-store sessions. Optional PREFILL_MS <ms> reports the client's measured cold prefill for the value receipt. See doc/shared_kv_cache.md. |
| KV.PREFIX.LOOKUP | SLOW | ❌ | Returns +HIT if both K and V sessions exist for the namespace, +MISS otherwise |
| KV.PREFIX.INFO | SLOW | ❌ | Global stats: registered prefix count, total tokens, total fetches |
| V.CREATE | SLOW | ❌ | Create V-store session for token-ID-indexed value storage (used by KV.PREFIX.* and externalized attention) |
| V.STOREBATCH | SLOW | ❌ | Append a batch of FP32 values quantized to session format (int8/turbo4/turbo3/turbo2/fp16) |
| V.FETCH | SLOW | ❌ | Fetch values by token ID list, or contiguous range via RANGE start end (single round-trip per layer; bypasses 64-token RESP frame limit) |
| KV.PREFIX.BLOCKS | SLOW | ❌ | Per-block visibility for a stored prefix |
| KV.PREFIX.MEMBERSHIP | SLOW | ❌ | Which blocks of a prefix are resident |
| KV.PREFIX.OWNER | SLOW | ❌ | Which worker holds a prefix. -w N is N independent keyspaces, so this is how a client finds the right one |
| KV.PREFIX.WARM | SLOW | ❌ | Promote a cold-tier prefix back to RAM |
| KV.PREFIX.SAVE | SLOW | ❌ | Persist a prefix to the cold tier |
| V.SNAPSHOT | SLOW | ❌ | Point-in-time V-store snapshot |
| V.COMMIT | SLOW | ❌ | Seal a snapshot |
| V.RESTORE | SLOW | ❌ | Reload a sealed snapshot |
| V.INFO | SLOW | ❌ | Per-session or global V-store statistics |
| ATTEND.PREFIX.STORE | SLOW | ❌ | Push K/V to native Metal SDPA session cache; resident until DROP / LRU eviction. Body: <session_id> <layer_id> <H> <N> <D> <K_blob> <V_blob>. See doc/shared_kv_cache.md Stage 2. |
| ATTEND.PREFIX.LOOKUP | SLOW | ❌ | Probe the Metal session cache for (sid, layer_id). Returns +HIT if slot exists with non-nil K/V, +MISS otherwise. Symmetric to KV.PREFIX.LOOKUP for V-store state. Required for stage-2-aware client PionPromptCache.lookup to avoid stale V-store hits causing skipped _stage2_push_cold. |
| ATTEND.PREFIX.QUERY | SLOW | ❌ | Run M=1 single-query attention on cached K/V (Q-only on wire). Body: <session_id> <layer_id> <H> <D> <top_k> <Q_blob>. Returns H*D float32. Native Metal SDPA path (sdpa_q1_fp32/fp16); supports D ∈ {32,64,96,128,160,192,256,512}. 146× faster than ATTEND.QUERYBATCH at H=8 N=2048. |
| ATTEND.PREFIX.QUERY_FUSED | SLOW | ❌ | M>=1 fused suffix-SDPA + prefix-merge in one dispatch. Body: <sid> <layer> <H_q> <D> <S_suf> <H_kv> <Q> <K_suf> <V_suf> <head_map> [<fa_window>]. GQA-aware. Returns merged attention H_q*M*D*4 bytes; no LSE trailer (merge fused). |
| ATTEND.PREFIX.QUERY_SPARSE | SLOW | ❌ | Sparse-mask M=1 attention with caller-supplied per-head indices. Body: <sid> <layer> <H> <D> <K_sparse_max> <Q> <indices> <counts> [<fa_window>]. For learned-router v2 consumers. Output H*D*4 bytes. |
| ATTEND.PREFIX.QUERY_SPARSE_AUTO | SLOW | ❌ | Sparse-mask M=1 with server-side block-mean top-K selection. Body: <sid> <layer> <H_q> <D> <B> <K_top> <H_kv> <Q> <head_map> [<fa_window>]. K_mean cached server-side per slot. GQA-aware. Output H_q*D*4 bytes. |
| ATTEND.PREFIX.QUERY_SPARSE_AUTO_FUSED | SLOW | ❌ | Sparse-AUTO + dense-suffix + online-softmax merge in one fused dispatch. Body: <sid> <layer> <H_q> <D> <B> <K_top> <H_kv> <S_suf> <Q> <K_suf> <V_suf> <head_map> [<fa_window>]. Closes the "ignore suffix" v1 wire-mode-sparse caveat. Output H_q*D*4 bytes. |
| SSM.PREFIX.STORE | SLOW | ❌ | Store opaque byte blob (serialized SSM recurrent state — Mamba [conv_state, ssm_state], RWKV-7 size=3 tuple, or any future linear-attention family). Body: <session_id> <layer_id> <state_blob>. Server never parses the blob. |
| SSM.PREFIX.FETCH | SLOW | ❌ | Retrieve previously-stored blob for (sid, layer_id). Returns bulk string or $-1 (miss). |
| SSM.PREFIX.DROP | SLOW | ❌ | Drop stored blob. Body: <session_id> [<layer_id>]. Layer omitted = drop all layers for the session. Idempotent. |
| ATTEND.CREATE | SLOW | ❌ | Create attention session with key_dim and value_dim |
| ATTEND.STORE | SLOW | ❌ | Stage token KV pairs (FP32 memcpy); 3.67M tok/s via binary protocol |
| ATTEND.FINALIZE | SLOW | ❌ | Batch build HNSW index from staged keys |
| ATTEND.QUERY | SLOW | ❌ | Top-k HNSW search over attention keys; 86us per layer at 128K tokens |
| ATTEND.INFO | SLOW | ❌ | Index statistics (sessions, tokens, queries) |
Binary protocol commands on port 1975 (0xCA5E framing):
| Command byte | Command | Description |
|---|---|---|
| 0x01 | LAYER.STORE | Store per-layer tensor (binary blob) |
| 0x02 | LAYER.FETCH | Fetch per-layer tensor |
| 0x03 | PING | Binary keepalive |
| 0x20 | ATTEND.CREATE | Create attention session (binary) |
| 0x21 | ATTEND.STORE | Stage token KV pairs (3.67M tok/s) |
| 0x22 | ATTEND.FINALIZE | Batch build HNSW from staged keys |
| 0x23 | ATTEND.QUERY | Top-k HNSW query (86us at 128K) |
Requires --kvcache flag. Python client: vllm-pion/ package (PionKVClient, PionAttentionClient, ExternalizedAttentionLayer).
Fixed-size state cache — STATE.*¶
A distinct surface from V-store: a fixed-size, overwrite-in-place buffer for
per-request recurrent state, not an evicting cache. Every command answers
-ERR state cache not enabled (use --kvcache) when the flag is absent.
| Command | Pion path | Notes |
|---|---|---|
STATE.ALLOC <id> <bytes> |
SLOW | Reserve a fixed-size slot |
STATE.WRITE <id> <offset> <data> |
SLOW | Overwrite in place; no growth |
STATE.READ <id> <offset> <len> |
SLOW | |
STATE.FREE <id> |
SLOW | |
| STATE.INFO | SLOW | Slot count and occupancy |