Skip to content

ATTEND.* and ATTEND.PREFIX.*

Stage 2 of the shared KV cache: instead of shipping K/V back to the model, Pion computes the attention over the cached prefix itself — natively on Metal with --metal-attention, or through the MLX bridge — and the consumer attaches only the suffix. This is the path behind the same-process TTFT row, the sparse long-context selector and the pion-exo hook.

Wire forms

ATTEND.PREFIX.STORE  <sid> <layer> <H> <N> <D> <K_blob> <V_blob>
ATTEND.PREFIX.LOOKUP <sid> <layer>                                       → +HIT / +MISS
ATTEND.PREFIX.QUERY  <sid> <layer> <H> <D> <top_k> <Q_blob> [<fa_window>] → bulk H*D fp32
ATTEND.PREFIX.QUERY_FUSED <sid> <layer> <H_q> <D> <S_suf> <H_kv>
                          <Q> <K_suf> <V_suf> <head_map> [<fa_window>]   → bulk H_q*M*D fp32 (suffix+merge fused)
ATTEND.PREFIX.QUERY_SPARSE <sid> <layer> <H> <D> <K_sparse_max>
                           <Q> <indices> <counts> [<fa_window>]          → bulk H*D fp32 (caller-supplied indices)
ATTEND.PREFIX.QUERY_SPARSE_AUTO <sid> <layer> <H_q> <D> <B> <K_top>
                                <H_kv> <Q> <head_map> [<fa_window>]      → bulk H_q*D fp32 (server picks indices via block-mean top-K)
ATTEND.PREFIX.QUERY_SPARSE_AUTO_FUSED
       <sid> <layer> <H_q> <D> <B> <K_top> <H_kv> <S_suf>
       <Q> <K_suf> <V_suf> <head_map> [<fa_window>]                       → bulk H_q*D fp32 (sparse-prefix + dense-suffix + merge, single dispatch)

Native Metal SDPA via src/ffi/metal_compute.metal + src/ffi/metal_wrap.m, selected by --metal-attention (or --metal-attention-fp16 for vanilla mlx-lm precision parity). No Python, no Unix socket. D ∈ {32, 64, 96, 128, 160, 192, 256, 512} (D=512 is for Gemma 4 full-attention layers; dynamic threadgroup memory scales s_o per-PSO). Six kernels per D-PSO: sdpa_q1_fp32/fp16, sdpa_batched_q_fp32/fp16, sdpa_batched_q_fused_fp32/fp16, sdpa_q1_sparse_fp32/fp16, sdpa_q1_sparse_fused_fp32/fp16. Multi-worker (per-worker session caches with linear probing + tombstones). End-to-end M=1 ATTEND.PREFIX.QUERY median 0.441 ms at H=8/N=2048/d=128, faster than MLX raw compute (0.489 ms). Bit-equivalent to vanilla mlx-lm on tests/test_mlx_lm_patch.py (20/20 token agreement).

Sparse-mask path: server picks block-mean top-K from resident K/V (block size B, K_top blocks), runs sparse SDPA over the picked indices. Optional fused variant also merges a caller-supplied dense suffix in the same dispatch — the "wire-mode sparse" consumer in pion-vllm-mlx/pion_vllm_mlx/mlx_lm_patch.py uses it. 100% NIAH at 64K on Gemma-4-E2B-4bit (in-proc lane, sparse on full layers, K_block=64 K_blocks=8, 326× warm TTFT vs vanilla). Wire-lane consumer end-to-end on Llama-3.2-1B (GQA): same magic-number answers as vanilla. Validation gates: tests/test_attend_sparse_kernel.py, tests/test_attend_d512.py, tests/test_attend_sparse_auto.py, tests/test_attend_sparse_auto_fused.py, tests/test_wire_sparse_consumer.py.

The batched-Q form of ATTEND.PREFIX.QUERY (Q shape (H, M, D), M derived from blob length) returns [output: H*M*D float32 | LSE: H*M float32] — the LSE trailer is required by online-softmax merge. M=1 callers see no wire-format change (no LSE trailer). See tests/test_attend_prefix_lse.py for the end-to-end verification.

For decode (M=1), the per-layer round-trip dominates. For TTFT (M=large), batched-Q is 22.9 ms wire vs 754 ms unbatched at Llama-3.2-3B shape — 32.9× faster. The monkey-patch uses M=1 in the simple case shipped here; TTFT-batched M>1 is the optimization that closes the small-N regime.

When to use Stage 1 vs Stage 2

  • Stage 1 (KV.PREFIX.* + V.FETCH RANGE + PionPromptCache) — the client's inference engine runs attention locally on fetched K/V. Right when the client is on the same host as Pion (Apple Silicon unified memory) or when the inference engine doesn't expose a hook for offloading attention.
  • Stage 2 (ATTEND.PREFIX.STORE/QUERY) — Pion's MLX sidecar runs attention on K/V that never leaves Pion's process memory after the initial push. Right when the inference engine accepts an external attention output (e.g., custom vLLM CacheEngine, exo's gpu_attention mode), or when many queries share the same K/V across sessions and the wire cost of Q+K+V re-marshaling dominates.

The two are complementary; the same prefix can live in both V-store (for clients that fetch K/V) and the MLX sidecar (for clients that offload attention).

Externalized attention commands

RESP commands on port 1974:

Command Pion path GLIDE Description
KV.STORE ⚠️ SLOW ❌ Store KV cache tensor with HNSW-indexed embedding key. Experimental — TTL/MODEL stubs, capacity-capped 1000 entries, no LRU. See the external-parametric-memory Stage-2 reframe results §5.
KV.FETCH ⚠️ SLOW ❌ Fetch nearest cached tensor by cosine similarity. Experimental — MODEL arg silently ignored on FETCH.
KV.INFO SLOW ❌ KV cache store statistics (entries, blob bytes, hits, misses, capacity)
~~KV.EVICT~~ none ❌ Not implemented — documented in source comments only, no handler.
KV.PREFIX.REGISTER SLOW ❌ Register namespace for shared prefix KV cache; creates <ns>_pk and <ns>_pv V-store sessions. Optional PREFILL_MS <ms> reports the client's measured cold prefill for the value receipt. See doc/shared_kv_cache.md.
KV.PREFIX.LOOKUP SLOW ❌ Returns +HIT if both K and V sessions exist for the namespace, +MISS otherwise
KV.PREFIX.INFO SLOW ❌ Global stats: registered prefix count, total tokens, total fetches
V.CREATE SLOW ❌ Create V-store session for token-ID-indexed value storage (used by KV.PREFIX.* and externalized attention)
V.STOREBATCH SLOW ❌ Append a batch of FP32 values quantized to session format (int8/turbo4/turbo3/turbo2/fp16)
V.FETCH SLOW ❌ Fetch values by token ID list, or contiguous range via RANGE start end (single round-trip per layer; bypasses 64-token RESP frame limit)
KV.PREFIX.BLOCKS SLOW ❌ Per-block visibility for a stored prefix
KV.PREFIX.MEMBERSHIP SLOW ❌ Which blocks of a prefix are resident
KV.PREFIX.OWNER SLOW ❌ Which worker holds a prefix. -w N is N independent keyspaces, so this is how a client finds the right one
KV.PREFIX.WARM SLOW ❌ Promote a cold-tier prefix back to RAM
KV.PREFIX.SAVE SLOW ❌ Persist a prefix to the cold tier
V.SNAPSHOT SLOW ❌ Point-in-time V-store snapshot
V.COMMIT SLOW ❌ Seal a snapshot
V.RESTORE SLOW ❌ Reload a sealed snapshot
V.INFO SLOW ❌ Per-session or global V-store statistics
ATTEND.PREFIX.STORE SLOW ❌ Push K/V to native Metal SDPA session cache; resident until DROP / LRU eviction. Body: <session_id> <layer_id> <H> <N> <D> <K_blob> <V_blob>. See doc/shared_kv_cache.md Stage 2.
ATTEND.PREFIX.LOOKUP SLOW ❌ Probe the Metal session cache for (sid, layer_id). Returns +HIT if slot exists with non-nil K/V, +MISS otherwise. Symmetric to KV.PREFIX.LOOKUP for V-store state. Required for stage-2-aware client PionPromptCache.lookup to avoid stale V-store hits causing skipped _stage2_push_cold.
ATTEND.PREFIX.QUERY SLOW ❌ Run M=1 single-query attention on cached K/V (Q-only on wire). Body: <session_id> <layer_id> <H> <D> <top_k> <Q_blob>. Returns H*D float32. Native Metal SDPA path (sdpa_q1_fp32/fp16); supports D ∈ {32,64,96,128,160,192,256,512}. 146× faster than ATTEND.QUERYBATCH at H=8 N=2048.
ATTEND.PREFIX.QUERY_FUSED SLOW ❌ M>=1 fused suffix-SDPA + prefix-merge in one dispatch. Body: <sid> <layer> <H_q> <D> <S_suf> <H_kv> <Q> <K_suf> <V_suf> <head_map> [<fa_window>]. GQA-aware. Returns merged attention H_q*M*D*4 bytes; no LSE trailer (merge fused).
ATTEND.PREFIX.QUERY_SPARSE SLOW ❌ Sparse-mask M=1 attention with caller-supplied per-head indices. Body: <sid> <layer> <H> <D> <K_sparse_max> <Q> <indices> <counts> [<fa_window>]. For learned-router v2 consumers. Output H*D*4 bytes.
ATTEND.PREFIX.QUERY_SPARSE_AUTO SLOW ❌ Sparse-mask M=1 with server-side block-mean top-K selection. Body: <sid> <layer> <H_q> <D> <B> <K_top> <H_kv> <Q> <head_map> [<fa_window>]. K_mean cached server-side per slot. GQA-aware. Output H_q*D*4 bytes.
ATTEND.PREFIX.QUERY_SPARSE_AUTO_FUSED SLOW ❌ Sparse-AUTO + dense-suffix + online-softmax merge in one fused dispatch. Body: <sid> <layer> <H_q> <D> <B> <K_top> <H_kv> <S_suf> <Q> <K_suf> <V_suf> <head_map> [<fa_window>]. Closes the "ignore suffix" v1 wire-mode-sparse caveat. Output H_q*D*4 bytes.
SSM.PREFIX.STORE SLOW ❌ Store opaque byte blob (serialized SSM recurrent state — Mamba [conv_state, ssm_state], RWKV-7 size=3 tuple, or any future linear-attention family). Body: <session_id> <layer_id> <state_blob>. Server never parses the blob.
SSM.PREFIX.FETCH SLOW ❌ Retrieve previously-stored blob for (sid, layer_id). Returns bulk string or $-1 (miss).
SSM.PREFIX.DROP SLOW ❌ Drop stored blob. Body: <session_id> [<layer_id>]. Layer omitted = drop all layers for the session. Idempotent.
ATTEND.CREATE SLOW ❌ Create attention session with key_dim and value_dim
ATTEND.STORE SLOW ❌ Stage token KV pairs (FP32 memcpy); 3.67M tok/s via binary protocol
ATTEND.FINALIZE SLOW ❌ Batch build HNSW index from staged keys
ATTEND.QUERY SLOW ❌ Top-k HNSW search over attention keys; 86us per layer at 128K tokens
ATTEND.INFO SLOW ❌ Index statistics (sessions, tokens, queries)

Binary protocol commands on port 1975 (0xCA5E framing):

Command byte Command Description
0x01 LAYER.STORE Store per-layer tensor (binary blob)
0x02 LAYER.FETCH Fetch per-layer tensor
0x03 PING Binary keepalive
0x20 ATTEND.CREATE Create attention session (binary)
0x21 ATTEND.STORE Stage token KV pairs (3.67M tok/s)
0x22 ATTEND.FINALIZE Batch build HNSW from staged keys
0x23 ATTEND.QUERY Top-k HNSW query (86us at 128K)

Requires --kvcache flag. Python client: vllm-pion/ package (PionKVClient, PionAttentionClient, ExternalizedAttentionLayer).


Fixed-size state cache — STATE.*

A distinct surface from V-store: a fixed-size, overwrite-in-place buffer for per-request recurrent state, not an evicting cache. Every command answers -ERR state cache not enabled (use --kvcache) when the flag is absent.

Command Pion path Notes
STATE.ALLOC <id> <bytes> SLOW Reserve a fixed-size slot
STATE.WRITE <id> <offset> <data> SLOW Overwrite in place; no growth
STATE.READ <id> <offset> <len> SLOW
STATE.FREE <id> SLOW
STATE.INFO SLOW Slot count and occupancy