SSM.PREFIX.* — recurrent-state companion¶
Hybrid Mamba / GatedDeltaNet / RWKV models carry state that is not K/V.
SSM.PREFIX.* stores that state under the same namespace contract as
KV.PREFIX.*, so a hybrid model's whole prefix — softmax layers via the
V-store, linear layers here — is shared across processes and survives a
restart.
The KV.PREFIX. substrate above handles transformer-attention layers. For hybrid Mamba+Transformer models (LoLCATs, MOHAWK, Mamba-in-Llama recipes; or any future external small hybrid), the SSM half of the model has per-layer recurrent state, not per-token K/V. SSM.PREFIX. is the companion substrate for that state.
SSM.PREFIX.STORE <sid> <layer> <state_blob> → +OK
SSM.PREFIX.FETCH <sid> <layer> → bulk string | $-1
SSM.PREFIX.DROP <sid> [<layer>] → +OK (layer omitted = drop all layers for the session)
state_blob is opaque to the server — pure host-side byte storage. The consumer chooses serialization. The reference format (see tests/test_ssm_prefix_roundtrip.py):
uint32 version = 1
uint32 n_arrays
for each array:
uint32 ndim
uint32[ndim] shape
uint32 dtype_code (0=fp32, 1=fp16, 2=bf16 — serialized as fp32, lossless)
raw bytes
This keeps the substrate model-family-agnostic: Mamba (size=2 [conv_state, ssm_state]), RWKV-7 (size=3), and future RetNet / Hedgehog / GLA all serialize differently but the server never parses. Adding a new family is ~1-2 days of consumer-side (de)serializer work.
Validation:
- Drift check (reproduced by tests/test_ssm_prefix_roundtrip.py): bit-perfect state hydration across Mamba-130M-f32 (2048 decode tokens), Mamba-370M-f16 (1024 decode tokens), and RWKV-7 168M (512 decode tokens). All show token agreement 100%, max-abs-diff = 0.000e+00 vs no-snapshot baseline. Deterministic-recurrence property holds.
- Wire round-trip (tests/test_ssm_prefix_roundtrip.py): end-to-end through pion-server. 64/64 bit-perfect Mamba-130M, 64/64 bit-perfect RWKV-7. 1 MB random blob round-trip + overwrite + multi-layer drop all OK.
Pickup cost for new families: each new family needs its own (de)serializer in tests/test_ssm_prefix_roundtrip.py's serialize_arrays_cache / deserialize_arrays_cache. The RWKV-7 serializer is the reference implementation.
End-to-end hybrid model. mlx-community/Qwen3.5-4B-MLX-4bit (24 GatedDeltaNet + 8 Qwen3NextAttention, 3:1 ratio, Apache 2.0), a publicly released hybrid, with measured wins:
| Prefix length | Vanilla cold | Pion warm | Speedup | Token agreement |
|---|---|---|---|---|
| 174 tokens | — | — | 3.3× | — |
| 2,048 tokens | 6,080 ms | 246 ms | 24.77× | 18/18 |
| 4,096 tokens | 12,757 ms | 390 ms | 32.7× | 16/18 |
| 8,192 tokens | 27,598 ms | 2,373 ms | 11.6× | 18/18 |
The speedup peaks at 4K. Warm TTFT is roughly constant up to 4K (246 → 390 ms) and then jumps to 2,373 ms at 8K, because the wire-fetch term grows with shipped bytes (max layer 33.55 MB at L=8192) while vanilla prefill grows only linearly — so 8K measures 11.6×, below 4K's 32.7×. Token agreement is 18/18 at 2K and 8K and 16/18 at 4K (greedy-argmax non-determinism present in both the vanilla and Pion paths).
18/18 layers bit-perfect across the cleanly-typed split path (24 GatedDeltaNet via SSM.PREFIX.*, 8 Qwen3NextAttention via KV.PREFIX.* + V.STOREBATCH). PionPromptCache is hybrid-aware — _classify_cache walks the cache list and routes per-slot. Drop-in for mixed-cache models: pc = PionPromptCache(model, vquant="fp16", port=1974); cache = pc.get_or_prefill(prefix_ids, namespace=ns). Reproducer: benchmarks/reproducers/sweep_qwen3_5_warm_ttft.py. Practical ceiling on the current wire is L=8192 (max layer 33.55 MB at 64 MB CLIENT_BUF_SIZE); past 8K needs streaming SSM.PREFIX.FETCH / V.FETCH RANGE.