Skip to content

SSM.PREFIX.* — recurrent-state companion

Hybrid Mamba / GatedDeltaNet / RWKV models carry state that is not K/V. SSM.PREFIX.* stores that state under the same namespace contract as KV.PREFIX.*, so a hybrid model's whole prefix — softmax layers via the V-store, linear layers here — is shared across processes and survives a restart.

The KV.PREFIX. substrate above handles transformer-attention layers. For hybrid Mamba+Transformer models (LoLCATs, MOHAWK, Mamba-in-Llama recipes; or any future external small hybrid), the SSM half of the model has per-layer recurrent state, not per-token K/V. SSM.PREFIX. is the companion substrate for that state.

SSM.PREFIX.STORE  <sid> <layer> <state_blob>     → +OK
SSM.PREFIX.FETCH  <sid> <layer>                  → bulk string | $-1
SSM.PREFIX.DROP   <sid> [<layer>]                → +OK (layer omitted = drop all layers for the session)

state_blob is opaque to the server — pure host-side byte storage. The consumer chooses serialization. The reference format (see tests/test_ssm_prefix_roundtrip.py):

uint32 version = 1
uint32 n_arrays
for each array:
    uint32 ndim
    uint32[ndim] shape
    uint32 dtype_code   (0=fp32, 1=fp16, 2=bf16 — serialized as fp32, lossless)
    raw bytes

This keeps the substrate model-family-agnostic: Mamba (size=2 [conv_state, ssm_state]), RWKV-7 (size=3), and future RetNet / Hedgehog / GLA all serialize differently but the server never parses. Adding a new family is ~1-2 days of consumer-side (de)serializer work.

Validation: - Drift check (reproduced by tests/test_ssm_prefix_roundtrip.py): bit-perfect state hydration across Mamba-130M-f32 (2048 decode tokens), Mamba-370M-f16 (1024 decode tokens), and RWKV-7 168M (512 decode tokens). All show token agreement 100%, max-abs-diff = 0.000e+00 vs no-snapshot baseline. Deterministic-recurrence property holds. - Wire round-trip (tests/test_ssm_prefix_roundtrip.py): end-to-end through pion-server. 64/64 bit-perfect Mamba-130M, 64/64 bit-perfect RWKV-7. 1 MB random blob round-trip + overwrite + multi-layer drop all OK.

Pickup cost for new families: each new family needs its own (de)serializer in tests/test_ssm_prefix_roundtrip.py's serialize_arrays_cache / deserialize_arrays_cache. The RWKV-7 serializer is the reference implementation.

End-to-end hybrid model. mlx-community/Qwen3.5-4B-MLX-4bit (24 GatedDeltaNet + 8 Qwen3NextAttention, 3:1 ratio, Apache 2.0), a publicly released hybrid, with measured wins:

Prefix length Vanilla cold Pion warm Speedup Token agreement
174 tokens — — 3.3× —
2,048 tokens 6,080 ms 246 ms 24.77× 18/18
4,096 tokens 12,757 ms 390 ms 32.7× 16/18
8,192 tokens 27,598 ms 2,373 ms 11.6× 18/18

The speedup peaks at 4K. Warm TTFT is roughly constant up to 4K (246 → 390 ms) and then jumps to 2,373 ms at 8K, because the wire-fetch term grows with shipped bytes (max layer 33.55 MB at L=8192) while vanilla prefill grows only linearly — so 8K measures 11.6×, below 4K's 32.7×. Token agreement is 18/18 at 2K and 8K and 16/18 at 4K (greedy-argmax non-determinism present in both the vanilla and Pion paths).

18/18 layers bit-perfect across the cleanly-typed split path (24 GatedDeltaNet via SSM.PREFIX.*, 8 Qwen3NextAttention via KV.PREFIX.* + V.STOREBATCH). PionPromptCache is hybrid-aware — _classify_cache walks the cache list and routes per-slot. Drop-in for mixed-cache models: pc = PionPromptCache(model, vquant="fp16", port=1974); cache = pc.get_or_prefill(prefix_ids, namespace=ns). Reproducer: benchmarks/reproducers/sweep_qwen3_5_warm_ttft.py. Practical ceiling on the current wire is L=8192 (max layer 33.55 MB at 64 MB CLIENT_BUF_SIZE); past 8K needs streaming SSM.PREFIX.FETCH / V.FETCH RANGE.