The memory planes¶
Pion is one process that holds the memory an inference stack otherwise spreads over four services: the prompt's K/V tensors, recurrent (SSM) state, MoE expert weights, vectors and embeddings, and the agent's own memory — behind one wire protocol, with one durability story.
The planes¶
| Plane | Commands | Reference |
|---|---|---|
| Prefix K/V cache | KV.PREFIX.*, V.* |
KV.PREFIX.* · V-store |
| Attention over cached K/V | ATTEND.*, ATTEND.PREFIX.* |
ATTEND.* |
| Recurrent state | SSM.PREFIX.* |
SSM.PREFIX.* |
| MoE expert paging | MOE.EXPERT.* |
MOE.EXPERT.* |
| Vectors and retrieval | FT.*, VADD/VSIM and the rest of VSET, BM25, hybrid |
FT.* and VSET |
| Semantic cache and routing | AI.*, RAG.*, AI.KNN_LM.*, NEURON.PKM.* |
AI.* · RAG and friends |
| Key-value | the Redis command surface | Redis-compatible commands |
Every plane speaks RESP2/RESP3, so any Redis client can reach it. The K/V planes also have a binary lane on port+1 for large tensor transfers.
What Pion is not¶
- Not an inference engine. Pion does not run forward passes. mlx-lm, MLX, vLLM and similar runtimes run the model and use Pion as their memory.
- Not a general vector database. The vector engine exists so that recall does not need a second database. It is one HNSW index per server (vector engine).
- Not an AI gateway. Pion Serve is an example proxy built on the planes above. It is not required to use them.
The integrations (pion-vllm-mlx, the MCP server, the LangGraph, AutoGen and
LlamaIndex adapters, pion-exo) are separate packages that talk to the server
over the wire.