Pion — the memory engine for AI inference¶
Pion remembers what your model already read: prefill a long prompt once, and every later request, from any process and even after a crash, starts at the first new token. It is a memory engine for AI inference — one small Mojo binary that holds the prompt's K/V cache, and the rest of a model's working memory, and serves it over the Redis wire.
On Apple Silicon it is four lines around mlx_lm:
from pion_vllm_mlx import PionPromptCache
pc = PionPromptCache(model, host="127.0.0.1", port=1974) # once per loaded model
cache = pc.get_or_prefill(prefix_ids, namespace="app|v1|llama-1b|system_v1") # MISS: prefill once + store · HIT: fetch
text = generate(model, tok, prompt=suffix_ids, prompt_cache=cache) # decode as usual, at native speed
| Time to first token, Apple Silicon | vanilla mlx-lm | Pion warm | |
|---|---|---|---|
| Llama-3.2-1B-4bit, 2,048-token prefix, same process | 1,530 ms | 30.2 ms | 50.6× |
| Same model and prefix, from a separate process, over the wire | 1,558 ms | 64.7 ms | 24× |
The two rows differ by where the attention runs; the second pays a wire hop
for a cache a separate process wrote, and reproduces the vanilla output at
BLEU 1.000. Each row has a reproducer:
tests/bench_ttft.py and
cross_process_ttft.py.
Read the limits before installing: it caches prefill, not decode, and it
only pays off when a long prefix is really reused.
What Pion is, and is not¶
Pion is a memory engine: the prompt's K/V, recurrent (SSM) state, MoE expert
weights, vectors and embeddings, and agent memory — behind one wire protocol,
with one durability story. Two properties are Pion's own: an acked write is
still there after SIGKILL, and the cache is read directly by a different
program over a protocol every language already has a client for.
It is not an inference engine: it feeds forward passes, it does not run them. The vector engine exists so that recall needs no second database, and Pion Serve is an example proxy, not a requirement.
Where to start¶
- Get started — Shared KV cache: namespaces, Stage 1 vs Stage 2, the wire protocol, what is measured.
- Reference — Command matrix for every Redis-compatible
command and its known divergences; Configuration for
profiles; Networking for the event-loop tiers and the
binary lane on
port+1; Client APIs. - Operations — Running in production: crash breadcrumbs, the status file, supervised serving, durability of every type; Persistence; Multi-tenant; Distributed systems (a preview — read its caveats).
- Internals — Architecture, Memory management, Vector engine.
- Licensing —
licensing.md: Pion is Apache-2.0, with one free closed binary library (libpion_vector) for the tuned vector kernels.
Also here¶
AI gateway · Embeddings · Pion Serve · Data types · Benchmarking guide · Development guide · Codebase search