0.9.0 · public preview · macOS Apple Silicon + Linux
Your model re-reads the same prompt on every request.
Pion remembers what your model already read: prefill a long prompt once, and every later request, from any process and even after a crash, starts at the first new token. It is a memory engine for AI inference — one small Mojo binary that holds the prompt's K/V cache, and the rest of a model's working memory, and serves it over the Redis wire.
curl -fsSL https://github.com/pavelhorak/pion/releases/latest/download/pion-macos-arm64.tar.gz | tar xz cd pion-*-macos-arm64 ./pion-server.sh # port 1974 redis-cli -p 1974 PING # +PONG
Roughly 2 MB: the binary, the three Mojo runtime dylibs it actually links
against, and the Metal shader library that --metal-attention needs. No
toolchain, no Python, no model download. Verify the download against the
release's SHA256SUMS if you care to (shasum -a 256 -c SHA256SUMS).
Runnable end to end, with the before/after timings printed:
examples/prompt_cache_demo.py. Install with
pip install -e 'pion-vllm-mlx/[mlx]' (not on PyPI yet).
The client for the prompt-cache path. It needs a running pion-server --kvcache --metal-attention -w 1 from the macOS tab.
docker build -t pion . # from source, ~20 min docker build --target runtime-prebuilt -t pion . # wraps ./pion-server, seconds # (Linux host only — the image # runs the binary in the context) docker run -p 1974:1974 -v pion-data:/data pion redis-cli -p 1974 PING # +PONG
No image is published to a registry yet, so build one first — hermetically from source, or in seconds around a binary you already have:
git clone https://github.com/pavelhorak/pion.git && cd pion pixi install pixi run build # produces ./pion-server ./pion-server # port 1974, one shared keyspace
Prerequisites: - pixi builds and runs everything here —
curl -fsSL https://pixi.sh/install.sh | bash, then restart your shell. It brings its own Mojo toolchain; nothing else needs installing. -redis-clifor the snippets below (brew install redis/apt install redis-tools). Any Redis client works — Pion speaks RESP2 and RESP3. - macOS: the Metal shader step (xcrun metal) needs full Xcode (not just Command Line Tools). Without it the build prints[Metal] skippedand reuses the trackedsrc/ffi/metal_compute.metallib, so--metal-attentionstill works — you only need Xcode if you editmetal_compute.metal. To rebuild the shader, install Xcode from the App Store, thenxcode-select -s /Applications/Xcode.app/Contents/Developer. - Linux (GPU):pixi run buildcalls/usr/local/cuda/bin/nvccand linkscudart. Install CUDA 12 + cudart first. - Linux (no GPU): usepixi run build-portable— GPU-free build (no nvcc, no cudart, portable x86-64-v2 binary). Disables--metal-attentionand--cuda-attentionbut all KV / vector / AI commands work. It produces./pion-server-dev, not./pion-server— substitute that name in every command below.
Loopback by default, no TLS; a password flips the bind address. Read the security model before the server sees a network.
A drop-in for make_prompt_cache
On Apple Silicon, PionPromptCache sits where mlx-lm's own prompt cache does. The first run warms it; every later run — this process or another — hits.
get_or_prefill returns what mlx_lm.make_prompt_cache(model) would, with the
prefix's K/V already in it. prefix_ids is the part every request shares (the
system prompt, the few-shot block, the retrieved chunk); suffix_ids is the
part that changes. The first process to ask pays the prefill once and stores
it; every later call — this process, another one, or after a restart — fetches
it. That is the cross-process row in the table below. The same-process row
adds the Stage-2 attention patch, in which Pion computes the attention over
the prefix itself (pion-vllm-mlx/README.md).
from pion_vllm_mlx import PionPromptCache pc = PionPromptCache(model, host="127.0.0.1", port=1974) # once per loaded model cache = pc.get_or_prefill(prefix_ids, namespace="app|v1|llama-1b|system_v1") # MISS: prefill once + store · HIT: fetch text = generate(model, tok, prompt=suffix_ids, prompt_cache=cache) # decode as usual, at native speed
Two numbers, and what each one measures
| Time to first token | vanilla mlx-lm | Pion warm | factor |
|---|---|---|---|
| Llama-3.2-1B-4bit, 2,048-token prefix, same process | 1,530 ms | 30.2 ms | 50.6× |
| Same model and prefix, from a separate process, over the wire | 1,558 ms | 64.7 ms | 24× |
| Gemma-4-E2B-4bit, 64K context, sparse mask — 100% needle recall attending 0.78% of the prefix | 326×NIAH-class only |
The two rows differ by where the attention runs, not by how much faster the same thing got. The first keeps model and cache in one process. The second pays a wire hop for a cache a separate process wrote, and produces BLEU 1.000 on a 50-token greedy completion against the vanilla output — a different socket and a different model object, same text.
A third number, because it is the one that surprised us: at 64K context with a sparse mask, 100% needle recall while attending 0.78% of the prefix, with bit-identical greedy decode. That is needle-in-a-haystack retrieval, not a claim about every long-context task — see the list below.
Each row has a script: tests/bench_ttft.py for
the first, cross_process_ttft.py
for the second. The reproducers list
what each one needs.
This section exists because the demo that ships with this project once printed 0.86× — slower than doing nothing — and we shipped it that way for a while.
- It caches prefill, not decode. If your bottleneck is tokens-per-second once generation starts, this changes nothing.
- The reuse has to be real, and the prefix has to be long. Time to first token from a separate process, only the prefix length varying: 34 tokens → 1.6×. 268 → 6.0×. 1,035 → 16.4×. 2,049 → 24×. The ratio holds up at short prefixes; the saving does not — at 34 tokens it is about 19 ms, not worth a second process. The demo defaults to 2,048, and
--prefix-tokens 33shows how little there is to save. - The sparse selector is NIAH-class. Factual QA over dense paragraphs needs a query-aware selector and collapses to 0.38×.
-w Nis N independent keyspaces, not one shared keyspace. The server refuses to start with-w N > 1unless you pass--independent-workers.- No TLS. Loopback by default; terminate TLS at a proxy.
- No eviction.
--maxmemoryrefuses writes with-OOMonce the process reaches the limit, but nothing is ever evicted to make room. - Cluster mode is single-node. Slot migration and replication exist; production multi-node does not.
- One maintainer, working evenings.
Every competitor here does something well, and most of them do more than a table can show. Cells are what each project ships today (checked 2026-09-19); the footnotes carry the sources.
| LMCache (+vLLM) | SGLang HiCache | oMLX | mlx-lm cache_prompt |
Pion | |
|---|---|---|---|---|---|
| Cache shared across processes | ✓ via remote store¹ | ✓ L3¹ | ✗ per-process² | manual file | ✓ wire-native |
| Survives restart | ✓ | ✓ | ✓ SSD tier | ✓ manual | ✓ |
| Crash-consistent (acked = durable) | ✗ | ✗ | ✗ | ✗ | ✓ WAL† |
| KV quantization | FP8; 4-bit demo³ | FP8 | ✓ TurboQuant | ✗ | fp16/int8/turbo4/INT3/INT2/mlx4g32 |
| Hybrid (Mamba/GDN) prefix state | in-process⁴ | host + storage tiers⁴ | in-process + SSD | ✗ | cross-process + persisted |
| Sparse long-context selector | ✗ | ✗ | ✗ | ✗ | ✓ block-mean, bit-identical decode |
| MoE expert paging | ✗ | ✗ | ✓ in-process (experimental) | ✗ | ✓ cross-process + histograms |
| Apple Silicon | ✗ (CUDA) | ✗ (CUDA) | ✓ | ✓ | ✓ (+ Linux CPU) |
| Redis wire / any RESP client | ✗ | ✗ | ✗ | ✗ | ✓ |
| Runs as | vLLM plugin + service | serving engine | menu-bar app / server | library | one server binary |
The two rows that are actually ours are crash-consistency and the shape of the sharing. Everything else on this list is a matter of degree.
Crash-consistent is not the same as persistent. LMCache's disk backend,
SGLang's L3 and oMLX's SSD tier all persist, and all of them recompute on a
miss — that is a perfectly good design. Pion's contract is narrower and
stronger: an acked write is there after a SIGKILL.†
Shared across processes is not the same as distributed. LMCache does cross-instance sharing through a remote store, and that is its entire pitch; oMLX clusters by routing a request to the node that already holds the prefix, which is a different and reasonable answer. Pion is one process on one port that a different program, a different model object, or a different machine reads directly — no connector, no serving stack, no CUDA.
¹ LMCache ships local_disk_backend.py, p2p_backend.py and NIXL/remote connectors; SGLang HiCache L3 backs onto Mooncake/3FS/NIXL.
² oMLX per-node caches are private to the process; its cluster scheduler scores nodes on prefix affinity and routes to the hot one. An export/import API was requested by a user on 2026-09-12 (jundot/omlx#3612) and is not shipped.
³ LMCache demoed 4-bit KV with AMD on 2026-08-28.
⁴ vLLM merged hybrid prefix caching 2026-07-12 (vllm#46384), vllm-metal 2026-08-10 (#584); both share that state within a process. SGLang's HiCache tiers hybrid GDN/Mamba state to host memory and its storage backends as of 2026-09-21 (sglang#37507), inside its own serving stack.
† WAL durability is verified on macOS and Linux, V-store included: SIGKILL → restart replays the WAL and V.FETCH returns bit-equal data (max |Δ| = 0.0), validated on EPYC 8124P. tests/test_vstore_wal.py runs in Gate 2c on every gate.
TTFT, for the case Pion is built for: 1,530 ms → 30.2 ms warm
(50.6×, Llama-3.2-1B-4bit, 2,048-token prefix, same process) and
24× from a separate process over the wire (1,558 → 64.7 ms). PionPromptCache is a 4-line drop-in for
make_prompt_cache; cross-instance output verified BLEU 1.0. If you run one
server and never restart it, vLLM's APC already does this and you do not need
Pion.
On oMLX specifically
oMLX (21.8K★) is the best way to serve models locally on a Mac today, and if that is what you want, use it — a menu-bar app, continuous batching, a RAM+SSD tiered KV cache, TurboQuant KV, and cache-aware cluster routing across machines.
Pion is not a serving app and does not compete with it. oMLX keeps each node's cache private to its own process and routes work to whichever node is warm; Pion is a cache substrate that a separate process reads over a wire protocol, holds hybrid/SSM state across processes rather than within one, and makes an acked write survive a crash. The two compose, and Pion as a shared or remote tier underneath oMLX is a thing we would like to build.
That the problem is real is not just our claim: an oMLX user recently asked for durable, portable prefix-cache artifacts so an agent session would stop re-prefilling 60–120k tokens after a reload (#3612). They proposed a different shape from ours — an agent-owned file with a manifest, exported through oMLX's own API, rather than a server — so read it as evidence for the problem, not as a vote for Pion. The constraints they arrived at are the four Pion had to solve: carry the recurrent/GDN state and not only KV, pin the quantization parameters in the manifest, key on the exact token sequence rather than the text, and fail closed on any mismatch.
The Redis wire, clients unchanged
RESP2 and RESP3. redis-py, go-redis, ioredis, Valkey GLIDE, RedisVL, LangChain and LMCache's RedisConnector talk to Pion as they are.
14.0M ops/sec (Linux w=32) · P99 0.30ms (P=1), 1.6ms (vector)
Peak numbers above one worker use --independent-workers: N independent keyspaces. At P=1 it is parity with Redis — the localhost TCP ceiling every engine hits.
Vector search in the same process
HNSW with SIMD INT8 kernels, FT.CREATE/FT.SEARCH, the Redis 8 VSET commands (one vector set per key), BM25 and hybrid retrieval, an in-process embedding model so the first request just works.
Recall@100 0.937 (INT8), 0.951 (INT4) · QPS 6,803 (Linux), 8,134 (macOS)
One index per server; L2 and COSINE, anything else refused at FT.CREATE.
One binary, durable by default
Zero-allocation hot path, mmap WAL with segment rotation, snapshots, blob arenas for large values, crash breadcrumbs and a 1 Hz status file that survives an OOM kill.
mmap WAL + HNSW snapshot + V-store WAL
Loopback by default. No TLS, no maxmemory — read the security model first.
Runnable examples ship in examples/, each with its prerequisites stated — the prompt-cache demo, hybrid retrieval, 64K sparse NIAH, MoE expert paging on a 16 GB Mac, agent memory over MCP.