First five minutes¶
The four Pion lines¶
On Apple Silicon, PionPromptCache sits where mlx-lm's own prompt cache does.
The first run warms it; every later run — this process or another — hits.
# No server yet? One curl — see "Prebuilt binary" below.
./pion-server --kvcache --metal-attention -w 1
from pion_vllm_mlx import PionPromptCache
pc = PionPromptCache(model, host="127.0.0.1", port=1974) # once per loaded model
cache = pc.get_or_prefill(prefix_ids, namespace="app|v1|llama-1b|system_v1") # MISS: prefill once + store · HIT: fetch
text = generate(model, tok, prompt=suffix_ids, prompt_cache=cache) # decode as usual, at native speed
get_or_prefill returns what mlx_lm.make_prompt_cache(model) would, with the
prefix's K/V already in it. prefix_ids is the part every request shares (the
system prompt, the few-shot block, the retrieved chunk); suffix_ids is the
part that changes. The first process to ask pays the prefill once and stores
it; every later call — this process, another one, or after a restart — fetches
it. That is the cross-process row in the table below. The same-process row
adds the Stage-2 attention patch, in which Pion computes the attention over
the prefix itself (pion-vllm-mlx/README.md).
Runnable end to end, with the before/after timings printed:
examples/prompt_cache_demo.py. Install with
pip install -e 'pion-vllm-mlx/[mlx]' (not on PyPI yet).
The Redis wire, vector search, the semantic cache¶
KV — redis-cli talks to Pion unmodified:
Vector search. A vector is DIM * 4 raw float32 bytes — 6144 for the
default DIM 1536 — so this needs a real client rather than redis-cli. Three
things the API requires, in this order:
FT.CREATEbefore anyHSET. Vectors are only routed into the HNSW once the index exists; documents added earlier are ordinary hashes and are not searchable.HSET <key> vec <blob> <field> <value>— at least two field/value pairs. The vector-ingest path is only taken forHSETwith 6 or more arguments.DISTANCE_METRICacceptsL2andCOSINE. Any other value —IPincluded — is refused atFT.CREATErather than silently answered with the wrong ordering. Omitting the keyword givesL2.
DIM is per-index in the RediSearch API but the blob must be DIM * 4 bytes of
float32 either way; the example below uses the server's default 1536.
pip install redis numpy # or: pixi run python -m pip install redis numpy
python3 - <<'PY'
import numpy as np, redis
r = redis.Redis(port=1974)
D = 1536 # blob is D*4 = 6144 bytes
r.execute_command("FT.CREATE", "idx", "SCHEMA", "vec", "VECTOR", "HNSW", "6",
"TYPE", "FLOAT32", "DIM", str(D), "DISTANCE_METRIC", "COSINE")
rng = np.random.default_rng(0)
vecs = {f"doc:{i}": rng.random(D).astype(np.float32) for i in range(20)}
for key, v in vecs.items():
r.execute_command("HSET", key, "vec", v.tobytes(), "title", key)
r.execute_command("FT.OPTIMIZE", "idx") # build the graph, then query
hits = r.execute_command("FT.SEARCH", "idx", "*=>[KNN 3 @vec $B]",
"PARAMS", "2", "B", vecs["doc:7"].tobytes(),
"DIALECT", "2")
print(hits[0], "hits; nearest =", hits[1].decode()) # 3 hits; nearest = doc:7
PY
Semantic cache. The cache's index is per worker, so it needs a
single-worker server (the default)
and an embedding backend. On macOS use Apple's NLEmbedding — no downloads, no
Python:
./pion-server --nle-embed # macOS: ANE-accelerated, 512-dim
# other platforms: pixi run install-inference # once; then ./pion-server
# # auto-embeds via MiniLM-L6-v2
redis-cli -p 1974 AI.SEMANTIC_CACHE SET "capital of France?" "Paris"
redis-cli -p 1974 AI.SEMANTIC_CACHE GET "What is the capital of France?"
# -> "Paris" — a different wording, matched by meaning
Server profiles¶
Run
./pion-server --helpfor the full flag reference (cluster / AI / inference / tuning groups), grouped and one-line-described.
./pion-server --profile kv # KV-only: ~50MB/worker, no HNSW
./pion-server --profile vector # KV + vector search
./pion-server --profile ai --flare # KV + vector + AI (semantic cache, RAG, FLARE)
./pion-server --kvcache -w 1 # + externalized attention (ATTEND.*)
./pion-server --metal-attention -w 1 # + native Metal SDPA on Mac (no Python; fp32, more accurate)
./pion-server --metal-attention-fp16 -w 1 # fp16 kernel — matches vanilla mlx-lm precision
./pion-server --metal-attention --fa-window 2048 -w 1 # sliding-window SDPA: scans only the last N tokens. Safe on Mistral SWA / Longformer; lossy on dense models (Llama, Gemma, GPT).
Where next¶
- Where Pion does not help — read this before you build on it.
- Shared KV cache — namespaces, lanes, Stage 1 vs Stage 2.
- Command index and the CLI flags.