Skip to content

First five minutes

The four Pion lines

On Apple Silicon, PionPromptCache sits where mlx-lm's own prompt cache does. The first run warms it; every later run — this process or another — hits.

# No server yet? One curl — see "Prebuilt binary" below.
./pion-server --kvcache --metal-attention -w 1
from pion_vllm_mlx import PionPromptCache

pc = PionPromptCache(model, host="127.0.0.1", port=1974)                       # once per loaded model
cache = pc.get_or_prefill(prefix_ids, namespace="app|v1|llama-1b|system_v1")  # MISS: prefill once + store · HIT: fetch
text = generate(model, tok, prompt=suffix_ids, prompt_cache=cache)             # decode as usual, at native speed

get_or_prefill returns what mlx_lm.make_prompt_cache(model) would, with the prefix's K/V already in it. prefix_ids is the part every request shares (the system prompt, the few-shot block, the retrieved chunk); suffix_ids is the part that changes. The first process to ask pays the prefill once and stores it; every later call — this process, another one, or after a restart — fetches it. That is the cross-process row in the table below. The same-process row adds the Stage-2 attention patch, in which Pion computes the attention over the prefix itself (pion-vllm-mlx/README.md).

Runnable end to end, with the before/after timings printed: examples/prompt_cache_demo.py. Install with pip install -e 'pion-vllm-mlx/[mlx]' (not on PyPI yet).

The Redis wire, vector search, the semantic cache

KV — redis-cli talks to Pion unmodified:

redis-cli -p 1974 SET user:1 '{"name":"alice"}'
redis-cli -p 1974 GET user:1

Vector search. A vector is DIM * 4 raw float32 bytes — 6144 for the default DIM 1536 — so this needs a real client rather than redis-cli. Three things the API requires, in this order:

  1. FT.CREATE before any HSET. Vectors are only routed into the HNSW once the index exists; documents added earlier are ordinary hashes and are not searchable.
  2. HSET <key> vec <blob> <field> <value> — at least two field/value pairs. The vector-ingest path is only taken for HSET with 6 or more arguments.
  3. DISTANCE_METRIC accepts L2 and COSINE. Any other value — IP included — is refused at FT.CREATE rather than silently answered with the wrong ordering. Omitting the keyword gives L2.

DIM is per-index in the RediSearch API but the blob must be DIM * 4 bytes of float32 either way; the example below uses the server's default 1536.

pip install redis numpy       # or: pixi run python -m pip install redis numpy
python3 - <<'PY'
import numpy as np, redis
r = redis.Redis(port=1974)
D = 1536                                        # blob is D*4 = 6144 bytes

r.execute_command("FT.CREATE", "idx", "SCHEMA", "vec", "VECTOR", "HNSW", "6",
                  "TYPE", "FLOAT32", "DIM", str(D), "DISTANCE_METRIC", "COSINE")

rng = np.random.default_rng(0)
vecs = {f"doc:{i}": rng.random(D).astype(np.float32) for i in range(20)}
for key, v in vecs.items():
    r.execute_command("HSET", key, "vec", v.tobytes(), "title", key)

r.execute_command("FT.OPTIMIZE", "idx")         # build the graph, then query

hits = r.execute_command("FT.SEARCH", "idx", "*=>[KNN 3 @vec $B]",
                         "PARAMS", "2", "B", vecs["doc:7"].tobytes(),
                         "DIALECT", "2")
print(hits[0], "hits; nearest =", hits[1].decode())   # 3 hits; nearest = doc:7
PY

Semantic cache. The cache's index is per worker, so it needs a single-worker server (the default) and an embedding backend. On macOS use Apple's NLEmbedding — no downloads, no Python:

./pion-server --nle-embed                       # macOS: ANE-accelerated, 512-dim
# other platforms: pixi run install-inference   # once; then ./pion-server
#                                               # auto-embeds via MiniLM-L6-v2

redis-cli -p 1974 AI.SEMANTIC_CACHE SET "capital of France?" "Paris"
redis-cli -p 1974 AI.SEMANTIC_CACHE GET "What is the capital of France?"
# -> "Paris"  — a different wording, matched by meaning

Server profiles

Run ./pion-server --help for the full flag reference (cluster / AI / inference / tuning groups), grouped and one-line-described.

./pion-server --profile kv            # KV-only: ~50MB/worker, no HNSW
./pion-server --profile vector        # KV + vector search
./pion-server --profile ai --flare    # KV + vector + AI (semantic cache, RAG, FLARE)
./pion-server --kvcache -w 1          # + externalized attention (ATTEND.*)
./pion-server --metal-attention -w 1  # + native Metal SDPA on Mac (no Python; fp32, more accurate)
./pion-server --metal-attention-fp16 -w 1  # fp16 kernel — matches vanilla mlx-lm precision
./pion-server --metal-attention --fa-window 2048 -w 1  # sliding-window SDPA: scans only the last N tokens. Safe on Mistral SWA / Longformer; lossy on dense models (Llama, Gemma, GPT).

Where next