Examples¶
Every example ships in the repository under examples/, with its
prerequisites stated. This index is the directory's own README, included at
build time.
Quick index. Start with the 5-minute demo at the top; the rest are research reproducers.
Start here¶
| Script | What it shows | Prerequisites |
|---|---|---|
prompt_cache_demo.py |
Shared KV Cache — the four PionPromptCache lines, vanilla cold prefill vs Pion warm hit, TTFT printed for both, outputs compared. Default --prefix-tokens 2048; --prefix-tokens 33 reproduces the case where Pion loses. |
pip install -e ../pion-vllm-mlx/[mlx] · mlx-lm · the default in-process path needs no server; --backend pion (cross-process reuse) needs Pion running --kvcache --metal-attention -w 1 |
cag_legal_demo/ |
CAG-hybrid pion-serve — 4 SCOTUS opinions, calibration, head-to-head vs Vector-RAG. Self-contained with corpus + calibration state. | pion-serve, Ollama |
Feature demos¶
| Script | Feature | Notes |
|---|---|---|
hybrid_retrieval_demo.py |
Hybrid Retrieval Cache | Chunk-id-keyed K/V, inproc + pion backends |
sparse_mask_64k_niah.py |
Sparse-mask SDPA at 64K context | 100% NIAH, 0.78% prefix budget; use --skip-vanilla on 16 GB Mac |
moe_expert_substrate_demo.py |
MoE expert paging | MOE.EXPERT.* — models beyond RAM |
agent_memory_demo.py |
AI agent memory via MCP | agent_remember / agent_recall / agent_forget |
ai_platform_scenario.py |
Multi-command AI gateway scenario | Semantic cache + RAG + routing |
cache_aware_router.py |
AI.ROUTE.* semantic load balancer |
88% accuracy, 0.14 ms |
prefix_routing_demo.py |
KV.PREFIX.* routing |
Cross-worker prefix directory |
multinode_prefix_routing.py |
Cross-host prefix routing | pion-exo cluster |
session_affinity_demo.py |
Session affinity routing | blake2b session-id hashing |
pion_vs_ollama_demo.py |
Pion Serve vs raw Ollama latency | Semantic cache hit rate |
coalescing_proxy.py |
Request coalescing proxy | Batch identical prompts |
exo_session_demo.py |
exo distributed inference hook | pion-exo + KV.PREFIX |
pion_glide_demo.py |
Valkey GLIDE client | GLIDE cluster client compat |
Research reproducers (step-series)¶
The step*.py files are an incremental research series exploring speculative/kNN-LM inference. They are not maintained as standalone demos.
| Script | Topic |
|---|---|
step6_flare.py, step6b_flare_fewshot.py |
FLARE mid-generation retrieval |
step7_rest_pion_drafter.py |
REST-based speculative draft |
step8_rest_13b.py, step9_knnlm.py |
kNN-LM baseline |
step10_vocab_bias.py, step11_rest_probabilistic.py |
Vocabulary bias; REST probabilistic draft |
step12_kv_cache_aligned.py, step13_flare_mojo.py |
KV-cache-aligned draft; FLARE in Mojo |
Modules, not demos¶
Two files here are imported by the demos rather than run directly:
pion_moe_tier.py— the MoE expert substrate, used bymoe_expert_substrate_demo.py.pion_moe_tier_client.py— its wire-backed client, also used by pion-serve'spion-moebackend.
Both were moved here from the private research tree so the demo and the backend actually run outside the development repo.