RAG.* · AI.KNN_LM.* · NEURON.PKM.*¶
Three substrate families that share one property: they are datastores an inference loop reads at every step, so latency is the whole spec.
RAG.* — speculative RAG¶
| Command | Pion path | GLIDE | Description |
|---|---|---|---|
| RAG.SPECULATE.ENABLE | SLOW | ❌ | Enable speculative RAG for a session (creates trajectory ring buffer of last 10 embeddings) |
| RAG.QUERY | SLOW | ❌ | Query with speculation: check pre-computed cache first (cosine > 0.9), fall back to live HNSW search |
| RAG.SPECULATE.INFO | SLOW | ❌ | Per-session stats: hit rate, trajectory length, predictions outstanding |
Prediction model: predicted = current + alpha * (current - previous), alpha in [0.5, 1.0, 1.5] (3 predictions per query). 75% hit rate on linear trajectory, 0.2ms prediction+lookup latency.
Requires --kvcache flag.
AI.KNN_LM.* and NEURON.PKM.*¶
AI.SEMANTIC_CACHE is the shipped, benchmarked one. The rest are real, wired and
tested, but they are not launch claims — measured on narrower workloads than the
numbers beside them suggest, so read each caveat as part of the entry.
- AI.SEMANTIC_CACHE — cosine cache, 60% hit rate, auto-embed.
- AI.COMPLETE — cache + LLM in one command.
- AI.CHAT — full RAG pipeline (retrieve + generate) in one command.
- AI.ROUTE — semantic load balancer, 88% accuracy at 0.14ms.
- RAG.SPECULATE — predict next query via embedding momentum, 75% hit rate on linear-trajectory query streams.
- AI.KNN_LM.* — token-id-tagged kNN datastore substrate for client-side kNN-LM augmentation (CREATE / STORE / STOREBATCH / QUERY / INFO / DROP). Up to 16 named datastores per worker. Brute-force scan below 5K entries; per-vector SQ8 + asymmetric INT8 SIMD HNSW above (M=32, heap-based PQ + per-vec batch-4 distance kernel + lazy reciprocal pruning). dim=768, n=30K, k=10: 0.90 recall, 1.94 ms median latency, 36 s bulk-build — fits inline in any inference loop. Use cases: code completion, log generation, in-domain text infill.
Product-key memory (NEURON.PKM.*) gives exact top-k over N = S² slots at
2·√N·(dim/2) MACs — 1M slots, dim 896, k=32 answers in 0.073 ms where the
kNN-LM HNSW takes 5.13 ms.