AI.* — semantic cache, embeddings, routing¶
AI.SEMANTIC_CACHE is the shipped, benchmarked one. The rest are real, wired
and tested, but they are not launch claims — read each caveat as part of the
entry.
| Command | Pion path | GLIDE | Description |
|---|---|---|---|
| FT.CREATE | SLOW | ❌ | Create vector index (HNSW, SCHEMA, VECTOR field) |
| FT.SEARCH | SLOW | ❌ | KNN vector search; PARAMS format; BM25 <q> [K k] [K1 f] [B f]; HYBRID (vector+text); TAG/FILTER predicates |
| FT.OPTIMIZE | SLOW | ❌ | Build l0_compact, compact_buffer; save HNSW to disk (pion.hnsw.0) |
| FT.DROPINDEX | SLOW | ❌ | Drop vector index; clears in-memory HNSW. An unknown name is refused, nothing dropped |
| FT.INFO | SLOW | ❌ | [index_name, <name>, num_docs, <n>]; checks the name against the shared registry, same answer on every worker; unknown name → Unknown index name |
| FT.HYBRID | SLOW | ❌ | FT.HYBRID <index> <text_query> <vector_blob> [K <k>] [ALPHA <a>] [K1 <f>] [B <f>] — BM25+vector fusion with Reciprocal Rank Fusion |
| FT.ADDTEXT | SLOW | ❌ | Add text doc: store in keyspace + BM25 doc set (integer ids), embed via HTTP → insert into semantic HNSW |
| FT.SEARCHTEXT | SLOW | ❌ | Semantic text search over docs added via FT.ADDTEXT |
| AI.COMPLETE | SLOW | ❌ | Semantic cache check → LLM call on miss → cache result |
| AI.SEMANTIC_CACHE GET | SLOW | ❌ | Embed query → HNSW search → return cached response if cos sim ≥ threshold |
| AI.SEMANTIC_CACHE SET | SLOW | ❌ | Embed query → insert into HNSW → store response |
| AI.CHAT | SLOW | ❌ | Full RAG in one command: retrieve context → build prompt → call LLM |
| AI.FLARE LOAD | SLOW | ❌ | Add text to in-Mojo FLARE knowledge base |
| AI.FLARE RUN | SLOW | ❌ | Mid-generation retrieval loop (FLARE algorithm in Mojo) |
| AI.FLARE INFO | SLOW | ❌ | FLARE KB stats |
| AI.KNN_LM.CREATE | SLOW | ❌ | Allocate token-id-tagged kNN datastore: <ds_id> <dim> [<max_entries>] (default max=100K). Up to 16 datastores per worker. Substrate enabled when --kvcache or --inference is on. |
| AI.KNN_LM.STORE | SLOW | ❌ | Append one (token_id, embedding) pair: <ds_id> <next_token_id> <emb_blob>. Auto-builds HNSW once count crosses 5000. |
| AI.KNN_LM.STOREBATCH | SLOW | ❌ | Bulk append: <ds_id> <n> <ids_blob> <emb_blob> (ids: n × Int32 LE; emb: n × dim × Float32 LE). |
| AI.KNN_LM.QUERY | SLOW | ❌ | Top-k kNN: <ds_id> <k> <emb_blob>. Returns k × 8 bytes packed as <Int32 LE token_id><Float32 LE distance>. Brute-force scan below 5K, focused FP32 HNSW above. Sub-linear scaling, ~0.94 ms median at 30K entries. |
| AI.KNN_LM.INFO | SLOW | ❌ | count=N dim=D max_entries=M |
| AI.KNN_LM.DROP | SLOW | ❌ | Free datastore + HNSW graph buffers |
| NEURON.PKM.CREATE | SLOW | ❌ | Allocate a product-key memory table: <table> <dim> <n_slots> [VDIM <v>] [VALTYPE F32\|F16]. n_slots must be a perfect square S² (S ≤ 4096); dim must be even. Up to 8 tables per worker. Enabled by --kvcache / --inference. |
| NEURON.PKM.SETKEYS | SLOW | ❌ | Load codebook half 0 or 1: <table> <half> <blob> — S × (dim/2) Float32 LE. Builds the INT8 mirror. Both halves required before QUERY. |
| NEURON.PKM.SETVALS | SLOW | ❌ | Write value rows: <table> <off> <n> <blob> — n × vdim in VALTYPE. Value matrix is allocated on the first call. |
| NEURON.PKM.QUERY | SLOW | ❌ | Exact top-k over all n_slots: <table> <k> <q_blob> [FAST]. nq = len(q_blob)/(dim·4) heads share one codebook pass. Returns nq·k × 8 bytes as <Int32 LE slot_id><Float32 LE score>, descending, padded (-1, -inf). dim=896 k=32 1M slots: 0.073 ms (vs 5.13 ms for AI.KNN_LM.QUERY at the same shape). |
| NEURON.PKM.FFN | SLOW | ❌ | Fused lookup + softmax-weighted value read: <table> <k> <q_blob> [FAST] [TEMP <t>] → nq × vdim × 4 bytes Float32 LE. Value rows never cross the wire. |
| NEURON.PKM.INFO | SLOW | ❌ | dim=D half=H s_rows=S n_slots=N d_pad=P keys_ready=0\|1 vdim=V valtype=f32\|f16 val_rows=R queries=Q |
| NEURON.PKM.DROP | SLOW | ❌ | Free codebooks, value matrix, and query scratch |
| AI.EMBED | SLOW | ❌ | Embed text with the in-process model. -ERR AI.EMBED requires --inference or --emb-enabled when no backend is active |
| AI.GENERATE | SLOW | ❌ | Generate via the sidecar; -ERR ... requires --inference otherwise |
| AI.LOADMODEL | SLOW | ❌ | Load a sidecar model; -ERR ... requires --inference otherwise |
| AI.MEMORY | SLOW | ❌ | AI.MEMORY ADD\|RECALL\|CONTEXT ... — agent-memory surface; a bare call returns the syntax line |
The gateway commands¶
AI.COMPLETE¶
One-command cache-augmented generation: checks the semantic cache, returns the cached response on a hit, calls the LLM and caches the result on a miss.
redis-cli AI.COMPLETE "What is the capital of France?" TOKENS 100 THRESHOLD 0.92
# Cache hit (≥ threshold cosine similarity): returns cached response in <1ms
# Cache miss: calls LLM, stores result, returns LLM response
Cache hit performance: 275× faster than a live LLM call (see examples/pion_vs_ollama_demo.py).
FT.ADDTEXT¶
Stores a text document and makes it immediately searchable:
- Stores
doc_id+textas a HASH in the keyspace - Registers
doc_idin the BM25 doc set, soFT.OPTIMIZEindexes it - Embeds
textviaEmbeddingClient(HTTP/1.0 POST to/v1/embeddings) - Inserts the embedding into the per-worker semantic HNSW index (
add_and_insert) - Returns
+OK
Steps 3–4 require EmbeddingConfig.enabled = True; steps 1–2 do not, so BM25 ingest works with no embedding backend at all.
Step 2 only applies to non-negative integer doc_ids — BM25 hits are returned as integer ext_ids. A string id like doc:1 is still stored and embedded, so FT.SEARCHTEXT finds it, but it stays invisible to FT.SEARCH … BM25.
# string ids — semantic search only
redis-cli FT.ADDTEXT products doc:1 "Wireless headphones with active noise cancellation"
redis-cli FT.ADDTEXT products doc:2 "Mechanical keyboard Cherry MX switches"
# integer ids — semantic search AND BM25 (after FT.OPTIMIZE)
redis-cli FT.ADDTEXT products 1 "Wireless headphones with active noise cancellation"
redis-cli FT.OPTIMIZE products
redis-cli FT.SEARCH products BM25 "noise cancellation" K 3
FT.SEARCHTEXT¶
Semantic search over text documents added via FT.ADDTEXT:
- Embeds
query_textviaEmbeddingClient - Searches the per-worker semantic HNSW (ef=32, default k=10)
- Returns matching
doc_idstrings as a RESP array
AI.CHAT¶
Full RAG pipeline in a single RESP3 command:
- (optional) Embeds
text, searches HNSW for top-k doc_ids, retrieves stored texts - Builds LLM prompt:
"Context:\n<docs>\n\nUser: <prompt>"(or just<prompt>if no CONTEXT) - Sends HTTP/1.0 POST to
/v1/chat/completions - Parses
"content":"..."from JSON response, returns as RESP bulk string
# Simple chat
redis-cli AI.CHAT "What is the capital of France?"
# RAG: retrieve context then generate
redis-cli AI.CHAT "Which audio product do you recommend?" CONTEXT products "audio" K 3
# → "Based on the available products, I recommend the wireless headphones..."
AI.SEMANTIC_CACHE¶
Caches LLM responses keyed by semantic similarity of the input query:
AI.SEMANTIC_CACHE SET "What is the capital of France?" "+Paris\r\n"
AI.SEMANTIC_CACHE GET "capital of France?" → +Paris
AI.SEMANTIC_CACHE GET "unrelated query" → $-1 (cache miss)
Workspace audit trail¶
A cache hit tells you what the model answered, not what it had in mind while
answering. WORKSPACE records the caller's snapshot — e.g. the retrieved context or
top-k, base64 — at cache-WRITE time, and EXPLAIN reads it back with no inference
infrastructure required. That is the point: a compliance reviewer can ask
what was in the workspace without being able to run the model.
AI.SEMANTIC_CACHE SET "capital of France?" "Paris" WORKSPACE "<base64 lens top-k>"
AI.SEMANTIC_CACHE EXPLAIN "capital of France?" → "<base64 lens top-k>"
AI.SEMANTIC_CACHE EXPLAIN "never cached" → nil
AI.SEMANTIC_CACHE GET "capital of France?" WITHWORKSPACE
→ %2 response/workspace (RESP3)
→ *4 flat pairs (RESP2)
The blob is opaque to the server — stored and returned, never interpreted —
so the lens format can change without a wire change. An entry written without
WORKSPACE reports nil rather than an empty string, so "no audit trail" is
distinguishable from "an empty one".
WITHWORKSPACE is opt-in and the default GET reply is unchanged, on both
protocols. Always returning a RESP3 map would break every RESP3 client already
reading GET as a bulk string, and a wire break to add observability is the wrong
trade. EXPLAIN needs no flag.
Implementation note: workspaces[] runs parallel to responses[], and other
callers (MCP memory in ai.mojo, RAG ingest in vector.mojo) append to
responses directly without knowing about workspaces. cache_set therefore
PADS to count before appending rather than assuming alignment — if the two
ever drift, EXPLAIN returns some other entry's audit trail, which is worse
than returning none.
See src/network/semantic_cache.mojo and src/network/embedding_client.mojo.
Test: tests/test_gh115_workspace.py. It ENV_SKIPs without an embedding backend, because every SET is a
silent no-op there and the failures would all be about the missing backend.
AI.ROUTE.* — semantic load balancer¶
| Command | Pion path | GLIDE | Description |
|---|---|---|---|
| AI.ROUTE.REGISTER | SLOW | ❌ | Register inference node with semantic centroid embedding + optional CAPACITY |
| AI.ROUTE.UPDATE | SLOW | ❌ | Update a node's centroid embedding (as KV cache evolves) |
| AI.ROUTE | SLOW | ❌ | Route query embedding to best node by cosine similarity; 0.14ms, 7K QPS |
| AI.ROUTE.REMOVE | SLOW | ❌ | Remove node from routing table |
| AI.ROUTE.INFO | SLOW | ❌ | Per-node stats: routed count, capacity, endpoint |
Routing strategy: FP32 brute-force cosine for <=16 nodes (perfect accuracy); HNSW O(log N) for >16 nodes. 88% routing accuracy with Ollama embeddings.
Requires --kvcache flag.
Embedding backends and how the first request just works: Embeddings.