Skip to content

AI.* — semantic cache, embeddings, routing

AI.SEMANTIC_CACHE is the shipped, benchmarked one. The rest are real, wired and tested, but they are not launch claims — read each caveat as part of the entry.

Command Pion path GLIDE Description
FT.CREATE SLOW ❌ Create vector index (HNSW, SCHEMA, VECTOR field)
FT.SEARCH SLOW ❌ KNN vector search; PARAMS format; BM25 <q> [K k] [K1 f] [B f]; HYBRID (vector+text); TAG/FILTER predicates
FT.OPTIMIZE SLOW ❌ Build l0_compact, compact_buffer; save HNSW to disk (pion.hnsw.0)
FT.DROPINDEX SLOW ❌ Drop vector index; clears in-memory HNSW. An unknown name is refused, nothing dropped
FT.INFO SLOW ❌ [index_name, <name>, num_docs, <n>]; checks the name against the shared registry, same answer on every worker; unknown name → Unknown index name
FT.HYBRID SLOW ❌ FT.HYBRID <index> <text_query> <vector_blob> [K <k>] [ALPHA <a>] [K1 <f>] [B <f>] — BM25+vector fusion with Reciprocal Rank Fusion
FT.ADDTEXT SLOW ❌ Add text doc: store in keyspace + BM25 doc set (integer ids), embed via HTTP → insert into semantic HNSW
FT.SEARCHTEXT SLOW ❌ Semantic text search over docs added via FT.ADDTEXT
AI.COMPLETE SLOW ❌ Semantic cache check → LLM call on miss → cache result
AI.SEMANTIC_CACHE GET SLOW ❌ Embed query → HNSW search → return cached response if cos sim ≥ threshold
AI.SEMANTIC_CACHE SET SLOW ❌ Embed query → insert into HNSW → store response
AI.CHAT SLOW ❌ Full RAG in one command: retrieve context → build prompt → call LLM
AI.FLARE LOAD SLOW ❌ Add text to in-Mojo FLARE knowledge base
AI.FLARE RUN SLOW ❌ Mid-generation retrieval loop (FLARE algorithm in Mojo)
AI.FLARE INFO SLOW ❌ FLARE KB stats
AI.KNN_LM.CREATE SLOW ❌ Allocate token-id-tagged kNN datastore: <ds_id> <dim> [<max_entries>] (default max=100K). Up to 16 datastores per worker. Substrate enabled when --kvcache or --inference is on.
AI.KNN_LM.STORE SLOW ❌ Append one (token_id, embedding) pair: <ds_id> <next_token_id> <emb_blob>. Auto-builds HNSW once count crosses 5000.
AI.KNN_LM.STOREBATCH SLOW ❌ Bulk append: <ds_id> <n> <ids_blob> <emb_blob> (ids: n × Int32 LE; emb: n × dim × Float32 LE).
AI.KNN_LM.QUERY SLOW ❌ Top-k kNN: <ds_id> <k> <emb_blob>. Returns k × 8 bytes packed as <Int32 LE token_id><Float32 LE distance>. Brute-force scan below 5K, focused FP32 HNSW above. Sub-linear scaling, ~0.94 ms median at 30K entries.
AI.KNN_LM.INFO SLOW ❌ count=N dim=D max_entries=M
AI.KNN_LM.DROP SLOW ❌ Free datastore + HNSW graph buffers
NEURON.PKM.CREATE SLOW ❌ Allocate a product-key memory table: <table> <dim> <n_slots> [VDIM <v>] [VALTYPE F32\|F16]. n_slots must be a perfect square S² (S ≤ 4096); dim must be even. Up to 8 tables per worker. Enabled by --kvcache / --inference.
NEURON.PKM.SETKEYS SLOW ❌ Load codebook half 0 or 1: <table> <half> <blob> — S × (dim/2) Float32 LE. Builds the INT8 mirror. Both halves required before QUERY.
NEURON.PKM.SETVALS SLOW ❌ Write value rows: <table> <off> <n> <blob> — n × vdim in VALTYPE. Value matrix is allocated on the first call.
NEURON.PKM.QUERY SLOW ❌ Exact top-k over all n_slots: <table> <k> <q_blob> [FAST]. nq = len(q_blob)/(dim·4) heads share one codebook pass. Returns nq·k × 8 bytes as <Int32 LE slot_id><Float32 LE score>, descending, padded (-1, -inf). dim=896 k=32 1M slots: 0.073 ms (vs 5.13 ms for AI.KNN_LM.QUERY at the same shape).
NEURON.PKM.FFN SLOW ❌ Fused lookup + softmax-weighted value read: <table> <k> <q_blob> [FAST] [TEMP <t>] → nq × vdim × 4 bytes Float32 LE. Value rows never cross the wire.
NEURON.PKM.INFO SLOW ❌ dim=D half=H s_rows=S n_slots=N d_pad=P keys_ready=0\|1 vdim=V valtype=f32\|f16 val_rows=R queries=Q
NEURON.PKM.DROP SLOW ❌ Free codebooks, value matrix, and query scratch
AI.EMBED SLOW ❌ Embed text with the in-process model. -ERR AI.EMBED requires --inference or --emb-enabled when no backend is active
AI.GENERATE SLOW ❌ Generate via the sidecar; -ERR ... requires --inference otherwise
AI.LOADMODEL SLOW ❌ Load a sidecar model; -ERR ... requires --inference otherwise
AI.MEMORY SLOW ❌ AI.MEMORY ADD\|RECALL\|CONTEXT ... — agent-memory surface; a bare call returns the syntax line

The gateway commands

AI.COMPLETE

AI.COMPLETE <query> [TOKENS <n>] [THRESHOLD <t>]

One-command cache-augmented generation: checks the semantic cache, returns the cached response on a hit, calls the LLM and caches the result on a miss.

redis-cli AI.COMPLETE "What is the capital of France?" TOKENS 100 THRESHOLD 0.92
# Cache hit (≥ threshold cosine similarity): returns cached response in <1ms
# Cache miss: calls LLM, stores result, returns LLM response

Cache hit performance: 275× faster than a live LLM call (see examples/pion_vs_ollama_demo.py).


FT.ADDTEXT

FT.ADDTEXT <index> <doc_id> <text>

Stores a text document and makes it immediately searchable:

  1. Stores doc_id + text as a HASH in the keyspace
  2. Registers doc_id in the BM25 doc set, so FT.OPTIMIZE indexes it
  3. Embeds text via EmbeddingClient (HTTP/1.0 POST to /v1/embeddings)
  4. Inserts the embedding into the per-worker semantic HNSW index (add_and_insert)
  5. Returns +OK

Steps 3–4 require EmbeddingConfig.enabled = True; steps 1–2 do not, so BM25 ingest works with no embedding backend at all.

Step 2 only applies to non-negative integer doc_ids — BM25 hits are returned as integer ext_ids. A string id like doc:1 is still stored and embedded, so FT.SEARCHTEXT finds it, but it stays invisible to FT.SEARCH … BM25.

# string ids — semantic search only
redis-cli FT.ADDTEXT products doc:1 "Wireless headphones with active noise cancellation"
redis-cli FT.ADDTEXT products doc:2 "Mechanical keyboard Cherry MX switches"

# integer ids — semantic search AND BM25 (after FT.OPTIMIZE)
redis-cli FT.ADDTEXT products 1 "Wireless headphones with active noise cancellation"
redis-cli FT.OPTIMIZE products
redis-cli FT.SEARCH products BM25 "noise cancellation" K 3

FT.SEARCHTEXT

FT.SEARCHTEXT <index> <query_text> [K <k>]

Semantic search over text documents added via FT.ADDTEXT:

  1. Embeds query_text via EmbeddingClient
  2. Searches the per-worker semantic HNSW (ef=32, default k=10)
  3. Returns matching doc_id strings as a RESP array
redis-cli FT.SEARCHTEXT products "noise cancelling audio" K 3
# → 1) "doc:1"

AI.CHAT

AI.CHAT <prompt> [CONTEXT <index> <text> [K <k>]]

Full RAG pipeline in a single RESP3 command:

  1. (optional) Embeds text, searches HNSW for top-k doc_ids, retrieves stored texts
  2. Builds LLM prompt: "Context:\n<docs>\n\nUser: <prompt>" (or just <prompt> if no CONTEXT)
  3. Sends HTTP/1.0 POST to /v1/chat/completions
  4. Parses "content":"..." from JSON response, returns as RESP bulk string
# Simple chat
redis-cli AI.CHAT "What is the capital of France?"

# RAG: retrieve context then generate
redis-cli AI.CHAT "Which audio product do you recommend?" CONTEXT products "audio" K 3
# → "Based on the available products, I recommend the wireless headphones..."

AI.SEMANTIC_CACHE

Caches LLM responses keyed by semantic similarity of the input query:

AI.SEMANTIC_CACHE SET "What is the capital of France?" "+Paris\r\n"
AI.SEMANTIC_CACHE GET "capital of France?"    → +Paris
AI.SEMANTIC_CACHE GET "unrelated query"       → $-1 (cache miss)

Workspace audit trail

A cache hit tells you what the model answered, not what it had in mind while answering. WORKSPACE records the caller's snapshot — e.g. the retrieved context or top-k, base64 — at cache-WRITE time, and EXPLAIN reads it back with no inference infrastructure required. That is the point: a compliance reviewer can ask what was in the workspace without being able to run the model.

AI.SEMANTIC_CACHE SET "capital of France?" "Paris" WORKSPACE "<base64 lens top-k>"
AI.SEMANTIC_CACHE EXPLAIN "capital of France?"      → "<base64 lens top-k>"
AI.SEMANTIC_CACHE EXPLAIN "never cached"            → nil
AI.SEMANTIC_CACHE GET "capital of France?" WITHWORKSPACE
                                                     → %2 response/workspace (RESP3)
                                                     → *4 flat pairs        (RESP2)

The blob is opaque to the server — stored and returned, never interpreted — so the lens format can change without a wire change. An entry written without WORKSPACE reports nil rather than an empty string, so "no audit trail" is distinguishable from "an empty one".

WITHWORKSPACE is opt-in and the default GET reply is unchanged, on both protocols. Always returning a RESP3 map would break every RESP3 client already reading GET as a bulk string, and a wire break to add observability is the wrong trade. EXPLAIN needs no flag.

Implementation note: workspaces[] runs parallel to responses[], and other callers (MCP memory in ai.mojo, RAG ingest in vector.mojo) append to responses directly without knowing about workspaces. cache_set therefore PADS to count before appending rather than assuming alignment — if the two ever drift, EXPLAIN returns some other entry's audit trail, which is worse than returning none.

See src/network/semantic_cache.mojo and src/network/embedding_client.mojo. Test: tests/test_gh115_workspace.py. It ENV_SKIPs without an embedding backend, because every SET is a silent no-op there and the failures would all be about the missing backend.


AI.ROUTE.* — semantic load balancer

Command Pion path GLIDE Description
AI.ROUTE.REGISTER SLOW ❌ Register inference node with semantic centroid embedding + optional CAPACITY
AI.ROUTE.UPDATE SLOW ❌ Update a node's centroid embedding (as KV cache evolves)
AI.ROUTE SLOW ❌ Route query embedding to best node by cosine similarity; 0.14ms, 7K QPS
AI.ROUTE.REMOVE SLOW ❌ Remove node from routing table
AI.ROUTE.INFO SLOW ❌ Per-node stats: routed count, capacity, endpoint

Routing strategy: FP32 brute-force cosine for <=16 nodes (perfect accuracy); HNSW O(log N) for >16 nodes. 88% routing accuracy with Ollama embeddings.

Requires --kvcache flag.


Embedding backends and how the first request just works: Embeddings.