Pops correctly from the first non-empty key. Does not block: the TIMEOUT argument is parsed and ignored, so an all-empty key set answers nil immediately instead of waiting. BLPOP k 0 will not wait for a producer
BRPOP
✅
✅
🟡
SLOW
✅
Right-hand form; same non-blocking caveat as BLPOP
BLMOVE
✅
✅
❌
—
✅
Blocking
BLMPOP
✅
✅
❌
—
✅
Blocking
LPUSHX
✅
✅
✅
SLOW
✅
Pushes only if the key exists; 0 otherwise
RPUSHX
✅
✅
✅
SLOW
✅
Right-hand form
RPOPLPUSH
✅
✅
✅
SLOW
✅
Deprecated in Redis 6.2 for LMOVE. Validates BOTH keys before moving anything
Streams are fully implemented and WAL-persisted (snapshot v2 + WAL cmd 23 XADD /
27 XDEL; INFO reports streams_persisted:1). Consumer groups are the one gap —
every X*GROUP/pending/claim command refuses explicitly rather than faking state.
Command
Redis 8
Valkey 8
Pion
Pion path
GLIDE
Notes
XADD
✅
✅
✅
SLOW
✅
Real IDs (ms-seq, * auto), stored and WAL-persisted (cmd 23)
XREAD
✅
✅
✅
SLOW
✅
Returns entries; BLOCK is parsed but never blocks (returns at once)
Message delivery works (RESP2 and RESP3 push). The one limitation is cross-worker:
with -w N > 1 a publisher and subscriber on different workers do not see each
other (shared-nothing keyspace), so pub/sub is coherent at -w 1 (the default)
or when both connections land on the same worker — see operations.md §2b.
Command
Redis 8
Valkey 8
Pion
Pion path
GLIDE
Notes
SUBSCRIBE
✅
✅
✅
SLOW
✅
Confirmation + live message delivery
UNSUBSCRIBE
✅
✅
✅
SLOW
✅
PUBLISH
✅
✅
✅
SLOW
✅
Returns the subscriber count; delivers within a worker
Engine: Lua 5.1.5 (PUC-Rio reference implementation), statically linked. One VM per worker (shared-nothing). Coroutine-based: redis.call() yields to host for dispatch — no callbacks.
Sandbox: Only base, table, string, math, cjson libraries loaded. No io, os, debug, package. Functions removed: dofile, loadfile, loadstring, print (use redis.log instead). Memory limit: 1 MB per execution. Instruction limit: 1M instructions per execution.
Per-connection tx state (tx_in_multi[fd]): queues commands, validates names at QUEUE time against the generated command table, and answers EXEC with -EXECABORT on an unknown one
EXEC
✅
✅
✅
SLOW
✅
Runs the queued commands atomically and returns their replies as an array; -EXECABORT if any was rejected at QUEUE time
DISCARD
✅
✅
🟡
SLOW
✅
Returns +OK
WATCH
✅
✅
✅
SLOW
✅
Monitors keys for changes between WATCH and EXEC; EXEC returns null if any watched key was modified. Per-fd version tracking via key_versions[65536] array, bumped by SET/DEL/HSET in fast path
Real auth: with --requirepass/--tenant, AUTH <pw> gates every command (-NOAUTH before, -WRONGPASS on a bad password); with no password set, AUTH replies the Redis error, not +OK
HELLO
✅
✅
✅
SLOW
✅
HELLO 3 switches the connection to RESP3 (map/push replies); HELLO/HELLO 2 stay RESP2
SHUTDOWN
✅
✅
✅
SLOW
✅
NOSAVE supported. A plain SIGTERM drains the WAL first
Add text doc: store in keyspace + BM25 doc set (integer ids), embed via HTTP → insert into semantic HNSW
FT.SEARCHTEXT
SLOW
❌
Semantic text search over docs added via FT.ADDTEXT
AI.COMPLETE
SLOW
❌
Semantic cache check → LLM call on miss → cache result
AI.SEMANTIC_CACHE GET
SLOW
❌
Embed query → HNSW search → return cached response if cos sim ≥ threshold
AI.SEMANTIC_CACHE SET
SLOW
❌
Embed query → insert into HNSW → store response
AI.CHAT
SLOW
❌
Full RAG in one command: retrieve context → build prompt → call LLM
AI.FLARE LOAD
SLOW
❌
Add text to in-Mojo FLARE knowledge base
AI.FLARE RUN
SLOW
❌
Mid-generation retrieval loop (FLARE algorithm in Mojo)
AI.FLARE INFO
SLOW
❌
FLARE KB stats
AI.KNN_LM.CREATE
SLOW
❌
Allocate token-id-tagged kNN datastore: <ds_id> <dim> [<max_entries>] (default max=100K). Up to 16 datastores per worker. Substrate enabled when --kvcache or --inference is on.
AI.KNN_LM.STORE
SLOW
❌
Append one (token_id, embedding) pair: <ds_id> <next_token_id> <emb_blob>. Auto-builds HNSW once count crosses 5000.
AI.KNN_LM.STOREBATCH
SLOW
❌
Bulk append: <ds_id> <n> <ids_blob> <emb_blob> (ids: n × Int32 LE; emb: n × dim × Float32 LE).
AI.KNN_LM.QUERY
SLOW
❌
Top-k kNN: <ds_id> <k> <emb_blob>. Returns k × 8 bytes packed as <Int32 LE token_id><Float32 LE distance>. Brute-force scan below 5K, focused FP32 HNSW above. Sub-linear scaling, ~0.94 ms median at 30K entries.
AI.KNN_LM.INFO
SLOW
❌
count=N dim=D max_entries=M
AI.KNN_LM.DROP
SLOW
❌
Free datastore + HNSW graph buffers
NEURON.PKM.CREATE
SLOW
❌
Allocate a product-key memory table: <table> <dim> <n_slots> [VDIM <v>] [VALTYPE F32\|F16]. n_slots must be a perfect square S² (S ≤ 4096); dim must be even. Up to 8 tables per worker. Enabled by --kvcache / --inference.
NEURON.PKM.SETKEYS
SLOW
❌
Load codebook half 0 or 1: <table> <half> <blob> — S × (dim/2) Float32 LE. Builds the INT8 mirror. Both halves required before QUERY.
NEURON.PKM.SETVALS
SLOW
❌
Write value rows: <table> <off> <n> <blob> — n × vdim in VALTYPE. Value matrix is allocated on the first call.
NEURON.PKM.QUERY
SLOW
❌
Exact top-k over all n_slots: <table> <k> <q_blob> [FAST]. nq = len(q_blob)/(dim·4) heads share one codebook pass. Returns nq·k × 8 bytes as <Int32 LE slot_id><Float32 LE score>, descending, padded (-1, -inf). dim=896 k=32 1M slots: 0.073 ms (vs 5.13 ms for AI.KNN_LM.QUERY at the same shape).
NEURON.PKM.FFN
SLOW
❌
Fused lookup + softmax-weighted value read: <table> <k> <q_blob> [FAST] [TEMP <t>] → nq × vdim × 4 bytes Float32 LE. Value rows never cross the wire.
Expert weights served from a tiered cache (per-worker RAM LRU → SSD → network)
so a MoE model larger than device RAM runs at interactive cache-hit latency.
Without --moe-cache every command answers -UNAVAILABLE naming the flag,
rather than a plausible empty result.
Asynchronous warm; returns before the fetch completes
MOE.EXPERT.PIN <model_id> <layer> <expert>
SLOW
+OK
Protects an expert from LRU eviction
MOE.EXPERT.UNPIN <model_id> <layer> <expert>
SLOW
+OK
MOE.EXPERT.INFO <model_id>
SLOW
JSON / -UNAVAILABLE
Per-model manifest
MOE.EXPERT.STATS
SLOW
JSON
Always available, even with the tier disabled — {"stage":1,"enabled":false,...}
MOE.EXPERT.HIST <model_id>
SLOW
JSON / -UNAVAILABLE
Per-distribution access histograms
MOE.EXPERT.PRUNE <model_id> <layer> <expert> [on]
SLOW
+OK / -ERR
HIST namespaces are per-distribution but PRUNE is global. A prune driven by one narrow histogram wrecks perplexity on other distributions — use the union-safe client policy
16c. Fixed-Size State Cache — STATE.* (--kvcache)¶
A distinct surface from V-store: a fixed-size, overwrite-in-place buffer for
per-request recurrent state, not an evicting cache. Every command answers
-ERR state cache not enabled (use --kvcache) when the flag is absent.
Store KV cache tensor with HNSW-indexed embedding key. Experimental — TTL/MODEL stubs, capacity-capped 1000 entries, no LRU. See the external-parametric-memory Stage-2 reframe results §5.
KV.FETCH ⚠️
SLOW
❌
Fetch nearest cached tensor by cosine similarity. Experimental — MODEL arg silently ignored on FETCH.
KV.INFO
SLOW
❌
KV cache store statistics (entries, blob bytes, hits, misses, capacity)
~~KV.EVICT~~
none
❌
Not implemented — documented in source comments only, no handler.
KV.PREFIX.REGISTER
SLOW
❌
Register namespace for shared prefix KV cache; creates <ns>_pk and <ns>_pv V-store sessions. Optional PREFILL_MS <ms> reports the client's measured cold prefill for the value receipt. See doc/shared_kv_cache.md.
KV.PREFIX.LOOKUP
SLOW
❌
Returns +HIT if both K and V sessions exist for the namespace, +MISS otherwise
KV.PREFIX.INFO
SLOW
❌
Global stats: registered prefix count, total tokens, total fetches
V.CREATE
SLOW
❌
Create V-store session for token-ID-indexed value storage (used by KV.PREFIX.* and externalized attention)
V.STOREBATCH
SLOW
❌
Append a batch of FP32 values quantized to session format (int8/turbo4/turbo3/turbo2/fp16)
V.FETCH
SLOW
❌
Fetch values by token ID list, or contiguous range via RANGE start end (single round-trip per layer; bypasses 64-token RESP frame limit)
KV.PREFIX.BLOCKS
SLOW
❌
Per-block visibility for a stored prefix
KV.PREFIX.MEMBERSHIP
SLOW
❌
Which blocks of a prefix are resident
KV.PREFIX.OWNER
SLOW
❌
Which worker holds a prefix. -w N is N independent keyspaces, so this is how a client finds the right one
KV.PREFIX.WARM
SLOW
❌
Promote a cold-tier prefix back to RAM
KV.PREFIX.SAVE
SLOW
❌
Persist a prefix to the cold tier
V.SNAPSHOT
SLOW
❌
Point-in-time V-store snapshot
V.COMMIT
SLOW
❌
Seal a snapshot
V.RESTORE
SLOW
❌
Reload a sealed snapshot
V.INFO
SLOW
❌
Per-session or global V-store statistics
ATTEND.PREFIX.STORE
SLOW
❌
Push K/V to native Metal SDPA session cache; resident until DROP / LRU eviction. Body: <session_id> <layer_id> <H> <N> <D> <K_blob> <V_blob>. See doc/shared_kv_cache.md Stage 2.
ATTEND.PREFIX.LOOKUP
SLOW
❌
Probe the Metal session cache for (sid, layer_id). Returns +HIT if slot exists with non-nil K/V, +MISS otherwise. Symmetric to KV.PREFIX.LOOKUP for V-store state. Required for stage-2-aware client PionPromptCache.lookup to avoid stale V-store hits causing skipped _stage2_push_cold.
ATTEND.PREFIX.QUERY
SLOW
❌
Run M=1 single-query attention on cached K/V (Q-only on wire). Body: <session_id> <layer_id> <H> <D> <top_k> <Q_blob>. Returns H*D float32. Native Metal SDPA path (sdpa_q1_fp32/fp16); supports D ∈ {32,64,96,128,160,192,256,512}. 146× faster than ATTEND.QUERYBATCH at H=8 N=2048.
Store opaque byte blob (serialized SSM recurrent state — Mamba [conv_state, ssm_state], RWKV-7 size=3 tuple, or any future linear-attention family). Body: <session_id> <layer_id> <state_blob>. Server never parses the blob.
SSM.PREFIX.FETCH
SLOW
❌
Retrieve previously-stored blob for (sid, layer_id). Returns bulk string or $-1 (miss).
SSM.PREFIX.DROP
SLOW
❌
Drop stored blob. Body: <session_id> [<layer_id>]. Layer omitted = drop all layers for the session. Idempotent.
ATTEND.CREATE
SLOW
❌
Create attention session with key_dim and value_dim
Enable speculative RAG for a session (creates trajectory ring buffer of last 10 embeddings)
RAG.QUERY
SLOW
❌
Query with speculation: check pre-computed cache first (cosine > 0.9), fall back to live HNSW search
RAG.SPECULATE.INFO
SLOW
❌
Per-session stats: hit rate, trajectory length, predictions outstanding
Prediction model: predicted = current + alpha * (current - previous), alpha in [0.5, 1.0, 1.5] (3 predictions per query). 75% hit rate on linear trajectory, 0.2ms prediction+lookup latency.
Requires --kvcache flag.
20. VSET Commands (Redis 8 vector sets — one set per key)¶
Vector sets are separate from FT.* indexes and need no FT.CREATE. Persisted like every other type: VADD/VREM/VSETATTR are WAL-logged and SAVE / BGREWRITEAOF serialize the set.
Note: "Total supported" counts all commands with ✅ or 🟡 status. Pion-native AI commands (sections 16-19) have no Redis equivalent and are counted separately.
Pion Fast Path (33 commands, zero-alloc dispatch)¶
Python redis-py sends INCRBY key 1, not INCR key. Both work; INCR takes the fast path.
Multi-worker + AI
When --flare / --emb-enabled is set, Pion auto-caps to -w 1. SemanticCache is per-worker and cannot share state across workers.
SCAN cursor semantics
SCAN is implemented but returns all keys in a single sweep (cursor always returns 0 on second call). Applications that rely on incremental cursor-based iteration may need adjustment.
Blocking commands
BLPOP/BRPOP pop correctly but do not block: an all-empty key set answers nil at once. The other blocking forms are not implemented. Poll with LPOP/RPOP/ZPOPMIN instead.
Lua scripting
EVAL/EVALSHA/SCRIPT + FUNCTION LOAD/LIST/DELETE/FLUSH/STATS + FCALL fully implemented with Lua 5.1.5 VM (sandboxed, cjson). redis.call() supports 30 commands (see table above). FUNCTION DUMP/RESTORE are stubs.
GLIDE connects in cluster or standalone mode. With Pion:
- Standalone mode works today: GlideClient.create(GlideClientConfiguration(..., port=1974))
- Cluster mode (GlideClusterClient) requires --cluster flag on Pion; sends CLUSTER SLOTS/SHARDS at connect time; Pion responds correctly
- RESP3 negotiation (HELLO 3): supported; the connection switches to RESP3 replies
- Auth (AUTH): supported when the server runs with --requirepass