Skip to content

FT.* and VSET — vector search

HNSW with SIMD INT8 kernels, the RediSearch-shaped FT.* protocol, BM25 and hybrid retrieval, and Redis 8 vector sets (one set per key — see below). One index per server; L2 and COSINE, anything else refused at FT.CREATE.

FT.*

Pion is wire-compatible with Redis Stack's vector search subset. Commands are dispatched in slow_path.mojo by matching tl (token length) and the first bytes of tp (token pointer).

FT.CREATE

FT.CREATE <index> ON HASH SCHEMA … <field> VECTOR HNSW 6
    TYPE FLOAT32 DIM <d> DISTANCE_METRIC L2 M <m> EF_CONSTRUCTION <ef>
- Captures <index> (first token after the command) as hnsw.index_name. - Scans remaining tokens for the VECTOR keyword; captures the immediately preceding token as vector_field_name. - Scans for EF_CONSTRUCTION keyword; parses the following integer as hnsw.ef_construction. - DISTANCE_METRIC — resolved in a pre-scan, before any index state is written, so refusing one is a true no-op. L2 → distance_metric = 0, COSINE → 1, anything else → -ERR unsupported DISTANCE_METRIC. Omitting the keyword gives L2.

The search itself is unchanged either way: the beam kernels compute query_norm + node_norm - 2·dot, i.e. squared L2 over affinely-quantized codes, so L2 is what this engine has always answered. COSINE is implemented by L2-normalizing the FP32 vector at ingest and the query at quantization time — on unit vectors, L2 ordering is cosine ordering.

It is done there and not in the distance function because the quantizer is affine with an offset: an offset cancels in a difference but not in a dot product, so a cosine computed from quantized dots would be wrong. The metric is persisted in index-file header word 24 (a warm restart that forgot it would query a normalized graph with a raw query) and published to SharedHNSWView.pre_distance_metric so a cross-worker HSET normalizes the same way the querying worker will.

  • Sets hnsw.index_ready = True.
  • Subsequent multi-field HSET calls will route the matching field to the HNSW index.

HSET (multi-field, fast path)

HSET <key> <f1> <v1> <f2> <v2> … <vec_field> <float32_bytes>
- Handled in fast_path.mojo when num_args >= 6 and (num_args - 2) % 2 == 0. - Each field is stored in the HASH in keyspace. - If a field's name matches vector_field_name and its byte length equals hnsw.dim * 4, it is routed to hnsw.add_vector(key_id, fp32_ptr). - Non-vector fields (id, metadata, …) are stored normally in the hash. - Returns :N\r\n where N = number of fields written.

FT.SEARCH

FT.SEARCH <index> "*=>[KNN <k> @<field> $blob EF_RUNTIME <ef> AS score]"
    PARAMS 2 blob <float32_bytes> [DIALECT 2] [SORTBY score] [LIMIT 0 <k>]
- Verifies <index> matches hnsw.index_name (case-sensitive byte comparison; skipped if no index has been created). - Extracts k by scanning the query string for KNN. - Extracts EF_RUNTIME <ef> from the query string if present; falls back to hnsw.ef_runtime (default 150). - Extracts LIMIT offset count from the outer token stream; if count < k, uses count as the result cap. - Finds the query vector blob in the PARAMS section by matching its byte size to hnsw.dim * 4 (size-based routing — no name coupling between PARAMS and schema). - Calls hnsw.search_fp32_scored(blob_ptr, k, dist_scores, ef_query) with the resolved ef. - Returns a RESP2 FT.SEARCH-format response:

*(1 + 2*N)\r\n  :N\r\n
[for each result:
  $<id_len>\r\n<id_str>\r\n   ← doc key
  *4\r\n
  $2\r\nid\r\n  $<id_len>\r\n<id_str>\r\n
  $5\r\nscore\r\n  $<len>\r\n<L2_distance>\r\n
]

This shape is emitted at every worker count. Doc-key resolution order: shared hk_keys_buf slot→key map (one direct load; written by the fast-path vector HSET), then the keyspace __hk__<slot> probe (keys >31B that don't fit the shared map's 32-byte slots), then the numeric slot id. The id value is the doc hash's own id field when the doc hash is resolvable (what RETURN 1 id consumers and tests/test_multiworker_recall.py parse), else the doc key, else the slot number.

FT.INFO

FT.INFO <index>
Returns a 4-element array (index_name, <actual_name>, num_docs, <count>) when index_ready, where <actual_name> is the exact index name stored at FT.CREATE time and <count> is hnsw.num_nodes. Returns -ERR Unknown index name when no index has been created.

FT.OPTIMIZE

FT.OPTIMIZE <index>
Triggers batch graph construction from the FP32 ingest buffer. Calls hnsw.build_index_from_shared() (if shared buffer has data) or hnsw.build_index() (local). After build: compact_vectors() reorders vectors in BFS order, publish_to_shared() makes the index available to other workers via SharedHNSWView. Returns +OK.

Ingest → optimize → search is a one-way contract. Every HSET must land before FT.OPTIMIZE; the FP32 ingest buffer is freed after the build, so an HSET issued after FT.OPTIMIZE is stored as a hash but not indexed (no error, and FT.SEARCH will not find it). For incremental inserts, FT.DROPINDEX → re-ingest everything → FT.OPTIMIZE again.

FT.DROPINDEX

FT.DROPINDEX <index>
Calls hnsw.reset_index() — zeroes num_nodes, resets entry_point_id = -1, clears node_map and visited_map, resets index_name_len = 0, resets global_min/max to calibration defaults, sets index_ready = False. Returns +OK.

FT.HYBRID <index> <text_query> <vector_blob> [K <k>] [ALPHA <a>]
Executes parallel BM25 full-text search and HNSW vector search in a single command, fusing results via Reciprocal Rank Fusion (RRF).

Parameters: - <index> — index name (must match a previously created FT.CREATE index) - <text_query> — plain-text query string for BM25 scoring - <vector_blob> — raw Float32 bytes (dim * 4 bytes) for HNSW nearest-neighbor search - K <k> — number of results to return (default 10) - ALPHA <a> — blending weight: 0.0 = full vector, 1.0 = full BM25, 0.5 = equal (default 0.5)

Fusion formula (RRF):

score = vec_weight / (60 + rank_vec) + bm25_weight / (60 + rank_bm25)
where vec_weight = 1 - alpha and bm25_weight = alpha. The constant 60 is the standard RRF damping factor.

Internal behavior: 1. HNSW search with K * 4 oversampling (retrieves 4x candidates for quality) 2. BM25 search with K * 4 oversampling 3. RRF fusion: candidates from both result sets are merged by document ID, scored, and sorted 4. Top-K results returned

Response format: Standard FT.SEARCH RESP2 format with RRF combined scores (same wire format as FT.SEARCH).

Example:

redis-cli FT.HYBRID myindex "machine learning transformers" \
    $(python3 -c "import struct; print(struct.pack('1536f', *[0.01]*1536))") \
    K 10 ALPHA 0.5

See also: doc/embeddings.md for text embedding generation.


BM25 and hybrid retrieval

Pion includes a built-in BM25 full-text search engine, constructed during FT.OPTIMIZE from the TEXT schema field declared in FT.CREATE. The schema's first TEXT field is the one indexed.

Ingest — two paths, both work

The doc set is union(HNSW nodes, FT.ADDTEXT registrations), so a corpus needs no vectors at all to be lexically searchable.

FT.CREATE idx SCHEMA body TEXT vec VECTOR HNSW 6 TYPE FLOAT32 DIM 512 DISTANCE_METRIC L2

# (a) text-only — no vector, no embedding sidecar needed
FT.ADDTEXT idx 7 "the --fa-window flag caps the sliding attention span"

# (b) alongside a vector — any key shape; the __hk__<slot> reverse map
#     written by the vector HSET is what points BM25 at the text
HSET doc:7 vec <512 float32 LE>
HSET doc:7 body "the --fa-window flag caps the sliding attention span"

FT.OPTIMIZE idx                # builds the inverted index — required after ingest
FT.SEARCH idx BM25 "sliding attention" K 10

FT.ADDTEXT doc ids must be non-negative integers to be BM25-indexable — hits are returned as integer ext_ids, so a string id has nothing to map back to. A string id still works for FT.SEARCHTEXT, which searches the semantic-cache HNSW rather than the inverted index.

FT.SEARCHTEXT runs against that separate, per-worker semantic-cache graph, and it comes with limits the BM25 side does not have: 10,000 documents per worker, no clearing on FT.DROPINDEX, and O(N²) bulk ingest. See doc/embeddings.md § "FT.ADDTEXT / FT.SEARCHTEXT limits and gotchas" before building a text knowledge store on it.

FT.SEARCH … BM25 against an index with no inverted index at all returns an error naming the missing step, not an empty array — a caller could not otherwise tell "never optimized" from "matched nothing". A built index that genuinely matches nothing still returns the empty form.

BM25 state is per-worker and not persisted: it is never published to the shared view and save_to_disk does not carry it, so re-run FT.OPTIMIZE after a restart, and prefer -w 1 for lexical workloads.

One index per server

A Pion server holds one index at a time. FT.CREATE overwrites the shared schema, FT.OPTIMIZE rebuilds the single HNSW graph and the single set of BM25 postings, and pion.hnsw.0 is written by whichever index optimized last. Running FT.OPTIMIZE on index B therefore replaces whatever was serving index A — including a concurrently-serving production index.

This is the contract, not a bug, and a query naming the displaced index gets an error rather than *0, which would be indistinguishable from a genuine miss:

FT.SEARCH clobA BM25 "alpha bravo"
-ERR index 'clobA' has no BM25 postings - this server holds one index at a time and the
     current postings belong to 'clobB' (a later FT.OPTIMIZE replaced them); re-ingest
     and FT.OPTIMIZE 'clobA' to serve it again

FT.SEARCH clobA *=>[KNN 2 @vec $q] PARAMS 2 q <blob>
-ERR index 'clobA' is not the index loaded on this server - it holds one index at a
     time and currently serves 'clobB' (a later FT.CREATE/FT.OPTIMIZE replaced it)

FT.HYBRID refuses the same way rather than quietly degrading to a vector-only result scored against another index's postings. Index names are compared byte-exactly (case-sensitive), matching the pre-existing KNN check. The BM25 [<index>]: N terms, M docs log line names its index so a cross-index rebuild is identifiable after the fact.

Recovery is what the error says: re-ingest the displaced index's documents and re-run FT.OPTIMIZE for it. Ownership is recorded at build time, so FT.CREATE B alone does not displace A's postings — only FT.OPTIMIZE does.

Operational consequence: do not point a scratch/probe index at a server that is serving a real one. A 2-document probe's FT.OPTIMIZE replaces a 300-passage index.

Indexing

  • Built automatically by FT.OPTIMIZE, from scratch each time
  • Tokenization: hard-split on whitespace and structural punctuation (" ' \ ( ) [ ] { } < > | \ , ; : ! ?); the remaining punctuation (- . _ + = * / & % @ # ~ $ ^) is trimmed from token edges but **kept inside** a token, so--fa-windowindexes asfa-windowrather thanfa+window`. Tokens are lowercased and hashed with FNV-1a. Index-time and query-time tokenization go through one function and cannot drift.
  • No caps. Vocabulary and per-document term count grow on demand, and token length is unbounded.

Scoring

  • BM25 with defaults k1 = 1.2, b = 0.75, tunable per query (see below)
  • IDF: log(1 + (N - df + 0.5) / (df + 0.5)) where N = documents in the BM25 doc set, df = document frequency. This is Lucene's non-negative form; rank_bm25's BM25Okapi instead uses log(N - df + 0.5) - log(df + 0.5) with an epsilon floor for negative values. The difference is measurable but small (see below).
  • Lookup: term hashes are stored sorted; query terms resolve by binary search, O(log V) per term. Up to 64 unique query terms are scored.

Querying

FT.SEARCH <index> BM25 "query text" [K <k>] [K1 <f>] [B <f>]
FT.HYBRID <index> "query text" <vector blob> [K <k>] [ALPHA <a>] [K1 <f>] [B <f>]

K1 (0–100) sets term-frequency saturation; B (0–1) sets length normalisation, where B 0 disables it entirely. Out-of-range values are rejected rather than silently clamped. Omitting both reproduces the documented defaults exactly.


Metadata filters

FT.SEARCH filters a KNN query by TAG and NUMERIC fields declared in the FT.CREATE schema. Every form below is applied; anything else is refused with an error, never ignored.

FT.SEARCH idx "*=>[KNN 10 @vec $v]"               FILTER @price:[10 50]   PARAMS 2 v <blob> DIALECT 2
FT.SEARCH idx "(@price:[10 50] @cat:{a|b})=>[KNN 10 @vec $v]"             PARAMS 2 v <blob> DIALECT 2
FT.SEARCH idx "*=>[KNN 10 @vec $v]"               FILTER price 10 50      PARAMS 2 v <blob>   # RediSearch legacy
FT.SEARCH idx "*=>[KNN 10 @vec $v]"               FILTER cat=a            PARAMS 2 v <blob>
FT.SEARCH idx "*=>[KNN 10 @vec $v]"               FILTER price [10 50]    PARAMS 2 v <blob>
  • NUMERIC @f:[lo hi] is inclusive; -inf / +inf work. Exclusive ( bounds are refused.
  • TAG @f:{a} matches the value exactly; @f:{a|b} matches either.
  • Up to 4 clauses, ANDed — several FILTER arguments and/or clauses in the query prefix. | between clauses (OR) and negation are refused.
  • Filtering is applied to KNN candidates, widening the candidate set (4×, 16×, … up to 16,384 or the whole index) until k rows pass — a selective filter costs more search, it does not silently return fewer rows while matches exist below that bound.
  • Field names are case-sensitive, as in the hash.

The query itself must be <prefilter>=>[KNN k @field $param [EF_RUNTIME n] [AS alias]] with the vector in PARAMS by name; an unknown form, an unknown vector field, a missing $param or a blob of the wrong size is an error.

Scores are the index metric's distance for each returned row — squared L2 under L2, 1 - cos under COSINE — computed from the stored vectors (FP32 re-rank copy, or the INT8 codes dequantized), not the beam's internal quantized distance. Rows are sorted by that score.


Vector sets (Redis 8 VSET commands)

One vector set per key, created by its first VADD, with its own dimension. Search is exact — every element is scored against the query — so there is no recall loss and nothing to tune, at a cost linear in the set's size. Scores are Redis's: (1 + cos) / 2, where 1 is identical. Options that only steer Redis's HNSW graph (Q8, NOQUANT, BIN, M, EF) are accepted and change nothing; REDUCE, FILTER, VEMB … RAW and VLINKS are refused with an error rather than ignored. Vector sets are WAL-persisted and snapshotted: they survive a restart.

Command Redis 8 Valkey 8 Pion Pion path Notes
VADD ✅ ❌ ✅ SLOW VADD key (FP32 <blob> \| VALUES n v…) <element> [SETATTR json] — :1 new, :0 updated; Q8/NOQUANT/BIN/M/EF/CAS accepted (no effect); REDUCE refused
VSIM ✅ ❌ ✅ SLOW VSIM key (ELE e \| FP32 blob \| VALUES n v…) [WITHSCORES] [WITHATTRIBS] [COUNT n] [EPSILON d] — exact scan; score (1+cos)/2; FILTER refused
VCARD ✅ ❌ ✅ SLOW Elements in the set; 0 for a missing key
VDIM ✅ ❌ ✅ SLOW The set's dimension (fixed by its first VADD)
VINFO ✅ ❌ ✅ SLOW Redis's nine fields; quant-type f32, graph fields 0 (no graph)
VISMEMBER ✅ ❌ ✅ SLOW 1 / 0
VSETATTR ✅ ❌ ✅ SLOW 1 set, 0 missing key or element; empty string removes
VGETATTR ✅ ❌ ✅ SLOW JSON string or nil
VEMB ✅ ❌ ✅ SLOW The stored vector as floats; RAW refused
VRANDMEMBER ✅ ❌ ✅ SLOW Random element(s); count>0 distinct, count<0 with repeats
VREM ✅ ❌ ✅ SLOW 1 / 0; the last element removes the key
VRANGE ✅ ❌ ✅ SLOW VRANGE key start end [count] — lexicographic (-, +, [x, (x)
VLINKS ✅ ❌ — SLOW Refused: an exact set has no HNSW graph

Vector sets are separate from FT.* indexes and need no FT.CREATE. Persisted like every other type: VADD/VREM/VSETATTR are WAL-logged and SAVE / BGREWRITEAOF serialize the set.


Engine internals and benchmarks: Vector engine.