Vector Engine: High-Performance Search¶
Pion is a unified, shared-nothing database engine built in Mojo. Its Vector Engine provides real-time approximate nearest-neighbor (ANN) search with Redis Stack-compatible FT.* commands, enabling drop-in use with any client that targets Redis as a vector store (VectorDBBench, LangChain, LlamaIndex, etc.).
Architecture Overview¶
Client (VectorDBBench / redis-py / any RESP client)
│ FT.CREATE / HSET / FT.SEARCH
▼
slow_path.mojo ← FT.* dispatcher
fast_path.mojo ← multi-field HSET router
│
├─ keyspace (SlabHashMap) — stores HASH docs (id, metadata, …)
└─ HNSWGraph — stores quantized vectors + graph
Each worker owns a private HNSWGraph (shared-nothing). A single shared listen fd distributes connections via accept() scheduling (no SO_REUSEPORT). After FT.OPTIMIZE, the building worker publishes read-only index pointers via SharedHNSWView; other workers lazy-borrow these pointers on the first FT.SEARCH. All mutable search state (visited_map, query_int8) remains per-worker.
HNSW Index (src/vector/hnsw.mojo)¶
The core index is an HNSW (Hierarchical Navigable Small Worlds) graph.
Graph structure¶
- Level 0 contains all vectors; higher levels contain exponentially fewer nodes (controlled by
Mandml = 1/ln(M)). - Each node stores a quantized
Int8vector + neighbor lists (capacity2*Mat level 0,Mat higher levels) allocated from a pre-allocated neighbor pool. node_map[id] → internal_idxallows O(1) lookup by user-supplied integer ID.
Insert (add_vector)¶
- Assign a random level (
_random_level). - Quantize the
Float32input vector toInt8via calibrated SQ8. - Greedy descent from
max_leveldown tolevel+1to find best entry point. - For each level from
min(level, max_level)down to 0: beam-search (_search_layer, ef =ef_construction), prune neighbors (heuristic), link bidirectionally.
Search (search_fp32, search_fp32_scored)¶
- Upper levels: greedy traversal using
_dist_fp32_int8(query FP32 × stored Int8). - Level 0: beam search with
MinHeap(candidates) +MaxHeap(results), ef-bounded. - Query is quantized to INT8 once; beam uses
_dist_int8_int8batch kernels (batch-8 and batch-4) — matches the graph build metric, no per-neighbor dequantization. - Slot-space addressing: the gather loop reads each neighbor's compact slot
from
l0_slots(a 33-u32/node mirror ofl0_compact) and computes the vector address ascompact_buffer + slot*compact_stride + compact_hdr, so nonodes[]dereference sits on the critical path.0xFFFFFFFFentries fall back to thenodes[]path. Applies to the INT8 beam and all three quant beams. - Staged prefetch (INT8 beam): gather issues 4 lines (norm header + 256B prefix) per neighbor; the full slot prefetch is issued one SDOT batch ahead inside the scoring loops.
search_fp32_scoredadditionally fills a caller-suppliedList[Float32]with L2 distances in nearest-first order — used byFT.SEARCHto return scores.
Visited-set optimisation¶
Epoch-stamped visited set: visited_epoch[UInt16] per node with a monotonically
increasing cur_epoch — a node is visited this query iff visited_epoch[nid] == cur_epoch,
so there is no per-query memset (full clear only on UInt16 wrap).
Key fields added for FT.* integration¶
| Field | Type | Purpose |
|---|---|---|
vector_field_name |
InlineArray[UInt8, 32] |
Schema field name holding vectors (set by FT.CREATE) |
vector_field_len |
Int |
Byte length of that name |
index_name |
InlineArray[UInt8, 64] |
Index name captured from FT.CREATE argument |
index_name_len |
Int |
Byte length of stored index name |
index_ready |
Bool |
True after FT.CREATE; gates vector routing in HSET |
ef_runtime |
Int |
Default ef for FT.SEARCH when no per-query override (default 150) |
Calibrated Scalar Quantization — SQ8¶
All stored vectors are quantized from Float32 → Int8 at insert time.
- Range: calibrated from the vectors being built, by every build (
HNSWGraph._calibrate): one global range over mean ± 8σ (Welford). The HSET-ingest build calibrates on up to the first 65,536 ingested vectors, streaming mode on the first 1,000. 8σ is wide enough that the few near-constant outlier dimensions typical of embedding data are clipped (they cancel in every L2 difference) while every component that really varies is kept. - Formula:
q = Int8(clamp((v - min) / range * 254) - 127) - Memory: 4× compression vs FP32 (1536 dims → 1536 bytes instead of 6144).
- FP32 buffer ingest:
add_vector()buffers raw FP32 vectors; graph is built in batch onFT.OPTIMIZEcall (6× faster insert — no graph ops during HSET). - Compact buffer: After FT.OPTIMIZE,
compact_vectors()reorders INT8 vectors in BFS traversal order into a contiguous buffer for cache-friendly beam search. Each build (or snapshot load) also derivesl0_slots— the slot-space adjacency mirror the beam kernel addresses vectors through — pluscompact_stride/compact_hdr(INT8:dim+8/8; quant variants: their own stride/0).
SIMD & Compile-Time-Fused Kernels (src/vector/kernels.mojo)¶
Distance computation is the hottest inner loop. Pion uses Mojo's comptime parameters to generate fully-unrolled, fused kernels for common dimensions:
fn l2_distance_fp32_int8_fused_jit[dim: Int](
v1: UnsafePointer[Float32],
v2: UnsafePointer[Int8],
min_val: Float32,
range_val: Float32
) -> Float32:
# Loop unrolled at compile time for dim ∈ {128, 384, 768, 1536}
# Dequantization fused into the accumulation: no separate pass
Specialised variants exist for dim ∈ {128, 384, 768, 1536} — the four most common embedding dimensions. Unknown dimensions fall back to a scalar loop.
Additional kernel variants:
- l2_distance_int8_int8_batch8_jit[dim] — INT8×INT8 batch-8, primary search kernel
- l2_distance_int8_int8_batch4_jit[dim] — INT8×INT8 batch-4, remainder loop
- l2_distance_int8_jit[dim] — scalar INT8×INT8, build-time and fallback
- l2_distance_int4 — 4-bit packed (when use_int4=True)
- hamming_distance_jit[dim_u64] — binary quantization (when use_bq_traversal=True)
Software prefetching (sys.intrinsics.prefetch) is applied to the next neighbor's vector and neighbor-list pointers inside _search_layer and add_vector.
Wire Protocol — FT.* Commands¶
Pion is wire-compatible with Redis Stack's vector search subset. Commands are dispatched in slow_path.mojo by matching tl (token length) and the first bytes of tp (token pointer).
FT.CREATE¶
FT.CREATE <index> ON HASH SCHEMA … <field> VECTOR HNSW 6
TYPE FLOAT32 DIM <d> DISTANCE_METRIC L2 M <m> EF_CONSTRUCTION <ef>
<index> (first token after the command) as hnsw.index_name.
- Scans remaining tokens for the VECTOR keyword; captures the immediately preceding token as vector_field_name.
- Scans for EF_CONSTRUCTION keyword; parses the following integer as hnsw.ef_construction.
- DISTANCE_METRIC — resolved in a pre-scan, before any index
state is written, so refusing one is a true no-op. L2 → distance_metric =
0, COSINE → 1, anything else → -ERR unsupported DISTANCE_METRIC.
Omitting the keyword gives L2.
The search itself is unchanged either way: the beam kernels compute
query_norm + node_norm - 2·dot, i.e. squared L2 over affinely-quantized
codes, so L2 is what this engine has always answered. COSINE is
implemented by L2-normalizing the FP32 vector at ingest and the query at
quantization time — on unit vectors, L2 ordering is cosine ordering.
It is done there and not in the distance function because the quantizer is
affine with an offset: an offset cancels in a difference but not in a dot
product, so a cosine computed from quantized dots would be wrong. The metric
is persisted in index-file header word 24 (a warm restart that forgot it
would query a normalized graph with a raw query) and published to
SharedHNSWView.pre_distance_metric so a cross-worker HSET normalizes the
same way the querying worker will.
- Sets
hnsw.index_ready = True. - Subsequent multi-field HSET calls will route the matching field to the HNSW index.
HSET (multi-field, fast path)¶
- Handled infast_path.mojo when num_args >= 6 and (num_args - 2) % 2 == 0.
- Each field is stored in the HASH in keyspace.
- If a field's name matches vector_field_name and its byte length equals hnsw.dim * 4, it is routed to hnsw.add_vector(key_id, fp32_ptr).
- Non-vector fields (id, metadata, …) are stored normally in the hash.
- Returns :N\r\n where N = number of fields written.
FT.SEARCH¶
FT.SEARCH <index> "*=>[KNN <k> @<field> $blob EF_RUNTIME <ef> AS score]"
PARAMS 2 blob <float32_bytes> [DIALECT 2] [SORTBY score] [LIMIT 0 <k>]
<index> matches hnsw.index_name (case-sensitive byte comparison; skipped if no index has been created).
- Extracts k by scanning the query string for KNN.
- Extracts EF_RUNTIME <ef> from the query string if present; falls back to hnsw.ef_runtime (default 150).
- Extracts LIMIT offset count from the outer token stream; if count < k, uses count as the result cap.
- Finds the query vector blob in the PARAMS section by matching its byte size to hnsw.dim * 4 (size-based routing — no name coupling between PARAMS and schema).
- Calls hnsw.search_fp32_scored(blob_ptr, k, dist_scores, ef_query) with the resolved ef.
- Returns a RESP2 FT.SEARCH-format response:
*(1 + 2*N)\r\n :N\r\n
[for each result:
$<id_len>\r\n<id_str>\r\n ← doc key
*4\r\n
$2\r\nid\r\n $<id_len>\r\n<id_str>\r\n
$5\r\nscore\r\n $<len>\r\n<L2_distance>\r\n
]
This shape is emitted at every worker count. Doc-key resolution order:
shared hk_keys_buf slot→key map (one direct load; written by the fast-path
vector HSET), then the keyspace __hk__<slot> probe (keys >31B that don't fit
the shared map's 32-byte slots), then the numeric slot id. The id value
is the doc hash's own id field when the doc hash is resolvable (what
RETURN 1 id consumers and tests/test_multiworker_recall.py parse), else
the doc key, else the slot number.
FT.INFO¶
Returns a 4-element array (index_name, <actual_name>, num_docs, <count>) when index_ready, where <actual_name> is the exact index name stored at FT.CREATE time and <count> is hnsw.num_nodes. Returns -ERR Unknown index name when no index has been created.
FT.OPTIMIZE¶
Triggers batch graph construction from the FP32 ingest buffer. Callshnsw.build_index_from_shared() (if shared buffer has data) or hnsw.build_index() (local). After build: compact_vectors() reorders vectors in BFS order, publish_to_shared() makes the index available to other workers via SharedHNSWView. Returns +OK.
Ingest → optimize → search is a one-way contract. Every HSET must land before FT.OPTIMIZE; the FP32 ingest buffer is freed after the build, so an HSET issued after FT.OPTIMIZE is stored as a hash but not indexed (no error, and FT.SEARCH will not find it). For incremental inserts, FT.DROPINDEX → re-ingest everything → FT.OPTIMIZE again.
FT.DROPINDEX¶
Callshnsw.reset_index() — zeroes num_nodes, resets entry_point_id = -1, clears node_map and visited_map, resets index_name_len = 0, resets global_min/max to calibration defaults, sets index_ready = False. Returns +OK.
FT.HYBRID — Hybrid Vector + BM25 Search¶
Executes parallel BM25 full-text search and HNSW vector search in a single command, fusing results via Reciprocal Rank Fusion (RRF).Parameters:
- <index> — index name (must match a previously created FT.CREATE index)
- <text_query> — plain-text query string for BM25 scoring
- <vector_blob> — raw Float32 bytes (dim * 4 bytes) for HNSW nearest-neighbor search
- K <k> — number of results to return (default 10)
- ALPHA <a> — blending weight: 0.0 = full vector, 1.0 = full BM25, 0.5 = equal (default 0.5)
Fusion formula (RRF):
wherevec_weight = 1 - alpha and bm25_weight = alpha. The constant 60 is the standard RRF damping factor.
Internal behavior:
1. HNSW search with K * 4 oversampling (retrieves 4x candidates for quality)
2. BM25 search with K * 4 oversampling
3. RRF fusion: candidates from both result sets are merged by document ID, scored, and sorted
4. Top-K results returned
Response format: Standard FT.SEARCH RESP2 format with RRF combined scores (same wire format as FT.SEARCH).
Example:
redis-cli FT.HYBRID myindex "machine learning transformers" \
$(python3 -c "import struct; print(struct.pack('1536f', *[0.01]*1536))") \
K 10 ALPHA 0.5
See also: doc/embeddings.md for text embedding generation.
BM25 Full-Text Search¶
Pion includes a built-in BM25 full-text search engine, constructed during FT.OPTIMIZE from the TEXT schema field declared in FT.CREATE. The schema's first TEXT field is the one indexed.
Ingest — two paths, both work¶
The doc set is union(HNSW nodes, FT.ADDTEXT registrations), so a corpus needs no vectors at all to be lexically searchable.
FT.CREATE idx SCHEMA body TEXT vec VECTOR HNSW 6 TYPE FLOAT32 DIM 512 DISTANCE_METRIC L2
# (a) text-only — no vector, no embedding sidecar needed
FT.ADDTEXT idx 7 "the --fa-window flag caps the sliding attention span"
# (b) alongside a vector — any key shape; the __hk__<slot> reverse map
# written by the vector HSET is what points BM25 at the text
HSET doc:7 vec <512 float32 LE>
HSET doc:7 body "the --fa-window flag caps the sliding attention span"
FT.OPTIMIZE idx # builds the inverted index — required after ingest
FT.SEARCH idx BM25 "sliding attention" K 10
FT.ADDTEXT doc ids must be non-negative integers to be BM25-indexable — hits are returned as integer ext_ids, so a string id has nothing to map back to. A string id still works for FT.SEARCHTEXT, which searches the semantic-cache HNSW rather than the inverted index.
FT.SEARCHTEXT runs against that separate, per-worker semantic-cache graph, and it comes with limits the BM25 side does not have: 10,000 documents per worker, no clearing on FT.DROPINDEX, and O(N²) bulk ingest. See doc/embeddings.md § "FT.ADDTEXT / FT.SEARCHTEXT limits and gotchas" before building a text knowledge store on it.
FT.SEARCH … BM25 against an index with no inverted index at all returns an error naming the missing step, not an empty array — a caller could not otherwise tell "never optimized" from "matched nothing". A built index that genuinely matches nothing still returns the empty form.
BM25 state is per-worker and not persisted: it is never published to the shared view and save_to_disk does not carry it, so re-run FT.OPTIMIZE after a restart, and prefer -w 1 for lexical workloads.
One index per server¶
A Pion server holds one index at a time. FT.CREATE overwrites the shared schema, FT.OPTIMIZE rebuilds the single HNSW graph and the single set of BM25 postings, and pion.hnsw.0 is written by whichever index optimized last. Running FT.OPTIMIZE on index B therefore replaces whatever was serving index A — including a concurrently-serving production index.
This is the contract, not a bug, and a query naming the displaced index gets an error rather than *0, which would be indistinguishable from a genuine miss:
FT.SEARCH clobA BM25 "alpha bravo"
-ERR index 'clobA' has no BM25 postings - this server holds one index at a time and the
current postings belong to 'clobB' (a later FT.OPTIMIZE replaced them); re-ingest
and FT.OPTIMIZE 'clobA' to serve it again
FT.SEARCH clobA *=>[KNN 2 @vec $q] PARAMS 2 q <blob>
-ERR index 'clobA' is not the index loaded on this server - it holds one index at a
time and currently serves 'clobB' (a later FT.CREATE/FT.OPTIMIZE replaced it)
FT.HYBRID refuses the same way rather than quietly degrading to a vector-only result scored against another index's postings. Index names are compared byte-exactly (case-sensitive), matching the pre-existing KNN check. The BM25 [<index>]: N terms, M docs log line names its index so a cross-index rebuild is identifiable after the fact.
Recovery is what the error says: re-ingest the displaced index's documents and re-run FT.OPTIMIZE for it. Ownership is recorded at build time, so FT.CREATE B alone does not displace A's postings — only FT.OPTIMIZE does.
Operational consequence: do not point a scratch/probe index at a server that is serving a real one. A 2-document probe's FT.OPTIMIZE replaces a 300-passage index.
Indexing¶
- Built automatically by
FT.OPTIMIZE, from scratch each time - Tokenization: hard-split on whitespace and structural punctuation (
" ' \( ) [ ] { } < > | \ , ; : ! ?); the remaining punctuation (- . _ + = * / & % @ # ~ $ ^) is trimmed from token edges but **kept inside** a token, so--fa-windowindexes asfa-windowrather thanfa+window`. Tokens are lowercased and hashed with FNV-1a. Index-time and query-time tokenization go through one function and cannot drift. - No caps. Vocabulary and per-document term count grow on demand, and token length is unbounded.
Scoring¶
- BM25 with defaults
k1 = 1.2,b = 0.75, tunable per query (see below) - IDF:
log(1 + (N - df + 0.5) / (df + 0.5))where N = documents in the BM25 doc set, df = document frequency. This is Lucene's non-negative form;rank_bm25'sBM25Okapiinstead useslog(N - df + 0.5) - log(df + 0.5)with an epsilon floor for negative values. The difference is measurable but small (see below). - Lookup: term hashes are stored sorted; query terms resolve by binary search, O(log V) per term. Up to 64 unique query terms are scored.
Querying¶
FT.SEARCH <index> BM25 "query text" [K <k>] [K1 <f>] [B <f>]
FT.HYBRID <index> "query text" <vector blob> [K <k>] [ALPHA <a>] [K1 <f>] [B <f>]
K1 (0–100) sets term-frequency saturation; B (0–1) sets length normalisation, where B 0 disables it entirely. Out-of-range values are rejected rather than silently clamped. Omitting both reproduces the documented defaults exactly.
Metadata Filters¶
FT.SEARCH filters a KNN query by TAG and NUMERIC fields declared in the FT.CREATE schema. Every form below is applied; anything else is refused with an error, never ignored.
FT.SEARCH idx "*=>[KNN 10 @vec $v]" FILTER @price:[10 50] PARAMS 2 v <blob> DIALECT 2
FT.SEARCH idx "(@price:[10 50] @cat:{a|b})=>[KNN 10 @vec $v]" PARAMS 2 v <blob> DIALECT 2
FT.SEARCH idx "*=>[KNN 10 @vec $v]" FILTER price 10 50 PARAMS 2 v <blob> # RediSearch legacy
FT.SEARCH idx "*=>[KNN 10 @vec $v]" FILTER cat=a PARAMS 2 v <blob>
FT.SEARCH idx "*=>[KNN 10 @vec $v]" FILTER price [10 50] PARAMS 2 v <blob>
- NUMERIC
@f:[lo hi]is inclusive;-inf/+infwork. Exclusive(bounds are refused. - TAG
@f:{a}matches the value exactly;@f:{a|b}matches either. - Up to 4 clauses, ANDed — several
FILTERarguments and/or clauses in the query prefix.|between clauses (OR) and negation are refused. - Filtering is applied to KNN candidates, widening the candidate set (4×, 16×, … up to 16,384 or the whole index) until
krows pass — a selective filter costs more search, it does not silently return fewer rows while matches exist below that bound. - Field names are case-sensitive, as in the hash.
The query itself must be <prefilter>=>[KNN k @field $param [EF_RUNTIME n] [AS alias]] with the vector in PARAMS by name; an unknown form, an unknown vector field, a missing $param or a blob of the wrong size is an error.
Scores are the index metric's distance for each returned row — squared L2 under L2, 1 - cos under COSINE — computed from the stored vectors (FP32 re-rank copy, or the INT8 codes dequantized), not the beam's internal quantized distance. Rows are sorted by that score.
Open build and the closed vector library¶
The tuned 1536-dim beam searches — INT8 and the three quantized variants —
and the tuned product-key-memory kernels are distributed as a static library,
libpion_vector, vendored under vendor/pion-vector/<platform>/. Everything
else is open source here: graph build, persistence, the other dimensions,
every quantizer and dequantizer, the Metal shaders, and this document.
| Build | Vector code | Use |
|---|---|---|
pixi run build |
links libpion_vector.a (-D PION_HELD_VECTOR) |
default; what releases ship |
pixi run build-open |
src/vector/reference/ |
auditors, ports, anyone who will not run a binary they cannot read |
any other build (iOS, a bare mojo build, Linux until the Linux archives ship) |
src/vector/reference/ |
— |
The references are the same algorithms with the tuning removed — same
traversal, same heaps, same per-lane float formula — and return bit-identical
results. What they leave out is prefetch staging, batch-of-8 scoring and the
ISA dot instructions (SDOT / VNNI). Two checks hold that to be true:
pixi run test-vector-differential feeds every closed routine and its
reference the same synthetic state (0 differences required, PKM FP32 scores
within 1e-4 because the tuned kernels sum in a different order), and
tests/test_vector_build_equivalence.py --binaries <lib> <open> warm-loads
one real index into both builds and requires identical keys and scores.
What the open build costs (gate config: Performance1536D50K, ef=150,
-w 10, Mac M4, 3 interleaved ABBA pairs each, 2026-09-25):
| Search | build QPS |
build-open QPS |
Open build | Recall@100 (both) |
|---|---|---|---|---|
| INT8 (default) | 8,822 | 6,044 | −30.0% (3/3 pairs: −30.0, −28.7, −31.5) | 0.958 |
| PolarQuant (INT4) | 6,789 | 4,491 | −33.8% (−33.8, −31.0, −35.4) | 0.965 |
| TurboQuant (INT3 + QJL) | 5,035 | 3,776 | −22.3% (−26.3, −20.6, −22.3) | 0.95 |
| NanoQuant (INT2) | 3,606 | 2,818 | −21.9% (−24.9, −21.9, +2.3) | 0.46 |
QPS is the median lib/open pair. Recall is the same because the results are
the same; per-run recall differs only through HNSW build nondeterminism.
Product-key memory (NEURON.PKM.*, 1M slots) was measured the same way on
2026-09-25: server compute +39% exact, +102% with 8 heads, single-query wire
latency +12%. Nothing else changes between the two builds — KV
throughput, prompt-cache TTFT and every other headline number come from open
code.
The interface is open too: src/vector/vector_abi.mojo (the calls),
beam_view.mojo and quant_beam_view.mojo (the argument structs, whose
field order is ABI). INFO reports pion_vector: and --version a
vector: line, both asked of the linked code; the server refuses to start
against a library with a different ABI version.
Quantized variants¶
Measured on Performance1536D50K, ef=150, -w 10, Mac, 2026-09-25:
| Variant | QPS (c=10) | Recall@100 | Compact bytes/vector |
|---|---|---|---|
| INT8 (default) | 8,021 | 0.960 | 1,600 |
| PolarQuant (INT4) | 6,635 | 0.965 | 868 |
| TurboQuant (INT3+QJL) | 4,924 | 0.953 | 676 + 192 QJL |
| NanoQuant (INT2) | 3,390 | 0.464 | 484 |
Every quant variant also keeps an FP32 re-rank copy (6 KB/vector) in memory and in the index file, so none of them saves memory or disk overall today; their compact buffers only shrink the beam's working set. NanoQuant is experimental: INT2 on these embeddings is coarse (recall@10 is 0.81 at ef=150).
The proof a mode really ran is its FT.OPTIMIZE log line —
[PolarQuant] Block-INT4 …, [TurboQuant] Block-INT3 …,
[NanoQuant] Block-INT2 …. Warm restart and multi-worker serving work for all
three.
PolarQuant — Block-INT4 Quantization¶
Enabled with the --polarquant server flag.
Encoding¶
- Block structure: 48 blocks x 32 dims = 1536 dimensions
- Per-block storage: FP16 scale factor + 32 x 4-bit packed values = 18B/block
- Total per vector: 48 x 18B = 868B (vs 1536B for INT8, 43% smaller)
- WHT rotation: Walsh-Hadamard Transform applied before quantization for robustness — spreads information across dimensions, reducing quantization error
Usage¶
TurboQuant — Block-INT3 + QJL Error Correction¶
Enabled with the --turboquant server flag.
Encoding¶
- Block-INT3: 48 blocks x 32 dims, 14B/block = 676B/vector
- QJL error correction: random sign flips + Walsh-Hadamard Transform, 192B sign storage per vector
- Total per vector: 676B + 192B = 868B
Search Kernel¶
- Batch-8
int3_dotkernel for distance computation - QJL Hamming correction applied as a fast error-correction step after initial INT3 distances
Usage¶
NanoQuant — Block-INT2 Quantization¶
Enabled with the --nanoquant server flag. The smallest compact footprint (484B/vector); experimental, see the recall figures above.
Encoding¶
- Block-INT2: 48 blocks × 32 dims, 10B/block (2B FP16 scale + 8B packed 2-bit data) = 484B/vector
- 4 levels: unsigned {0,1,2,3} → signed {-3,-1,1,3} for SDOT; stored scale = absmax / 3, so the outer levels decode to ±absmax
- No QJL needed: FP32 re-rank suffices to maintain recall — simplest quantization pipeline
Search Kernel¶
- Batch-8
int2_dotkernel: unpack 2-bit → Int8, asymmetric SDOT with INT8 query - 8 cache lines per vector (484B) vs 14 (INT4/INT3) — lower memory bandwidth
Usage¶
Streaming Ingest¶
Auto-enabled for indices with more than 1M elements. Eliminates the FP32 staging buffer that would otherwise be required during bulk ingest.
- Memory savings: at 5M vectors x 1536 dims x 4 bytes = 33.8GB FP32 staging buffer eliminated
- Mechanism: vectors are quantized on arrival and inserted directly into the index, bypassing the intermediate FP32 buffer
- TurboQuant streaming compaction: when
--turboquantis active, streaming ingest performs INT8 to FP32 reconstruction on-the-fly for the INT3 quantization pipeline
Default Configuration¶
| Parameter | Value | Source |
|---|---|---|
dimensions |
1536 | config.mojo — matches VectorDBBench Performance1536D50K |
max_elements |
600,000 | config.mojo |
M |
16 | config.mojo (desktop/cloud profiles) |
ef_construction |
100 (desktop), 200 (cloud); overridden by FT.CREATE | config.mojo / slow_path.mojo |
ef_runtime |
150 | hnsw.mojo field default |
use_int4 |
False (all profiles — V13 INT4 abandoned, recall 0.74) | config.mojo |
use_bq |
False (all profiles — Hamming too coarse on 1536-dim) | config.mojo |
Benchmarking — VectorDBBench¶
Harness (benchmarks/VectorDBBench/vectordb-benchmark.py)¶
Runs VectorDBBench sequentially against Redis, Valkey, Pion (all via the redis subcommand of the vectordbbench CLI), then zvec (via the zvec subcommand — no server required):
# Single-engine quick run (builds Pion first):
pixi run bench-vdb # ef_construction=128, ef_runtime=150
# 4-way comparison (Redis, Valkey, Pion, zvec):
cd benchmarks/VectorDBBench
python vectordb-benchmark.py
The script auto-starts and stops each server. For each Redis-compatible engine it uses:
- redis-server for Redis (vanilla; no FT. support → will fail gracefully)
- valkey-server for Valkey (vanilla; no FT. support → will fail gracefully)
- ./pion-server -p 6379 -w 1 for Pion (native FT.*)
- vectordbbench zvec for zvec (embedded library, no server)
Constants (top of script; edit to match your test run):
CASE = "Performance1536D50K" # VectorDBBench dataset identifier
M = 16 # HNSW graph degree
EF_CONSTRUCTION = 128 # Build-time beam width
EF_RUNTIME = 200 # Query-time beam width (matches Pion default)
PORT = 6379
The redis subcommand sends the full FT.CREATE → HSET (bulk ingest) → FT.SEARCH loop. Pion handles all three natively; no Lua scripts or modules required.
CLI entry point: vectordbbench redis (installed by pixi run install-vdbbench). Use vectordbbench --help to list available backends.
Install VectorDBBench¶
GPU Vector Search (Metal)¶
Metal GPU brute-force search with async pipelining and adaptive CPU/GPU routing. Enabled with --gpu flag.
Architecture¶
- Metal compute shader (
src/ffi/metal_compute.metal): INT8 L2 distance, 1 query vs N candidates, char4×4 unrolled + threadgroup shared query cache; optimized multiquery kernel for batch dispatch - Obj-C++ wrapper (
src/ffi/metal_wrap.m): persistent Metal context, per-worker buffers (16 workers), asyncdispatch_semaphorecompletion,dispatchThreadgroups - Adaptive routing: rolling GPU latency tracking → auto-fallback to CPU HNSW when GPU busy (LLM/MAX)
- FP32 re-rank buffer: BFS-ordered FP32 vectors built during
compact_vectors()for optional GPU oversample re-ranking - Native Mojo kernel (
src/vector/gpu_search.mojo) ready for Xcode-enabled builds (-D ACCELERATOR=apple-m4)
Results (Performance1536D50K, ef=150, w=10→4 capped, macOS M4)¶
| Metric | GPU v2 (Metal) | CPU (HNSW) |
|---|---|---|
| Peak QPS (c=1) | 1,665 | 1,520 |
| Peak QPS (c=5) | 5,949 | 4,486 |
| Peak QPS (c=10) | 8,030 | 7,722 |
| Recall@100 | 0.937 | 0.937 |
| P99 Latency | 0.8ms | 0.9ms |
GPU +26% over CPU at c=5; converges at c=10 (4 P-core macOS cap). Multiquery kernel positioned for Linux io_uring batch dispatch.