Skip to content

Workers and keyspaces

The single most important operational fact about Pion: -w N is N independent keyspaces, not one keyspace served by N threads. The default is -w 1, and the server refuses -w N > 1 unless you acknowledge the semantics with --independent-workers.

Pion is shared-nothing. A worker owns a private hash map, WAL, HNSW graph and allocators, and there is no cross-worker request bus. Workers compete for accept() on one listen fd, so a connection is bound to whichever worker won the race — and:

conn A  ->  worker 1   SET user:1 alice   -> +OK
conn B  ->  worker 3   GET user:1         -> (nil)

That is not a bug report, it is the model. -w N gives you N independent keyspaces, not one keyspace served by N threads. Cross-worker PUBLISH delivers to zero subscribers, FLUSHALL clears one worker's slice, and replication covers worker 0 only.

Because every default connection pool (redis-py, Jedis, go-redis, ioredis) opens more than one connection, this breaks read-your-writes silently — no error, just nils. So:

  • The default is -w 1. One worker, one coherent keyspace, every Redis client correct out of the box.
  • -w N > 1 refuses to start unless you also pass --independent-workers, which prints the semantics at startup.

Multi-worker is the right choice when each client pins a single connection (or shards keys across per-worker ports itself), and for read-mostly vector search — the HNSW graph is published across workers via SharedHNSWView, so FT.SEARCH is coherent at -w N even though the KV keyspace is not: every worker returns the same ranking and the same keys. The documents' hash fields still live on the worker that ingested them, so a result answered elsewhere carries its key and score but not its id field — read the key. That is the configuration behind the w=16 / w=32 benchmark numbers above.

./pion-server -w 1                              # default: one coherent keyspace
./pion-server -w 16 --independent-workers       # 16 keyspaces, client pins a connection

If you want one logical keyspace across cores today, run N single-worker Pions on N ports and shard client-side, the same way you would shard Redis.


What the server does about it

The default is -w 1, and -w N > 1 refuses to start without --independent-workers. This is an operational fence, not a tuning knob.

Workers are shared-nothing: each owns a private hash map, WAL, HNSW graph and allocators, there is no cross-worker request bus, and connections are assigned by accept() race on one shared listen fd. So a write acknowledged +OK on one connection is invisible to a read on another. Every default connection pool (redis-py, Jedis, go-redis, ioredis) opens more than one connection, which makes the violation silent — nils, no error, nothing in any log. Measured at -w 4 with 16 concurrently-opened connections: 41 of 90 GETs of a just-acked key returned nil.

./pion-server -w 1                          # default — one coherent keyspace
./pion-server -w 16 --independent-workers   # 16 keyspaces; prints the semantics at startup
./pion-server -w 16                         # FATAL, exit 1, names the flag

Deploy -w N > 1 only when:

  • every client pins a single connection for its whole lifetime, or shards keys across per-worker ports itself; or
  • the workload is read-mostly vector search — the HNSW graph is published across workers via SharedHNSWView, so FT.SEARCH stays coherent at -w N (same ranking, same keys on every worker) even though the KV keyspace does not. Hash fields are per worker: a result answered by a worker that did not ingest the doc omits its id field rather than inventing one.

Otherwise run N single-worker Pions on N ports and shard client-side.

Three related splits follow the same line and are not fixed by the flag: cross-worker PUBLISH delivers to zero subscribers, FLUSHALL clears one worker's slice, and replication covers worker 0 only.

Diagnostic note. Connections opened serially all land on one worker and mask the split completely — a serial probe reports zero nils on a -w 16 server. Any check of cross-worker behaviour must open its connections concurrently. tests/test_gh253_multiworker_fence.py is the regression test.

Caps interact with the fence in the safe direction: --flare/--inference/ --profile ai collapse any -w N back to 1, and macOS caps to 4. The fence reads the final count, so a -w 8 that a cap already reduced to 1 starts normally.