Skip to content

Docker and deployment

No image is published to a registry yet, so build one first — hermetically from source, or in seconds around a binary you already have:

docker build -t pion .                              # from source, ~20 min
docker build --target runtime-prebuilt -t pion .    # wraps ./pion-server, seconds
                                                    # (Linux host only — the image
                                                    #  runs the binary in the context)

docker run -p 1974:1974 -v pion-data:/data pion
redis-cli -p 1974 PING     # +PONG

The from-source target is verified end to end as of 0.985 on linux/arm64: a 161 MB image that passes Gate 1 (114/114) and the full Redis parity suite from outside the container, and keeps its keyspace across docker restart through the /data volume.

Data (WAL, snapshots, blob arenas) lives in the /data volume, so it survives container restarts. Three things worth knowing:

  • The default command passes --epoll, because Docker's default seccomp profile blocks the io_uring syscalls (Docker ≥ 25). For the faster io_uring path, run with --security-opt seccomp=unconfined and drop --epoll.
  • The slim image ships no Python, so it runs with --no-auto-embed. Features that need an embedding model (semantic cache, auto-embed) want a host install or the AI image variant.
  • Use a named volume (-v pion-data:/data) as shown. The server runs as the unprivileged pion user, and a named volume inherits /data's ownership from the image; a bind mount (-v $(pwd)/data:/data) keeps the host directory's owner instead, so the WAL cannot be created. For a bind mount, chown the host directory to the image's pion uid first.

Prebuilt tarballs and the release pipeline

Official binaries are built by CI from a v* tag (.github/workflows/release.yml): macOS arm64, Linux x86_64 and Linux arm64, each tested before upload, with the versioned tarballs, byte-identical pion-<platform>.tar.gz aliases and SHA256SUMS.

Prebuilt tarball:

tar -xzf pion-<version>-macos-arm64.tar.gz && cd pion-<version>-macos-arm64
./pion-server.sh                    # wrapper sets the dyld/ld library path

Run the wrapper, not bin/pion-server: the tarball bundles the Mojo runtime libraries the binary needs, and the wrapper points the loader at them.

Metal shader library (macOS). --metal-attention loads metal_compute.metallib, searched relative to the executable: <exe>/, <exe>/../lib, <exe>/../share/pion, <exe>/../src/ffi, then the working directory for a source checkout. PION_METAL_LIB overrides the search. The banner prints Metal Attn: requested, and the engine then prints Metal Attn: ACTIVE (<path>) or Metal Attn: NOT ACTIVE under the same key, so one grep answers whether it is running.

Inference sidecar. When the PyTorch sidecar (--inference, or auto-embed) cannot be found — a release tarball ships none — the server says so and starts without it rather than waiting. When it can be found, the server listens first and the sidecar starts concurrently, so clients queue instead of being refused.

Building a tarball yourself — scripts/package_release.sh (run on the target platform; add --build to build first):

scripts/package_release.sh          # → dist/pion-<version>-<os>-<arch>.tar.gz

It ships ./pion-server as it finds it, so build at HEAD first; it refuses a working tree with uncommitted changes (--allow-dirty overrides for local use). On macOS it refuses when src/ffi/metal_compute.metallib is missing or older than metal_compute.metal — building the shaders needs full Xcode, not only the Command Line Tools. To test a tarball the way a user will run it, hide .pixi/envs/default/lib first: the binary's baked rpath points there, so an unpacked tarball resolves its libraries on the build machine whether or not the bundle is complete.

Docker — root Dockerfile, multi-arch, GPU-free:

docker buildx build --platform linux/amd64,linux/arm64 -t pion .
docker run -p 1974:1974 -v pion-data:/data pion
  • amd64 builds with build-portable (--target-cpu x86-64-v2) so the image runs on any x86-64 host; arm64 uses the default GPU-free build.
  • Default CMD is --epoll --no-auto-embed: Docker's default seccomp profile blocks io_uring (Docker ≥ 25), and the slim image ships no Python/torch for the embedding sidecar. For io_uring: docker run --security-opt seccomp=unconfined pion --no-auto-embed.
  • State (WAL, snapshots, blob arenas) lands in the /data volume.
  • docker build --target runtime-prebuilt -t pion:dev . wraps an already-built host ./pion-server and its runtime libraries in seconds, for local use. The default target builds from source (~20 min cold on 4 cores, toolchain layer cached; budget ~5 GB of build cache, docker builder prune -f reclaims it).
  • Dockerfile.linux + docker/run-linux.sh are an interactive Linux dev shell, not the shippable image.

CUDA images are a separate lane (the linux-64 default build requires nvcc; see pixi.toml).

Supervised serving with auto-restart, crash breadcrumbs and the status file: Running in production.