Indicatrix GitHub
Networking

Distributed rendering

Offloading sample tracing to a coordinator and its joined workers over mutual TLS.

1Preview-then-handoff

Preview-then-handoff: the desktop renders a local preview while the camera moves, then a remote worker traces sample-index ranges that are summed into the same accumulator. DESKTOP · indicatrix-cut camera moving local CPU / GPU preview, few samples, immediate camera settled hand sample tracing off to the worker WORKER · indicatrix-worker GPU / CPU path tracer traces sample-index ranges [0, 4096) [4096, 8192) … not screen tiles no seams, no per-tile load imbalance on a small stone mTLS 1.3 · token enrolment finished sample ranges, any order ACCUMULATOR local and remote samples are summed per pixel; the image converges as ranges arrive, whichever side produced them
Fig. 1 — Preview-then-handoff: local preview while the camera moves, remote sample-index ranges once it settles, one shared accumulator.

Spectral path tracing through a high-index gemstone with dozens of internal bounces is expensive, and streaming every interactive frame to a remote machine would add visible input latency. Indicatrix instead renders locally while the camera moves, and hands off to its one configured remote endpoint once the camera settles: the client sends a compact SceneState (serialized with postcard) and the remote streams traced radiance back, which is folded into the client's own accumulation buffer. That endpoint is always a coordinator, indicatrix-worker serve, which renders the request itself, splits it across render workers that have joined it, or both (section 3); the desktop app never talks to a worker directly.

The settings dialog's Live Compute choice decides what happens at the settle: Local only, Remote only, or Local + Remote. With Local + Remote both sides trace the same image against one shared sample budget — the live view's single target sample count, there is no separate remote setting — through one sample cursor per settle (LiveEpoch, apps/indicatrix-cut/src/bridge/sample_cursor/live.rs), so their samples are combined, never duplicated. The remote side claims short chunks sized to about 1.5 s of its measured rate (8 samples for the first chunk, before a rate is known; LIVE_CHUNK_TARGET_SECS, bridge/remote/live_lane), so a camera drag wastes at most a second or two of remote time. A chunk that fails keeps the prefix the remote already traced and returns the untraced remainder to the local tracer; after two consecutive failures the remote sits out the rest of that settle. Each settle is stamped with the scene it was dispatched for, so a scene change during the handoff can never merge samples of the old scene into the new image.

2Sample-range additivity

Splitting a frame into screen-space tiles load-balances poorly for a gemstone render: background tiles finish almost instantly while tiles covering the stone itself carry the full bounce budget. Because each sample's RNG seed is a deterministic function of pixel index and sample number, sample contributions are additive and order-independent, so Indicatrix instead partitions work by sample-index range across the whole frame — one node traces samples [0, 64), another [64, 128), and so on — with no tile-seam stitching and linear load balancing across heterogeneous machines.

Disjointness is the whole correctness condition, so it is enforced in one place: a SampleCursor in the GUI-free crates/indicatrix-dispatch hands out disjoint absolute sample ranges to every contributor of one image — the export's local and remote lanes, the live view's local tracer and remote chunks, and the coordinator's own lane and joined workers alike. Every backend also applies the same non-finite rule (add_finite_sample, crates/indicatrix/src/optics/raytracer/sampling.rs, with its twin in the GPU reduction): a sample with a NaN or infinite component is dropped but still counted, so one bad sample can never poison a pixel's sum and the divisor of a merge is always the number of samples actually traced. A statistical test renders one range on the CPU and the next on the GPU, merges them the production way, and checks the result against an independent CPU reference (renderer/gpu/merge_tests.rs).

3The coordinator and joining workers

indicatrix-worker serve is the coordinator (apps/indicatrix-worker/src/serve, src/coordinator). It always serves the read-only design library to viewers. On a build with the worker feature it also opens a separate worker port that render workers dial into, and it traces samples with its own CPU/GPU only when started with --render (--only-gpu/--only-cpu imply it); without that flag it acquires no GPU and traces nothing itself. Its WELCOME advertises exactly that: no render capability for a bare coordinator with no joined workers (the desktop app then sees a library-only remote and renders locally), the plain CPU or GPU backend for --render alone, and a Coordinator { workers, threads, gpus } backend as soon as a worker has joined. A capability change between requests is announced to the viewer with CAPABILITY_CHANGED.

RELEASE NOTE

serve no longer renders by default. A single-machine setup that should keep rendering must add --render (or --only-gpu/--only-cpu, which imply it).

A worker is indicatrix-worker join <coordinator-host:port> (apps/indicatrix-worker/src/join): it connects out to the coordinator's worker port over mutual TLS with a worker certificate, reports its render capability in its HELLO, and then serves the coordinator's requests on that connection. No inbound port is needed on the worker, so a machine behind NAT or a cloud VM can join. --slots K opens K parallel connections (default 1, at most 64), each serving one request stream at a time; every slot reconnects forever with jittered backoff, 1 s doubling to 60 s.

Default portListenerWho connectsAllowlist
7878viewer port (--bind)indicatrix-cutallowlist-viewers.txt
7879viewer enrollment (--enroll-bind)a viewer redeeming a token—
7880worker port (--worker-bind)indicatrix-worker joinallowlist-workers.txt
7881worker enrollment (--worker-enroll-bind)join --token—

The worker port sits two ports above --bind and each enrollment listener one above its own port, on --bind's host unless overridden (WORKER_PORT_OFFSET, apps/indicatrix-worker/src/cli/mod.rs). Every listener binds loopback only unless --allow-remote is passed. --no-workers skips the worker port and its enrollment listener; --no-enroll skips both enrollment listeners.

How a request runs depends on its kind (src/coordinator/job). An export-type request — a still export, a tilt-video frame — becomes a job: a lane pool (indicatrix-dispatch's LanePool) over the request's sample range, with one lane for every idle joined worker whose pixel cap accepts the image plus the coordinator's own lane, each claiming chunks from the shared sample cursor sized to its measured rate. A live-view request stays on the own lane alone, the lowest-latency path; a coordinator without --render gives it the single fastest idle worker instead. If a worker dies mid-job, the samples it had not finished go back to the cursor and the remaining lanes trace them, so the viewer still receives every sample of the request exactly once. Only when no lane at all is left does the viewer get ALL_WORKERS_LOST, and it then renders the unfinished part locally. The coordinator pings every idle joined connection after 10 s without traffic and drops it after 30 s of silence (LivenessConfig, src/coordinator/registry.rs); a joined worker that hears nothing for 45 s reconnects (JOIN_IDLE_TIMEOUT).

Two advanced flags change the live-view rule: --interactive-workers N also gives each live-view request up to N of the fastest idle joined workers (default 0), and --pin-interactive-worker <label> prefers one named worker while it is connected, idle and accepts the image. Throughput jobs — an export streamed final-only, or any final-picture request — wait in a per-viewer queue, one active job per viewer certificate, and all in-flight jobs share one memory budget, --max-job-memory-mib (default 2048; each job is charged width × height × 48 bytes, and a job past the cap is refused, not queued).

4What travels on the wire

5Mutual TLS, roles and token enrollment

The wire protocol lives in crates/indicatrix-net, typed message framing with no networking of its own. indicatrix-worker's serve and join commands carry the actual sockets and TLS.

6The indicatrix-worker CLI

indicatrix-worker has four subcommands: render (one-shot scene-to-PNG, no networking), serve (the coordinator: the design-library protocol over TLS, the worker port, and with --render its own render lane), join (a render worker dialling a coordinator), and cert (manages the private CA: init, issue-server, issue-client, issue-token, claim). render, join and serve --render need a build with the worker feature (gpu implies it). Flags below are copied from the CLI's own --help text in apps/indicatrix-worker/src/cli/mod.rs.

terminal — render
indicatrix-worker render \
  --scene scene.json \
  --out render.png \
  --width 3840 \
  --height 2160 \
  --samples 8192 \
  --threads 16

--only-gpu / --only-cpu force a backend; the default is a hybrid split across both when it measures as worthwhile.

terminal — serve (coordinator)
indicatrix-worker serve \
  --bind 0.0.0.0:7878 --allow-remote \
  --ca pki/ca.pem --cert pki/server.pem --key pki/server.key \
  --render

--bind defaults to loopback only (127.0.0.1:7878); --allow-remote is required for any other address, on any of the four listeners. --db points at the design-library SQLite file (default facet_diagrams.sqlite in the working directory). Leave out --render for a coordinator that only distributes work to joined workers. --threads and --only-gpu/--only-cpu govern the own lane; --max-connections defaults to 64.

terminal — join (worker)
indicatrix-worker join coordinator.local:7880 --slots 2

The address is the coordinator's worker port (also accepted as --coordinator). --cert-dir names the worker certificate bundle (ca.pem, client.pem, client.key; default worker-cert in the working directory); --threads and --only-gpu/--only-cpu work as for render. With one GPU, more than one slot mainly helps CPU-only machines.

7Setting up a coordinator, its workers and the desktop app

terminal — on the coordinator machine
# 1. create a CA
indicatrix-worker cert init --dir pki/

# 2. issue the coordinator's own certificate (used on the viewer and worker ports)
indicatrix-worker cert issue-server --dir pki/ --host coordinator.local --ip 192.0.2.10

# 3. start the coordinator; drop --render if this machine should not trace samples itself
indicatrix-worker serve --bind 0.0.0.0:7878 --allow-remote \
  --ca pki/ca.pem --cert pki/server.pem --key pki/server.key --render

# 4. mint one-time, 180-second enrollment tokens (loopback only)
indicatrix-worker cert issue-token --ca pki/ca.pem --admin-addr 127.0.0.1:7879 --name my-laptop
indicatrix-worker cert issue-token --ca pki/ca.pem --admin-addr 127.0.0.1:7881 --name gpu-box --role worker
terminal — on each worker machine
# first run: claim the worker token into worker-cert/, then join
indicatrix-worker join --coordinator coordinator.local:7880 --token GW1-...

# every later run reuses the claimed bundle
indicatrix-worker join coordinator.local:7880

serve's viewer enrollment listener defaults to the same host as --bind, one port up, and the worker enrollment listener to one port above the worker port (--enroll-bind and --worker-enroll-bind override them), but cert issue-token only ever succeeds from a loopback peer regardless of that address — the operator runs it on the coordinator itself, not remotely. join --token claims one port above the worker port it was given (7881) unless --enroll-addr says otherwise; a token is single-use, so later runs leave it out. For an air-gapped machine, cert issue-client --role worker writes the same bundle for copying by hand.

In indicatrix-cut, open the Remote panel (Remote Coordinator) and choose Set up remote coordinator: enter the coordinator's viewer address (coordinator.local:7878), paste the enrollment address (coordinator.local:7879) and the printed token, and press Redeem. The client redeems it (cert claim's wire equivalent), stores the returned certificate bundle, fills in the certificate folder, and completes the mutual-TLS handshake on the next Test connection or render. The same form holds the live stream preferences and the default export Transfer; the app keeps exactly one such endpoint, and a settings file from an older build that listed several workers keeps only the first.