Documentation

Operations

Health checks, node replacement, networking, monitoring, and recovery rules for production celld.

Health and readiness

Use celld’s public health endpoint:

GET /.well-known/celld/health

A healthy node returns HTTP 200. A draining or unready node returns 503.

The first healthy response is a rollout gate. A fresh process waits for its live lease, no active donor, sufficient fleet memory headroom, and acceptable ownership distribution before it reports healthy.

Graceful shutdown

Graceful shutdown is celld’s node drain.

SIGTERM, SIGINT, and private POST /shutdown all enter the graceful shutdown path. celld immediately stops public admission, finishes already accepted public requests, and hands owned cells to peers in paced batches.

For Coolify-managed containers, use the normal stop/redeploy operation so Docker sends SIGTERM.

The default complete shutdown bound is 40 seconds. The orchestrator stop grace must exceed that bound. Xicar’s baseline is 60 seconds.

See Upgrades for the complete rolling and stopped-fleet workflows.

Single-node replacement

The single-node profile is disposable by design:

CELLD_DURABILITY=bucket
no CELLD_WATCH
no persistent volume

A replacement node restores state from the bucket. Expect colder startup and more bucket reads than a node with retained local state.

Because there is only one serving node, planned replacement creates an availability interruption unless an external migration strategy is used.

Fleet node replacement

A fleet node uses:

CELLD_DURABILITY=fleet
CELLD_WATCH=/var/lib/celld/state

Keep the node’s local volume when restarting or recreating the celld container on the same server.

Do not replicate that volume to the other servers. Every node has its own local copy, and celld performs logical replication through its follower protocol and the shared bucket.

Private peer network

Never expose the internal listener to the public internet.

public traffic
Cloudflare → ingress → celld :8080

peer traffic
celld A :8081 ↔ celld B :8081 ↔ celld C :8081
private network only

Use WireGuard, Tailscale, a private VPC, or another trusted encrypted path when the underlying network is not private.

Monitoring priorities

At minimum monitor health status, node count, memory pressure, restore backlog, repeated ownership recovery, object-store errors, request latency, and local disk usage on fleet nodes.

Incident principle

For business-critical state, celld remains the execution and coordination layer rather than the only system of record. Canonical financial, booking, payout, and accounting state stays in Postgres so recovery procedures can reason about runtime state and business records independently.