Documentation

Upgrades

Rolling and stopped-fleet upgrade workflows for production celld, including ingress drain and graceful handoff.

Every celld upgrade starts by reading the release notes for the exact source and target versions. The release decides whether the fleet can run mixed versions during the transition.

Xicar recognizes two upgrade classes:

Upgrade type Old and new versions overlap? User-visible downtime
Rolling update Yes, when the release explicitly supports it Normally none
Stop and update No Expected during the all-stopped window

Do not infer upgrade compatibility from semantic versioning. A release can change peer protocols, bucket formats, recovery formats, or ownership behavior.

Two different drains

An upgrade has two separate drain concerns.

Ingress drain stops Cloudflare from sending new public requests to a node.

celld drain is the runtime’s graceful shutdown. It stops new public admission, finishes requests already accepted, proves and snapshots owned cell state, releases ownership, and hands cells to compatible peers.

These are complementary. Removing a server from ingress does not hand off Durable Objects, and graceful celld shutdown does not guarantee that a simple DNS round-robin configuration will stop selecting that origin immediately.

celld graceful shutdown

celld has no separate long-lived “drain mode” command for routine rollout. Graceful shutdown is the drain.

Any of these begin the same graceful shutdown path:

  • SIGTERM;
  • SIGINT;
  • POST /shutdown on the private internal operator listener.

For Coolify, prefer the normal container stop/redeploy path so Docker sends SIGTERM.

During graceful shutdown celld immediately:

  1. reports 503 from /.well-known/celld/health;
  2. refuses new public requests with 503;
  3. finishes public requests that crossed the admission gate before shutdown;
  4. continues peer traffic for cells it has not handed off yet;
  5. hands owned cells to compatible successors in paced batches;
  6. exits after handoff or when the configured shutdown bound is reached.

The default complete shutdown bound is 40 seconds. The orchestrator stop grace must be longer than that. Xicar uses a 60-second stop grace as its baseline.

The internal operator endpoint is useful for manual operations:

curl -X POST http://<private-node-address>:8081/shutdown

The operator API is alpha and unauthenticated. It must stay on the trusted private network.

Do not use POST /shutdown?handoff=preserve for an ordinary rolling version upgrade. That mode is for a clean same-node reload and keeps ownership records.

Do not use POST /rebalance/pause as a drain mechanism. It pauses balancing fleet-wide; it does not replace graceful node shutdown.

Ingress drain with Xicar’s current Cloudflare setup

Xicar currently uses multiple proxied A records for simple Cloudflare round-robin origin selection.

With that setup, Cloudflare chooses one of the configured origin addresses. Basic Zero-Downtime Failover only retries another origin for specific Cloudflare origin errors such as 521, 522, 523, 525, and 526. celld’s intentional drain response is HTTP 503, so a plain proxied multi-record setup should not rely on that response to remove the node from rotation.

For the current setup, perform an ingress drain before stopping a node:

  1. remove the node’s proxied A record from the serving hostname;
  2. verify that normal public request traffic to that origin has stopped or fallen to zero;
  3. gracefully stop celld;
  4. upgrade and start the replacement;
  5. wait for celld health to return 200;
  6. add the proxied A record back.

Do not use a fixed DNS TTL as the only drain signal. Verify traffic at the origin before stopping it.

A future Cloudflare Load Balancer with an HTTP health monitor pointed at /.well-known/celld/health can automate origin health and removal more cleanly. Even with a health-aware load balancer, explicitly disabling the origin before planned maintenance is the most deterministic ingress drain.

Rolling update

Use this workflow only when the celld release notes say the source and target versions can coexist.

For a three-node fleet:

A old   B old   C old
  ↓
A new   B old   C old
  ↓
A new   B new   C old
  ↓
A new   B new   C new

Upgrade one node at a time.

Before the rollout

  1. Pin the target image tag.
  2. Read the release notes for the exact upgrade path.
  3. Confirm the fleet is healthy.
  4. Confirm R2 is healthy.
  5. Confirm the remaining nodes have enough capacity to absorb one node’s cells.
  6. Keep the existing node-local persistent volumes.
  7. Make sure Coolify’s stop grace exceeds CELLD_SHUTDOWN_TOTAL_MS.

For each node

  1. Remove the node from ingress. With the current Cloudflare setup, remove its proxied A record.
  2. Verify ingress has drained. Confirm the node is no longer receiving ordinary public traffic.
  3. Gracefully stop celld. Stop/redeploy the Coolify resource or send SIGTERM. Do not kill the process.
  4. Let celld hand off ownership. The node’s health endpoint becomes 503; existing accepted requests finish and cells move to peers.
  5. Upgrade the image. Keep the same fleet bucket, private peer address, configuration, and persistent local volume.
  6. Start the replacement.
  7. Wait for /.well-known/celld/health to return 200. The first-readiness gate waits for the live lease, no active donor, adequate memory headroom, and acceptable fleet ownership distribution.
  8. Return the node to ingress. Re-add its proxied A record.
  9. Verify traffic and fleet health.
  10. Only then continue to the next node.

A three-node fleet normally keeps two nodes running while one is upgraded, so fleet durability can continue using follower proofs.

A two-node fleet temporarily becomes a one-node fleet during each restart. It remains correct, but writes can fall back to bucket proofs until the second node returns, increasing write latency and R2 pressure.

Stop and update

Use this workflow when the release notes say mixed versions must not share the serving fleet.

The rule is:

No new-version node starts while any old-version node can still participate in the fleet.

The safest workflow has an explicit all-stopped boundary.

Phase 1: prepare

  1. Pin the target image on every node, but do not start it yet.
  2. Read the release notes and migration notes.
  3. Back up the fleet bucket.
  4. Back up the node-local celld data when the upgrade guidance calls for it. With fleet durability, follower disks can contain acknowledged writes that have not reached the bucket yet.
  5. Prevent automatic restart of the old image during the all-stopped boundary.
  6. Stop deployment writers or other control-plane processes when the release guidance requires it.

Phase 2: reduce serving capacity

When the release notes do not require traffic to stop before the first old node shuts down, Xicar can reduce the old fleet one node at a time:

A old   B old   C old
          ↓
        B old   C old
                  ↓
                C old

For each node being removed:

  1. remove it from ingress;
  2. verify public traffic has drained;
  3. gracefully stop celld;
  4. leave it stopped.

Do not start the new image yet.

This reduces capacity during the maintenance window. Once only one old node remains, fleet durability has no follower ensemble and durable writes fall back to the bucket.

If the release notes require application traffic to stop before any node shutdown, skip this gradual-serving phase and enter maintenance mode first.

Phase 3: create the all-stopped boundary

Before stopping the final old node:

  1. stop or gate public application traffic;
  2. stop deployment writers if required;
  3. gracefully stop the final old node;
  4. verify every old celld process and supervisor is stopped;
  5. wait for every old node lease to expire;
  6. make sure the old binaries cannot restart.

For an HTTP/API service, a controlled Cloudflare maintenance response is preferable to leaving clients pointed at dead origins. A temporary edge response can return 503 Service Unavailable with an appropriate retry policy while the fleet is stopped.

Phase 4: start the new fleet

After every old node is gone:

  1. start the new-version celld processes with the same fleet bucket and required node-local data;
  2. allow any startup migration to complete;
  3. wait for the new nodes to report healthy;
  4. verify peer connectivity and fleet state;
  5. restore public ingress;
  6. resume deployment writers;
  7. monitor errors, restore activity, and write latency closely after traffic returns.

Do not roll back by casually starting an old binary after a new-version format migration. Follow the release-specific rollback guidance.

Example: v0.5.1 to v0.6.0

For CELLD_DURABILITY=fleet, celld v0.6.0 explicitly requires a stopped-fleet upgrade from v0.5.1. A v0.6.0 node expects a follower recovery tail format that a v0.5.1 follower does not provide, so the versions must not overlap in a serving fleet.

For CELLD_DURABILITY=bucket, that specific transition can use a rolling update because bucket durability has no follower ensemble.

This is why the upgrade type belongs to the release transition, not merely to whether a deployment has one node or many.

References