Skip to content

Node Maintenance

A node in maintenance takes no new containers, and the controller stops recovering, restarting, and health-checking the containers already on it. With cs-agent 3.4.0 or later, the node's agent also pauses backups and other background work, so you can upgrade Docker, the kernel, or the operating system without the agent starting work underneath you.

Containers are not moved, with one exception: in an availability zone with clustered storage, entering maintenance from the admin also migrates the node's containers elsewhere in the availability zone.

Agent version

Maintenance is enforced on the node only by cs-agent 3.4.0 or later. With an older agent, maintenance is a flag on the controller alone, as it was before: the controller stops placing and recovering containers, but backups, restores, and cleanup keep running on the node. The node page warns when this is the case.

Two holds

Maintenance can be placed from either side, and each side owns its own hold:

Hold Placed by Released by
Controller The admin node page, or the Manage API The same places
Node cs-agent maintenance on, run on the node cs-agent maintenance off on the node, or Release all holds

The node is in maintenance while either hold is in place, and each side releases only its own. Releasing maintenance in the admin therefore can't take a node out of maintenance in the middle of a Docker upgrade that someone started on the node.

The controller sends its hold to the agent, and the agent reports both holds back. Every couple of minutes the controller re-reads each node's state and corrects a controller hold that the agent missed.

Holds survive agent restarts and reboots. Clear them explicitly when the work is done.

From the node

Run the maintenance subcommand on the node as root:

sudo cs-agent maintenance on --reason "docker upgrade" --wait --timeout 2h   # place the node hold and wait until idle
sudo cs-agent maintenance status                                            # show the state and any work still running
sudo cs-agent maintenance off                                               # release the node hold
Command Description
on --reason TEXT Places the node hold. --reason is required (up to 512 bytes).
off Releases the node hold. A controller hold stays in place, and the command says so.
status Shows the maintenance state and the work still in flight. Read-only.

Flags:

Flag Description
--wait With on, block until nothing is running on the node and the controller has acknowledged the hold.
--timeout DUR How long --wait may block, such as 30m or 2h. Required with --wait.
--settle DUR Once the node is idle, wait this long and check again. Defaults to 30s.
--no-controller With --wait, don't wait for the controller to acknowledge the hold. Avoid it while you're restarting Docker: until the controller has seen the hold, it may still treat the node as online.
--json Print a single JSON object on stdout. Works with every command.

--wait counts running backup and restore tasks, repository prune and compact runs, and any running backup container as work in flight. Its exit codes tell you how it fell short:

Exit code Meaning
0 The node hold is in place, nothing is running, and the controller has acknowledged the hold.
1 A usage error, the hold was released while waiting, or the wait was interrupted (Ctrl-C or SIGTERM). An interrupted wait leaves the node hold in place.
2 Timed out with work still running, or the agent's database stayed unreadable.
3 The agent's database could not be opened.
4 The controller hasn't acknowledged the hold. on --wait also exits with 4 before placing a hold if the controller hasn't polled the node with maintenance support in the last two minutes, whether because it is too old, down, or lagging.

The provisioner's rolling Docker upgrade (make upgrade-docker) uses the node hold this way: it holds each node with cs-agent maintenance on --wait, does its work, and releases the hold when it finishes.

From the controller

Open the node from its availability zone in the admin, then use the Maintenance panel:

  • Enter Maintenance places the controller hold. The reason is optional and internal; customers never see it. The message to customers is optional and becomes the customer notice.
  • Release Controller Hold releases only the controller hold. If the node also holds itself, it stays in maintenance.
  • Refresh from Agent reads the agent's current state immediately, instead of waiting for the next check.

Every change is audited. The same actions are available through the Manage API, which also accepts a reason.

What the node page shows

The node page (and the node's card on the availability zone page) shows:

  • Who holds the node: the controller, the node, or both, with each hold's reason and since when. A controller hold the agent hasn't confirmed yet is marked as such.
  • Background work: whether the node has quiesced, or how many jobs are still running and which ones (backups and restores still draining).
  • Skipped backups: the number of scheduled backups the node skipped during the window.
  • Agent status: when the agent's state was last read.
  • Warnings:
Warning Meaning
Controller and node disagree about the controller hold The agent's copy of the controller hold still disagrees with the controller after a correction attempt.
Held for more than 12 hours A hold has been in place for more than 12 hours. Usually a forgotten hold.
Agent older than 3.4 The controller hold is in place, but the agent doesn't enforce it, so backups and other work keep running on the node.
Agent status is stale The controller hasn't read the agent's state in the last 10 minutes.
Override not yet delivered to the node A Release all holds hasn't reached the agent yet.

Stale holds and persistent disagreements also raise a system event.

The availability zone page summarizes how many of its nodes are in maintenance, how many have quiesced, and how many have warnings.

Release all holds

When a node is stuck in maintenance under a node hold that nobody is around to release, Override: Release All Holds on the node page releases both holds at once. The button appears only while a node hold exists. The override is audited and raises a warning-level system event.

The override only releases a node hold placed before it. If a new node hold is placed while the override is on its way to the agent (for example, a Docker upgrade that just started), that hold is left in place.

What pauses on the node

While either hold is in place, the agent:

  • Starts no new backup or restore tasks.
  • Skips scheduled backups and repository prune and compact runs.
  • Cancels restores that hadn't started yet.

Work that is already running finishes. The firewall, the metadata service, and changes sent from the controller keep working. When maintenance ends, the backup slots that were missed are skipped rather than all run at once; backups resume on their next scheduled slot.

On the controller side, while the node is in maintenance:

  • No new containers are placed on it.
  • Its containers aren't health-checked, recovered, or restarted, and missed heartbeats can mark it disconnected but don't take it offline, evacuate it, or raise a node offline event. If maintenance ends while the node is still unreachable, it shows as offline straight away.
  • When maintenance ends, the controller checks the node's containers again and re-sends the load balancer configuration.

Container actions during maintenance

Start, stop, restart, and rebuild requests for containers on the node are deferred: the request waits, its event says so, and it runs once maintenance ends. This includes the restarts that apply an already-saved change, such as a custom load balancer picking up a new domain or an SSH container switching password login on or off, and builds and rebuilds for existing projects. A newer power request for the same container replaces a waiting one; the replaced one is canceled as superseded. The automatic recovery after maintenance never replaces a waiting request.

These are refused while the node is in maintenance, with a message saying why:

  • Backups, restores, backup deletion, and backup downloads.
  • Adding containers to a service. Removing containers is always allowed.
  • Rotating an SSH (SFTP) container's password.

Configuration changes, such as project metadata, SSH keys, and firewall rules, are still sent to the node during maintenance.

On a node that is offline (not in maintenance, but unreachable), power actions are refused as well, and configuration changes are sent once the node is back; see Pending agent syncs.

What customers see

Customers see a node's state, never its holds or the internal reason. The customer portal shows a banner on project and service pages, and a label in the project list and on container pages. The User API carries the same state, including which operations are available, deferred, or refused.

  • A node is shown as in maintenance while either hold is in place, even when it is also unreachable (as during a reboot). Otherwise it's offline when disconnected, and online otherwise.
  • Across several nodes, such as a project or an availability zone, the worst state wins. An availability zone's state counts its enabled nodes plus any disabled node that still hosts containers or SFTP containers, since disabling a node only stops new placements on it.

Customer notice

The customer notice is a short, plain-text message shown alongside the node's state. It is separate from the internal maintenance reason.

  • Set it with Message to customers when entering maintenance, or on its own from the Customer Notice panel on the node page, for example during an unplanned outage.
  • It can only be set while the node is in maintenance or offline, and it clears itself when the node is back online. You can also clear it by hand.
  • Every change is audited.

Warning

Treat the customer notice as public. It is shown to every customer with something in that availability zone, and every API user can read it through the locations endpoint.

Pending agent syncs

Changes the controller sends to a node agent (project metadata, SSH keys, SFTP host keys, project setup and removal, firewall rules, and a volume's backup settings) used to be sent once and lost if the node couldn't be reached. They are now recorded and re-sent, built from current data, once the node is back.

The availability zone page shows the number of agent syncs still waiting, how old the oldest is, and how many are dead. A sync becomes dead after the agent answered and refused it 24 times, or 48 hours after its first refused attempt; a node that simply can't be reached doesn't count against that limit. Dead syncs raise a system event. Once you've fixed the cause, click Retry dead on the availability zone page to try them again.

Agent configuration

agent.yml on the node has one maintenance setting:

maintenance:
  stale_hold_hours: 12

A hold older than stale_hold_hours is logged and reported every hour, so a forgotten hold doesn't silently stop backups. 0 disables the warning. This only affects the agent's own reporting; the controller's "Held for more than 12 hours" warning is separate.

Rolling back the agent

Agent versions before 3.4.0 ignore maintenance holds and run work normally. Release every hold before downgrading the agent: a hold left in place takes effect again when the agent is upgraded.