Skip to content

Fledge v0.5.2.1 ​

Automatic failover. When a node dies, its servers come back on another node from their newest backup, and you can watch, test and tune it from a new Resilience page.

Highlights ​

  • Automatic failover. A node that has been silent for the waiting time (5 minutes by default) has its servers rebuilt on other nodes and restored from their newest backups. In testing, a server was running again about 45 seconds after the decision, and two servers recovered side by side.
  • No double running. A node that comes back removes the copies of servers that now live elsewhere. Optionally, nodes stop their protected servers when they lose the panel for too long, so a node that is cut off but alive can't keep a game running that has been started elsewhere.
  • Planned moves. Move one server, or empty a whole node, with a final backup taken while the server is stopped. Nothing is lost, and the old copy is removed only after the new one works.
  • Resilience page. Which servers are ready to recover and why not, a failover countdown per node, recovery history with timings and the age of the data used, node up/down history, and Simulate failure to see where every server of a node would go without changing anything.
  • Settings for all of it, under Settings → Panel → Automatic failover. No .env changes.

Added ​

  • Failover settings: on/off; wait before failing over; recoveries at once; oldest usable backup; pause between failovers of one server; same location only; recover servers without a backup (with empty data); keep backups fresh on an interval; stop servers if a node loses the panel; remove the old copy when a node returns; how long removed data is kept; notification webhook with a test button. Failover is off by default and refuses to be turned on without object storage unless you allow recovery with empty data.
  • Keep backups fresh: takes backups of protected running servers on an interval, only on nodes with disk volumes because those backups don't stop the game.
  • Per-server switch to exclude a server from failover.
  • Webhook notifications (Slack, Discord and Mattermost compatible) when a node goes offline or returns, and when a server is recovered, can't be recovered, or fails. The URL is stored encrypted.
  • Monitoring: GET /api/failover/status, /api/failover/events, a dry-run POST /api/failover/plan, and metrics fledge_failover_events{state} and fledge_servers_without_backup on /api/metrics.
  • Nodes record when they go offline and return, shown as history on the Resilience page.
  • API: POST /api/servers/:id/failover (run now; for a node that is still online it needs force), POST /api/servers/:id/migrate, POST /api/nodes/:id/evacuate.
  • Agents report which servers' data they hold, and handle a new evict job. Removed data is moved to .evicted/ on the node and deleted after the retention you set.

Changed ​

  • Nodes are marked offline by the failover engine, which runs on one API replica at a time (a PostgreSQL advisory lock).
  • Jobs queued in one step for the same server now run in the order they were queued. Jobs inserted together previously shared a timestamp and could run in either order.
  • Agent version 0.5.2.1.

Fixed ​

  • Planned moves no longer leave a server stopped if the final backup fails; it is started again where it was.
  • A recovery is only reported done when every step (create, restore, stop or start) succeeded.

Verified ​

Run against real components, no mocked data: PostgreSQL, the API, the panel, two Go agents, each with its own Docker daemon, real containers, and an S3-protocol test server.

  • Killed a node (agent, containers and volumes gone): detected after 35 s, failed over after the configured 2 minutes, restored from the newest backup and running on the other node; the webhook received each step.
  • Two servers recovered concurrently within the limit; a server with no backup stayed down with a clear reason, and recovered with empty data once that was allowed.
  • The returning node removed its stale copies and left the other node's containers alone.
  • A planned move carried a file written seconds earlier, ended with the server running on the new node, and removed the old copy afterwards.
  • Automatic backups appeared on schedule for a running server; the dry run named the right target and backup age.
  • With stop-on-panel-loss on, a panel outage stopped the protected servers on the node and they restarted when the panel returned.
  • Type checks, Go tests (including rollback), API tests and the production build.

Known limits ​

  • Failover restores from backups. Everything written since the newest backup is lost, and recovery takes as long as a restore. It is not replication.
  • It reacts to a node going silent. It does not detect a frozen game on a healthy node, a node that is reachable but broken, or the panel being down.
  • Without Stop servers if a node loses the panel, a node cut off from the panel keeps its game running until it reconnects, so a server can run on two nodes in that window. Turning it on has a cost: if the panel itself goes down, protected servers stop on every node after the fencing time.
  • A recovery that fails part-way leaves the server on the new node in a failed state with the old data kept aside; retry it by hand.
  • Moves have downtime (a few minutes); live migration is not implemented.
  • Verified with two agents on one machine, not on separate hosts or with a real game. Real AWS S3 retention, a Valheim boot, third-party SFTP clients, load testing, and rootless operation remain unvalidated.

Upgrading ​

Update from Updates in the panel or with update.sh / update.ps1. The database gains failover tables and a per-server switch automatically; failover stays off until you turn it on. Then:

  1. Make sure object storage is on and servers have recent backups (turn on Keep backups fresh if you want it automatic).
  2. Update agents from Nodes to 0.5.2.1. Older agents don't report what they hold, so stale copies on them are not removed.
  3. Use Simulate failure on each node before relying on it.

Released under the AGPL-3.0-only license.