Fledge v0.5.2.1
Automatic failover. When a node dies, its servers come back on another node from their newest backup, and you can watch, test and tune it from a new Resilience page.
Highlights
- Automatic failover. A node that has been silent for the waiting time (5 minutes by default) has its servers rebuilt on other nodes and restored from their newest backups. In testing, a server was running again about 45 seconds after the decision, and two servers recovered side by side.
- No double running. A node that comes back removes the copies of servers that now live elsewhere. Optionally, nodes stop their protected servers when they lose the panel for too long, so a node that is cut off but alive can't keep a game running that has been started elsewhere.
- Planned moves. Move one server, or empty a whole node, with a final backup taken while the server is stopped. Nothing is lost, and the old copy is removed only after the new one works.
- Resilience page. Which servers are ready to recover and why not, a failover countdown per node, recovery history with timings and the age of the data used, node up/down history, and Simulate failure to see where every server of a node would go without changing anything.
- Settings for all of it, under Settings → Panel → Automatic failover. No
.envchanges.
Added
- Failover settings: on/off; wait before failing over; recoveries at once; oldest usable backup; pause between failovers of one server; same location only; recover servers without a backup (with empty data); keep backups fresh on an interval; stop servers if a node loses the panel; remove the old copy when a node returns; how long removed data is kept; notification webhook with a test button. Failover is off by default and refuses to be turned on without object storage unless you allow recovery with empty data.
- Keep backups fresh: takes backups of protected running servers on an interval, only on nodes with disk volumes because those backups don't stop the game.
- Per-server switch to exclude a server from failover.
- Webhook notifications (Slack, Discord and Mattermost compatible) when a node goes offline or returns, and when a server is recovered, can't be recovered, or fails. The URL is stored encrypted.
- Monitoring:
GET /api/failover/status,/api/failover/events, a dry-runPOST /api/failover/plan, and metricsfledge_failover_events{state}andfledge_servers_without_backupon/api/metrics. - Nodes record when they go offline and return, shown as history on the Resilience page.
- API:
POST /api/servers/:id/failover(run now; for a node that is still online it needsforce),POST /api/servers/:id/migrate,POST /api/nodes/:id/evacuate. - Agents report which servers' data they hold, and handle a new
evictjob. Removed data is moved to.evicted/on the node and deleted after the retention you set.
Changed
- Nodes are marked offline by the failover engine, which runs on one API replica at a time (a PostgreSQL advisory lock).
- Jobs queued in one step for the same server now run in the order they were queued. Jobs inserted together previously shared a timestamp and could run in either order.
- Agent version 0.5.2.1.
Fixed
- Planned moves no longer leave a server stopped if the final backup fails; it is started again where it was.
- A recovery is only reported done when every step (create, restore, stop or start) succeeded.
Verified
Run against real components, no mocked data: PostgreSQL, the API, the panel, two Go agents, each with its own Docker daemon, real containers, and an S3-protocol test server.
- Killed a node (agent, containers and volumes gone): detected after 35 s, failed over after the configured 2 minutes, restored from the newest backup and running on the other node; the webhook received each step.
- Two servers recovered concurrently within the limit; a server with no backup stayed down with a clear reason, and recovered with empty data once that was allowed.
- The returning node removed its stale copies and left the other node's containers alone.
- A planned move carried a file written seconds earlier, ended with the server running on the new node, and removed the old copy afterwards.
- Automatic backups appeared on schedule for a running server; the dry run named the right target and backup age.
- With stop-on-panel-loss on, a panel outage stopped the protected servers on the node and they restarted when the panel returned.
- Type checks, Go tests (including rollback), API tests and the production build.
Known limits
- Failover restores from backups. Everything written since the newest backup is lost, and recovery takes as long as a restore. It is not replication.
- It reacts to a node going silent. It does not detect a frozen game on a healthy node, a node that is reachable but broken, or the panel being down.
- Without Stop servers if a node loses the panel, a node cut off from the panel keeps its game running until it reconnects, so a server can run on two nodes in that window. Turning it on has a cost: if the panel itself goes down, protected servers stop on every node after the fencing time.
- A recovery that fails part-way leaves the server on the new node in a failed state with the old data kept aside; retry it by hand.
- Moves have downtime (a few minutes); live migration is not implemented.
- Verified with two agents on one machine, not on separate hosts or with a real game. Real AWS S3 retention, a Valheim boot, third-party SFTP clients, load testing, and rootless operation remain unvalidated.
Upgrading
Update from Updates in the panel or with update.sh / update.ps1. The database gains failover tables and a per-server switch automatically; failover stays off until you turn it on. Then:
- Make sure object storage is on and servers have recent backups (turn on Keep backups fresh if you want it automatic).
- Update agents from Nodes to 0.5.2.1. Older agents don't report what they hold, so stale copies on them are not removed.
- Use Simulate failure on each node before relying on it.
