Skip to main content

Fleet Management

June 2026 | ControlForge v1.0.1069

Fleet management is the supervisory layer for overseeing many independent ControlForge nodes from one place: discover them on the network, register them, poll their health, push config or stored snapshots out to them, and detect when their projects have drifted out of agreement. It is a fleet operator's surface — pure HTTP API, no ST functions. Each node stays a fully independent PLC; the fleet manager just watches and coordinates them.

This is not clustering. See Fleet vs. Cluster below — they solve different problems and can be used together.


Fleet management supervises many independent ControlForge PLCs over HTTP — discover, register, poll health, push config and snapshots, and detect when node projects drift out of agreement.

Fleet vs. Cluster​

These two words get confused constantly. They are orthogonal:

Fleet (/api/fleet/*)Cluster (/api/cluster/*)
What it isA supervisory view over many separate PLCsOne logical PLC spread across a boss + minions
Node independenceEach node runs its own project, scans on its ownMinions execute slices of one shared program
CouplingLoose — a node works fine if the manager is offlineTight — minions coordinate with the boss at runtime
Failure blast radiusOne node down ≠ others affectedBoss/minion topology is a single control system
Typical use40 CRAC units / UPS / PDUs across a buildingOne big machine whose I/O is physically distributed
Access patternManager polls/pushes over HTTP per nodeAll minion access proxied through the boss

Rule of thumb: fleet = many machines you supervise; cluster = one machine made of many boxes. A node can be a standalone PLC seen by the fleet manager, a cluster boss seen by the fleet manager, or both.


The workflow at a glance​

discover (mDNS) → nodes appear in the registry
register/edit → PUT a node so it persists with role/family/tier metadata
poll → pull fresh health from a node on demand
drift → compare every healthy node's project hash
push config → send a rendered YAML config to a node
push snapshot → send a stored project snapshot to one or many nodes
collect → pull snapshots from nodes into the manager's store

All examples below assume a bearer token. Grab one:

TOKEN=$(curl -s -X POST http://localhost:8302/api/auth/login \
-H 'Content-Type: application/json' \
-d '{"username":"goplc","password":"goplc"}' | python3 -c 'import sys,json;print(json.load(sys.stdin)["token"])')

1. Discover nodes (mDNS)​

Scan the local network for ControlForge nodes advertising over mDNS. The scan window is timeout seconds (1–30, default 3). Discovered nodes are added to the registry automatically.

curl -s -H "Authorization: Bearer $TOKEN" \
"http://localhost:8302/api/fleet/discover?timeout=3"
{
"added": 0,
"discovered": 2,
"nodes": [
{
"id": "james-Precision-7740-8309",
"name": "ControlForge@james-Precision-7740:8309",
"hostname": "james-Precision-7740",
"host": "192.168.1.251",
"port": 8309,
"cluster_role": "standalone",
"source": "mdns",
"status": "healthy",
"version": "1.0.1001",
"uptime": "16h14m6s",
"task_count": 2,
"program_count": 2,
"project_hash": "47ac7231",
"last_seen": "2026-06-22T00:44:01-04:00",
"last_checked": "2026-06-22T00:44:01-04:00"
}
]
}

discovered is how many responded; added is how many were new to the registry. source: "mdns" marks nodes the manager found itself versus ones added by hand (PUT).


2. List and inspect the registry​

List every known node. Optional query filters narrow by role, family, and tier — the metadata you attach when you register a node (see below).

# all nodes
curl -s -H "Authorization: Bearer $TOKEN" \
http://localhost:8302/api/fleet/nodes

# only tier-1 gateways in the "crac" family
curl -s -H "Authorization: Bearer $TOKEN" \
"http://localhost:8302/api/fleet/nodes?role=gateway&family=crac&tier=1"
{
"count": 2,
"fleet_manager": false,
"nodes": [ /* node objects as above */ ]
}

Fetch a single node by its id:

curl -s -H "Authorization: Bearer $TOKEN" \
http://localhost:8302/api/fleet/nodes/james-Precision-7740-8309

A node's status (healthy / unhealthy) and project_hash are the two fields you watch most: status says it's reachable, project_hash says what it's running.


3. Register, update, or remove a node​

Use PUT to add a node by hand or to attach/overwrite metadata (role, family, tier, friendly name) on a discovered one. The node id in the path is the key; the body carries the fields to set.

curl -s -H "Authorization: Bearer $TOKEN" \
-X PUT http://localhost:8302/api/fleet/nodes/crac-3 \
-H 'Content-Type: application/json' \
-d '{
"name": "CRAC Unit 3 (Row B)",
"host": "10.0.0.73",
"port": 8302,
"role": "gateway",
"family": "crac",
"tier": "1"
}'

Remove a node from the registry (does not touch the node itself — only the manager's record of it):

curl -s -H "Authorization: Bearer $TOKEN" \
-X DELETE http://localhost:8302/api/fleet/nodes/crac-3

4. Poll a node's health on demand​

Discovery refreshes health on its own schedule, but you can force an immediate health poll of one node:

curl -s -H "Authorization: Bearer $TOKEN" \
-X POST http://localhost:8302/api/fleet/nodes/james-Precision-7740-8309/poll

The response is the freshly-polled node object (updated status, version, uptime, project_hash, last_checked). Poll after a push to confirm the node picked up the change.


5. Detect drift​

The single most useful fleet query: are all my healthy nodes running the same project? drift groups every healthy node by its project_hash. One group = everything agrees. More than one group = drift.

curl -s -H "Authorization: Bearer $TOKEN" \
http://localhost:8302/api/fleet/drift
{
"drifted": true,
"unique_hashes": 2,
"healthy_nodes": 2,
"groups": [
{ "hash": "47ac7231", "count": 1, "node_ids": ["james-Precision-7740-8309"] },
{ "hash": "(empty)", "count": 1, "node_ids": ["james-Precision-7740-8302"] }
]
}

drifted: true with unique_hashes > 1 is your signal to reconcile. Each group's node_ids tells you exactly which nodes are on which version. A (empty) hash means that node didn't report a project hash (e.g. nothing loaded yet).


6. Push configuration to a node​

Send a config YAML directly to a node. The body is { yaml, vars } — vars is optional and feeds Go text/template substitution into the YAML before it's applied (same engine as template/render).

curl -s -H "Authorization: Bearer $TOKEN" \
-X POST http://localhost:8302/api/fleet/nodes/crac-3/config \
-H 'Content-Type: application/json' \
-d '{
"yaml": "tasks:\n - name: Main\n type: periodic\n scan_time_ms: {{ .scan_ms }}\n",
"vars": { "scan_ms": "20" }
}'

7. Push and collect snapshots​

List what's available to push to a node​

curl -s -H "Authorization: Bearer $TOKEN" \
"http://localhost:8302/api/fleet/nodes/crac-3/snapshots?limit=50"
{
"node_id": "crac-3",
"snapshots": [
{
"hash": "abaaf88a",
"name": "counter1",
"created_at": "2026-06-18T20:17:35Z",
"source": "download",
"size_raw": 14686,
"size_gz": 2283
}
]
}

Push one stored snapshot to a single node​

Identify the snapshot by its hash:

curl -s -H "Authorization: Bearer $TOKEN" \
-X POST http://localhost:8302/api/fleet/nodes/crac-3/push \
-H 'Content-Type: application/json' \
-d '{ "hash": "abaaf88a" }'

Push one snapshot to many nodes at once​

push-bulk fans the same snapshot out concurrently. Omit node_ids (or send null) to target every node; provide a list to target a subset.

curl -s -H "Authorization: Bearer $TOKEN" \
-X POST http://localhost:8302/api/fleet/push-bulk \
-H 'Content-Type: application/json' \
-d '{
"hash": "abaaf88a",
"node_ids": ["crac-3", "crac-4", "crac-5"]
}'

This is the reconcile step after drift flags a split: pick the good hash, push-bulk it to the drifted node_ids, then re-run drift to confirm one group.

Collect snapshots from nodes into the manager​

collect pulls project snapshots from fleet nodes into the manager's local snapshot store, so they're available to push elsewhere:

curl -s -H "Authorization: Bearer $TOKEN" \
-X POST http://localhost:8302/api/fleet/snapshots/collect

Export and purge the manager's snapshot store​

# bundle collected snapshots for export
curl -s -H "Authorization: Bearer $TOKEN" \
-X POST http://localhost:8302/api/fleet/snapshots/export

# purge the store — requires explicit confirm
curl -s -H "Authorization: Bearer $TOKEN" \
-X POST http://localhost:8302/api/fleet/snapshots/purge \
-H 'Content-Type: application/json' \
-d '{ "confirm": true }'

purge is guarded by confirm: true — without it, nothing is deleted.


8. Render a config template​

Before pushing, you can render a templated YAML to verify what a node will actually receive. The render endpoint takes the same { yaml, vars } shape and returns the substituted string — no node touched, no config applied.

curl -s -H "Authorization: Bearer $TOKEN" \
-X POST http://localhost:8302/api/fleet/template/render \
-H 'Content-Type: application/json' \
-d '{
"yaml": "node_name: {{ .name }}\nscan: {{ .scan_ms }}",
"vars": { "name": "crac-3", "scan_ms": "20" }
}'
{ "rendered": "node_name: crac-3\nscan: 20" }

Templating uses Go text/template: {{ .key }} pulls key from vars. Render first to eyeball the result, then send the same { yaml, vars } to /config to apply it for real.


9. Per-node history​

Each node accumulates a snapshot/health history you can page through:

curl -s -H "Authorization: Bearer $TOKEN" \
"http://localhost:8302/api/fleet/nodes/crac-3/history?limit=50"
{ "node_id": "crac-3", "history": [] }

Notes & limits​

  • Authentication: every fleet endpoint requires a bearer token. The examples assume $TOKEN from the login call at the top.
  • The manager is loosely coupled. It supervises over HTTP; nodes keep running independently if the manager goes away. This is the structural difference from a cluster, where the boss/minion topology is one control system.
  • (empty) project hash in drift means a node reported no project — treat it as "unknown / not configured", not as agreeing with anything.
  • source distinguishes origin: mdns nodes were auto-discovered; hand-added nodes (PUT) persist regardless of whether mDNS sees them.
  • Template substitution is Go text/template, shared by /template/render, /nodes/{id}/config, and push paths that take vars. Render before you push to catch a bad variable before it lands on a node.
  • purge requires confirm: true. It clears the manager's collected-snapshot store, not the nodes.
  • Snapshots are identified by hash, not by name — two snapshots can share a name across versions, but the hash is what push/push-bulk resolve.