HydraPipeline

HydraCluster Runbook

Render node fleet management.

Release-line note: v2.0.30 exists on releases.experiencenet.com but not in git — it was a phantom build accidentally produced from a misplaced v2.0.30 tag (intended for hydraheadflatscreen) that was deleted seconds later, after CI had already cut a binary. The v2.0.30 binary is byte-identical to v2.0.26. v2.0.31 is the next clean release-line version on master and supersedes both.

Infrastructure

Resource Value
Server hydracluster.experiencenet.com (46.224.29.125)
Config /root/.hydracluster/config.yaml
Data /root/.hydracluster/nodes.yaml
Service systemctl status hydracluster
Logs journalctl -u hydracluster -f

Overview

HydraCluster manages render node enrollment, provisioning, and monitoring. Render nodes are Windows machines running LarkXR Standalone (cloud rendering). They are enrolled via a web UI and managed through the node agent (hydranode).

Key Concepts

SSH Access

ssh root@hydracluster.experiencenet.com
# or
ssh root@46.224.29.125

Health Check

curl -s https://hydracluster.experiencenet.com/api/v1/health

Live Event Stream

The admin UI includes a real-time event stream at /admin/events (requires admin login). All significant cluster events — eligible body discovery, sunshine PIN submissions, stream start/stop, head status changes — appear here as they happen, colour-coded by type.

Use this first when debugging iPad/kiosk pairing and streaming issues. It is faster than SSHing in and tailing journalctl.

stream_stop is emitted for every session close that had an active session record: body self-reports idle, body goes offline mid-stream, admin stop, iPad cancel, and flatscreen stop. It is NOT emitted as a duplicate if the session was already closed by a prior path (e.g. iPad cancel closes the session immediately; when the body then reports idle, no second event fires).

The SSE endpoint (GET /api/v1/events/stream) is also accessible directly with an admin bearer token — useful for scripted monitoring. The server buffers the last 200 events; new connections receive the full buffer before live events begin.

Troubleshooting

Service not responding

  1. Check the live event stream at https://hydracluster.experiencenet.com/admin/events first — the stream going dead confirms the service is down.
  2. SSH to the server: ssh root@46.224.29.125
  3. Check service status: systemctl status hydracluster
  4. Check logs: journalctl -u hydracluster --since '10 min ago' --no-pager
  5. Restart if needed: systemctl restart hydracluster

Symptoms of a self-deadlock from new auth middleware

If POST endpoints (e.g. /api/v1/nodes/{id}/exec) start hanging with 0 bytes received while GETs still respond, and the journal shows lots of GETs but no recent POSTs to the slow endpoint, suspect an auth-middleware self-deadlock.

This was the v2.0.32 → v2.0.33 regression: requireAdminOrNodeToken used defer s.mu.Unlock() and then called next(). The downstream heads handlers (handleListHeads et al.) also acquired s.mu → non-reentrant mutex → goroutine deadlocked on its own lock → mutex pinned → every other lock acquirer queued forever. hydranode's network-recovery routine on the same box then misread the wedged API as a network failure and rebooted the cluster.

Rule for any new server middleware:

If you suspect this is happening live, the fingerprint in journalctl -u hydracluster is many GET /api/v1/body/shell/check lines but no recent matching POST log line for the slow request — the POST is queued waiting for the lock and never reaches its log statement.


Remote Command Execution

HydraCluster supports two exec modes for running commands on body machines.

Structured Exec (default)

Commands are sent as JSON via the API, executed directly by the node agent (PowerShell on Windows, bash on Linux), with clean stdout/stderr separation and proper exit codes. No PTY involved.

# Basic command
hydracluster exec <nodeId> "Get-Process | Select-Object -First 5"

# With timeout
hydracluster exec <nodeId> "choco list --local-only" --timeout 60s

# From a local file (avoids all shell escaping)
hydracluster exec <nodeId> --file ./setup.ps1

# JSON output (for scripting)
hydracluster exec <nodeId> "hostname" --json

The command is queued on the server and picked up by the node agent on its next heartbeat (up to 30 seconds). The CLI polls for the result until it arrives or the timeout expires.

Shell Exec (WebSocket PTY — always responsive)

hydracluster exec --shell <nodeId> "top -bn1 | head -5" --timeout 10s

Async queue vs WebSocket shell — when to use which

These are two distinct transports. Understanding the difference matters when the node is slow to respond or the queue is backed up.

Async exec queue (default) WebSocket shell (--shell)
Transport HTTP poll: hydranode fetches commands on its tick (up to 30 s delay) WebSocket PTY: direct connection, responds immediately
Queue Commands pile up in-memory on hydracluster; processed FIFO No queue — each --shell call opens a fresh PTY session
Use when Scripting, JSON output, fire-and-forget, parallel commands Queue is backed up; need immediate response; binary data (base64)
Failure mode Times out if queue is long or hydranode stopped polling Fails fast if WebSocket connection can't be established

When the async queue backs up: this happens when many exec requests accumulate (e.g. after a debugging session with many rapid execs, or after a node reboot where hydranode was offline and the queue grew). The node heartbeat still arrives (node shows online) but exec results stay pending indefinitely. Switch to --shell until the queue drains.

Detecting a backed-up queue: use the queue status endpoint to see exactly what is waiting:

curl -s -H "Authorization: Bearer $ADMIN_TOKEN" \
  https://hydracluster.experiencenet.com/api/v1/nodes/<ID>/exec/queue | jq .
# {"queued": 4, "in_flight": 1, "items": [...]}

Clearing a backed-up queue: drop all pending (not yet dequeued) execs:

curl -s -X DELETE -H "Authorization: Bearer $ADMIN_TOKEN" \
  https://hydracluster.experiencenet.com/api/v1/nodes/<ID>/exec/queue | jq .
# {"cleared": 4}

To cancel a single exec by ID:

curl -s -X DELETE -H "Authorization: Bearer $ADMIN_TOKEN" \
  "https://hydracluster.experiencenet.com/api/v1/nodes/<ID>/exec/queue/<execId>" | jq .
# {"cancelled": "exec-1716123456789"}

In-flight execs (already dequeued to the agent) cannot be cancelled — only queued ones. A hydracluster restart is the nuclear option to drop in-flight execs as well.

The queue is in-memory on hydracluster — it does not survive a hydracluster restart.

Design principle: exec invokes named commands, not raw shell magic

Exec is for ad-hoc, one-off operations. When a workflow is used regularly by operators (stream control, status checks, etc.), the receiver binary must expose a proper named Cobra subcommand for it. Exec then invokes that named command — no curl+JSON quoting, no piped commands, no inline shell scripting.

Bad (fragile):

hydracluster exec node-X "curl -s -X POST -H 'Content-Type: application/json' -d '{\"experience\":\"mercator-talks\"}' http://127.0.0.1:9740/api/v1/stream/start"

Good (named command on the receiver):

hydracluster exec node-X "hydraheadflatscreen stream-start mercator-talks"

A named command on the binary is testable, readable, and unambiguous. If an exec command starts looking like a shell script (quoting, pipes, flags), that is a signal to add a Cobra subcommand to the receiver instead.

The exec shell does not include ~/.hydranode/bin in PATH. Always use the full binary path when calling receiver-side commands:

# macOS and Linux kiosk heads (hydraheadflatscreen)
hydracluster exec node-X "~/.hydranode/bin/hydraheadflatscreen stream-start mercator-talks"
hydracluster exec node-X "~/.hydranode/bin/hydraheadflatscreen stream-stop"

# Disable the self-service kiosk on a head whose machine doubles as a
# workstation (agent v2.2.3+ idles instead of relaunching the kiosk UI):
curl -sk -X POST https://hydracluster.experiencenet.com/api/v1/nodes/node-X/kiosk-mode \
  -H "Authorization: Bearer <admin-token>" -H "Content-Type: application/json" \
  -d '{"disabled":true}'

Linux heads (Arch/omarchy) provision via recipes/hydraheadflatscreen-linux.yaml: assign the hydraheadflatscreen role and a district; the recipe installs packages, the agent binary, config, a WireGuard sudoers rule, and the systemd --user unit. See the hydraheadflatscreen runbook's "Linux heads (omarchy)" section for head-side details.

API Endpoints

Endpoint Auth Purpose
POST /api/v1/nodes/{id}/exec Admin Submit command, returns {"id": "exec-..."}
GET /api/v1/nodes/{id}/exec/{execId}/result Admin Poll for result (200=done, 202=pending)
GET /api/v1/nodes/{id}/exec/queue Admin Inspect queue: {"queued": N, "in_flight": N, "items": [...]}
DELETE /api/v1/nodes/{id}/exec/queue Admin Clear all queued (not yet dequeued) execs, returns {"cleared": N}
DELETE /api/v1/nodes/{id}/exec/queue/{execId} Admin Cancel one queued exec by ID; 404 if in-flight or already complete
GET /api/v1/body/exec Node token Agent fetches next queued command
POST /api/v1/body/exec/result Node token Agent reports result
POST /api/v1/body/scales Node token hydraskin node reports its Incus inventory (403 unless the node has the hydraskin role)
GET /api/v1/nodes Admin or read-only Node list, including scale inventory

Read-only token

server.read_only_token in the config grants GET /api/v1/nodes and nothing else. It exists so hydrascalerouter — which runs on a public-facing district server — can read which scales claim which domains without holding a credential that could mutate the fleet.

Scales may report a domain and endpoint, carried through verbatim on GET /api/v1/nodes. hydracluster applies no routing policy: it does not validate the domain or resolve conflicts between scales claiming the same one. That belongs to hydrascalerouter, which is the only place that sees every node at once. Both fields are capped at 253 bytes because they are node-supplied and render into the dashboard.


Node Management API

Role Assignment

Assign roles to a node via API (roles must be in the RoleCatalog):

curl -s -X POST https://hydracluster.experiencenet.com/api/v1/nodes/<ID>/roles \
  -H "Authorization: Bearer <ADMIN_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"roles": ["hydraguard-air", "hydraheadflatscreen"]}'

The node agent picks up new roles on its next provision poll and executes the matching recipe.

Renaming a node

curl -s -X POST https://hydracluster.experiencenet.com/api/v1/nodes/<ID>/name \
  -H "Authorization: Bearer <ADMIN_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"name": "ipad-head-lobby"}'
# {"status":"ok","name":"ipad-head-lobby"}

Names must be unique across nodes that are not denied, the same rule enrollment applies. Renaming a node to the name it already holds is a no-op, not a conflict. Errors are 400: name is required for an empty name, node "<name>" already exists for a duplicate, node "<ID>" not found for an unknown node.

Renaming does not propagate to HydraGuard. A head iPad's WireGuard peer is created under the name the node held at provisioning time, and nothing renames it afterwards. The tunnel keeps working — WireGuard matches on public key, not name — but the two systems now disagree. If that head is later re-provisioned, hydracluster asks HydraGuard for a peer under the new name, which allocates a fresh address and orphans the old peer. Removing the orphan is a manual edit today (hydraguard #457).

Renaming is also the fix when a device enrolls under a colliding name — see Recovering a head with no WireGuard config below.

Removing a node

Use this to retire a node hydracluster should forget — decommissioned hardware, a device you no longer possess, or a duplicate/mis-enrolled entry.

curl -s -X DELETE https://hydracluster.experiencenet.com/api/v1/nodes/<ID> \
  -H "Authorization: Bearer <ADMIN_TOKEN>"
# {"status":"ok"}   — or 404 {"error":"node \"<ID>\" not found"}

Three equivalent paths, all of which do the same thing under the hood (Store.Remove(id) + Store.Save()):

What it does and does not do. Removal only deletes the node's entry from hydracluster's YAML store. It does not tear down WireGuard peers, close active sessions, uninstall the agent, or send anything to the device. Practical implications:

Head Management

Head devices (hydraheadflatscreen, hydraheadwindows) stream from body nodes via Moonlight/Sunshine. In production, heads pick their body through the eligibility discovery procedure (GET /api/v1/bodies/eligible). See body-selection.md for the selection discipline and head-identity design principles.

Assign a body to a head (admin override, not the normal path)

The endpoint below is an admin override for exceptional situations. For normal operations, do not manually pin a body to a head. Instead, change the inputs to selection (body district/venue/owner, eligibility rules, drain flag). See body-selection.md.

curl -s -X POST https://hydracluster.experiencenet.com/api/v1/nodes/<HEAD_ID>/head-assignment \
  -H "Authorization: Bearer <ADMIN_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"body_id": "<BODY_NODE_ID>", "app_id": "Desktop"}'

Get head config (polled by agent)

curl -s https://hydracluster.experiencenet.com/api/v1/heads/<HEAD_ID> \
  -H "Authorization: Bearer <ADMIN_TOKEN>"

Returns stream config with both WireGuard (stream_url) and LAN (stream_url_lan) addresses resolved from the body node. The hydraheadflatscreen agent probes LAN first (TCP 47990, 1s timeout) and falls back to WireGuard. Also returns sunshine_username and sunshine_password from the head's district provider config.

Moonlight pairing PIN submission: Sunshine's web UI (port 47990) is localhost-only, so remote heads cannot submit the PIN directly. Use the sunshine-pin proxy endpoint instead — it execs the submission on the body machine where localhost:47990 is always reachable:

curl -s -X POST https://hydracluster.experiencenet.com/api/v1/nodes/<BODY_NODE_ID>/sunshine-pin \
  -H "Authorization: Bearer <ADMIN_OR_NODE_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"pin":"<PIN>"}'

HydraHeadiPad calls this endpoint automatically during self-service pairing.

List all heads

curl -s https://hydracluster.experiencenet.com/api/v1/heads \
  -H "Authorization: Bearer <ADMIN_TOKEN>"

Returns an array of head objects. Each entry includes node_status ("online"/"offline" — node heartbeat state), last_seen (RFC3339 timestamp of the last hydranode heartbeat), assigned_body_name (the name of the body currently in use — see priority below), and diagnostics (a map of agent-reported key/value pairs, omitted when empty) in addition to the head-level status ("idle"/"streaming"/"error" — reported by the kiosk agent). Use node_status to distinguish a kiosk that is genuinely offline from one that is online but idle. Use assigned_body_name to correlate heads to bodies without resolving IP addresses. Diagnostic keys set by hydraheadflatscreen include version, wireguard ("up"/"down"), app ("kiosk"/"moonlight"/"none"), routing ("lan"/"wireguard"/"unknown"), and latency_ms (TCP RTT in milliseconds to Sunshine port 47990, present only when streaming).

assigned_body_name resolution priority: (1) live_body_id — the body node ID reported in the most recent hydraheadflatscreen heartbeat (PUT /api/v1/heads/{id} body_id field); clears automatically when no stream is active. (2) head_body_id — the admin-configured body assignment (set via POST /api/v1/nodes/{id}/head-assignment). Self-service kiosk heads populate the live field; centrally-assigned heads populate the admin field. The stream config (stream_url, stream_url_lan) is always derived from the admin assignment only.

iPad fleet enrollment QR (HydraHeadiPad)

iPad enrollment is a two-step self-registration flow. No head entry needs to exist before scanning.

Step 1 — display the QR code (admin web UI, recommended):

  1. Open https://hydracluster.experiencenet.com/enroll
  2. Click the iPad tab
  3. The QR code is shown inline — point the iPad at the screen

Step 1 (alternative) — get the raw payload via API:

curl -s https://hydracluster.experiencenet.com/api/v1/enroll-qr \
  -H "Authorization: Bearer <ADMIN_TOKEN>"

Returns {"server_url":"...","enrollment_token":"..."}. Encode this JSON as a QR code and print/display it — one QR serves all iPads in the fleet.

Step 2 — iPad scans the QR, app self-registers:

The app POSTs to POST /api/v1/heads using the enrollment token:

curl -s -X POST https://hydracluster.experiencenet.com/api/v1/heads \
  -H "Authorization: Bearer <FLEET_ENROLLMENT_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"name":"ipad-lobby","district":"bxl1","venue":"cloud-seven"}'

Returns {"head_id":"node-...","token":"...","server_url":"..."}. The iPad saves these as its permanent identity. The head is auto-approved with role hydraheadipad and OS ios, and immediately starts heartbeating. It appears in GET /api/v1/heads with type: "hydraheadipad".

If HydraGuard is configured (hydraguard_url + hydraguard_token in config.yaml), enrollment also calls POST /api/v1/headipad/provision on HydraGuard and stores the resulting WireGuard config and peer IP on the node. Failure is non-fatal — the head is still enrolled.

Config: fleet_enrollment_token must be set in config.yaml under server:. The token is the same for all iPads — treat it as a moderately sensitive credential (anyone with it can enroll new heads).

Fetch iPad WireGuard config

After enrollment, the head's WireGuard configuration can be retrieved with the admin token:

curl -s https://hydracluster.experiencenet.com/api/v1/heads/<HEAD_ID>/wireguard-config \
  -H "Authorization: Bearer <ADMIN_TOKEN>"

or with that head's own node token, which is what the HydraHeadiPad app uses to set up its own tunnel without an operator scanning a QR off a screen (hydracluster #449):

curl -s https://hydracluster.experiencenet.com/api/v1/heads/<HEAD_ID>/wireguard-config \
  -H "Authorization: Bearer <THAT_HEADS_NODE_TOKEN>"

Returns the WireGuard config as text/plain (the full [Interface] + [Peer] block), with Cache-Control: no-store. Returns 404 if the head was enrolled without HydraGuard provisioning.

An unknown node ID returns 404 for the admin token, but 403 for a node token: the middleware refuses before any lookup, so it never reveals whether a given ID exists. A denied node also gets 403 even for its own ID, so denying a node revokes its access to its own tunnel credential.

Auth model: admin credential, or the head's own node token, and nothing else. Any other valid node token (another head, any body, any skin) gets 403, and a token that is not enrolled at all gets 401. The route deliberately does not use requireAdminOrNodeToken, which authenticates any enrolled node and never compares it to {id} (hydracluster #456): behind that wrapper every node token in the fleet could read every head's tunnel private key. The self-scoped wrapper lives in pkg/api/selfscope.go and compares the authenticated node's ID to the path ID.

The config embeds the tunnel's private key. HydraGuard shows a private key only at creation and never stores it, so this endpoint and the QR below are the only copies. Treat the output accordingly: do not paste it into tickets or chat, and do not publish it anywhere reachable without auth.

Show the WireGuard QR for an iPad

The WireGuard iOS app imports a tunnel by scanning a QR. The node detail page renders one for any head that has a stored config:

  1. Open /admin/nodes/<HEAD_ID> in hydracluster admin (signed in).
  2. Expand Show enrollment QR under the WireGuard heading.
  3. In the WireGuard iOS app: Add tunnel → Create from QR code, and scan the screen.

The image itself is at GET /admin/heads/{id}/wireguard-qr.png (session-cookie auth, Cache-Control: no-store). It returns 404 when the head has no stored config.

The section only appears when a config exists, and the QR is collapsed by default so that opening a head does not put a VPN credential on screen. For the same reason it is behind cookie auth: it must never be exposed publicly.

Recovering a head with no WireGuard config

Symptom: wireguard-config returns 404 and the node page shows no WireGuard section, on a head that should have a tunnel.

WireGuard provisioning at enrollment is non-fatal. If HydraGuard rejects it, the head still enrolls and nothing surfaces the failure — the only trace is a line in the hydracluster log at enrollment time:

journalctl -u hydracluster --since '-1h' | grep -i 'WireGuard provisioning'
# WireGuard provisioning for <name> failed (non-fatal): hydraguard returned 409:
#   head iPad peer "<name>" already exists on 10.10.200.N/32; its private key was
#   shown only at creation and is not stored.

The common cause is a name collision. The iPad app sends no name, so handleFleetEnrollHead defaults every device to ipad-head; if a peer of that name survives from a previous device, HydraGuard returns 409 and the new head gets no tunnel.

There is no re-provision endpoint — provisionHeadIPadWireGuard is only called from the enrollment handler — so recovery means enrolling again under a name that is free on both sides:

  1. Confirm the collision: check the log line above for the peer name and address.
  2. Free the name in HydraGuard by removing the stale peer, or pick a different name. Removal has no supported command today (hydraguard #457); until it lands, follow the safe manual procedure recorded there rather than editing mesh.yaml freehand.
  3. Delete the affected head (DELETE /api/v1/nodes/{id}) so its name is free here too.
  4. Re-enroll the iPad from /enroll, then rename it to something device-specific (see Renaming a node), noting that the HydraGuard peer keeps the enrollment-time name.

Verify with GET /api/v1/heads/{id}/wireguard-config returning 200, and the peer present on the hub: wg show wg0 dump | awk '$4 ~ /10.10.200./'.

Decommissioning. Deleting a head does not remove its WireGuard peer. When a device leaves — returned rental, retired hardware — revoke the peer on the hub as well, or it keeps routed access to the mesh (AllowedIPs covers 10.10.0.0/16 and 10.0.0.0/8). See hydraguard #457.

Get experience catalog for a head

curl -s https://hydracluster.experiencenet.com/api/v1/heads/<HEAD_ID>/experiences \
  -H "Authorization: Bearer <ADMIN_TOKEN>"

Returns the catalog of experiences a kiosk should display, sourced from hydraexperiencelibrary's /api/v1/experiences/live endpoint keyed by the head's district + venue. If the head has a non-empty AllowedExperiences whitelist, the result is intersected with it.

This endpoint is the source of truth for kiosk grids and is intentionally independent of body availability — a venue with no body online still returns its catalog (or [] if nothing is planted there yet). Body discovery only runs at stream-start time (POST /api/v1/stream/start on the agent's local API).

Status codes:

Fetch kiosk desktop screenshot

Returns image/jpeg directly. The mechanism differs by head type:

TOKEN="<admin-token>"
curl -sf -H "Authorization: Bearer $TOKEN" \
  "https://hydracluster.experiencenet.com/api/v1/heads/<HEAD_ID>/screenshot" \
  > /tmp/kiosk.jpg

Flatscreen heads (hydraheadflatscreen): fetches via exec channel → curl http://127.0.0.1:9740/api/v1/screenshot | base64 on the head. Requires the Terminal screenshot loop to be running (see hydraheadflatscreen runbook → Screenshots). Times out after 20 s.

Status codes: 200 JPEG on success; 502 if screenshot file is stale or absent (Terminal loop not running); 504 if the exec channel doesn't respond within 20 s.

If 504 (exec queue backed up): use the WebSocket shell path instead — it bypasses the async exec queue:

"$HC_BIN" exec <node-id> "base64 -i /tmp/hydra-live-screenshot.jpg" \
  --shell --server "$HC_SERVER" --admin-token "$HC_TOKEN" --timeout 30s \
  > /tmp/raw.txt
grep -E '^[A-Za-z0-9+/]+=*$' /tmp/raw.txt | tr -d '\n' | base64 -d > /tmp/kiosk.jpg

Note: macOS base64 requires -i <file> not a positional argument. The grep strips shell prompts before decoding.

iPad heads (hydraheadipad): no exec channel — uses a push flow instead. The server sets a pending flag and waits up to 15 s for the iPad to poll GET /api/v1/heads/{id}/commands (every 3 s), capture via RPScreenRecorder, and upload to POST /api/v1/heads/{id}/screenshot. Full Metal framebuffer is captured (Moonlight video included). Times out after 15 s.

Status codes: 200 JPEG on success; 504 if the iPad does not respond within 15 s (app backgrounded, device offline, or still booting).

The admin node detail page also exposes a Take Screenshot button for iPad heads — click it and the JPEG appears inline within ~5 s.

Fetch kiosk agent logs

Returns the last 200 lines of the head agent log as JSON ({"log_path":"...","lines":[...]}):

curl -sf -H "Authorization: Bearer $TOKEN" \
  "https://hydracluster.experiencenet.com/api/v1/heads/<HEAD_ID>/logs" | jq '.lines[-20:]'

Status codes: 200 JSON on success; 504 if the head exec channel doesn't respond within 15 s.

XR Sessions (Quest heads)

Quest heads (role hydraheadquest) stream immersively through ALVR instead of Moonlight (issue #556). The cluster brokers one XR session per body. The session record lives on the body's entry in nodes.yaml, so it survives a cluster restart. The flat session store never sees XR sessions.

How it works:

  1. The head calls POST /api/v1/heads/{id}/xr-session with a body it picked from GET /api/v1/bodies/eligible?head_id=<id>&stream_mode=xr.
  2. The body learns about the session in its next POST /api/v1/body/status response (xr_session directive). It stops Sunshine, registers the ALVR driver, starts the chain, and reports xr_state back on the same channel.
  3. The head polls GET /api/v1/heads/{id}/xr-session until state is armed, then launches the ALVR client. Arming takes 45 to 90 seconds, worst case about 110.
  4. DELETE /api/v1/heads/{id}/xr-session ends the session. The body tears down, unregisters the driver, restarts Sunshine, and returns to the flat pool.

Session states: requested, arming, armed, active, draining, ending, failed. Failure reasons appear in the reason field (arm_timeout, connect_timeout, chain_unstable). Failed sessions are held for 60 seconds so the head can read the reason, then cleared.

Cluster timers (10 second cadence, in the session watchdog):

Enable a body for XR:

# Add the alvr role next to hydrabody (ALVR must be staged at C:\hydra\alvr,
# see the hydrabody runbook). Then restart hydrabody or set Reprovision:
# body config only propagates on first boot or reprovision.
curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"roles":["hydrabody","alvr"]}' \
  "https://hydracluster.experiencenet.com/api/v1/nodes/<BODY_ID>/roles"

Set the head's ALVR client hostname (also editable on the node detail page):

curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"hostname":"0529.client"}' \
  "https://hydracluster.experiencenet.com/api/v1/nodes/<HEAD_ID>/xr-client-hostname"

Enroll a Quest head through fleet enrollment by adding "type":"hydraheadquest" to the POST /api/v1/heads body. Without the field the endpoint keeps today's iPad behavior.

Exclusivity rules:

XR session endpoints

Endpoint Auth Purpose
POST /api/v1/heads/{id}/xr-session Admin or node token Create an XR session. Body: {"body_id","experience","client_hostname","client_ip"}. Returns 201 with the GET shape. 409 {"error":"body_busy"} when the body is taken, 409 {"error":"not_xr_capable"} when it cannot serve XR.
GET /api/v1/heads/{id}/xr-session Admin or node token Read the head's session: {"session_id","state","body_id","experience","body_host","reason"}. body_host is the body's WireGuard IP. 404 when the head has no session.
DELETE /api/v1/heads/{id}/xr-session Admin or node token End the session. Idempotent, always 204.
POST /api/v1/nodes/{id}/xr-client-hostname Admin Set the head's ALVR client hostname, e.g. 0529.client.
GET /api/v1/bodies/eligible?head_id=<id>&stream_mode=xr Admin or node token Filter to XR-capable idle bodies. Entries gain xr_drivers and xr_state.

API Endpoints (Node Management)

Endpoint Auth Purpose
POST /api/v1/nodes/{id}/roles Admin Set roles on a node
DELETE /api/v1/nodes/{id} Admin Remove a node from the store. Returns {"status":"ok"}, or 404 if the ID is unknown. Store-only: no WireGuard peer teardown, no session cleanup, no signal to the device. See "Removing a node" below.
POST /api/v1/nodes/{id}/update-node Admin Flag node for immediate hydranode self-update
POST /api/v1/nodes/{id}/update-services Admin Flag node to update all provisioned service binaries
POST /api/v1/nodes/{id}/head-assignment Admin Assign body to head
POST /api/v1/nodes/{id}/allowed-experiences Admin Set experience filter for kiosk head
DELETE /api/v1/nodes/{id}/stream Admin Close the active Sunshine session on a body and clear stream_status (orphaned stream recovery). Sends exec-based ForceStop (synchronous, 15 s timeout) AND sets TerminatePending so the body receives TerminateStream=true on its next heartbeat (≤5 s) as a fallback if exec fails. Session closed as admin_stop. See body-recovery.md.
GET /api/v1/heads Admin List all head nodes with stream config
GET /api/v1/heads/{id} Admin Get head config (includes experience_library_url, allowed_experiences)
GET /api/v1/heads/{id}/experiences Admin Get the kiosk experience catalog for a head (proxied from hydraexperiencelibrary)
GET /api/v1/heads/{id}/screenshot Admin Fetch kiosk screen as JPEG — flatscreen: exec channel (20 s timeout); iPad: push flow via commands poll (15 s timeout)
POST /api/v1/heads/{id}/screenshot Admin or node token iPad uploads captured JPEG after seeing screenshot: true in commands response
GET /api/v1/heads/{id}/commands Admin or node token Lightweight fast-poll for pending iPad commands (iPad polls every 3 s); returns {"screenshot": true/false}
GET /api/v1/heads/{id}/logs Admin Fetch last 200 lines of head agent log as JSON via exec channel
GET /api/v1/heads/{id}/wireguard-config Admin or that head's own node token Fetch stored WireGuard config (text/plain, Cache-Control: no-store) provisioned at iPad enrollment. Self-scoped: any other node's token gets 403, an unenrolled token 401. Used by HydraHeadiPad to set up its own tunnel
POST /api/v1/heads/{id}/stream Admin Start stream on flatscreen head (exec hydraheadflatscreen stream-start <experience>)
DELETE /api/v1/heads/{id}/stream Admin or node token Stop stream. Flatscreen: exec hydraheadflatscreen stream-stop on the head (blocks up to 15 s). iPad: immediately marks body idle in cluster memory, sets TerminatePending so body receives TerminateStream=true on its next heartbeat (≤5 s), closes session as head_stop, and fire-and-forgets a body exec as belt-and-suspenders; returns immediately. Request body {"body_id":"<id>"} is optional for iPads (falls back to last heartbeat live_body_id).
PUT /api/v1/heads/{id} Admin Update head status

API Endpoints (Sessions and Log Access)

Endpoint Auth Purpose
GET /api/v1/sessions Admin List all currently active streaming sessions
GET /api/v1/sessions/history Admin List closed sessions (last 500), most recent first
GET /api/v1/sessions/{id}/logs Admin Fetch correlated body + head logs for a session; see below
GET /api/v1/nodes/{id}/body-logs Admin Fetch last 500 lines of hydrabody.log from a body node; accepts ?since=<RFC3339>&until=<RFC3339> for time-window filtering

Fetch body logs

Returns last 500 lines of C:\Windows\System32\config\systemprofile\.hydrabody\hydrabody.log from the body node via the exec channel:

curl -sf -H "Authorization: Bearer $TOKEN" \
  "https://hydracluster.experiencenet.com/api/v1/nodes/<BODY_ID>/body-logs" | jq '.lines[-20:]'

Optional time filter (RFC3339):

curl -sf -H "Authorization: Bearer $TOKEN" \
  "https://hydracluster.experiencenet.com/api/v1/nodes/<BODY_ID>/body-logs?since=2026-05-22T10:00:00Z&until=2026-05-22T10:30:00Z" | jq '.lines[]'

Status codes: 200 JSON on success; 400 if node is not a body node; 504 if the exec channel doesn't respond within 30 s.

Fetch session logs (correlated)

Fetches both the body log and the head log for the duration of a specific session, returned in a single response. Session ID comes from GET /api/v1/sessions or GET /api/v1/sessions/history.

curl -sf -H "Authorization: Bearer $TOKEN" \
  "https://hydracluster.experiencenet.com/api/v1/sessions/<SESSION_ID>/logs" | jq '{body_logs: .body_logs[-10:], head_logs: .head_logs[-10:]}'

Response shape:

{
  "session_id": "abc12345-7",
  "body_id": "node-abc12345",
  "body_name": "cosmic-pretzel-98",
  "head_id": "node-xyz99",
  "started_at": "2026-05-22T10:15:00Z",
  "ended_at": "2026-05-22T10:45:00Z",
  "body_logs": ["2026/05/22 12:15:00 [stream] ...", "..."],
  "head_logs": ["...", "..."]
}

head_logs is empty ([]) if the session has no associated head ID (session opened before head heartbeat was received). ended_at is null for active sessions.

Status codes: 200 on success (even if exec failed — errors are logged server-side and the affected log array is empty); 404 if session not found.


Render Node Provisioning

Render nodes are Windows machines that run LarkXR Standalone (cloud rendering). They are enrolled in HydraCluster, which provisions them via the node agent.

Prerequisites

Fresh Enrollment

  1. Open https://hydracluster.experiencenet.com/enroll on the machine
  2. Select the Windows / Linux tab (default), fill in a name, set owner, submit
  3. Copy the PowerShell install command and run it as Admin
  4. The node agent installs, starts, and the node goes online in the dashboard
  5. In the admin dashboard, assign the render-node role and set district/venue

The node agent will:

Verification

After provisioning, verify these ports are listening:

Check via remote shell:

Get-Process *lark*,*java*,*mysql*,*redis*,*nginx* | Format-Table Name,Id
netstat -an | Select-String '8181|8282|13306'

The dashboard should show provider_status: running.

Reprovision (Clean Reinstall)

When LarkXR needs to be fully reinstalled:

  1. Stop all LarkXR processes:

    Stop-Process -Name LarkXRLauncher,LarkXRServer,CloudLarkRenderServer,java,mysqld,nginx,redis-server -Force -ErrorAction SilentlyContinue
    
  2. Remove old installation and cached state:

    Remove-Item C:\LarkXR -Recurse -Force
    Remove-Item C:\Windows\System32\config\systemprofile\.hydranode\provider_version.txt -Force
    Remove-Item C:\Windows\System32\config\systemprofile\.hydranode\config_cache.yaml -Force
    Remove-Item C:\Windows\System32\config\systemprofile\.hydranode\downloads -Recurse -Force
    
  3. The node agent will detect the missing installation on its next heartbeat (30s) and reprovision automatically.

Alternatively, trigger reprovision from the dashboard (node detail page > Reprovision button). As of v0.23.10+, the Reprovision button triggers a forced reinstall -- the node agent stops the provider and re-downloads regardless of existing files.

Recovery (Node Agent Down)

If the node is offline (node agent not heartbeating):

Windows:

# Check if the scheduled task exists
schtasks /query /tn HydraNode

# Start it
schtasks /run /tn HydraNode

# If missing, reinstall
C:\hydranode\hydranode.exe install
schtasks /run /tn HydraNode

Recovery instructions are also available at https://hydracluster.experiencenet.com/enroll (no admin login needed).

Node Agent Update

The node agent auto-updates from the release server. As of v0.23.9:

To force an update on a remote machine via API:

curl -s -X POST https://hydracluster.experiencenet.com/api/v1/nodes/<ID>/update-node \
  -H "Authorization: Bearer <ADMIN_TOKEN>"

The node picks up the flag on its next heartbeat and triggers the self-updater immediately. To update service binaries (not hydranode itself):

curl -s -X POST https://hydracluster.experiencenet.com/api/v1/nodes/<ID>/update-services \
  -H "Authorization: Bearer <ADMIN_TOKEN>"

Manual update on Windows (last resort):

# Download new binary
Invoke-WebRequest -Uri 'https://releases.experiencenet.com/hydranode/production/latest/hydranode-windows-amd64.exe' -OutFile C:\hydranode\hydranode-new.exe

# Stop, replace, reinstall, start
schtasks /End /TN HydraNode
Start-Sleep 3
Copy-Item C:\hydranode\hydranode-new.exe C:\hydranode\hydranode.exe -Force
Remove-Item C:\hydranode\hydranode-new.exe
C:\hydranode\hydranode.exe install
schtasks /Run /TN HydraNode

Warning: Stopping the node agent on a remote-only machine (private LAN) means you lose remote access until it restarts. The task's repetition interval (1 minute) should auto-restart it, but if the binary is locked, the replacement will fail silently.

Key Paths

Path Purpose
C:\hydranode\hydranode.exe Node agent binary
C:\hydranode\enroll.yaml Enrollment token
C:\Windows\System32\config\systemprofile\.hydranode\ SYSTEM profile data dir
...\.hydranode\config.yaml Node config (server URL, token)
...\.hydranode\config_cache.yaml Cached config from server
...\.hydranode\provider_version.txt Installed provider version
...\.hydranode\hydranode.log Node agent log
C:\LarkXR\larkxr-standalone\ LarkXR installation
C:\LarkXR\larkxr-standalone\log\ LarkXR Launcher logs

Known Issues

HydraSkin (Container Hosts)

Nodes with the hydraskin role run Incus as container hosts. A container on such a node is called a scale.

The operational runbook lives in the hydraskin repo: hydraskin/docs/runbooks/hydraskin.md — instance kinds (OCI / system container / VM), scale defaults and overrides, disk caps and the btrfs quota requirement, hydraskin project and hydraskin expose (publishing a scale on the node's LAN address so the mesh can reach it), GOMEMLIMIT guidance, debugging the reporter, port 8443, and reprovisioning. Moving a service onto a scale is covered in hydraskin/docs/runbooks/service-cutover.md.

What lives in this repo:

Note the recipe is loaded at startup from /root/.hydracluster/recipes/ and is deployed separately from the binary — editing it in git is not enough, it must be copied to the server and hydracluster restarted.

Backup & Recovery

What is backed up

Hetzner automated daily server snapshots are enabled on hydracluster (46.224.29.125), context hydraexperiencenet. Backup window: 14:00–18:00 UTC, 7-day retention.

The snapshot covers the entire server disk, including:

Restore procedure

  1. In the Hetzner Cloud Console (hydraexperiencenet project), open Servers → hydracluster → Backups.
  2. Identify the snapshot to restore from. Restoring overwrites the current disk — all changes since the snapshot are lost.
  3. Power off the server, restore the snapshot, then power on.
  4. Verify the service came back up:
    curl https://hydracluster.experiencenet.com/api/v1/health
    
  5. Check node count matches expectations:
    curl -H "Authorization: Bearer $TOKEN" https://hydracluster.experiencenet.com/api/v1/nodes | jq length