Monitoring and status
How the public status page gets its numbers, what each state means, and what the page deliberately does not show.
The status page is a static page that reads one JSON file. A small probe on the server produces that file every minute.
The probe
A systemd timer starts a Python script once a minute. The script:
- Runs a check for every component, in parallel.
- Retries a failed network check once after a moment, so one lost packet does not count as an outage.
- Stores each result in a local SQLite history.
- Compares the new state with the previous one and records a change event.
- Rebuilds the public JSON file and replaces it atomically, so the page never reads a half-written file.
Checks run from the server itself. Network checks go through the real public hostname when there is one, so a pass means Caddy, TLS and the application all work.
What is checked
| Kind | Used for | A pass means |
|---|---|---|
| Public HTTPS request | JobHunter, PhoneGate, the web edge | The expected status code and body came back through Caddy |
| Local service check | MCP pools, relays, PostgreSQL | The systemd unit is active and its port accepts connections |
| Container health | JobHunter workers, database and queue | Every expected container reports healthy |
| Private network | Android nodes, Tailscale | The node is online in the tailnet |
| Tunnel listener | A14 egress tunnel | The reverse tunnel is up, meaning the phone is connected |
States
| State | Meaning |
|---|---|
| Operational | The check passed. |
| Degraded | Partly working. For example, some workers are healthy and others are not. |
| Down | The check failed after a retry. |
| Offline | Used for a phone that is not online in the tailnet. |
Uptime numbers
Every probe adds a score: 1 for operational, 0.5 for degraded and 0 for down. A day's bar is the average of that day's scores, and the percentage on the right is the average over the last 60 days. A bar is green at 99.9% or higher, amber from 95%, and red below that. Grey means no data.
The headline
The banner counts only core components. Dev and tool components, such as the development phone, can be offline without turning the banner amber. When that happens, a short note under the headline says so.
- All core components up: All core systems operational.
- Some core components down or degraded: N core components affected.
- Half or more of the core components down: Major outage.
What the page leaves out
The feed holds names, states, latencies and uptime. It never contains internal hostnames, ports, addresses, versions or error text. A failing check shows as a state, and the reason stays in the server logs.
Reading the feed
The same data is available as JSON at https://status.nodeforege.win/status.json, with permissive cross-origin access, which is how the live dots on the home page work.