NodeForEdge

Documentation / Platform

Infrastructure

One VPS, Caddy at the edge, Docker and systemd for services, and a private tailnet for everything that should not be public.

The whole lab runs on a single small VPS. That is a deliberate choice: one machine is easy to reason about, cheap to run and quick to rebuild, and the edge phones supply the capabilities a VPS lacks.

The machine

  • A small cloud VPS running Ubuntu 24.04 LTS with a modest disk.
  • Public IPv4 and IPv6, with DNS at Cloudflare. Records point straight at the server; Cloudflare is used for DNS only, not as a proxy.
  • Services are plain systemd units for single-process programs, and Docker Compose for JobHunter, which brings its own PostgreSQL and Redis.
  • A host PostgreSQL 16 cluster serves the services that are not in Docker.

Request path

 Internet ─▶ Caddy (TLS, routing, auth rules)
               ├─▶ nodeforege.win           static landing
               ├─▶ status.nodeforege.win    static status page + JSON feed
               ├─▶ docs.nodeforege.win      static documentation
               ├─▶ jobhunter.nodeforege.win ─▶ API on loopback
               ├─▶ phonegate.nodeforege.win ─▶ Web Studio / MCP on loopback
               └─▶ other hosts ─▶ MCP pools, relays, gateways on loopback

 Tailnet (private) ─▶ SSH, debug links to the phones, updates, health checks

Applications listen on loopback or on a private Docker bridge. Only Caddy faces the internet.

Caddy

Caddy terminates TLS with automatic certificates, routes by hostname and applies rules that belong at the edge: bearer checks in front of MCP endpoints, network restrictions on privileged paths, mutual TLS for the A14 tunnel, and security headers on the static sites.

Configuration is split into a main file plus one small file per project, so a change to one site cannot break another. It is validated before every reload.

Private network

Tailscale connects the VPS and the phones into one private network. Management traffic uses it, and some paths, such as device updates, are accepted from it only.

Hardening and resource safety

  • fail2ban blocks repeated failed logins.
  • earlyoom kills the biggest offender before the kernel runs out of memory, instead of letting the machine freeze.
  • A pressure watchdog watches CPU and memory pressure and acts only on sustained load, to protect the services that matter most.
  • SSH is key-based, and public admin surfaces need their own authentication on top.

Observability

  • Netdata shows real-time host metrics.
  • Applications write structured logs, and JobHunter exports Prometheus metrics.
  • The public status page shows component health.

Backups

Database backups run as a scheduled job and are checked with a restore script. Cold data is archived to cloud storage, with a matching restore script, so that a rebuild starts from stored data rather than from memory.

Rebuilding

Services live in /srv, each in its own directory with its own virtual environment or Compose file. Configuration for the web edge is plain text. Rebuilding means recreating the base system, restoring data, and bringing the units back up.

Built 2026-10-06. Addresses, tokens and ports are left out on purpose.