The whole of Kidslen — the Kotlin API, the parent dashboard, the ops console, the marketing site, PostgreSQL, Prometheus, Grafana, Alertmanager, and the CI that builds all of it — runs on one Proxmox LXC container in a room in my flat. Public HTTPS, four hostnames, a working push-to-deploy pipeline, and not a single inbound port open on my router.
People assume this means “cheap but fragile”. It is cheap. It is not fragile, because the fragile parts were designed out rather than paid away. This post is the infrastructure chapter of the series: what the box actually does, the trick that makes deploys safe, and what the same setup would cost on a cloud provider if I stopped being stubborn.
One container, not twelve
The instinct with a stack this size is to spread it: a VM per service, or Kubernetes, or a managed database because databases are scary. I did none of that.
There is one unprivileged LXC container. Inside it: systemd units for the Spring Boot API, nginx serving three static bundles (dashboard, ops console, marketing site), PostgreSQL, and the monitoring trio. Everything talks over localhost. There is no service mesh because there is no mesh — there is a loopback interface.
The reason is not laziness, it is failure-surface arithmetic. Every network hop between two components of my own system is a hop that can time out, retry, half-open, or fail in a way I have to write code to handle. A solo builder’s most limited resource is attention. Distributing a system you do not need to distribute spends that attention on plumbing instead of on the product.
Proxmox gives me the two things I actually wanted from virtualisation: snapshots before risky changes, and a container boundary so that a runaway build cannot take the host down with it. Backups are pg_dump plus a snapshot, off the box, nightly.
The runner that builds the machine it lives on
This is the part that raises eyebrows. Kidslen’s CI runs on GitHub Actions, but the Linux jobs execute on a private runner inside the same container that hosts production.
So the sequence on a push to main is:
- GitHub receives the push and dispatches the workflow.
- The runner agent in the LXC picks up the job.
- It builds the Kotlin backend with Gradle, runs the 57 backend tests, builds the three frontends, runs their unit suites and the Playwright end-to-end specs.
- If everything is green, it writes the new artefacts beside the running ones and restarts the systemd units.
No artefact ever leaves the machine. No registry, no SSH deploy step, no secret in GitHub that can push code onto my box — GitHub cannot reach the runner at all; the runner reaches out to GitHub. That inversion is the whole security story. The CI system has no credential for my infrastructure because it never connects to it.
The obvious objection is that a build competing for CPU with production is asking for trouble. It is, and I handled it the boring way: the runner service is nice-ed and memory-capped by its systemd slice, Gradle’s worker count is pinned well below the core count, and the API has a health endpoint that Alertmanager watches during builds. In several months the only latency I have measured during a build is on the ops console’s heavier report queries, and the ops console has one user.
iOS builds cannot happen in a Linux container, so there is a second private runner on a Mac mini, labelled macos, that only ever picks up the companion app’s iOS jobs. It is a separate machine with a separate label and nothing else on it. Apple’s toolchain is the one part of this stack I cannot containerise, and pretending otherwise would have cost me a week.
The gate: a smoke test that walks the real product
Green unit tests do not mean the product works. I learned this expensively, and it is the subject of the next post in this series. The defence I built is a deploy gate that does not test code — it tests the running system.
After the units restart and pass their health checks, the workflow runs a script that performs the entire core journey against the live API, as a real client would:
| Step | What it proves |
|---|---|
| Register a parent account | Auth, Argon2id hashing, token issue |
| Create a child profile | Ownership model, deny-by-default authorisation |
| Generate an enrollment code | Short-lived code issue and storage |
| Enroll a simulated device with that code | The pairing handshake, device identity |
| Send a heartbeat | Device liveness, refresh-token rotation |
| Ingest a batch of activity events | The queued JSONB pipeline, dedup |
| Read the activity feed back | Processor, projections, date handling |
| Fetch the ops overview | Cross-tenant aggregation, ops auth |
If any step fails, the workflow fails loudly and I get an alert. The script uses a dedicated test tenant that the retention job cleans up, and it runs against production because a smoke test against a staging environment proves things about staging.
That table is deliberately the same list as the product’s core promise. A parent installs the companion app on their child’s phone, the child sees the consent screen and the permanent banner, and the parent sees activity appear. If that chain works, Kidslen works. If it does not, nothing else matters — so that chain is exactly what the gate walks, on every deploy, before I am allowed to call it shipped.
Adding this gate was the single highest-value day of infrastructure work in the project. It has caught a broken database migration, a misconfigured CORS origin, and a serialisation change that compiled fine and returned the wrong shape.
Cloudflare Tunnel: four public hostnames, zero open ports
My router forwards nothing. Not 443, not 80, not SSH. Instead a cloudflared daemon in the container opens outbound connections to Cloudflare’s edge and holds them open; Cloudflare terminates TLS at the edge and sends requests back down those connections.
Four hostnames map to four local services:
kidslen.app— the Astro marketing siteapp.kidslen.app— the parent dashboardapi.kidslen.app— the Spring Boot APIops.kidslen.app— the ops console
The security properties are better than anything I would have achieved by hand. There is no listening port for anyone to scan. My home IP address is not in DNS. Certificate renewal is not my problem. DDoS absorption happens at an edge with far more capacity than my uplink.
The ops console gets a second layer: Cloudflare Access sits in front of ops.kidslen.app and requires an identity-provider login before a request is ever proxied to my container. The console has its own authentication too — Access is not a replacement for authorisation, it is a bouncer at the door of the building. Someone who finds the hostname gets an identity challenge from Cloudflare, not an HTML page to probe.
One honest caveat: this makes Cloudflare a dependency. If the tunnel is down, Kidslen is unreachable, and I cannot fix that by SSHing somewhere. I accepted it because the alternative — my own edge, my own certificates, my own IP exposed — is a larger risk surface operated by one person who also has to write the product.
What it costs, honestly
Here is the real bill, and then the comparison people actually want.
What I pay now: the two domains (kidslen.com and kidslen.app, bought for two years), electricity for a box that was already running, and my home internet connection, which I would pay for anyway. Cloudflare Tunnel and Access are on plans that cost me nothing at this scale. The Mac mini was already on the desk.
The marginal monthly cash cost of running Kidslen’s infrastructure is, rounded honestly, the domains amortised — and nothing else.
What the cloud equivalent would be. I am not going to invent a number, so let me be precise about the shape instead. To reproduce what the container does, I would be paying for:
| Component | Cloud equivalent |
|---|---|
| API + frontends | A small always-on compute instance |
| PostgreSQL | A managed database instance with backups |
| Prometheus / Grafana / Alertmanager | A managed observability plan, or a second instance |
| CI minutes | Hosted Linux runners, plus macOS runners for iOS at several times the Linux rate |
| Egress | Metered, and SSE connections are long-lived |
Every one of those is individually modest and collectively is a real monthly subscription — the kind of number that is trivial for a funded team and meaningfully annoying for a product with no revenue. The macOS runner line is the one that surprises people: hosted macOS minutes are the most expensive thing in ordinary mobile CI, and a Mac mini on a desk removes that line item entirely.
What I pay instead of money. This is the part that infrastructure posts skip. I pay in:
- Operations attention. When Postgres needs an upgrade, that is my evening.
- A single point of failure. One machine, one flat, one internet connection. Snapshots and off-box backups bound the damage; they do not make it zero.
- A hard ceiling. This setup is correct for an early-access product with a small number of families. It is not correct at ten thousand devices, and I know which metrics will tell me — ingest queue depth and SSE connection count. When those numbers move, the architecture moves. Nothing in the design assumes the container is permanent: the API is stateless apart from the database, the frontends are static bundles, and the tunnel does not care what is behind it.
That last point is what makes the whole thing defensible rather than merely cheap. The infrastructure is small on purpose and portable by construction.
Why this matters for a product about children
A product that monitors minors is a product that holds sensitive data about people who cannot consent for themselves in any legal sense. Every additional place that data is copied — a hosted CI cache, a third-party log aggregator, a build artefact registry — is another place it can leak from, and another vendor whose breach becomes my breach.
Running the stack in one container I control is not a cost optimisation dressed up as a principle. It genuinely reduces the number of systems that ever touch a child’s activity data, which is the same reason the product collects titles and durations rather than screenshots. Small blast radius, everywhere, on purpose.
The build artefacts never leave the machine. The database never leaves the machine. The only things that cross the boundary are the outbound tunnel connections and the push notifications, and both of those carry the minimum they can.
Next in this series: the week everything broke — a threading bug that only appeared on the slow machine, a native build that refused to compile, and a date format that returned HTTP 400 in front of a camera.
If you want to see what this infrastructure actually serves, it is at kidslen.app.