The most useful sentence in Kidslen’s design document is not about what it contains. It is a constraint:
The largest realistic tenant is one family with four children and six connected platforms.
Everything below follows from that. Part 2 ended with the observation that most of the complexity I had reviewed in a legacy system existed to compensate for something simple that had not been done. This post is my attempt to not repeat that, by writing down the things I deliberately left out and why.
The stack, briefly
- Backend: Kotlin 2, Spring Boot 3.4, PostgreSQL. One service, one container,
api.kidslen.app. - Parent dashboard: React 19 + Vite,
app.kidslen.app. - Ops console: React 19 + Vite,
ops.kidslen.app, behind Cloudflare Access. - Companion app: React Native, installed on the child’s phone.
- Marketing site: Astro,
kidslen.app. - Infrastructure: a single Proxmox LXC at home, GitHub Actions runners on my own hardware (a Linux runner in the LXC, a macOS runner on a Mac mini for iOS builds), deploys triggered by push behind a smoke-test gate, everything exposed through a Cloudflare Tunnel with no inbound ports open.
- Observability: Prometheus, Grafana, Alertmanager.
That is the whole map. It fits in a paragraph, which is the point.
What I did not build
No Kafka, no message broker
The legacy system I reviewed had grown a streaming layer. Kidslen’s ingest path is a table in PostgreSQL.
Events arrive from the companion app, get inserted into a queue table as JSONB with a status column, and a processor picks them up. That is the entire queue. It gives me durability, transactional inserts alongside the rest of my data, SELECT ... FOR UPDATE SKIP LOCKED for concurrent workers, and the ability to inspect the backlog with a query instead of a CLI tool I have to remember the flags for.
A broker earns its keep when you have multiple independent consumers, cross-service fan-out, replay across teams, or throughput a database cannot absorb. I have one consumer, one service, and an event volume that a laptop would not notice. Adding Kafka would have added a second stateful system to operate, monitor, back up, upgrade and reason about — in exchange for solving a problem I do not have.
No Parquet/S3 batch pipeline
This is the direct lesson from Part 2. The legacy platform’s nightly export existed because reporting queries against the main store were too slow, and they were too slow because of a missing index.
Kidslen’s reports — watch time by platform, by day, by child — are SQL queries against the live tables, with indexes designed alongside the queries in the same migration. A family’s year of viewing history is a small number of rows by database standards. There is no analytics store, no nightly job, no reconciliation script, and therefore no possibility of the cleanup-versus-export race that was my favourite finding of the review.
Removing a subsystem removes every bug that subsystem could ever have had. That is the cheapest reliability work available.
No microservices
One deployable backend. Auth, ingest, processing, the policy engine, alerts, export and the admin surface all live in the same codebase, behind module boundaries rather than network boundaries.
Splitting them would buy independent scaling I do not need and independent deployment I do not want, and would cost me distributed transactions, cross-service tracing, version skew between services, and five more things to deploy at 11pm. The boundary I care about — “this code path must not see another family’s data” — is enforced by deny-by-default authorisation and query scoping, not by a network hop. A network hop was never an authorisation mechanism anyway.
No self-built auth, no self-built crypto
Rotating refresh tokens with reuse detection, Argon2id password hashing, rate limiting on auth endpoints. Standard, boring, from libraries. The interesting part of this product is not the login form.
No Kubernetes
One container and one managed database. Deploys are a push that triggers my own build runner, which builds, runs tests, deploys, and runs a smoke test before the deploy counts as done. Rollback is redeploying the previous image. I can explain the entire production environment to another engineer in five minutes, which is the actual metric I care about when I am the only person on call.
The mechanism that does the work
What I did build is one pipeline, and it is worth walking through because everything the product does is a step in it.
companion app
│ batched events (POST, idempotency key per event)
▼
queue table ──────────► processor ──────────► policy engine
(JSONB payload, (dedup, retry, (limits, bedtime,
status, attempts) dead-letter) blocks, keywords)
│ │
▼ ▼
viewing records alerts
│ │
└─────────┬─────────────┘
▼
push + SSE ──► parent dashboard
1. Ingest. The companion app batches events and posts them. Each event carries a client-generated key. The endpoint’s only job is to validate, authorise, and insert into the queue table as JSONB. It does no interpretation, which means a schema change on a platform cannot break ingestion — the raw payload lands safely even if I do not fully understand it yet. That has already saved me once.
2. Dedup. The event key is unique. A retried upload — flaky mobile network, app killed mid-request, backend restart — inserts nothing the second time. This is the fix for one of the legacy findings: retries without idempotency mean a retry silently doubles someone’s watch time, and in a product where a number crossing a threshold sends an alert, double-counting is not cosmetic. It sends a false accusation.
3. Process. A worker claims rows with SKIP LOCKED, normalises the payload into viewing records, and marks them done. Failures increment an attempt counter and back off; after a threshold the row goes to a dead-letter state and stops retrying forever. Dead-letter depth is a metric with an Alertmanager rule on it, because the failure mode that scares me is not “processing crashed loudly”, it is “processing has been quietly discarding one platform’s events for nine days”.
4. Policy engine. Each new viewing record is evaluated against that child’s active policies: daily screen-time limit, bedtime window, blocked platform, keyword alert. This is deliberately a small, pure, heavily tested function — given records and policies, produce violations. It is the piece most worth testing, because it is the piece that generates accusations.
5. Alerts. A violation produces an alert, delivered as a push notification and over SSE to any open dashboard. The child’s app is notified too. In a transparent product, an alert about you is also an alert to you.
6. SSE, not WebSockets. The dashboard needs server-to-client updates and nothing else; the parent’s actions are ordinary HTTP requests. Server-sent events are one-directional, ride on plain HTTP, reconnect automatically in the browser, and pass through the Cloudflare Tunnel without special handling. A WebSocket would have given me bidirectionality I have no use for, plus its own connection lifecycle to get wrong.
7. Retention. One job, the only thing in the system permitted to delete viewing data, operating on explicit age thresholds and coupled to no other job’s progress. Plus an on-demand full export and an on-demand full delete, because the transparency promise from Part 1 is only real if data can actually leave.
Why PostgreSQL with JSONB
The queue payloads are JSONB; the normalised viewing records are ordinary relational columns.
That split is intentional. The raw event shape is dictated by six third-party platforms that change without telling me, so it must be schemaless at the edge. The data I query, report on and enforce policies against is mine, stable, and relational. JSONB lets me accept messy input without letting mess into the core model, and I can still query it with an index when I need to investigate.
One database also means one backup, one restore procedure, one place where “everything about this family” lives. When a parent asks me to delete their data, I want that to be a transaction, not a distributed cleanup across four systems that might half-succeed.
Why family scale changes every decision
The trap in this kind of product is designing for the scale you hope for rather than the scale you have. Designing for imaginary scale is not free — you pay for it immediately, in operational surface and in every future change — and you pay in the currency of a one-person team.
At family scale:
- A year of viewing history is small, so reporting can be live queries.
- Ingest is bursty but tiny, so the queue can be a table.
- There is one consumer, so there is no broker.
- There is one deployable, so there is no service mesh.
- There is one operator, so simplicity is a safety feature, not an aesthetic.
If the shape of the load genuinely changes, every one of these decisions is reversible: the queue table becomes a broker, the reports become materialised views, the monolith splits along the module boundaries that already exist. What is not reversible is the complexity you adopt on day one and then have to carry while you are also trying to ship.
The legacy system I reviewed did not fail because it chose the wrong tools. It accumulated four layers of machinery around one missing index, and then could not remove any of them. I would rather add a layer later, with evidence, than spend years explaining one I never needed.
Next in the series: the core capture mechanism — WebView plus injected JavaScript, why it is inherently fragile, and the one mitigation that makes it survivable.
The product is at kidslen.app.