Before I wrote a line of Kidslen, I spent several weeks doing a technical due-diligence review of a legacy OTT media-analytics platform. Different domain, adjacent shape: a system that ingests enormous volumes of viewing events from set-top boxes and apps, stores them, and reports on them.
I cannot name the system, the vendor, its clients, or show you its code. That is fine, because the interesting part was never the code. The interesting part is that by the end of the review I had a catalog of roughly sixty concrete problems, and when I sat down to design a product of my own, I found I was not starting from a blank page. I was starting from a list of things not to do.
This is the most useful artefact I have ever gotten out of a code review, and it is worth explaining how it happened.
What a due-diligence review actually produces
A due-diligence review is not a code review. Nobody wants your opinion on naming. The question you are answering is: if we own this system tomorrow, what will hurt, how much, and how soon?
So you read differently. You go looking for the things that cost money later:
- What happens on a bad deploy?
- What happens when the biggest table gets twice as big?
- Who can read production data, and is there a record of who did?
- If the person who built this left, what stops working?
- What is the oldest test, and does it still run?
I logged every finding with a severity and a one-line consequence. Sixty-ish findings, ranging from cosmetic to genuinely alarming. Below are the ones that ended up shaping Kidslen. I have kept them abstract enough that they could describe a thousand systems, because honestly, they do.
The findings that mattered
1. Credentials committed to the repository
Real, working credentials in tracked files. Not in a .env.example. In the repo, in the history, readable by anyone who ever had a clone — including former contractors.
Rotating a secret does not fix this; git history is forever, and nobody ever rewrites it because nobody wants to be the person who broke everyone’s branches. The consequence line I wrote was: “any past or present clone holder has production access.”
2. Zero database indexes on the largest table
The main events collection was taking roughly 31 million rows per day. It had no indexes beyond the primary key.
It worked, in the sense that ingestion worked. Reporting did not really work; queries that should be milliseconds were minutes, and the fix that had evolved was to pre-compute batch exports overnight instead of making the database fast. An architectural layer had grown specifically to route around a missing CREATE INDEX.
3. A cleanup job racing an export job
This one is my favourite finding of my career so far.
There was a retention job that deleted raw events older than a threshold, and a separate export job that read those raw events and wrote them out. They ran on independent schedules with no coordination and no shared notion of “this range has been safely exported”.
Under normal timing, nothing happened. Under a slow export — a big day, a retry, a restart — the cleanup job could delete data that the export had not yet written. Silent, permanent, invisible in every dashboard, because both jobs reported success. The data was just quietly less complete than anyone believed.
Nobody had noticed. That is the part that stays with you: the failure mode was indistinguishable from working correctly.
4. Manual SSH deploys
Deployment was a person with SSH access, a sequence of commands, and institutional memory. No pipeline, no record of what version was running, no repeatable rollback. “What is in production?” was a question you answered by logging in and looking.
5. Essentially no tests
A handful of tests existed. Most did not run. There was no CI to notice that they did not run. Every change was validated by deploying it and watching.
6. The rest
The remainder of the sixty were variations on the same themes: no audit trail of who accessed what; error responses that were HTML stack traces; unbounded queries with no pagination; one giant service where everything could reach everything; logs with personal data in them; retries without idempotency, so a retry meant a duplicate; monitoring that watched whether the process was alive but not whether it was doing anything.
None of it was incompetence. All of it was a system that grew for years under delivery pressure with nobody’s job being “say no”.
Turning findings into rules
Here is the mapping I actually used. Left column: what I found. Right column: the rule I wrote into Kidslen’s design before implementing anything.
| Legacy problem | Kidslen decision |
|---|---|
| Credentials committed to the repo | No secret ever lives in the repository. Secrets are injected as environment variables at runtime; CI has a secret-scanning step; a build fails rather than ships a key |
| No indexes on a 31M-rows/day collection | Every query the product actually runs has a supporting index, written in the same migration as the table. Schema changes are versioned migrations, never hand-applied |
| Batch export layer built to dodge slow queries | No Parquet/S3 batch pipeline at all. Query the primary database directly; at family scale it is fast enough, and “fast enough” removes an entire subsystem |
| Cleanup job deleting unexported data | The retention job is the only thing that deletes, it operates on explicit age thresholds, and deletion is never coupled to another job’s progress. Export reads a consistent snapshot and never removes anything |
| Manual SSH deploys | Push-triggered deploys through GitHub Actions runners on my own hardware, with a smoke-test gate. If the smoke test fails the deploy does not complete. Nobody types deployment commands by hand |
| No tests, no CI | Tests gate every merge. Backend 57 tests at roughly 96% line coverage; parent dashboard 81 unit and 6 end-to-end; ops console 81 unit and 7 end-to-end; companion app 113 tests; marketing site unit plus Playwright |
| No audit trail | An append-only audit log records privileged actions — who viewed what, who changed a policy, who exported or deleted data. The audit log is a product feature, not just an ops tool |
| HTML stack traces as error responses | RFC 7807 problem-details responses everywhere, with a documented OpenAPI contract. Errors are typed, machine-readable, and leak nothing |
| Unbounded queries | Every list endpoint is paginated with an enforced maximum page size. There is no endpoint that can return “everything” |
| Everything can reach everything | Deny-by-default security: an endpoint is unreachable unless it is explicitly authorised, and every query is scoped to the requesting family |
| Personal data in logs | Structured logs carry identifiers, never content. Titles and viewing detail live in the database under retention, not in a log file with a different lifecycle |
| Retries without idempotency | The event pipeline deduplicates on a client-supplied event key, with retry and a dead-letter queue. A retried upload cannot double-count watch time |
| Liveness-only monitoring | Prometheus metrics on the work itself — queue depth, processing lag, dead-letter count, policy evaluations — with Grafana dashboards and Alertmanager rules. “The process is up” is not a health signal |
| Long-lived, never-rotated sessions | Rotating refresh tokens with reuse detection; Argon2id for password hashing; rate limiting on authentication endpoints |
| Data kept forever by default | Explicit retention windows, a documented deletion path, and full export. Collecting a minor’s data obliges you to be able to delete it |
Sixty findings do not map one-to-one to sixty rules — many collapsed into the same principle. But every row above traces back to something I actually saw break, or was one slow night away from breaking.
The rule I did not expect to write
The single most valuable thing I took from that review was not on the list. It was this:
Most of the complexity was there to compensate for something simple that had not been done.
The batch pipeline existed because of a missing index. The overnight export existed because of the batch pipeline. The reconciliation scripts existed because of the export. The cleanup race existed because of the reconciliation scripts. One skipped CREATE INDEX had, over several years, grown a four-layer subsystem with a data-loss bug in it.
That observation is the reason Part 3 of this series is mostly about things I did not build. When I was tempted by Kafka, by microservices, by a separate analytics store, I kept asking the same question: am I adding this because the problem requires it, or because I skipped something simpler?
At family scale, the honest answer was almost always the second one.
Why this belongs in a product about children’s data
It would be easy to read the above as engineering hygiene. It is not, in this product.
Kidslen holds viewing records belonging to minors. The cleanup-versus-export race in the legacy system was, in its own domain, a reporting inaccuracy. The same class of bug here means a family’s data is silently wrong, or silently gone, or silently retained past the window I promised. Every one of those breaks the only thing the product actually sells.
The transparency rule from Part 1 — the child can see exactly what is collected — only means something if the system genuinely knows what it collected, kept it correctly, and can prove it deleted what it said it deleted. Indexes, audit logs, idempotency and retention jobs are not back-office concerns here. They are the ethical commitment, implemented.
Next in the series: the architecture, and the much longer list of things I deliberately left out.
The product is at kidslen.app if you want to see what these rules turned into.