Kidslen is four applications: a Kotlin/Spring Boot backend, a React 19 parent dashboard, a React 19 ops console, and a React Native companion app. Plus a marketing site. One person, evenings and weekends.

They were built in parallel by multiple AI agents, in days rather than months. That sentence is the part people want to hear about, so let me get the boring truth out of the way first: the agents were not the hard part. The contract was.

The bottleneck was never typing

If you have used agents seriously, you know the failure mode. One agent builds a login endpoint that returns { accessToken, refreshToken }. Another builds a login screen that expects { token, refresh }. Both are correct. Both have passing tests. Neither works with the other, and you find out at integration time, which is the most expensive possible moment.

Running four agents in parallel multiplies that problem by more than four, because every pair of components has an interface to disagree about. Speed on each component does not help if the components do not meet.

So before any agent wrote code, I wrote the API contract.

The contract document

One document. It was the single most valuable artefact in the project, and it existed before the first line of implementation. It specified:

  • Every endpoint: method, path, auth requirement, request schema, response schema, status codes.
  • Error format: RFC 7807 problem details, with the type URIs enumerated. Not “return an error” — the exact shape, for every failure case.
  • Auth semantics: rotating refresh tokens, what reuse detection does, what happens on a refresh with a rotated token, token lifetimes.
  • Pagination: the exact parameters, the exact envelope, the enforced maximum page size.
  • The event ingest schema: the batched upload shape, the idempotency key, what the server guarantees about duplicates.
  • The SSE stream: event names, payload shapes, reconnection and replay semantics.
  • The policy model: what a policy is, the four types, how a violation is represented.
  • Naming and date conventions: camelCase, UTC ISO-8601 timestamps, IDs as opaque strings. Boring, and it eliminates an entire class of disagreement.

Then I handed the relevant slice of that document to each agent, with a hard instruction: the contract is the source of truth; if something is ambiguous, stop and ask, do not decide.

That last clause did more work than any prompt-engineering trick. The default agent behaviour when a spec is unclear is to make a reasonable choice and continue confidently, and two agents making reasonable choices independently is exactly how you get accessToken versus token.

Why parallelism actually worked

With the contract in place, the four workstreams became genuinely independent:

  • The backend agent implemented endpoints to match the contract, with tests asserting the contract’s shapes.
  • The parent dashboard agent built against the contract’s responses, mocked, before the backend existed.
  • The ops console agent did the same for the admin surface.
  • The companion app agent built the enrollment flow, the transparency screen, the capture harness and the upload queue against the ingest contract.

Integration day was not the disaster it usually is. Not because the agents were brilliant, but because the interface had been decided in advance by a human, and disagreements about it were contract bugs — fixable in one place, with all four sides updated from the same edit.

The thing I would tell anyone attempting this: the contract document is the parallelism. The agents are just throughput. If you skip the document, you have not parallelised anything; you have created four incompatible codebases very quickly.

What the agents got right

Credit where it is due. They were genuinely good at:

  • Breadth of boilerplate. Controllers, DTOs, mappers, migrations, form validation, table components, loading and empty states. The unglamorous volume that makes an app feel finished, done consistently across four codebases.
  • Following an explicit convention once it is written down. Every endpoint returning RFC 7807 errors, every list paginated the same way, every timestamp UTC. Consistency is where agents outperform a tired human at midnight.
  • Test scaffolding. Given a specified behaviour, generating the test cases around it — including edge cases I had not listed, particularly around pagination boundaries and empty collections.
  • Cross-language translation. “Here is the contract’s TypeScript type; produce the Kotlin data class” is mechanical and error-prone for humans and near-free for an agent.
  • UI iteration speed. Restructuring a dashboard layout is minutes, so I could try three versions instead of defending the first one.

What I had to verify myself

Here is the part that matters, and the part I want on the record.

I re-ran every test suite myself. I did not trust a single report that said tests passed.

This is not paranoia about dishonesty. It is that “the tests pass” is a claim with several failure modes that all look identical in a summary: the suite passed but half of it was skipped; the test asserts the implementation rather than the behaviour; the test mocks the exact thing it claims to verify; the suite was never actually executed in the final state of the code. A report is a sentence. A green run in my own terminal is evidence. Part 2 of this series was, after all, about a system where the tests existed and nobody had noticed they did not run.

So the posture across the project, all of it verified by me running it:

ComponentTests
Backend (Kotlin / Spring Boot)57 tests, roughly 96% line coverage
Parent dashboard81 unit, 6 end-to-end
Ops console81 unit, 7 end-to-end
Companion app113 tests
Marketing siteUnit tests plus Playwright

Beyond re-running, the things I reviewed line by line rather than skimming:

Security boundaries. Every authorisation check. Agents are good at writing an endpoint and less reliably careful about whether the requesting user is entitled to this specific family’s rows. Deny-by-default is only real if every query is actually scoped, and that is a per-endpoint human read. This is the single place I trusted nothing.

The policy engine. It generates accusations. A false violation is a parent confronting their child over my bug. I read that logic personally and wrote extra tests around the boundaries — midnight rollovers, timezone edges, a bedtime window that crosses into the next day, a limit reached exactly.

Data deletion and retention. The promise from Part 1 — full export, full delete — has to be true. I verified deletion myself against a real database, checking that it removed what it claimed across every table, rather than trusting a test that asserted a function had been called.

Anything touching money or secrets. Configuration, token handling, what gets logged. I grepped for secrets in the history myself, because the first finding of the due-diligence review in Part 2 was credentials committed to a repo, and it would be humiliating to repeat it.

The concurrency in the pipeline. SKIP LOCKED claiming, retry counters, the dead-letter threshold. Concurrency bugs pass tests and fail in production; I reasoned through the paths on paper.

What this actually felt like

Less like managing junior engineers and more like being a technical lead on a team that works at very high speed, never gets tired, has read every library’s documentation, and has no judgement about what matters.

The judgement is the job. Which of these sixty findings become rules. Whether to add Kafka. Whether a partial capture failure should render as an empty chart or as a warning. Whether the child gets a hidden mode. No agent produced any of those decisions, and none of them are code.

The honest summary of the speedup: the agents compressed the implementation, which was maybe half the work. They did not compress the design, the contract, the verification, or the decisions. And they created a new kind of work that did not exist before — reviewing a large volume of plausible code, which is a genuinely different and more tiring skill than reviewing a small volume of code from someone whose failure modes you know.

The rule I ended up with

Agents write the code. I own the claims.

Every number in this series is one I measured. Every test count above is from a run I executed. The product ships on trust — it holds children’s viewing data and it tells parents what their kids did — and the day I forward a claim I have not verified is the day the product’s premise is a lie.

That is the end of this five-part series on building Kidslen: the idea, the due-diligence review that became the spec, the architecture and its omissions, the capture mechanism, and how it was built. More build-log posts to come as things break, because they will.

The product is at kidslen.app.

Export for reading

Comments