Every build-in-public series has a week it does not want to write about. This is mine. Three failures in five days, none of which my test suite caught, all of which were in front of a user — even if that user was me.

I am writing them up in detail because a product that asks families to trust it with a child’s activity data has no business publishing only its wins. Also because each bug taught a lesson that generalises past Kidslen, and two of them are lessons I had read before and did not believe until they cost me a day.

Bug one: the events that only went missing on the slow machine

Kidslen’s parent dashboard shows a live activity feed. When a child’s device reports that a video started, the parent sees it appear without refreshing. That is Server-Sent Events: the browser holds an open HTTP connection, the server writes events into it as they happen.

On the Spring side, each connected dashboard gets an SseEmitter. Two things write to it:

  • the event processor, when a newly ingested activity event passes the policy engine and is ready to display;
  • a scheduled heartbeat, every fifteen seconds, which writes a comment frame to keep proxies from closing an idle connection.

On my development machine, this worked flawlessly for weeks. On the CI runner in the LXC — a slower, busier machine — events started going missing. Not all of them. Not reproducibly. The end-to-end spec would pass four times and fail the fifth with an activity item that never arrived.

The kind of failure that makes you blame the test.

I did blame the test, for about half a day. I added waits. I increased the timeout. The flake rate went down and did not go to zero, which is the signature of a race, not a timing problem — if waiting longer helps but never fixes it, you are not waiting for something slow, you are losing a coin toss.

Then I read the documentation properly. SseEmitter is not thread-safe. That is stated plainly, and I had built a system where the single most concurrent path in the application wrote to one from two independent threads: the processor thread and the scheduler thread.

Two threads calling send() on the same emitter can interleave. What gets written down the socket is not “event A then event B”; it can be fragments of both, and the client-side parser, encountering something that is not a well-formed event frame, does the only reasonable thing: it discards it. The events did not fail to send. They were sent as garbage and thrown away by the browser.

The fix is unglamorous: serialise all writes per emitter. Every emitter is wrapped in a small holder with its own lock, and every path that sends — processor, heartbeat, close — goes through it. Sends to different emitters remain fully parallel, so the fix costs nothing at any realistic connection count; the lock is only ever contended by the two writers of one connection.

Two things fell out of this that I now treat as rules:

Wrap, do not remember. The bug was not that I did not know emitters needed care. It was that “remember to synchronise” was a rule living in my head, and my head was not present when I added the heartbeat scheduler three weeks after the processor. The fix that works is the one where the unsafe object is not reachable — the raw emitter is private, and the only thing any other class can obtain is the synchronised holder.

A slower machine is a better test than a faster one. My laptop was fast enough to hide the race almost every time. The CI runner, competing for CPU with a Gradle build, widened the window enough to expose it. I had been treating the CI machine’s flakiness as a defect of the CI machine. It was the most valuable piece of test hardware I had. I now deliberately run the end-to-end suite with the runner under load before a release, and I stopped adding sleeps as a first response to intermittent failure.

Bug two: the native build that would not build

The same week, the companion app stopped compiling. Not the JavaScript — React Native’s JavaScript layer is the easy half. The native half.

Three separate problems, all the same underlying shape.

A dependency still asking for a repository that no longer exists. One transitive library declared jcenter() in its Gradle configuration. JCenter has been retired; under Gradle 9 the build does not degrade gracefully, it fails to resolve. The library itself was fine — it was published on Maven Central too — but its build script pointed at a dead address.

The fix was patch-package: a checked-in patch rewriting jcenter() to mavenCentral() in that dependency’s build file, applied automatically after install. Patching a dependency feels wrong the first few times. It is far better than the alternatives, which are forking the library, or pinning Gradle to a version that is itself on the way out.

An unmaintained library that would not compile against a current Xcode. I had been using a certificate-pinning library for React Native. It had not seen a release in a long time, it did not build under Android Gradle Plugin 8, and it did not build against a current Xcode either. I spent most of a day trying to patch my way through, and then stopped and asked the question I should have asked first: what does this dependency actually give me that I cannot get underneath it?

The answer was: nothing. Certificate pinning is a platform capability. Android has network security configuration; iOS has App Transport Security and its own pinning affordances. The library was a convenience wrapper over things the operating systems already do, and it was a wrapper that had stopped being maintained while the platforms kept moving.

So I removed it and did pinning at the native layer on both platforms. The dependency count went down, the build went green, and the security property is identical — arguably stronger, since it is now enforced by the platform rather than by JavaScript that could be bypassed if the bridge were compromised.

Firebase refusing Swift Package Manager with static linkage. The combination of Firebase’s SPM distribution and the static linkage the rest of the project needed did not work together. This one has no clever lesson. I moved that dependency to the other supported integration path and moved on. Some days the fix is to stop being right about the dependency manager.

The general lesson from all three: every native dependency in a React Native project is a bet that its maintainer will keep up with two platform vendors. Android Gradle Plugin and Xcode both move on their own schedule and both make breaking changes. A library with no release in eighteen months is not stable, it is a scheduled outage. Before adding one now, I ask whether the platform already provides the capability — and for pinning, telemetry, storage and notifications, it very often does.

Bug three: the HTTP 400 that only a demo video found

The third one is the one that actually embarrassed me.

I was recording a product demo — screen capture of the real dashboard, doing the real things a parent would do. Register, add a child, enroll a device, watch activity appear, open the weekly report.

The weekly report screen showed an error. HTTP 400.

Everything was green. Backend tests: green. Dashboard unit tests: green. Playwright end-to-end: green. Smoke test gate on the last deploy: green. And the activity screens, when a human clicked through them in sequence, returned 400.

The cause was a date format mismatch. The dashboard’s date-range picker serialised its value in one format; the API’s query parameter binding expected another. The API rejected it, correctly, with a 400 and a well-formed RFC 7807 problem document that nobody was reading because the screen just showed “something went wrong”.

Why did nothing catch it?

  • The backend tests sent the format the backend expected. They were testing the parser, and the parser was correct.
  • The dashboard unit tests mocked the API. They asserted the component rendered the rows it was given. It was given rows.
  • The end-to-end tests used a fixed date range set up by the fixture — which happened to be produced in the format the API wanted, because the fixture was written by the same person on the same day as the API.

Three test layers, all green, all testing around the seam rather than across it. The bug lived exactly in the gap between two things that each had excellent coverage. This is the most common shape of production bug I have ever met, and I walked straight into it in my own product.

The fix was five minutes: one shared formatting helper, used by every caller, with a test that asserts the emitted string matches what the API’s binding accepts. The interesting part is not the fix, it is what it changed about how I verify.

Your test suite is not your product. Watch the real thing.

I now do a recorded walkthrough of the full parent journey before every release that touches the dashboard. Not a test run — a screen recording, watched back at normal speed, with sound. It is slow and it feels unproductive and it has caught things no assertion caught: a spinner that never ends, a Vietnamese string overflowing a button, a consent banner that briefly flickered out of view during a re-render, which on this product is not a cosmetic bug at all.

There is a reason this works. Automated tests assert things you already thought of. Watching the product shows you things you did not. The demo video was not a marketing task that happened to find a bug; it was the best QA session of the month, and it found a bug that every other layer of verification was structurally incapable of finding.

What I changed permanently

FailurePermanent change
Concurrent writes to one SseEmitterUnsafe object made unreachable behind a synchronised holder; end-to-end suite run against a loaded machine before release
Dead repository in a transitive dependencyChecked-in patch-package patch; dependency-age review before adding native libraries
Unmaintained pinning libraryRemoved; pinning moved to the platform layer on both OSes
Date format across the dashboard/API seamOne shared helper, one contract test asserting the exact serialised form
All threeA recorded human walkthrough of the real product before any release touching the dashboard

None of these are clever. Every one of them is the thing I would have told a junior engineer to do, and did not do myself until it cost me a week.

If there is a single thread running through all three, it is this: the parts of a system that break are the parts where two correct things meet. A correct processor and a correct scheduler, sharing an emitter. A correct library and a correct build tool, disagreeing about a repository. A correct dashboard and a correct API, disagreeing about a date. Coverage inside each component tells you nothing about the seams. Only running the real thing does.


Next in this series: privacy as the product — consent records, the banner a child cannot dismiss, retention that actually deletes, and what app-store review asks of software that monitors minors.

Kidslen is at kidslen.app, bugs, fixes and all.

Export for reading

Comments