Every part of Kidslen so far has been a design choice I am comfortable defending. This post is about the part that keeps me up at night.
To show a parent what their child actually watched, I need viewing history from six platforms — YouTube, Netflix, TikTok, Instagram, Disney+, Prime Video. None of them offer a “let a parental-monitoring product read my child’s watch history” API. So the mechanism is the one everybody in this category ends up at, whether or not they say so publicly: a WebView, the child’s own logged-in session, and injected JavaScript.
Let me explain how it works, and then be honest about what is wrong with it.
The consent frame first
Before any mechanism: the connection flow is explicit and two-sided. The child logs into their own account in a visible in-app browser, having been shown — on screen, before the browser opens — what will be read from that platform. The companion app carries a permanent “monitored by <parent>” banner and a transparency screen listing every connected platform and what is collected from it. Either side can disconnect a platform at any time, and disconnecting revokes the stored session.
This matters here more than anywhere else in the product, because what follows is a description of reading someone’s account data. The only thing that separates it from the spyware I complained about in Part 1 is that the account holder is looking at it while it happens, and can stop it.
What “connect an account” actually does
When a child connects YouTube, the app opens a WebView pointed at the platform’s normal web login. Nothing is proxied and nothing is faked — it is the platform’s real login page, the real 2FA, the real consent screens. The child types their password into the platform’s own page.
The app injects two scripts into that WebView:
- a common script, shared by all six platforms: the state machine, the message protocol, error reporting, timing helpers;
- a per-platform script: everything that knows what “logged in” looks like on this specific site, which endpoint holds viewing history, and where in the response the fields live.
The split matters. The common half is stable — I have changed it rarely. The per-platform half is the part that rots, and keeping it isolated means a break in one platform is a small, contained edit rather than a surgery on shared code.
The state machine lives in the page
The connection flow is a state machine with roughly these states:
INIT → AWAITING_LOGIN → LOGGED_IN → PROBING → HARVESTING → COMPLETE
plus FAILED and NEEDS_REAUTH.
The non-obvious decision: the state is persisted in the page’s own localStorage, not in native memory.
The reason is navigation. A login flow is not one page. It is a login page, a password page, an interstitial, a 2FA challenge, a consent screen, a redirect, and finally the app itself. Every one of those is a fresh document, and my injected script is executed from scratch each time with no memory of what came before. Keeping the state in native and pushing it down on every navigation means a race on every single page load. Keeping it in localStorage — same origin as the platform, so it survives every in-site navigation — means the script’s first action on any page is to read where it already was.
INIT is therefore not “start the flow”. It is “find out whether this flow is already in progress, and if so, resume it”.
The state record is small: current state, attempt counters, a session id that ties it back to the native side, and timestamps for timeouts. No credentials, ever. Nothing sensitive is stored in the page.
The message protocol
Communication with native is a JSON message over the WebView bridge, one shape for everything:
{
"v": 1,
"sessionId": "…",
"type": "STATE_CHANGED",
"platform": "youtube",
"payload": { "from": "AWAITING_LOGIN", "to": "LOGGED_IN" },
"ts": 1757400000000
}
Message types, roughly:
| Type | Direction | Meaning |
|---|---|---|
READY | page → native | Script injected and running on this document |
STATE_CHANGED | page → native | State machine transition, for UI and diagnostics |
NEEDS_USER_ACTION | page → native | 2FA or a challenge — native shows the WebView and stops any overlay |
SESSION_CAPTURED | page → native | Login succeeded; session material is available |
DATA | page → native | A batch of normalised viewing items |
ERROR | page → native | Selector missed, endpoint shape changed, request failed |
CONFIG | native → page | Per-platform metadata: endpoints, key paths, limits |
ABORT | native → page | User cancelled, or a timeout fired — tear down cleanly |
Versioning the envelope (v) was worth it within the first month. Because the scripts are served from the backend (see below) and the app on the phone is whatever version the child last updated to, the two sides drift. The version field lets native reject a message shape it does not understand instead of misinterpreting it.
A typical successful sequence:
native: open WebView → inject common + platform script
page: READY (login page)
page: STATE_CHANGED INIT → AWAITING_LOGIN
… user types password …
page: NEEDS_USER_ACTION { reason: "2fa" }
… user completes 2FA …
page: STATE_CHANGED AWAITING_LOGIN → LOGGED_IN
page: SESSION_CAPTURED
native: CONFIG { endpoints, keyPaths, pageSize }
page: STATE_CHANGED LOGGED_IN → PROBING
page: STATE_CHANGED PROBING → HARVESTING
page: DATA { items: [ … ] } × n
page: STATE_CHANGED HARVESTING → COMPLETE
native: queue items for upload → close WebView
The cookie problem
Here is the part that is genuinely interesting.
The session cookies that matter are HttpOnly. That flag exists precisely so that page JavaScript cannot read them — it is the browser’s defence against a script stealing a session. My injected script is page JavaScript. document.cookie shows it nothing.
The resolution is that the injected script never needs to see them. The native layer owns the WebView’s cookie store, and on both platforms native code can read cookies the page cannot — that is a normal capability of an app that hosts a WebView. So when the page posts SESSION_CAPTURED, native reads the session material out of the WebView cookie store and stores it in the OS keychain/keystore on the device.
Two consequences I want to state plainly:
- Session material stays on the child’s device. It is used to make requests from that device. Getting the cookies off the phone would make me a very attractive target for very little product benefit.
- The scripts never handle credentials. They observe navigation and DOM state and they call endpoints that the browser authorises with cookies it attaches itself. Even a fully compromised capture script cannot exfiltrate a password, because it never touches one.
Then: call the platform’s own endpoints
Once the session exists, scraping the DOM is the wrong tool. Rendered pages are lazy, virtualised, infinite-scrolled and A/B tested; reading them is slow, noisy and breaks constantly.
Instead, the page-context script calls the same internal endpoints the platform’s own web app calls — the ones behind the history page. Same origin, same cookies, same headers the page would send, because the request is being made by the page. The response is JSON, paginated, complete.
The part I am mildly proud of is that the script does not hardcode where the fields are. It receives metadata-driven key paths in the CONFIG message:
{
"endpoint": "/…/history",
"keyPaths": {
"items": "contents.sectionList.items",
"title": "meta.title.text",
"id": "meta.videoId",
"watchedAt":"meta.watchedTimestamp",
"duration": "meta.lengthSeconds"
},
"pageSize": 50
}
The script walks those paths generically and emits a normalised item. When a platform reshuffles its response — and they do — the fix is usually a string in a config row, not a code change. When the shape changes more deeply, it is an edit to one platform script.
Being honest: this is fragile
It is worth saying plainly, because a post that presented this as robust engineering would be dishonest.
- Every platform can change its DOM, its endpoint, its response shape or its auth flow at any time, without notice, and none of them owe me anything.
- A selector that has worked for months can break on a Tuesday afternoon for a subset of users on an A/B test I cannot see.
- Anti-automation measures can escalate.
- Failures are often partial — one platform stops reporting while the other five are fine, which is exactly the silent failure mode I described fearing in Part 3.
I did not find a way around this. Nobody in this category has. What I did instead was design for the fact that it will break.
The mitigation that actually matters
The capture scripts are served from the backend, versioned, and fetched by the companion app at runtime. They are not compiled into the app binary.
That one decision changes the economics of every breakage:
| Scripts in the binary | Scripts served from backend | |
|---|---|---|
| Fix a broken selector | Code change, build, app-store review, wait, hope users update | Edit script, deploy, devices pick it up on next fetch |
| Time to recovery | Days at best, and only for users who update | Hours |
| Users still broken after the fix | Everyone who has not updated | None |
| Roll back a bad fix | Another release | Redeploy the previous version |
App-store review latency is not a detail here; it is the difference between a bug that lasts an afternoon and a bug that lasts two weeks for half your users. Shipping the volatile code through my own deploy pipeline instead of someone else’s release process is the single most important structural decision in the companion app.
Supporting pieces:
- Scripts are versioned and pinned per platform, so I can roll a fix out and roll it back.
ERRORmessages carry which key path or selector missed, so a breakage shows up in metrics and in Grafana as “YouTube capture error rate”, not as a silent gap in someone’s history.- A platform that fails degrades alone; the other five keep reporting.
- The parent dashboard shows a platform’s connection health honestly, including “this platform stopped reporting”, rather than showing an empty chart that implies the child watched nothing.
That last one is a product rule, not a technical one. In a monitoring product, “no data” and “no activity” must never render the same way. A silent capture failure that looks like a quiet evening is how you get a parent making an accusation based on my bug.
What I would tell someone building this
The mechanism is not the clever part — a WebView with injected scripts is the obvious answer once you look at the problem. The clever part is admitting the mechanism is unstable and building the whole system around fast recovery: isolate the volatile code, serve it from somewhere you control, version the protocol, instrument every failure, and make partial failure visible instead of invisible.
Next in the series: how the backend, parent dashboard, ops console and companion app were built in parallel by AI agents against one written API contract — and what I re-verified by hand rather than trusting.
The product is at kidslen.app.