On 18 September I read thirty competitors and decided the one thing worth building in Hive was a Routine: one recurring job, a producer on a cheap model, a checker on a strong one, arbitration only when they disagree, a budget that pauses itself. Writing that was easy. The test was whether I would actually run my own morning through it.

So on the 20th I did. Not a demo account. My real AWS account, my real GitHub PRs, on the one LXC container (CT137) that runs Hive: standard-library Python, SQLite, a Preact frontend with no build step. This is the log of that morning — four jobs that ran, one thing that broke, the fix.

08:30 — the ops brief

The first Routine is a room I named “Reveal · Ops sáng” — Reveal is the product I ship, ops sáng is morning ops. It runs at 08:30 with three seats.

The producer is on grok-4.6, cheap and fast, with exactly two tools: aws_status and daily_brief. It writes one file, outputs/OPS.md. The checker doesn’t read that file and nod. It calls the same two tools itself and checks every number against what they return. The lead only speaks if producer and checker disagree. That is the conditional-arbitration shape the debate papers recommend: no discussion unless there is a dispute.

aws_status is read-only: CloudWatch alarms, the RDS and EC2 inventory, month-to-date cost by service — including the Amazon Bedrock lines, because that’s where AI spend hides in an AWS bill — and the Budgets API. daily_brief covers GitHub across my two accounts: PRs waiting for my review, team review requests, my own open PRs, assigned issues.

What the brief said this morning, and what the checker confirmed:

  • AWS month-to-date: $1,521. RDS $804, EC2 $567, EC2-Other $119. Bedrock lines total about $0.15. The checker’s verdict on the whole file was one line: “all numbers and alerts match aws_status + daily_brief exactly; files verified.” The run cost $0.13.
  • 8 EC2 instances, 1 Aurora PostgreSQL instance.
  • 4 CloudWatch warnings. Two are RDS alarms stuck in INSUFFICIENT_DATA, pointing at a DB identifier that no longer exists. I had scrolled past those in the console for weeks. The snapshot flags it rds_target_missing; the producer listed it and the checker confirmed the alert list matched. That single line paid for the checker seat.
  • GitHub: 1 PR waiting for my review (a Dependabot flask bump). 12 open PRs of my own across hivedashboard and ampd-reveal. Twelve is too many. The brief will say so every morning until it isn’t.

The brief also reports on Hive itself. AI spend this month is about $134 — but $125 of that is a worst-case estimate for 25 legacy runs from before per-run pricing. I price unpriced runs at the ceiling, not zero, so the number errs high on purpose. Caps unchanged: $20 per day for the workspace, $5 per run.

The planning room

The second job has no schedule. “Reveal · Kế hoạch tính năng” (feature planning) is a standing room with four seats: a Product Lead who leads, a Tech Lead and a Cursor Engineer as workers, and Claude Opus via the Brain as the checker.

I type /plan. The lead plans, the workers run in parallel waves, the checker reviews, and the room writes outputs/REPORT.md: one row per feature, feature · value · effort · risk · owner · week. A table, not an essay. The point is that I read a table over coffee instead of running three chat sessions and merging them in my head.

This morning’s run: the Tech Lead assessed 8 features (12–18 person-days), the Cursor Engineer marked 6 of 8 as automatable by a Cursor Cloud agent, and Claude’s critique cut the scope to 5 (about 13 person-days including review), deferring the Lambda and multi-model work. The room cost $0.59.

The PR that waits for me

The third job is new today: a scheduled Automation of kind “Origin task.” Origin is Hive’s autonomous coding agent — clone, edit, test, open a PR. Until this morning I launched it by hand. Now it runs on a schedule.

Weekly, on ampd-reveal: check whether pipecat-ai and google-genai have newer releases, bump if so, run the tests, open a PR. The configuration is three header lines at the top of the prompt:

repo: https://github.com/luonghongthuan/ampd-reveal
base: main
approval: push

approval: push is the whole idea. The run does everything up to the push and stops. Nothing lands without my click. I didn’t build a new gate; Origin already had plan and push approvals. The scheduled task just defaults to the push gate. A dependency PR opened on a Sunday, waiting for Monday’s click, is exactly the amount of autonomy I’m comfortable with.

Budgets at four heights

The least visible piece, and the one I’d sell first. Every layer this morning has a ceiling: the Routine has a monthly budget, the room has a budget, each agent is capped at $5, the workspace at $20 a day. Hitting any of them pauses the thing that hit it and makes no model call. The ledger reserves before a call and settles after, so the pause happens before the spend, not after the invoice. Paperclip’s pause at 100 %, at four heights instead of one.

What broke

Now the honest part.

Yesterday my chief-of-staff rooms started failing with “Model error 502.” Nothing pointed at a room; the rooms looked fine. The root cause was a fallback. When a seat had no model set, the code took the first model in the pool list. That happened to be xp-claude-opus-5, an alias on a prepaid gateway whose upstream now answers upstream_unsupported. Every room that relied on the default was quietly routing to a model that didn’t exist any more; the 502 was the gateway saying so.

Three moves fixed it:

  1. One default_model() function that reads HIVE_DEFAULT_MODEL, then the pool list. It replaced every hardcoded gpt-5.5 fallback in flows.py, rooms.py, evals.py and routes.py. There were more than I’d like to admit.
  2. Pool list reordered, grok-4.6 first. The default should be the cheap model that works, not the expensive one that might.
  3. Dead aliases removed: the xp-* family, gpt-5.5, gemini-2.5-pro (removed upstream), grok-3-mini-fast (out of credits).

The audit turned up more than the 502:

  • gpt-5.5 and cdx-gpt-5.5 via the gateway return an HTML login page instead of JSON — the codex route needs a re-login. Hive now routes around it rather than parsing a login form as a completion.
  • deepseek-chat: 402, insufficient balance.
  • Linear MCP token expired; needs a reconnect in Apps.
  • Google Calendar not connected yet. One OAuth click. Not clicked.
  • Telegram delivery of the brief needs my chat id on my profile. Until then the brief is a file in the room — fine, but not a notification.

None of these are agent bugs. They are the boring failure modes of real integrations on real accounts: a token expires, a balance runs out, a vendor removes a model, a route needs a login. What bothers me is that the agents didn’t fail loudly on any of them.

The lesson

Agents die quietly when a model alias dies. A room doesn’t crash. It returns a 502 and moves on, and if I’m not reading the room it looks like nothing happened. The fix isn’t better error text. It’s structural: the pool list is the single source of truth for which models exist, and every fallback reads it. A hardcoded model name anywhere in the code is a future outage with a date I don’t know yet.

It’s also why the checker earns its seat. The producer on the ops brief could have been routed to a dead alias too. If it had, the checker — on a different model, re-calling the tools itself — would have refused to sign off an empty OPS.md, and the lead would have been woken. Two models on two routes is cheap insurance against one route going dark.

Where this leaves me

On the 18th I said I wouldn’t build a canvas or a fourth orchestrator, and that the thing to sell is one recurring, verified job: cheap producer, strong checker, arbitration only on dispute, a budget that stops itself. Two days of running my own morning through it hasn’t changed my mind. It has made the pitch concrete enough to say in one breath:

Every morning at 08:30 it reads my AWS and my GitHub, writes a one-page brief, has a second model verify every number, and tells me about the two alarms pointing at a database that doesn’t exist. Once a week it bumps my dependencies and opens a PR that waits for my click. It cannot spend more than I told it to.

That’s the job. Everything I fixed today was so it runs tomorrow while I’m not watching.

Export for reading

Comments