I self-host an agent platform called Hive: standard-library Python, SQLite, a Preact frontend with no build step, one process on one Proxmox container. It already has rooms where several agents talk to each other, a debating council, cron, webhooks, budgets, evals. Honestly, it has a lot. The question I hadn’t answered was: if a five-person team wanted to buy it, what would they be buying?

On 18 September I spent a full day reading competitors to answer that. Around thirty products in six groups. Pricing pages, complaints on G2, Reddit and Hacker News, and two research papers on multi-agent debate. Then I re-read my own code — flows.py, rooms.py, llm.py, triggers.py — so I wouldn’t lie to myself about what Hive already had.

The result was tidier than I expected.

The conclusion first

The 2026 market does not pay for orchestration. The open-source core is free everywhere: n8n, Dify, CrewAI, LangGraph, Paperclip. Nobody pays to have “agents wired together” anymore. They pay for a repeated job that runs correctly, with someone approving it, without wrecking the invoice.

The four biggest complaints, repeated everywhere:

  1. Cost surprises. Zapier sits at 1.4 stars on Trustpilot, mostly over billing (Zapier Agents complaints). Lindy’s credits are opaque. Make charges 50 credits per agent run.
  2. Brittle agents. CrewAI has a whole thread on stopping infinite loops from eating tokens. n8n has a memory failure in the AI Agent node. LangGraph shipped a self-hosted flaw chain in June.
  3. Nobody checks the output. The quietest complaint and the most painful. Human-in-the-loop, where it exists, is a gate on a tool call (n8n 2.6’s HITL tools), not a second model reading the deliverable and saying whether the numbers are right.
  4. Self-hosting is hard. Dify and LibreChat are a stack of containers (a six-month Dify review). Open WebUI changed its license. OpenClaw exposed 21k instances to the internet.

Two more signals told me where the floor and the ceiling are. OpenAI is shutting AgentKit down on 30 November 2026 (source) — the closed orchestrator from the biggest vendor didn’t survive a year. And Paperclip went to 30k GitHub stars in three weeks on exactly one idea: agents organised like an org chart, each with a budget that pauses the agent at 100 %. People aren’t short of frameworks. They want to know where the bill stops.

The competitor table (condensed)

GroupWhoWhat to learnWhat not to copy
No-code canvasn8n, Dify, Flowise, Langflow, Make, Zapier AgentsHITL at tool level; webhook/cron/Slack are table stakes; Dify shows cost per runThe canvas. They’ve won; Hive composes work with Rooms, not nodes
Dev frameworksCrewAI, LangGraph, Agent Framework (AutoGen), Agno, MastraA lead that keeps a ledger and re-plans (Hive already has the board); suspend/resumeA fourth orchestrator
”AI employee” SaaSLindy, Relevance, GumloopPricing by model tier; manager agentOpaque credits, always on
Control planesPaperclip, Block Buzz, OdysseusPer-agent budgets that pause; agents as an org chartRuntime-agnostic everything
Closed vendorsClaude Cowork / routines / Managed Agents, OpenAI AgentKit, Gemini EnterpriseRoutine = recurring job with a deliverable; 20–30 USD/seat is the tolerance ceilingDepending on one vendor
Local chatOpen WebUI, AnythingLLM, LibreChat, OpenClawLocal-first is coming backChannel sprawl, no verification

The “not to copy” column turned out to matter as much as the “learn” column. It’s the wall between me and the thing I’m best at: adding systems.

The paper that changed my mind

I used to think of the council — several agents debating over several rounds — as Hive’s “premium” mode. Two papers say the opposite. The Cost of Consensus (arXiv, May 2026) and the earlier problem drift paper both find that free debate between homogeneous models performs worse than isolated self-correction, because the agents drift off the problem together. The winning pattern is a verifier triggered on condition: one writer, one checker, and a heavier session only when the checker objects.

Hive has both shapes. Rooms run /goal (producer, then checker, two rounds) and a full four-phase council at 2 + 2N calls for N seats, up to six seats. The research says: make the first one the default, and stop treating the second as the grown-up option.

What Hive has, what it lacks

Reading the code honestly:

  • A checker already in /goal; /review runs two parallel reviewers and mediates.
  • Two-layer budgets: agent 5 USD, workspace 20 USD/day, plus room, project and API key, with a reserve/settle ledger.
  • Triggers: cron, HMAC webhooks, email/ICS pollers, room schedules on a 30-second tick, OAuth for GitHub, Linear, Notion, Sentry, Cloudflare.
  • Handover: docx/pptx/xlsx, SHA-256 signoff, Telegram out.
  • BYO endpoint: any OpenAI-compatible base_url works, so Ollama, LM Studio and vLLM already run; RAG embeddings are already local on nomic-embed-text.

And what’s missing, which turned out to be plumbing rather than architecture:

  • Triggers land in Chat or Origin, not in a room’s /goal — so a scheduled job has never had a checker.
  • The checker returns prose, not a structured verdict, so nothing downstream can branch on it.
  • The council always runs all three rounds, even for easy questions. Council and delegate both set _no_cache, so the two most expensive paths never hit the cache. (I expected missing Anthropic cache_control too; the gateway, cli-proxy-api, already injects it for Claude — false alarm.)
  • No outbound webhook when a run finishes. No Telegram in. No local preset, no endpoint health check.
  • No packaged “one job” to point a customer at.

The feature I chose: Routines

A Routine is one recurring job, end to end, under a budget that stops itself:

trigger → room /goal → producer on a cheap/local model → checker on a strong model → conditional escalation → deliverable + signoff → delivered to the team’s channel

In customer language: “every morning at 8, collect Linear tickets and GitHub PRs, write a one-page brief, have the checker verify the numbers, post to the team Telegram.” Or: “order form webhook → extract → match against the price list → export xlsx → wait for accounting to approve.”

Step by step:

  1. Trigger. Cron, HMAC webhook, email/ICS, later Telegram in. The payload attaches to the run as a SHA-256’d input (attach_input exists).
  2. Produce on the cheap tier. The producer runs on the lowest rung of the ladder, preferring local. route_down is currently opt-in and only applies to the first attempt; for Routines it becomes the default.
  3. Verify on condition. The checker runs a strong model in prompt_mode=task and returns a JSON verdict: pass | revise | escalate, a list of faults with quotations, a confidence. On revise the producer gets only the verdict and the rejected passage, not the whole draft — roughly half the tokens of round two. Two rounds maximum (already GOAL_ROUNDS). On escalate, or two rejections, two arbiters who were neither producer nor checker rule on exactly the disputed criteria: blind, in parallel, one call each. A defect clears only when both say it holds, with evidence; the merge is deterministic, so no extra Lead call. No cross-examination.
  4. Hand over. The deliverable is a file in the room, signed off against a manifest, with docgen if it’s a document. The human gate is the existing 300-second tool approver on Today, later an Approve button in Telegram.
  5. Deliver. Telegram out (exists), plus an outbound webhook POST url {run, status, files, cost} so the customer’s n8n / Make / Zapier / Google Sheet receives it, and the OpenAI-compatible /v1 API so their existing tools can call Hive.
  6. Budget that pauses. Each Routine has a monthly budget_usd. Hit the ceiling and it pauses, notifies Telegram, and makes no model call. Reuses the existing ledger.

The cost arithmetic

This is the whole argument in numbers. A six-seat council costs 2 + 2N = 14 calls on a strong model, every time, easy question or hard. A Routine expects 2–3 calls: one producer on a cheap/local model, one checker on a strong model, occasionally one revision. Worst case — two revisions and an escalation to the mini-council — about 7. The heavy part runs on the cheap tier; the strong model reads, it doesn’t write.

The rule I’ll write into the docs: local for production, extraction and classification; cloud for judgment and verification. Hive’s Brain (smart engine, graph, Origin) still needs Claude and stays that way. Routines don’t go through the Brain.

Not building

I wrote this list before the to-do list, and I think it’s worth more.

  • A drag-and-drop canvas — n8n, Dify, Flowise have won.
  • A fourth orchestrator next to CrewAI, LangGraph and Agent Framework (AutoGen is already in maintenance mode).
  • 70 connectors.
  • An agent message bus, presence.
  • Council as the default.
  • A local Brain.
  • Marketplace, per-seat billing, multi-tenancy.
  • A2A (150 organisations) waits until a customer asks. Wrapping Hive as an MCP server so Claude Code, Cursor or n8n can call a Routine as a tool comes first.

I’m also dropping the 12 USD/seat price I once wrote down. Sell one Routine that runs: a self-hosted deployment on the customer’s machine, the first three Routines configured, a monthly operations fee. The sales proof is a Usage page showing what each Routine costs against running fully in the cloud.

Order of work

Three waves, each locked by an eval before the next.

  1. Ollama / LM Studio / vLLM presets with a health check; a local rung on the ladder; triggers targeting room:/goal; outbound webhook. Acceptance: a real cron runs /goal on Ollama, checker on cloud, a file is produced, the webhook receives the payload.
  2. Structured verdict and a fix round scoped to the failed criteria; arbitration on dispute; drop _no_cache for delegate; parallel delegates. Acceptance: token counts before/after on 5 eval cases, cache hits > 0, two-seat council in ≤ 5 calls.
  3. Telegram in with an Approve button; Routine budget that pauses and notifies; a Usage page per Routine; a landing page that sells Routines. Acceptance: hit the ceiling → paused, no model call; restore drill.

What I took away

I started the day expecting to find a feature competitors had and Hive didn’t. I ended it with the opposite: Hive already stands in the four gaps the market complains about — a checker, two budget layers, one stdlib process, evals — but hadn’t joined them into something buyable. The work isn’t adding systems. It’s connecting the pipes and removing options.

Reading thirty competitors to learn what not to build turned out to matter as much as learning what to build.

Export for reading

Comments