I self-host an agent platform called Hive: standard-library Python, SQLite, a Preact frontend with no build step, one process on one Proxmox container. It already has rooms where several agents talk to each other, a debating council, cron, webhooks, budgets, evals. Honestly, it has a lot. The question I hadn’t answered was: if a five-person team wanted to buy it, what would they be buying?
On 18 September I spent a full day reading competitors to answer that. Around thirty products in
six groups. Pricing pages, complaints on G2, Reddit and Hacker News, and two research papers on
multi-agent debate. Then I re-read my own code — flows.py, rooms.py, llm.py,
triggers.py — so I wouldn’t lie to myself about what Hive already had.
The result was tidier than I expected.
The conclusion first
The 2026 market does not pay for orchestration. The open-source core is free everywhere: n8n, Dify, CrewAI, LangGraph, Paperclip. Nobody pays to have “agents wired together” anymore. They pay for a repeated job that runs correctly, with someone approving it, without wrecking the invoice.
The four biggest complaints, repeated everywhere:
- Cost surprises. Zapier sits at 1.4 stars on Trustpilot, mostly over billing (Zapier Agents complaints). Lindy’s credits are opaque. Make charges 50 credits per agent run.
- Brittle agents. CrewAI has a whole thread on stopping infinite loops from eating tokens. n8n has a memory failure in the AI Agent node. LangGraph shipped a self-hosted flaw chain in June.
- Nobody checks the output. The quietest complaint and the most painful. Human-in-the-loop, where it exists, is a gate on a tool call (n8n 2.6’s HITL tools), not a second model reading the deliverable and saying whether the numbers are right.
- Self-hosting is hard. Dify and LibreChat are a stack of containers (a six-month Dify review). Open WebUI changed its license. OpenClaw exposed 21k instances to the internet.
Two more signals told me where the floor and the ceiling are. OpenAI is shutting AgentKit down on 30 November 2026 (source) — the closed orchestrator from the biggest vendor didn’t survive a year. And Paperclip went to 30k GitHub stars in three weeks on exactly one idea: agents organised like an org chart, each with a budget that pauses the agent at 100 %. People aren’t short of frameworks. They want to know where the bill stops.
The competitor table (condensed)
| Group | Who | What to learn | What not to copy |
|---|---|---|---|
| No-code canvas | n8n, Dify, Flowise, Langflow, Make, Zapier Agents | HITL at tool level; webhook/cron/Slack are table stakes; Dify shows cost per run | The canvas. They’ve won; Hive composes work with Rooms, not nodes |
| Dev frameworks | CrewAI, LangGraph, Agent Framework (AutoGen), Agno, Mastra | A lead that keeps a ledger and re-plans (Hive already has the board); suspend/resume | A fourth orchestrator |
| ”AI employee” SaaS | Lindy, Relevance, Gumloop | Pricing by model tier; manager agent | Opaque credits, always on |
| Control planes | Paperclip, Block Buzz, Odysseus | Per-agent budgets that pause; agents as an org chart | Runtime-agnostic everything |
| Closed vendors | Claude Cowork / routines / Managed Agents, OpenAI AgentKit, Gemini Enterprise | Routine = recurring job with a deliverable; 20–30 USD/seat is the tolerance ceiling | Depending on one vendor |
| Local chat | Open WebUI, AnythingLLM, LibreChat, OpenClaw | Local-first is coming back | Channel sprawl, no verification |
The “not to copy” column turned out to matter as much as the “learn” column. It’s the wall between me and the thing I’m best at: adding systems.
The paper that changed my mind
I used to think of the council — several agents debating over several rounds — as Hive’s “premium” mode. Two papers say the opposite. The Cost of Consensus (arXiv, May 2026) and the earlier problem drift paper both find that free debate between homogeneous models performs worse than isolated self-correction, because the agents drift off the problem together. The winning pattern is a verifier triggered on condition: one writer, one checker, and a heavier session only when the checker objects.
Hive has both shapes. Rooms run /goal (producer, then checker, two rounds) and a full four-phase
council at 2 + 2N calls for N seats, up to six seats. The research says: make the first one the
default, and stop treating the second as the grown-up option.
What Hive has, what it lacks
Reading the code honestly:
- A checker already in
/goal;/reviewruns two parallel reviewers and mediates. - Two-layer budgets: agent 5 USD, workspace 20 USD/day, plus room, project and API key, with a reserve/settle ledger.
- Triggers: cron, HMAC webhooks, email/ICS pollers, room schedules on a 30-second tick, OAuth for GitHub, Linear, Notion, Sentry, Cloudflare.
- Handover: docx/pptx/xlsx, SHA-256 signoff, Telegram out.
- BYO endpoint: any OpenAI-compatible
base_urlworks, so Ollama, LM Studio and vLLM already run; RAG embeddings are already local onnomic-embed-text.
And what’s missing, which turned out to be plumbing rather than architecture:
- Triggers land in Chat or Origin, not in a room’s
/goal— so a scheduled job has never had a checker. - The checker returns prose, not a structured verdict, so nothing downstream can branch on it.
- The council always runs all three rounds, even for easy questions. Council and delegate both set
_no_cache, so the two most expensive paths never hit the cache. (I expected missing Anthropiccache_controltoo; the gateway, cli-proxy-api, already injects it for Claude — false alarm.) - No outbound webhook when a run finishes. No Telegram in. No local preset, no endpoint health check.
- No packaged “one job” to point a customer at.
The feature I chose: Routines
A Routine is one recurring job, end to end, under a budget that stops itself:
trigger → room
/goal→ producer on a cheap/local model → checker on a strong model → conditional escalation → deliverable + signoff → delivered to the team’s channel
In customer language: “every morning at 8, collect Linear tickets and GitHub PRs, write a one-page brief, have the checker verify the numbers, post to the team Telegram.” Or: “order form webhook → extract → match against the price list → export xlsx → wait for accounting to approve.”
Step by step:
- Trigger. Cron, HMAC webhook, email/ICS, later Telegram in. The payload attaches to the run
as a SHA-256’d input (
attach_inputexists). - Produce on the cheap tier. The producer runs on the lowest rung of the ladder, preferring
local.
route_downis currently opt-in and only applies to the first attempt; for Routines it becomes the default. - Verify on condition. The checker runs a strong model in
prompt_mode=taskand returns a JSON verdict:pass | revise | escalate, a list of faults with quotations, a confidence. Onrevisethe producer gets only the verdict and the rejected passage, not the whole draft — roughly half the tokens of round two. Two rounds maximum (alreadyGOAL_ROUNDS). Onescalate, or two rejections, two arbiters who were neither producer nor checker rule on exactly the disputed criteria: blind, in parallel, one call each. A defect clears only when both say it holds, with evidence; the merge is deterministic, so no extra Lead call. No cross-examination. - Hand over. The deliverable is a file in the room, signed off against a manifest, with docgen if it’s a document. The human gate is the existing 300-second tool approver on Today, later an Approve button in Telegram.
- Deliver. Telegram out (exists), plus an outbound webhook
POST url {run, status, files, cost}so the customer’s n8n / Make / Zapier / Google Sheet receives it, and the OpenAI-compatible/v1API so their existing tools can call Hive. - Budget that pauses. Each Routine has a monthly
budget_usd. Hit the ceiling and it pauses, notifies Telegram, and makes no model call. Reuses the existing ledger.
The cost arithmetic
This is the whole argument in numbers. A six-seat council costs 2 + 2N = 14 calls on a
strong model, every time, easy question or hard. A Routine expects 2–3 calls: one producer on
a cheap/local model, one checker on a strong model, occasionally one revision. Worst case — two
revisions and an escalation to the mini-council — about 7. The heavy part runs on the cheap
tier; the strong model reads, it doesn’t write.
The rule I’ll write into the docs: local for production, extraction and classification; cloud for judgment and verification. Hive’s Brain (smart engine, graph, Origin) still needs Claude and stays that way. Routines don’t go through the Brain.
Not building
I wrote this list before the to-do list, and I think it’s worth more.
- A drag-and-drop canvas — n8n, Dify, Flowise have won.
- A fourth orchestrator next to CrewAI, LangGraph and Agent Framework (AutoGen is already in maintenance mode).
- 70 connectors.
- An agent message bus, presence.
- Council as the default.
- A local Brain.
- Marketplace, per-seat billing, multi-tenancy.
- A2A (150 organisations) waits until a customer asks. Wrapping Hive as an MCP server so Claude Code, Cursor or n8n can call a Routine as a tool comes first.
I’m also dropping the 12 USD/seat price I once wrote down. Sell one Routine that runs: a self-hosted deployment on the customer’s machine, the first three Routines configured, a monthly operations fee. The sales proof is a Usage page showing what each Routine costs against running fully in the cloud.
Order of work
Three waves, each locked by an eval before the next.
- Ollama / LM Studio / vLLM presets with a health check; a local rung on the ladder; triggers
targeting
room:/goal; outbound webhook. Acceptance: a real cron runs/goalon Ollama, checker on cloud, a file is produced, the webhook receives the payload. - Structured verdict and a fix round scoped to the failed criteria; arbitration on dispute;
drop
_no_cachefor delegate; parallel delegates. Acceptance: token counts before/after on 5 eval cases, cache hits > 0, two-seat council in ≤ 5 calls. - Telegram in with an Approve button; Routine budget that pauses and notifies; a Usage page per Routine; a landing page that sells Routines. Acceptance: hit the ceiling → paused, no model call; restore drill.
What I took away
I started the day expecting to find a feature competitors had and Hive didn’t. I ended it with the opposite: Hive already stands in the four gaps the market complains about — a checker, two budget layers, one stdlib process, evals — but hadn’t joined them into something buyable. The work isn’t adding systems. It’s connecting the pipes and removing options.
Reading thirty competitors to learn what not to build turned out to matter as much as learning what to build.