JOURNAL
8. Price a Chatbot Product Around Real Unit Economics
Calculate cost per successful task, account for cache writes and capacity, and offer transparent client pricing.

On this page
Chatbot Engineering · Part 8 of 9 · Research checked October 7, 2026. Proposed designs and assumptions are distinguished from measured implementation results.

Charge for a dependable product, not an unexplained multiplier on tokens. Clients buy answered questions, verified analysis, secure access and support. Token accounting is still important because it reveals whether the promised service has sustainable margins.
Calculate cost per successful task
model_cost = (
uncached_input_tokens * input_rate
+ cache_read_tokens * cache_read_rate
+ cache_write_tokens * cache_write_rate
+ output_tokens * output_rate
) / 1_000_000
task_cost = model_cost + retrieval + tools + retry_cost
fully_loaded_cost = task_cost + allocated_infra + support + training_amortization
cost_per_success = total_cost_for_period / successful_tasks_for_period
Map provider usage categories carefully. Depending on the provider, writing a cache can be an additional charge or a separately priced input category. Normalize the invoice into your internal accounting instead of counting the same base input twice by accident. Add storage, embedding ingestion, observability and payment processing where they apply.
Use the exact model, region, context tier and billing mode. OpenAI’s official pricing page is the source for live rates; the numerical example below is deliberately hypothetical. It is not a quotation for an OpenAI, Anthropic or local-hosting plan.
A worked example with explicit assumptions
Suppose one task uses 2,000 uncached input tokens, 3,000 cache-read tokens and 500 output tokens. Assume illustrative rates of $1, $0.10 and $4 per million tokens respectively, with no cache write on that particular task. Its model cost is $0.0043. If tools and retrieval add $0.002, variable cost becomes $0.0063.
At 10,000 such tasks, variable cost is $63. Add a hypothetical $120 monthly infrastructure allocation and $100 support allocation: total monthly cost is $283. At a target gross margin of 60%, required revenue is $283 divided by 0.40, or $707.50 before tax and other unmodeled costs.
This is margin pricing, not a 60% markup. Multiplying cost by 1.6 produces a much smaller margin. Also distinguish tasks from messages: a task may require clarification, a query and an explanation. Use observed distributions instead of treating every conversation as the example above.
Self-hosting changes the cost structure
A local model avoids a third-party per-token invoice but still consumes compute, electricity, hardware life, storage and operator time. Include idle capacity, backup, monitoring and upgrades. Amortize training and evaluation over a realistic useful lifetime rather than assuming the adapter never changes.
An illustrative comparison: if dedicated local capacity costs $400 monthly and avoids $0.01 of net external cost per task, simple break-even is 40,000 tasks per month. That arithmetic omits differences in quality, concurrency, additional support and hardware risk. It is a screening estimate, not a purchasing decision.
Use a predictable commercial package
For this kind of B2B assistant, a practical initial structure is a one-time onboarding fee, a monthly workspace subscription with included usage, and published overage rates. Price private infrastructure and unusual connectors separately because they create real setup and capacity costs.
| Component | What the client receives | What must be defined |
|---|---|---|
| Onboarding | Connector setup, metric mapping and acceptance testing | Included sources and change scope |
| Subscription | Workspace, monitoring and standard support | Seats, retention and included usage |
| Overage | Additional completed tasks or published credit units | Rate, warning thresholds and hard cap |
| Private deployment | Dedicated approved processing boundary | Capacity, availability and maintenance responsibilities |
Set actual prices after a pilot establishes support load, workload mix and willingness to pay. A low token bill does not justify a low-price promise if connector maintenance dominates cost. Conversely, do not disguise ordinary internal retries as extra customer actions.
Make billing inspectable
Define whether usage is charged per accepted task, successful result or weighted credit. Document how a clarification, retry, cancellation, cached answer and failed tool call count. As a starting policy, absorb internal service failures and show any user-requested additional analysis before charging it.
Provide a usage dashboard with included balance, overage estimate, budget alerts and an administrator-controlled hard cap. Offer exportable usage records with request identifiers and service categories, without exposing conversation content. Require explicit opt-in for overage rather than surprising the client.
Outcome pricing needs a defensible definition
Pricing by “resolved ticket” can work only when both parties agree what resolution means, how reopenings are handled and how attribution is measured. Otherwise billable outcomes become disputes. Start with a clear subscription and metered capacity while measuring outcomes separately.
The first commercial goal is a trustworthy invoice and enough margin to maintain quality. Revisit rates using actual successful-task cost and support data, not a vendor’s token price alone.
Workshop: turn unit economics into a client quotation
Use the earlier hypothetical $0.0063 variable cost per ordinary task and $220 monthly fixed allocation for infrastructure and support. At 10,000 ordinary tasks, fully loaded cost is $283. An illustrative $799 monthly package would leave $516 before tax and unmodeled expenses, a gross margin of approximately 64.6%. This is a pricing worksheet, not an offered commercial plan or current vendor rate.
The package needs explicit limits: one workspace, an agreed connector scope, an included task allowance, a concurrency ceiling, data retention and a defined support window. An enterprise client asking for dedicated private capacity and an uptime commitment is buying a different service. Quote its infrastructure, recovery and support obligations separately.
Test the expensive tail, not just the average
Suppose 5% of tasks cost $0.12 rather than $0.0063 because they require long evidence or expensive tools. Weighted variable cost becomes 0.95 × 0.0063 + 0.05 × 0.12 = $0.011985 per task. At 10,000 tasks, variable cost is $119.85 and total cost is $339.85. The same $799 package now has approximately 57.5% gross margin, below the illustrative 60% target.
This is why “10,000 messages included” can be commercially dangerous. A message is not a workload unit. A plan may distinguish standard analysis from a disclosed extended-analysis class, or use credits with published weights. If a heavy task consumes twenty credits, the ordinary-cost equivalent is $0.126, roughly covering the assumed $0.12 variable cost before other allocations. The classification must be visible before the task runs, not invented after the bill.
| Illustrative monthly case | Variable cost | Total with $220 allocation | Margin at $799 |
|---|---|---|---|
| 10,000 ordinary tasks | $63.00 | $283.00 | 64.6% |
| 5% heavy, 95% ordinary | $119.85 | $339.85 | 57.5% |
| No tasks, capacity retained | $0.00 | $220.00 | 72.5% |
Define a billable event and reconcile it
Create one usage record per accepted task ID, with task class, quoted credits, outcome and provider/tool usage. Record attempts separately. An internal retry should update the same task’s operating cost, not create a new client charge. A user-requested new comparison is a new task when the product has disclosed that policy.
A failed internal tool call can consume provider tokens even when client usage is refunded. Keep a provider-cost ledger and a client-billing ledger distinct so you can explain the difference. Reconcile the provider invoice, your request records and the client’s credit statement monthly. Set an explicit policy for canceled work that has already incurred external cost.
Price onboarding from a work breakdown
Estimate connector setup, metric-definition workshops, permission mapping, data-quality repair, acceptance testing and client training separately. For example, a hypothetical 24 engineering hours at an internal loaded cost of $30/hour plus eight product/testing hours at $25/hour costs $920 before overhead. That is a cost estimate, not automatically a selling price. Apply the desired margin and account for change scope, payment fees and project risk.
Define which changes count as configuration and which require implementation. Adding a new synonym may be configuration; creating a new metric with joins and reconciliation is engineering. A subscription promising unlimited connectors and custom analytics can create an unbounded services obligation even with cheap inference.
Use a capacity reservation for private deployment
A dedicated local service should be priced around reserved capacity and operations, with a tested workload envelope. Describe the hardware or resource allocation, approved models, maximum context, concurrency and maintenance responsibility. Avoid selling “unlimited local tokens”; unlimited demand is incompatible with a finite server.
Your completed quotation should include scope, onboarding, monthly reservation, included usage, overage policy, hard cap, failure credits, retention, support, availability and an example invoice. Have a client administrator explain that example back to you. If they cannot predict the bill for a standard and a heavy task, simplify the package before launch.
From this chapter to a runnable experiment
Benchmark and training use author-created synthetic data with correlated templates. They do not establish equivalence to larger models or production customer quality. Generative planners are evaluated offline; the public demo uses the controlled baseline. Benchmark: four candidates completed. LoRA: completed, 64/64 steps.
Reading path and measured evidence · Try the synthetic-data demo · Download example code
What this implementation demonstrates
Article pricing is illustrative. Self-hosting removes per-token API charges but retains compute, memory, operations and capacity costs. Price the value and task budget; measure heavy-request frequency before promising a usage allowance.


Discussion
Comments are reviewed before publication. Your email is kept private.