JOURNAL

1. Analyze the Work Before You Build the Chatbot

Define chatbot jobs, risk boundaries, quality targets and a measured path toward fewer model dependencies.

From user work to a quality contract. Ask: September revenue / USD · Clarify: month or year is missing · Refuse: SQL, other tenants, profit · Complete: number + source + period. Commerce Assist · classify before querying

Read with AI

Choose content to copy and paste into your AI assistant. Nothing is sent automatically. CMS content is converted to Markdown; original Markdown is used when available.

Chatbot Engineering · Part 1 of 9 · Research checked October 7, 2026. Proposed designs and assumptions are distinguished from measured implementation results.

From user work to a quality contract. Ask: September revenue / USD · Clarify: month or year is missing · Refuse: SQL, other tenants, profit · Complete: number + source + period. Commerce Assist · classify before querying
From user work to a quality contract. Ask: September revenue / USD · Clarify: month or year is missing · Refuse: SQL, other tenants, profit · Complete: number + source + period. Commerce Assist · classify before querying

A chatbot becomes useful when it completes a recognizable job: explain a product, answer a documented question, compare business metrics or prepare an action for approval. “Use the strongest model” is not a product requirement. It leaves success, cost and acceptable failure undefined.

This series develops one reference product: a bilingual business assistant that answers approved documentation questions, analyzes synthetic sales data and proposes actions without executing them automatically. The target is fewer operational dependencies and predictable quality. Matching a frontier model across arbitrary conversation is outside that target.

Start with an inventory of actual work

Collect consented, redacted requests from the intended workflow. Interview both users and the people who currently resolve their questions. For each request, record the desired outcome, required evidence, permissions, frequency, acceptable delay and consequence of an incorrect answer. Keep an ambiguous request as ambiguous instead of manufacturing a confident label.

For example, “How were sales last month?” needs clarification: gross or net revenue, which timezone, completed or created orders, and which business unit? “Summarize the refund policy” needs a versioned source. “Refund this order” requires identity, an authorized tool and a separate execution decision.

Job Evidence Completion condition
Documentation answer Authorized, current documents Supported answer with source links
Business analysis Approved metric and data snapshot Correct result, filters and freshness
Action preparation Validated record and permissions Reviewable proposal; no unintended execution
Unclear request Missing scope A useful clarifying question

Separate decisions, answers and actions

A router selects a task. A generative model writes an answer or proposes a structured plan. Application code authorizes and runs tools. These roles can share infrastructure, but their responsibilities should remain explicit.

Jev is relevant to the decision role: its official interface exposes typed choices, scores and yes/no decisions. It does not generate conversational answers. A Jev-like classifier therefore replaces some routing calls, not the whole chatbot. TypeSafe AI model documentation.

Use deterministic rules for exact conditions such as authentication, account ownership, numeric limits and irreversible-action approval. A model may recognize intent, but it must not grant permissions. A low model score also cannot be interpreted as a calibrated likelihood of failure without validation.

Write a quality contract

Define acceptance criteria before comparing models. Illustrative targets for the reference product are at least 95% exact metric correctness, at least 90% evidence-supported documentation answers, zero observed cross-tenant disclosures in the release test suite, and p95 complete-answer latency below five seconds at a stated concurrency. These are proposed gates, not measured achievements or guarantees.

Split results by language and task. An aggregate score can conceal weak Vietnamese performance or a failure on high-risk workflows. Record when the system should abstain. A chatbot that asks one relevant question can be better than one that supplies an impressive answer to the wrong metric.

Reduce dependencies in the right order

Begin with one answer-model interface, one retrieval strategy and a small typed tool catalog. Establish a strong-model baseline. Then test a smaller model on the same work. Add a separate classifier only if routing measurably improves cost or reliability. Add reranking, OCR, vision or a specialist only for a demonstrated gap.

“One model” can mean one generative backbone while a lightweight encoder handles retrieval. It need not mean sending every deterministic task through that backbone. Conversely, using different adapters for every client adds training, versioning and memory-management work even when the base weights are shared.

Define the first release

The smallest useful release has source-grounded answers, two or three approved business metrics, clarification, visible failures and user feedback. Defer autonomous writes, arbitrary SQL and connector marketplaces until there is evidence users need them. Measure task completion rather than messages sent.

Have a user complete a real scenario: choose a workspace, ask for monthly revenue, correct an assumption, inspect the result and export it. Observe where they lose trust or repeat work. Those failures provide better backlog items than a generic list of AI features.

What the existing experiment proves

The Private Decision Lab already demonstrates bilingual routing on a CPU LXC. Its frozen encoder and trained classifier heads return templates or simulated recommendations. It has no free-form generation, real business database or executable action tools.

The implementation case study reports small, same-author test fixtures and observed failures. It is useful evidence for feasibility, not proof of production quality or equivalence to Jev. The remaining chapters explain how to turn that starting point into a measured chatbot engineering program.

Workshop: turn a vague request into a product decision

We will call the fictional reference product Commerce Assist. Its first customer is an operations manager who checks monthly revenue and explains policy changes. The reference dataset is synthetic, uses USD and reports calendar periods in UTC. These constraints deliberately remove currency conversion and ambiguous business calendars from the first experiment. A real client must approve its own definitions.

Start with this request: “How did we do last month?” There are at least four decisions hidden inside it: what “we” means, what “did” measures, which month the user intends and what comparison is useful. The authenticated workspace resolves the tenant. A workspace-level reporting preference may resolve the metric. If that preference is absent, ask for the metric. Resolve relative dates from the request timestamp and reporting timezone, then display the exact period. Never let the model guess a tenant from conversation.

Request family Preferred path Why this path When to escalate
“Open the products page” Validated navigation template No generated explanation needed Destination is unclear
“Net revenue in September 2026” Structured metric plan + tool The database owns the value Metric or reporting calendar is missing
“What changed in the refund policy?” Authorized versioned retrieval + answer Needs evidence from two document versions A required version is unavailable
“Why did revenue fall?” Comparison tools + qualified explanation Can establish changes, not causal proof Required dimensions are unsupported
“Refund all those orders” Prepare a scoped action proposal Financial side effects require separate authority Writes are disabled in the release

Quantify the cost of a wrong answer

Do not apply the same success metric to navigation and financial decisions. An incorrect navigation answer may cost a few seconds. A fabricated revenue number can contaminate a management report. A cross-tenant disclosure cannot be repaired by a more accurate average score. Use a risk gate before any weighted utility calculation.

For an illustrative workload of 1,000 monthly questions, suppose 600 are routine metrics, 250 policy questions, 100 unclear and 50 action requests. If the system answers all 100 unclear requests rather than asking, a high overall completion rate may actually measure overconfidence. Maintain separate counters for correct completion, useful clarification, appropriate refusal and wrong completion. Useful clarification should not be labeled a successful numerical answer, but it should not automatically be labeled failure either.

A practical release worksheet asks: who signs off the metric catalog; who owns document freshness; which requests are blocked; what happens when no source is available; and who investigates a disputed answer? These are product ownership decisions. A technical lead can implement limits, but cannot invent a finance team’s definition of revenue.

Choose the first slice using an explicit trade-off

Score candidate jobs by frequency, business value, evidence availability and consequence of error. For this reference product, monthly net revenue has a clear query and repeat demand, so it is a good first slice. Open-ended recommendations about business strategy have less objective ground truth and are a poor first acceptance test. Refund execution introduces a new security and reconciliation problem, so keep it outside the initial read-only release.

The first deliverable is one vertical scenario: identify workspace, resolve the metric and period, run a scoped query, render evidence and permit an authorized export. Complete that path before adding many connectors. Test it with a nontechnical user who does not know the underlying schema. If they cannot tell whether the number covers September or the current rolling 30 days, the product contract is still incomplete.

Artifacts to take back to your project

Produce a one-page scope, five representative scenarios per request family, a signed metric definition, an answer acceptance rubric and a list of blocked capabilities. Assign an owner to each artifact. For every feature proposal, identify the scenario it improves and the metric that would reveal improvement.

Your completion exercise is to rewrite “build an AI chatbot for sales” into: “An authenticated operations user can request net revenue for a calendar month, inspect its definition and effective period, and export only permitted aggregate results; absent scope triggers clarification, and unavailable data produces an explicit unavailable state.” That requirement is implementable, testable and suitable for comparing model sizes.

From this chapter to a runnable experiment

Benchmark and training use author-created synthetic data with correlated templates. They do not establish equivalence to larger models or production customer quality. Generative planners are evaluated offline; the public demo uses the controlled baseline. Benchmark: four candidates completed. LoRA: completed, 64/64 steps.

Reading path and measured evidence · Try the synthetic-data demo · Download example code

What this implementation demonstrates

The runnable contract accepts net revenue in USD and explicit monthly periods. Missing time requires clarification; profit, customer filtering and writes are refused. This is a testable product boundary rather than unrestricted conversation.

Discussion

Comments are reviewed before publication. Your email is kept private.

← Back to allĐọc tiếng Việt