JOURNAL

4. Prove Whether a Smaller Chatbot Is Good Enough

Build independent evaluation sets, score grounded answers and measure quality, cost and coverage together.

Evaluate each layer independently. Planner: correct outcome and periods? · Validator: valid schema and allowlist? · Executor: read-only, server-owned tenant · Answer: verified numbers and sources?. A wrong planner is separate from an executor breach

Read with AI

Choose content to copy and paste into your AI assistant. Nothing is sent automatically. CMS content is converted to Markdown; original Markdown is used when available.

Chatbot Engineering · Part 4 of 9 · Research checked October 7, 2026. Proposed designs and assumptions are distinguished from measured implementation results.

Evaluate each layer independently. Planner: correct outcome and periods? · Validator: valid schema and allowlist? · Executor: read-only, server-owned tenant · Answer: verified numbers and sources?. A wrong planner is separate from an executor breach
Evaluate each layer independently. Planner: correct outcome and periods? · Validator: valid schema and allowlist? · Executor: read-only, server-owned tenant · Answer: verified numbers and sources?. A wrong planner is separate from an executor breach

“Comparable quality” must mean comparable performance on a defined workload, with a stated tolerance. A local model can equal a stronger model on a constrained metric-selection task and still fail badly at unfamiliar reasoning. Evaluation should make that boundary visible.

Create an independent test set

Separate training, development and locked test sets before model tuning. Split by customer, source document, workflow family and time where appropriate. Near-duplicate paraphrases of one source request belong in the same split. Otherwise the test mostly measures recognition of examples the model has effectively seen.

Use an independent author or redacted real requests for the locked set. Include short Vietnamese requests, code-switching, missing dates, inconsistent terminology, stale documents, impossible requests and permission boundaries. Keep a separate challenge suite for adversarial behavior. A challenge suite tests specific failures; its class distribution is not representative of production traffic.

Grade the job, not the prose

For a data-agent case, compare metric identifiers, filters, grouping, result values and authorized scope. For a document answer, verify that cited passages support each material claim and were accessible to that user. For action preparation, compare the intended target and parameters without executing the action.

Measure Question it answers
Exact task correctness Did the system complete the requested job?
Grounding and citation validity Can the claims be checked against authorized evidence?
Coverage at a quality threshold How much work can it complete without escalation?
False accept / false refuse Does it proceed incorrectly or block useful work?
p95 latency and successful-task cost Is it viable under the stated load?

Use paired comparisons

Run each candidate on the same input, evidence snapshot and tool results. Record model revision, prompt revision, sampling settings, context size and runtime. Give the grader anonymized outputs in randomized order. Human reviewers should inspect disagreements and high-impact failures rather than accept an LLM judge as ground truth.

An illustrative non-inferiority rule is that the smaller model’s task success rate must be within two percentage points of the reference, with a prespecified confidence interval and enough examples to support that conclusion. Choose sample size using expected error rates and the intended statistical test. A dozen examples cannot substantiate a two-point difference.

Evaluate abstention and routing together

A router that sends every request to a stronger model achieves little cost reduction. A router that sends everything to the cheap path may hide expensive correctness failures. Plot accepted coverage against error rate while varying the threshold on the development set, then lock the threshold before the final test.

Report raw predictions and post-policy outcomes separately. Model errors, clarifying questions and security-policy interventions have different causes. Calibrate probabilities only on representative held-out data; a margin between classifier scores is not automatically a probability that the answer is correct.

The existing demo’s failure cases matter

The CPU case study achieved 11 correct raw chatbot routes out of 12 held-out examples and seven raw agent routes out of eight. After the margin policy, the LXC agent result was six of eight. Those fixtures were written by the same author as the training data, so the results are illustrative.

“Vẽ sơ đồ microservices” was classified as a diagram request but was then diverted to clarification because its margin was small. “Giúp tôi với” should have asked for clarification but went to a product route. This shows why a protective threshold can reduce utility and why ambiguity needs its own evaluation. Measured implementation case study.

Test the full system

Model-only scores omit retrieval, authorization, tool timeouts, frontend retries and rendering. Test missing evidence, an unavailable database, an expired session, a canceled browser request and a quota rejection. A useful failure should explain what happened and offer an appropriate recovery without fabricating a result.

Use synthetic tenant canaries in non-production isolation tests. Verify that inaccessible content never reaches the prompt, cache or export. Also test a warm connection reused after another tenant’s query. Include malformed structured outputs and valid-looking requests for forbidden dimensions.

Promote a release only with evidence

Maintain a release report listing dataset hashes, candidate versions, task-level scores, language slices, failures, latency distribution, cost assumptions and an explicit go/no-go decision. Shadow traffic first when permission and privacy controls allow it. Then canary a limited share with an immediate rollback path.

# Proposed report fields; fill only from actual runs
model_revision: ...
prompt_revision: ...
dataset_sha256: ...
hardware_and_concurrency: ...
schema_valid_count / total: ...
correct_tasks / total: ...
en_success / en_total: ...
vi_success / vi_total: ...
accepted_coverage_and_error_rate: ...
unauthorized_disclosures_observed: ...
p50_ms / p95_ms / peak_memory_mb: ...
total_cost_usd / successful_tasks: ...
known_failures_and_release_decision: ...

Do not repeatedly tune against the locked test. When a new failure becomes a regression example, reserve new independent examples for the next release. Evaluation is a maintained product asset, not a one-time screenshot of a favorable score.

Workshop: build the Commerce Assist evaluation pack

Give each case an immutable ID and separate the input from the expected result. A case needs language, user scope, permitted sources, source snapshot, request timestamp, expected plan or clarification, expected result and severity. The grader should not see the model name. Keep the expected tenant in the harness rather than in model-editable text.

{
  "case_id": "revenue-compare-vi-001",
  "language": "vi",
  "authenticated_workspace": "alpha",
  "request": "So sánh doanh thu thuần tháng 9/2026 với tháng Tám",
  "expected_metric": "net_revenue_v1",
  "expected_periods": [
    ["2026-09-01", "2026-10-01"],
    ["2026-08-01", "2026-09-01"]
  ],
  "expected_cents": [20000, 10000],
  "expected_difference_cents": 10000,
  "severity": "financial_accuracy",
  "fixture_kind": "synthetic"
}

Add a paired English case with the same expected plan, not merely the same keyword. Add a genuinely ambiguous variant such as “Compare last month’s performance” with no default metric. Its expected result is clarification. A model that produces the numerical fixture for that variant should fail the case even if the number happens to be correct.

Use a deterministic grader where the output is objective

Parse the proposed plan, validate its schema and compare normalized dates, metric version and allowed grouping. Compare integer cents exactly. Allow approved aliases only through a versioned normalization map, not an ad hoc grader that rescues a preferred candidate. For answers, evaluate the material claims separately: correct amount, correct period, supported comparison and no unsupported causal claim.

Grade a citation by resolving its source and checking the cited passage against the claim. Merely including a URL is not grounding. Grade authorization using the actual evidence handed to the model and the rows returned by tools. A correct final sentence can still be a failed security test if forbidden evidence entered the prompt.

Calculate uncertainty instead of reporting a perfect tiny score

A useful statistical illustration is the “rule of three”: with zero observed failures in n independent representative trials, an approximate one-sided 95% upper bound on failure probability is 3/n. At n=100, zero failures is still consistent with a failure rate near 3%; at n=1,000, roughly 0.3%. This approximation does not turn an adversarial fixture suite into representative production sampling and does not guarantee security.

For comparing two models on the same cases, retain paired outcomes: both correct, only A correct, only B correct and both wrong. Use an appropriate paired confidence procedure or test selected in advance. Count disagreement cases for human adjudication. Do not use two independent-looking averages to discard the advantage of paired evaluation.

Turn threshold selection into a decision table

Hypothetical threshold Accepted cases Wrong accepted Coverage Error among accepted
Low 90 of 100 9 90% 10%
Medium 70 of 100 2 70% 2.86%
High 40 of 100 0 40% 0% observed; uncertain

These values are fabricated to illustrate a trade-off. If the target error among accepted requests is below 3%, the medium threshold is a candidate for further validation; the high threshold sacrifices substantial coverage and still lacks a guarantee. Choose on development data, then evaluate the fixed policy on an independent test. Also inspect the rejected cases: repeatedly asking for clarification on clear Vietnamese requests is a product regression.

Create a release packet an operator can use

The packet should contain the case manifest, source and artifact hashes, configuration, raw predictions, grader output, adjudicated disagreements and a concise failure taxonomy. Distinguish planner errors, query errors, retrieval errors, unsupported answer claims and UI failures. That separation tells you whether to change the model, data layer or interface.

Finish by giving a colleague five failed cases and asking them to reproduce the result from the manifest. If they cannot reproduce which model, prompt and snapshot were used, the evaluation is not ready to govern a deployment.

For formal interval calculations, see NIST exact binomial confidence limits.

From this chapter to a runnable experiment

Benchmark and training use author-created synthetic data with correlated templates. They do not establish equivalence to larger models or production customer quality. Generative planners are evaluated offline; the public demo uses the controlled baseline. Benchmark: four candidates completed. LoRA: completed, 64/64 steps.

Reading path and measured evidence · Try the synthetic-data demo · Download example code

What this implementation demonstrates

Keep outcome, exact-plan and computed-result correctness separate. Incorrect dates can still produce identical empty results. Reports retain raw outputs to make that difference inspectable. Planner evaluation does not measure executor breaches; tenant and read-only protections are checked in integration tests.

Measured results on the synthetic test set
Candidate Exact plans Correct results / metric Invalid outputs p50 / ms RSS / MiB
baseline 120/120 60/60 0 0.043 25
qwen25 8/120 7/60 37 12113.0 618
qwen3 29/120 24/60 9 7200.4 920
adapter 80/120 50/60 2 7888.2 918

Separate challenge: zero previous-period baseline

These two EN/VI cases are reported separately from the main 120. September versus June has a zero previous total; percentage change must be null. On both EN/VI cases, all three models discarded comparison_period despite the comparison request. Protected arithmetic cannot fix this interpretation failure.

Candidate Exact plans / 2
baseline 2/2
qwen25 0/2
qwen3 0/2
adapter 0/2

A valid but incorrect plan

Qwen2.5 · test-002-en · Expected and actual on the same test case:

{
  "expected": {
    "type": "metric_plan",
    "payload": {
      "metric": "net_revenue",
      "period": {
        "start": "2026-08-01",
        "end_exclusive": "2026-09-01"
      },
      "comparison_period": null,
      "group_by": null,
      "currency": "USD"
    }
  },
  "actual": {
    "type": "metric_plan",
    "payload": {
      "metric": "net_revenue",
      "period": {
        "start": "2026-08-01",
        "end_exclusive": "2026-09-30"
      },
      "comparison_period": null,
      "group_by": null,
      "currency": "USD"
    }
  }
}

Discussion

Comments are reviewed before publication. Your email is kept private.

← Back to allĐọc tiếng Việt