PRODUCT × ENGINEERING

Chatbot Engineering: A Practical Field Guide

CHATBOT ENGINEERING / FIELD GUIDE

From product analysis to data agents, model evaluation, fine-tuning and operations. Follow Commerce Assist, inspect measured evidence and try the EN/VI demo.

Try Commerce Assist ↗ · Đọc tiếng Việt

Choose your reading path

  • Product: chapters 1 → 8 → 9 for value, costs and UX.
  • Engineering: chapters 2 → 4 → 3 → 5 → 6 → 7, from contract to operations.

Nine chapters

  1. Analyze the Work Before You Build the Chatbot

    Define chatbot jobs, risk boundaries, quality targets and a measured path toward fewer model dependencies.

  2. Build a Data Agent That Can Explain Its Numbers

    Design a governed data agent with a semantic layer, read-only tools, tenant isolation and verifiable answers.

  3. Choose a Chatbot Model You Can Actually Operate and Train

    Compare practical local model candidates, CPU routing and GPU fine-tuning without confusing size with production quality.

  4. Prove Whether a Smaller Chatbot Is Good Enough

    Build independent evaluation sets, score grounded answers and measure quality, cost and coverage together.

  5. Train the Behavior, Retrieve the Knowledge: A LoRA Playbook

    Prepare supervised examples, configure an adapter experiment and decide whether fine-tuning earns its operational cost.

  6. Make the Chatbot Faster Without Quietly Making It Worse

    Optimize context, output, caching, retrieval and inference using an end-to-end latency and token budget.

  7. Run a Private Chatbot as a Production Service

    Plan tenant isolation, bounded tools, deployment gates, failure recovery and observability for a production chatbot.

  8. Price a Chatbot Product Around Real Unit Economics

    Calculate cost per successful task, account for cache writes and capacity, and offer transparent client pricing.

  9. Design a Chatbot People Can Trust, Configure and Recover From

    Specify answer and analysis workflows, safe configuration, accessible interactions and a practical demo walkthrough.

Experiments and evidence

Benchmark and training use author-created synthetic data with correlated templates. They do not establish equivalence to larger models or production customer quality. Generative planners are evaluated offline; the public demo uses the controlled baseline. Benchmark: four candidates completed. LoRA: completed, 64/64 steps.

Measured results on the synthetic test set
Candidate Exact plans Correct results / metric Invalid outputs p50 / ms RSS / MiB
baseline 120/120 60/60 0 0.043 25
qwen25 8/120 7/60 37 12113.0 618
qwen3 29/120 24/60 9 7200.4 920
adapter 80/120 50/60 2 7888.2 918

Did the adapter help?

On the same 120 cases, base Qwen3 produced 29 exact plans and the adapter produced 80, a difference of +51 cases. This is one seed and one narrow fixture, not an estimate of general chatbot quality.

Invalid outputs fell from 9 to 2; metric proposals on labeled refusal cases fell from 25 to 11. Adapter warm p50 was 7.89 seconds versus 7.20 for the base. Contract compliance improved, but the adapter has not met public-release acceptance; the demo continues to use the baseline. This does not establish equivalence to larger models.

Per-case comparison: 51 improved, 0 regressed, 69 unchanged. Case IDs and before/after outputs are included in results.json.

Separate challenge: zero previous-period baseline

These two EN/VI cases are reported separately from the main 120. September versus June has a zero previous total; percentage change must be null. On both EN/VI cases, all three models discarded comparison_period despite the comparison request. Protected arithmetic cannot fix this interpretation failure.

Candidate Exact plans / 2
baseline 2/2
qwen25 0/2
qwen3 0/2
adapter 0/2

p50 is measured on the workstation CPU, excluding public HTTP/Cloudflare; 18 warm subset runs are a small sample. RSS is a snapshot, not peak; baseline RSS is the standalone evaluator, not the entire LXC web service. Inference uses Q8_0 GGUF via llama.cpp; CPU training uses float32.

Reproduce the experiment

Download code, tests and instructions · Dataset JSON · Measured results JSON · Results CSV

Weights are not bundled; the script downloads pinned official revisions. No inference API key is needed. Follow the README to distinguish demo, benchmark and training.