PRODUCT × ENGINEERING
Chatbot Engineering: A Practical Field Guide
CHATBOT ENGINEERING / FIELD GUIDE
From product analysis to data agents, model evaluation, fine-tuning and operations. Follow Commerce Assist, inspect measured evidence and try the EN/VI demo.
Try Commerce Assist ↗ · Đọc tiếng Việt
Choose your reading path
- Product: chapters 1 → 8 → 9 for value, costs and UX.
- Engineering: chapters 2 → 4 → 3 → 5 → 6 → 7, from contract to operations.
Nine chapters
- Analyze the Work Before You Build the Chatbot
Define chatbot jobs, risk boundaries, quality targets and a measured path toward fewer model dependencies.
- Build a Data Agent That Can Explain Its Numbers
Design a governed data agent with a semantic layer, read-only tools, tenant isolation and verifiable answers.
- Choose a Chatbot Model You Can Actually Operate and Train
Compare practical local model candidates, CPU routing and GPU fine-tuning without confusing size with production quality.
- Prove Whether a Smaller Chatbot Is Good Enough
Build independent evaluation sets, score grounded answers and measure quality, cost and coverage together.
- Train the Behavior, Retrieve the Knowledge: A LoRA Playbook
Prepare supervised examples, configure an adapter experiment and decide whether fine-tuning earns its operational cost.
- Make the Chatbot Faster Without Quietly Making It Worse
Optimize context, output, caching, retrieval and inference using an end-to-end latency and token budget.
- Run a Private Chatbot as a Production Service
Plan tenant isolation, bounded tools, deployment gates, failure recovery and observability for a production chatbot.
- Price a Chatbot Product Around Real Unit Economics
Calculate cost per successful task, account for cache writes and capacity, and offer transparent client pricing.
- Design a Chatbot People Can Trust, Configure and Recover From
Specify answer and analysis workflows, safe configuration, accessible interactions and a practical demo walkthrough.
Experiments and evidence
Benchmark and training use author-created synthetic data with correlated templates. They do not establish equivalence to larger models or production customer quality. Generative planners are evaluated offline; the public demo uses the controlled baseline. Benchmark: four candidates completed. LoRA: completed, 64/64 steps.
| Candidate | Exact plans | Correct results / metric | Invalid outputs | p50 / ms | RSS / MiB |
|---|---|---|---|---|---|
| baseline | 120/120 | 60/60 | 0 | 0.043 | 25 |
| qwen25 | 8/120 | 7/60 | 37 | 12113.0 | 618 |
| qwen3 | 29/120 | 24/60 | 9 | 7200.4 | 920 |
| adapter | 80/120 | 50/60 | 2 | 7888.2 | 918 |
Did the adapter help?
On the same 120 cases, base Qwen3 produced 29 exact plans and the adapter produced 80, a difference of +51 cases. This is one seed and one narrow fixture, not an estimate of general chatbot quality.
Invalid outputs fell from 9 to 2; metric proposals on labeled refusal cases fell from 25 to 11. Adapter warm p50 was 7.89 seconds versus 7.20 for the base. Contract compliance improved, but the adapter has not met public-release acceptance; the demo continues to use the baseline. This does not establish equivalence to larger models.
Per-case comparison: 51 improved, 0 regressed, 69 unchanged. Case IDs and before/after outputs are included in results.json.
Separate challenge: zero previous-period baseline
These two EN/VI cases are reported separately from the main 120. September versus June has a zero previous total; percentage change must be null. On both EN/VI cases, all three models discarded comparison_period despite the comparison request. Protected arithmetic cannot fix this interpretation failure.
| Candidate | Exact plans / 2 |
|---|---|
| baseline | 2/2 |
| qwen25 | 0/2 |
| qwen3 | 0/2 |
| adapter | 0/2 |
p50 is measured on the workstation CPU, excluding public HTTP/Cloudflare; 18 warm subset runs are a small sample. RSS is a snapshot, not peak; baseline RSS is the standalone evaluator, not the entire LXC web service. Inference uses Q8_0 GGUF via llama.cpp; CPU training uses float32.
Reproduce the experiment
Download code, tests and instructions · Dataset JSON · Measured results JSON · Results CSV
Weights are not bundled; the script downloads pinned official revisions. No inference API key is needed. Follow the README to distinguish demo, benchmark and training.