JOURNAL
3. Choose a Chatbot Model You Can Actually Operate and Train
Compare practical local model candidates, CPU routing and GPU fine-tuning without confusing size with production quality.

On this page
Chatbot Engineering · Part 3 of 9 · Research checked October 7, 2026. Proposed designs and assumptions are distinguished from measured implementation results.

The best model for this project is the smallest supported candidate that passes the project’s quality and latency gates. There is no justified universal winner without an EN/VI benchmark on the intended tasks. Training convenience, inference speed and answer quality are different selection criteria.
My practical starting shortlist
| Candidate | Role in the experiment | Main consideration |
|---|---|---|
| Frozen multilingual encoder + classifier heads | Local intent routing | CPU-friendly; cannot generate answers |
| Qwen3-4B | Text-only LoRA baseline | Established causal-LM tooling; older comparison baseline |
| Qwen3.5-4B | Current small-model challenger | Validate hybrid/multimodal architecture support |
| Gemma 4 E4B IT | Independent small-model challenger | Validate processor, adapter and serving compatibility |
| A strong hosted model | Quality reference and optional escalation | External processing, variable cost and availability |
This is an engineering shortlist, not a leaderboard. I would start the first reproducible text training experiment with Qwen3-4B, then compare Qwen3.5-4B before committing production infrastructure. That recommendation reflects toolchain simplicity; it does not claim the older model produces better answers.
What the current model cards establish
Qwen3-4B provides open weights under Apache 2.0, a text-generation interface and documented thinking controls. Use its own chat template and test non-thinking operation for short structured tasks. A task that needs reasoning may require a different budget.
Qwen3.5-4B is also Apache 2.0 and has a newer multimodal/hybrid architecture. Do not copy a causal-LM training recipe across architectures blindly: verify the appropriate loader, attention kernels, supported quantization and adapter targets.
Gemma 4 E4B IT is an Apache 2.0 open-weight alternative with multimodal tooling. It is a current challenger worth testing, not interchangeable with Gemma 3’s older license or model loader. Multilingual claims on a model card do not establish Vietnamese quality for your business data.
Which is easiest to train?
For routing, a frozen encoder with a small classifier head is the easiest experiment on the hardware already used here. It requires far less compute than adapting a generative backbone and has a narrow, inspectable output space. The existing demo proves that this route can run locally, although its tiny evaluation set is not enough for production.
For generated answers, a small text-only instruction model with a supported Transformers/PEFT stack is the simpler starting point. Train an adapter rather than the entire backbone. First confirm that the base model can already follow the output schema and language requirements; fine-tuning is a poor rescue plan for a fundamentally unsuitable model.
Fit the actual hardware
The demo LXC has 4 GiB of RAM and CPU inference. It suits the existing routing experiment; it is not a credible multi-user generative service or a QLoRA training machine. The inspected laptop and that Proxmox host have no NVIDIA GPU. This says nothing about uninspected hosts elsewhere in the cluster.
Weight-only memory is roughly parameter count multiplied by bytes per parameter: four billion parameters at four bits are approximately two decimal gigabytes before quantization metadata. Runtime memory additionally includes KV cache, activations, framework overhead and concurrent sequences. Training adds gradients and optimizer state for trainable parameters. Weight size is therefore not a VRAM guarantee.
Do not apply the four-billion-parameter example to every model labeled “4B.” Gemma 4 E4B’s official card distinguishes approximately 4.5 billion effective parameters from eight billion including embeddings. Budget resident weights using the actual stored tensors and their precision, not the effective-parameter label.
Test the exact context length, batch size and quantization on the intended hardware. Record peak allocated memory and out-of-memory failures. Rent a short-lived compatible GPU for a measured experiment before purchasing a dedicated server; do not promise that “4B fits” means production concurrency is solved.
Choose with a scorecard, not a name
Compare schema validity, exact metrics, grounded answers, EN/VI quality, useful abstention, p95 latency, cost per successful task and operational complexity. Eliminate candidates that violate a critical permission requirement, regardless of their average score. Evaluate quantized and unquantized variants separately.
Keep provider-specific behavior behind an adapter, but preserve capability checks. An OpenAI-compatible endpoint does not guarantee identical structured-output enforcement, token accounting or tool semantics. Pin model revisions, tokenizer files and runtime versions once a candidate passes.
Hosted models are a reference, not a self-host plan
Use a hosted model to establish an answer-quality ceiling for the selected tasks, with permission to process the test data. OpenAI’s model-selection guidance supports evaluating quality, latency and cost against the actual workload. Its current pricing documentation also states that its fine-tuning platform is winding down and is not accessible to new users. Do not make a new customer fine-tuning plan depend on old tutorials describing universal availability.
Workshop: a model decision memo for Commerce Assist
Start by separating three resource budgets: inference memory, training memory and operating effort. A model that is easy to fine-tune may still be too slow at your required concurrency. A model with a strong public benchmark may require an unsupported kernel on your server. For this project, the decision memo should identify the supported hardware, the task suite and the person responsible for upgrades.
| Deployment context | First experiment | Why | Do not infer |
|---|---|---|---|
| Existing 4 GiB CPU LXC | Current encoder + classifier | Already demonstrated within this scope | That it can also host a fast generative chatbot |
| CPU development machine with more RAM | Supported quantized small generator | Offline functional evaluation | That interactive performance holds for many users |
| Compatible single GPU | 4B text baseline and one challenger | Bounded adapter and serving comparison | A guaranteed VRAM fit without measurement |
| Client demands strict local processing | Local model + explicit abstention | Preserves the processing boundary | Permission to escalate externally |
Compare candidates under identical operating conditions
Use one versioned evidence snapshot and tool catalog. Give each model the correct native chat template, but the same business instructions and output contract. Record whether thinking is enabled, the output cap and sampling settings. A comparison of one model with 8,000 context tokens and another with 2,000 is a comparison of configurations as well as models; label it accordingly.
Run the generator first without a router, then add routing as a separate ablation. Otherwise you cannot tell whether the cheap model improved or whether a router quietly sent its difficult cases elsewhere. Report local-path quality, escalated-path quality, overall quality and external-call rate. A “single local model” claim is misleading if the system needs a hosted model for half its accepted work.
Use a results sheet that cannot hide an important failure
candidate,revision,quantization,template_version,thinking
task,language,test_count,correct_count,clarify_count,wrong_accept_count
schema_valid_count,forbidden_tool_attempts,unauthorized_results
context_tokens,output_tokens,p50_ms,p95_ms,peak_memory_mb
concurrency,external_escalation_count,cost_per_success_usd
These are report columns, not measured results. Start with a hard gate for tenant isolation and permitted tools. Then compare task success and useful abstention. Only then optimize latency and cost. A weighted aggregate can be useful among eligible candidates, but it must not compensate for a forbidden disclosure with faster generation.
Consider an illustrative pair: A completes 92 of 100 tasks and B completes 91. If A’s one additional success is a formatting fix while B avoids a cross-tenant answer A produced, B may be the only eligible release. Conversely, one hundred synthetic questions still provide limited evidence for small quality differences. Preserve the individual failures rather than publish only the totals.
Estimate memory without mistaking the label for capacity
For a conventional attention model, a rough KV-cache estimate depends on layers, KV heads, head dimension, bytes per value and retained tokens, multiplied by two for keys and values and by concurrent sequences. Grouped-query attention and hybrid architectures change these terms. Use the deployed runtime’s measured allocation for a purchase decision; do not transfer a generic formula unchanged to Qwen3.5 or every Gemma variant.
Even a small increase in context can reduce the number of concurrent sessions that fit in memory. Evaluate at 1, 2, 4 and 8 active requests if those levels match your intended traffic. An out-of-memory event at four requests is a capacity finding, not merely a model-quality failure. Record rejected requests and queue wait as well as successful latency.
Make a decision you can reverse
Select one default generator revision, one adapter if it earned its place, and an explicit escalation policy. Keep the model gateway interface stable while checking capabilities at startup. Save the rejected-candidate report so a future upgrade can target a known deficiency. Re-run the locked workload when changing quantization, tokenizer or prompt template.
The deliverable is a decision memo saying: “We selected candidate X for these tasks and hardware because it passed these gates at this load; these failures remain; this is the fallback policy.” Until that report exists, Qwen3-4B is a proposed training baseline and Qwen3.5/Gemma are challengers, not established winners.
From this chapter to a runnable experiment
Benchmark and training use author-created synthetic data with correlated templates. They do not establish equivalence to larger models or production customer quality. Generative planners are evaluated offline; the public demo uses the controlled baseline. Benchmark: four candidates completed. LoRA: completed, 64/64 steps.
Reading path and measured evidence · Try the synthetic-data demo · Download example code
What this implementation demonstrates
Select on exact plans before comparing speed. The separate 20-case validation split selected Qwen3 for this pilot; held-out scores are reported separately. The baseline has an advantage on this narrow, author-designed contract. Its score does not establish general language understanding.
| Candidate | Exact plans | Correct results / metric | Invalid outputs | p50 / ms | RSS / MiB |
|---|---|---|---|---|---|
| baseline | 120/120 | 60/60 | 0 | 0.043 | 25 |
| qwen25 | 8/120 | 7/60 | 37 | 12113.0 | 618 |
| qwen3 | 29/120 | 24/60 | 9 | 7200.4 | 920 |
| adapter | 80/120 | 50/60 | 2 | 7888.2 | 918 |


Discussion
Comments are reviewed before publication. Your email is kept private.