JOURNAL

5. Train the Behavior, Retrieve the Knowledge: A LoRA Playbook

Prepare supervised examples, configure an adapter experiment and decide whether fine-tuning earns its operational cost.

LoRA pilot and held-out evidence. Train: 64 records · CPU float32 · Validation: 20 · select before test · Test: 120 EN/VI · unchanged · completed · 64/64 · 1123s. Not QLoRA · family split · repeated templates

Read with AI

Choose content to copy and paste into your AI assistant. Nothing is sent automatically. CMS content is converted to Markdown; original Markdown is used when available.

Chatbot Engineering · Part 5 of 9 · Research checked October 7, 2026. Proposed designs and assumptions are distinguished from measured implementation results.

LoRA pilot and held-out evidence. Train: 64 records · CPU float32 · Validation: 20 · select before test · Test: 120 EN/VI · unchanged · completed · 64/64 · 1123s. Not QLoRA · family split · repeated templates
LoRA pilot and held-out evidence. Train: 64 records · CPU float32 · Validation: 20 · select before test · Test: 120 EN/VI · unchanged · completed · 64/64 · 1123s. Not QLoRA · family split · repeated templates

Fine-tuning is most useful for stable behavior: selecting tools, following a schema, asking the right clarifying question and using a consistent response format. Retrieval is usually the better place for changing policies and product knowledge. Business arithmetic and permissions belong in application code.

Decide whether training addresses the failure

If the answer is wrong because retrieval supplied an outdated policy, repair ingestion and version selection. If it is wrong because the query used gross revenue instead of net revenue, repair the metric contract. If a capable base model repeatedly violates a stable output format despite clear prompts, a supervised adapter experiment is justified.

Compare four stages: unchanged base model, improved prompt and tools, adapter model, and quantized deployed adapter. Keep evidence and evaluation inputs fixed. Count training, evaluation and serving complexity in the decision; a small quality gain may not justify another artifact to operate.

Prepare supervised examples

A training record should include the user request, relevant authorized context and the desired assistant response. For a data agent, train a typed metric request rather than an invented numerical answer. Include ambiguity, unsupported metrics, unavailable evidence and appropriate refusal alongside successful tasks.

{"messages":[
 {"role":"system","content":"Return an approved metric request or ask for missing scope."},
 {"role":"user","content":"Show September 2026 net revenue in USD by channel."},
 {"role":"assistant","content":"{\"metric\":\"net_revenue\",\"period\":{\"start\":\"2026-09-01\",\"end_exclusive\":\"2026-10-01\"},\"group_by\":\"channel\",\"currency\":\"USD\"}"}
]}

Remove credentials, identifiers and private business content unless there is a lawful, approved training basis. Record provenance and license for each source. Synthetic examples need human checking; a teacher model’s confident mistake becomes a training target otherwise. Check provider agreements before using hosted outputs for distillation.

Use LoRA or QLoRA for the first experiment

LoRA trains a smaller set of adapter parameters while preserving the base weights. QLoRA uses a quantized frozen backbone to reduce memory requirements during adapter training. Neither means the resulting model has learned a secure database permission system. QLoRA research paper.

The PEFT documentation describes four-bit loading, NF4, double quantization and preparation for adapter training. Use a compatible compute dtype for the actual GPU and verify the installed backend. PEFT quantization guide.

A concrete experiment configuration

For the Qwen3-4B text baseline, start with a short context, batch size one, gradient accumulation and a modest adapter rank. The settings below are proposed initial conditions, not an experimentally optimal configuration or a completed training run.

base_model: Qwen/Qwen3-4B
revision: PIN_THE_REVIEWED_COMMIT
quantization: 4bit_nf4
adapter:
  rank: 16
  alpha: 32
  dropout: 0.05
  targets: [q_proj, k_proj, v_proj, o_proj]
training:
  max_length: 1024
  per_device_batch: 1
  gradient_accumulation: 16
  learning_rate: 0.0001
  epochs: 1
  gradient_checkpointing: true
  seed: 42
evaluation:
  development_each_epoch: true
  locked_test_after_selection: true

Inspect the model’s module names before selecting adapter targets. The same names may not cover a newer hybrid architecture. Start with a smoke batch, inspect the trainable parameter count and verify that loss is computed on the intended answer tokens. Check whether long examples truncate away the assistant response.

Keep the training template and serving template aligned

Use the model’s chat template and explicitly decide how thinking mode is represented. A mismatch can create unwanted reasoning output or malformed JSON at inference. TRL supports conversational supervised training and assistant-only loss when the template provides the required generation markers. Versioned TRL SFT Trainer documentation.

Pin an environment only after its smoke test passes on the target machine. Store the package lock, CUDA/driver details, tokenizer hash, dataset hashes and training arguments alongside the adapter. This article does not supply an untested dependency lock or claim a GPU run happened.

Evaluate before merging or quantizing

Run the independent task suite against the base and adapter. Check schema adherence, useful abstention, bilingual quality and unrelated regressions. Lower training loss is not evidence of better task completion. Stop or revise the experiment if it learns format but loses correct clarification behavior.

If merging an adapter for serving, preserve the original base and adapter artifacts. Evaluate the merged and quantized model again, using the exact production template and context limits. An inference optimization can change behavior even when the model name stays the same.

Prefer a reusable backbone before client-specific adapters

Use client configuration and retrieval for branding, metric definitions and documents. Introduce client-specific adapters only when repeated failures demonstrate a stable behavioral difference that configuration cannot solve. Versioning dozens of adapters creates rollout, isolation, routing and cold-load costs.

The first milestone is a reproducible, narrow improvement over the base model, supported by independent evaluation. Reproducing Jev’s proprietary training process or achieving frontier-wide equivalence is a different research objective and has not been demonstrated here.

Workshop: prepare an adapter experiment that answers a real question

The experiment question is: “Does an adapter improve Commerce Assist’s structured metric planning and appropriate clarification over the unchanged Qwen3-4B baseline?” It is not “Can we make training loss decrease?” Hold retrieval, tool execution and the metric catalog constant. Use bilingual examples with the same output contract; language is a data slice, not a reason to change the schema.

A proposed pilot dataset might contain 400 approved metric requests, 200 requests requiring clarification, 200 unsupported or forbidden requests and 200 document-answer formatting cases. These counts are design assumptions, not a validated minimum. Preserve independent development and test sets from different request families or authors. If production traffic is dominated by metrics, evaluate the real mix as well as a balanced diagnostic suite.

Audit examples before spending GPU time

For each record, check source provenance, the required fields, valid assistant JSON, correct semantics and no secret or client identifier. The earlier metric-only record teaches a raw plan; the mixed-task experiment needs an outcome envelope: metric_plan with a payload, clarify with a question and missing fields, refuse with a reason and supported alternative, or document_answer with text and citations. Validate each outcome with its own schema. Only metric-plan payloads go through the sample’s tool-contract validator; clarification, refusal and document answers must never be sent to the query executor. Tokenize with the selected chat template and inspect examples near the maximum length. A sample truncated before its target is not a useful supervised example.

Include contrastive task variants: an explicit September request, a missing-period request and a request for a forbidden customer-level export. Human reviewers should confirm that the clarification asks only for missing information and that the refusal explains the supported alternative. Avoid teaching the model to refuse all mentions of sensitive terminology; that would damage legitimate aggregate analysis.

A concrete Python training scaffold

The following scaffold targets the documented TRL 0.26 SFT API and a compatible Transformers/PEFT/bitsandbytes environment. It is a proposed GPU recipe, not a GPU-tested dependency lock. Pin the reviewed model revision through an environment variable. The datasets are JSONL records with prompt and completion fields, produced from reviewed examples. Completion-only loss avoids relying on an unverified assistant-mask chat template.

import os
import torch
from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, prepare_model_for_kbit_training
from trl import SFTConfig, SFTTrainer

name = "Qwen/Qwen3-4B"
revision = os.environ["MODEL_REVISION"]
dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16
tokenizer = AutoTokenizer.from_pretrained(name, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
    name, revision=revision, device_map={"": 0},
    quantization_config=BitsAndBytesConfig(
        load_in_4bit=True, bnb_4bit_quant_type="nf4",
        bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=dtype
    )
)
model = prepare_model_for_kbit_training(model)
model.config.use_cache = False
data = load_dataset("json", data_files={"train": "train.jsonl", "validation": "dev.jsonl"})
trainer = SFTTrainer(
    model=model, processing_class=tokenizer,
    train_dataset=data["train"], eval_dataset=data["validation"],
    peft_config=LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
        target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], task_type="CAUSAL_LM"),
    args=SFTConfig(output_dir="adapter", max_length=1024,
        per_device_train_batch_size=1, gradient_accumulation_steps=16,
        learning_rate=1e-4, num_train_epochs=1, completion_only_loss=True,
        gradient_checkpointing=True, bf16=(dtype == torch.bfloat16),
        fp16=(dtype == torch.float16), eos_token="<|im_end|>", eval_strategy="epoch", seed=42)
)
batch = next(iter(trainer.get_train_dataloader()))
assert (batch["labels"] != -100).any(), "No supervised completion tokens remain"
trainer.train()
trainer.save_model("adapter")

Render the prompt with the intended non-thinking chat template and ensure the completion contains the matching assistant ending token. Test that convention on a smoke batch before training. This scaffold is specific to the text baseline; it is not a drop-in loader for multimodal Qwen3.5 or Gemma. Syntax checking cannot prove GPU compatibility or useful learning.

Build one JSONL record with the correct template

The example below produces one formatting illustration, not a usable training dataset. Repeat preparation for independently reviewed records and keep development cases separate. The application adapter validates the outcome envelope, then unwraps metric_plan.payload before calling the metric validator.

import json
from transformers import AutoTokenizer
import os

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B", revision=os.environ["MODEL_REVISION"])
messages = [
    {"role": "system", "content": "Return a supported outcome envelope. Ask if scope is missing."},
    {"role": "user", "content": "Show September 2026 net revenue in USD by channel."}
]
target = {"kind": "metric_plan", "payload": {
    "metric": "net_revenue", "period": {"start": "2026-09-01", "end_exclusive": "2026-10-01"},
    "group_by": "channel", "currency": "USD"
}}
prompt = tokenizer.apply_chat_template(messages, tokenize=False,
    add_generation_prompt=True, enable_thinking=False)
completion = json.dumps(target, ensure_ascii=False, separators=(",", ":")) + "<|im_end|>"
assert len(tokenizer(prompt + completion, add_special_tokens=False)["input_ids"]) <= 1024
with open("example-record.jsonl", "w", encoding="utf-8") as handle:
    handle.write(json.dumps({"prompt": prompt, "completion": completion}, ensure_ascii=False) + "\n")

Inspect the actual smoke batch: prompt labels should be masked, completion labels retained and padding ignored. The assertion in the scaffold checks that some supervised tokens survive; it does not replace checking correct boundaries on several long records. Inspect both EN and VI examples and reject or restructure overlength records instead of silently losing their targets.

Run a small ablation, then decide

Keep the base-model result. Run one initial adapter configuration. If it improves development metrics, compare one lower learning rate or one lower rank, not dozens of trials against the locked test. Record effective batch size, tokenized example lengths, trainable parameters, peak GPU memory and elapsed time. Select using task correctness and useful abstention, not training loss.

After selection, run the independent test once and compare the deployed quantized artifact too. A failure analysis should identify whether the adapter learned a formatting shortcut, over-refused Vietnamese questions or memorized a date pattern. If the gain disappears on independent cases, keep the simpler base model.

Store the experiment as an auditable artifact

Save base revision, adapter weights, tokenizer files, package lock after a passing smoke test, data hashes, template, seed and evaluation packet. Keep raw private examples in approved storage rather than a public model repository. The acceptance deliverable is an independently supported improvement on the stated tasks, not the existence of a downloadable adapter.

From this chapter to a runnable experiment

Benchmark and training use author-created synthetic data with correlated templates. They do not establish equivalence to larger models or production customer quality. Generative planners are evaluated offline; the public demo uses the controlled baseline. Benchmark: four candidates completed. LoRA: completed, 64/64 steps.

Reading path and measured evidence · Try the synthetic-data demo · Download example code

What this implementation demonstrates

The pilot updates 1,146,880 LoRA parameters. Prompt labels are masked while completion and EOS receive supervision. The 64 training records repeat templates, so low training loss can reflect memorization. Promote an adapter only after held-out and representative customer evaluation establish quality, safety and latency.

Measured results on the synthetic test set
Candidate Exact plans Correct results / metric Invalid outputs p50 / ms RSS / MiB
baseline 120/120 60/60 0 0.043 25
qwen25 8/120 7/60 37 12113.0 618
qwen3 29/120 24/60 9 7200.4 920
adapter 80/120 50/60 2 7888.2 918

Did the adapter help?

On the same 120 cases, base Qwen3 produced 29 exact plans and the adapter produced 80, a difference of +51 cases. This is one seed and one narrow fixture, not an estimate of general chatbot quality.

Invalid outputs fell from 9 to 2; metric proposals on labeled refusal cases fell from 25 to 11. Adapter warm p50 was 7.89 seconds versus 7.20 for the base. Contract compliance improved, but the adapter has not met public-release acceptance; the demo continues to use the baseline. This does not establish equivalence to larger models.

Per-case comparison: 51 improved, 0 regressed, 69 unchanged. Case IDs and before/after outputs are included in results.json.

Discussion

Comments are reviewed before publication. Your email is kept private.

← Back to allĐọc tiếng Việt