JOURNAL
Building a Private Decision Layer for Chatbots on a CPU Homelab
A CPU-only chatbot and agent routing demo: trained classifier heads, bilingual tests, measured LXC latency, failure cases and a live homelab deployment.

On this page
A useful chatbot has decisions to make before it writes an answer: what the user wants, which information it may access, and whether a requested action needs approval. I built a small local decision layer for those questions, then deployed an interactive version on a dedicated homelab LXC.
The first deployed run classified 11 of 12 chatbot requests and 7 of 8 agent requests correctly before policy overrides. Warm model inference reached roughly 14 ms at p95 on CPU. Those are results from a tiny synthetic fixture set, not evidence of production reliability. The failures are part of the case study.
Try it: Private Decision Lab accepts English and Vietnamese requests and shows the selected route, raw model route, uncalibrated score and inference latency. Agent actions are simulations. The chatbot replies with templates, not generated prose.
What this has to do with Jev
In my introduction to Jev, I explored the distinction between making a typed decision and generating a conversational answer. That distinction is the starting point here. Jev evaluates questions against a state; it does not replace the component that writes an answer.
TypeSafe’s current model documentation says hosted Jev is not adapted with customer-specific fine-tuning or LoRA. This experiment therefore uses an independent open encoder and two classical classifier heads. It is neither a Jev checkpoint nor a reproduction of TypeSafe’s proprietary training process. I have not benchmarked this model against the hosted Jev API.
The aim is narrower: establish whether a small, private inference component can route a few real project categories before spending time on a larger model. Avoiding hosted inference calls removes provider token billing for this component. It does not remove hardware, electricity, maintenance or data-labeling costs.
Separate routing, answers and permissions

The deployed application has one Python process. It loads a frozen multilingual MiniLM encoder and two trained logistic-regression heads at startup. The encoder maps a request to a 384-dimensional embedding. A head estimates scores over a fixed set of labels. The model card lists Apache 2.0 licensing and a 128-token default sequence limit; that limit is retained and disclosed in the demo.
| Mode | Decisions | What happens next |
|---|---|---|
| Chatbot | Products, blog, architecture, collaboration, private information or out of scope | A curated template response, link or clarification |
| Agent routing | Blog draft, diagram, read status or approval required | A simulated skill recommendation; no action executes |
These fixed labels make the task inexpensive and inspectable, but restrict what the system can do. There is no private RAG index, conversational memory or generative answer model. Asking it to write a complete architecture will select a topic or ask for context, not produce a design.
More importantly, a label never grants permission. The deployed service has no production credentials, private documents or executable tools. A future agent would still need deterministic authorization, document access checks and human approval where appropriate. The classifier’s confidence is not a substitute for any of those controls.
What I actually trained
I wrote 48 bilingual chatbot training examples across six labels and 24 agent examples across four labels. I encoded them with the existing pretrained model, then fitted one logistic-regression head per task. The encoder weights stayed frozen. The heads used C=5, max_iter=1000 and random_state=42.
The evaluation fixtures contain 12 chatbot questions, eight agent questions and eight additional ambiguity/credential-request challenges. Evaluation strings are separate from the training strings. However, the same author wrote both sets, so this is not an independent blind evaluation. I have not tuned the heads on the held-out examples.
The deployment retains the same encoder revision and classifier files. An artifact manifest records SHA-256 hashes, and startup verifies them. The data fixture also has a recorded hash. The model revision is e8f8c211226b894fcb81acc59f3b34ba3efd5f42; the dataset hash is 4c6febee23e0aee6db415a43086660d6b8b75cfa14dedd63c07f03d4fbc40a74. Serialized classifier files are loaded only from trusted, operator-provisioned storage.
Measurements: laptop versus deployed LXC
The initial experiment ran on a Windows laptop with an Intel Core i5-1245U and 64 GiB RAM. The deployed application runs in an unprivileged Debian 13 container with two vCPU and 4 GiB RAM on the shared homelab. Both runs use CPU inference. The LXC evaluation loads the existing heads; it does not retrain them.
| Metric | Initial Windows run | Deployed LXC run |
|---|---|---|
| Chatbot raw classification | 11/12 (91.7%) | 11/12 (91.7%) |
| Chatbot raw macro F1 | 0.911 | 0.911 |
| Agent raw classification | 7/8 (87.5%) | 7/8 (87.5%) |
| Agent exact match after demo gates | Not measured in the initial spike | 6/8 (75.0%) |
| Warm chatbot inference p50 | 26.2 ms | 12.7 ms |
| Warm chatbot inference p95 | 35.0 ms | 14.1 ms |
| Process RSS after evaluation | About 1,299 MiB | About 1,113 MiB |
| Classifier artifact verification and model load | Not measured | 10.5 seconds |
| Hosted inference API calls | 0 | 0 |
The timing sample comprises the 12 chatbot evaluation requests and eight challenge requests, evaluated sequentially after a warm-up. It includes encoding and classifier inference, but excludes browser, HTTP, Cloudflare transit and cold startup. The two runs use different thread counts and shared system conditions; they are not a controlled hardware comparison. Process RSS is a snapshot, not peak memory or the same metric as systemd’s cgroup memory.
A character-based TF-IDF plus logistic-regression baseline classified eight of the 12 chatbot questions correctly in the initial run. The embedding model did better on this small fixture set. The sample is too small to claim that advantage will hold for real traffic.
The failures are more useful than the percentage
“Sketch the approval pipeline” was classified as approval required rather than diagram creation. The model associated the approval vocabulary with an action policy, although the user was asking to draw that policy. This is the difference between recognizing keywords and understanding the requested operation.
“Giúp tôi với” — “Help me” — was routed to products. The top-two margin rule did not catch it. There was insufficient evidence for that topic, but a model can still prefer one of its available labels. A clarification label, representative ambiguous data and a calibrated abstention policy need more work.
“Vẽ sơ đồ microservices” received the correct raw diagram label but a very small score margin. The deployed rule changed it to clarification. That makes the raw classification correct and the final exact match incorrect at the same time. Reporting only the raw 7/8 agent result would hide the actual behavior users see.
Explicit credential requests are handled by a simple template-refusal override. It is a bypassable illustrative denylist, not a prompt-injection defense. The meaningful protection in this demo is that no private data or executable tools exist behind the model.
Identical chatbot totals also conceal two different outcomes. “Please disclose the production credential” was incorrectly routed to architecture by the raw classifier and corrected to private information by the refusal override. “Giá vé máy bay đi Hà Nội bao nhiêu?” — asking the price of a flight to Hanoi — had the correct raw out-of-scope label, but the margin rule changed it to clarification. One correction and one abstention leave the exact-match total unchanged.
Scores are signals, not permission or proof
The demo exposes the maximum classifier score and the difference between the top two scores. A margin below 0.15 asks for clarification. That threshold was chosen as a demonstration heuristic, not fitted on an independent calibration split. Displaying a percentage does not turn it into a probability that the answer is correct.
TypeSafe’s confidence documentation similarly explains confidence as a statistic derived from a distribution and recommends testing thresholds against the domain and action risk. The demo’s margin is a different statistic, not an implementation of Jev’s confidence contract.
For a real application, I would measure both coverage and error: how often the system acts, how often it abstains, and which mistakes remain among accepted requests. A route that opens a documentation page can tolerate a different error profile from a route that publishes content or changes production infrastructure.
Deploying a small public experiment responsibly
The service runs under a dedicated non-root account with read-only source/model storage, offline model-loading flags and a systemd memory limit. The origin accepts the web port only from loopback or the existing tunnel connector. It uses no cloud inference key. The pinned runtime is Python 3.12.12, CPU PyTorch 2.14.1, sentence-transformers 6.1.0 and scikit-learn 1.9.1.
The API accepts at most 2,000 characters and a 16 KiB body. One inference runs at a time, with a bounded waiting queue. Per-client and global request limits constrain this public demo. The encoder still sees at most 128 tokens, so a long accepted request may be truncated; the interface discloses that limit.
# Inside the provisioned application directory
.venv/bin/python -m pytest tests -q
# Re-evaluate the fixed dataset, without training on it
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 PYTHONPATH=. \
.venv/bin/python ops/benchmark.py /tmp/decisionlab-results.json
# A typed request; agent actions remain simulated
curl https://decisionlab.luonghongthuan.com/api/decide \
-H 'Content-Type: application/json' \
-d '{"mode":"agent","text":"Are the services running?"}'
Application logs do not retain request bodies or chat history. Public requests still pass through Cloudflare before reaching the homelab; local inference does not mean the entire network path is private. Visitors should use public or synthetic text and avoid submitting sensitive information.
Verification covered both modes, input limits, model-unavailable responses, rate limits, offline loading, restart behavior, EN/VI browser flows and mobile layouts. User and model text are rendered as text rather than HTML. The full model service is separate from the WordPress container and production NanoClaw workflows.
What I would train next
The next improvement is a better dataset, not a larger checkpoint by default. I would collect consented, redacted requests from actual workflows, define labeling guidelines, use independent annotators and split by conversation or project. A blind test set should remain untouched while thresholds are calibrated on separate data.
Then I would compare this frozen-encoder baseline with an open decision model and, where data handling and API terms permit it, hosted Jev. Backbone fine-tuning would need to show a useful gain on the blind evaluation, not simply reproduce training examples. A local generative model for grounded free-form answers is another experiment with its own latency, memory and quality requirements.
This is the first implementation case study in From Jev to a Private Decision Model. Future articles will cover minimal state and atomic questions, calibration and abstention, private dataset design, model fine-tuning, and shadow-mode operation. Those are planned investigations, not completed experiments.
Try the live experiment, inspect the raw route and policy override, and start with a request where the model has a reason to ask for clarification.


Discussion
Comments are reviewed before publication. Your email is kept private.