JOURNAL

7. Run a Private Chatbot as a Production Service

Plan tenant isolation, bounded tools, deployment gates, failure recovery and observability for a production chatbot.

Demo trust boundaries. Browser → Cloudflare → LXC 151 · API: body bounds, quotas, queue · Plan → validator → fixed query · No hosted fallback · synthetic data. A public demo does not prove production authentication

Read with AI

Choose content to copy and paste into your AI assistant. Nothing is sent automatically. CMS content is converted to Markdown; original Markdown is used when available.

Chatbot Engineering · Part 7 of 9 · Research checked October 7, 2026. Proposed designs and assumptions are distinguished from measured implementation results.

Demo trust boundaries. Browser → Cloudflare → LXC 151 · API: body bounds, quotas, queue · Plan → validator → fixed query · No hosted fallback · synthetic data. A public demo does not prove production authentication
Demo trust boundaries. Browser → Cloudflare → LXC 151 · API: body bounds, quotas, queue · Plan → validator → fixed query · No hosted fallback · synthetic data. A public demo does not prove production authentication

A production chatbot is an application with an uncertain component inside it. Models produce proposals; the surrounding service provides identity, authorization, bounded execution, persistent state and recovery. Self-hosting changes who operates that service, not whether those controls are needed.

Choose a deployment mode per data class

Private-local mode keeps inference and sensitive tools inside the approved boundary. Hybrid mode uses local processing for private data and a permitted external service for eligible requests. Hosted mode simplifies inference operations but requires an explicit data-processing and cost decision.

Make the active mode visible to administrators and enforce it server-side. If external processing is disabled, a failed local model must not silently fall back to a hosted endpoint. Explain an unavailable capability instead. Public Cloudflare transit also means a public self-hosted demo is not a network-isolated system.

Keep authority outside model output

Derive user and tenant identity from a validated session. Resolve connector credentials from a server-side vault. Give tools narrow capabilities such as “read monthly approved revenue,” rather than a general shell or database credential. The model’s JSON cannot expand that capability.

For a write action, render the exact target and expected change before approval. Bind approval to a request identifier, user, tenant, action parameters and expiry. If parameters change, require a new approval. Use idempotency keys so an accidental retry does not create duplicate effects.

Isolate data in every layer

Database rows, document indexes, conversation storage, exports and caches all need tenant boundaries. A correct database policy does not protect a vector index that returns another customer’s documents. Test authorization before any evidence reaches a model.

Avoid logging raw prompts or tool results by default. Capture request identifiers, task type, latency, token counts, status and sanitized errors. If a support investigation needs content, make collection controlled, limited and disclosed, with retention and deletion procedures.

Bound the service

Enforce input size, output size, tool duration, maximum tool rounds, queue length and per-tenant budgets. Limit concurrency based on actual peak memory and acceptable queue delay. A request timeout should cover the entire operation, not only the model call.

A service should become ready only after model artifacts and required dependencies are available. Separate liveness from readiness. Under overload, reject work before memory exhaustion and explain the recovery path. Use a circuit breaker for a failing connector rather than retrying it for every conversation.

Deploy reproducible artifacts

Keep the base-model revision, adapter hash, tokenizer, dependency lock, prompt revision and metric-catalog version in a release manifest. Verify artifact hashes before loading. Run the process as a non-root user with read-only code and model files, a bounded memory allocation and restricted network access.

An LXC can host a CPU router, API and UI, as the existing demo shows. GPU model serving may need a separate appropriately configured host. For a single-host product, document outage and backup recovery honestly; containerization does not supply high availability.

Observe product health and model health

Track task success, unsupported-request rate, useful clarification, customer feedback and failed exports alongside service metrics. For generative serving, observe queue wait, time to first token, generation latency, running requests and memory pressure. vLLM exposes production metrics for instrumenting its serving layer. vLLM production-metrics documentation.

Label metrics with bounded categories rather than raw user text or high-cardinality identifiers. Alert on sustained service degradation and task regressions. An average latency graph can look healthy while a small customer experiences repeated timeouts.

Test failure and rollback

Before a rollout, exercise a dead database, stale index, malformed plan, model crash, canceled request, expired approval and an exhausted budget. Verify that the UI shows a failed or partial state and does not imply the requested work completed. Check that a retry cannot replay a privileged action.

Promote through offline evaluation, shadow testing where authorized and a canary release. Keep the previous application and model artifacts available. Record rollback compatibility for conversation schemas and metric definitions; a model rollback alone may not undo an incompatible application migration.

What is production-ready here?

The current public routing demo has bounded input, queue and quotas, an isolated CPU service and no real action tools. Those controls support its demonstration scope. It remains a small classifier experiment with synthetic evaluation, not a validated client-facing analytics product.

A billable release would additionally need authenticated workspaces, real connector authorization, independent quality evidence, recovery exercises, commercial terms and an agreed support policy. Treat this chapter as the release checklist to implement and verify, not as a claim those capabilities already exist.

Workshop: define a request state machine before deploying

A Commerce Assist request should have explicit states: received, authorized, planned, validated, executing, completed, failed and canceled. A write proposal adds awaiting_approval and approved before execution. Store the request ID and policy version with state transitions. “The model said done” is not a valid completed state; completion follows a confirmed tool result and successful delivery of the permitted output.

Cancellation has two dimensions: the user no longer wants the answer, and the backend has actually stopped consuming resources. Record both. If the backend cannot interrupt a running query safely, suppress delivery after cancellation but continue accounting until the operation finishes. For future write actions, cancellation must never be represented as rollback unless the tool confirmed rollback.

Map boundaries to separate components

Component Responsibility Secrets/data available
Gateway Session, tenant scope, budget admission Validated identity; no broad database password in prompts
Planner Propose supported structured work Allowed catalog and authorized context
Tool executor Validate and execute fixed operations Scoped connector credentials
Answer renderer Explain a verified result Permitted result envelope
Usage ledger Meter accepted/completed work Request ID, task class, usage; no raw conversation required

These are logical boundaries, not a requirement for five microservices. A single process can enforce them with separate interfaces and limited credentials. Split deployment only when isolation or scaling requires it. Adding distributed services introduces timeouts, authentication between components and more failure modes.

Write an incident runbook with concrete decisions

Suppose the model service stops responding. The gateway rejects new generation work, preserves submitted text and marks existing requests unavailable; deterministic navigation or already-valid metric templates may remain usable. It must not switch to an external provider in a private-only workspace. The operator checks readiness, memory pressure and artifact verification before restart, then uses a known fixture to confirm recovery.

If a tenant leak is suspected, stop the affected connector and cache path first. Preserve sanitized request identifiers and release versions, revoke suspect cached artifacts, and investigate the authorization boundary. Do not keep the service open simply because the model’s average score is good. Define notification and investigation responsibilities with the client before an incident, not during one.

Use a minimum observability envelope

{
  "request_id": "opaque-id",
  "workspace_bucket": "bounded-internal-category",
  "task_class": "metric_compare",
  "release": "application+model+prompt-manifest",
  "queue_ms": 120,
  "tool_ms": 80,
  "input_tokens": 1200,
  "output_tokens": 180,
  "outcome": "completed",
  "external_processing": false
}

These values illustrate a schema, not a collected trace. Keep private identifiers in access-controlled audit storage; monitoring labels should remain bounded. Include error category and retry count for failures. A usage ledger needs a durable idempotent update, so retrying a billing event cannot double charge a client.

Practice recovery before promising an SLO

Define backup contents: configuration, connector references, metric catalog, approved conversation storage if enabled and release manifests. Model weights may be restored from a verified artifact store; do not assume public downloads will always be reachable. Keep secrets in the approved secret system and document restoration separately.

Specify recovery-point and recovery-time objectives as proposed service commitments, then measure a restore drill. A successful backup job is not evidence that restoration works. Run a fresh instance, restore the catalog and policy, execute tenant-isolation fixtures, then validate the permitted answer path.

Release gate and handover

The handover includes the runbook, capacity report, rollback instructions, support owner and known limitations. A read-only single-host pilot can be commercially useful if its availability and maintenance window are honestly described. Do not market it as highly available merely because the LXC restarts automatically.

The acceptance exercise is a deliberately interrupted request followed by recovery without duplicated work, unintended external processing or a misleading success state. That exercise tests the production service around the model, where many real failures occur.

From this chapter to a runnable experiment

Benchmark and training use author-created synthetic data with correlated templates. They do not establish equivalence to larger models or production customer quality. Generative planners are evaluated offline; the public demo uses the controlled baseline. Benchmark: four candidates completed. LoRA: completed, 64/64 steps.

Reading path and measured evidence · Try the synthetic-data demo · Download example code

What this implementation demonstrates

The public baseline runs on an LXC with shared quotas, bounded bodies and an execution gate. Canceling an HTTP request does not release capacity before its background worker finishes. These are real resource controls; customer authentication, RBAC and business audit remain separate production work.

Discussion

Comments are reviewed before publication. Your email is kept private.

← Back to allĐọc tiếng Việt