Practical AI · Published 9/29/2026 · By Eric Njuguna
Put an Operating Model Around Your LLM
Production LLM features need bounded inputs, observable decisions, recovery paths, and clear human review—not just a prompt that worked in a demo.
A language model can make a convincing first demo. Then the first customer asks a question outside the happy path, a source document changes, or the model returns an answer that sounds certain and is wrong.
The production question is not only “Can the model answer?” It is “What does the surrounding system do when the answer is incomplete, unsupported, slow, or unsafe to act on?”
Bound the job
Give the model a task with a defined input and output. If the task is to classify an incoming request, define the allowed labels and a way to say “uncertain.” If the system uses retrieved documents, retain which sources informed the answer and make missing evidence visible.
Structured output helps a service validate shape; it does not prove the content is true. Validate both the response format and the business rules that must hold before the next step.
Keep consequential actions reviewable
Let the model draft, classify, or summarize where that helps. Require human approval before consequential actions when the cost of an error warrants it. The reviewer needs enough context to disagree: the original request, relevant sources, the model’s output, and a clear control to accept or reject it.
Human review is not a magic safety label. If queues grow, response times stretch, or reviewers rubber-stamp suggestions, measure that behavior and change the workflow.
Observe the whole feature
Track more than latency and API errors. Collect a privacy-conscious record of which task ran, whether required sources were present, whether validation passed, whether a person changed the result, and whether the workflow completed. Sample outputs against a versioned evaluation set before changing models or prompts.
Set a fallback for timeouts and low confidence. That may be a conventional workflow, a request for more information, or a clear message that the system cannot complete the task. Returning an uncertain answer as if it were verified is not a fallback.
Let the use case earn its complexity
Compare the AI path with a simpler rule or existing process. Does the model improve an outcome that matters to users? Can the team explain failures, estimate running cost, and remove or replace the feature if it stops helping?
I treat AI as one component in a dependable system: useful when it has a bounded job, observable behavior, and a recovery path. That is how AI earns its keep.