By the Onsites AI team · Last updated · 4-minute read
AI hallucination is not a mystery; it is a defined failure — the machine states, fluently and in your voice, something that is not true. The response is not to fear the tools (the value is too large) but to interpose guardrails between the draft and the customer: a stack this guide walks through, from grounding to signature to audit. Onsites AI is built around that stack as an architectural choice rather than a settings page — the copilot drafts from the handbook and CRM, Harness answers carry their sources visible, nothing auto-sends — because a small business's name is not recoverable from a bad screenshot.
Precise definitions keep the guardrails aimed. Hallucination is asserted fabrication: the copilot stating a refund window nobody set, a price nobody published, a promise nobody made — in fluent prose, in your voice, with total confidence. It is not vague hedging (that is undertrained caution, and the fix is feed, not fear), and it is not the model being "creative" — it is the machine filling a gap in its grounding with something plausible. The cost accounting follows the stakes: an ungrounded draft caught at the reading step costs seconds; the same line sent costs the thread, the screenshot, and however much a public "actually, our policy says…" thread costs in trust — which is why the stack's order matters: grounding keeps confabulation unlikely, sources make it visible, the signature makes it harmless, and the audit makes it expensive to hide. Teams that interpose all four rarely meet the expensive case at all.
The machine speaks from what it is given; give it your truth. The copilot's grounding is the workspace's written reality — the handbook's five pages, the CRM record of the actual account, the thread's own history — and the hallucination risk falls as that feed tightens. Two operational consequences follow. Write the truth down: prices with numbers, refund windows with days, discount authority with names; an ungrounded model invents defaults, and the price it invented is the one that ships. Keep the feed fed: every Harness miss and every drifted draft is a documentation task in the update ritual — the system's confusion is your syllabus.
The single design lever against confabulation is visible provenance: answers that carry the line they stand on — the handbook section quoted, the CRM field cited, the thread message referenced. Sources change the conversation's arithmetic in three ways at once: the reviewer's job becomes checking the citation instead of doubting everything; the wrongness becomes visible (an answer with no source is speculation, on its face, before it costs anything); and the audit trail inherits substance — "on what basis did we tell customer 302 X?" resolves to a click, not a memory contest. A draft without a visible source is treated exactly like a junior agent's unverified claim: fixable before it ships, cheap if caught, expensive if not.
The guardrail that survives every vendor pitch: nothing auto-sends to a customer. The copilot proposes; a human reads — tone, stakes, the edge the machine can't smell — and their name goes on the send. This is the layer that converts hallucination from brand risk into editing time, and it is why the desk's AI pricing (drafts at 3–6 credits, tool runs at 24) exists: speed where stakes are low, a person where the brand is on the line. Keep the exceptions deliberately narrow and written down — the off-hours acknowledgement that promises only an update, the status line with a fixed number — and keep everything with discretion (refunds, disputes, pricing exceptions, anything a manager would need to explain) behind the signature, permanently. The signature is not caution; it is the product the customer is paying attention to.
Most guardrail effort is spent where the money is, and money threads earn a doubled protocol. Before send, always: the reading human checks three points — the number against the record (a refund computed is a refund owed), the promise against the handbook's exact rule, and the deadline against what the desk can actually deliver. The copilot's draft makes those three checks faster — the number and the cited rule are right there in the answer — which is the honest case for AI on money threads: not fewer humans, but a reviewable sentence instead of a blank page. Two standing rules close the layer. Any draft citing money and without a visible source is returned to draft, not edited forward — it is cheaper to regenerate grounded than to defend a guess. And any money answer the customer might forward to an accountant, a court, or their own procurement gets re-read by whoever can sign it, because the thread outlives the shift and the screenshot outlives the thread.
The audit trail closes the loop: who sent what, drafted by which model, grounded on which record — recorded on both the cloud and self-hosted deliveries, so the question "what exactly did we promise this account" has an answer that survives the people involved. Two drills keep the stack honest, per the scorecard habit on the monthly sample (ten real threads, drafts scored against sources — accuracy, tone, the next-step promise) and the post-change eval (any model update reruns the same ten threads; quality drift shows in the diff, not in customers). The two failure shapes worth a standing watch: the fluent invention (an answer with no source, or a source the record doesn't support) and the drifting voice — an AI that becomes slightly more apologetic or more generous than your policy each month. Both are caught by drill, and both are cheaper to catch on Tuesday than in a review thread.
How does a help desk keep AI answers from being wrong?
Four layers: grounding — the copilot reads only your handbook, CRM and thread history, never a general internet; visible sources on grounded answers so every line can be checked; a human signature on every outbound message, with nothing auto-sent; and an audit trail recording who sent what, grounded on which record.
What is "grounding," in one paragraph?
Feeding the model your workspace's written truth — handbook, CRM, history — so drafts compose from real facts instead of a general model's priors. It also makes the failure mode visible: an unsupported line is either an ungrounded draft (caught by the reading step) or a sourced answer whose citation doesn't check out (caught by source review).
Why is the human signature layer non-negotiable?
Because everything with stakes — refunds, disputes, exceptions — needs a name on it that can be asked, blamed and taught. AI proposes, human disposes; nothing auto-sends. The copilot's value is speed on the un-staked half, not presence on the staked one.
How do we catch drift before customers do?
Two drills, monthly: score ten real drafts against their sources the way QA scores tickets, and rerun the eval whenever the model changes. Watch for the two failure shapes: fluent lines whose citations don't check out, and a gradually drifting voice — both cheaper to fix on Tuesday than after a screenshots thread.
What does the audit trail record during all this?
Who sent what, grounded on which handbook section or CRM record — on both cloud and self-hosted deliveries. Audits resolve "who told customer 302 what, and based on what?" to a click rather than a memory contest.