By the Onsites AI team · Last updated · 5-minute read
The copilot writes; agent tools do. It is the difference between "draft a reply to this" (a text problem, 3–6 credits) and "check every open order above five thousand dollars for a missing reply, then list the accounts for follow-up" (a work problem — the model must reach into records, states and history, then compose, at 24 credits per run ≈ $0.24). Tools are also where the trust question shifts character: a draft can be reread; an operation must be governed. This guide is the practical constitution — what tools may do on a support desk, what they may not, and the paper trail that makes the delegation defensible to every audience a team ever owes an answer to.
Tools shine on tasks with three properties: they are defined (the success condition says what done means), they are cross-record (the answer lives across accounts, documents and threads, not inside one message), and they are tedious (a human could do it, and would hate it). The recurring families on real desks: the weekly sweeps — unanswered high-value orders, dormant accounts, quotes aging without follow-up; the compliance sweeps — threads mentioning refunds that lack a documented decision, documents stuck between states; the preparation jobs — summarize this account's last quarter of activity before the call; and the bulk-reading jobs — "find every case where we promised a date" before the ops review. What all of these share is the shape of inspection: tools read and assemble, and a human takes the list and decides. That division is not a limitation; it is the design that lets tools run at all.
Three no-lines keep delegation trustworthy, and they are worth stating as policy rather than vibes. No unreviewed outbound: a tool may assemble the draft, but a human's name signs every send — the copilot article's rule, now extended to operations. No irreversible action: refunds, document state changes, deletions and anything with a money footprint stay human; tools may gather the case for the decision, never make the decision. No scope beyond the seat: the run inherits the requesting seat's access, which makes seat hygiene a tool-safety control (the residency guide's joiners-and-leavers checklist applies here too). When a team is tempted past these lines — "let the agent just send it" — the honest question is not capability but answerability: who is on the hook when the machine is wrong? The audit trail answers "what happened"; only a human can answer "who meant it."
Tool runs inherit prompting discipline, with two additions that matter. Define done in the task: "list every account whose last order was 90+ days ago and whose last reply from us is missing" — a condition the run can check, not "find some accounts to follow up on." Vague tasks produce plausible lists that a senior must quietly redo; defined tasks produce lists that get worked. Name the output shape: "a table with account, last order, days since last contact, and the next step" turns a run into a handoff — and the table lands in the thread or the run record where anyone on shift can pick it up. The last rule is cadence: delegate rhythms, not one-offs. The first time you run "find unanswered high-value orders," it is a question; the fifth time, schedule it — recurring sweeps are where tools compound, and the desk's process matures exactly where they keep repeating. A good test before any task: could you hand this to a careful new hire on their second week, with a written page of what done means? If yes, the tools version will be cheaper, faster and every bit as auditable.
Every tool run records what it did: the task in, the records it touched, the result out — with sources visible, in the same audit fabric as the rest of the desk. That is what separates agent tools from the "AI somewhere in the pipeline" pattern that compliance teams rightly distrust: not that it is automated, but that the automation is legible. Two governance habits follow from the trail and take an hour a quarter. The run review: read a sample of the month's tool runs the way the scorecard habit reads tickets — are the right jobs being delegated, is anything drifting into deciding where it should be looking. The ledger habit: when the run reveals a recurring manual task ("we check this three times a week"), that is the next candidate for a canned flow or handbook entry — tools surface the workflow's shape, and the desk's process grows where the machine showed it.
Week one: delegate nothing that touches records; run the ask layer on everything else so the team learns what the machine knows. Week two: the first inspection runs — one sweep, one account-prep — reviewed against known answers. Week three: the list-based jobs graduate from "we asked" to scheduled — the unanswered-orders sweep, the dormant-accounts check — and the team stops doing them by hand. Week four: the run review; the gaps it shows become handbook entries, and the first genuinely new automation (a recurring sweep nobody had time for before) is born. Teams that follow this shape end the month with tools the desk trusts, because the trust was built by inspection first and operations second. Teams that invert it — automation first, review later — spend the same month learning to distrust the layer. The lesson is uncomfortable but reliable: the tool run you reviewed before you relied on is the one that will still be relied on next quarter.
Tools price at 24 credits — a quarter per run — and the honest arithmetic of ownership is against time. The five-person desk above (see the Harness guide's worked week) runs the ask layer daily and tools weekly; a quarter of that pattern is a modest credit-pack line, against hours of senior time returned every month — sweeps done on Sunday instead of never, preparation read in four lines instead of forty minutes. The self-hosted team prices the same runs against its own endpoint: fewer cents, more care. Either way, the pattern is the point: tools make the desk's memory operational, and the desk's humans stay the ones it works for.
What are agent tools in a support desk?
The operations layer of the desk's AI: a 24-credit run (~$0.24) that lets the model actually reach into workspace records — accounts, documents, thread history — to do cross-record jobs: weekly sweeps of unanswered high-value orders, dormant accounts, compliance lists, account prep. Tools read and assemble; humans decide and send.
Why the 24-credit price over the 3-credit draft?
It prices the job, not the text: a run that must find, cross-reference and compose across records is more compute than a draft. It is also the honest boundary marker — drafts are casual, runs are decisions — and the price keeps the expensive layer used deliberately.
What should agent tools never be allowed to do?
Three lines: no unreviewed outbound — a human's name signs every message; no irreversible action — refunds, deletions, state changes stay human; and no more scope than the requesting seat's access. Tools may look and list; deciding and acting remain human.
How is tool use audited?
Every run is recorded in the audit trail: task in, records touched, result out, sources visible. Reviewing a monthly sample of runs (like the QA scorecard habit) catches drift early — and recurring discovered tasks become handbook entries or canned flows.
How much do tool runs cost in practice?
A typical five-person team running daily look-ups plus weekly cross-record sweeps spends well inside a $20 pack per quarter. The honest comparison is against senior-hours returned: sweeps that happened at all, prepared in minutes instead of the afternoon they used to take.