Onsites AI is 100% free forever — 3 seats and 100 MB included. Start free →
Learning Center

A QA scorecard: review 10 conversations a week

By the Onsites AI team · Last updated · 4-minute read

THE SCORECARD: A CARD, NOT A SURVEILLANCE SUITE ASK OF EVERY REVIEWED THREAD: 1 real question answered, not the words 2 next step + date stated 3 buyer's vocabulary, not the desk's 4 tone earns a second message 5–8: accuracy • promise kept • one-thread complete • would YOU accept it? THE CADENCE: 5 threads/week, 15 min Friday, rotate reviewers one pattern named → one fix shipped THE AI BAR: drafts reviewed like any reply — same card, same weekly read the card keeps drafts honest, not feared QA is a mirror the desk holds up to itself weekly — or it is theater.

Quality assurance on small support teams has a bad reputation, earned honestly: at most desks it is either surveillance (a manager scoring agents against a spreadsheet nobody agreed to) or theater (a quarterly audit nobody reads). The version that works on a desk of three to ten is neither: it is a card — eight questions asked of five threads a week, rotationally, by the team itself — whose only product is one named fix per week. This guide builds that scorecard: the eight questions and why each earns its slot, the cadence that keeps it honest, the AI-review bar that keeps drafts trustworthy without turning review into suspicion, and the tone-and-completeness standards that separate a card that changes a desk from one that decorates a wiki.

The eight questions, and why each

One — the real question: did the reply answer the buyer's actual question, not the words on the surface?(the FCR half-answer test — the most common failure, so it leads). Two — the next step: does the reply state what happens next and when ("refunded by Friday," "viewing offer by noon")? A reply without a next step is a promise-less sentence. Three — the vocabulary: buyer words or desk words? ("shipped Thursday" vs "status updated to fulfilled" — the template voice carries it). Four — the tone: would the reply earn a second message, or merely end one? (measured by the honest reader, not a rubric). Five — accuracy: every fact checkable — price, date, policy — is right; wrong facts are reopens wearing correct grammar. Six — the promise: any promise made has a date, an owner and a follow-up scheduled. Seven — thread completeness: the whole thread read — the reply answering only the latest message mid-saga misses the earlier half. Eight — the acceptance test: would you, receiving this as a buyer, feel handled? One question, asked honestly, catches what rubrics miss.

The cadence: five threads, fifteen minutes, rotated

The card works only at a cadence small teams can keep forever. Five threads a week: not sixty — five; sampled blindly (last Tuesday's closes, every fourth thread) so cherry-picking can't settle in; the mix deliberately includes one chat, one email, one saga — the failure modes live in the mix. Rotate reviewers: this week's reviewer reads next week's fix; the card belongs to the team, and every agent reviews before every agent is reviewed — the rotation is what keeps the card a mirror rather than a manager's lens. Fifteen minutes, Friday: the five threads, the eight questions, aloud where possible; the product is the week's one named fix (a macro rewritten, a policy handbook-ed, a vocabulary translated) — one fix, honestly, beats a scorecard of forty noted. The red-flag override: a thread where a promise was made with no follow-up, or a fact was wrong to the buyer's cost, goes same-day — the card's cadence is for the trend; the flags are for now.

The AI bar: drafts reviewed like replies

The scorecard's most important modern job is keeping AI drafts trustworthy without making review distrustful. The same card applies: a copilot-drafted reply reviewed through the eight questions is exactly as good as the review — the copilot discipline (human eyes, human send, human name on it) makes drafts a first draft, and the card is how the desk checks the draft-quality line. Patterns beat verdicts: when five reviewed threads show the same draft-failure — half-answered invoices, tone that reads terse in chat, dates hallucinated forward — the card's product is the pattern fix (handbook example written, draft instruction updated, template voice corrected), not a note on one thread. The trust boundary, kept explicit: the AI share of replies (the leverage line) rises only alongside a card that keeps catching its failures — a desk whose drafts go unread is a desk whose QA is about to find them; the card is the contract between speed and trust, and it is why drafts at credit prices are worth their cent.

What QA is not, on a small desk

It is not a productivity metric: replies per hour, time-per-reply, activity counts — the vanity list — have no slot on the card; a card that grades effort turns readers into checked boxes. It is not person-level scoring: at five threads a week, per-agent averages are noise with consequences; the card reviews threads, names patterns, and fixes systems — individuals learn from the weekly read, not from a leaderboard. It is not the customer's voice: CSAT (the survey mirror) reads what buyers felt; the card reads what the desk did — they are different mirrors for different questions, and the desk needs both (the card catches the half-answer the buyer hasn't noticed yet). It is not permanent: the card's questions evolve — a quarter of clean accuracy scores retires question five's weekly slot in favor of whatever the reopens now name. The card that never changes is a card that stopped reading.

A worked quarter of the card

One desk's scorecard through three months. Q1 weeks 1–4: the card's five reads surface three patterns — invoice replies half-answer (the real question is "why different from the quote"), two promises lacked follow-ups, one chat reply read terse. Fixes: the invoice macro rewritten around the quote-in-thread answer (the threaded pattern), the follow-up rule dated, chat drafts warmed. Weeks 5–8: the reads go clean on invoices and promises; a new pattern — shipping replies that answer the latest message of a saga and miss the first half — trains the thread-completeness habit; the handbook's delay-note example gets the worked fix. Weeks 9–12: five clean reads; the card retires accuracy's weekly slot (twelve clean weeks) in favor of the new reopens pattern (deposit questions). Quarter's product: three systemic fixes, one habit trained, one card slot rotated — twenty minutes weekly, zero subscriptions, no leaderboard. That is QA's whole honest shape at small scale: a mirror with a cadence, a fix per week, and a card that earns its place every quarter or changes.

Frequently asked questions

What should a support QA scorecard contain for a small team?
Eight questions asked of each reviewed thread: the real question answered (not the surface words), a next step with a date, buyer vocabulary not desk vocabulary, tone that earns a second message, factual accuracy, promises with owners and scheduled follow-ups, whole-thread completeness, and the acceptance test — would you, as a buyer, accept this reply?

How often should a small desk run QA reviews?
Five threads a week, fifteen minutes on Friday, reviewers rotating weekly. Sample blindly (every fourth closed thread, mixed across channels) so cherry-picking can't settle in. The weekly product is one named fix — more than that is theater, none is stagnation.

How should AI-drafted replies be reviewed?
Through the same eight-question card: drafts are first drafts with human eyes and a human name on the send. Watch for patterns across the week's reads — repeated half-answers or tone drift mean a handbook example or draft instruction needs updating, not a note on one thread.

What should QA never be on a small desk?
Not a productivity metric (no time-per-reply or activity counts), not person-level scoring (five threads a week makes per-agent averages noise), not a substitute for CSAT (the card reads what the desk did; surveys read what buyers felt), and not permanent — clean quarters retire slots in favor of new patterns.

What does a working QA cadence actually produce?
One systemic fix a week and a card that evolves: rewritten macros, handbook examples, authority granted, vocabulary translated. Twenty minutes weekly compounds into the queue habits buyers feel — response times, fewer reopens, and drafts that earn their credit prices.

Create your free workspace →  See the pricing