By the Onsites AI team · Last updated · 5-minute read
A three-person support team does not need a metrics suite; it needs a scoreboard. The difference: metrics suites measure everything and move nothing, while a scoreboard puts a dozen numbers on a wall that two weeks of work can actually shift. This guide picks those twelve for a tiny desk — the KPIs that respond to real changes in queue habits, macros, AI leverage and escalation discipline — and names the five to ignore, the vanity numbers whose only honest function is guilt. Every metric here is computable from the desk itself (Onsites' queues and reports, or any shared inbox with tags), needs no survey tool for the satisfaction line (that guide's ratings tag does it), and takes under ten minutes a week to review as a team.
1 — First response time: median, not mean (one weekend-old ticket wrecks an average), per channel, against the published benchmarks — email under two business hours, chat under two minutes. It is the KPI buyers describe in reviews and the one routing and drafts move first. 2 — Full resolution time: how long until the customer's problem is genuinely done — the number satisfaction follows; watch its tail (the 10% slowest) as much as the median. 3 — Backlog age: the age of your oldest open thread, measured daily — the single best predictor of a desk drifting into neglect; a team that watches backlog age never has Monday archaeology. 4 — Open at close: threads open at day's end, tracked as a trend — a team that closes the queue daily holds the number under twenty; the trend line rising says "we need help," before the complaints say it louder.
5 — First-contact resolution: the share of threads solved in one exchange — the metric that buys the team its time back (the FCR guide works the measurement); its enemy is the reply that answers words rather than problems. 6 — Reopen rate: threads reopened within a week — the honest audit of "resolved"; under 5% is the small-team floor, and every reopen is a review of what the first answer missed. 7 — CSAT: the ratings tag (0–5 or thumbs) attached to resolved threads monthly — a response rate of 15–25% on a small desk is honest sampling, and the number exists for the trend, not the decimal (the survey guide sets the ritual in an afternoon). 8 — Promise-kept rate: follow-ups and callbacks delivered when promised — the least-discussed KPI and, per the sequences guide, the one buyers feel most; the desk's dated follow-ups make it countable.
9 — Channel mix: the share of volume per channel, monthly — it tells you where the desk's hours go and, per the channel guide, where a channel gap is quietly pushing buyers to whatever inbox they can find. 10 — Busiest hour: the hour your volume peaks — staffing one extra overlap hour where the peak lives moves first response more than any tool. 11 — Macro and handbook coverage: the share of replies that drew on saved templates or the handbook — growing coverage is how a three-person desk absorbs growth without hiring (the handbook guide works the compounding). 12 — AI share: the share of replies the copilot drafted first — the leverage line that, at credit prices, costs about a cent per assist; watch it rise with review discipline intact (the quality lens keeps drafts honest).
Vanity one — total tickets ever: a number that only grows, rewards nothing, and says nothing; the shape of volume (line 9) matters, the odometer does not. Vanity two — replies per agent per day: speed theater — it optimizes for short answers, punishes the 40-thread sagas, and correlates with reopen rate in the wrong direction. Vanity three — admin activity: logins and status shuffles measure presence, which three-person teams do not need surveilled; the scoreboard reads the queue's outcomes instead. Vanity four — average time per reply: context-free — a 12-minute reply that closes a saga beats a 2-minute reply that reopens one; FCR and resolution time own this question. Vanity five — deflection without verified satisfaction: the metric AI marketing loves, and the only one that can flatter while buyers churn (the deflection guide explains how deflection lies without a CSAT floor). Ignore all five without guilt; the scoreboard's twelve carry the desk.
Watch one desk's scoreboard through a quarter to see why twelve beats fifty. A three-person team starts at: first response — median 5.2 hours; backlog age — 9 days; reopens — 14%; promises kept — they never measured it. Weeks 1–4: routing rules assign by channel, drafts go live, the morning list (oldest five threads) becomes a ritual; first response drops to 1.8 hours, backlog age to 3 days, reopens to 9% (the replies being reviewed found the macro that half-answered). Weeks 5–8: the handbook covers the top eight questions, the overlap hour moves to the channel-mix peak, promises-kept starts measured at 71%; it ends the quarter at 96%. Weeks 9–12: FCR climbs from 54% to 68% as macro coverage passes 60%, and the CSAT tag's trend (62% satisfied, 4.4/5) starts responding — one bad fortnight traced to a rewritten shipping policy before reviews said it. The scoreboard moved the desk because each number had a named lever; the team never opened a dashboard deeper than these twelve, and never surveilled anyone — the numbers were the team's own mirror, not a manager's lens. That is the entire philosophy of the small-desk scoreboard: few numbers, nameable levers, weekly motion.
The scoreboard works only as a ritual: same time weekly, ten minutes, the whole team. Pull the twelve (the desk's reports, or the tags' export — the metrics ritual stays honest only if pulling it costs nothing). Read four numbers aloud: first response, backlog age, reopen rate, promise-kept — the four that went wrong somewhere. One fix per week: a macro written, a routing rule adjusted, the overlap hour moved, the handbook's weakest article rewritten — small teams compound one fix weekly; suites that fix nothing move nothing. Red flags, instant: backlog age crossing a week, reopens over 10%, or promises slipping — those three override the ritual and get a same-day fix. The scoreboard's purpose is not surveillance; it is the three-person team's way of watching its own systems decay before buyers do it for them — and the reason a tiny desk can outperform a suite-covered one is precisely that it reads twelve honest numbers instead of fifty flattering ones.
Which KPIs should a 3-person support team track?
Twelve: first response time, full resolution time, backlog age, open-at-close count, first-contact resolution, reopen rate, CSAT, promise-kept rate, channel mix, busiest hour, macro/handbook coverage, and AI share of replies. Review them weekly in ten minutes; fix one thing.
Which metrics should small teams ignore?
Total tickets ever, replies per agent per day, admin logins, average time per reply, and unverified deflection. They measure effort or flatter the desk without improving outcomes — several actively mislead (reply speed against reopen rate; deflection without CSAT).
How do you measure CSAT without a survey tool?
A ratings tag attached to resolved threads and a two-question follow-up email per the CSAT guide — a 15–25% response rate on a small desk is honest sampling. Track the monthly trend, not the decimal.
What's the fastest-moving KPI for a small desk?
First response time: routing plus instant drafts move it in the first week, and it is the number buyers mention in reviews. Backlog age is the fastest early-warning: when the oldest open thread passes a week, the desk needs a fix, not a metric.
How often should the scoreboard be reviewed?
Weekly, ten minutes, whole team: read the twelve, say four aloud (first response, backlog age, reopens, promises), ship one fix, escalate the three red flags (week-old backlog, 10%+ reopens, slipped promises) same-day. Ritual beats dashboard.