By the Onsites AI team · Last updated · 5-minute read
CSAT is the support metric everyone wants and nobody small can afford — or so the survey-tool industry has convinced desks. The truth: honest customer satisfaction measurement needs a ratings tag, a two-question email and a monthly ritual, all of which exist in a free workspace and an afternoon. What it does not need is a survey subscription, a third NPS platform, or the false precision of a decimal point. This guide builds CSAT from the desk itself: the ratings methodology (tag on resolved threads), the two-question email that gets answered, the sampling math that keeps the number honest at small volumes, and the review ritual that turns verbatims into fixes. It also says plainly what CSAT can and cannot tell a small team — which numbers to trust, which to ignore, and why the trend answers the only question worth asking.
One — the ratings tag: every resolved thread gets rated; the desk's resolution flow asks "how did we do?" (1–5 or thumbs) either on the buyer's side of the thread or tallied by hand in the monthly pass — the desk's own reports or tags carry the count, no tool required. Two — the two-question email: sent with or just after the resolution to every fifth thread — "How did we do, 1 to 5?" and "What one thing would you fix?" — the second question is the survey's verbatim engine, and the sequence sends it automatically with the human's name on it. Three — the honest sampling ritual: every fifth thread (never "everyone," which exhausts buyers and biases toward the annoyed; never "the ones we choose," which flatters), tallied monthly — on a 200-thread month that is 40 surveys and, at typical response rates, 30–50 answers. The whole build: an afternoon, the free workspace, zero subscription.
Small numbers deserve honest error bars, and the math is simple enough to respect. 30 replies on 200 resolved threads is real sampling — but its band is wide: a 4.2/5 average from 30 replies carries roughly a ±0.5 band; the difference between 4.2 and 4.5 in a month is noise. Which is why the ritual reads the trend, not the decimal: three months of 4.1 → 4.3 → 4.4 is a signal; one month's 4.6 is not. The verbatims carry the weight: the one-thing-would-you-fix answers are the survey's actual product — a month's 25 verbatims tell you what to fix; the average tells you mostly that you are likable. The response-rate floor: 15–25% is the honest small-desk band; if responses fall under 10%, the survey is too long or too frequent — one question, one click, every fifth thread keeps it answerable. The bias to name: angry buyers respond most (their verbatims are gold), delighted ones respond politely, the lukewarm majority stays silent — which is exactly why the verbatims are read and the decimal is not.
It can tell: whether the desk's systemic fixes are felt (the quarter's macro and handbook work — the handbook's quality — shows up in the trend); where the friction concentrates (verbatims clustering on shipping, on tone, on one agent's handoffs); and whether a change broke something (the CSAT dip that follows a policy change is the trend earning its keep). It cannot tell: whether you are "good" (the decimal means nothing at small samples); who is the best agent (person-level sampling at small volumes is unfair theater); or whether loyalty is rising — that is NPS's question, and the metrics-comparison guide keeps the three apart. The discipline that matters: CSAT is the desk's mirror, not the desk's grade — a mirror you check monthly for fog, not a grade you publish on the website.
The survey's fate is decided by its wording, length and moment. Length: two questions, never more — the second ("what one thing would you fix?") is the verbatim engine and the only open text buyers reliably answer; a third question halves the response rate and adds noise. The moment: with or just after resolution — the buyer who just had the problem solved is the most answer-inclined person they will ever be; three days later they are someone else. The sender: a person — "Mia from the Onsites team," not "support@onsites.ai"; the template discipline applies to surveys exactly as to support replies, and the reply-to must be a monitored inbox because some verbatims start conversations. The subject line, honest: "How did we do?" beats "We'd love your feedback!" because it asks rather than begs. One click to answer: the rating is a link or a reply, never a login. Written this way, the every-fifth email reads as care rather than extraction — which is why the 15–25% response band opens, and why the verbatims arrive warm enough to act on.
The survey earns its keep as a ritual: pull the month (responses, verbatims, the average with its band); read the verbatims aloud — all of them, as a team, fifteen minutes; the fix-list writes itself (the shipping-policy change, the macro that half-answers, the tone drift on one channel); ship one fix and note which verbatims it answers — next month's trend is the fix's verdict; thank the outliers: a delighted verbatim deserves a reply, an angry one deserves a same-week thread — the survey is also a queue of its own. The red flags: response rate collapsing (survey fatigue — lengthen the gap), verbatims clustering on tone (the quality lens goes on the drafts), a month with three 1-star verbatims naming the same gap — that is not a metric, that is a ticket. The ritual's cost: fifteen minutes monthly and an afternoon's setup once; its product: the only feedback loop that reads the desk's own quality without subscribing to anything.
One desk's first surveyed quarter shows the shape. A three-person team, 200 threads monthly: setup — the ratings tag plus the every-fifth two-question email (40 sent, 11 answered in month one — 27%, honest). M1: average 4.1, verbatims — six mention slow refunds, two mention tone on chat. Fix: the refund ladder written (the refund guide's), chat drafts reviewed against tone. M2: 4.3, refund mentions drop to one, response rate 21%. M3: 4.4, the trend visible, the fix-list shorter — the quarter's remaining verbatims name a shipping-delay gap, which becomes next quarter's first fix (the delay-note playbook). The cost across the quarter: zero software, maybe 60 credits of drafted survey sends, an afternoon's setup and forty-five minutes of rituals. What the team bought with it: a mirror that caught a refund-policy hurt before it reached the reviews — which is the whole return on CSAT done honestly, small, and without a subscription.
Can we measure CSAT without a survey tool?
Yes: a ratings tag on resolved threads plus a two-question follow-up email ("How did we do, 1–5?" and "What one thing would you fix?") sent to every fifth thread. Setup is an afternoon on a free workspace; the desk's tags and reports carry the tally.
What response rate makes CSAT honest at small volumes?
15–25% is the honest small-desk band — 30 replies on 200 resolved threads. Below 10%, shorten the survey (one question, one click) and lengthen the interval. Read the trend monthly, never the decimal: at that sample size, one month's 4.6 is noise.
Should we survey every resolved thread?
No — every fifth thread. Surveying everyone exhausts buyers and biases toward the angry; cherry-picking flatters the desk. A fixed sampling rule is honest and keeps the ritual cheap.
What does CSAT actually tell a small team?
Whether systemic fixes are felt (the trend moves after macro, handbook and policy changes), where friction concentrates (verbatims), and when a change broke something. It cannot grade agents (sampling too small) or measure loyalty — that is NPS's question.
How should we act on the results?
Monthly, fifteen minutes: pull the month, read all verbatims aloud as a team, ship one fix, reply to the outliers, and treat three 1-star verbatims naming the same gap as a ticket, not a metric. The verbatims are the product; the average is only the wrapper.