By the Onsites AI team · Last updated · 5-minute read
NPS, CSAT and CES get sold as a metric family, but they answer three different questions about three different moments: CSAT asks how was this interaction (the support desk's mirror), CES asks how hard did the process make you work (the systems scanner), and NPS asks would you recommend us (the year-scale loyalty gauge). Small teams that adopt all three at once measure everything and move nothing; small teams that adopt them in the right order — CSAT first, CES when the trend stabilizes, NPS when the sample size earns it — get each metric's full value. This guide defines the three without the marketing gloss, shows when each one lies at small sample sizes, and sets the adoption order for a desk of three to ten people. All three can run from a free workspace; the survey ritual carries the first two without any subscription.
The question: "How did we handle this?" (1–5) attached to resolved threads, read monthly as a trend, acted on through verbatims — the full methodology lives in the no-tool CSAT guide. What it measures: the desk's interaction quality, month over month — the trend responds to macro, handbook and policy fixes within weeks, which is why it starts first: it is the only one of the three that connects a change to a verdict inside a quarter. When it lies: person-level comparisons at small samples (one agent's 4.6 against another's 4.1 is noise wearing a leadership meeting); response-rate collapse below 10% (bias toward the angry); and decimal worship (±0.5 bands at 30 replies). Its honest use: the monthly fifteen minutes — verbatims aloud, one fix shipped, trend noted. CSAT's whole value is speed-to-action; that is also why it is the starter metric and not the finish line.
The question: "How easy was it to get your issue resolved?" (1–7 or 1–5, low = hard), sent as a quarterly pulse — the metric that finds friction CSAT's per-interaction lens misses: the refund that took three touches, the size question whose answer lives on a page the buyer couldn't find, the channel gap where buyers route themselves to a worse experience (the channel guide's evidence). What it measures: the buyer's work — effort correlates with churn more tightly than delight does, because effort is remembered and delight is assumed. When it lies: asked per-interaction it just duplicates CSAT; it needs quarterly cadence and systemic reads (cross-referenced to repeat contacts — the FCR lens) to measure process rather than mood. Its honest use: the quarter's friction audit — low-easy scores clustered on return flows, on hours, on one channel — each nameable, each fixable, each verifiable the following quarter. CES is the metric that tells you the desk's systems built their own walls.
The question: "How likely are you to recommend us?" (0–10; promoters 9–10, passives 7–8, detractors 0–6; the score is promoters minus detractors). What it measures: the relationship across the whole year — marketing, product, pricing and support together; support moves it slowly, which is precisely why small teams measure it last: a quarter of desk fixes cannot move a year-scale gauge. When it lies for small teams: at 40 responses the score's band is wider than the differences teams celebrate; below roughly 200 responses a quarter, NPS is a vanity number in a suit — the deflection lesson applies (metrics that flatter without samples are theater). Its honest use at small scale: a sampled semi-annual pulse read as a trend, with the detractor verbatims read as a ticket queue — never as a score to report. If the team cannot get 100+ responses per pulse, park NPS and let CSAT and CES carry the feedback load; it will still be there when the volume earns it.
Start with CSAT (an afternoon, every-fifth sampling, monthly ritual): it connects fixes to verdicts fastest and builds the survey muscle the others will use. Add CES in the second quarter: quarterly pulse, read against repeat-contact data; its friction points feed the same fix-list CSAT's verbatims write. Earn NPS: when sampling passes 100 responses per pulse, read it as the year-scale trend it is — until then it is decoration. The three traps to avoid: adopting all three at once (measure everything, move nothing); person-level dashboards (unfair at small samples and corrosive to teams); and — the one the industry sells hardest — treating the scores as grades for buyers rather than mirrors for the desk. The KPI scoreboard keeps all of this in proportion: CSAT is one number of twelve; the desk's systems — routing, response clocks, the handbook — are what the scores eventually read.
Because the three can blur, watch them read one buyer's year. A trading buyer, Acme: their March requote thread gets CSAT'd at 5/5 (the desk's interaction mirror: quote #1042, requote, proforma — fast and human). Their April refund question gets CSAT'd at 3/5 — the interaction was fine but the wait was not. Their quarterly CES pulse scores the same April experience 4/7 — not "were you pleased," but "you had to send three messages and wait a day between each; that is work." Their year-end NPS, at the volume only scale brings: a 9 — promoter, but the verbatim names April's waits. The three instruments, three verdicts, zero contradictions: the interaction shone, the process made them work, the relationship held. That is also the correct reading order for any desk: CSAT tells you how today went, CES tells you what today cost the buyer in effort, NPS — eventually — tells you whether the relationship survived the year. Three different questions, one buyer, no confusion once the cadences are kept apart (monthly, quarterly, twice-yearly) and no instrument is asked another's question.
One team's year, metric by metric. Q1: CSAT built (every fifth thread, 11–14 answers monthly); first trend — 4.1 → 4.2, verbatims name slow refunds and chat tone; both become fixes (the refund ladder, the draft review). Q2: trend 4.3 → 4.4; first CES pulse (one quarter, every 8th thread): 6.2/7 overall, but the return-flow cluster scores 4.8 — the self-service gap becomes the quarter's fix (the help center's returns article). Q3: CES rebound 6.6/7; CSAT stable; the team considers NPS, counts 35 responses and parks it — honestly. Q4: CSAT 4.5, CES 6.7, the desk's fixes visible in both; NPS rescheduled for the year the volume triples. The total spent on the ecosystem: zero subscriptions; the insights: two policy fixes, one help-center gap, one tone drift — each caught in the quarter it appeared. That is the metrics' real purpose at small scale: not a scoreboard for buyers, but an early-warning system the desk runs on itself, cheap enough to keep forever.
What is the difference between NPS, CSAT and CES?
They ask different questions at different moments: CSAT — "how was this interaction?" (monthly, per resolved thread, acted on via verbatims); CES — "how hard was the process?" (quarterly pulse, acts on systemic friction); NPS — "would you recommend us?" (year-scale loyalty, needs large samples).
Which should a small team start with?
CSAT: an afternoon of setup, an honest monthly ritual, and the fastest connection between fixes and verdicts. Add CES once the CSAT trend is stable (quarterly pulse). Park NPS below roughly 100 responses per pulse — below that, it flatters rather than measures.
When does each metric lie?
CSAT lies at person-level comparisons and small samples (±0.5 bands at 30 replies); CES lies when asked per-interaction (it duplicates CSAT); NPS lies below ~200 responses — a score without volume is vanity. All three lie when used as grades for buyers instead of mirrors for the desk.
Is NPS worth tracking for a 3–10 person team?
Not yet: a quarter of desk fixes cannot move a year-scale relationship gauge, and the sample size at small desks (tens of responses) makes the score noise. Park it, read detractor verbatims as tickets, and revisit when sampling earns it.
How do the three connect to desk operations?
CSAT verifies the interaction (the refund ladder, tone on drafts); CES scans the systems (touches per resolution, channel gaps, self-service articles); NPS, when earned, reads the year. All three feed the same fix-list the KPI scoreboard reviews weekly.