← back to portfolio

Building a QA grid that argues with the dashboard

Quality framework ยท a full evaluation and scoring pipeline for a client that doesn't exist

Meridian Fleet Systems is fictional. I made up a client so I could show a whole framework end to end without touching anything covered by an NDA. The company does not exist. The grid, the math, the sample data and the reasoning are real work.

The problem with the obvious metric

Meridian sells fleet maintenance software to small trucking outfits. The program I built this for is their inbound demo booking desk, fourteen bilingual reps whose whole job is one call. Qualify the lead, book a demo with an Account Executive. They don't close and they don't quote.

The number leadership watches is demos booked. You can game that number in about four seconds. Book everyone. Let the AE find out on the call that the prospect runs nine vehicles against a fifteen vehicle floor and has no budget. The rep's dashboard looks great, the pipeline fills up with junk, and nobody notices for two weeks.

So the first thing I decided was that the grid had to be allowed to disagree with the dashboard. A rep can book a demo and still score badly. A rep can turn a prospect away and score a 96. If the grid can't produce that result, you've just measured booking rate twice and given it a fancier name.

Weighting

Five pillars, eighteen criteria, everything scored 1 to 5. Discovery carries the most weight, and that's the part doing the pushing back:

PillarWeightWhy it sits there
A. Discovery & Qualification30%The thing the headline metric misses, so it gets the biggest share.
B. Sales Execution25%Value framing, honest product talk, actually asking for the demo.
C. Compliance & Documentation20%Disclosure, consent, and CRM notes the AE can work from.
D. Communication & Rapport15%Matters, but it's the easiest thing to score generously and the least predictive.
E. Next-Step Integrity10%Scored on its own, separately from whether a demo got booked.

Communication sits low for a reason. It's the pillar evaluators inflate, because a warm confident rep sounds like a good rep. There's a rep in the sample data built to expose that. Charming, high D scores, discovery so thin his qualified leads keep falling apart later.

Auto-fails skip the math entirely

Six triggers drop the total to zero. People argue with this one, so here's my defence. Most calls on this desk carry no compliance risk at all, and a handful carry a lot of it. Average those together and you end up with a nice looking 84 that quietly contains a consent violation. I would rather the number be ugly and honest.

CodeTrigger
AF-01Substantive discussion before the recording disclosure
AF-02Quoting a rate, term or discount the rep isn't authorised to offer
AF-03Presenting a roadmap feature as if it ships today
AF-04Continuing after an opt-out, or not logging a consent withdrawal
AF-05Falsified disposition, so inflating booked demos or burying a refusal
AF-06Disclosing another client's data, volumes or pricing

Two of those exist only because the desk isn't allowed to close. A rep who invents a discount is doing the AE's job badly, and it's the most tempting thing in the world to do when you can feel a deal slipping away from you.

The math, and the one clever bit

pillar_% = earned / (5 × applicable_criteria) weighted = pillar_% × pillar_weight × (100 / live_weight) total = sum(weighted) # 0 if any auto-fail fires

live_weight is the interesting part. Objection handling doesn't come up on every call, sometimes the prospect just says yes. Scoring that a zero punishes the rep for something the prospect did, so criteria marked NA fall out completely and the remaining pillars rescale to 100. Small thing, and it removes a huge amount of arguing from the feedback conversation.

What a scored call looks like

Two calls from the sample month. I picked these two because telling them apart is the entire reason the framework exists.

The first one booked a demo.

PillarRawWeighted
A. Discovery & Qualification16/2519.2
B. Sales Execution10/2012.5
C. Compliance & Documentation12/2012.0
D. Communication & Rapport12/1512.0
E. Next-Step Integrity5/105.0
Total0, AUTO-FAIL

On paper that's a 60.7. It scores zero, on two triggers, and they feed each other. At 04:12 the rep answered price pressure with "for a fleet your size we can probably do around fifteen percent off", which is a discount the desk has zero authority to hand out AF-02. Then dispositioned the lead as Demo Booked, Qualified, when the prospect had eleven vehicles against a fifteen vehicle floor and had mentioned twice that budget sat with a parent company nobody had talked to yet AF-05.

The coaching note on that one doesn't open with the discount. It opens with the qualification, because the discount got invented to rescue a deal that never should have been qualified in the first place. Fix the qualification and the discount problem mostly goes away on its own.

The second one turned the prospect away.

PillarRawWeighted
A. Discovery & Qualification25/2530.0
B. Sales Execution14/1523.3
C. Compliance & Documentation20/2020.0
D. Communication & Rapport14/1514.0
E. Next-Step Integrity10/1010.0
Total97.3, Exemplary

Outcome on that call: disqualified. The rep ran full discovery, got the pain down to an actual number (fourteen unplanned roadside events last quarter, roughly $2,300 each once you count the customer penalty), found a real deadline in a failed CVOR audit with an October re-inspection, asked for the demo, and then took it back once the fleet turned out to sit under the floor. Logged the October re-inspection as a nurture trigger on the way out the door.

B is 14/15 instead of 15/15 only because the booking slot never ended up being needed. Watch the denominator there too. Objection handling was NA, so the pillar is out of 15 and the weights rescaled around it. Best call on the desk that month, and it booked nothing.

Where this actually points

Scoring calls one at a time is the input. The output is what happens when you roll twenty evaluations across six reps together:

RepAvgAuto-failsWeakest pillar
L. Fontaine42.71E, Next-Step Integrity (52%)
J. Aliyev61.5–C, Compliance (47%)
M. Tremblay76.7–A, Discovery (59%)
K. Persaud85.8–D, Communication (76%)
R. Okonkwo90.8–C, Compliance (85%)
T. Nakamura96.8–D, Communication (95%)

Tremblay is the charmer. 76.7 overall, discovery sitting at 59%. Listen to one of his calls and he sounds fine, which is the whole argument for putting the weight where I put it.

The stats bug I walked into

First time I ran the numbers, the outlier detection flagged Fontaine and nobody else. Aliyev at 61.5, compliance at 47%, clearly struggling, went straight through unflagged.

The auto-fail did it. That zero pulled Fontaine's average down to 42.7 and pushed the standard deviation out to 20.3, and once the spread is that wide, a threshold of one deviation below the mean catches basically nobody. The failing rep was hiding the weak one.

Fix was to run the outlier baseline on scored calls only. The zeros stay in the reported average where the client should absolutely see them, they just come out of the distribution, since they measure a compliance event rather than how the rep is performing. Deviation drops to 16.1, both reps get flagged, and they get two different conversations.

You only trip over this by running real numbers through the thing. A grid that has never been executed against a full month of data can look perfectly fine on paper forever.

Calibration, and blaming the grid first

Monthly, three calls, every evaluator scores blind. Tolerance is 5 points on the total and 1 point on any single criterion. The sample set has a call where four evaluators landed 10.1 points apart, well out of tolerance, and the split traces back to two criteria:

CriterionRangeThe disagreement
A4 Authority & decision process2–4Prospect volunteered "I'd loop in ops". Does that count as authority surfaced, or does it only count once the rep gets a name and maps the approval path?
B1 Value framing tied to stated pain3–5Three capabilities named, one of them tied back to something the prospect said. Two evaluators scored the energy, two scored the tie-back.

Nobody was being careless there. Both criteria were worded loosely enough to support two honest readings, so my working rule is that a split past tolerance counts as a defect in the grid until proven otherwise. Rewrite the wording before you correct the evaluator. A4 now reads "mapped who else signs off and what the internal approval path looks like", which makes "I'd loop in ops" a clean 2.

Get this part wrong and scores stop being comparable between evaluators, which quietly poisons every trend you build on top of them, including that outlier analysis further up the page.

The pipeline

One Python file, standard library, no dependencies. The grid lives in JSON as the single source of truth and the readable grid document gets generated from it, so the document can't drift away from the math the way a Word file always does eventually.

$ python3 scoring.py scored 24 evaluations (20 live, 4 calibration) out/grid.md, out/summary.md, out/roster.csv, out/evaluations/*.md

Four outputs pointed at four different readers. A rendered grid for the client, a per-call feedback report written so it can go straight to the rep, a flat CSV with every criterion as its own column for whoever wants to pivot it, and a summary carrying rep trends, desk-wide pillar weakness, the five weakest criteria and the calibration spread.

That last table is the one that changes what anybody does on Monday. Across the sample month the weakest criteria are CRM notes accuracy at 3.60 and authority mapping at 3.70, both of them desk-wide, both sitting in the process rather than in any one person. Coaching six reps individually won't move either number. Training and workflow will, and pointing at the right one is most of this job.