I backtested our lead score against real won and lost deals. The number-one predictor of a win was not even in the model.
Everyone trusted the fit score. So I tested it against the deals we actually won and lost, and the signals that really separated them were behavioral, not firmographic, and the strongest one was not weighted at all.
The picture, before the words.
Every scoring model is a set of assumptions until you test it against reality. Ours assumed fit predicted wins.
But when I pulled the won and lost deals and measured what actually separated them, the fit score barely moved the needle, and several accounts we lost had scored in the 90s on fit. A model that scores your losses as highly as your wins is not scoring. It is decorating.
Solve the seam, wire the stack, split the drag from the judgment.
For every signal, I compared how often it showed up in won deals versus lost ones. Lift, not vibes. A signal that appears equally in wins and losses is noise, no matter how good it feels.
The strongest discriminators were behavioral, how the buyer actually worked, not the firmographics we had been scoring on. The single best one carried a 2.67x lift and was not weighted at all.
One behavioral signal plus an ICP threshold stacked to a 3.56x lift. The combination beat either alone, which is the whole argument for keeping signals separate so they can compound.
The fit score and CRM presence both hovered near 1.1x, barely better than a coin flip. I stopped letting them carry weight they had not earned.
What actually predicts a win, backtested.
I scored the model against real won and lost deals. Here is what discriminated, what I threw out, and how a top account adds up.
An ICP-fit base, the product-engagement escalator, and two live triggers. Capped at 100, top of the Strike band.
| Signal | Weight | Evidence |
|---|---|---|
| Runs Salesforce or HubSpot | .35 | 1.64x, won 52% vs lost 32% |
| Hiring sales right now | .30 | 1.55x, active sales-hiring signal |
| Sells the way our buyers do | .13 | 1.22x, CTA + pricing separation |
| Can reach a decision-maker | .12 | Outreach feasibility |
| Right size (25 to 2,000 FTE) | .10 | Won median 108 vs lost 46 FTE |
What was true after that was not true before.
The result, and the detail behind it.
Lead with what changed, then break it down.
The strongest behavioral signal separated wins from losses at 2.67x, and stacked with an ICP threshold it reached 3.56x, while the fit score everyone defended sat at 1.13x, barely discriminating at all. The model was rebuilt to weight what the data said predicted a win, not what we assumed did.
Before you trust a scoring model, backtest it against the deals you actually won and lost. Measure lift per signal. The ones that feel important and the ones that are important are rarely the same list.
The receipts, and what the field says.
From the people I built with.
“In four and a half years and eleven managers at Uber, Heath stood out. Data-driven, leads by example, and made sure every action had a purpose and the numbers to back it.”
“Heath is one of the rarest sales leaders I've learned from. He takes a holistic approach to go-to-market and understands every customer-facing department has a role in growth. I believe he's creating the blueprint for sales leaders to become holistic revenue leaders.”
Same method, on your workflow.
The playbooks and skills in Operator School are these builds, genericized so you can run the loop yourself. Get the next one in your inbox.