← BUILDSSCORING EVIDENCE
── THE BUILD · THE WIN/LOSS BACKTEST

I backtested our lead score against real won and lost deals. The number-one predictor of a win was not even in the model.

Everyone trusted the fit score. So I tested it against the deals we actually won and lost, and the signals that really separated them were behavioral, not firmographic, and the strongest one was not weighted at all.

A Built GTM systemThe method, genericized so you can run the same loop on your own workflow.
One build like this in your inbox each week. Free, receipts only.
── THE SHIFT

The picture, before the words.

Behavioral signal, stacked with an ICP threshold3.56x
The strongest single behavioral signal2.67x
The fit score everyone trusted1.13x
Hover a bar to isolate it.
── THE RECEIPTS
2.67x
win-lift from the strongest behavioral signal
3.56x
that signal stacked with an ICP threshold
16 to 0
won-lost record of every account that reached Hot or Strike
1.13x
the fit score everyone trusted, barely discriminating
── THE PROBLEM

Every scoring model is a set of assumptions until you test it against reality. Ours assumed fit predicted wins.

But when I pulled the won and lost deals and measured what actually separated them, the fit score barely moved the needle, and several accounts we lost had scored in the 90s on fit. A model that scores your losses as highly as your wins is not scoring. It is decorating.

── HOW I RAN THE LOOP

Solve the seam, wire the stack, split the drag from the judgment.

01
Measure lift, not opinion

For every signal, I compared how often it showed up in won deals versus lost ones. Lift, not vibes. A signal that appears equally in wins and losses is noise, no matter how good it feels.

02
Find the predictors that were not in the model

The strongest discriminators were behavioral, how the buyer actually worked, not the firmographics we had been scoring on. The single best one carried a 2.67x lift and was not weighted at all.

03
Stack the signals that compound

One behavioral signal plus an ICP threshold stacked to a 3.56x lift. The combination beat either alone, which is the whole argument for keeping signals separate so they can compound.

04
Retire the flatterers

The fit score and CRM presence both hovered near 1.1x, barely better than a coin flip. I stopped letting them carry weight they had not earned.

THE RECEIPT · THE WIN-PREDICTOR BACKTEST

What actually predicts a win, backtested.

I scored the model against real won and lost deals. Here is what discriminated, what I threw out, and how a top account adds up.

THE SCORE, IN ACTION
Globex Industries
100
Strike
Strike threshold at 90
ICP fit59Product engagement+30New sales leader+6Competitor in stack+5

An ICP-fit base, the product-engagement escalator, and two live triggers. Capped at 100, top of the Strike band.

The fit weights, backtested against won vs lost.
SignalWeightEvidence
Runs Salesforce or HubSpot.351.64x, won 52% vs lost 32%
Hiring sales right now.301.55x, active sales-hiring signal
Sells the way our buyers do.131.22x, CTA + pricing separation
Can reach a decision-maker.12Outreach feasibility
Right size (25 to 2,000 FTE).10Won median 108 vs lost 46 FTE
The main pages of the backtest. The full model has the data-sourcing map and the overrides.
Open the full scoring model
── THE OUTCOME

What was true after that was not true before.

The result, and the detail behind it.

Lead with what changed, then break it down.

3.56x
the fit score sat at 1.13x; the behavioral stack hit 3.56x
WIN-LIFT, WEAKEST SIGNAL TO STRONGEST
BeforeAfter

The strongest behavioral signal separated wins from losses at 2.67x, and stacked with an ICP threshold it reached 3.56x, while the fit score everyone defended sat at 1.13x, barely discriminating at all. The model was rebuilt to weight what the data said predicted a win, not what we assumed did.

GET THE NEXT BUILD
One system like this in your inbox each week. Free, receipts only.
── STEAL THIS

Before you trust a scoring model, backtest it against the deals you actually won and lost. Measure lift per signal. The ones that feel important and the ones that are important are rarely the same list.

── GO DEEPER

The receipts, and what the field says.

THE MARKET ON THIS
── WHAT PEOPLE SAID

From the people I built with.

In four and a half years and eleven managers at Uber, Heath stood out. Data-driven, leads by example, and made sure every action had a purpose and the numbers to back it.
Cullen Trevino
Uber Eats
Heath is one of the rarest sales leaders I've learned from. He takes a holistic approach to go-to-market and understands every customer-facing department has a role in growth. I believe he's creating the blueprint for sales leaders to become holistic revenue leaders.
Arthur Castillo
── RUN THE SAME LOOP

Same method, on your workflow.

The playbooks and skills in Operator School are these builds, genericized so you can run the loop yourself. Get the next one in your inbox.

free · weekly · receipts only