Insight

Lead scoring that sales will actually use

Most scoring models are abandoned quietly within a quarter. The ones that survive share a property: sales helped build them and can see why a score moved.

Published by Somnium Digital

The quiet failure mode

Lead scoring rarely fails loudly. Nobody announces that the model is wrong. What happens is that a salesperson works a lead the model rated low, closes it, and stops looking at scores. Then another does the same. Within a quarter the field is still populated, still on the dashboard, and no longer consulted by anyone making a decision.

This is worth naming because it means the usual metric — whether the model is in place — tells you nothing. The only useful measure is whether the order in which leads are worked has actually changed, and that is observable in the CRM if anyone looks.

The causes are consistent. The model was built by marketing without sales, it scores engagement rather than fit, its outputs cannot be explained, and it was never tested against deals that actually closed. Each is avoidable and each is common.

Fit and intent are different axes and collapsing them is the original error

A single score conflates two things that behave completely differently. Fit is whether this organisation is the kind of customer you can serve profitably: sector, size, structure, geography, the systems they run. It changes slowly and it is largely knowable at the point of enquiry. Intent is whether they are moving now: pages viewed, documents downloaded, a demo requested, a second person from the same company appearing.

Collapsed into one number, a perfect-fit account browsing casually looks identical to a poor-fit account reading everything. Those two leads need entirely different handling — the first is a nurture case for a business development conversation, the second is usually a competitor, a student, or someone who will never buy.

Two axes make the handling obvious, and the grid is what sales can actually act on. High fit and high intent goes to sales immediately. High fit and low intent goes into a relationship programme. Low fit and high intent gets a self-service route rather than a person's time. Low fit and low intent gets nothing, which is a legitimate answer.

High fit, high intent
Contact now, with a person, quickly. This is the only quadrant worth interrupting a salesperson for.
High fit, low intent
Relationship building over months. The mistake here is treating it as a lead and burning the relationship early.
Low fit, high intent
Self-service, documentation, a clear price if you have one. Enthusiasm is not qualification.
Low fit, low intent
No action. A model that never says no is not scoring anything.

Build it backwards from deals that closed

The right starting material is your last two years of closed-won and closed-lost opportunities, not a workshop about the ideal customer. The question is empirical: what did the accounts that became good customers have in common at the point they entered the pipeline, and what did the ones that wasted six months have in common?

This exercise is uncomfortable in a useful way. It regularly contradicts the stated ideal customer profile — the sector everyone believes is core turns out to have the worst win rate, or the deal size everyone chases turns out to have the longest cycle and the highest churn. Those findings are more valuable than the scoring model itself.

It also produces the fit criteria honestly, weighted by what actually correlated rather than by what sounds strategic. A criterion that appears in eighty per cent of wins and thirty per cent of losses earns its weight. One that appears equally in both is noise, however important it feels.

What to do when you have twenty deals, not two thousand

Most B2B businesses do not have enough closed deals for anything statistical, and pretending otherwise produces a model that is precise about nothing. This is the normal case and it has a reasonable answer: build the fit criteria from structured judgement rather than from correlation, and be explicit that is what you are doing.

The method is to have the two or three people who know the customer base best independently list the characteristics of the accounts they would most want more of, then reconcile the lists. Disagreements are the interesting part — they usually reveal that the business is serving two different segments and calling them one.

Then hold the model loosely and review it against outcomes every quarter. With small numbers the review is the model: after three quarters you have real evidence about which criteria predicted anything, and the model earns its precision gradually rather than claiming it on day one.

Explainability is what makes it survive contact with sales

A score a salesperson cannot interrogate is a score they will ignore, and they are right to. The practical requirement is that any lead can be opened and the reason for its rating read in a sentence: what fit criteria it met, what it did, when.

This is a straightforward argument against opaque predictive scoring for most businesses. A machine-learned score may be more accurate in aggregate and still be less useful in practice, because the person deciding whether to call cannot see whether the model knows something they do not or has simply mistaken a job applicant for a buyer.

It also matters for improvement. When a salesperson can see why a lead scored high, they can tell you the reason is wrong — that the download everyone treats as intent is actually read mostly by consultants. That feedback is how the model gets better, and an opaque model cannot receive it.

Scoring decays, and decay is where most models drift

Behaviour has a half-life. Someone who read three pages nine months ago is not in market now, but many models keep the points, and over time every long-standing contact accumulates a high score simply by existing. The result is a hot list dominated by the least current records, which is precisely backwards.

Decay does not need to be sophisticated: intent points expiring over a defined window is usually enough. Fit points should not decay at all, since an organisation's sector and size did not change because time passed — though they should be refreshed when the underlying data is.

The related discipline is resetting after an outcome. A lead that went to sales and was disqualified should not sit at a high score waiting to be routed again, and a closed-won account should score in a completely different system aimed at expansion rather than acquisition.

The handover is the part that gets skipped

Scoring only matters if crossing the threshold causes something to happen reliably. In practice this is where implementations quietly fail: the score updates, a notification fires into a channel nobody reads, and the lead sits.

The mechanics that work are unremarkable. Crossing the threshold creates a task with a named owner and a due time, not a notification. The record carries what the person needs before they call — what the enquiry said, what was viewed, which fit criteria matched. And there is a defined answer to what happens if the task is not actioned within the window, because otherwise nothing happens and no one knows.

Agreeing that the threshold is a commitment, in both directions, is the actual project. Marketing commits to what a qualified lead means; sales commits to working every one within a stated time and recording why it was disqualified if it was. Without the second half, the model has no feedback loop and will decay into decoration.

A first version worth shipping

The version to start with is smaller than most proposals. Five to seven fit criteria drawn from closed-won evidence or structured judgement. Three or four intent signals that genuinely indicate buying rather than browsing. A two-axis grid rather than a single number. A decay window. One threshold that creates one owned task.

Run it in parallel for a quarter without changing anyone's behaviour, and compare what it would have prioritised against what actually converted. That comparison is cheap, it is the only honest validation available before deployment, and it regularly saves a business from rolling out a model that would have deprioritised its best deals.

Then adjust and turn it on, with a standing quarterly review against outcomes. A scoring model is not a deliverable that is finished; it is a hypothesis that gets corrected. The businesses that get value from scoring are the ones that treat the review as part of the system rather than as a project that ended.

Questions

Do we have enough data for lead scoring?

Probably not enough for anything statistical, and that is normal. Build fit criteria from structured judgement, be explicit that is what you have done, and let quarterly review against outcomes earn the precision over time rather than claiming it upfront.

Should we use predictive AI scoring?

For most mid-sized B2B businesses, no. It needs volume you probably do not have, and it produces a number your salespeople cannot interrogate — which is the fastest route to a model nobody consults. Explainability is worth more than aggregate accuracy here.

Why two scores instead of one?

Because fit and intent behave differently and need different handling. A perfect-fit account browsing casually and a poor-fit account reading everything produce the same combined number and require opposite responses.

How do we know if it is working?

Not by whether it is implemented. By whether the order in which leads are worked has changed, and whether conversion from qualified lead to opportunity improved. Both are observable in the CRM if anyone checks.

Sales says the scores are wrong. Now what?

Take the specific examples, because they are usually right about something concrete — a download that signals research rather than intent, a job title that means something different in their sector. That feedback is the improvement mechanism, and a model that cannot receive it will not survive.

What does this cost to set up?

Quoted per phase after we have looked at your closed-won and closed-lost history. The analysis is worth doing regardless of whether you build a model, because it frequently contradicts the stated ideal customer profile in ways that change more than scoring does.

Where this sits in what we do

This article covers one decision inside a wider engagement. The solution page sets out how that engagement runs, what it includes and what it costs to find out.

Building or fixing a scoring model?

We will start with your last two years of closed deals and tell you what actually predicted a win — including if the answer is that your current model would have deprioritised your best customers.

Get in touch

Tell us what you are trying to change

Describe the problem rather than the service — the two frequently differ, and working out which is which is the useful part of a first conversation. We reply within one working day, and if it is outside what we do well you will hear that in the reply rather than after a call.

We use what you send to reply to you. Nothing else, and no list.

WhatsApp