Growth-stage B2B SaaS

Robbie Maltby

One operator, whole engine

Lead scoring fails when fit and intent collapse into one number

In brief: A lead scoring model needs to answer two separate questions: is this the right company and person, and are they showing signs of buying now? I score fit and engagement separately, require both for MQL, route qualified leads quickly, and use sales disposition codes to improve the model over time.

Every scoring audit turns up some version of the same lead.

It has a score of 95, built from twelve email opens, three PDF downloads and a webinar attendance. The company is a seven-person agency that’s never going to buy the product.

Meanwhile, a VP of Engineering at a 400-person target account sits at 40 because she visited the pricing page twice this week and has never clicked an email.

Sales calls the 95. Nothing happens. They try again the next week. After enough of those, they stop trusting the MQL queue.

Nobody has necessarily done anything wrong. The scoring model is just mixing two different questions into one number.

I separate them.

Fit and engagement are different things

Fit is how closely the person and company match the customers you actually want.

That might include:

  • industry
  • company size
  • geography
  • role or seniority
  • account type

It changes relatively slowly and usually comes from form data, CRM data and enrichment.

I take the criteria from voice-of-customer research and closed-won data rather than deciding that “VP at a SaaS company” sounds like a good buyer. The job titles, company types and segments that keep appearing in real opportunities give me something much more useful to work with.

Engagement is different. It’s behavioural and recent.

Pricing-page visits, demo-page visits, product documentation, webinar attendance and email clicks can all tell me something about current interest.

And unlike fit, engagement should fade.

Some platforms already make this distinction fairly naturally. Salesforce Account Engagement has separate grading and scoring concepts, with grade used for profile fit and score for activity.1 Marketo implementations commonly keep demographic and behavioural scoring separate before combining them into qualification logic.2 HubSpot supports multiple scoring properties, so the same model can be built there too.

I don’t care much whether the UI calls them score, grade or properties. I care that the system can answer both questions independently.

A working scoring model

Here’s an illustrative version for a mid-market SaaS product.

The numbers aren’t a template I’d copy between companies. I’d expect to change them once there’s enough closed-won and closed-lost data to see which signals actually correlate with buying.

SignalTypePoints
Industry on ICP listFit+20
100–1,000 employeesFit+15
Title matches buying roles from researchFit+15
Target geographyFit+5
Free email domainFit−10
Competitor domainFitDisqualify
Student / academic domainFitDisqualify
Pricing page viewEngagement+15
Demo page viewEngagement+10
Webinar attendedEngagement+10
Docs or integration pages, 3+ in a sessionEngagement+5
Email clickEngagement+3
Email openEngagement0
No activity for 30 daysEngagement−15

Then I set the MQL threshold across both axes.

For example:

Fit ≥ 40 AND Engagement ≥ 25

In Account Engagement, that might be expressed as a grade threshold plus a score threshold.

A high-fit person with very little current activity stays in nurture.

A highly active person who clearly doesn’t fit the customer profile doesn’t become an MQL just because they consumed a lot of content.

That distinction removes quite a lot of noise on its own.

I score email opens at zero.

Apple’s Mail Privacy Protection prevents senders from reliably using tracking pixels to determine whether an Apple Mail user actually opened an email.3 Opens are still useful for some aggregate reporting, but I wouldn’t use them as evidence that an individual buyer is becoming more sales-ready.

Clicks tell me more.

Where scoring usually goes wrong

Scoring is mostly arithmetic. That’s useful because it means I can inspect the rules and work out why somebody qualified.

The common problems are fairly predictable.

Fit and engagement are blended together. This is the one I see most. Enough low-value activity can compensate for very poor fit.

Nothing decays. If behavioural points only accumulate, somebody who spent an afternoon researching the product nine months ago can still look highly engaged today. I normally use decay rules or rolling activity windows.

Nothing subtracts. Unsubscribes, obvious non-buyers, competitor domains and other negative signals need somewhere to go. A scoring model that only adds points gets increasingly generous with age.

The model is never calibrated. I want to compare last quarter’s closed-won, closed-lost and recycled records against the scores they held when they became MQLs.

If most of the eventual customers never crossed the threshold, the model is missing something.

If most MQLs are being recycled for poor fit, the fit criteria are too loose.

The weights are hypotheses. Closed-won and closed-lost data are how I check them.

The model was built entirely in the settings screen. Scoring software makes it very easy to invent numbers. That doesn’t mean the numbers represent buying behaviour.

I want the fit criteria tied back to research and historical customers, and the behavioural criteria tied to actions that plausibly indicate progress towards a buying decision.

What happens to everybody who is not ready

Most real leads aren’t ready for sales when they first enter the database.

That’s what lifecycle automation is for.

I build the nurture inside the marketing automation platform and trigger it from lifecycle stages, list membership and property changes. I don’t want somebody exporting a CSV on Friday so another person can upload it into a newsletter tool.

A basic map might look like this:

StageGoalContentCadenceExit
SubscriberStay familiarEditorial newsletterWeekly or fortnightlyHand-raise or scoring threshold
New leadUnderstand the problem and establish fitProblem-led emails, case studies4–6 emails over 3–4 weeksMQL threshold or timeout to subscriber
MQL recycled by salesAddress why the lead was recycledContent matched to reason codeMonthlyRe-MQL on new engagement
Closed-lost / no decisionStay present for the next buying cycleProof, useful product updatesQuarterlyOpportunity reopened
CustomerAdoption and expansionOnboarding and behavioural triggersEvent-drivenn/a

The newsletter matters more than it sometimes gets credit for.

The same 95:5 logic I discussed in the positioning post applies to your own database. Most of the people in it aren’t actively buying at any given moment.4

For those people, I’d rather send something they actually choose to read than put them through an endless sequence of product emails.

Email is also the part of this system I enjoy most, and I suspect that shows in the results. On my most recent engagement, the subscriber list grew from zero to 16,000 and became a consistent source of demo demand over the following two years.

Routing starts when somebody qualifies

Once somebody crosses the MQL threshold, the scoring job is finished for the moment.

Now the routing needs to work.

I normally want the system to:

  1. assign the lead to the correct person or queue based on segment or territory
  2. notify the owner somewhere they will actually see it
  3. record when the MQL was created
  4. start the response-time clock
  5. require a disposition if sales accepts, recycles or disqualifies the lead

The response-time evidence is old, but useful if handled carefully.

A 2011 Harvard Business Review study looked at responses to test leads sent to 2,241 US companies. Thirty-seven percent responded within an hour, 24% took more than 24 hours, and 23% never responded at all.5

In a separate dataset of 1.25 million sales leads, companies that attempted contact within an hour were nearly seven times as likely to qualify a lead as those that waited an hour longer, and more than sixty times as likely as those waiting 24 hours or more.5

The often-quoted five-minute figures come from earlier Lead Response Management research associated with James Oldroyd and InsideSales.com, not from the Harvard study. That research found a very steep drop in contact and qualification odds as response time moved from five minutes towards thirty.6

I wouldn’t take the historical multipliers as a forecast for a modern B2B SaaS funnel.

The useful part is the direction: somebody who has just requested a demo is easier to reach while the problem is still in front of them than a day later.

For most teams, I start with a written SLA of first contact within one business hour and then adjust it around the actual sales motion.

The rest of the SLA matters too.

For example:

  • first attempt within one business hour
  • several attempts across more than one channel
  • a defined working window
  • mandatory disposition when the lead is recycled or disqualified

I usually keep the disposition list short enough that sales will actually use it:

  • bad data
  • no ICP fit
  • no response
  • timing
  • competitor
  • other, with a note

The automation can chase the missing field. It shouldn’t rely on somebody remembering to complete it at the end of the week.

Sales outcomes improve the marketing model

The disposition codes are useful to marketing because they tell me why the model was wrong.

If a quarter of recycled MQLs are coming back as no ICP fit, I look at the fit scoring and the campaign targeting.

If no response dominates, I look at lead sources, response time and whether the engagement threshold is meaningful.

If timing dominates, that’s a different problem. Those leads may belong in a better recycle programme rather than being treated as bad demand.

I also send useful downstream outcomes back to the advertising platforms where the integration supports it.

A demo request tells an ad platform that somebody filled in a form.

An accepted SQL or qualified opportunity tells it considerably more.

That’s why offline conversion imports matter. The campaign can eventually optimise against stages closer to revenue rather than treating every form submission as equally valuable.

The CRM is where I keep the stage timestamps, scores, routing outcome, SLA performance and reason codes. That’s also where the reporting layer gets the funnel numbers that matter financially.

On my last engagement, the wider lifecycle layer included more than 500 automated lifecycle and event emails. Alongside changes to scoring and routing, digital SQL conversion on the high-intent surfaces moved from 3.72% to 6.90%.

There was nothing particularly clever about the mechanism. Better-fit leads reached sales faster, and obviously poor-fit records stopped taking up the same amount of attention.

If sales has stopped trusting the MQL queue

I’d start with the last quarter of leads.

Pull every MQL, the fit and engagement data it had when it qualified, what sales did with it, how quickly they responded and the final disposition.

Then group the failures.

You normally don’t need a new marketing automation platform to see what went wrong. You need to know whether the bad leads are coming from loose fit criteria, inflated engagement, missing decay, campaign targeting, slow routing or something after the handoff.

Once that’s visible, the scoring model becomes much easier to change.

And when sales starts seeing a queue where the people actually look like buyers again, trust tends to come back with it.

I’m engaged with client work just now, so none of this is a pitch. But if you’re building growth-stage B2B SaaS and want this kind of work when capacity opens, tell me what you’re building. I read every note.

Sources

  1. Salesforce, “Scores and Grades in Account Engagement (Pardot).” help.salesforce.com (opens in new tab)

  2. Adobe, Marketo Engage scoring documentation. experienceleague.adobe.com (opens in new tab)

  3. Apple, “Mail Privacy Protection” announcement, June 2021. apple.com (opens in new tab)

  4. Dawes, J., “Advertising Effectiveness and the 95-5 Rule,” LinkedIn B2B Institute / Ehrenberg-Bass Institute, 2021. business.linkedin.com (opens in new tab)

  5. Oldroyd, J.B., McElheran, K., Elkington, D., “The Short Life of Online Sales Leads,” Harvard Business Review, March 2011. hbr.org (opens in new tab) 2

  6. Oldroyd, J.B. / InsideSales.com, “Lead Response Management Study,” 2007.

← All operating guides