America/Chicago
ProjectsJune 20, 2026

Designing an ML Lead Scoring Architecture People Will Actually Use

image
Nearly every B2B SaaS company starts with a hand-weighted points model. Demo request, twenty points. Whitepaper download, five. Title contains "Director", ten. It persists for a good reason: a sales leader can read it, argue with it, and predict what it will do. It is also fit to nobody's actual data. The weights encode what a room believed about buying behavior on the afternoon they were set. Replacing it with a model fit to closed-won outcomes is straightforward. Replacing it with a model people trust is the real work. Tiers, not scores. A continuous 0–100 score invites arguments about whether 73 is meaningfully different from 71. A small set of named tiers maps to actions instead: the top tier gets worked today, the bottom one goes to nurture. Reps act on tiers. They litigate scores. Weighted ensembles over a single model. Different signal families (firmographic fit, engagement recency, intent, product usage) behave differently and decay at different rates. An ensemble lets each family carry its own logic, and lets you retire a misbehaving signal without refitting everything. Explainability as a hard requirement. Every score ships with its top contributing signals. This is not a nice-to-have: it is the mechanism by which reps develop trust, and the mechanism by which you find out the model has latched onto something absurd. The first time a rep says "this is top-tier because of a careers page visit," you have learned something no validation metric would have told you. Feedback capture. If a rep disagrees with a tier, that disagreement needs somewhere to go. Rep overrides are training data and an early warning that the motion has shifted.
  • Leakage from post-conversion fields. A feature that only populates after a deal is qualified will produce a model with spectacular validation numbers and no predictive value. Audit every feature for when it becomes available relative to the decision point.
  • Fitting on a motion that no longer exists. If the go-to-market changed eighteen months ago, historical closed-won data is describing a different company.
  • Optimizing the metric instead of the outcome. A model that maximizes AUC on conversion can quietly prioritize small, fast deals over the enterprise motion the company is actually betting on.
  • Shipping without the sales conversation. A technically correct model that reps ignore has zero value. Adoption is part of the system, not a rollout step.
  1. Run the new model in shadow mode alongside the existing points model. Compare what each surfaces before anything changes for reps.
  2. Bring two or three reps in during shadow mode, not after launch. They will find the absurd signals faster than any validation suite.
  3. Cut over one segment at a time.
  4. Instrument tier-to-close rates from day one, so the retraining conversation is evidence-based rather than vibes-based.
If your scoring model is not being used, the problem is rarely the math. It is usually that nobody can see why it says what it says, and a model a rep cannot interrogate is a model a rep will route around.

Related projects

Governing AI Operations Inside a GTM Stack

Governing AI Operations Inside a GTM Stack

Agentic automation reaches revenue-path systems long before anyone writes a policy for it. The registry, write policy, and audit program that turn informal AI operations into a governed function.
Building a Unified Marketing Measurement Practice

Building a Unified Marketing Measurement Practice

How to combine multi-touch attribution, marketing mix modeling, and incrementality testing into one measurement system that survives contact with a finance team.