A 90% accurate lead score can still miss the jobs that matter
An accuracy percentage does not tell you which good inquiries were missed. Ask what the score predicts, when it was made and what happened afterward.
TL;DR
- AI lead scoring for home services needs an outcome, a scoring time and a comparison. An accuracy percentage alone cannot tell you whether the score finds inquiries that become useful work.
- In the invented example below, a model catches eight of ten inquiries that become booked visits. A rule that predicts no bookings catches none. Both are 90% accurate.
- Test on records the model did not learn from, using only information available when the score was made. Give every inquiry the same time to reach the outcome.
- Keep unanswered, unresolved and not-yet-mature inquiries visible. Download the open audit worksheet and compare the score with a simple rule before using it to change a process.
What should AI lead scoring for home services predict?
A lead score should estimate a named outcome at a named moment. A score made when an inquiry arrives could estimate whether a visit will be booked within a defined window. A score made after a call can use information from that call. Those are different tasks.
For an HVAC shop, the distinction matters. An emergency repair may book quickly; a replacement inquiry may need a consultation. A booked visit, a completed job and collected revenue are different outcomes. A model evaluated on bookings has not thereby been evaluated on profitable work.
Write down the target before reviewing the chart. The worked example here uses a visit booked within 14 days of the inquiry. That is an invented review window, not a recommended deadline for every trade. Your shop should choose the window that fits the task, then apply it consistently.
The sold jobs and CRM hub covers the records behind marketing reports. This audit focuses on a predictive score, rather than a system that simply classifies what a completed call already contains.
How can two systems both be 90% accurate?
Take 100 invented inquiries with fully observed outcomes. Ten become booked visits within the chosen window; ninety do not. These counts describe a teaching example, not a client account or an industry benchmark.
A baseline rule predicts that none will book. It is right about ninety inquiries and wrong about ten: 90% accuracy. It identifies none of the inquiries that book.
Now apply an invented model and a fixed threshold for flagging an inquiry. The result is:
| Model decision | Visit booked | No visit booked | Total |
|---|---|---|---|
| Flagged for review | 8 | 8 | 16 |
| Not flagged | 2 | 82 | 84 |
| Total | 10 | 90 | 100 |
The model is also right ninety times: eight correctly flagged bookings plus eighty-two correctly unflagged non-bookings. Its accuracy is 90%.
But it found eight of the ten bookings. Its recall is 8 ÷ 10 = 80%: the share of actual bookings it caught. Its precision is 8 ÷ 16 = 50%: the share of flagged inquiries that booked. Two bookings were missed, and eight flagged inquiries did not book.
The scikit-learn evaluation documentation, checked October 9, 2026, explains these classification measures. The calculations above apply their definitions to our invented counts.
This model separates some useful inquiries from the rest. The table still cannot show that deploying it will increase bookings. It does not say whether a dispatcher can act on the flags, what the extra review costs or whether every inquiry already receives prompt attention.
Ask what decision the score changes. If it adds a review queue, count the workload as well as the bookings found. Keep the shop’s normal response standards in place while evaluating it. A low score does not establish that a real customer cannot become good work.
Did the model see the answer before making the prediction?
A model intended to score a new inquiry cannot use information created afterward. A booked appointment flag, an invoice or a later follow-up note would reveal part of the answer.
This problem is called data leakage. The scikit-learn guide to common pitfalls, checked October 9, 2026, explains how information unavailable at prediction time can make evaluation look overly optimistic.
Ask for the time each input became available. Keep the score, model version and input cutoff with the inquiry record. If the test reconstructs scores later, confirm that later information was excluded. A final CRM export can contain fields that did not exist when the phone first rang.
For this audit, add one column naming the prediction time and another naming the input cutoff. They make an impossible early prediction easier to spot without needing to read model code.
Are the test records actually separate from training?
Testing on the same records used to build a model cannot establish performance on new inquiries. The scikit-learn cross-validation guide, checked October 9, 2026, explains held-out evaluation and why the split must fit the data’s structure.
Home service records can include repeat calls about one job, returning customers and nearby time periods affected by the same weather. Ask how those relationships were handled. Putting one household’s first call into training and its follow-up into testing may create an easier test than predicting a new household.
A future-period test can answer a useful business question: how did a frozen model perform on inquiries that arrived afterward? Keep the model and threshold fixed for that evaluation. If either changes during the period, preserve the versions and report them separately.
Also give inquiries equal follow-up. An inquiry scored yesterday has not had the same opportunity to book as one scored two weeks ago. Mark records whose outcome window has not finished as pending. Do not count pending or unresolved outcomes as failures merely to fill the table.
Does a score of 80 mean an 80% chance of booking?
Only if that meaning has been defined and checked. Some scores are rankings or points; they are not probabilities.
For scores presented as booking probabilities, group comparable predictions and compare them with the observed booking rate. The scikit-learn calibration guide, checked October 9, 2026, explains this relationship between predicted probabilities and observed outcomes.
Keep the number of records beside each group. A few inquiries cannot support a precise claim about an entire market. A score may rank inquiries usefully while overstating their chance of booking; ranking and probability accuracy answer different questions.
Check whether the relationship changes by job type or season before applying one interpretation across the whole shop. The weather and operations review shows why demand and available crew capacity belong beside a marketing result.
Should predicted scores go back into the ad platform?
Keep predicted outcomes and observed outcomes separate. Sending an estimate does not make it a real booking or sale.
Google’s offline conversion import guide, checked October 9, 2026, describes importing later outcomes after ad interactions. Before changing the signal an account optimizes toward, ask which event is being imported, how it is matched and whether it is observed or predicted.
A platform-attributed conversion still does not prove that advertising created a job that otherwise would not have happened. The brand and customer-history worksheet keeps that distinction visible. A predictive model and an advertising incrementality test answer different questions.
What this doesn't cover
- A claim about Lonta’s model accuracy. Every count in the worked example is invented. No proprietary model, client lift or industry booking rate is reported.
- A universal threshold or minimum sample size. The right review threshold depends on errors, workload and job economics. This worksheet does not calculate statistical uncertainty or select a model for your account.
- Proof of additional jobs. Predicting an outcome is different from measuring whether acting on the score improves that outcome.
Do this in your account
- Download the AI lead-score audit worksheet. It is open text, with no signup.
- Name the outcome, scoring moment, follow-up window and decision the score would change.
- Freeze the model and threshold. Record the input cutoff and ask how training and test records were separated.
- Build the four-count table on mature, resolved test records. Keep pending, unmatched and unresolved records in separate counts.
- Calculate precision and recall alongside accuracy. Repeat the same calculation for a simple baseline rule.
- If the score claims to be a probability, compare score groups with observed outcomes and show the group sizes.
- Write down the workload, missed opportunities and unresolved questions before changing the client-facing process or advertising signal.
For the wider marketing plan, see HVAC marketing. For the service workflow around data and decisions, see Lonta AI’s technology page.
FAQ
Is 90% accuracy good for an HVAC lead score? It depends on the outcomes and errors. In this article’s invented example, predicting no bookings is already 90% accurate. Compare the model’s useful inquiries found, missed bookings and review workload with a baseline.
Can the score replace a dispatcher’s judgment? The audit does not establish that it can. A score needs a defined use and evidence that the resulting process helps the shop. Staff still need the customer’s situation, service availability and normal response standards.
How much data is enough? This worksheet sets no universal minimum. Show the number of inquiries and outcomes, keep incomplete follow-up visible and ask for an uncertainty assessment appropriate to the task. A small table is a review of recorded counts, not proof of reliable future performance.
Does a high score prove the advertising worked? No. It estimates an outcome; it does not estimate the extra work caused by an ad. A customer may have hired the shop through another route, and attribution alone does not settle that question.
Sources
Checked October 9, 2026. The worksheet is our review framework. All example counts are invented; the sources explain methods and platform capabilities, not the performance of a Lonta model.
- scikit-learn: common pitfalls, for information leakage and preprocessing.
- scikit-learn: cross-validation, for held-out evaluation and data-dependent splits.
- scikit-learn: model evaluation, for classification metrics.
- scikit-learn: calibration, for predicted probabilities and observed outcomes.
- Google Ads Help: offline conversion imports, for importing later conversion events.