LONTA AI — LEAD-SCORE AUDIT WORKSHEET Version: October 9, 2026 Guide: https://lonta.io/blog/ai-lead-scoring-home-services/ Service workflow: https://lonta.io/technology/ This is a reusable review template. It contains no client results or industry benchmarks and does not establish that using a score causes additional jobs. 1. DEFINE THE TASK BEFORE LOOKING AT RESULTS Account / trade: Outcome (booking, completed job, collected revenue, other): Scoring moment: Input cutoff (what was available at that moment): Follow-up window, and why it fits this task: Decision the score would change: Normal customer-response standards that remain in place: 2. FREEZE THE COMPARISON Model and version: Threshold for flagging an inquiry: Baseline rule and threshold: Training period: Test period: Were test records held out of training and tuning? How were repeat inquiries, households and related jobs kept from leaking? What changed in tracking, capacity, service area or weather during the test? 3. ACCOUNT FOR THE RECORDS Total eligible inquiries: Mature and resolved outcome records included in the four-count table: Pending (outcome window not finished): Unresolved / missing outcome: Unmatched / duplicate record: Excluded, with stated reason: Do not silently convert unknown or pending outcomes to no booking. 4. COUNT FOUR OUTCOMES ON THE SAME TEST RECORDS TP: flagged, outcome happened: FP: flagged, outcome did not happen: FN: not flagged, outcome happened: TN: not flagged, outcome did not happen: N = TP + FP + FN + TN: Accuracy = (TP + TN) / N Precision = TP / (TP + FP) Recall = TP / (TP + FN) Flagged review workload = TP + FP Missed observed outcomes = FN If a denominator is zero, report the metric as undefined, not zero. Repeat this table and these calculations for the baseline. Invented teaching example only: TP=8, FP=8, FN=2, TN=82, N=100. Accuracy=90%, precision=50%, recall=80%, flagged workload=16. An always-no-booking baseline has TP=0, FP=0, FN=10, TN=90: accuracy=90%, precision undefined, recall=0%, flagged workload=0. These are invented counts, not an HVAC benchmark or Lonta model result. 5. CHECK A PROBABILITY CLAIM Does the score actually claim to be a probability? Outcome and window associated with that probability: For each comparable score group, record: group / count / average predicted probability / observed outcome rate Keep group sizes and missing outcomes visible. A few records do not support a precise claim about an entire market. This template does not compute uncertainty. 6. WRITE THE BUSINESS CONCLUSION Which recorded outcomes did the model find that the baseline missed? Which useful opportunities did it miss? What review workload would the threshold create? Which errors or data limits remain unresolved? Who will review the proposed process or advertising-signal change? What evidence would justify a controlled test of acting on the score? Keep observed events, predicted events and incremental impact separate. The four-count table evaluates a prediction against observed outcomes; it does not show that changing the workflow will create extra jobs or revenue. METHOD REFERENCES, CHECKED OCTOBER 9, 2026 https://scikit-learn.org/stable/common_pitfalls.html https://scikit-learn.org/stable/modules/cross_validation.html https://scikit-learn.org/stable/modules/model_evaluation.html https://scikit-learn.org/stable/modules/calibration.html https://support.google.com/google-ads/answer/2998031 Reuse: you may copy and adapt this Lonta AI worksheet, including commercially, with attribution and a link to the guide. Preserve the distinction between an invented example and real performance. Third-party sources retain their own terms.