Performance measurement

Measure what happened, not what the assistant generated.

A useful scorecard connects conversation activity to verified next steps. It also makes skipped records, failed handoffs and incomplete evidence visible instead of burying them in one success number.

Define the denominator before the result

Specify the account, date range, lead sources and eligible conversation set. Paused officers, excluded stages, stop requests and records already handled by a person should not automatically be counted as unanswered assistant failures.

Keep distinct outcomes distinct: generated reply, accepted send, delivered message, borrower response, recorded appointment, attended consultation, application and funded loan. One event does not prove the next. If delivery or attendance evidence is missing, label it unknown rather than silently treating it as success.

Build a small evidence-led scorecard

Build a small evidence-led scorecard
MeasureEvidenceCaution
CoverageEligible inquiries and unresolved casesSeparate exclusions and incomplete history
Response speedInquiry time to a defined relevant response eventShow slow cases as well as an average
Conversation qualityContext use, relevance and human overlapReview a consistent sample, including failures
AppointmentsVerified bookings and separately verified attendanceDo not count a link or suggested time as booked
EconomicsComplete costs for the same period and scopeDisclose bundled charges and omitted labor

Avoid unfair before-and-after claims

A new campaign, different lead source or expanded team can change outcomes independently of the assistant. Compare equivalent scopes and disclose remaining differences. A few promising conversations can guide a pilot but do not establish a conversion-rate improvement for every officer.

Read the exceptions: a repeated question can harm quality even when a text was delivered. A lower subscription bill can still be expensive if the officer must repair every handoff. Check outcomes alongside contact boundaries, opt-outs and the time required for human review.

Keep a review record people can reproduce

  1. Document the period

    Use one date range for expenses, eligible inquiries and outcomes.

  2. Retain outcome evidence

    Keep the relevant send, appointment and attendance records within approved systems.

  3. Record uncertainty

    Show missing delivery or attendance evidence instead of filling gaps with estimates.

  4. Review changes

    Compare the same definitions after configuration or workload changes.

Questions before you choose

Performance measurement: common questions.

What is the best metric for mortgage AI?
There is no single sufficient metric. Combine eligible coverage, relevant delivery, conversation quality, verified appointments, attendance and complete costs.
Are generated messages the same as delivered messages?
No. Generation, send acceptance and delivery are distinct states. Report the evidence you actually have.
Should every record in a pipeline count as an unanswered lead?
No. Define eligibility and separate pauses, exclusions, contact preferences and completed human replies from unresolved eligible conversations.
Can a small pilot prove funded-loan ROI?
Not by itself. Funded-loan attribution needs verified downstream outcomes, adequate observation time and a credible comparison with other contributing activity.

Test the conversation before you activate outreach.