← All posts

Measuring an AI Front Office: from conversations answered to jobs completed

Support reporting counts replies. These numbers count what became work — requests qualified, visits booked, jobs completed — plus the correction rate that gates autonomy.

Conversation reporting answers a question that stops short of the business. A workspace can reply to every message, resolve every thread, score well on response time, and book nothing. For a field-service company the outcome that matters is whether an inbound message turned into a visit somebody completed, so the Front Office report measures the funnel rather than the conversation.

The funnel: created, qualified, booked, completed

The first half covers qualification: requests created in the period, how many are still qualifying, how many reached ready for scheduling, and how many went to handoff. Those four say whether structured intake is working at all. A high handoff count is not automatically a failure — sending an ambiguous or unsupported case to a person is the correct outcome — but it should be explainable by the reasons listed further down the same panel.

The second half covers what happened after qualification: slots held while a customer decided, appointments confirmed, jobs created, and jobs completed. The panel used to stop at ready for scheduling, which is the point where a request could become work rather than the point where it did. Confirmed bookings are counted from when they were confirmed rather than from their current status, so a visit that happened and was later cancelled is not quietly erased from the conversion.

Counted on creation, so the same report answers the same question

Everything is counted on when a row was created inside the window rather than where it sits today. A booking made this month and cancelled next month belongs to this month either way. Without that rule the same report would answer a slightly different question every time it was opened, which is exactly the property you do not want in a number a team uses to decide whether to give an agent more authority.

Four derived figures sit underneath: the share of created requests that ended as confirmed bookings, the share that reached readiness, the median time from creation to first readiness, and the correction rate. The correction rate is the important one. It compares AI-extracted answers against the ones a person later overrode, which makes it the closest thing the product has to an accuracy signal on real traffic rather than on an evaluation set.

The correction rate is the promotion gate

That signal is meant to gate autonomy. The rollout plan makes promotion from Observe to Assist to Qualify conditional on measured behaviour, and a climbing correction rate means Ven is guessing rather than reading — the moment to fix a template or narrow a service, not the moment to grant a wider permission. Escalation reasons are read the same way: a reason rising up the list usually points at one policy or one badly worded question, not at the model.

Queue health and catalog readiness sit beside the numbers

Two health checks sit beside the funnel because nothing else in the product can raise them. Queue health reports pending, processing, and dead-lettered events plus the age of the oldest event still waiting, so a stalled worker is visible before a customer notices a confirmation that never arrived. Catalog readiness counts active services, how many define a required qualification field, and how many of those give Ven a question it may actually ask — because a service role armed over an empty catalog looks perfectly healthy and answers nothing.

Review the funnel weekly against one decision: promote, hold, or fix

Start from the two ends. Requests created tells you whether intake is finding real demand; jobs completed tells you what that demand became. Everything between them is diagnosis — a healthy total with a thin booked share is a scheduling or capacity problem, while a thin ready share with plenty of created requests is a template or catalog problem.

Read the correction rate before deciding anything about autonomy. Falling or flat corrections alongside a rising ready share is the pattern that earns a promotion from Assist to Qualify. Rising corrections mean the opposite regardless of how good the volume looks, because every correction is a person repairing a fact a customer already saw applied to their request.

Check queue health and catalog readiness in the same pass. Dead-lettered events or an oldest-pending age measured in hours means confirmations may not have reached customers, and a service role armed over a catalog with no askable services will produce escalations that look like model failure but are really an empty template list.

Frequently asked question

Which metric decides whether to increase AI autonomy?

The correction rate — the share of AI-extracted qualification answers a person later overrode — read together with the share of requests reaching ready for scheduling. Rising corrections mean the agent is guessing rather than reading, which is a reason to fix a service template rather than to grant a wider permission. Escalation reasons and catalog readiness explain most of the rest.