Measuring whether your AI chat is working
Most chat dashboards lead with the number that is easiest to make look good. Here are the ones worth watching.
ยท 1 min read
Deflection rate is the wrong headline
It counts conversations that did not reach a human โ including everyone who gave up. A bot that frustrates people into leaving scores beautifully.
Watch these instead
- Conversations that ended in a cart or an order
- Second questions โ did the visitor trust the first answer enough to ask again?
- Handoffs that arrived with context, as a share of all handoffs
- Questions it could not answer, grouped โ this is your content backlog
Read twenty transcripts a week
No metric replaces this. Twenty conversations tells you more about what your assistant is doing to your customers than any dashboard, and it takes fifteen minutes.
The number that decides renewal
Revenue from conversations, against what the tool costs. If it is not obviously positive after a month of real traffic, the content is wrong or the tool is.
Segment by channel before you conclude anything
Website chat, WhatsApp and helpdesk traffic behave nothing alike. Averaged together they hide the one that is broken. A shop whose overall numbers look fine is often carrying an excellent widget and a WhatsApp assistant nobody has read since launch.
Give it a month before you judge it
The first week measures your setup, not the assistant. Content gets fixed, questions shift, and the numbers move a lot. Judge it on the fourth week, against the same period before you had it.