How to Measure a Conversational AI Agent (and What to Do With the Data)
I tried to buy a flight using an AI agent and couldn't. Then I tried Amazon's Rufus and Netflix: same story. Here's what to track in a conversational agent, and what to actually do with that data.
I needed to book a flight to Lisbon a few weeks ago. I opened the airline booking platform I use most and noticed they'd just launched an AI agent. I figured, why not, let's try it.
I typed: I need a flight to Lisbon in August, help me find the best price.
We went back and forth a bit, I gave it more context, and the agent came back with a few flight options.
I asked it to book the flight for me. It couldn't, which I expected.
But when I tried to book the first flight by clicking the link it shared, the page said "sorry, something went wrong." I tried the second flight. Same thing. By the third one, it didn't even redirect me: "there was a problem redirecting you to checkout. Please try a new search."
I asked what was going on. It told me it was having technical issues.
Fifteen minutes later, I hadn't bought anything. And I was pretty frustrated.
It wasn't just that platform
I wondered: was this just that one platform? I started testing agents from other well known companies.
Most hadn't even launched one yet.
Of the ones that had:
Amazon has an agent called Rufus. I asked for book recommendations and the answers were solid. But the purchase links were broken. I couldn't complete a single order.
Netflix flat out gave me bad recommendations. I ended up searching for a movie manually, like always.
Two fairly simple conclusions:
- Few companies have launched AI agents so far.
- The ones that did are stuck halfway there. They still don't replace traditional search, and they don't help the user the way they should.
The questions that stuck with me
That whole experience left me with more questions than answers:
Are the agents that already exist actually helping their users?
How do you measure the experience of someone interacting with an agent?
Do users who interact with the agent convert or retain more than those who don't?
And once you start measuring all of this, what do you actually do with that data?
The number that confirmed what I lived through
I did some research and found the number that put data behind my frustration.
77% of companies that launched AI agents have no idea whether they're helping their users.
93% plan to implement agents in the next two years.
That's the paradox: everyone talks about AI, almost nobody knows if theirs actually works.
And if you don't measure it, you can't improve it. It's that simple.
Why you have to measure an agent (no way around it)
There are three concrete reasons:
- Understand adoption → how many users are actually using the agent?
- Identify friction → where does it fail? where does it frustrate people?
- Prove impact → do users who interact with the agent convert and retain more than those who don't? If you can't answer that, why did you launch it in the first place?
And there's a fourth variable almost nobody looks at: internal impact. If your agent is used by employees (support, sales, ops), the question is the same, just pointed inward. Is it making them more productive, or is it just another tool adding friction?
Clicks and pageviews won't tell you any of this
Here's a common mistake: measuring a conversational agent with the same metrics you'd use for a web page. Clicks, pageviews, time on screen.
An agent isn't a page. It's a conversation. And a conversation isn't measured in clicks, it's measured by whether it solved what the user actually needed.
Surveys and NPS aren't enough either. They give you a general picture, but they won't tell you exactly which message broke the conversation, or which intent the agent couldn't resolve.
You need to look at the conversation itself. And in Latin America that includes channels that barely get used in the US, like WhatsApp: if your agent lives there, your way of measuring it has to live there too.
What to actually track
Group your metrics into three buckets.
Adoption
→ Adoption rate: what percentage of your users interact with the agent
→ Emerging use cases: what are people asking for that you didn't anticipate
→ Who uses it vs. who doesn't, and what sets them apart
Friction
→ Rage prompts: when a user repeats or rephrases because they're not satisfied with the answer
→ Issues: specific agent errors (direct actionable: adjust the prompt)
Impact
→ Retention impact: do users who interact with the agent come back more?
→ Conversion impact: do they complete the action they were after more often?
→ Time to complete: how long it takes someone to get what they need with the agent, versus the traditional flow
How we tested this: the workshop and the live demo
A week ago we ran a workshop showing these exact same problems live. We walked through real examples of how these errors show up day to day, and we used a tool called Pendo Agent Analytics, which does exactly that: measures your agents' performance, shows you the basic metrics, and tells you whether they're actually helping your users.
For the demo we built a simple travel platform: you can buy packages and flights manually, or by talking to an agent. We left it online so you can try it yourselves: bildung-travel.lovable.app. Go in, interact with the agent, and see what errors you run into. It's a good exercise before you keep reading.
Here's the conversation we had live:
"I'm thinking about traveling to Japan in March, what packages do you have?"
→ The answer was so long it was impossible to skim.
"Can you find me flights from Buenos Aires departing March 23rd for two people, and tell me the current price?"
→ Long answer again. And the agent doesn't handle the sale, it's purely informational.
"Okay, can you at least open Google Flights or Kayak with that search already filled in?"
→ I clicked. The search didn't load.
"I'll go look for the flight myself. Either way I want to book the package for two people. Can you email me the itinerary and payment link?"
→ The email never arrived.
Up to that point, all I had was a feeling. Five minutes of frustration, with no numbers behind it.
What we did next is something very few companies do: we went and looked at that same conversation inside Pendo Agent Analytics. And that's where all that frustration turns into concrete data:
→ Rage prompts: the exact moment I pushed back because the email hadn't arrived was flagged. That's measured frustration, not just my impression.
→ Use cases: the tool automatically clusters every conversation by use case, so you can see whether the problem the agent is supposed to solve (in this case, helping you book a package) is actually getting solved, and new use cases nobody anticipated show up too.
→ Issues: the specific errors in how the agent was built get flagged, like not being able to handle the sale or not sending the email. The actionable there is direct: adjust the prompt or check the integration.
→ Individual conversations: you can see each full conversation, connected to session replay, to understand exactly what was happening on screen while the user was talking to the agent.
Every single failure I ran into (the endless answer, the broken link, the email that never showed up) is right there, already classified. Without a tool like this, all of that stays as "a bad feeling." With it, it becomes a prioritized list of things to fix.
From data to action: two levers
Once you're measuring, you literally have two levers to pull.
1. Adjust the agent
The data tells you exactly where it fails: which questions it doesn't answer well, which intents it doesn't cover, which prompts trigger frustration.
Example: users ask "does the price include the flight?" but the agent doesn't have pricing info loaded.
→ Actionable: add context to the system prompt.
Example: a lot of rage prompts show up on comparison questions like "which is better, Patagonia or New Zealand?" and the agent gives a generic answer.
→ Actionable: improve the recommendation logic.
A side benefit nobody mentions: every improvement that reduces back and forth also reduces tokens consumed. That's a direct cost saving for the company, not just a better experience.
2. Adjust the product
Sometimes the agent doesn't actually have a problem. The problem is that it's covering up a hole that already existed in the product.
Example: the agent recommends a destination, the user goes to the page, but comes back to the chat to ask "what exactly is included?"
→ The problem isn't the agent. It's that the detail page doesn't show pricing clearly.
→ Actionable: redesign the destination card.
Example: a lot of users ask the agent "what destinations do you have under $3,000?", but that filter doesn't exist on the destinations page. If that intent shows up 40 times in Agent Analytics, it's not a fluke, it's a clear signal.
→ Actionable: add the price filter that should have been there from the start.
The agent is just another feature
Here's the idea I keep coming back to after all of this: an AI agent is just another feature of your product. It's not magic, it's not separate, it's not exempt from the same rules you'd apply to any other part of your platform.
And that means one thing: you have to measure it.
If you're not measuring whether the users who use your agent convert and retain more than the ones who don't, you have a nice feature that cost the company money. That's it.
We already wrote about the 6 levels of product analytics maturity. The logic for agents is the same, it's just that almost nobody's applying it yet. That's where the opportunity is.
If you're about to launch an agent, or you already launched one and aren't sure what's going on with it, we offer a free 30-minute assessment to walk through what to measure and what actionables you already have available.
If you need help measuring your agent, email me at guido@bildungdata.com.

