Last Updated: 3 September 2026

There is one number Keith Zubchevich thinks every company running an AI agent should be able to recite on demand, and almost none of them can: turns to understanding intent. How many exchanges it takes before it's confirmed with the user that the agent understands what the person actually came for. Everything the agent does after that point is built on it. Get that wrong and you're off to a bad start, if not lost already.
Across major e-commerce and travel sites, Conviva found it took agents an average of 3.4 turns to understand intent. And this took two minutes and 32 seconds. Three and a half exchanges should not consume two and a half minutes. When an average stops making sense, it's usually because it's an average of two very different things. Which is why you can’t watch one number.
Conviva calls the method contrastive analysis, and the principle is that you never rely on a single signal to tell you something is wrong. Alert on the metric itself. Alert on the time spent in that step, the two minutes and 32 seconds. Alert on abandonment, which was only 8.5% in this example. When any of them moves, you contrast the full population across dimensions until it resolves to the same place: not a bad agent, a bad cohort.
Zubchevich, president and CEO of Conviva, says the industry keeps reading this as an accuracy problem. It isn't. Every one of those turns is a correct, well-formed, eval-passing response. The agent is doing exactly what it was built to do, and the customer is walking out anyway, because nobody is measuring the thing the customer actually feels, which is how long it took to be understood. And you will never see it in an average or in a single transcript. It only appears when you contrast the full population across dimensions.
In our interview, Zubchevich discusses what Conviva has learned watching AI agents in production, where the current tooling falls short, and the metrics brands actually need to measure agent experience.
Q1. Conviva calls itself an "agent experience company." For readers who may not be familiar with the term, what does that actually mean? And what made AI-agent experience important enough to become central to how the company positions itself?
Agent experience is the discipline of measuring whether an AI agent actually worked for the user. It considers if the agent got to the desired outcome efficiently, and if the process was pleasant and effective for the user.
Here's what that means in practice: You ask a shopping agent for a wide-fit running shoe because you're training for a marathon. It comes back fast with three great options for long-distance running. It passes every eval you'd run against it. But none of those shoes come in wide. So you minimize the chat, go search the site yourself, and find your own pair. You bought something, but the agent failed you.
Solving this is Conviva’s mission. Deliver the full promise of agents by making everyone's agentic experience as efficient, as effective, and as pleasant as possible. All three, or you haven't delivered anything. Not whether the response looked right. Not even whether the outcome happened. Whether the person got what they came for, and what it took out of them to get there.
Q2. Conviva's platform analyzes what it calls the intent, sentiment, and behavioral signals inside AI agent interactions. In concrete terms, how does the AI identify those signals? And how confident is Conviva that it can accurately figure out what a customer is actually trying to do?
We start on the outside and work our way in. Almost everyone else starts inside out.
Mechanically, it's three signals joined into one story.
Almost everyone in this market has exactly one of those. The observability vendors have the trace. Product analytics has the clickstream. Nobody joins them, which means nobody can tell you the only thing worth knowing: that the shopper asked for a wide fit, the width parameter never made it into the tool call, and ninety seconds later they opened a new tab and searched the site themselves.
We don't score sentiment off the words in the chat box. Words lie. Behavior doesn't. Did the customer restate the request? Did they narrow what they originally asked for because the agent couldn't handle the real version? Did they leave the conversation and search the site manually? Did they come back later to double-check the outcome? Every one of those is a named, measurable signal, and we run them across the full population of sessions in real time. Not a sample. A pattern that shows up in 2% of conversations is still real, and it's still costing you money. If you're sampling, you will never see it.
That's how you get from "the agent responded correctly" to "the agent worked for this person." You stop reading transcripts one at a time and start finding the cohort where the trace looked perfect and the behavior said otherwise.
Am I confident in the approach? It's the same one we've used for twenty years to measure quality of experience for the world's biggest publishers, at live-event scale, on the nights when tens of millions of people are watching at once. We didn't learn it in a lab. We learned it watching real people, at real scale, in real time.
Q3. Conviva captures every single interaction rather than a statistical sample, a method the company calls full-census telemetry, processing roughly five trillion events per day. Why does capturing everything matter specifically for AI? Does it actually change what the AI can figure out, or is it more about making sure nothing slips through the cracks?
We don't believe in statistically sampling anything, if you have a choice. Nobody samples because sampling is better. No one has ever stood up in a meeting and argued that seeing less of your data produces a more accurate answer. Sampling exists because the vendor's architecture can't hold the full population. It's a constraint sold to you as a feature. If you're watching one conversation in twenty, you'll see the common failure a hundred times and completely miss the segment quietly leaking revenue, because nobody happened to sample it. You have to see everything to find the pattern that's invisible until it isn't.
But finding it is only half the problem. The other half is sizing it, and that's the half nobody talks about. Say your sample does surface a broken interaction. Now what? Is that one unlucky customer, or is it the visible edge of eleven thousand people hitting the identical wall this week? You cannot tell. So your team works on the issue that showed up most vividly instead of the one bleeding the most money. Full census means the issue shows up with its size already attached. This segment, this many sessions, this much revenue. Now there's no debate about what to fix first. You work the list in order.
And yes, it changes what the analytics can find. Go back to the wide running shoe. Capture every conversation and a pattern surfaces: every time a shopper asks about fit, one dimension, out of dozens, they don't buy what the agent recommended. Now dig into the tool calls. Turns out fit isn't a tag in the catalog the agent can see, so it has no idea whether it's recommending wide or narrow. That's a small fix with a large number attached to it. And you cannot find it in a sample.
Q4. AI models and agent frameworks are evolving fast, and many companies build on top of specific large language models. Does Conviva's value depend on any particular model or framework? Or is the goal to act as a layer of intelligence that works underneath whatever AI stack a customer already runs?
No, and that's deliberate. The model is not the differentiator. What decides whether an agent works is everything wrapped around it: the tools, the context, the permissions, the data it can see, and whether anyone is watching what real users do with it and feeding that back in. Putting a better driver in the car doesn't help if the car is a Prius.
So we built Conviva to sit underneath whatever stack a customer already runs. Traces reach us two ways: over standard OpenTelemetry, so it makes no difference which framework built the agent, or through our own sensor, the same client-side technology we've had running at global scale for twenty years. That sensor also picks up what's happening on the surrounding site or app. Most customers end up using both paths. The point is that we meet your instrumentation where it already is instead of asking you to rebuild it first. Model-agnostic, framework-agnostic, on purpose.
And being model-agnostic isn't just a compatibility statement. It's what makes the next thing possible. Because we sit across all of them, we can tell you which model is actually producing the better experience for your users. Not on a public benchmark, on your traffic, with your customers, on the tasks you actually run.
Now attach price to it. Every model has a token cost. Once you can see experience and outcomes by model, with cost per model sitting right next to it, you've stopped choosing on reputation and started choosing on experience per dollar. And here's what falls out more often than people expect: two models land within noise of each other on the experience metrics, and one of them costs materially less. That's not a hard call. Take the cheaper one and spend the difference somewhere it moves a number. The expensive assumption in this market is that the newest, largest model has to be the right one for every task. Sometimes it is. Frequently it isn't, and you're paying a premium for a difference your customers cannot feel.
Q5. Conviva published research showing that consumer-facing AI agents spend an average of two minutes and 32 seconds gathering basic context before they even begin helping a customer. The same study found that 8.5% of shoppers abandoned mid-session, either during that context-setting phase or right after it. How was that measured, and how many deployments did the research cover?
It was measured using full-census, client-side telemetry across customer-facing agent sessions on e-commerce and travel booking sites. No sampling.
The measurement happens inside the conversation. We start the clock the moment the session opens and run it until the user confirms the agent has understood what they came for. Every exchange in that window, in both directions, is the 2:32. So it's the same span as turns to understanding intent, one measured in exchanges, the other in time. And we measure it both ways on purpose, because the two come apart. Turns tell you how much work the customer had to do. Time tells you how long they sat there doing it. An agent can get to intent in three turns and still burn ninety seconds on each one, and the customer feels the clock, not the turn count.
Abandonment we take from behavior rather than from the transcript, because a transcript can't tell you the difference between a conversation that ended and a person who quit. We split it by where it happened: people who walked while still being asked questions, and people who got all the way through, got a response, and left anyway. Two different failures, two different fixes.
Four failure modes, measured four separate ways. That's deliberate. Track one and you only ever find the failures that one metric is shaped to catch.
Q6. That same research found that in nearly two thirds of AI agent sessions, customers were asked to re-explain information the agent should have already had. Things like pages they visited, tasks they attempted, errors they ran into. What does Conviva's solution to that problem look like when it's running in production? And is there a customer where it's already working?
The moment the conversation starts, Conviva measures how much of it is spent re-establishing ground the customer already covered, which segments it happens to, and what it costs (in turns, time,and abandonment). Clickstream events make that measurement possible.
In the shoe example I gave earlier, I said that the agent recommends three pairs and then the shopper opens a new tab, searches the site manually, and buys. If you look at the agent's logs in isolation, it looks like it did a good job. But when you combine it with the clickstream data you see the purchase happened despite the agent’s recommendations, not because of them.
That's the loop we're building, and we're building it alongside customers who are pushing us to move faster on it.
Q7. An AI shopping agent can complete a purchase and still leave the customer frustrated. At the same time, Conviva's A2A Commerce Report found that AI-referred buyers convert at 3.5 times the site average, and at 45.5% when they ask an on-site agent about price. How should companies think about that gap between a successful outcome and a genuinely good customer experience?
Outcome and experience are not the same question, and that gap is exactly where agentic commerce is going to get overconfident. There are three rungs on this ladder: Quality of service, which tells you if the agent is up and responding; Quality of response, which tells you if an individual reply looks accurate and on-topic. Most teams are stacked on those two, and a 45.5% conversion rate clears both without breaking a sweat. It tells you nothing about the people who didn't convert, or whether the ones who did are coming back.
The third rung is Quality of Experience, which tells you if the shopper got what they came for, what it cost them, and if the process is one they'd willingly go through again.
That last one you can only answer by watching what happens after the outcome. Did they come back to double-check the order? Did they open a ticket about the thing the agent said was resolved? A completed purchase that generates a correction, a complaint, or a return isn't a clean win. It's a win with a bill attached, and the bill shows up in someone else's department.
Q8. A Gartner report cited in Conviva's research identifies the absence of a dedicated context layer as a root problem in most AI agent deployments. Conviva's answer involves something it calls a context graph. For readers who know what a CRM or customer profile does: how is a context graph different? What makes it something those existing tools can't replicate?
Gartner is right that the layer is missing. But we found a bigger problem sitting in front of it, and that's the one we've chosen to solve first.
Before you ask whether an agent knows enough about the consumer, ask a simpler question: is the agent delivering a good experience with everything it has already been given? Because it has a great deal already. Whatever context graphs it's wired into. The tool calls available to it. The prompts it's running. The reasoning steps it takes to get from the question to the answer. All of that is producing an experience right now, for every user, today, and almost nobody is measuring whether that experience is any good. That's the fundamental problem, and it sits upstream of the consumer context question.
So, to the CRM question. A CRM tells you who someone is and what they've bought: declared, historical, name and tier and past orders. It’s useful, but it tells you nothing about how this conversation is going. What we build is the behavioral record of the agent's own work, what the customer did in response to each thing the agent said. Did they accept it or restate it? Did they quietly narrow what they were asking for? Did they leave and do it themselves? A CRM knows your loyalty tier. This knows the agent just lost you, and why.
Is a full consumer context graph valuable? Yes, and it's coming. But it's the second problem. This industry has a habit of reaching for the ambitious thing before it has solved the fundamental one. We're doing these in the order that actually works. Same behavior, opposite meaning, depending on who's doing it. That's the part a customer profile will never see.
Q9. If an AI agent is going to respond well from the first message, it needs some understanding of what the customer has already done. What kind of behavioral context can Conviva give an agent in real time, and how does that actually change the response the customer receives?
Let me push on the premise a little, because I think the industry has this backwards. Everybody is fixated on the opening message. That's the part I'm least focused on, and honestly the part I'm least confident anyone solves cleanly in the near term.
The signal that's real, available, and actionable today is what happens during and after. The search you ran while the chat was still open. The step you were stuck on when you asked. An error you hit that the agent was never told about. Whether you closed the conversation and started a fresh one because the first went nowhere. Whether you came back two hours later to check on the thing it told you was handled. None of that lives in the transcript. It lives in everything around the transcript. And it changes what happens on turn four, which matters considerably more than turn one, because turn four is where people leave.
If you capture those signals across your entire population and you stop fixing one conversation at a time, you start seeing which segments the agent systematically fails, and you fix the agent.
Q10. Switching to Nexa, Conviva's AI analytics agent. Nexa lets people ask questions about their company's data in natural language and get instant insights, dashboards, or issue investigations. What actually happens under the hood between someone typing a question and Nexa returning an answer?
Someone types something like "why did conversion drop for mobile users in the Southeast last week." Nexa has to work out what that's actually asking, which dimensions, which time window, which metric definition, translate it into a query against the full-census event stream, run it, and hand back an answer, a chart, or a dashboard.
We hold Nexa to the exact standard we're selling everyone else. We don't just track whether Nexa responded, how fast, and whether the answer was correct. Those are table stakes, and they're not sufficient. We measure things like Clarification Ratio — how often a user has to restate the question because Nexa missed their intent, and Correction Rate — how often a user has to fix the answer. And we slice those across dozens of dimensions.
Here’s an example: Clarification Ratio runs higher on queries with very little detail. That's not an excuse for Nexa. It has to work for everybody, including the person who types six words. It's a signal that when the query is thin, Nexa should ask a clarifying question before it runs the analysis. That insight goes straight into our coding agent and becomes a product change.
We're customer zero. If the methodology didn't work on our own agent, I wouldn't be selling it to you.
Q11. Nexa combines retrieval-augmented generation with fine-tuning rather than relying on just one approach. What does each technique contribute, and what did you learn that convinced the team both were necessary?
Retrieval keeps Nexa grounded in what actually happened instead of what a general-purpose model assumes probably happened. It pulls from our own full-census event data rather than inventing a plausible-sounding number. That solves "is this fact real."
It does not solve "does this system speak our language." Knowing that "SPI" means Streaming Performance Index, or that "a session" means something specific in our schema and not the dictionary definition, that's fluency, and retrieval doesn't give it to you.
What convinced us was watching a model get every fact right and still answer the wrong question, because it was reasoning in generic terms instead of ours. Grounding without fluency gives you a correct answer to a slightly wrong question. Fluency without grounding gives you a confident answer to the right question that's completely made up. Pick one and you're trusting the model to fill in the rest on its own. Bad idea.
Q12. Every AI system has boundaries. What types of questions can Nexa answer reliably right now? And where does a human analyst still need to step in because the AI can't safely make the call on its own?
Nexa is reliable today on anything with a clear definition and a well-formed question against data we already hold: metric lookups, slicing a known dimension, generating a standard report, catching a threshold breach the second it happens. Ask it for SPI by device model in the last hour and it'll nail that all day.
And we don't grade that on vibes. We track a composite task success rate for Nexa itself, whether someone accepts the analysis, digs deeper, or shares the report out, versus abandoning it or asking again.
Where a human still steps in is judgment. Deciding that a technically anomalous pattern is fine because of a known one-off event. Deciding what the business should do about a finding, not just what the finding is. Nexa will tell you a segment is failing and hand you the trace. It doesn't get to decide what "acceptable" means for your business, and it shouldn't. That call belongs to whoever owns the outcome.
Q13. During the 2026 World Cup, Conviva monitored tens of millions of concurrent streams across more than 20 platforms, and Nexa was being used by operations teams in real time. What did the AI find during a live match that would have been especially hard or slow for a human team to catch?
One broadcast partner built an MCP-based dashboard tracking Streaming Performance Index by device model, live, during the tournament. The data showed SPI dropping 10 to 22% on specific smart TV models the moment 4K was enabled.
They toggled 4K off for those models with a feature flag and stabilized the stream before the match ended. The whole thing was built in about twenty minutes, live, at World Cup scale.
Could a human team have found that? Eventually, combing through device-level logs after the fact. But nobody is doing that during a live match with tens of millions of concurrent viewers, and by the time they do, the goal has been scored and the viewer is gone. That's the entire argument for population-level analysis in real time: the pattern only exists at a scale and a speed that no manual review process operates at.
Q14. Conviva reported that Nexa adoption among operations teams grew more than 20% week over week during the tournament. One streaming provider in the Middle East used it to build a full executive quality-of-experience report with country-level, platform-level, and bilingual output, in the time it would previously have taken just to set up the query. Is that a speed improvement for people who already do analytics work? Or does it change who inside a company can do that kind of analysis in the first place?
Both. But they're not equally interesting.
For the people already doing this work, it's a real speed gain, the report now gets built in the time it used to take to open the query. The bigger shift is who gets to ask the question at all. Natural language means an operations lead who has never written a query in their life can pull a country-level, bilingual executive report without waiting in an analyst's queue.
That's the part I care about. Speed helps the people who are already doing analysis. Removing the query-writing bottleneck changes how many people in the company can do analysis, period. We're seeing C-suite executives at Fortune 500 companies show up in our top-user dashboards week after week. Those are not people who were writing SQL last year.
Q15. When Conviva detects a pattern linked to poor AI agent performance, how much of the diagnosis can the system handle on its own? Can it identify why the agent is failing, or does a human still need to interpret the data and establish the root cause?
It gets most of the way to "here's exactly where and how this is failing" on its own. This is the part of the video business I'm most confident carries over, because we already run automated baseline-deviation detection across tens of thousands of user cohorts for our largest streaming customers. Same technique, new domain.
At baseline we track six recurring root causes: missing context, unclear tool instructions, no clarifying question asked, an ambiguous default, no check against the original ask, and stale or partial retrieval. Population-level contrastive analysis exists specifically to isolate which segment is failing and surface the likely mechanism, automatically, rather than waiting for somebody to go looking. Teams can extend that list for their own use case.
Today Nexa finds the root cause and recommends the fix. Most teams still keep a human in the loop to route that into their coding agent and ship the change. We're driving toward fully self-healing agents, and we'll be the first to show real results from it, because we're already doing it with Nexa.
Q16. Taking that one step further: how far can Conviva close the loop today between detecting that an AI agent is performing badly and actually improving that agent? Can the platform recommend or push changes to prompts, tools, or workflows on its own?
Today we close the loop on the evidence, not the fix.
Through MCP, that evidence, the failing segment, the trace, the likely root cause, goes straight into the customer's own coding agent. Their engineers act on it inside the workflow where the agent actually gets built, instead of someone copying findings into a ticket that sits for three weeks.
The technology will be fully solved before the people and process are, at least in large enterprises. We work with companies at very different stages of agentic maturity. Some want to flip this on tomorrow. Others need a governance conversation first. Both are fine. But the ones who move first are going to compound the advantage, exactly like they did in the streaming wars.
Q17. More and more customer interactions are moving from clicking through an interface to asking an AI agent to handle things on the customer's behalf. As that shift happens, what new types of behavioral data start to matter that companies could safely ignore in the era of apps and websites?
In the app-and-website era the signal was structural: which page, which button, how far down the funnel. Clean, unambiguous. You either clicked "add to cart" or you didn't.
Conversation has no such thing. The signal that matters now is turn-level and sequential. How many exchanges it took before it was confirmed with the shopper that the agent understood what they came for, turns to understanding intent, out of the gate. And note the word confirmed. Not when the agent internally decided it knew. When the customer saw it back and agreed. Those are different moments, and the distance between them is where people quit. How often someone had to restate themselves. Whether they quietly opened a second tab. Whether they downgraded what they actually wanted because the agent couldn't handle the real version.
Companies could safely ignore all of that when the interface was a funnel. They can't now. The words in a conversation can look completely fine while the customer is quietly settling for less, and that only shows up in the behavior around the conversation, never in the eval score.
Q18. Conviva currently processes over 12 billion sessions per year and supports more than $20 billion in customer revenue. If this interview happened again in two or three years, what would AI be doing inside the platform that it cannot do today? Particularly when it comes to the shift from understanding customer behavior to actively improving how AI agents make decisions, without waiting for a human to intervene.
I'll do you one better. One year, not three. I want self-healing agents in production, and I think we get there. The shift is from "here's what's wrong" to "we saw what was wrong, we fixed it, and here's what it did to your business."
But the bigger change isn't inside our platform. It's that AI in production finally gets held to a real standard.
We watched this happen in streaming. Think about how much buffering you tolerated in 2010. You'd sit there and wait. If that happened to you today you'd abandon the show in thirty seconds and never open that app again. Nothing changed about your patience. What changed is that the industry got good, and once it got good, viewers stopped accepting anything less. Conviva coined the Streaming Performance Index that the industry now benchmarks itself against.
The same reckoning is coming for AI agents, and it's coming fast. Right now people are still charmed that the thing talks back. That grace period ends. And when it does, the companies who were measuring the experience the whole time will be the ones still standing.
We didn't learn that in a lab. We learned it watching real people, at real scale, in real time. That's still the only place this gets learned, and it's still where we're putting everything.