Why AI Products Feel Impossible To Measure And What To Do About It

Why AI Products Feel Impossible To Measure And What To Do About It

Udit Mehrotra is a distinguished product leader who currently serves as Head of Product for North America Languages Experience at Amazon.

getty

​A few months after one of the AI features I shipped went live, I was sitting in a business review when someone asked a version of the question I had been quietly dreading: Is this actually helping customers?

I had data. Engagement was strong, and the model was responding to the vast majority of queries. But the honest answer to what they were really asking (Are customers making better decisions because of this?), I couldn’t give cleanly. Not because we hadn’t measured anything, but because we had measured the wrong things.

I’ve shipped several AI-powered features over the past few years: a conversational shopping assistant, a bilingual generative AI review system and a personalized size recommendation engine. Each one was different, but they shared a problem. We kept defaulting to the same measurement instincts we’d developed for non-AI products; things like engagement rate, click-through and session metrics, and those instincts didn’t translate.

This isn’t a niche problem. According to McKinsey, only 39% of companies report that AI investments have delivered significant financial returns, even as adoption continues to climb. Part of that gap is technical, but a meaningful part is measurement. Teams are shipping AI features without a clear framework for knowing whether they’re working.

AI products behave like relationships, but we keep measuring them like transactions.

With a deterministic feature (say, a new checkout button or a redesigned filter), the output is fixed. The feature behaves consistently, and measurement is straightforward: did customers use it or not? A/B testing works cleanly because the treatment is stable.

AI changes that. A model’s output varies with input, context and the distribution of queries it encounters. An AI assistant might handle 90% of queries well and quietly fail on the 10% that matter most. If engagement is your primary signal, you won’t know. Klarna found this directly when they dug into their AI customer service agent: aggregate satisfaction scores looked acceptable, but when they segmented by query type, failure rates on complex issues were higher than the headline number suggested. The model was passing a test it wasn’t actually taking.

Our bilingual review system ran into the same thing. Engagement was healthy, but the useful question wasn’t whether customers were engaging. It was whether customers who used the bilingual summaries made better purchase decisions, and we hadn’t built the instrumentation to answer that. After a few cycles of this, I started forcing myself to answer four questions before any AI experiment, and it changed how my teams approached measurement.

Before measuring AI performance, follow these four steps.

​1. Define

First is whether you’ve defined what “correct” looks like before the experiment starts. For deterministic features, “correct” is obvious: the button worked or it didn’t. For AI, you need to build a test set, a collection of representative inputs with expected outputs, and evaluate your model against it independently of engagement.

This is the step teams skip most often. You can’t trust a lift in engagement to tell you the model improved. Engagement goes up for all kinds of reasons, and some of them have nothing to do with the model getting better.

2. Grade

Second is whether you have a way to evaluate output quality at scale, not just user behavior. This is the step that requires the most investment and helps deliver the most diagnostic value.

For our size recommendation engine, we built an evaluation harness that compared model recommendations against actual return outcomes. Not just click rates, but whether the size the model suggested turned out to be the right one. The difference between what customers clicked and what they kept was where the real signal lived. This is what separates teams that can actually diagnose why a model got worse from teams that are mostly guessing between launches.

3. Trace

Third is whether you can name the actual customer outcome you’re trying to move; not the proxy metric, but the real behavior change you believe this feature enables.

For a shopping assistant, engagement isn’t the outcome. A customer who asks five questions and buys confidently is different from a customer who asks five questions, gets confused and abandons. Both look identical in session data. Return rate, repurchase rate and post-purchase satisfaction are harder to measure but are far more honest.

4. Translate

Last is whether you can explain the result to someone who has never heard the word “model.” If your only answer to ‘Did this work?’ is a lift in a metric that requires three slides to define, you’ll lose credibility with every business review you sit in.

A checkout optimization moved conversion. A recommendation module moved attach rate. You could draw a straight line from the feature to the number. AI teams need to do the same, and that means choosing outcome metrics that connect to decisions the business actually cares about.

Measuring right is what keeps AI products funded.

None of this is simple. AI measurement requires more upfront work than most teams budget for, and the payoff isn’t immediately visible. But the alternative is a growing portfolio of features that look fine in dashboards and disappoint in business reviews.

The teams shipping AI that actually sticks aren’t necessarily measuring more. They’re measuring in the right sequence: define correctness before you ship, evaluate quality independently of engagement, trace outcomes rather than proxies and translate results into language that builds credibility with the people who fund the next experiment. That’s the infrastructure that makes the difference between AI products that get funded for the next iteration and AI products that quietly get deprioritized after the first review.​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?


Read More

Zaļā Josta - Reklāma