Vivek Ahuja, VP-IT at rSTAR, spearheading business and IT transformation with a focus on manufacturing, energy/utilities and construction.

getty
I have sat in dozens of enterprise AI evaluation meetings over the past two years. Almost every meeting starts the same way: Someone pulls up a comparison of GPT-4 versus Claude versus Gemini, walks through benchmark scores and asks the room which model the company should standardize on.
In my opinion, this is the wrong starting place, and it is consuming an enormous amount of executive attention that should be directed elsewhere. A foundation model is only one component of an enterprise AI solution, yet most organizations are evaluating AI the way they once evaluated servers: by specs.
While 88% of organizations have adopted AI in at least one business function, according to McKinsey research from 2025, only about one-third have begun scaling AI across the enterprise. In my experience, companies struggling to scale are often trying to solve a model-selection problem when they really have a measurement problem.
To address it, companies must stop focusing only on benchmark scores and start focusing on business outcomes.
The Benchmark Trap
Model benchmarks tell you how well a model performs on standardized tests, but they do not tell you how well it will perform on your data, within your workflows, or against your business objectives.
Even Amazon has questioned how much traditional AI benchmarks matter for enterprise deployments (subscription required). As Rohit Prasad, Amazon’s SVP of AGI, explained: "The only way to do real benchmarking is if everyone conforms to the same training data and the evals are completely held out. That’s not what’s happening. The evals are frankly getting noisy, and they’re not showing the real power of these models."
Instead of chasing the highest benchmark scores, organizations create more value by adapting AI solutions to proprietary data and business workflows.
For example, I've seen enterprises pick a model based on leaderboard rankings, deploy it into a contact center and watch it underperform a simpler model that was better tuned to their specific knowledge base and customer language.
The reason is that enterprise AI performance depends on the full stack: data quality, retrieval architecture, prompt design, integration latency and guardrails. The model is only one variable, and the rest often determine whether the solution delivers business value.
What A CFO Actually Wants To Know
Focusing on outcomes is also the fastest way to secure sustained AI funding. This means leaders must speak in outcomes and in experience. No CFO is going to approve a second year of AI investment because the model scored 92% on a reasoning benchmark if you can't also explain what changed in the business.
The switch is to focus on business metrics, rather than AI metrics, when beginning a pilot. For example, when my organizations deploy AI-powered solutions in utility contact centers, we measure five things: first-contact resolution rate, average handle time, customer satisfaction score, cost per resolution and agent ramp time.
This reframing shifts the conversation from "which model is best" to "which architecture delivers the best outcome." And that is a much more productive question for an enterprise to be asking.
A Three-Layer Measurement Framework
After deploying AI across multiple utility and manufacturing programs, I have settled on a framework that keeps teams focused on value rather than vendor hype. It works across use cases, from knowledge bases to agent assist to email automation.
The first layer is task accuracy. For any AI-driven workflow, what percentage of outputs are correct without human correction? This is not the model's benchmark accuracy. It is accuracy measured in production, on your data, with your edge cases. In the deployments I've worked on, there is often a 15-to-25-point gap between vendor-reported accuracy and what you see in a live environment.
The second layer is operational efficiency. Did handle time go down? Did resolution rates go up? Did the number of escalations decrease? These are the metrics that justify continued investment. If your AI deployment cannot show movement on at least one of these within 90 days, something in the stack needs to change.
The third layer is financial impact. What is the cost per AI-assisted interaction versus a fully manual one? What is the payback period? In my experience, the programs that survive budget cycles are the ones that can point to a dollar figure.
Where To Start This Quarter
If your AI evaluation is still centered on model comparison, here is how to shift the conversation:
• Define three to five business KPIs for every AI use case before selecting a model or vendor. If you cannot articulate what "success" looks like in operational terms, you are not ready to build. As mentioned above, I typically track first-contact resolution, average handle time, customer satisfaction, cost per resolution and agent ramp time in my projects.
• Measure production accuracy, not benchmark accuracy. Stand up an evaluation pipeline that tests your AI outputs against your actual data and edge cases on an ongoing basis. We validate AI using real customer interactions and subject matter expert-reviewed responses, which provides a much more accurate picture of production performance.
• Build a 90-day outcome scorecard for every AI deployment. Track task accuracy, operational efficiency and cost per interaction monthly. Sharing those results with finance can help shift the conversation from model performance to measurable ROI.
• Treat the model as a swappable component. Design your architecture so you can test a different model without rebuilding the entire solution. I've found that separating the model from the application architecture makes it much easier to improve performance while controlling costs.
Using this framework, you will likely see that selecting the newest model is often less important for succeeding with AI adoption than consistently measuring, improving and proving business value. Models will continue to change. A disciplined approach to measurement will always endure.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?