Why the most advanced AI models are not always the best choice

Gartner forecasts that over 40% of agentic AI projects will be scrapped by 2027, citing unclear business value, rising costs and inadequate risk controls. Image: Daniil Komov/Unsplash
- More than 40% of agentic AI projects are forecast to be scrapped by 2027, with unclear value and rising costs among the reasons.
- In controlled tests of an enterprise system handling supplier invoices, the most capable and most expensive configuration performed worst.
- Matching models to the specific task, and testing components together rather than one at a time, matters more than buying the highest capability tier.
Agentic artificial intelligence – software that plans and acts on its own rather than waiting for instructions at each step – is the enterprise technology story of 2026. It may also be the most expensive disappointment of 2027.
Gartner forecasts that over 40% of agentic AI projects will be scrapped by 2027, citing unclear business value, rising costs and inadequate risk controls. Organizations everywhere are now choosing where to place their first serious agentic bet. Very few are choosing on the basis of evidence, because very little evidence exists in public. Vendor claims are abundant. Controlled comparisons are rare.
Research led by myself and a co-author at the Massachusetts Institute of Technology set out to produce some. Working with a global enterprise that processes over $75 billion in supplier invoices each year, we built a system of AI agents that investigates and resolves invoice errors on its own, and escalates to a person when it cannot. We then ran four model configurations against the same 44 test invoices, measuring accuracy, computation consumed and time taken.
The results ran against the assumption behind most deployment decisions today: that reaching for the most capable model available is the safe choice. Call it the capability trap.
The capability trap is not hypothetical
The premium configuration – which paired the strongest reasoning model available with a premium document-reading model – came last. It resolved 91% of invoices without human help. A budget configuration reached 95%. A well-matched mid-tier configuration resolved all 44, using models priced at a fraction of the premium tier.
Every failure occurred during one task: identifying duplicate invoices. That specific task required the AI to evaluate five separate, competing rules at the same time to decide if two invoices matched. The most capable model was, if anything, too deliberative. It hedged on borderline matches that a more constrained model simply flagged. Raw model capability is easy to evaluate on paper, but fit for the specific job is not – and job fit is what actually drove our results.
A sample of 44 invoices is admittedly small, and these results are not a final ranking. But they do establish that performance cannot be taken for granted. Indeed, the setup everyone would have picked on reputation alone actually came in last.
Three forces behind the trap
Why do high-end setups fall into this trap? Our findings point to three primary drivers: model over-thinking, mismatched components and unnecessary structural complexity.
Elaboration where none is needed
Larger models tend to generate longer and more heavily qualified reasoning. On open-ended problems that is a strength. On straightforward, rule-based checks, it introduces doubt into decisions that simpler models settle cleanly. And that is exactly what went wrong when we tested duplicate detection.
Components that do not match
Our most useful result came from a configuration designed to save money: it kept the premium reasoning agent, but swapped in a cheaper model to read the documents. It backfired. Running the identical agent, that setup consumed roughly 68% more computation per invoice, because the agent spent its reasoning on cleaning up messier extracted text before it could start the actual job. Savings at one stage became losses at the next. As one industry analysis put it, garbage in, agentic out. Autonomous systems inherit the quality of everything upstream of them.
Unnecessary structural complexity
A third result points the same way. We tested two designs on identical inputs: a simple pipeline where agents run in a fixed order, and a more sophisticated one where a supervisor checks each output and can send work back for another attempt. The sophisticated design was slower, consumed more computation and produced the same accuracy. Not one case was resolved by one design and missed by the other.
One caveat matters here: our test invoices contained clearly defined errors, and on the genuinely ambiguous exceptions that human teams find hardest, the ability to try again may well prove decisive. But complexity should be added because a measured problem demands it, not because it looks serious.
Where the value actually sits
In the global enterprise we studied, one regional division processes 97% of its invoices without human intervention, while another manages just 67%. Closing that 30-point gap would move hundreds of thousands of invoices a year into hands-free processing, and the obstacles are mundane: suppliers holding several bank accounts across currencies, inconsistent tax treatment, credit notes that existing systems cannot handle. These are the contextual judgements that rule-based automation has never been able to make, and they will multiply as jurisdictions, including the European Union, move towards mandatory electronic invoicing.
The wider evidence points the same way. A study from MIT’s Project NANDA, reported by Fortune, examined 300 public deployments alongside interviews and surveys of business leaders, and found that only around one pilot in 20 produced rapid revenue gains. Its authors attribute the gap not to weak models but to poor fit with the workflows the tools were meant to change. The finding has been debated, and the exact percentage is contested. The direction is harder to argue with, and it matches what we saw: success depended less on the raw power of the AI model itself, and far more on the system built around it.
The finance leaders we interviewed were clear that they did not want an autonomous organization. They wanted an AI-assisted one. Our system was built to that brief: it logs every reasoning step for audit, and when it cannot resolve an issue, it drafts a summary and a recommended action for a person to review rather than acting alone. Governance built in from the start costs far less than governance retrofitted after an incident.
Testing beats buying
The capability trap is not an argument against adopting agentic AI. It is an argument against treating capability as a substitute for evidence. Two habits would change most deployment decisions:
- Benchmark on specific tasks, not on published rankings. The model tier that leads general leaderboards may finish third on the narrow job that actually needs doing; the only way of knowing is by running it.
- Test the chain, not the link. An agent is a sequence of components, and improving one in isolation says almost nothing about the cost or reliability of the whole system.
Both habits depend on a third choice made earlier: scoping the problem narrowly enough that the result can be measured at all. An ambitious pilot produces a story. A narrow one produces evidence, and evidence is what survives a budget review.
The organizations that avoid Gartner’s predicted 40% failure rate will not be the ones that moved fastest or bought the most capable models. They will be the ones that measured, tested their components together, and kept people at the points where judgement genuinely matters. Our own results came from a deliberately small problem, and all three findings went against what we expected. Deploying agents without that kind of test is not a technology decision. It is a guess with a budget attached.
This article draws on master's capstone research conducted with a co-author at the MIT Center for Transportation and Logistics. The sponsoring organization is not named at its request.
Don't miss any update on this topic
Create a free account and access your personalized content collection with our latest publications and analyses.
License and Republishing
World Economic Forum articles may be republished in accordance with the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Public License, and in accordance with our Terms of Use.
The views expressed in this article are those of the author alone and not the World Economic Forum.
Stay up to date:
Artificial Intelligence
Forum Stories newsletter
Bringing you weekly curated insights and analysis on the global issues that matter.
More on Artificial IntelligenceSee all
Niklas Mortensen
August 27, 2026




