
In August 2025, MIT’s Project NANDA published a number that should have rattled every CFO holding an AI budget. After analysing 300 enterprise deployments, surveying 350 employees, and interviewing 150 executives, the researchers concluded that 95% of enterprise AI pilots delivered no measurable P&L impact. The total spend they were tracking sat somewhere between thirty and forty billion dollars.
The instinctive response to a number like this is to blame the technology. The models hallucinate. The data is messy. The use cases are wrong. All of that is true, and none of it is the real story.
The real story is that almost nobody in the value chain has any commercial reason to find out whether the AI worked.
The buyer signs the contract, the vendor delivers the system, the system goes live, and at that point the contract is fulfilled. Whether the customer support deflection rate actually moved, whether the resolution time actually dropped, whether the agent actually paid for itself, all of this becomes the buyer’s private problem. The vendor has already booked the revenue. The buyer is left holding a tool and hoping for an outcome.
This article is about how to close that gap. It is also about why most of the industry has chosen not to.
McKinsey’s State of AI 2025 survey of nearly 2,000 organisations found that only 39% of organisations could link any EBIT impact to AI, and for most of those, the impact was below 5%. Just 5.5% of respondents reported that AI was generating significant value and contributing more than 5% of EBIT. The pattern is consistent across every major survey of the last twelve months. Adoption is universal. Measurable enterprise impact is rare.

The MIT report introduced a useful phrase for this divide. They called it the GenAI Divide, and they were unambiguous about its cause. The 95% failure rate, the report stated, was the clearest manifestation of the GenAI Divide, and the core issue was not the quality of the AI models but the learning gap between tools and organisations.
That framing is correct, but it is incomplete. The learning gap is real. So is the integration gap, the data gap, and the change management gap. But sitting underneath all of them is a contracting gap. The way enterprise AI is bought and sold today does not require anyone to prove that it worked. The default engagement model is a transaction, not a partnership. Vendors are paid for what they ship, not for what their work produces.
This is the gap Accubits has been building Outcome as a Service to close.
The premise is straightforward. If the buyer is paying for an outcome, the vendor should be paid based on whether that outcome is achieved. If the support cost per ticket does not drop, the vendor does not get paid the upside. If customer retention does not improve, no fee is collected on the improvement that did not happen. The vendor stops being a builder and starts being a co-owner of the result.
That sounds obvious when you say it out loud. It is, in fact, not how the industry works.
Three structural reasons keep AI vendors away from outcome measurement, and they are worth naming clearly.
The first is commercial. Time and materials, fixed bid, and milestone delivery contracts all share one feature. They pay the vendor for delivery, not for impact. The vendor’s commercial incentive ends the day the system goes live. Asking that vendor to also measure outcomes is asking them to volunteer for accountability they were never paid to take on. Most refuse politely. Some refuse by quietly never producing the report.
TCS’s CEO K Krithivasan acknowledged this directly on a recent earnings call. He noted that some clients prefer to start on a time and materials basis because as the project evolves they want to see how they benefit from the results, and only later do they consider moving to fixed pricing. The honest reading of that statement is that the customer is the one trying to drag the engagement towards outcome alignment, while the vendor is comfortable with the meter running.
The second reason is access. Even when a vendor wants to measure outcomes, they often cannot. The data that would prove or disprove value lives inside the customer’s environment. Production logs. CRM records. Support ticket histories. Revenue attribution data. Most vendor contracts do not include the data sharing rights, the instrumentation, or the long term observability needed to build a defensible ROI calculation. The vendor delivered the system. The data sits behind the customer’s firewall. The measurement never happens.
The third reason is the messiness of the AI category itself. Gartner has called this out plainly. Many vendors are engaging in agent washing, the rebranding of existing products such as AI assistants, robotic process automation and chatbots without any substantial agentic capabilities. Gartner estimates only about 130 of the thousands of agentic AI vendors are real. When most of the market is selling old software with a new label, no vendor has any incentive to introduce a measurement standard that would expose how little their tool actually changed.
This is why Gartner’s prediction that over 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value or inadequate risk controls reads less like a forecast and more like a description of a market that has been allowed to operate without measurement discipline for too long.
Anushree Verma, the Gartner Senior Director Analyst behind that prediction, put the underlying point bluntly. She said that to get real value from agentic AI, organisations must focus on enterprise productivity rather than just individual task augmentation, and that it is about driving business value through cost, quality, speed and scale. Productivity, cost, quality, speed, scale. Those are outcomes. None of them get measured in a delivery contract.
Even when buyers and vendors both want to measure outcomes, AI makes it genuinely difficult in ways traditional software did not.
Traditional enterprise software either worked or it did not. The CRM either captured the lead or it broke. The ledger either reconciled or it threw an error. Failures were visible, loud, and easily attributed.
AI fails quietly. A model can drift over months without anyone noticing that its accuracy on the most valuable cases has degraded. Hallucinations are stochastic. Adoption can be technically high while business impact is low because users are not trusting the outputs. Edge cases compound silently until a regulator or a high value customer surfaces one. None of these failures show up in a green status indicator.
Then there is the multi-stream value problem. A traditional software ROI calculation could often be reduced to a single number, a cost saved or a revenue gained. AI value tends to arrive across four streams at once. Cost reduction in operations. Speed improvement in cycle times. Quality improvement in outputs. New capability that was previously impossible. Each stream needs its own baseline, its own metric, and its own attribution logic. Most ROI exercises measure one of the four and quietly discard the rest.
The counterfactual problem makes attribution harder still. When revenue goes up after an AI deployment, was it the AI, the new pricing, the seasonal lift, or the senior sales hire who started in the same quarter? Without a control group or a holdout cohort, the AI gets credit for everything that happened in its general vicinity, or none of it, depending on the politics of the moment.
And then there are the hidden costs. AI ROI calculations must include hidden costs such as data cleaning, which can consume 30 to 50% of budgets, alongside model monitoring and change management expenses that often exceed initial technical costs. Inference costs in production routinely exceed the projections from pilot phase. Human review overhead, which is genuinely necessary in regulated workflows, often gets treated as a temporary expense even when it is structural. Hidden costs like oversight and maintenance can add 40 to 60% to total project costs, yet most ROI models omit them entirely.

Time horizons compound the problem. Software ROI compresses into a quarter. AI ROI follows a J-curve. The first six months often look worse than the baseline because the organisation is paying integration, training, and change management costs without yet seeing the productivity dividend. Most internal business cases are evaluated against a quarterly review that the J-curve cannot survive.
This is why measurement has to be designed in from the start. By the time anyone asks the question, the data needed to answer it is already gone.
A defensible AI ROI calculation has six steps. None of them are exotic. All of them are routinely skipped.
Define the outcome before you scope the system. This is the step that gets reversed most often. A vendor walks in with a capability, the customer agrees to deploy it, and the question of what business outcome it should move is left for later. Later never comes. The right sequence is to name the outcome first. Reduce average ticket resolution time by 40%. Cut churn in the SMB segment by three percentage points. Lift conversion on the trial flow by 15%. The system is then designed backwards from that outcome. If the outcome cannot be named, the project should not be funded.
Establish a baseline that will hold up under scrutiny. A baseline is not a vague memory of how things used to be. It is a documented measurement of the current state, taken with the same instrumentation and definitions you intend to use after deployment. If the baseline is sloppy, the post-deployment number is meaningless, because there is no defensible delta. Baselines should be locked in before the system is built, not reconstructed afterwards.
Pick metrics across cost, revenue, quality, and adoption. A single number will mislead you. Cost metrics tell you what you saved. Revenue metrics tell you what you gained. Quality metrics tell you whether the cost saving came at a customer experience price you are not yet seeing. Adoption metrics tell you whether the system is actually being used or quietly being routed around. The four together describe whether the deployment is real value or just busy work that happens to involve a model.
Account for the full cost stack. Inference costs. Storage and retrieval costs. Human review and escalation costs. Integration and middleware costs. Change management. Training. Ongoing model monitoring and red-teaming. Vendor licence fees. Compliance overhead. The temptation to leave out the costs that are inconvenient is enormous, particularly when a project is being defended internally. Resist it. The hidden costs are where ROI quietly evaporates.
Attribute properly. Use control groups. Run staggered rollouts where one region or one team gets the system first. Maintain holdouts long enough to see the steady state, not just the launch bump. If your AI initiative cannot survive a comparison with a comparable cohort that did not get it, you do not have evidence of value. You have a hopeful narrative.
Measure sustained value, not launch value. The most dangerous number in AI ROI reporting is the one that comes from the first month after launch. Adoption is high because the system is new. Resolution times look great because the easy cases are getting routed first. Customer satisfaction is up because the novelty is real. Six months later, the easy wins are exhausted, drift has set in, the operations team has stopped tuning the prompts because nobody owns it, and the original business case is quietly forgotten. Sustained ROI is what matters, and sustained measurement is what catches it.
The Klarna case is instructive on this last point. In February 2024, Klarna announced that its OpenAI powered AI assistant had handled 2.3 million conversations in its first month, equivalent to the work of 700 full time agents, with average resolution times falling from 11 minutes to under 2 and an estimated $40 million USD profit improvement for 2024. It was the most cited AI customer service success story of the year.
A year later, the picture had shifted. Klarna confirmed it was hiring humans back for customer service, with the CEO acknowledging that the firm had initially embraced AI with an eye toward cost savings and efficiency, but had perhaps underestimated the tradeoff. The launch numbers were real. The sustained outcome was more nuanced. If anyone had stopped measuring after the first month, they would have a wildly inflated view of what the deployment actually delivered over a full operating cycle.
This is not a criticism of Klarna. It is the opposite. They were one of the few companies that actually measured. Most enterprises never get the data to know whether their initial deployment story held up.
The right metrics depend on the function. Support workloads care about deflection rate, average resolution time, first contact resolution, escalation rate, customer satisfaction on AI handled cases, and cost per ticket. Sales and marketing workloads care about lead quality, conversion rate, time to first response, pipeline velocity, and revenue per rep. Operations workloads care about cycle time, error rate, throughput per FTE, and cost per transaction. Knowledge work cares about hours saved per role, quality of output as judged by the receiver, and the rate at which AI generated work is accepted versus reworked.
The most common trap is mistaking activity for outcome. Number of model calls is not a metric of value. Number of users with access is not adoption. Volume of conversations handled is not deflection. These are inputs that vendors love to put in dashboards because they are easy to grow. Salesforce learned this the hard way with the early Agentforce pricing. The original $2 per conversation model produced impressive activity numbers and disappointing revenue, because a single customer query could trigger eight backend processes, buyers could not model their costs, and the model immediately priced out anyone without a large AI budget. They moved to a per action consumption model instead, and the resulting structure has driven Agentforce to $540 million ARR by Q3 FY2026, growing 330% year over year. The lesson is that even well intentioned proxy metrics can mislead the people designing the contract, let alone the people consuming the dashboard.
The second trap is pilot ROI. A pilot deployment runs in a clean data environment, with a forgiving user group and a hand picked use case. Production requires 99.9% uptime, messy real world data, robust integration and security frameworks. ROI numbers from a pilot do not translate to production. Buyers who use pilot ROI to justify a wider rollout are usually disappointed, and the disappointment is the vendor’s reputational problem only briefly, because the contract is already signed.
The third trap is optimising the wrong proxy. If you measure your support AI on deflection rate alone, the system will route increasingly complex tickets away from humans even when those tickets need humans, and your CSAT will fall in ways that take a quarter to surface. If you measure your sales AI on number of leads contacted, the quality of contact will decay until the brand starts to suffer. The right approach is to track the proxy metric you optimise against alongside the customer outcome metric you actually care about. When the two diverge, you are looking at a system that is gaming itself.
The fourth trap is hidden cost erosion. A deployment that looks profitable in month three can be loss making by month nine because inference costs scaled non linearly, the human review team grew faster than expected, and the prompt engineering work that started as a side project became a permanent function nobody had budgeted. The cost side of the ROI equation needs the same instrumentation as the value side, and most enterprises only instrument one of the two.
Everything in the previous sections leads to one structural conclusion. If measurement is hard, if vendors are not paid to do it, if the data needed to do it sits with the buyer, and if the timelines required to do it honestly are longer than any standard contract, then the contract model itself is the problem.
The market is starting to feel its way towards this. Salesforce has now run through three pricing models for Agentforce in eighteen months. Sierra has built its entire AI agent business around outcome based pricing, arguing publicly that the per seat model that defined SaaS does not survive the agent era. Salesforce’s own research found that 90% of CIOs report that managing AI costs is limiting their ability to drive value, and that organisations are increasingly seeking pricing models aligned to how AI agents deliver business outcomes.
The IT services industry is feeling the same pressure. Accenture booked $1.5 billion in new GenAI bookings in a single quarter, while TCS reported a $1.5 billion AI and GenAI pipeline as of the end of June, with management noting that some engagements are now being structured around outcomes rather than time and materials. If AI makes one consultant as productive as five, revenue per project shrinks, and firms are experimenting with fixed fee or outcome based pricing to adapt.
These moves are awkward, partial, and often hedged. They are also unmistakably in one direction.
What Accubits has done is to build the operating model around this end state from the start, rather than retrofit it onto an existing services business. Outcome as a Service is not a pricing tweak on top of a software development practice. It is a different commercial commitment. The engagement begins with the buyer naming the outcome they want to move. Accubits then designs the agentic system, deploys it, runs it, measures it, and tunes it, with fees structured against the realised outcome rather than against deliverables shipped.

What makes this commercially possible is the underlying technology shift. Agentic systems can be observed continuously, tuned without rebuilds, and improved against measured outcomes in production. The marginal cost of iteration is low compared to traditional software development cycles. A vendor who is paid on outcome can afford to keep improving the system for as long as the outcome metric has room to move, because every percentage point of improvement is shared. The vendor’s economic interest aligns with the buyer’s operational reality. That alignment did not exist in the old build and bill world, because the only thing the vendor was paid to ship was the build.
The contractual implications are real. Outcome based engagement requires data sharing, joint instrumentation, jointly agreed baselines, jointly agreed measurement windows, and dispute resolution mechanisms when the numbers are contested. None of that is trivial. The reason most vendors avoid it is that it is harder than billing for hours. The reason it is worth doing is that it is the only structure that aligns both parties around the question every CFO is now asking, which is what did we get for this.
The era of vendors saying we delivered the system and walking away from the question of whether it worked is ending. It is ending slowly, against vendor resistance, but it is ending.
The test for any AI partner now is simple. Ask them what fee they will put at risk against a measurable business outcome. Ask them what baseline they accept and what measurement window they will be evaluated against. Ask them how they will share data and instrumentation with you. Ask them what happens if the outcome does not move. Ask them how their compensation changes if it moves more than expected. The answers will tell you very quickly whether you are talking to a builder who wants to be paid for delivery, or a partner who is willing to be paid for results.
The 95% failure number from MIT will probably look different by 2027. The reason will not be that the models got dramatically better, because they were never the bottleneck. It will be that the contracts got more honest, the measurement got more disciplined, and the vendors who survived the shakeout were the ones willing to be paid for what they actually changed.
That shift is overdue. It is what Outcome as a Service was built for.