The Future of IT Services Is Measured in Results, Not Hours

Ask most companies how they buy IT services and the answer comes back in units of effort: a team of eight developers, a block of 2,000 consulting hours, a monthly retainer for a managed service, a statement of work priced from an estimate of person-days. The invoice describes how hard the vendor worked. It says almost nothing about whether the business is better off.

That arrangement made sense for a long time. Software was hard to measure, effort was a reasonable proxy for value, and nobody had a better way to price uncertain work. But three things have changed at once. AI has broken the link between effort and output. Finance teams now want every technology line item tied to a number they can defend. And the systems we build are instrumented well enough that the number can actually be measured.

The result is a shift that is already under way: the future of IT services will be bought, priced and judged on results, not hours. In this article we look at why the hourly model is running out of road, what an outcome-based engagement actually looks like on paper, how it played out in two of our own Outcome as a Service case studies, and what both buyers and vendors need to change to make it work.

How We Ended Up Buying Software by the Hour

The hourly model was not an accident. It grew out of a real problem: software projects are uncertain, and someone has to carry that uncertainty. Each generation of IT service contract has been a different answer to the question of who that someone should be.

  • Staff augmentation rents people. The client directs the work, and the vendor’s only obligation is to supply capable hands. All delivery risk stays with the client.
  • Time and materials (T&M) rents a team plus some management. It is flexible and honest about uncertainty, but the client still pays for every hour regardless of what those hours produce.
  • Fixed-price, fixed-scope projects move some risk to the vendor, but only the risk of building the agreed specification. If the specification was the wrong thing to build, the vendor still gets paid in full.
  • Managed services price by service level: uptime, ticket response times, resolution SLAs. They measure the health of the system, not the value it creates.

Every one of these models measures an input (hours, people) or an output (features delivered, tickets closed). None of them measures the outcome: the change in the business that justified spending the money in the first place. That gap is where most of the frustration with IT services comes from.

Why the Hourly Model Is Running Out of Road

1. The incentives point the wrong way

When a vendor is paid by the hour, every hour is revenue. That does not make vendors dishonest; most good teams work hard to deliver value. But the commercial structure quietly rewards the wrong things. A problem solved in two weeks earns less than the same problem solved in two months. A project that reveals early on that it should be stopped earns less than one that carries on. A clever reuse of an existing component earns less than building it again.

Economists call this a principal–agent problem: the person doing the work (the agent) is rewarded for something other than what the person paying (the principal) actually wants. Hourly billing does not create bad vendors, but it never asks good ones to put anything at stake either.

This is the change that makes the shift unavoidable rather than merely attractive. When an AI coding assistant writes a first draft of a service in an afternoon, when an agent triages thousands of documents overnight, or when a screening model narrows billions of candidate molecules to a few thousand in a week, the number of hours stops being a meaningful measure of anything.

A vendor that bills by the hour faces an awkward choice. It can use AI fully and watch its revenue shrink as work gets faster. Or it can under-use AI to protect its billable hours, which means the client pays more for a slower result. Neither is sustainable. The only pricing model that lets both sides benefit from AI-driven productivity is one that pays for what the work achieves, not how long it took.

3. The buyer carries the risk but has the least control

In a T&M engagement the client pays for every wrong turn, yet the vendor usually knows more about which turns are likely to be wrong. The party with the most technical knowledge carries the least financial risk. That is backwards, and it is one reason why so many technology programmes end with a working system and no measurable business impact. We looked at that pattern for AI specifically in The 3 Most Common Reasons AI Initiatives Fail: the demo trap, the measurement trap and the ownership gap.

4. Status reports are not proof

An effort-based engagement is managed through activity: velocity charts, burn-down charts, sprint demos, percentage complete. These tell you the work is happening. They do not tell you it is working. When a CFO asks what the last eighteen months of technology spend returned, “we shipped 340 story points” is not an answer that survives a budget review.

5. Measurement is finally cheap

The old objection to outcome pricing was that results in software are too hard to measure. That was true when systems were opaque. It is much less true today. Modern platforms are instrumented end to end, product analytics and observability are standard, and AI systems in particular can be evaluated continuously against golden datasets and live traffic. If a metric matters to the business, there is now almost always a way to baseline it and track it.

Five Ways to Buy IT Services, Compared

It helps to put the models side by side. The question in each row is the one a buyer should ask before signing anything.

QuestionStaff augmentationTime & materialsFixed scopeManaged serviceOutcome as a Service
What you buyPeopleHours from a teamA specified deliverableA service levelOne agreed business result
Who decides howYouYou, with vendor inputThe specificationThe vendor, within the SLAThe vendor, within agreed guardrails
Who carries delivery riskYou, entirelyMostly youShared; scope changes revert to youShared, for operations onlyMostly the vendor, with fee at stake
If the first approach failsYou pay for the reworkYou pay for the reworkA change requestNot coveredThe vendor absorbs the cost of changing course
What settles the invoiceTimesheetsTimesheetsAcceptance of deliverablesSLA reportsA verified reading of the metric
Vendor is rewarded forHours suppliedHours workedFinishing to specKeeping things runningMoving the number

None of the older models is wrong in every situation. Staff augmentation is still the right answer when you have strong internal leadership and need extra capacity, and our guides on hiring a remote team effectively and working with an offshore development team cover how to do that well. But when the goal is a business result rather than extra hands, the last column is the only one whose incentives point at that result.

What “Measured in Results” Actually Means

“Outcome-based” is easy to say and easy to fake. A contract that pays a bonus for “successful delivery” is not outcome-based; it is fixed-scope with a nicer name. A real outcome contract has a recognisable anatomy, and every part of it is written down before work begins.

  1. One primary metric. The business result the engagement exists to move: cost per claim processed, weeks to a confirmed drug hit, conversion rate, first-contact resolution, absorption efficiency. One number, not a dashboard.
  2. A baseline. Where that metric stands today, measured the same way it will be measured at the end. Without a baseline there is nothing to improve on.
  3. A target and a measurement window. How far the number must move, and over what period it will be read, so a lucky week cannot settle the contract.
  4. A source of truth. The system the reading comes from, ideally one the client controls: the ERP, the CRM, the lab’s own assay results, the finance ledger.
  5. Attribution logic. How the effect of the work is separated from everything else going on, using phased rollouts, control groups, A/B tests or cohort comparisons.
  6. Gates. Intermediate, pass-or-fail thresholds that let both sides stop early and cheaply if the approach is not working.
  7. Guardrails. The things that must not get worse while the main metric improves: quality, safety, compliance, customer satisfaction, cost ceilings.
  8. Fee at risk. The part of the vendor’s fee that is paid only when the verified reading clears the target.

Once those eight things are agreed, something important happens: the vendor gets to own the how. The client is no longer approving line items, choosing frameworks or managing a backlog. It is holding the vendor to a number. That frees the vendor to choose the fastest, cheapest and most reliable route to the result, including routes that look nothing like the original plan.

Two Examples from Our Own Work

Abstract principles are easier to judge against real engagements. Here are two, one completed and one hypothetical, both from discovery-heavy domains where the hourly model works especially badly because nobody knows in advance how many hours the answer will take.

A supplement the body absorbs 10 times better

A nutraceutical brand came to us with a long-standing industry problem. Many valuable ingredients dissolve poorly, break down during digestion or pass through the body unabsorbed, and typical absorption efficiency sits at about 20–30%. Most of every dose is wasted. The client did not want a research report or a block of consulting hours. They wanted a credible answer to one question: is there a better way to deliver these ingredients?

Under Outcome as a Service, the engagement was defined by that answer. We turned the goal into four measurable gates linked to absorption: encapsulation efficiency, active loading, particle size and bioavailability. Then, rather than testing formulations one at a time in a lab, we built an AI-based formulation evaluation system that compared a wide range of candidate processes, including established methods such as liposomal delivery, against those gates.

The system recommended a new nano-formulation approach with projected performance of 92% encapsulation, 23% loading, 10× bioavailability and 48 nm particles, clearing every gate and theoretically delivering about 25% more absorption efficiency than current industry-standard methods. The client now takes one strong, data-backed approach into lab validation instead of working through many options. You can read the full story in How Our Outcome as a Service Helped a Nutraceutical Brand Build a Supplement the Body Absorbs 10X Better.

The point for this article is not the chemistry. It is that nobody priced this engagement by asking how many hours a formulation scientist would need. The method, AI, was chosen because the contract rewarded the answer rather than the effort. Under a T&M contract, a long programme of physical experiments would have been the commercially safer choice for the vendor and the worse one for the client.

What if drug discovery screening took 8 weeks instead of 40?

Our second example is a hypothetical case study on drug discovery screening, built on published industry benchmarks to show how an outcome contract would work on a hard biotech problem. The scenario: a mid-sized oncology biotech with a validated target and a plan of record for a conventional high-throughput screen of 500,000 compounds, taking around 40 weeks and $2.4M to reach confirmed hit series.

Instead of selling a platform licence or a team of data scientists, the contract is written around six gates the client can verify without us: time to confirmed hit series, number of chemically distinct series, compounds physically tested, model enrichment on held-out actives, selectivity and total screening spend. Part of the fee is held back and paid only when those gates are met.

The projected result of an AI-first virtual screening pipeline is striking: 8 weeks instead of 40, 1,200 compounds tested instead of 500,000, a confirmed hit rate around 670 times higher, and screening cost down by about 74%, from $2.4M to $0.62M.

But the most instructive part of that case study is not the headline. It is gate 4. Before a single compound is ordered, the model has to prove on data it has never seen that it ranks known actives well above chance. If it cannot, the programme stops in week two, cheaply, before the expensive lab work begins. That kind of early, honest kill switch almost never appears in an hourly contract, because stopping early is the one result an hourly vendor is never paid for. In an outcome contract, it protects both sides.

What Makes a Good Outcome?

Not every piece of technology work can or should be bought as an outcome. Before taking on any Outcome as a Service engagement, we test the proposed metric against six questions:

  • Is it measurable? Can it be expressed as a number, from a system both sides trust?
  • Can it be baselined? Do we know where it stands today, measured the same way?
  • Is it attributable? Can the effect of the work be separated from seasonality, pricing changes, marketing campaigns and everything else?
  • Is it within reach of the work? Can the engagement realistically move it, or does it depend mostly on factors outside anyone’s control?
  • Is it worth moving? Is the value of the improvement clearly larger than the cost of the engagement?
  • Is it time-bound? Will the result be visible within a measurement window short enough to settle a contract?

Here are some outcomes that usually pass those tests, by function:

FunctionEffort-based askOutcome-based ask
Customer service“Build us a chatbot.”Raise first-contact resolution from 54% to 70% without lowering CSAT.
Operations“Automate the invoice process.”Cut cost per invoice processed by 40%, with an error rate below 0.5%.
Legal & compliance“Give us an AI document tool.”Reduce time to first case summary from three days to four hours, at reviewer-agreed accuracy.
HR“Implement an applicant tracking system.”Reduce time-to-shortlist from 15 days to 3 for high-volume roles.
Research & development“Provide a team of data scientists.”Reach confirmed hit series in 8 weeks instead of 40, at under 35% of baseline cost.
Sales & marketing“Redesign the website.”Increase qualified demo requests per month by 30% over the next two quarters.
Infrastructure“Migrate our AI workloads.”Reduce GenAI inference cost per request by 50% at equal or better quality.

Several of these map directly to work we have delivered. Our LegalDoc AI assistant organises and summarises more than 150 document types; our AI recruitment automation solution targets shortlisting time; our business process automation solution targets processing cost; and our GenAI cost optimisation service helps clients save up to 60% on GenAI infrastructure. Each of those results is easier to buy, and easier to defend internally, when it is the thing the contract is written around.

And some work is a poor fit. Open-ended exploration with no agreed definition of success, platform work whose value only appears years later, or anything where the metric depends mainly on decisions the vendor cannot influence is usually better bought another way. A good outcome partner will tell you so rather than force a contract onto the wrong problem.

How Outcome-Based Pricing Works

“Pay only for results” sounds simple, but almost no serious outcome contract is pure contingency. The vendor still has real costs, and the client still wants predictability. In practice, outcome contracts combine a few common structures:

  • Base fee plus fee at risk. The most common shape. A base fee covers a portion of the vendor’s costs, and a meaningful share is held back and paid only when the verified reading clears the target. This is how our Outcome as a Service engagements are typically structured.
  • Gated payments. The engagement is split into stages, each ending in a pass-or-fail gate. The client can stop after any gate. The drug discovery case study uses this to put the expensive lab work behind a model-quality gate.
  • Gain-share. The vendor receives a percentage of the measured value created, such as cost saved or revenue gained, over an agreed period. It works best when the value is easy to quantify in money.
  • Per-outcome unit pricing. The client pays per unit of result: per claim processed correctly, per qualified lead, per resolved ticket. It is common in AI-driven operations where the volume is high and each unit is easy to verify.

The right shape depends on how quickly the outcome becomes visible, how easy it is to put a money value on it and how much risk each side is comfortable carrying. What every version shares is that the vendor’s payment moves with the client’s result.

The Hard Questions, Answered Honestly

Outcome contracts are not magic, and buyers are right to be sceptical. These are the objections we hear most often.

“Isn’t this just fixed-price with a different name?”

No. A fixed-price contract pays when the specified deliverable is accepted, whether or not it changes anything. An outcome contract pays when the business metric moves. If the vendor delivers exactly what was specified and the number does not move, a fixed-price vendor is paid in full; an outcome vendor is not. That difference is the whole point.

“Won’t vendors just pick easy targets?”

They will try, which is why the baseline and target are agreed jointly and measured from the client’s own systems. An honest baseline, a source of truth the client controls and a target tied to a real business case make sandbagging hard. Gates help too: they are set so that each stage has to prove something specific, not just show progress.

“What about gaming the metric?”

Goodhart’s law is real: when a measure becomes a target, it can stop being a good measure. A support team measured only on handling time will rush customers off the phone. That is what guardrails are for. Every outcome contract should name the things that must not get worse, such as satisfaction, accuracy, compliance and cost ceilings, and the at-risk fee should depend on the guardrails holding as well as the main metric moving.

“What if something on our side causes the miss?”

Outcome contracts need clear dependencies. If the result depends on the client providing data access, running lab assays on time or rolling out a new workflow to staff, those obligations are written into the contract alongside the vendor’s. In the drug discovery scenario, the client runs the assays and owns the chemistry decisions; Accubits owns the models, compute and approach. Each side is accountable for what it controls.

“How do we know the AI did it?”

This is the attribution question, and it should be answered before work begins, not argued over at the end. Phased rollouts, control groups and before-and-after cohorts all work. For AI systems in particular, transparency matters as well: if people can see why the system made a recommendation, they can trust the result it is credited with. Our post on why explainable AI matters goes deeper into that.

“Isn’t it more expensive?”

Sometimes the headline fee is higher, because the vendor is pricing in the risk it is taking on. But the comparison that matters is cost per result, not cost per hour. A T&M engagement that runs for a year and moves nothing is far more expensive than an outcome engagement that costs more per month and hits its target in a quarter. In the hypothetical drug discovery case, the outcome-based programme costs about a quarter of the conventional approach, because the vendor is free to use the cheapest route that works.

What Changes for the Vendor

Selling results instead of hours is not a pricing change. It changes how the work is done, and vendors who treat it as a sales tactic will lose money quickly. In our experience, five habits separate vendors who can deliver outcomes from those who cannot.

  1. Measure first. The first weeks of an outcome engagement go into baselining, instrumentation and agreeing how success will be read. Building starts once there is something to measure against.
  2. Prove the approach before scaling the spend. Early gates and offline evaluation catch a weak approach while it is still cheap to change. For AI work, this is the principle at the heart of our Agent Development Lifecycle: evaluation is the product, and an accuracy threshold is a deployment gate.
  3. Be willing to change course. When the vendor absorbs the cost of a failed approach, it has every reason to notice failure early and switch. That flexibility is worth more to the client than any fixed plan.
  4. Use the most productive tools available. AI-assisted engineering, reusable components and automation are no longer a threat to the vendor’s revenue; they are how it protects its margin. The client gets the benefit as speed.
  5. Stay until the number moves. A system that works but is not adopted does not move the metric. Outcome vendors care about training, workflow change, monitoring and post-launch improvement, because that is where the result is actually won. We covered this “last mile” in Why Scaling AI Is Fundamentally Different from Building AI, and the operational side in What is LLMOps?

This also changes the shape of the vendor’s team. The classic pyramid, with a few senior people billing out large numbers of junior hours, makes less sense when the work is not priced by headcount. Outcome delivery rewards smaller teams of senior engineers, domain specialists and evaluation experts, amplified by AI tools. Expect IT services firms to look less like staffing agencies and more like product companies over the next few years.

What Changes for the Buyer

Buyers have work to do as well. An outcome contract asks the client to be clearer about what it wants than most statements of work ever do.

  • Know the business case. Be able to say what moving the metric is worth, so the price of the outcome can be judged against its value.
  • Share real data early. Baselines need honest numbers. Outcome partners cannot price risk they cannot see.
  • Name an owner. Someone on the client side must own the metric, the dependencies and the measurement, and have the authority to act on them.
  • Let go of the how. The point of buying an outcome is to stop managing the method. Set guardrails, then let the vendor choose the route.
  • Change procurement. Many procurement processes are built to compare day rates. Comparing outcome proposals means comparing cost per result, risk carried and the credibility of each vendor’s measurement plan.

Our older guide to choosing outsourcing partners still applies here. Add one more question to the list: what is this vendor willing to put at stake?

A Checklist Before You Sign an Outcome Contract

Whether you are buying from us or from anyone else, a credible outcome-based proposal should let you answer “yes” to each of these:

  1. Metric: Is there one primary business metric, defined precisely enough that two people would measure it the same way?
  2. Baseline: Has the current value been measured from your own systems, over a representative period?
  3. Target and window: Is the target specific, and is the measurement window long enough to rule out luck?
  4. Source of truth: Will the reading come from a system you control?
  5. Attribution: Is there an agreed method for separating the effect of the work from everything else?
  6. Gates: Are there early, cheap points where either side can stop if the approach is not working?
  7. Guardrails: Are the things that must not get worse named and measured?
  8. Fee at risk: Is a meaningful share of the vendor’s fee tied to the verified result?
  9. Dependencies: Are your own obligations written down as clearly as the vendor’s?
  10. After the reading: Is it clear who runs, monitors and improves the system once the contract settles?

If a proposal cannot answer most of those, it is probably an hourly contract wearing an outcome-based label.

The Transition Will Be Gradual, and Then Sudden

Hourly contracts will not disappear overnight. There will always be work that is best bought as capacity, and many organisations will run a mix: staff augmentation for steady engineering capacity, managed services for operations, and outcome contracts for the initiatives that matter most to the business.

But the direction is clear. As AI keeps compressing the time it takes to do the work, the hour becomes a less and less meaningful unit to buy. Clients will increasingly ask why they are paying for effort that a vendor’s own tools have made unnecessary. Vendors that can price, deliver and prove results will win those conversations. Vendors that can only sell capacity will find themselves competing on day rate against software.

The best way to prepare is to start small. Pick one initiative with a clear metric, an honest baseline and a business case everyone agrees on, and buy it as an outcome. The discipline of defining that contract, deciding what success means and how you will know, is valuable even before any work begins.

How Accubits Can Help

Accubits has delivered more than 500 projects, including over 130 AI and data engagements, across more than 12 years. We built Outcome as a Service because we believe the future of IT services is measured in results, and we would rather be judged on that than on our timesheets.

You can see more of our work in our case studies, including a regulatory intelligence solution for a pharmaceutical company. If you have a result you need and you are tired of paying for hours that do not add up to it, talk to our team.

Written by

Accubits

Accubits Technologies is a full-service software provider enabling Federal agencies, Fortune 500 companies, Tech startups, and Enterprises to accelerate their business growth with bleeding-edge technology and solutions. Specializing in Artificial Intelligence and Blockchain technologies, Accubits helps organizations to​ be future-proof​ through data-driven solutions for mobile, cloud, and web platforms.

More from Accubits →