Loading
This is a hypothetical case study. The client, target and figures below are illustrative. They are modelled on published industry benchmarks and on how we run Outcome as a Service engagements, to show how an outcome contract would work on a real drug discovery problem. No client data is used. For a completed engagement, read how Outcome as a Service helped a nutraceutical brand build a supplement the body absorbs 10× better.

Early drug discovery spends most of its time searching. Before a team can optimise a molecule, it has to find one worth optimising, and that search, known as hit finding or screening, usually means physically testing hundreds of thousands of compounds to find a few dozen that do something useful.

This case study asks one question: what if a mid-sized biotech could go from an assay-ready target to confirmed, chemically distinct hit series in 8 weeks instead of 40, and only pay in full if that happened? Below we set out the challenge, the outcome contract, the AI system we would build, the projected results and the risks that could stop it working.

At a glance · projected
80%
less time from assay-ready target to confirmed hit series
8 vs 40 wks
elapsed screening time, AI-first vs conventional HTS
1,200
compounds tested in the lab, down from 500,000
~670×
higher confirmed hit rate (8% vs 0.012%)
−74%
screening cost: $0.62M against a $2.4M baseline

The challenge: screening is slow because it is physical

The client in this scenario is a 120-person oncology biotech with a validated kinase target, a working biochemical assay and a board that wants a development candidate within two years. Its plan of record was a conventional high-throughput screening (HTS) campaign against a 500,000-compound library, followed by hit confirmation, counter-screens and triage.

That plan works, and it has worked for decades. It is also slow and expensive. Industry analyses put the capitalised cost of bringing one new drug to market at around $2.6 billion, and the hit-to-lead stages alone at well over a year of elapsed time per programme (Paul et al., Nature Reviews Drug Discovery). Every month spent screening is a month the patent clock runs without a candidate to show for it.

When we broke the client's 40-week plan into stages, most of the time was not spent running the screen. It went on getting compounds into plates, and then on working out which of the hits were real.

Where the 40 weeks of a conventional HTS campaign go
Elapsed weeks per stage, assay-ready target to confirmed hit series. Hover a segment for detail.
Library preparation Primary screen Hit picking & retest Confirmation & counter-screens Triage & series selection
Illustrative baseline for a single-target biochemical HTS campaign at a mid-sized biotech using an external screening partner. 62% of the elapsed time sits after the primary screen.

Three problems stood out:

  • Most of the library is irrelevant to the target. A diverse HTS library is built to cover many targets, so almost all of it has no chance against any one of them.
  • Hit rates are tiny and noisy. A typical primary screen flags around 0.1% of compounds, and many of those turn out to be assay artefacts. In our baseline, 500 primary hits shrink to about 60 confirmed actives, a confirmed hit rate of 0.012%.
  • Chemical space is far bigger than any library. The number of drug-like molecules is estimated at around 1060 (chemical space). A 500,000-compound deck samples a vanishingly small corner of it, while make-on-demand catalogues such as Enamine REAL now list billions of synthesisable compounds that no one could screen physically.

The client had looked at AI-driven virtual screening before and was sceptical, for reasons we hear often: a vendor's model does well on a benchmark, the programme pays for the platform and the people, and at the end nobody can say whether the hits are better than HTS would have found. Our blog post on the three most common reasons AI initiatives fail covers that pattern in more detail.

Why Outcome as a Service fits this problem

The client did not want a platform licence or a team of data scientists billed by the hour. It wanted confirmed hit series, sooner. That is a result you can baseline and verify, which is the test we apply before taking on any Outcome as a Service engagement.

So the contract would be written around the result. Accubits takes responsibility for the approach, the models, the compute and the work of changing course if the first method does not perform. The client keeps what it is best placed to do: run the assays, own the chemistry decisions and verify the readings independently. Part of our fee is held back and paid only when the outcome gates are met.

How the same programme looks under three ways of buying it
Time & materials AI teamFee-for-service CRO screenOutcome as a Service
What is boughtData scientists’ hoursA screen of N compoundsConfirmed hit series by a date
Who owns the methodClient directs the workFixed protocolAccubits, within agreed guardrails
If the first approach failsClient pays for the reworkClient pays for a second screenAccubits absorbs the cost of changing approach
What settles the invoiceTimesheetsPlates runGates verified in the client's own lab

Defining the outcome: six measurable gates

Before any modelling starts, the goal is turned into numbers that both sides agree on and that the client can measure without us. Each gate is a pass or fail threshold.

#GateBaseline (HTS plan)TargetVerified by
1Time to confirmed hit series40 weeks≤ 8 weeksContract dates
2Chemically distinct series with IC50 < 1 µM in an orthogonal assay~3≥ 3Client's assay team
3Compounds physically tested500,000≤ 2,000Purchase and plate records
4Model enrichment on held-out actives (EF1%)1 (random)≥ 15Blinded retrospective test
5Series with ≥ 10× selectivity over the two closest off-target kinasesUnknown until later≥ 2Client's selectivity panel
6Total screening spend$2.4M≤ 35% of baselineInvoices

Gate 4 matters most for trust. Before a single compound is ordered, the model has to show on data it has never seen that it ranks known actives well above chance. If it cannot, the programme stops early and cheaply, before the expensive part begins.

What Accubits would build

The system is not one model. It is a pipeline of specialised models, each removing a different kind of bad candidate, followed by a decision layer where medicinal chemists stay in charge. We have built this kind of layered system before for regulated life-sciences work, for example our regulatory intelligence solution that automates literature analysis for a pharmaceutical company.

Architecture of the AI-first screening pipeline
Data flows left to right; lab results flow back into the models after every wave.
1 · Data
Target structuresPDB co-crystals + AlphaFold models
Public bioactivityChEMBL, PubChem
Client assay historystays in the client's cloud
Make-on-demand space~4.2 billion compounds
2 · Filter
Property & liability modelsADMET, PAINS, reactivity
Synthesisability scoringdelivery in ≤ 4 weeks
3 · Score
Active-learning dockingML surrogate docks ~1% of space
Rescoring ensembleML + physics; free-energy on the top 200
Selectivity modelvs closest off-target kinases
4 · Decide
Multi-objective rankingpotency, selectivity, novelty, cost
Chemist-in-the-loop reviewevery pick explained
Wet lab (client)dose-response, counter-screens
↺ Feedback loop: each wave of lab results retrains the scoring models before the next order is placed
Built on open, well-tested components where they exist, such as RDKit for cheminformatics and Open Free Energy for free-energy calculations, with the orchestration, active-learning and ranking layers built by Accubits.

1. Assemble the evidence. Target structures come from the Protein Data Bank and, where co-crystals are missing, from AlphaFold predictions. Known actives and inactives for the target family come from public bioactivity data and the client's own assay history, which never leaves the client's environment.

2. Filter out what could never become a drug. Property and liability models remove compounds with poor drug-likeness, known assay-interference motifs or reactive groups. Synthesisability scoring keeps only compounds that can be delivered within four weeks. This takes the space from billions to hundreds of millions without any docking.

3. Score intelligently, not exhaustively. Docking billions of molecules is too slow. An active-learning loop docks a sample, trains a fast surrogate model on the results and uses it to decide what to dock next, so only about 1% of the filtered space is ever docked for real. A rescoring ensemble of ML and physics-based methods re-ranks the best 25,000, with free-energy calculations on the top 200.

4. Let chemists decide. The final ranking balances potency, predicted selectivity, novelty and cost. Every pick comes with the reasons behind it, following the principles in our post on why explainable AI matters, so chemists can overrule the model and have that decision fed back into it.

From 4.2 billion possible molecules to 5 hit series
Candidates remaining after each stage. Bar length is on a log scale; hover for detail.
4.2B
Make-on-demand spaceEnumerated, never physical
310M
Property & liability filters93% removed
2.1M
Active-learning docking0.7% docked for real
25K
Rescoring ensembleML + physics
1,200
Chemist-reviewed picksordered in two waves
↓ in silico above · in the lab below ↓
96
Confirmed actives8% hit rate
5
Hit seriesIC50 < 1 µM, orthogonal assay
Projected counts for the hypothetical programme. For comparison, the HTS plan physically tests 500,000 compounds to reach roughly 60 confirmed actives and 3 series.

Proving the model before spending on the lab

Gate 4 is checked in the first two weeks. The team holds back a set of known actives for the target family that the models never see in training, mixes them into a large set of decoys, and asks the pipeline to rank everything. The enrichment curve below shows how many of the hidden actives appear as you move down the ranked list, compared with picking at random.

Retrospective enrichment: known actives recovered vs share of library screened
Blinded hold-out test, first 20% of the ranked list. Hover a point for its value.
AI pipeline rankingRandom selection
At 1% of the library screened, the AI ranking recovers 38% of known actives against 1% for random selection. At 5% it recovers 71%, at 10% 84% and at 20% 93%. 0% 25% 50% 75% 100% 0% 5% 10% 15% 20% Share of ranked library screened EF1% = 38 38% of actives in the top 1% Random: 20% AI: 93%
An enrichment factor (EF1%) of 38 means the top 1% of the ranked list holds 38 times more actives than a random 1% would. The gate was ≥ 15. Projected result, hypothetical data.

The 8-week plan

With the model proven, the programme runs as a tight sequence. Compound synthesis is the longest single step, so it starts as early as possible, and lab results from the first wave retrain the models while the second wave is still being made.

Engagement timeline, weeks 0–8
Red: Accubits-led work. Grey: client lab work. Hover a bar for detail.
W1W2W3W4W5W6W7W8
Data assembly & model training
Gate 4: enrichment verified
Virtual screen & ranking
Chemist review & wave 1 order
Make-on-demand synthesis
Active-learning round 2
Dose-response, counter & selectivity assays
Series selection & handover
Outcome reading & settlement
Week 0 is the day the outcome contract is signed and the assay is confirmed ready. Outcome definition and baselining happen before that date and are not counted.

Projected results

Time: 40 weeks down to 8

Every stage gets shorter, but not equally. The largest savings come after the primary screen: with 1,200 well-chosen compounds instead of 500 noisy primary hits, confirmation and triage take a few weeks instead of five months.

Elapsed weeks per stage: conventional HTS vs AI-first screening
Hover a bar for what drives the difference.
Conventional HTSAI-first (Outcome as a Service)
Library preparation
6
1.5
Primary screen
8
1
Compound sourcing / hit picking
6
3
Confirmation & counter-screens
10
2
Triage & series selection
10
0.5
Total40 weeks → 8 weeks (−80%)
Projected, hypothetical. The AI-first times assume compounds ordered from a make-on-demand supplier with a 3–4 week delivery window.

Cost: about a quarter of the baseline

The savings come from not buying, storing, plating and screening half a million compounds. The AI-first programme spends more on computation and specialist people, but those are small numbers next to a full HTS campaign.

Screening cost breakdown, USD millions
Bars are drawn to the same scale. Hover a segment for its value.
Conventional HTS$2.40M
Library & logisticsPrimary screenConfirmationTriage staff
AI-first (Outcome as a Service)$0.62M (−74%)
Compute $0.06MAI models & team $0.16MCompounds $0.18MLab assays $0.22M
Illustrative costs for a single-target campaign. Real figures vary with the target class, assay format and supplier pricing.

Every gate cleared

Projected result against each outcome gate
Red bar: projected result. Black tick: contract target. Scales run from 0 to twice the target, or from the baseline where lower is better.
1 · Time to hit series · 8 weeks vs target ≤ 8Met
040 weeks (baseline)
2 · Sub-µM series · 5 vs target ≥ 3Met
06
3 · Compounds tested · 1,200 vs target ≤ 2,000Met
04,000
4 · Enrichment EF1% · 38 vs target ≥ 15Met
030+
5 · Selective series (≥ 10×) · 3 vs target ≥ 2Met
04
6 · Screening spend · 26% of baseline vs target ≤ 35%Met
0100% ($2.4M)
For gates 1, 3 and 6 lower is better, so the red bar should end left of the tick. For gates 2, 4 and 5 higher is better, so it should end to the right. Projected, hypothetical.

What could go wrong

A hypothetical case study that only shows the upside is not much use. These are the risks we would write into the contract, and what happens if each one occurs.

The model does not beat chanceIf gate 4 fails in week 1.5, the programme stops before any compounds are bought. Client exposure: about two weeks and a small setup fee.
Poor structural dataSome targets have no good co-crystal and a low-confidence predicted structure. We would switch to ligand-based models trained on the target family, and say so at the baseline stage.
Hits that are real but not novelModels trained on public data can rediscover known chemotypes. The novelty term in the ranking and a patent-landscape check on every series reduce this risk.
Supplier delaysMake-on-demand synthesis is the longest AI-first step. Ordering in two waves, from more than one supplier, keeps a late batch from moving the gate date.
Assay drift in the client's labBecause the client runs the verifying assays, controls and reference compounds are agreed up front so the reading can be trusted by both sides.
Data governanceProprietary assay data stays in the client's cloud tenancy. Models are trained there and the weights belong to the client. Our post on cloud and big data in pharma covers the setup.

What this would mean for the client

  • Over seven months back on the programme clock. Thirty-two weeks saved at the start of discovery moves every later milestone forward, including patent filing and the first data room for investors.
  • Better starting points, not just faster ones. Five series with selectivity data already attached give lead optimisation more options and fewer surprises.
  • A reusable asset. The trained models and the active-learning pipeline stay with the client for the next target. Running them repeatably is an engineering problem in its own right, which we cover in why scaling AI is fundamentally different from building it.
  • Risk shared, not just moved. The client pays in full only when the result is verified in its own lab.

Key takeaways

  • Most screening time is spent after the screen. AI saves the most time by sending fewer, better compounds into confirmation and triage, not by running the primary screen faster.
  • Prove the model before the lab spends money. A blinded enrichment gate in week two turns an uncertain AI bet into a cheap go/no-go decision.
  • Outcome contracts suit discovery problems. “Confirmed hit series by week 8” is measurable, verifiable and valuable, which is what an outcome needs to be.
  • Chemists stay in charge. The system ranks and explains. People decide what gets made.

Frequently asked questions

Is this a real client engagement?

No. This is a hypothetical case study. The client, target and numbers are illustrative and are based on published industry benchmarks and on how Accubits structures Outcome as a Service contracts. For a completed engagement, see the nutraceutical absorption case study.

How can AI cut drug discovery screening time by 80%?

Most of a conventional screening campaign is spent handling physical compounds and ruling out false positives. An AI-first pipeline searches billions of virtual compounds computationally, filters out likely artefacts before anything is made, and sends only around 1,200 high-probability candidates to the lab. Confirmation and triage then take weeks instead of months.

Does AI virtual screening replace wet-lab experiments?

No. It decides what is worth testing. Every hit in this scenario is confirmed in dose-response, orthogonal and selectivity assays run by the client. The AI changes how many compounds reach the lab and how good they are, not whether the lab is needed.

What is Outcome as a Service?

Outcome as a Service is Accubits’ engagement model for problems where the result matters more than the process. The client and Accubits agree a measurable outcome and how it will be verified. Accubits chooses and builds the approach, and part of the fee depends on the verified result.

What happens if the AI model does not perform?

The contract includes an early enrichment gate. If the model cannot rank known actives well above chance in a blinded test, the programme stops in about two weeks, before compounds are bought or lab time is used. The at-risk part of Accubits’ fee is not paid.

Who owns the models and the data?

The client. Proprietary assay data stays in the client's cloud environment, models are trained there, and the trained models and pipeline are handed over at the end of the engagement.

Related reading

References

  1. DiMasi, J. A., Grabowski, H. G. & Hansen, R. W. (2016). Innovation in the pharmaceutical industry: New estimates of R&D costs. Journal of Health Economics, 47, 20–33.
  2. Paul, S. M. et al. (2010). How to improve R&D productivity: the pharmaceutical industry's grand challenge. Nature Reviews Drug Discovery, 9, 203–214.
  3. Jumper, J. et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589.
  4. RCSB Protein Data Bank and the AlphaFold Protein Structure Database.
  5. ChEMBL (EMBL-EBI) and PubChem (NCBI) bioactivity databases.
  6. Enamine REAL make-on-demand compound collections.
  7. RDKit open-source cheminformatics and Open Free Energy.

Have a discovery bottleneck you can put a number on?

If you can measure it and verify it, we can talk about contracting against it. Tell us the outcome you need, and we will tell you whether we would put part of our fee behind it.

Scope an outcome with us

Or read how the model works on the Outcome as a Service page.