Early drug discovery spends most of its time searching. Before a team can optimise a molecule, it has to find one worth optimising, and that search, known as hit finding or screening, usually means physically testing hundreds of thousands of compounds to find a few dozen that do something useful.
This case study asks one question: what if a mid-sized biotech could go from an assay-ready target to confirmed, chemically distinct hit series in 8 weeks instead of 40, and only pay in full if that happened? Below we set out the challenge, the outcome contract, the AI system we would build, the projected results and the risks that could stop it working.
The challenge: screening is slow because it is physical
The client in this scenario is a 120-person oncology biotech with a validated kinase target, a working biochemical assay and a board that wants a development candidate within two years. Its plan of record was a conventional high-throughput screening (HTS) campaign against a 500,000-compound library, followed by hit confirmation, counter-screens and triage.
That plan works, and it has worked for decades. It is also slow and expensive. Industry analyses put the capitalised cost of bringing one new drug to market at around $2.6 billion, and the hit-to-lead stages alone at well over a year of elapsed time per programme (Paul et al., Nature Reviews Drug Discovery). Every month spent screening is a month the patent clock runs without a candidate to show for it.
When we broke the client's 40-week plan into stages, most of the time was not spent running the screen. It went on getting compounds into plates, and then on working out which of the hits were real.
Three problems stood out:
- Most of the library is irrelevant to the target. A diverse HTS library is built to cover many targets, so almost all of it has no chance against any one of them.
- Hit rates are tiny and noisy. A typical primary screen flags around 0.1% of compounds, and many of those turn out to be assay artefacts. In our baseline, 500 primary hits shrink to about 60 confirmed actives, a confirmed hit rate of 0.012%.
- Chemical space is far bigger than any library. The number of drug-like molecules is estimated at around 1060 (chemical space). A 500,000-compound deck samples a vanishingly small corner of it, while make-on-demand catalogues such as Enamine REAL now list billions of synthesisable compounds that no one could screen physically.
The client had looked at AI-driven virtual screening before and was sceptical, for reasons we hear often: a vendor's model does well on a benchmark, the programme pays for the platform and the people, and at the end nobody can say whether the hits are better than HTS would have found. Our blog post on the three most common reasons AI initiatives fail covers that pattern in more detail.
Why Outcome as a Service fits this problem
The client did not want a platform licence or a team of data scientists billed by the hour. It wanted confirmed hit series, sooner. That is a result you can baseline and verify, which is the test we apply before taking on any Outcome as a Service engagement.
So the contract would be written around the result. Accubits takes responsibility for the approach, the models, the compute and the work of changing course if the first method does not perform. The client keeps what it is best placed to do: run the assays, own the chemistry decisions and verify the readings independently. Part of our fee is held back and paid only when the outcome gates are met.
| Time & materials AI team | Fee-for-service CRO screen | Outcome as a Service | |
|---|---|---|---|
| What is bought | Data scientists’ hours | A screen of N compounds | Confirmed hit series by a date |
| Who owns the method | Client directs the work | Fixed protocol | Accubits, within agreed guardrails |
| If the first approach fails | Client pays for the rework | Client pays for a second screen | Accubits absorbs the cost of changing approach |
| What settles the invoice | Timesheets | Plates run | Gates verified in the client's own lab |
Defining the outcome: six measurable gates
Before any modelling starts, the goal is turned into numbers that both sides agree on and that the client can measure without us. Each gate is a pass or fail threshold.
| # | Gate | Baseline (HTS plan) | Target | Verified by |
|---|---|---|---|---|
| 1 | Time to confirmed hit series | 40 weeks | ≤ 8 weeks | Contract dates |
| 2 | Chemically distinct series with IC50 < 1 µM in an orthogonal assay | ~3 | ≥ 3 | Client's assay team |
| 3 | Compounds physically tested | 500,000 | ≤ 2,000 | Purchase and plate records |
| 4 | Model enrichment on held-out actives (EF1%) | 1 (random) | ≥ 15 | Blinded retrospective test |
| 5 | Series with ≥ 10× selectivity over the two closest off-target kinases | Unknown until later | ≥ 2 | Client's selectivity panel |
| 6 | Total screening spend | $2.4M | ≤ 35% of baseline | Invoices |
Gate 4 matters most for trust. Before a single compound is ordered, the model has to show on data it has never seen that it ranks known actives well above chance. If it cannot, the programme stops early and cheaply, before the expensive part begins.
What Accubits would build
The system is not one model. It is a pipeline of specialised models, each removing a different kind of bad candidate, followed by a decision layer where medicinal chemists stay in charge. We have built this kind of layered system before for regulated life-sciences work, for example our regulatory intelligence solution that automates literature analysis for a pharmaceutical company.
1. Assemble the evidence. Target structures come from the Protein Data Bank and, where co-crystals are missing, from AlphaFold predictions. Known actives and inactives for the target family come from public bioactivity data and the client's own assay history, which never leaves the client's environment.
2. Filter out what could never become a drug. Property and liability models remove compounds with poor drug-likeness, known assay-interference motifs or reactive groups. Synthesisability scoring keeps only compounds that can be delivered within four weeks. This takes the space from billions to hundreds of millions without any docking.
3. Score intelligently, not exhaustively. Docking billions of molecules is too slow. An active-learning loop docks a sample, trains a fast surrogate model on the results and uses it to decide what to dock next, so only about 1% of the filtered space is ever docked for real. A rescoring ensemble of ML and physics-based methods re-ranks the best 25,000, with free-energy calculations on the top 200.
4. Let chemists decide. The final ranking balances potency, predicted selectivity, novelty and cost. Every pick comes with the reasons behind it, following the principles in our post on why explainable AI matters, so chemists can overrule the model and have that decision fed back into it.
Proving the model before spending on the lab
Gate 4 is checked in the first two weeks. The team holds back a set of known actives for the target family that the models never see in training, mixes them into a large set of decoys, and asks the pipeline to rank everything. The enrichment curve below shows how many of the hidden actives appear as you move down the ranked list, compared with picking at random.
The 8-week plan
With the model proven, the programme runs as a tight sequence. Compound synthesis is the longest single step, so it starts as early as possible, and lab results from the first wave retrain the models while the second wave is still being made.
Projected results
Time: 40 weeks down to 8
Every stage gets shorter, but not equally. The largest savings come after the primary screen: with 1,200 well-chosen compounds instead of 500 noisy primary hits, confirmation and triage take a few weeks instead of five months.
Cost: about a quarter of the baseline
The savings come from not buying, storing, plating and screening half a million compounds. The AI-first programme spends more on computation and specialist people, but those are small numbers next to a full HTS campaign.
Every gate cleared
What could go wrong
A hypothetical case study that only shows the upside is not much use. These are the risks we would write into the contract, and what happens if each one occurs.
What this would mean for the client
- Over seven months back on the programme clock. Thirty-two weeks saved at the start of discovery moves every later milestone forward, including patent filing and the first data room for investors.
- Better starting points, not just faster ones. Five series with selectivity data already attached give lead optimisation more options and fewer surprises.
- A reusable asset. The trained models and the active-learning pipeline stay with the client for the next target. Running them repeatably is an engineering problem in its own right, which we cover in why scaling AI is fundamentally different from building it.
- Risk shared, not just moved. The client pays in full only when the result is verified in its own lab.
Key takeaways
- Most screening time is spent after the screen. AI saves the most time by sending fewer, better compounds into confirmation and triage, not by running the primary screen faster.
- Prove the model before the lab spends money. A blinded enrichment gate in week two turns an uncertain AI bet into a cheap go/no-go decision.
- Outcome contracts suit discovery problems. “Confirmed hit series by week 8” is measurable, verifiable and valuable, which is what an outcome needs to be.
- Chemists stay in charge. The system ranks and explains. People decide what gets made.
Frequently asked questions
Is this a real client engagement?
No. This is a hypothetical case study. The client, target and numbers are illustrative and are based on published industry benchmarks and on how Accubits structures Outcome as a Service contracts. For a completed engagement, see the nutraceutical absorption case study.
How can AI cut drug discovery screening time by 80%?
Most of a conventional screening campaign is spent handling physical compounds and ruling out false positives. An AI-first pipeline searches billions of virtual compounds computationally, filters out likely artefacts before anything is made, and sends only around 1,200 high-probability candidates to the lab. Confirmation and triage then take weeks instead of months.
Does AI virtual screening replace wet-lab experiments?
No. It decides what is worth testing. Every hit in this scenario is confirmed in dose-response, orthogonal and selectivity assays run by the client. The AI changes how many compounds reach the lab and how good they are, not whether the lab is needed.
What is Outcome as a Service?
Outcome as a Service is Accubits’ engagement model for problems where the result matters more than the process. The client and Accubits agree a measurable outcome and how it will be verified. Accubits chooses and builds the approach, and part of the fee depends on the verified result.
What happens if the AI model does not perform?
The contract includes an early enrichment gate. If the model cannot rank known actives well above chance in a blinded test, the programme stops in about two weeks, before compounds are bought or lab time is used. The at-risk part of Accubits’ fee is not paid.
Who owns the models and the data?
The client. Proprietary assay data stays in the client's cloud environment, models are trained there, and the trained models and pipeline are handed over at the end of the engagement.
Related reading
References
- DiMasi, J. A., Grabowski, H. G. & Hansen, R. W. (2016). Innovation in the pharmaceutical industry: New estimates of R&D costs. Journal of Health Economics, 47, 20–33.
- Paul, S. M. et al. (2010). How to improve R&D productivity: the pharmaceutical industry's grand challenge. Nature Reviews Drug Discovery, 9, 203–214.
- Jumper, J. et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589.
- RCSB Protein Data Bank and the AlphaFold Protein Structure Database.
- ChEMBL (EMBL-EBI) and PubChem (NCBI) bioactivity databases.
- Enamine REAL make-on-demand compound collections.
- RDKit open-source cheminformatics and Open Free Energy.
Have a discovery bottleneck you can put a number on?
If you can measure it and verify it, we can talk about contracting against it. Tell us the outcome you need, and we will tell you whether we would put part of our fee behind it.
Scope an outcome with usOr read how the model works on the Outcome as a Service page.

