An AI Agent That Designs Drug Formulations and Runs Its Own Experiments

Finding the right formulation for a poorly soluble drug is one of the most tedious problems in pharmaceutical development. You need to blend oils, surfactants, and cosolvents in precise ratios so the drug dissolves in the gut, stays dissolved, and reaches the bloodstream. The combinatorial space is enormous, and traditional methods like design-of-experiments (DoE) explore it slowly, testing one batch at a time with statistical guidance but little accumulated knowledge. Andromeda 2, a new agentic system from Intrepid Labs, attacks this problem differently: it reasons over the lab's entire history of past formulation experiments, proposes candidates with scientific rationale, checks them against hard constraints, and runs them on an automated miniaturized laboratory. For paclitaxel, a notoriously difficult chemotherapy drug to deliver orally, the system found high-performing formulations at nearly three times the rate of its predecessor and over twenty times the rate of traditional DoE.

Why Paclitaxel Is a Hard Test Case

Paclitaxel is a high molecular weight molecule with extremely poor aqueous solubility. Getting it through the gut and into the bloodstream requires formulations that achieve therapeutically relevant drug loading while keeping the drug dissolved or supersaturated after the formulation is diluted in gastrointestinal fluid. The drug also faces biological barriers like P-glycoprotein efflux and first-pass metabolism, though those are outside the scope of this study.

The clinical relevance is real. DHP107, an oral lipid-based paclitaxel formulation, demonstrated noninferior efficacy compared to intravenous paclitaxel in Phase III trials for gastric cancer and is now in Phase III for breast cancer. So oral paclitaxel is not just an academic exercise. It is a formulation objective with direct patient impact, and it is genuinely hard.

How Andromeda 2 Works

Andromeda 2 is an agentic system that orchestrates multiple components: proposal generation, critique, review, planning, and constraint checking, all operating over a shared formulation context. It can invoke the same probabilistic models used by Andromeda 1, plus additional models, to construct batches of 16 formulations with scientific rationale.

The key difference from Andromeda 1 is evidence grounding. Andromeda 2 is initialized with Intrepid Labs' entire structured in-house formulation history: prior compositions, measured dissolution, dispersion, droplet size, PDI outcomes, excipient behavior, and platform-specific measurement history. Critically, this historical data contains no paclitaxel formulations. The system uses knowledge accumulated across other APIs and formulation programs to make decisions about a drug it has never seen before. After each batch, the newly generated paclitaxel measurements are added to its working context, and it adapts its strategy for the next batch.

A deterministic feasibility engine sits downstream of the agentic layer. It enforces hard constraints: valid excipient selections, sum-to-one compositional requirements, uniqueness relative to previously tested formulations, intra-batch distinctiveness, and robotic readiness. This separation is deliberate. The agentic layer directs scientific exploration, while the engineering layer ensures that submitted compositions are actually executable on the liquid-handling robots.

A formulation scientist reviews each proposed batch before execution as a safety check. The reviewer does not redesign, substitute, or retune the compositions. They verify that the batch is scientifically reasonable and experimentally executable. The human is in the loop for safety, not for scientific judgment.

The Numbers: 50% vs 17% vs 2%

All three strategies (Andromeda 2, Andromeda 1, and experimental DoE) received the same budget: 96 unique formulations, executed in six batches of 16 on the same miniaturized automated laboratory, using the same FaSSIF dissolution assay. The primary endpoint was AUC from 10 to 240 minutes in fasted-state simulated intestinal fluid, measuring the formulation's ability to maintain apparently solubilized paclitaxel.

Andromeda 2 achieved a median AUC of 70.1 mg·min/mL, compared with 12.0 for Andromeda 1 and 3.5 for DoE. The high-AUC hit rate (formulations reaching the pooled upper quartile) was 50% for Andromeda 2, 17% for Andromeda 1, and 2% for DoE. Andromeda 2 found 12 formulations meeting all four target product profile (TPP) objectives (minimum dissolution AUC, late-time concentration floor, late-time fraction-dissolved floor, and precipitation resistance ratio), versus 6 for Andromeda 1 and 0 for DoE.

The maximum AUC was comparable between Andromeda 2 and Andromeda 1 (157.7 vs 156.4 mg·min/mL). This is important: the agentic system's advantage was not finding a higher isolated peak. It was producing substantially more high-performing formulations across the campaign. The highest-AUC formulation met only two of the four TPP criteria, confirming that peak AUC and balanced formulation quality are not the same thing.

The Evidence Effect: 34% More AUC from Historical Data

A controlled ablation isolated the contribution of in-house experimental evidence. The reduced-evidence configuration withheld all historical formulation data while keeping everything else constant: formulation objective, TPP, executable design space, feasibility constraints, batch size, and wet-lab workflow. Both configurations still received the paclitaxel measurements generated during their own campaigns.

Full-evidence Andromeda 2 achieved a mean AUC of 62.0 mg·min/mL versus 46.2 for the reduced-evidence configuration, a 34% improvement. The high-AUC hit rate was 50% versus 31%. The batch-resolved trajectories reveal something more interesting than a simple head start. The reduced-evidence configuration matched or outperformed full-evidence Andromeda 2 in early batches but subsequently regressed. Full-evidence Andromeda 2 improved across successive batches and sustained its performance. Both configurations received the same new measurements during the campaign. The divergent trajectories suggest that historical evidence helps the system interpret and use newly generated results more effectively, not just providing better initial proposals.

Where the Formulations Concentrated

The Lipid Formulation Classification System (LFCS) categorizes formulations by their relative proportions of oil, water-insoluble surfactant, water-soluble surfactant, and hydrophilic cosolvent. Within the tested design space, high-performing formulations clustered in the Type IIIB/IV region: water-soluble surfactants with low or no oil content. The hit rate was 0% for Type II, 6% for Type IIIA, 35% for Type IIIB, and 32% for Type IV.

Andromeda 2 directed most of its budget toward this productive region and largely avoided the unproductive Type II systems. DoE allocated a substantial fraction of its budget to Type II formulations, which contributed to its low hit rate. But the advantage was not just about finding the right region. Within the Type IIIB/IV region, Andromeda 2 achieved a 49% hit rate versus 29% for Andromeda 1 and 3% for DoE. The system both concentrated search in the right neighborhood and selected better candidates within it.

The composition trajectory showed adaptation across batches. Andromeda 2 began with broad exploration, sampling 12 distinct oil-surfactant families in its first batch. Subsequent batches increasingly concentrated on productive formulations. It sampled fewer distinct families overall than DoE (16 vs 90) but generated far more high-performing formulations, while retaining several chemically distinct productive families for downstream development.

Comparison with Published Formulations

A selected formulation meeting all four TPP criteria achieved an apparent effective paclitaxel loading of 19 ± 5% w/w at the first FaSSIF measurement, approximately 3.3-fold higher than the 5.7% w/w loading reported for the S-SEDDS formulation of Gao et al. The authors emphasize that this comparison is contextual rather than head-to-head: endpoints, formulation types, and measurement methods differ across studies. The FaSSIF AUC measured here is not oral bioavailability or systemic exposure.

Vitamin E TPGS was prominent among the high-performing formulations, consistent with its known paclitaxel-solubilizing properties. It also exhibits independent effects on P-glycoprotein-mediated paclitaxel transport, which may be advantageous for future in vivo studies, though intestinal transport was not measured here.

What Andromeda 1 vs 2 Tells Us About Agentic Systems

Andromeda 1 is a probabilistic optimization model that has been deployed across dozens of live formulation development projects for diverse APIs and formulation modalities. It does not access historical data at campaign initialization; it learns the formulation-performance landscape from measurements generated during the current campaign. It is a strong baseline.

The fact that Andromeda 2 reached a comparable maximum AUC but a much higher median and hit rate suggests that the agentic approach's value is in consistency and breadth, not in finding a single outlier. In a pharmaceutical development context, having 12 full-TPP formulations to choose from is more valuable than having one exceptional formulation and a long tail of mediocre ones. The ability to select among multiple viable candidates gives formulators options for downstream optimization, stability testing, and regulatory strategy.

Limitations

Each strategy was evaluated in a single independently initialized campaign, limiting campaign-level statistical inference. Replicate autonomous campaigns are needed to establish reproducibility. The FaSSIF assay measures apparent solubilization in simulated intestinal fluid but does not capture digestion, intestinal permeability, metabolism, or in vivo performance. Digestion-coupled testing and pharmacokinetic studies would be required to determine whether these formulation advantages translate to oral exposure.

The experimental DoE comparator explored the shared composition space directly, without the expert pre-screening and prioritization that typically accompanies conventional formulation development. This is a fair comparison for the purpose of this study but should be noted when interpreting the DoE performance relative to real-world pharmaceutical development workflows where experienced formulators guide the search.

The Bigger Pattern: Evidence as Compound Interest

The most interesting implication of this work is the evidence-ablation result. Andromeda 2 had no paclitaxel data at the start. It used knowledge from other drugs and other formulation programs to make better decisions about paclitaxel. The reduced-evidence configuration still worked, still produced valid formulations, still improved across batches. But it could not sustain the same trajectory because it lacked the context to interpret new measurements as effectively.

This suggests a compounding effect. As Intrepid Labs' in-house experimental database grows across more APIs and formulation types, Andromeda 2's starting position for each new drug improves. The laboratory's accumulated knowledge becomes a reusable asset that makes every subsequent campaign more efficient. If this pattern generalizes, evidence-grounded autonomous laboratories could make formulation development progressively faster as institutional experimental knowledge accumulates. Each experiment does not just answer a question about the current drug. It adds context that improves future experiments on different drugs.

Read the paper on arXiv