The purpose of preclinical toxicology has never simply been to predict what happens in mice. It has always been to predict what happens in patients.

That distinction has driven decades of innovation in drug discovery, and it also exposes one of the greatest challenges facing modern toxicology today. Despite increasingly sophisticated preclinical testing, drug-induced liver injury (DILI) remains one of the leading causes of late-stage clinical attrition, regulatory restriction, and post-marketing withdrawal. Every year, candidates that clear conventional preclinical safety studies go on to reveal unexpected liver toxicity once they reach human trials or clinical use.

The problem is not a lack of scientific effort, and it does not diminish the immense contribution animal models have made to modern medicine. Animal studies have played a foundational role in developing therapies that have transformed patient care, and they continue to provide valuable insight into complex biological systems. But they remain, by definition, models of non-human biology. Differences in hepatic metabolism, transporter expression, immune response, disease progression, and genetic regulation mean that findings in animals do not always translate directly to patients.

Recognising this is not an argument against animal research. It is an argument for continuing to improve the evidence base that drug development decisions rest on. Over the past decade, real progress has been made toward that goal: advances in computational toxicology, human induced pluripotent stem cell (iPSC)-derived models, high-content phenotypic profiling, organ-on-chip systems, and artificial intelligence have expanded the toolkit considerably. Collectively, they are starting to change how the field thinks about preclinical evidence itself.

That broader shift is what makes the FDA's Roadmap to Reducing Animal Testing in Preclinical Safety Studies significant.

When the roadmap was published in April 2025, most of the discussion centred on one question: is the FDA moving away from animal testing? One year on, that framing feels too narrow. FDA's first-year progress report, Reducing Animal Testing in Nonclinical Studies: Year One Progress and the Path Forward (FDA, April 2026), suggests the roadmap is about something more consequential than animal use alone: building the scientific and regulatory foundations for a future in which human-relevant evidence is the cornerstone of safety assessment.

As FDA Commissioner Marty Makary put it when the report was released, the agency set out “an ambitious roadmap to eliminate unnecessary animal testing” a year earlier, and the Year One report was framed as evidence that first-phase progress had been met (Fierce Biotech, April 2026). Not every observer agrees on how far that progress goes; the National Association for Biomedical Research has cautioned that animal models and NAMs are both surrogates for human biology, and that context of use, not a blanket preference, should determine which is appropriate for a given question (same source). That caution is worth taking seriously rather than waving away.

Liver safety appears to be one of the first places this transition is beginning to take concrete shape.

Why Liver Safety Is Leading the Transition

Few areas of drug development illustrate the need for more predictive models as clearly as DILI. Unlike many toxicological endpoints, DILI is not driven by a single mechanism. A 2025 review in Drug Discovery Today compared the performance of in vitro NAMs, animal studies, and microphysiological systems (MPS) across a shared set of drugs and found considerable variability among in vitro NAMs, with too little MPS data yet available for a fully confident comparison, motivating the authors' proposal of DILIference, a curated reference drug list for benchmarking future NAMs against a common standard.

This complexity has challenged toxicologists for decades. Animal studies have identified genuine safety liabilities, but they have also produced false positives that halted promising therapeutics and false negatives that only surfaced once compounds reached clinical development. Species differ in drug metabolism, cytochrome P450 activity, bile acid transport, immune regulation, and regenerative capacity, and even closely related mammalian species can respond differently to the same compound.

Traditional in vitro models have their own limits. Immortalised hepatic cell lines often lack the metabolic competence to reproduce clinically relevant toxicity mechanisms, while primary human hepatocytes, though considered a gold standard for many applications, are constrained by donor variability, limited availability, and declining functionality in culture. No single model captures the full complexity of human liver biology, and treating any one system as the eventual 'perfect' replacement has arguably slowed progress by discouraging complementary approaches.

It is not a coincidence that liver safety features so prominently in FDA's Year One report. Liver microphysiological systems, liver-on-chip technologies, and AI-assisted biomarker tools such as AIM-NASH (FDA's first qualified AI drug development tool, built for scoring MASH/NASH liver biopsies rather than DILI specifically) all receive substantial attention. That is because liver safety has become one of the clearest illustrations of why human biology needs to sit closer to the centre of preclinical decision-making, not because hepatotoxicity is the only area benefiting from NAMs.

pixlbio has its own stake in this specific evidence base: our Pixl Bio Joins Landmark FDA–3Rs Collaborative Cross-Platform DILI Study — Now ISTAND-Accepted post covers our own iPSC-derived liver MPS data, our role among the eight commercial platforms evaluated, and what the study's January 2026 ISTAND acceptance means in detail. This piece takes a step back from that specific study to look at the roadmap as a whole.

The Roadmap Is Not the Destination

One common misreading of FDA's roadmap is that it marks the beginning of the end for animal testing. What the document actually says is both more cautious and more ambitious than that.

FDA is not suggesting animal studies disappear overnight. It recognises that meaningful replacement only happens once researchers can generate evidence that is demonstrably more predictive of human outcomes than what it replaces. That distinction matters, because it reframes the conversation from substitution to scientific progress.

FDA's first year of implementation points away from a one-for-one search for replacements and toward a small set of technology-agnostic principles set out in its March 2026 draft guidance, General Considerations for the Use of New Approach Methodologies in Drug Development: context of use, human biological relevance, technical characterization, and fit-for-purpose. These principles don't ask which model is being used. They ask whether the evidence generated is reliable enough to answer the specific scientific question at hand.

That is a real shift in regulatory thinking. The benchmark is no longer whether a new method reproduces an animal model's results. Increasingly, it's whether the method improves confidence in predicting human biology.

The guidance is also more specific than most commentary on it suggests. For hepatotoxicity models in particular, it sets out what human biological relevance actually looks like in practice: the model should contain the relevant cell types, such as hepatocytes, stellate cells, and Kupffer cells, and the biological features needed to recapitulate hepatocellular physiology, including albumin production and metabolic competence. On technical characterization, it recommends assays that measure albumin and urea secretion over a repeated-exposure timeframe, alongside functional endpoints such as CYP450 activity and ALT/AST release. That is, point for point, the same endpoint set liver MPS developers, including pixlbio, are already generating in cross-platform studies like the one we describe in our companion post on the FDA–3Rs Collaborative DILI project.

The guidance also draws a distinction worth sitting with: validation and qualification are not the same thing. Validation establishes the accuracy, reliability, and relevance of a method for a specific context of use; qualification is a separate, formal determination under FDA's drug development tool program (which includes the ISTAND pathway). A NAM does not need to be qualified to be used in a submission, and it does not even need to be fully validated to be considered, provided it's fit-for-purpose for the specific weight-of-evidence question being asked. That is a meaningfully lower bar to entry than many sponsors assume, and it reframes fit-for-purpose not as a consolation prize next to full validation, but as one of three legitimate ways a NAM can contribute: replacing a traditional method outright, filling a gap traditional models cannot address, or confirming and complementing what a traditional method already shows.

On technical characterization specifically, the guidance calls out biological variability, including genetics and donor variability, as something sponsors need to describe and account for as it relates to the context of use, not treat as background noise. That principle sits closely alongside the reasoning in our own post on why every pixlbio study runs across multiple donors.

For decades, progress in toxicology has largely meant refining the existing paradigm: better animal models, better assays, better endpoints, better translation, on the assumption that incremental improvement would eventually overcome the system's structural limits. FDA's roadmap points to a different premise. The opportunity is not a better wheel. It's a different vehicle: predictive AI flagging liabilities before a compound is even synthesised, human iPSC-derived models providing biologically relevant systems for functional testing, high-content phenotypic profiling catching subtle cellular responses conventional endpoints miss, microphysiological systems recreating tissue architecture that traditional culture cannot, and clinical data anchoring all of it back to real patient outcomes.

From Roadmap to Reality

One year on, the conversation has moved from aspiration toward implementation. According to FDA's own summary and independent coverage of the Year One report (FASEB Washington Update, May 2026; Spencer Fane, June 2026), the agency made the ISTAND qualification pathway permanent, published a NAMs Acceptability Database for transparency on what the agency currently accepts, updated monoclonal antibody guidance to remove the six-month non-human primate study requirement for certain products, and qualified AIM-NASH as its first AI drug development tool.

An independent academic analysis of the roadmap, published in Stem Cell Research & Therapy (Wu et al., 2026), found that a 15-year retrospective analysis of NAM submissions to FDA showed 93% falling into in silico or in vitro categories, information the authors argue should help focus validation resources on the platforms already carrying the most regulatory weight.

The more interesting signal here isn't the list of initiatives. It's what they collectively suggest: that regulatory willingness was never really the primary barrier to adoption. The harder, remaining problem is generating evidence robust, reproducible, and context-specific enough to support confident decisions across the drug development pipeline, which is especially difficult for a liability as biologically heterogeneous as DILI.

Building Evidence, Not Replacing Models

The most important message in FDA's roadmap may be the one easiest to miss: the future of toxicology will not be one technology replacing another. FDA's guidance returns again and again to context of use, weight of evidence, and fit-for-purpose validation, because each technology contributes a different kind of information, not a competing version of the same information.

Computational models can prioritise compounds and flag liabilities before synthesis. Human iPSC-derived cells provide biologically relevant systems for functional testing. High-content phenotypic profiling picks up subtle cellular changes conventional endpoint assays miss. Microphysiological systems recreate tissue architecture and dynamic biology no dish-based culture can. Clinical datasets anchor all of it to real outcomes. None of them is meant to answer every question on its own.

For liver safety, that layered approach is especially compelling, precisely because hepatotoxicity so rarely results from a single mechanism. Combining complementary technologies lets researchers interrogate multiple biological processes at once, improving mechanistic understanding and predictive confidence together rather than trading one for the other.

Why Toxicologists Have Been Skeptical, and What Actually Changes That

It's worth being direct about something the field doesn't always say out loud: individually, a lot of NAMs have earned a real trust problem among working toxicologists, and that skepticism is rational rather than reflexive. A single in vitro readout, one Cell Painting profile, or an isolated organ-chip result rarely tells a toxicologist anything they can act on, because none of those things predicts a clinical outcome by itself. Asking someone to make a go/no-go call on the strength of one disconnected assay, however sophisticated, is a hard sell, and it should be.

What actually changes that isn't a better version of any single method. It's continuity. The genuine unlock is being able to follow one signal end to end: a structural liability flagged early, confirmed or overturned in a human-relevant cell model, explained mechanistically, and eventually connected back to a real clinical outcome, so a toxicologist can see for themselves that the early signal meant something. That thread from early discovery through to clinical relevance is closer to what the field has actually been missing than any incremental gain in assay sensitivity or imaging throughput.

This is effectively what FDA's weight-of-evidence framing is describing from the regulatory side, too. The guidance doesn't ask whether any one method is accurate in isolation. It asks whether the methods used, taken together, build a defensible chain of evidence connecting an early signal to a human-relevant conclusion. That's a different, and in some ways higher, bar than validating a single assay, but it's also a more honest description of what it actually takes to earn a toxicologist's confidence.

The Next Challenge Is Integration

The scientific foundations for this shift already exist. Human-relevant models keep improving in physiological relevance, computational methods keep getting more sophisticated, and AI is letting researchers analyse biological complexity at a scale that wasn't practical even a few years ago.

The next challenge, in other words, is not invention. It's integration. Platforms that connect predictive computational models with human cellular biology, high-content phenotypic analysis, and mechanistic interpretation are likely to play a growing role throughout drug discovery, identifying safety liabilities earlier, improving confidence in go/no-go decisions, and reducing dependence on long, sequential testing strategies that front-load animal studies before human biology is ever explored directly.

Science Alone Will Not Drive Adoption

Despite the pace of technical progress, one of the clearest messages in FDA's report is that implementation isn't only a scientific challenge. It's a cultural one, too.

Developers need to feel confident proposing human-relevant approaches early in a program. Regulators need to keep building expertise in evaluating novel forms of evidence. Pharmaceutical companies need confidence that data generated through integrated NAM strategies will hold up across multiple regulatory jurisdictions, not just one. International harmonisation matters just as much: a safety strategy accepted by one regulator has limited practical value if additional animal studies are still required elsewhere. Organisations including ICCVAM, the International Council for Harmonisation (ICH), and regulatory agencies across Europe, North America, and Asia all have a role in making sure validation frameworks evolve consistently across jurisdictions.

Looking Ahead

FDA's roadmap is often described as a strategy for reducing animal testing. A year on, that description feels incomplete. At its core, the roadmap is about improving the quality of evidence used to predict patient safety. Replacing animal studies was never the endpoint in itself; the objective is evidence that's more informative, more reproducible, and more predictive of human outcomes. For liver safety, that future is already starting to take shape.

The opportunity in front of the field isn't to build a faster version of the existing system. It's to build a better one; one defined not by choosing between animal models, artificial intelligence, or advanced human cell systems, but by integrating them into evidence that better reflects human biology and supports more confident decisions throughout drug development.

One year after FDA published its roadmap, one conclusion is becoming difficult to ignore.

The future of liver safety is human.

pixlbio develops human-relevant DILI modelsusing iPSC-derived hepatocytes, Cell Painting phenomics, and AI-driven ToxSuite prediction. Our pixHep hepatocyte models and pixCellPaint platform arebuilt specifically for early-stage hepatotoxicity detection, aligned with thehuman biological relevance principle at the centre of FDA's evolving NAMvalidation framework.

To discuss howour approach could strengthen your safety programme, booka call with our team.

Tags: NAMs · FDARoadmap · DILI · Liver Safety · ISTAND · Animal Testing Alternatives · iPSCHepatocytes · Cell Painting · Regulatory Affairs · Preclinical Toxicology

Conclusion

FDA's Year One report will likely be read, correctly, as a story about regulatory momentum: more ISTAND submissions, a permanent qualification pathway, a validation framework with real teeth. But the more useful way to read it, for anyone working in liver safety specifically, is as a description of what toxicologists have been asking for all along, not a faster assay or a cheaper alternative to an animal study, but a way to trust an early signal because it has been followed, consistently, from structure through to a clinically relevant conclusion. That is a harder thing to build than any single NAM, and it is also the only thing that actually changes minds in a field that has good reason to be skeptical of point solutions; the technologies to do it already exist, and what is left is the discipline to connect them, and the patience to let the evidence, not the novelty, do the convincing.
References

FDA. New Approach Methodologies (NAMs). FDAScience and Research Special Topics

FDA (April 2026). Reducing Animal Testing inNonclinical Studies: Year One Progress and the Path Forward. FDAPress Announcements

FDA, CDER (March 2026). General Considerationsfor the Use of New Approach Methodologies in Drug Development. Guidance forIndustry (Draft). FederalRegister

Zubulake Z (March 2026). What FDA's NAM GuidanceMeans for Pharmaceutical Development. PharmaceuticalTechnology

FDA. Innovative Science and TechnologyApproaches for New Drugs (ISTAND) Program. FDADrug Development Tool Qualification Programs

FDA qualifies first AI drug development tool forreading MASH liver biopsy images (AIM-NASH). FierceBiotech

FDA celebrates progress to end animal testingbut experts warn there's still a long way to go. FierceBiotech,April 2026

FDA Releases Retrospective Report on Roadmap toReduce Animal Testing. FASEBWashington Update, May 2026

FDA Reports Meeting Year One Goals for ReducingAnimal Drug Testing. SpencerFane, June 2026

Wu X, Wu MA, Zou J, Kleinstreuer N, Wu JC(2026). FDA roadmap to reducing animal testing: a regulatory and scientificparadigm shift in nonclinical safety assessment. StemCell Research & Therapy

New approach methodologies (NAMs) fordrug-induced liver injury (DILI): Where are we now? (2025). DrugDiscovery Today

Watkins PB (2011). Drug safety sciences and thebottleneck in drug development. Clinical Pharmacology andTherapeutics

FDA. Liver Toxicity Knowledge Base (LTKB). FDABioinformatics Tools

pixlbio. Pixl Bio Joins Landmark FDA–3RsCollaborative Cross-Platform DILI Study — Now ISTAND-Accepted. pixlbio Blog