The Field: What Toxicogenomics Is Actually For

Who pays for this, what they need, and what would count as a contribution.

Everything else in this book is about methods. This chapter is about why anyone cares. Read it before deciding what to work on — because “improve PCC by 5%” is not a contribution to this field, and knowing why is the difference between a paper that lands and one that doesn’t.

All 2025–2026 claims sourced; links inline and in 17_SOURCES.md.


2.1 The problem, in one number

~350,000 chemicals are on the global market. Most are “data-poor” — no clear information on their mechanisms or health effects.

Not “under-studied.” Data-poor. Nobody knows what they do.

And the traditional way to find out:

traditional toxicity testing + human health assessment8+ years per chemical
animals per chemicalhundreds to thousands
costmillions

Do the arithmetic. At 8 years per chemical, 350,000 chemicals is not a backlog — it’s a mathematical impossibility. The field is not slowly catching up. It is permanently, structurally behind, and falling further behind every year as new chemicals enter commerce.

This is the problem toxicogenomics exists to solve. Not “understand biology better” — clear an impossible queue.


2.2 The deadline that changes everything

This is no longer aspirational. There are dates.

EPA

The EPA’s goal: eliminate ALL mammalian animal testing by 2035.

On June 2, 2026, EPA announced two major actions to replace animal-based testing with alternatives for chemical assessments under TSCA (Toxic Substances Control Act) and FIFRA (pesticides). It also introduced a streamlined process to nominate NAMs for use in pesticide and chemical assessments. (source)

FDA

April 2025: a roadmap for a stepwise transition from animal testing to validated New Approach Methodologies — beginning with monoclonal antibodies, expanding to other biologics, then new chemical entities.

March 2026: draft guidance, “General Considerations for the Use of New Approach Methodologies in Drug Development” — a validation framework and regulatory expectations for replacing animal toxicology studies. (FDA)

What “NAMs” means

New Approach Methodologies — FDA’s definition spans complex in vitro, 2D in vitro, in chemico, and in silico studies.

Read that last one again. In silico is a NAM.

A model that predicts a kidney profile from a liver profile is a candidate NAM. That is not a metaphor or a stretch — it is the regulatory category this work falls into, and there is now a formal process to nominate things into it.

The gap this creates

Animal testing is being phased out on a deadline. The replacements must exist. They must be validated. They must be accepted by regulators.

That’s the demand side. This is a field with a legal mandate and a shortage of supply. That is an unusual and favourable place to be doing methods research.


2.3 What regulators actually want (and it isn’t a gene profile)

This is the single most important thing for an ML person to understand, and it’s where most ML papers in this space miss.

A regulator does not want your predicted 8,565-dimensional vector. They want one number:

The Point of Departure (POD)

The dose below which nothing bad happens.

Everything downstream — safety factors, reference doses, exposure limits, whether a chemical can be sold — derives from that number.

Traditionally the POD is an apical POD: dose animals at several levels for months or years, look for actual damage (liver necrosis, tumours, death), find the dose where damage starts. This is what takes 8 years.

NLP framing: you spent a year building a beautiful sequence-to-sequence model, and the customer wants a scalar. Everything you optimize should be judged by whether it improves that scalar’s accuracy or the confidence around it. PCC on the vector is at best a proxy, and 16_MATH_NOTES.md §3 shows it’s a bad proxy.


2.4 The transcriptomic POD — why this field exists at all

Here is the finding the entire enterprise rests on:

Apply benchmark-dose analysis to gene expression from a SHORT-term exposure, and you get a POD that is concordant with the apical POD from a LONG-term study — including chronic toxicity and cancer.

Read that carefully. Days of exposure predict years of outcome.

tPOD vs apical POD agreementoften within 3-fold (Frontiers 2024)
exposure neededdays, not years
what it protects against“thought to be highly protective of all potential adverse toxicological effects”

Why it works, mechanistically: damage doesn’t appear from nowhere. Before a tumour, before necrosis, the cell’s transcriptional program shifts — stress response, repair, proliferation. The transcriptome is the early warning. Gene expression moves first; pathology follows.

NLP framing: it’s a leading indicator. You don’t wait for the user to churn; you read the signal that precedes churn.

This is the whole bet of toxicogenomics, and it has held up well enough that regulators are building products on it.


2.5 ETAP — the actual regulatory product

Not a proposal. A thing EPA is building.

EPA Transcriptomic Assessment Product (ETAP): a human health assessment for chemicals lacking traditional toxicity data, using a standardized short-term in vivo study + standardized transcriptomic analysis to derive a reference value. (EPA)

traditionalETAP
time8+ yearsmonths
chemicals it targetswell-studieddata-poor

Reviewed by EPA’s Board of Scientific Counselors in July 2023. Real assessments exist — e.g. for Perfluoro-3-Methoxypropanoic Acid.

This is the destination. A standardized short in-vivo study → transcriptomics → a number a regulator can act on, in months rather than a decade.

Everything in this book is upstream of that pipe.


2.6 Where DrugMatrix fits — and who is actually behind these papers

This is worth knowing because it explains the program’s real objective.

Scott S. Auerbach, Ph.D., DABT leads the Toxicoinformatics Group in the Predictive Toxicology Branch, Division of Translational Toxicology (DTT), NIEHS. He:

  • oversees the NTP DrugMatrix resource — the dataset all four papers are built on
  • built BMDExpress 2 and 3, the standard tools for genomic dose-response analysis
  • co-authored S1500+, the landmark gene set (01_BACKGROUND.md §2.3)
  • is corresponding author on both TransPlatformer and TransTissueFormer

(NIEHS profile)

So this is not an academic ML program that happens to use tox data.

It is a regulatory toxicology program that is using ML. The people who own DrugMatrix, who built the dose-response tooling, and who are driving transcriptomic PODs toward regulatory acceptance are the authors.

That should change what you consider a contribution. A method that improves MAE but produces profiles a toxicologist can’t act on is worth less than an honest negative result that de-risks a regulatory submission.


2.7 The keystone result — and the uncomfortable question it raises ⭐⭐

Johnson, Auerbach & Costa (2020), Toxicological Sciences 176(1):86 — “A Rat Liver Transcriptomic Point of Departure Predicts a Prospective Liver or Non-liver Apical Point of Departure”

The setup: 79 molecules from Open TG-GATEs. Derive a liver transcriptomic POD from short-term exposure. Compare it to the systemic apical POD.

The finding:

A liver tPOD predicted the systemic apical POD within 10× — even when that apical POD came from a non-liver endpoint.

Following subacute (29-day) dosing, liver BEPOD and systemic apical POD agreed within 10×.

Note the author list. Scott Auerbach — corresponding author on TransTissueFormer.

The uncomfortable question ⭐

If the liver transcriptome already predicts non-liver toxicity within 10×, what does cross-tissue translation add?

This is the question a reviewer will ask, and it deserves a real answer rather than a dodge. As far as I can tell there are two, and they’re different in kind:

Answer 1 — different products. The tPOD is a dose: “below 5 mg/kg, nothing bad happens anywhere.” Protective, conservative, and deliberately silent on mechanism. Cross-tissue translation produces a profile: “at 50 mg/kg the kidney shows p53 activation and oxidative stress.” You cannot get the second from the first. POD answers how much; the profile answers what and where.

Answer 2 — the 10× is doing a lot of work. Ten-fold is a wide band. For a protective screening threshold that’s fine — err low, nobody gets hurt. For deciding which organ to look at or what the mechanism is, it’s useless.

And the honest caveat: if what you need is a protective threshold, the 2020 result already delivers it, cheaply, from liver alone — and cross-tissue translation is solving a problem that is already solved. The contribution has to be on the mechanism side, or it isn’t a contribution.

This is the framing question for the whole program, and I’d want it answered before committing a year. It’s also exactly the kind of thing the corresponding author can answer in five minutes, since he wrote the 2020 paper.


2.8 The fifth paper you should know about

There is a ToxCompl successor that isn’t in your four PDFs:

Nguyen & Cong (2025), BioKDD’25“Completion of the DrugMatrix Toxicogenomics Database using 3-Dimensional Tensors”ToxiTenCompl.

The idea, and it’s a good one: DrugMatrix isn’t naturally a matrix. “each probe measurement is made with a compound, a dose, a duration, for a gene within a tissue” — that’s a 3D or 4D tensor. The 2D view is a flattening that throws away structure.

ToxCompl (2D):     rows = (platform, tissue, gene)   x   cols = treatments
ToxiTenCompl (3D): tissue  x  treatment  x  gene
methodgradient descent, not alternating minimization (their ToxiTenCompl)
beatsCP decomposition and 2D matrix factorization, on MSE and MAE
lossweighted MSE — higher weights on over/under-expressed genes
bonusnon-negative version yields interpretable tissue factors
stated aim“drugs that may cross species barriers, for example, from rats to humans”

⭐ Why this matters for everything in Chapters 2–5

The tensor formulation partly fixes a problem I spent three chapters complaining about.

In ToxCompl’s 2D layout, Cyp1a1-in-liver and Cyp1a1-in-kidney are different rows with independent parameters. The model has no idea they’re the same gene.

In the 3D tensor, gene is its own mode. The gene factor is shared across tissues by construction. That’s structurally the right move, and it’s the same instinct as “make the gene side inductive” (06_GENTOX.md §6.7) — arrived at from a different direction.

It also changes my obs/param analysis (04_TOXCOMPL.md §4.6). CP rank- on a tensor has parameters, not where . The tensor is far more parameter-efficient — which is likely why it wins, and it’s a better explanation than the paper gives.

Read this paper. It is the current state of the program’s completion work, and it partly pre-empts the criticism in Chapter 4.


2.9 What the field actually struggles with

The honest list, from the literature. These are the open problems — pick from here, not from the ML leaderboard.

problemstatus
regulatory acceptance of omics“implementation of omics-based approaches in AOPs and their acceptance by the risk assessment community is still a challenge”
AOPs aren’t actually used“regulatory agencies show limited formal adoption of AOPs despite their potential”; a “disconnect between AOP research and regulatory implementation”
cross-species applicabilityexplicitly named as a key gap
quantitative data insufficiencynamed as a key gap
mixture toxicitynamed as a key gap
workflow standardization for tPODsactive; “rigorous and reproducible wet and dry laboratory methodologies are required”
read-across is the current crutchEPA, Health Canada, and EU rely on it for data-poor chemicals — it’s chemical-similarity guessing

Two terms you’ll hit constantly

AOP (Adverse Outcome Pathway) — a structured causal chain: molecular initiating event → key events → adverse outcome. E.g. “binds receptor → activates transcription → cell proliferation → tumour.” It’s how toxicologists formalize mechanism. Regulators like the idea and barely use it. That gap is a research opportunity.

Read-across“this new chemical looks like that known chemical, so assume similar toxicity.” It’s the default for data-poor chemicals because there’s nothing better. It’s structure-based guessing with expert judgment on top.

Notice that GenTox’s compound GNN is a learned read-across. Pretrain on 1M molecules, embed, predict response. That is exactly the read-across problem, done with representation learning instead of expert judgment — and their own ablation shows learned beats hand-crafted (06_GENTOX.md §6.3).

That framing would land much better with a tox audience than “inductive matrix factorization” does. Same method, different sentence, completely different reception.


2.10 What counts as a contribution — and to whom

You are choosing between three audiences. They want different things, publish in different places, and reward different work.

Audience A — the ML community (NeurIPS, ICML, ICLR, KDD)

wantsnovelty, benchmarks, SOTA
where the program already goesBioKDD (the tensor paper)
your leveragethe NMT connection (12_TRANSLATION_TRANSFER.md) is genuinely novel to them
the riskthey don’t have the data and don’t care about rats. A tox-specific result is a poster, not an oral

Audience B — the toxicology community (Toxicological Sciences, Arch Toxicol, SOT)

wantsdoes it predict the apical endpoint? is it protective? can a regulator use it?
does not care aboutyour PCC, your architecture, your parameter count
does care aboutthe gemfibrozil→PPARα validation, the F1 0.636→0.718 downstream result
your leveragealmost nobody there can build these models
the riskyou must speak POD, BMD, AOP, concordance — not attention and embeddings

Audience C — the regulatory/NAM community (EPA ORD, NIEHS DTT, ICCVAM)

wantsvalidated, documented, reproducible, defensible
does not care aboutnovelty at all
the currencya method that survives a Board of Scientific Counselors review
your leveragethere is a formal nomination process and a 2035 deadline
the riskglacial; validation is measured in years

The strategic read

The program’s authors sit in B and C. DrugMatrix’s owner is the corresponding author. The destination is ETAP-shaped.

You arrive with A’s toolkit. That’s the arbitrage — and it cuts both ways. Your NMT knowledge is genuinely rare here. But if you optimize for A’s metrics you’ll produce something B and C cannot use, and the people who’d have to champion it are the ones you’d be failing.


2.11 The contributions that would actually matter

Re-derived from the field’s priorities rather than the ML leaderboard. Ordered by how much a toxicologist would care.

1. Enrichment-consistency as the metric ⭐⭐ — the best fit between your skills and their needs

15_FRONTIER.md F6 argued for this on ML grounds. The field argument is stronger.

Toxicologists do not evaluate profiles by correlation. They run the profile through enrichment analysis, get a mechanism, and act on the mechanism. That is the actual downstream task. TransTissueFormer did it by hand for 3 of 44 compounds and never automated it.

“Our model recovers the correct mechanism of action for 31 of 44 compounds” is a sentence a regulator can act on. “PCC 0.793” is a sentence a regulator will ignore.

And it’s the field’s own live problem: “The Metric Picks the Winner” (June 2026) shows rankings invert with metric choice, and GenTox proved the pathology in 2024. You’d be closing a loop the program itself opened.

2. Predict the tPOD, not the profile ⭐⭐ — the reframing

Everything in this book optimizes the wrong target.

The regulatory product is a dose. So evaluate cross-tissue translation on the question that matters:

Does a predicted kidney profile yield the same kidney tPOD as the measured one?

That’s a single number, comparable to apical data, directly ETAP-shaped, and BMDExpress already computes it — built by the corresponding author.

If predicted profiles give tPODs within 2–3× of measured ones, that is a NAM candidate, and there is a formal nomination process. If they don’t, the model isn’t fit for the purpose it exists for, and that’s worth knowing in a week rather than a year.

Nobody has done this. It’s the natural bridge between the ML and the regulation, it uses existing tooling, and it’s the experiment I’d want run before any of 15_FRONTIER.md.

3. Frame the compound GNN as learned read-across 🟢 — free

Same method, a sentence a tox audience recognizes. Read-across is the current regulatory crutch for data-poor chemicals. GenTox already beats hand-crafted descriptors with it. That’s a Tox Sciences paper written in the framing, not the method.

4. Cross-species — the field’s own named gap 🟡

“Limited cross-species applicability” is listed as a key gap. The tensor paper names it as the goal. UCE’s ESM2 tokenization is the principled route (11_SC_FOUNDATION_MODELS.md §6). And PLOS One 2020 already did rat→human hepatocytes with a bottleneck DNN — “circumventing the current reliance on orthologs” — and isn’t cited in any of the four papers.

Rat data is only a proxy for humans. Everything in this program is one translation short of the thing anyone actually wants.

5. Uncertainty 🟡 — what regulators need and models don’t give

A regulator cannot act on a point estimate with no error bar. Every model here outputs bare numbers. A calibrated interval on a predicted tPOD is worth more than a better mean, because it’s the difference between “usable in an assessment” and “interesting.”

This also connects to the MaxAE/sign-flip failure (04_TOXCOMPL.md §4.7): a model that’s confidently wrong about direction is worse than one that says it doesn’t know.

6. The baselines 🟢 — still first

Unchanged from 14_RESEARCH_AGENDA.md. Everything above is conditional on the mean predictor not matching 0.793. The field’s own flagship competition ran 1,200 teams and found models not consistently beating naive baselines. This isn’t pedantry — it’s the field’s central methodological problem.


2.12 What to read, in order

The field context (none of this is ML — read it anyway)

  1. ETAP — About — the destination, 3 pages
  2. Johnson, Auerbach & Costa (2020), Tox Sci 176(1):86liver tPOD predicts non-liver apical. The keystone. Read §2.7’s question while you do.
  3. FDA NAMs page — the regulatory frame, and note in silico is a NAM
  4. Bioinformatic workflows for deriving tPODs (Tox Sci 2024) — current status, data gaps, research priorities. A ready-made list of open problems.
  5. An AOP primerPMC5805086, concise, for toxicologists

The program’s own work you don’t have

  1. ToxiTenCompl (BioKDD’25) — the tensor successor to ToxCompl. §2.8.
  2. ToxCompl on bioRxiv — the published version
  3. BMDExpress 2/3 — the tool that turns your profile into the number regulators want. If you do contribution #2, this is the tool.

2.13 The summary

THE PROBLEM ~350,000 chemicals, most data-poor. Traditional assessment: 8+ years each. That queue can never clear.

THE DEADLINE EPA: eliminate all mammalian animal testing by 2035, with actions announced June 2026 and a formal NAM nomination process. FDA: NAM guidance March 2026. In silico methods are NAMs. A cross-tissue model is a candidate NAM — a regulatory category, not a metaphor.

WHAT REGULATORS WANT Not your vector. One number — the Point of Departure, the dose below which nothing bad happens.

WHY THE FIELD EXISTS Benchmark-dose analysis on short-term gene expression gives a POD concordant (often within 3×) with chronic apical PODs. Days predict years. The transcriptome moves before the pathology.

THE DESTINATION ETAP — short in-vivo study → transcriptomics → a regulatory reference value in months instead of 8 years, for data-poor chemicals. Real, reviewed, deployed.

WHO’S BEHIND THIS Scott Auerbach — leads Toxicoinformatics at NIEHS DTT, oversees DrugMatrix, built BMDExpress, corresponding author on two of your four papers. This is a regulatory toxicology program using ML, not an ML program using tox data.

THE UNCOMFORTABLE QUESTION His own 2020 paper shows a liver tPOD already predicts non-liver apical PODs within 10×. So what does cross-tissue translation add? Best answer: the POD is a dose, the profile is a mechanism — different products. But if a protective threshold is all you need, that’s already solved from liver alone. The contribution has to be on the mechanism side. Ask before committing a year.

THE FIFTH PAPER ToxiTenCompl (BioKDD’25) — 3D tensor completion, beats ToxCompl and CP, non-negative version gives interpretable tissue factors, explicitly aims at rat→human. It partly fixes the shared-gene problem Chapters 2–5 complain about, because gene becomes its own tensor mode.

WHAT WOULD ACTUALLY COUNT

  1. Enrichment-consistency as the metric — the real downstream task, and the mean predictor can’t game it
  2. Predict the tPOD, not the profile — the regulatory product is a dose; BMDExpress already computes it; nobody has tried
  3. Reframe the compound GNN as learned read-across — free, and it’s the field’s current crutch
  4. Cross-species — the field’s own named gap; rat is only a proxy for human
  5. Uncertainty — regulators can’t act on a bare point estimate
  6. The baselines — still first, still gating everything

Regulatory and field claims verified by search July 2026; links inline. The §2.7 question and the §2.8 parameter-efficiency reading are mine. Audience analysis in §2.10 is judgment, not fact — weigh it against what you learn from the people actually in the room.