The Field: What Toxicogenomics Is Actually For
Who pays for this, what they need, and what would count as a contribution.
Everything else in this book is about methods. This chapter is about why anyone cares. Read it before deciding what to work on — because “improve PCC by 5%” is not a contribution to this field, and knowing why is the difference between a paper that lands and one that doesn’t.
All 2025–2026 claims sourced; links inline and in 17_SOURCES.md.
2.1 The problem, in one number
~350,000 chemicals are on the global market. Most are “data-poor” — no clear information on their mechanisms or health effects.
Not “under-studied.” Data-poor. Nobody knows what they do.
And the traditional way to find out:
| traditional toxicity testing + human health assessment | 8+ years per chemical |
| animals per chemical | hundreds to thousands |
| cost | millions |
Do the arithmetic. At 8 years per chemical, 350,000 chemicals is not a backlog — it’s a mathematical impossibility. The field is not slowly catching up. It is permanently, structurally behind, and falling further behind every year as new chemicals enter commerce.
This is the problem toxicogenomics exists to solve. Not “understand biology better” — clear an impossible queue.
2.2 The deadline that changes everything
This is no longer aspirational. There are dates.
EPA
The EPA’s goal: eliminate ALL mammalian animal testing by 2035.
On June 2, 2026, EPA announced two major actions to replace animal-based testing with alternatives for chemical assessments under TSCA (Toxic Substances Control Act) and FIFRA (pesticides). It also introduced a streamlined process to nominate NAMs for use in pesticide and chemical assessments. (source)
FDA
April 2025: a roadmap for a stepwise transition from animal testing to validated New Approach Methodologies — beginning with monoclonal antibodies, expanding to other biologics, then new chemical entities.
March 2026: draft guidance, “General Considerations for the Use of New Approach Methodologies in Drug Development” — a validation framework and regulatory expectations for replacing animal toxicology studies. (FDA)
What “NAMs” means
New Approach Methodologies — FDA’s definition spans complex in vitro, 2D in vitro, in chemico, and in silico studies.
Read that last one again. In silico is a NAM.
A model that predicts a kidney profile from a liver profile is a candidate NAM. That is not a metaphor or a stretch — it is the regulatory category this work falls into, and there is now a formal process to nominate things into it.
The gap this creates
Animal testing is being phased out on a deadline. The replacements must exist. They must be validated. They must be accepted by regulators.
That’s the demand side. This is a field with a legal mandate and a shortage of supply. That is an unusual and favourable place to be doing methods research.
2.3 What regulators actually want (and it isn’t a gene profile)
This is the single most important thing for an ML person to understand, and it’s where most ML papers in this space miss.
A regulator does not want your predicted 8,565-dimensional vector. They want one number:
The Point of Departure (POD)
The dose below which nothing bad happens.
Everything downstream — safety factors, reference doses, exposure limits, whether a chemical can be sold — derives from that number.
Traditionally the POD is an apical POD: dose animals at several levels for months or years, look for actual damage (liver necrosis, tumours, death), find the dose where damage starts. This is what takes 8 years.
NLP framing: you spent a year building a beautiful sequence-to-sequence model, and the customer wants a scalar. Everything you optimize should be judged by whether it improves that scalar’s accuracy or the confidence around it. PCC on the vector is at best a proxy, and
16_MATH_NOTES.md§3 shows it’s a bad proxy.
2.4 The transcriptomic POD — why this field exists at all
Here is the finding the entire enterprise rests on:
Apply benchmark-dose analysis to gene expression from a SHORT-term exposure, and you get a POD that is concordant with the apical POD from a LONG-term study — including chronic toxicity and cancer.
Read that carefully. Days of exposure predict years of outcome.
| tPOD vs apical POD agreement | often within 3-fold (Frontiers 2024) |
| exposure needed | days, not years |
| what it protects against | “thought to be highly protective of all potential adverse toxicological effects” |
Why it works, mechanistically: damage doesn’t appear from nowhere. Before a tumour, before necrosis, the cell’s transcriptional program shifts — stress response, repair, proliferation. The transcriptome is the early warning. Gene expression moves first; pathology follows.
NLP framing: it’s a leading indicator. You don’t wait for the user to churn; you read the signal that precedes churn.
This is the whole bet of toxicogenomics, and it has held up well enough that regulators are building products on it.
2.5 ETAP — the actual regulatory product
Not a proposal. A thing EPA is building.
EPA Transcriptomic Assessment Product (ETAP): a human health assessment for chemicals lacking traditional toxicity data, using a standardized short-term in vivo study + standardized transcriptomic analysis to derive a reference value. (EPA)
| traditional | ETAP | |
|---|---|---|
| time | 8+ years | months |
| chemicals it targets | well-studied | data-poor |
Reviewed by EPA’s Board of Scientific Counselors in July 2023. Real assessments exist — e.g. for Perfluoro-3-Methoxypropanoic Acid.
This is the destination. A standardized short in-vivo study → transcriptomics → a number a regulator can act on, in months rather than a decade.
Everything in this book is upstream of that pipe.
2.6 Where DrugMatrix fits — and who is actually behind these papers
This is worth knowing because it explains the program’s real objective.
Scott S. Auerbach, Ph.D., DABT leads the Toxicoinformatics Group in the Predictive Toxicology Branch, Division of Translational Toxicology (DTT), NIEHS. He:
- oversees the NTP DrugMatrix resource — the dataset all four papers are built on
- built BMDExpress 2 and 3, the standard tools for genomic dose-response analysis
- co-authored S1500+, the landmark gene set (
01_BACKGROUND.md§2.3) - is corresponding author on both TransPlatformer and TransTissueFormer
So this is not an academic ML program that happens to use tox data.
It is a regulatory toxicology program that is using ML. The people who own DrugMatrix, who built the dose-response tooling, and who are driving transcriptomic PODs toward regulatory acceptance are the authors.
That should change what you consider a contribution. A method that improves MAE but produces profiles a toxicologist can’t act on is worth less than an honest negative result that de-risks a regulatory submission.
2.7 The keystone result — and the uncomfortable question it raises ⭐⭐
Johnson, Auerbach & Costa (2020), Toxicological Sciences 176(1):86 — “A Rat Liver Transcriptomic Point of Departure Predicts a Prospective Liver or Non-liver Apical Point of Departure”
The setup: 79 molecules from Open TG-GATEs. Derive a liver transcriptomic POD from short-term exposure. Compare it to the systemic apical POD.
The finding:
A liver tPOD predicted the systemic apical POD within 10× — even when that apical POD came from a non-liver endpoint.
Following subacute (29-day) dosing, liver BEPOD and systemic apical POD agreed within 10×.
Note the author list. Scott Auerbach — corresponding author on TransTissueFormer.
The uncomfortable question ⭐
If the liver transcriptome already predicts non-liver toxicity within 10×, what does cross-tissue translation add?
This is the question a reviewer will ask, and it deserves a real answer rather than a dodge. As far as I can tell there are two, and they’re different in kind:
Answer 1 — different products. The tPOD is a dose: “below 5 mg/kg, nothing bad happens anywhere.” Protective, conservative, and deliberately silent on mechanism. Cross-tissue translation produces a profile: “at 50 mg/kg the kidney shows p53 activation and oxidative stress.” You cannot get the second from the first. POD answers how much; the profile answers what and where.
Answer 2 — the 10× is doing a lot of work. Ten-fold is a wide band. For a protective screening threshold that’s fine — err low, nobody gets hurt. For deciding which organ to look at or what the mechanism is, it’s useless.
And the honest caveat: if what you need is a protective threshold, the 2020 result already delivers it, cheaply, from liver alone — and cross-tissue translation is solving a problem that is already solved. The contribution has to be on the mechanism side, or it isn’t a contribution.
This is the framing question for the whole program, and I’d want it answered before committing a year. It’s also exactly the kind of thing the corresponding author can answer in five minutes, since he wrote the 2020 paper.
2.8 The fifth paper you should know about
There is a ToxCompl successor that isn’t in your four PDFs:
Nguyen & Cong (2025), BioKDD’25 — “Completion of the DrugMatrix Toxicogenomics Database using 3-Dimensional Tensors” — ToxiTenCompl.
The idea, and it’s a good one: DrugMatrix isn’t naturally a matrix. “each probe measurement is made with a compound, a dose, a duration, for a gene within a tissue” — that’s a 3D or 4D tensor. The 2D view is a flattening that throws away structure.
ToxCompl (2D): rows = (platform, tissue, gene) x cols = treatments
ToxiTenCompl (3D): tissue x treatment x gene
| method | gradient descent, not alternating minimization (their ToxiTenCompl) |
| beats | CP decomposition and 2D matrix factorization, on MSE and MAE |
| loss | weighted MSE — higher weights on over/under-expressed genes |
| bonus | non-negative version yields interpretable tissue factors |
| stated aim | “drugs that may cross species barriers, for example, from rats to humans” |
⭐ Why this matters for everything in Chapters 2–5
The tensor formulation partly fixes a problem I spent three chapters complaining about.
In ToxCompl’s 2D layout, Cyp1a1-in-liver and Cyp1a1-in-kidney are different rows with independent parameters. The model has no idea they’re the same gene.
In the 3D tensor, gene is its own mode. The gene factor is shared across tissues by construction. That’s structurally the right move, and it’s the same instinct as “make the gene side inductive” (06_GENTOX.md §6.7) — arrived at from a different direction.
It also changes my obs/param analysis (04_TOXCOMPL.md §4.6). CP rank- on a tensor has parameters, not where . The tensor is far more parameter-efficient — which is likely why it wins, and it’s a better explanation than the paper gives.
Read this paper. It is the current state of the program’s completion work, and it partly pre-empts the criticism in Chapter 4.
2.9 What the field actually struggles with
The honest list, from the literature. These are the open problems — pick from here, not from the ML leaderboard.
| problem | status |
|---|---|
| regulatory acceptance of omics | “implementation of omics-based approaches in AOPs and their acceptance by the risk assessment community is still a challenge” |
| AOPs aren’t actually used | “regulatory agencies show limited formal adoption of AOPs despite their potential”; a “disconnect between AOP research and regulatory implementation” |
| cross-species applicability | explicitly named as a key gap |
| quantitative data insufficiency | named as a key gap |
| mixture toxicity | named as a key gap |
| workflow standardization for tPODs | active; “rigorous and reproducible wet and dry laboratory methodologies are required” |
| read-across is the current crutch | EPA, Health Canada, and EU rely on it for data-poor chemicals — it’s chemical-similarity guessing |
Two terms you’ll hit constantly
AOP (Adverse Outcome Pathway) — a structured causal chain: molecular initiating event → key events → adverse outcome. E.g. “binds receptor → activates transcription → cell proliferation → tumour.” It’s how toxicologists formalize mechanism. Regulators like the idea and barely use it. That gap is a research opportunity.
Read-across — “this new chemical looks like that known chemical, so assume similar toxicity.” It’s the default for data-poor chemicals because there’s nothing better. It’s structure-based guessing with expert judgment on top.
⭐ Notice that GenTox’s compound GNN is a learned read-across. Pretrain on 1M molecules, embed, predict response. That is exactly the read-across problem, done with representation learning instead of expert judgment — and their own ablation shows learned beats hand-crafted (
06_GENTOX.md§6.3).That framing would land much better with a tox audience than “inductive matrix factorization” does. Same method, different sentence, completely different reception.
2.10 What counts as a contribution — and to whom
You are choosing between three audiences. They want different things, publish in different places, and reward different work.
Audience A — the ML community (NeurIPS, ICML, ICLR, KDD)
| wants | novelty, benchmarks, SOTA |
| where the program already goes | BioKDD (the tensor paper) |
| your leverage | the NMT connection (12_TRANSLATION_TRANSFER.md) is genuinely novel to them |
| the risk | they don’t have the data and don’t care about rats. A tox-specific result is a poster, not an oral |
Audience B — the toxicology community (Toxicological Sciences, Arch Toxicol, SOT)
| wants | does it predict the apical endpoint? is it protective? can a regulator use it? |
| does not care about | your PCC, your architecture, your parameter count |
| does care about | the gemfibrozil→PPARα validation, the F1 0.636→0.718 downstream result |
| your leverage | almost nobody there can build these models |
| the risk | you must speak POD, BMD, AOP, concordance — not attention and embeddings |
Audience C — the regulatory/NAM community (EPA ORD, NIEHS DTT, ICCVAM)
| wants | validated, documented, reproducible, defensible |
| does not care about | novelty at all |
| the currency | a method that survives a Board of Scientific Counselors review |
| your leverage | there is a formal nomination process and a 2035 deadline |
| the risk | glacial; validation is measured in years |
The strategic read
The program’s authors sit in B and C. DrugMatrix’s owner is the corresponding author. The destination is ETAP-shaped.
You arrive with A’s toolkit. That’s the arbitrage — and it cuts both ways. Your NMT knowledge is genuinely rare here. But if you optimize for A’s metrics you’ll produce something B and C cannot use, and the people who’d have to champion it are the ones you’d be failing.
2.11 The contributions that would actually matter
Re-derived from the field’s priorities rather than the ML leaderboard. Ordered by how much a toxicologist would care.
1. Enrichment-consistency as the metric ⭐⭐ — the best fit between your skills and their needs
15_FRONTIER.md F6 argued for this on ML grounds. The field argument is stronger.
Toxicologists do not evaluate profiles by correlation. They run the profile through enrichment analysis, get a mechanism, and act on the mechanism. That is the actual downstream task. TransTissueFormer did it by hand for 3 of 44 compounds and never automated it.
“Our model recovers the correct mechanism of action for 31 of 44 compounds” is a sentence a regulator can act on. “PCC 0.793” is a sentence a regulator will ignore.
And it’s the field’s own live problem: “The Metric Picks the Winner” (June 2026) shows rankings invert with metric choice, and GenTox proved the pathology in 2024. You’d be closing a loop the program itself opened.
2. Predict the tPOD, not the profile ⭐⭐ — the reframing
Everything in this book optimizes the wrong target.
The regulatory product is a dose. So evaluate cross-tissue translation on the question that matters:
Does a predicted kidney profile yield the same kidney tPOD as the measured one?
That’s a single number, comparable to apical data, directly ETAP-shaped, and BMDExpress already computes it — built by the corresponding author.
If predicted profiles give tPODs within 2–3× of measured ones, that is a NAM candidate, and there is a formal nomination process. If they don’t, the model isn’t fit for the purpose it exists for, and that’s worth knowing in a week rather than a year.
Nobody has done this. It’s the natural bridge between the ML and the regulation, it uses existing tooling, and it’s the experiment I’d want run before any of 15_FRONTIER.md.
3. Frame the compound GNN as learned read-across 🟢 — free
Same method, a sentence a tox audience recognizes. Read-across is the current regulatory crutch for data-poor chemicals. GenTox already beats hand-crafted descriptors with it. That’s a Tox Sciences paper written in the framing, not the method.
4. Cross-species — the field’s own named gap 🟡
“Limited cross-species applicability” is listed as a key gap. The tensor paper names it as the goal. UCE’s ESM2 tokenization is the principled route (11_SC_FOUNDATION_MODELS.md §6). And PLOS One 2020 already did rat→human hepatocytes with a bottleneck DNN — “circumventing the current reliance on orthologs” — and isn’t cited in any of the four papers.
Rat data is only a proxy for humans. Everything in this program is one translation short of the thing anyone actually wants.
5. Uncertainty 🟡 — what regulators need and models don’t give
A regulator cannot act on a point estimate with no error bar. Every model here outputs bare numbers. A calibrated interval on a predicted tPOD is worth more than a better mean, because it’s the difference between “usable in an assessment” and “interesting.”
This also connects to the MaxAE/sign-flip failure (04_TOXCOMPL.md §4.7): a model that’s confidently wrong about direction is worse than one that says it doesn’t know.
6. The baselines 🟢 — still first
Unchanged from 14_RESEARCH_AGENDA.md. Everything above is conditional on the mean predictor not matching 0.793. The field’s own flagship competition ran 1,200 teams and found models not consistently beating naive baselines. This isn’t pedantry — it’s the field’s central methodological problem.
2.12 What to read, in order
The field context (none of this is ML — read it anyway)
- ETAP — About — the destination, 3 pages
- Johnson, Auerbach & Costa (2020), Tox Sci 176(1):86 — liver tPOD predicts non-liver apical. The keystone. Read §2.7’s question while you do.
- FDA NAMs page — the regulatory frame, and note in silico is a NAM
- Bioinformatic workflows for deriving tPODs (Tox Sci 2024) — current status, data gaps, research priorities. A ready-made list of open problems.
- An AOP primer — PMC5805086, concise, for toxicologists
The program’s own work you don’t have
- ToxiTenCompl (BioKDD’25) — the tensor successor to ToxCompl. §2.8.
- ToxCompl on bioRxiv — the published version
- BMDExpress 2/3 — the tool that turns your profile into the number regulators want. If you do contribution #2, this is the tool.
2.13 The summary
THE PROBLEM ~350,000 chemicals, most data-poor. Traditional assessment: 8+ years each. That queue can never clear.
THE DEADLINE EPA: eliminate all mammalian animal testing by 2035, with actions announced June 2026 and a formal NAM nomination process. FDA: NAM guidance March 2026. In silico methods are NAMs. A cross-tissue model is a candidate NAM — a regulatory category, not a metaphor.
WHAT REGULATORS WANT Not your vector. One number — the Point of Departure, the dose below which nothing bad happens.
WHY THE FIELD EXISTS Benchmark-dose analysis on short-term gene expression gives a POD concordant (often within 3×) with chronic apical PODs. Days predict years. The transcriptome moves before the pathology.
THE DESTINATION ETAP — short in-vivo study → transcriptomics → a regulatory reference value in months instead of 8 years, for data-poor chemicals. Real, reviewed, deployed.
WHO’S BEHIND THIS Scott Auerbach — leads Toxicoinformatics at NIEHS DTT, oversees DrugMatrix, built BMDExpress, corresponding author on two of your four papers. This is a regulatory toxicology program using ML, not an ML program using tox data.
THE UNCOMFORTABLE QUESTION His own 2020 paper shows a liver tPOD already predicts non-liver apical PODs within 10×. So what does cross-tissue translation add? Best answer: the POD is a dose, the profile is a mechanism — different products. But if a protective threshold is all you need, that’s already solved from liver alone. The contribution has to be on the mechanism side. Ask before committing a year.
THE FIFTH PAPER ToxiTenCompl (BioKDD’25) — 3D tensor completion, beats ToxCompl and CP, non-negative version gives interpretable tissue factors, explicitly aims at rat→human. It partly fixes the shared-gene problem Chapters 2–5 complain about, because gene becomes its own tensor mode.
WHAT WOULD ACTUALLY COUNT
- Enrichment-consistency as the metric — the real downstream task, and the mean predictor can’t game it
- Predict the tPOD, not the profile — the regulatory product is a dose; BMDExpress already computes it; nobody has tried
- Reframe the compound GNN as learned read-across — free, and it’s the field’s current crutch
- Cross-species — the field’s own named gap; rat is only a proxy for human
- Uncertainty — regulators can’t act on a bare point estimate
- The baselines — still first, still gating everything
Regulatory and field claims verified by search July 2026; links inline. The §2.7 question and the §2.8 parameter-efficiency reading are mine. Audience analysis in §2.10 is judgment, not fact — weigh it against what you learn from the people actually in the room.