NeurIPS 2026 CompetitionDraft · under review

The Virtual Embryo Challenge

Generative modelling of mouse embryogenesis across space, scale and time — under genetic perturbation.

~1M
cells
11
time points
3
tasks
2
tracks

Embryogenesis is fundamental — and largely unmodelled

A single fertilised cell becomes a complete organism through spatiotemporally coordinated gene regulation, cell-fate transitions, tissue morphogenesis and organ formation. Disruptions cause congenital defects, which still affect 1 in 33 newborns and remain a leading cause of infant mortality.

Large embryo atlases and spatial-transcriptomics datasets give us snapshots, but they don't reveal how cell states transition, how local molecular changes propagate to tissue- and organ-level phenotypes, or how development responds to perturbation.

The Virtual Embryo Challenge establishes a standardised benchmark for predictive embryogenesis: a curated dataset, an evaluation pipeline, baseline models, and three tasks that jointly stress spatial context, multiscale reasoning, temporal dynamics and perturbation response.

Three tasks, one shared atlas

Each task uses staged train / validation / hidden-test splits over the same whole-embryo and heart-focused resource. Validation and test ground truth are both withheld; final rankings reflect generalisation to held-out stages, embryos and genotypes.

Human-designed vs agent-designed, scored side by side

Both tracks address the same three tasks and are scored on the same metrics and hidden test sets. Prizes are awarded separately so the leaderboards directly contrast the two approaches.

Track 1, Human Team: Conventional ML-competition workflow.
Track 1Human Team
Conventional ML-competition workflow.

Methods designed and supervised by human participants. Algorithm/model design → submission → evaluation. Standard NeurIPS competition track.

Track 2, Agent Team: Coding agents / LLM-driven recursive systems.
Track 2Agent Team
Coding agents / LLM-driven recursive systems.

Methods produced by coding agents or LLM-based evolutionary algorithms that iteratively select, mutate, and evaluate candidate solutions without human supervision. The agent system and the complete evolutionary trace must be shared before prize evaluation.

Multimodal whole-embryo perturbation resource

~1 million cells across 11 developmental time points, spanning early gastrulation through cardiac progenitor emergence, heart-tube formation, looping and later morphogenesis.

Single-cell
Whole-embryo per-cell RNA across staged embryos from E6.75 to E12.5. Whole transcriptome, not the MERFISH panel.
Spatial
Sections decoded into per-cell 3D positions plus measured RNA on a 500-gene panel across the same developmental window.
Annotation
Per-cell cell-type, tissue-domain and anatomical-region calls. Released with training data but never submitted.
Perturbations
Mab21l2 (E9.5), Gata4 and β-catenin (E8.75) — three cardiac developmental regulators with paired wild-type controls.

Per-task file inventory →

Scored against a measured floor and an attainable ceiling

Every metric answers one biological question, and each is rescaled against two published anchors before the questions are weighted together. Read every score against those anchors — on Task 3, predicting no change at all already reaches 0.956 absolute pseudobulk correlation.

  • Did the right genes change? — Whether the model moved the genes development or a knockout actually moves, in the right direction and — for perturbations — by the right amount.
  • Is the distribution right? — Whether the predicted population of cells matches the observed one: the right cell states in the right proportions, with the gene-gene covariance that per-gene summaries cannot see.
  • Is the tissue the right shape? — For the spatial tasks, whether the predicted embryo has the right form, size and local neighbourhood structure — scored independently of expression.

The questions differ by task and are not equally weighted. Each metric's skill is mapped so that the floor sits at 50 and the attainable ceiling at 100, then combined at the declared weights — no clipping anywhere.

Every metric, its derivation, and how they combine →

Three scoring questions — differential expression, cell-state distribution and tissue shape — above a scale running from a floor of 50 to a ceiling of 100.
T1
DE gene recovery 25% · Change direction 25% · Cell-state distribution 30% · Gene-gene co-variation 20%
T2
Expression change 25% · Cell-state distribution 25% · Tissue shape and growth scale 25% · Local spatial organisation 25%
T3
Response gene recovery 30% · Response direction 25% · Response magnitude 25% · Cell-state distribution 20%

Launch → development → final

2026-06-30
Site, submission portal and evaluation platform live
2026-07-20
Starter kit released; website opens to participants
2026-07-30
P1 · Test phase begins
2026-08-15
P2 · Development phase begins; validation leaderboard opens

Full timeline through the NeurIPS announcement →

$104K from the Laude Institute Moonshots Seed Grant

$54K
Winner prizes
Per track ($27K × 2): one $8K first prize, two $5K second prizes, three $3K third prizes. Tracks are scored on the same hidden tests but awarded separately.
$30K
Travel awards
15–20 grants for early-career researchers to attend the NeurIPS workshop.
$20K
Outreach & education
Website, starter-kit repo, tutorials, reproducible walkthroughs, baseline documentation, participant communication channels.

Prize breakdown and compute →

Baselines, metrics and tutorials are public

The starter kit ships the exact evaluation code, a floor, a simple method and one dynamics reference model per task, plus a standalone scorer for any submission file. The baselines are transparent reference points, not competitive upper bounds. The same atlas you can browse on this site is what the tasks are built on.

Common questions

Who can participate?

Anyone — academic, industry, independent — who can submit a prediction file conforming to the required format. Each person joins one team; each team works on one track.

What exactly do I submit?

A predicted set of cells for the target condition: expression for Task 1, expression plus 3D coordinates for Tasks 2 and 3. You never submit cell-type labels — the organisers assign types with a frozen classifier applied identically to every submission, so hidden labels are never exposed.

Is the data really released?

Training data is released for method development. Validation and test ground truth are both withheld: a validation submission returns a leaderboard score, and the answers are released only when the final test set is.

What are the floor and the ceiling?

Two anchors published with every score. The floor is copy_last — predict the preceding stage verbatim (wt_identity for Task 3). The ceiling is half the target scored against its other, disjoint half — deliberately not the target resubmitted verbatim, which is a number no honest generative model can reach.

All 15 questions →

Organisers

Hosted by the Qiu Lab at Stanford University, in collaboration with researchers at Harvard, UC San Diego, MBZUAI, CMU and GenBio. The full organising committee and contributor roles will be listed alongside the starter-kit release.

Reach the organisers at [email protected] for questions about the competition, data, prize logistics or partnerships. A public submission portal and discussion forum go live around two weeks before the P1 test phase.

This page summarises the NeurIPS 2026 competition proposal currently under review. Dates, datasets, prize amounts and exact metric formulations are subject to change between proposal acceptance and launch. Task splits, the submission contract and the scoring table follow the public baselines repo; where the two differ, the repo is authoritative.