Evaluation

T1 — Temporal gene-expression distribution prediction. single-cell RNA · predicts expression.

Submissions

A submission is a predicted set of cells for the target condition. You submit expression, and for the spatial tasks 3D coordinates — never cell-type labels. The organisers assign types with a frozen classifier applied identically to every entry, so hidden labels are never exposed and no submission can influence how it is typed.

Requirements for a valid file · Task 1

The starter kit's Task 1 tutorial ↗ walks through building a prediction file end to end, and score_h5ad.py applies every check below to a local file before you ever upload it. A file that breaks one of these comes back as a validation error naming what was wrong — not as a low score.

  • var_names must be exactly the task gene panel — 32,285 genes — in panel order. The check is element by element; a mismatch reports the expected and received counts plus the first missing and extra names. If you have the same genes in a different order, pass --allow-reorder and the scorer reindexes for you.
  • .X must be a 2D cells × genes matrix, finite, and non-negative. Sparse and dense both load — the scorer densifies and casts to float32 either way — so float32 is worth writing yourself only to avoid a surprise in your own pipeline.
  • Not caught by validation —Values must already be log-normalised. A raw count matrix is non-negative and finite, so it passes every check and is then scored as though it were on the log scale. Nothing will tell you; the score will simply be wrong. Do not submit counts that are normalised but not log-transformed either.
  • obs["celltype"] is optional. Cells with no label are read as NA, and no label in your file affects how it is typed: the scorer trains its own probe on the held-out truth and applies it identically to every entry.
  • The number of cells is free. There is no minimum, no cap, and no correspondence to the target — every metric is distributional, so a submission need not preserve cell identity or count.
  • No coordinates. Task 1 is dissociated single cells; obsm is not read at all.

File contract · Task 1

Format
AnnData .h5ad, .X = [n_cells, n_genes]
Genes
Task gene panel, in panel order
Cells
Need not match the target count
Coordinates
Not applicable — single-cell modality
Labels
Never submitted

What the scorer reads

A submission is one AnnData file. This is the layout the scorer opens it expecting for Task 1; anything not listed is free.

AnnData object with n_obs × n_vars = <your cells> × <task gene panel>
    X       float32, log-normalised, finite, non-negative
    var     index = the task gene panel, in panel order
    obs     'celltype' optional and ignored

What is deliberately not constrained

  • Cell count. Nothing requires n_pred = n_true — a genuine growth or proliferation model may predict a different number of cells than the target has. Where a count difference would confound a comparison, both point clouds are subsampled to a shared size first.
  • Coordinate frame. Every spatial metric is invariant to translation and rotation, and outside the laterality blind spot to reflection, so no registration to the atlas is expected.
  • Cell ordering and identity. No metric assumes predicted cell i corresponds to target cell i.

Score a file locally

The starter kit exposes the same scoring path the task runners use, standalone — no baseline model involved. It loads the task’s real target itself, validates your file against the expected gene panel and order, and prints the full metric panel as JSON.

python score_h5ad.py --task T1 --input pred.h5ad
  • Task 1 is whole-transcriptome. Assuming the 500-gene MERFISH panel is the most common schema error here.
  • A collapsed, near-deterministic prediction clears the mean-based diagnostic and then loses the entire distribution group.