Leaderboard
How the score works
Each task is scored on a handful of weighted biological questions. Every metric is rescaled against two published anchors — the floor (copy_last, or wt_identity for Task 3) and the attainable ceiling, which is half the held-out target scored against its other half. The mapping is hyperbolic, so the result is bounded by construction rather than by clipping: 50 is the floor, 100 the ceiling, and a model worse than doing nothing lands below 50 without the scale running out of room.
A metric a submission cannot produce, or that returns NaN, counts as zero skill inside its question — never dropped from the average. Full derivation →
Submission rules
- One team per person, one track per team. Ranking uses each team's best score on a task.
- Submissions are rate-limited per team per task per day, and the limit tightens in the final phase to reduce leaderboard probing.
- Validation and test ground truth are both withheld. A validation submission comes back as a score, never as the answers.
- In the P1 test phase the validation answers are not released at all — submissions are format-checked and queued, and no scores are published.
- Duplicate registrations, unauthorised data use, code sharing across teams, or falsified results are grounds for disqualification.
Reference baseline scores are published in the starter kit ↗ and regenerated on every scoring run, so they move whenever the panel does.