Results.
Our work, project by project.
What our contributors produced, compared with the projects they worked on.Reasoning tasks
Multi-hop reasoning evaluation
Tasks designed to make a frontier model fail across several connected steps. A group of four on a 9,000-person project.
Latest report · October 2026181×average tasks per expert vs. per enrollee
Terminal tasks and harnesses
Terminal benchmarks
Runnable terminal tasks with reference solutions and verifiers. A group of 16 on a 21,000-person project.
Earlier site figures · September 2026≈43×the project average per person
RL environments and rubrics
Verifiable task authoring
Graded RL tasks with reward functions and reference solutions. A group of 12 on a 2,400-person project.
Earlier site figures · September 202612×output share relative to people share
Evals and red-teaming
Adversarial evaluation
Adversarial prompts and preference comparisons targeting model failure modes. A group of 28 on a 160-person project.
Earlier site figures · September 2026≈2.6×output share relative to people share