Results.

Our work, project by project.

What our contributors produced, compared with the projects they worked on.

Reasoning tasks

Multi-hop reasoning evaluation

Tasks designed to make a frontier model fail across several connected steps. A group of four on a 9,000-person project.

Latest report · October 2026
Share of the project9,000 people
Completed output8.04%
People0.0444%
181×average tasks per expert vs. per enrollee
How this is calculated

(402 ÷ 4) ÷ (5,000 ÷ 9,000) = 180.9×. The baseline includes all 9,000 enrollees, including anyone with no completed tasks. Active participation and hours are unknown.

The latest count is completed tasks; acceptance has not been reconfirmed. These are reported individual contributions, not independently audited company results.

Earlier Frontier agent benchmark record

The original site reported 16 of 9,000 people producing 10% of accepted output: 10% ÷ (16 ÷ 9,000 × 100) = 56.25×. Its relationship to the latest four-expert report is unconfirmed, so it is retained here rather than counted as another project.

Terminal tasks and harnesses

Terminal benchmarks

Runnable terminal tasks with reference solutions and verifiers. A group of 16 on a 21,000-person project.

Earlier site figures · September 2026
Accepted tasks per person21,000 people
Our experts≈150
Project average3.5
≈43×the project average per person
How this is calculated

About 150 accepted tasks per expert ÷ 3.5 accepted tasks per project contributor ≈ 42.9×.

Restored from our original site. These historical self-reported figures are awaiting reconfirmation and have not been independently audited. The project population was reported at joining; active participation and hours were not recorded.

RL environments and rubrics

Verifiable task authoring

Graded RL tasks with reward functions and reference solutions. A group of 12 on a 2,400-person project.

Earlier site figures · September 2026
Share of the project2,400 people
Accepted output6%
People0.5%
12×output share relative to people share
How this is calculated

12 ÷ 2,400 × 100 = 0.5% of people. 6% of accepted output ÷ 0.5% of people = 12×.

Restored from our original site. These historical self-reported figures are awaiting reconfirmation and have not been independently audited. The project population was reported at joining; active participation and hours were not recorded.

Evals and red-teaming

Adversarial evaluation

Adversarial prompts and preference comparisons targeting model failure modes. A group of 28 on a 160-person project.

Earlier site figures · September 2026
Share of the project160 people
Accepted output45%
People17.5%
≈2.6×output share relative to people share
How this is calculated

28 ÷ 160 × 100 = 17.5% of people. 45% of accepted output ÷ 17.5% of people ≈ 2.57×.

Restored from our original site. These historical self-reported figures are awaiting reconfirmation and have not been independently audited. The project population was reported at joining; active participation and hours were not recorded.