Four experts, 402 reasoning tasks
How four experts delivered 8% of a 9,000-person project’s output.

In August 2026, four experts from our group were invited onto a project with a specific assignment: write questions that would make a new frontier model fail across several connected reasoning steps.
By the end, those four had contributed 402 of the project’s 5,000 completed tasks. There were 9,000 people enrolled. For me, their contribution was a proof of concept for the kind of expert training we are building at IsoData.
The work
Multi-hop questions require a solver to reach an intermediate conclusion, use it to find the next piece of information, and continue until there is enough evidence to answer. In this project, that often meant moving between URLs and following a trail across several sources. It was a form of red-teaming focused on reasoning and research.
The work tested whether a model could retain useful information while also using its web-search tools effectively. Finding one relevant page was only part of the task. The model had to recognize what that page revealed and decide where to look next.
In the project’s tests, each of our 402 tasks caused the model to fail on at least two of three attempts.
How we started
The project initially onboarded 1,000 people. In the first week, all four of our contributors were in the top ten. Those ten contributors together produced more than half of the tasks submitted that week.
Our experts arrived with experience and training in this kind of work. They understood how to investigate model failures, so they could begin applying that experience from day one. Within two or three days, they had developed a clearer picture of the mistakes this particular model was making.
As they learned which reasoning steps and research paths caused trouble, they could find useful failures more quickly and improve the questions they submitted.
How we found the failures
The experts were in constant conversation. When they began finding edge cases, they put in extra hours to investigate them and compare what they were seeing.
Some questions led through older sources or PDFs to an answer that appeared in only one place online. A person could follow the clues from one source to the next, but the model sometimes lost that trail. Requiring those intermediate steps made the questions difficult even when the final answer could be checked directly.
Embedded PDFs were another recurring problem. In some runs, the browser could reach the document, yet the model struggled to locate or extract the relevant passage. Some failures appeared to involve the tools themselves. Finding these cases gave the project concrete examples to investigate, both for tool fixes and for training on difficult research workflows.
By the end of the project, all four remained near the top of the rankings, with a substantial lead. Their conversations helped them build on one another’s findings, and their prior experience helped them recognize patterns sooner.
Where this leaves us
I have kept coming back to LIMA: Less Is More for Alignment since I first read it. The researchers fine-tuned an already pretrained 65-billion-parameter LLaMA model on just 1,000 carefully selected prompts and responses, with strong instruction-following results. What stayed with me was how much care went into choosing those examples.
AI needs large amounts of data, but the quality of an individual example still matters. This project made that tangible for me. An expert who understands the model’s mistakes can design a question that exposes a failure other questions miss.
Finding people who can do that can take months of onboarding and sifting through work. I think of it as looking for gold: you need people who can recognize the useful examples and know how to bring them out. At IsoData, we train experienced contributors to do that work and bring what they have learned to the next project.
Our aim is to have experts ready to contribute from day one, then train them on the specifics of your brief. A team that gets productive earlier can take on more of the work at the beginning and give a project a better chance of finishing ahead of schedule. Sharing what those experts discover can also help other contributors improve. That is the kind of team I want IsoData to build.