Research notes · Agent evaluations
Testing agents beyond the examples
A small sample dataset we built: five expert-written tasks testing how AI-generated programs handle unfamiliar inputs.

This is a very small sample dataset of five tasks we built. We asked AI agents to build programs from expert-written requirements and examples. We then tested their code on unseen cases to find where it worked—and where it broke.
On this deliberately difficult set, none of the 25 submitted programs passed every check. In follow-up tests across three tasks, all 15 programs improved when the same test records were reworded to resemble the examples, without asking the AI to fix its programs.
Five experts contributed one executable task each: reproduce a CNC interpreter, reconcile pharmacy records, reduce river field books, certify a fare engine and audit expense claims.
These are red-team evaluations of reliability. They target unfamiliar wording, rare software behavior and interactions between rules. The task files include instructions, examples, environments and graders.
From task to checked result
- 01Task briefRules, inputs and public examples
- 02Agent writes codeA program or test suite
- 03Hidden casesNew records or software builds
- 04Grader checksOutputs against the required answers
The supplied Claude Opus 5.5 submissions made substantial progress. The closest CNC programs matched 590 of 600 hidden programs. Every fare suite caught 30 of 31 faulty builds and accepted all five correct builds. Yet none of the 25 main submissions earned a full solve.
That result reflects a deliberately difficult selection: solved candidates were dropped during development. Several submissions were regraded after task revisions. The scores describe these failure cases, rather than an average success rate across professional work.
When familiar wording matters
One pharmacy log says “quantity charted 2 in error, actual 1.” Three submitted parsers kept the original quantity; two treated the correction as a void. All five matched the ten public example shifts. The strongest reconciled only five of 48 hidden shifts completely.
To examine that gap, we regraded saved programs on the same hidden events, rewritten in the phrasing used by the public examples. Expected answers stayed fixed. The agent didn’t get another attempt to repair its solver.
Same programs, different wording
Hidden cases correct · five programs per task
Pharmacy notes
- Program 1: original wording, 0 of 48 correct shifts; familiar wording, 41 of 48.
- Program 2: original wording, 0 of 48 correct shifts; familiar wording, 39 of 48.
- Program 3: original wording, 5 of 48 correct shifts; familiar wording, 42 of 48.
- Program 4: original wording, 2 of 48 correct shifts; familiar wording, 41 of 48.
- Program 5: original wording, 0 of 48 correct shifts; familiar wording, 41 of 48.
Stream fieldbooks
- Program 1: original wording, 0 of 60 correct fieldbooks; familiar wording, 60 of 60.
- Program 2: original wording, 5 of 60 correct fieldbooks; familiar wording, 60 of 60.
- Program 3: original wording, 5 of 60 correct fieldbooks; familiar wording, 60 of 60.
- Program 4: original wording, 0 of 60 correct fieldbooks; familiar wording, 58 of 60.
- Program 5: original wording, 0 of 60 correct fieldbooks; familiar wording, 60 of 60.
Expense claims
- Program 1: original wording, 42 of 56 correct claims; familiar wording, 49 of 56.
- Program 2: original wording, 42 of 56 correct claims; familiar wording, 56 of 56.
- Program 3: original wording, 27 of 56 correct claims; familiar wording, 51 of 56.
- Program 4: original wording, 23 of 56 correct claims; familiar wording, 35 of 56.
- Program 5: original wording, 36 of 56 correct claims; familiar wording, 50 of 56.
Read the exact values
| Task | Program | Original | Familiar wording |
|---|---|---|---|
| Stream fieldbooks | sg-1 | 0 / 60 | 60 / 60 |
| Stream fieldbooks | sg-2 | 5 / 60 | 60 / 60 |
| Stream fieldbooks | sg-3 | 5 / 60 | 60 / 60 |
| Stream fieldbooks | sg-4 | 0 / 60 | 58 / 60 |
| Stream fieldbooks | sg-5 | 0 / 60 | 60 / 60 |
| Expense claims | ec-1 | 42 / 56 | 49 / 56 |
| Expense claims | ec-2 | 42 / 56 | 56 / 56 |
| Expense claims | ec-3 | 27 / 56 | 51 / 56 |
| Expense claims | ec-4 | 23 / 56 | 35 / 56 |
| Expense claims | ec-5 | 36 / 56 | 50 / 56 |
| Pharmacy notes | cn2-1 | 0 / 48 | 41 / 48 |
| Pharmacy notes | cn2-2 | 0 / 48 | 39 / 48 |
| Pharmacy notes | cn2-3 | 5 / 48 | 42 / 48 |
| Pharmacy notes | cn2-4 | 2 / 48 | 41 / 48 |
| Pharmacy notes | cn2-5 | 0 / 48 | 41 / 48 |
Four stream-gauging programs then passed all 60 hidden books; the remaining program passed 58. One expense program passed all 56 hidden claims. Pharmacy improved to 39–42 of 48 shifts.
Our experts noticed that the programs often depended on familiar phrasing. Code that handled the examples could still miss the same kinds of information written differently. In this small sample, rewording the test records helped the same programs get more answers right.
The pharmacy diagnostics preceded a policy clarification about converting drug amounts to units. Six remaining misses involved that form, so they don’t establish failure under the clarified rule.
A different way to present the same work
The pharmacy database control presents the same shifts and answers as structured data. All three fresh attempts solved it. The successful attempts show that these rules were solvable through a structured interface. The log and database versions used separate programs, so the comparison cannot isolate wording from every other difference.
Same 48 pharmacy shifts, two formats
| Input format | Full solves |
|---|---|
| Written logs Corrections and events in prose | 0 / 5 attempts |
| Structured database Events organized into fields | 3 / 3 attempts |
The practical question is which parts of a workflow survive a change in its inputs. The examples below show a specific failure from each task, alongside the required behavior. A fresh evaluation on a frozen release would test whether new attempts reproduce these failures.
Inside a task
CNC machining
G-code toolpath parity
Reproduce the exact tool path of the LinuxCNC 2.9.4 rs274 interpreter in standard-library Python, cutter compensation and canned cycles included.
- Deliverable
simulate(program, setup)- Examples
- 59 public
- Hidden set
- 600 programs
edge-0020G1 X7.298 Y0 F400 G2 X-7.298 Y0 I-7.298 I7.298 G1 X14.596
Submitted program behavior
All five attempts skipped the I7.298 block and returned 5 moves.
What was required
6 moves. Cut a full circle back to the same point, then feed to X14.596.
G2 is still active. A block with only an I word is another clockwise arc; without an end point, it ends where it started.
Behavior summarized from the dataset’s failure notes. Move counts refer to the full program. Full transcripts and submitted programs are not included in the download.
Inspect the full release
Instructions, environments, graders and hidden tests.
Five task families, including both pharmacy formats.
Method, versions and limits
The supplied records identify Claude Opus 5.5 (claude-opus-5-5) and fresh session subagents in Claude Code with shell access, on 9–10 October 2026. Restrictions on test, solution and web access were instruction-enforced and audited, rather than technically isolated. These are development evaluations; the public release contains summaries, not full session transcripts.
Five main task families have five submitted programs each. The pharmacy database control has three separate, fresh attempts. Expense, stream and fare programs were written for earlier drafts and regraded on the released tasks; pharmacy received a policy clarification after its attempts. This is not a set of 25 fresh attempts against one frozen release.
The wording diagnostics use saved programs and preserve the underlying events and expected answers. Expense diagnostics adapt hard-coded table paths for local execution. They measure output correctness, which may omit file, format or other full-task conditions. Source artifact hashes and exact values are included in the diagnostic data; the underlying saved programs are not in the public bundle.
Recorded main-attempt times range from 20 to 150 minutes. Summaries list a nominal 90-minute budget; downloadable configurations allow eight hours. These settings differ. No timing or cost comparison is inferred. Results cover a selected set, one reported model and one agent setup; they have not been independently rerun for this article.
The task archive uses Harbor layout. CNC and fare reference solutions are held back; graders and hidden tests are included. The other four task folders include reference solutions. The archive does not pin a Harbor version or configure the reported network restrictions. Its reproduction commands have not been revalidated for this article.
Apache 2.0 is declared in the supplied dataset. Preserve the canary strings and exclude these evaluation tasks from training data. People and companies in the pharmacy and expense records are fictional.
Build evaluations for your domain. Work with IsoData