{"task_id": "gcode-toolpath-parity", "title": "G-code toolpath parity", "domain": "CNC machining", "summary": "Reproduce the exact tool path of the LinuxCNC 2.9.4 rs274 interpreter in standard-library Python, cutter compensation and canned cycles included.", "deliverable": "simulate(program, setup)", "path": "gcode-toolpath-parity", "instruction": "Our web program checker must predict the exact tool-center path a LinuxCNC 2.9 mill will run, to check fixture clearance, on a server without LinuxCNC. Write `/app/toolpath.py`, standard library only, exposing `simulate(program: str, setup: dict) -> dict`. Ground truth is the `rs274` standalone interpreter from Debian's `linuxcnc-uspace` package, version `1:2.9.4-2+deb13u1`, installed here: `python3 /app/data/probe/run_probe.py <program.ngc> <setup.json>` runs it with our metric machine configuration and prints its canonical calls; add `--number` to tag each line with its line number so a failing block shows which line it is.\n\n`setup` is `{\"tools\": {\"<number>\": {\"diameter\": mm, \"length\": mm}}, \"work_offsets\": {\"G54\": [x, y, z], ..., \"G59\": [...]}, \"g28_position\": [x, y, z], \"g30_position\": [x, y, z]}`, all machine coordinates in mm. The program starts at machine zero with no tool loaded, in LinuxCNC's power-on modal state.\n\nReturn `{\"moves\": [...], \"error\": null}`. `moves` lists every motion and dwell the interpreter commands, in order, including zero-length motions and inserted corner arcs:\n\n- `{\"type\": \"rapid\", \"to\": [x, y, z]}` for a straight traverse\n- `{\"type\": \"feed\", \"to\": [x, y, z], \"feed\": f}` for a straight feed\n- `{\"type\": \"arc\", \"to\": [x, y, z], \"plane\": \"XY\"|\"XZ\"|\"YZ\", \"center\": [a, b], \"turns\": n, \"feed\": f}`, where `center` holds the plane's two axes in that order (XY: x, y; XZ: x, z; YZ: y, z) and `turns` is the signed integer rs274 prints: positive is counterclockwise seen from the positive end of the plane's normal axis (Z, Y, X), magnitude 1 for a single turn or partial arc\n- `{\"type\": \"dwell\", \"seconds\": s}`\n\nPositions are machine coordinates of the controlled point in mm: work position plus the active G54..G59 offset, the G92/G52 offset and the tool length offset, converted to mm in a G20 program. `feed` is the feed rate in effect for that motion, in mm/min. No other keys.\n\nIf rs274 stops with an error, return the motions it printed before stopping and `\"error\": {\"line\": n, \"code\": c}`, where `n` is the 1-based line of `program` holding the block rs274 reports as failing and `c` maps from its message:\n\n- `Radius to end of arc differs from radius to start` \u2192 `ARC_RADIUS_MISMATCH`\n- `Arc radius too small to reach end point` \u2192 `ARC_RADIUS_TOO_SMALL`\n- `Current point same as end point of arc` \u2192 `ARC_SAME_ENDPOINT`\n- `Length of cutter compensation entry move is not greater than the tool radius` or `Radius of cutter compensation entry arc is not greater than the tool radius` \u2192 `COMP_ENTRY_TOO_SHORT`\n- `Tool radius not less than arc radius with comp` \u2192 `COMP_ARC_TOO_SMALL`\n- any message ending in `without gouging` \u2192 `COMP_GOUGE`\n- `The move just after exiting cutter compensation mode must be straight, not an arc` \u2192 `COMP_EXIT_ARC`\n- `R less than z in cycle in xy plane` \u2192 `CYCLE_R_BELOW_Z`\n- `Cannot use axis values without a g code that uses them` \u2192 `AXIS_WORDS_WITHOUT_MOTION`\n\nInput domain: G0 G1 G2 G3 G4 G10 L2/L20 (P0 to P6) G17 G18 G19 G20 G21 G28 G30 G40 G41 G42 G43 G43.1 G43.2 G49 G52 G53 G54 to G59 G64 G73 G80 G81 G82 G83 G90 G91 G90.1 G91.1 G92 G92.1 G92.2 G92.3 G94 G98 G99; M2 M3 M5 M6 M8 M9 M30; words D F H I J K L N P Q R S T X Y Z; `( )` and `;` comments; any letter case and spacing. G20 or G21 appears, if at all, only in the first block that has a G or M word. Every T, D and H number is in the tool table, and a diameter can be negative. Canned cycles and cutter compensation are used only in G17. M6 never shares a block with G43/G43.1/G43.2/G49 and never follows a tool length offset change without a motion in between. Each program ends with M2 or M30 or is stopped by one of the errors above. Programs whose outcome hinged on a last-bit floating point tie were dropped.\n\nGrading: the grader has no LinuxCNC. Each program gets a fresh process, running as an unprivileged user that cannot start processes or open sockets, which imports a copy of `/app/toolpath.py` and makes one `simulate` call with 2 seconds of CPU. The 600 hidden programs are job-shop programs, randomized programs across the domain, and short corner-case programs. Discrete values must match exactly; positions, centers and feeds within 0.01 (mm or mm/min); dwells within 0.001 s. The task passes only if every hidden program matches. Examples with expected answers are in `/app/data/examples/`.\n\nYou have 28800 seconds to complete this task.\n", "public_examples": 59, "hidden_cases": 600, "hidden_unit": "programs", "grading": "Discrete fields exact, coordinates within 0.01 mm", "answer_key": "held_back", "answer_key_sha256": "b95ccb1b52361f7014326e699b4550dc148c31b7af14d6b514a296f8164989f5", "model": "claude-opus-5-5", "blind_attempts": 5, "blind_solved": 0, "closest_attempt": "10 of 600 hidden programs wrong"}
{"task_id": "controlled-count-close-notes", "title": "Controlled-substance count close", "domain": "Hospital pharmacy", "summary": "Reconcile automated dispensing cabinets at shift close from the controlled-substance log that nurses, technicians and pharmacists write.", "deliverable": "reconcile(shift_dir)", "path": "controlled-count-close/notes", "instruction": "The inpatient pharmacy closes the controlled-substance books for its dispensing cabinets at the end of every shift, by hand, and wants that automated. /app/shifts/ holds ten past shifts that were reconciled by hand; their reports are in /app/examples/. Each shift folder holds log.txt, the shift's controlled-substance log as the nurses, technicians and pharmacists write it, one numbered entry per line (`#<entry> <time written> <author>: <text>`); roster.csv, the staff on the shift with their roles; and start_counts.csv, the count of every cabinet pocket when the shift opened.\n\nThe rules are pharmacy policy CS-14 in /app/policy.md, which also explains the floor shorthand and how entries point back to earlier ones. Products, count units, unit strengths, volumes and the names staff use for each product are in /app/formulary.csv.\n\nMost of the work is reading the log correctly. Entries are free text: one transaction can be worded many ways, and a word that often goes with one kind of entry can turn up in another. Decide what each entry records from what it says as a whole, the way a pharmacist reading the log would, then apply the policy.\n\nWrite /app/reconcile.py, a single Python 3.13 file that uses only the standard library and defines `reconcile(shift_dir)`. Given the path of one shift folder like those in /app/shifts, it returns the shift-close report as a dict:\n\n```\n{\"on_hand\": {\"<cabinet>\": {\"<product code>\": <int>, ...}, ...},\n \"open_discrepancies\": [\"<entry number>:<product code>\", ...]}\n```\n\n`on_hand` has one integer for every cabinet and product in the shift's starting counts: the expected count at shift close. `open_discrepancies` lists the id of every discrepancy that is open at shift close, as the policy defines it, each once, in any order. The function reads /app/formulary.csv and the shift folder only. It must not use the network or start other processes.\n\nAlso run it on every folder in /app/shifts and save each report as JSON at /app/reports/<folder name>.json.\n\nGrading: reconcile() is called on 48 hidden shift folders and on the 10 public shifts. The public shifts are the mild end of the set: two cabinets each, about 35 entries, few corrections and late entries, clean typing, and only a small share of the phrasings staff use. The hidden shifts come from the same units, products, policy and log format, but most cover three cabinets and run to about 65 entries (up to 80), with more corrections, late entries and back references. Their logs use many phrasings, abbreviations and number formats (such as `.5mg`) that the public logs do not show, and typos. Every entry is still clear to a careful pharmacist who knows the policy. Each call runs as an unprivileged user with no network and a 10 second CPU limit, with /app/formulary.csv in place. You pass only if, for every shift, every on-hand count and the exact set of open discrepancy ids match. Your saved reports for the public shifts must match as well.\n\nYou have 28800 seconds to complete this task.\n", "public_examples": 10, "hidden_cases": 48, "hidden_unit": "shifts", "grading": "Exact counts and exact set of open discrepancy ids", "answer_key": "included", "answer_key_sha256": null, "model": "claude-opus-5-5", "blind_attempts": 5, "blind_solved": 0, "closest_attempt": "5 of 48 hidden shifts right, 97.5% of pocket counts right"}
{"task_id": "controlled-count-close-db", "title": "Controlled-substance count close (database control)", "domain": "Hospital pharmacy", "summary": "The same shifts and answers as a clean SQLite export.", "deliverable": "reconcile(shift_dir)", "path": "controlled-count-close/db", "instruction": "The inpatient pharmacy closes the controlled-substance books for its dispensing cabinets at the end of every shift, by hand, and wants that automated. /app/shifts/ holds ten past shifts that were reconciled by hand; their reports are in /app/examples/. Each shift folder holds shift.sqlite, the cabinet system's export for that shift: the shift window, the roster with roles, the starting count of every cabinet pocket, and every numbered log entry with its transactions, co-signs, corrections and pharmacist resolutions.\n\nThe rules are pharmacy policy CS-14 in /app/policy.md. Products, count units, unit strengths and volumes are in /app/formulary.csv.\n\nHow the export stores things: times are local `YYYY-MM-DD HH:MM`. `entry.late_for` is set only on a late entry, and holds the time it states. In `txn`, `units` is the number of count units moved; `dose_mcg` and `amount_mcg` are in micrograms, and a NULL `dose_mcg` on a removal means no dose was written; `ref_entry` is the entry a waste, return or receipt is against; `other_cabinet` is the other end of a transfer. A `cosign` row is a co-sign by the entry's author. A `correction` row either voids or amends `target_entry`; for an amendment `field` is one of `qty`, `dose`, `amount`, `witness`, and `new_value` is written in the same units as the column it replaces (units, mcg or an employee id). A `resolution` row names the discrepancy by `target_entry` and `product`.\n\nWrite /app/reconcile.py, a single Python 3.13 file that uses only the standard library and defines `reconcile(shift_dir)`. Given the path of one shift folder like those in /app/shifts, it returns the shift-close report as a dict:\n\n```\n{\"on_hand\": {\"<cabinet>\": {\"<product code>\": <int>, ...}, ...},\n \"open_discrepancies\": [\"<entry number>:<product code>\", ...]}\n```\n\n`on_hand` has one integer for every cabinet and product in the shift's starting counts: the expected count at shift close. `open_discrepancies` lists the id of every discrepancy that is open at shift close, as the policy defines it, each once, in any order. The function reads /app/formulary.csv and the shift folder only. It must not use the network or start other processes.\n\nAlso run it on every folder in /app/shifts and save each report as JSON at /app/reports/<folder name>.json.\n\nGrading: reconcile() is called on 48 hidden shift folders exported by the same system and on the 10 public shifts. The public shifts are the mild end of the set: two cabinets each, about 35 entries, few corrections and late entries. The hidden shifts come from the same units, products and policy, but most cover three cabinets and run to about 65 entries (up to 80), with more corrections, late entries and back references. Each call runs as an unprivileged user with no network and a 10 second CPU limit, with /app/formulary.csv in place. You pass only if, for every shift, every on-hand count and the exact set of open discrepancy ids match. Your saved reports for the public shifts must match as well.\n\nYou have 28800 seconds to complete this task.\n", "public_examples": 10, "hidden_cases": 48, "hidden_unit": "shifts", "grading": "Exact counts and exact set of open discrepancy ids", "answer_key": "included", "answer_key_sha256": null, "model": "claude-opus-5-5", "blind_attempts": 3, "blind_solved": 3, "closest_attempt": "solved in 12 to 25 minutes"}
{"task_id": "stream-gauging-fieldbook", "title": "Stream gauging field books", "domain": "Hydrology", "summary": "Reduce hydrographers' transcribed field books to discharge, area, mean velocity and mean gauge height by the midsection method.", "deliverable": "reduce.py BOOK.txt OUT.json", "path": "stream-gauging-fieldbook", "instruction": "Our hydrometric section is clearing a backlog of discharge measurements. The hydrographers' field books have been transcribed to plain text exactly as written, each in that hydrographer's own shorthand, and the office now has to compute every measurement under our SOP without anyone retyping the notes.\n\nWrite `/app/gauging/reduce.py`:\n\n    python3 /app/gauging/reduce.py BOOK.txt OUT.json\n\nIt reads one transcribed field book plus the current-meter ratings in `/app/data/meters.csv` and writes:\n\n    {\"discharge_m3s\": Q, \"area_m2\": A, \"mean_velocity_ms\": V, \"mean_gauge_height_m\": GH}\n\nHow a measurement is computed is set out in `/app/docs/SOP_discharge_midsection.md`: the midsection method, edges, walls and piers, 0.6, 0.2/0.8 and three-point verticals, angle coefficients, meter ratings and meter changes, unit changes, struck-out and repeated readings, later corrections, and the time-weighted gauge height. The book is the only source for a measurement; the meter file covers every meter we use.\n\n`/app/examples/` holds ten measurements from our own books with the office's computed results (`book.txt`, `expected.json`). These are the tidy end of what comes in. The held-out books come from other hydrographers on other rivers, wading and from bridges, in metres or feet, and use the full range of field-book habits and shorthand seen across the agency.\n\nGrading: your program is run on 60 held-out books, one book per run, and must get every one right. Every held-out book was checked to determine its measurement completely under the SOP. Discharge, area and mean velocity must each be within 0.1%, and the mean gauge height within 0.001 m. The books fix every value exactly, so these allowances only cover floating point and unit conversion. Each run gets 20 seconds of CPU and Python 3.13 with the standard library and numpy; it may read the book it is given and the meter file, and write only its output file.\n\nYou have 28800 seconds to complete this task.\n", "public_examples": 10, "hidden_cases": 60, "hidden_unit": "books", "grading": "Discharge, area and velocity within 0.1%, gauge height within 0.001 m", "answer_key": "included", "answer_key_sha256": null, "model": "claude-opus-5-5", "blind_attempts": 5, "blind_solved": 0, "closest_attempt": "5 of 60 hidden books right"}
{"task_id": "fare-engine-certification", "title": "Fare engine certification", "domain": "Transit fares, test engineering", "summary": "Write a pytest certification suite from a written fare policy, with no fare engine to copy from. Graded against hidden correct and faulty vendor builds.", "deliverable": "a pytest suite", "path": "fare-engine-certification", "instruction": "Tollmere Transit is certifying fare engines from outside vendors against its 2027 card fare rules. Write the certification suite that the agency's test rig will run against each vendor build.\n\nThe rules are in /app/docs/fare_rules.md and are the only authority. Every build implements one function, `farecalc.price_account(rider, taps)`, described in section 9 of the rules. /app/docs/examples.json has a few accounts priced correctly. No engine is installed here; you may write your own to develop against.\n\nDeliverable: a pytest suite in /app/suite/. That is one or more `test_*.py` files, optionally a `conftest.py` and `.json` data files, no subdirectories, at most 40 files and 512 KB in total. The suite must `import farecalc` and use only `farecalc.price_account`. It may use pytest 9.1.1 and the Python standard library, nothing else.\n\nHow it is run: `python3 /app/rig/run_suite.py /app/suite --engine DIR` runs the suite exactly as certification does, against the engine package at DIR/farecalc (default DIR is /app/engine). The build runs in a separate process, and inside the suite `farecalc` is a thin client that forwards `price_account` calls to it. Nothing else from the package is available and the build's source cannot be read. One run allows at most 80 `price_account` calls, at most 60 taps per call, and 60 seconds. A call past a limit raises RuntimeError. The suite cannot start processes, open sockets or load native code (ctypes, cffi and similar modules will not import). For every build it runs under a new user with a fresh copy of /app/suite and an empty temp dir, and the builds come in a random order, so nothing carries over from one build to the next.\n\nThe declared input domain:\n\n- `rider` is \"adult\" or \"concession\".\n- `taps` is a list of tap dicts as in section 9, in strictly increasing time order, with unique non-empty string ids, stops from the table in section 1, and times within calendar year 2027.\n- Each reversal names the id of an earlier tap-on, and no tap-on is reversed twice.\n\nOutside this domain behaviour is unspecified: a build may reject such input with any exception, or price it any way it likes. Inside it, a build may vary only what section 9 leaves open: the text of journey labels (only which trips share a label is specified) and the order of keys in `days`. Results are plain dicts, lists, strings, ints and None.\n\nGrading: the suite is run against a hidden set of builds. Some are correct, independent implementations of the rules; the suite must pass on every one of them. The rest each misread the fare rules somewhere; the suite must fail (at least one failing or erroring test) on every one of them. Each faulty build gives a wrong statement for some input in the declared domain. The result is all or nothing: failing one correct build, or passing one faulty build, scores zero.\n\nYou have 28800 seconds to complete this task.\n", "public_examples": 9, "hidden_cases": 31, "hidden_unit": "faulty builds caught", "grading": "Pass all 5 correct builds, reject all 31 faulty builds", "answer_key": "held_back", "answer_key_sha256": "7c5abd2a31770ebd41afc65fb16d20a1b0b6bf01393809cdd017783a05003f0a", "model": "claude-opus-5-5", "blind_attempts": 5, "blind_solved": 0, "closest_attempt": "1 of 31 faulty builds missed"}
{"task_id": "expense-claim-audit", "title": "Expense claim audit", "domain": "Accounts payable", "summary": "Audit travel and expense claims from a corporate card feed and the employee's own notes under a one-page policy.", "deliverable": "audit(claim_dir)", "path": "expense-claim-audit", "instruction": "Build `/app/audit.py`: one Python 3.13 file, standard library only, with a function `audit(claim_dir)` that audits a single travel and expense claim under policy T&E-7 (`/app/policy.md`).\n\n**Input.** `claim_dir` is a folder shaped like the ones under `/app/claims/`:\n\n- `trip.csv` - who travelled, where, and when they left and came back;\n- `card.csv` - the corporate card transactions included in the claim;\n- `notes.txt` - the employee's free-text lines, one per annotated card transaction or added item;\n- `prior_reimbursed.csv` - out-of-pocket items paid on earlier claims.\n\nDestination tiers come from `/app/cities.csv` and daily dollar rates from `/app/fx_rates.csv`. Read nothing else. The grader blocks network access, starting processes, and importing `socket`, `ssl`, `ctypes` or `multiprocessing` (directly or through another module), so keep to the rest of the standard library.\n\n**Output.** A dict with exactly these three keys and no others, for example\n\n```\n{\"pay_employee\": \"312.07\", \"recover\": \"48.10\", \"flags\": [\"C1042\", \"L4\"]}\n```\n\nDollar amounts are strings with two decimal places. `flags` holds each flagged transaction or line id once, in any order.\n\n**Reports.** Once the module works, run it over each folder in `/app/claims/` and write the dict to `/app/reports/<folder>.json`. The accepted audits of those ten claims are in `/app/examples/` for comparison.\n\n## About the notes\n\nEmployees write their notes however they like, so the same expense turns up in different words. Keywords are a poor guide: \"airport\" appears in taxi, train and parking lines alike, and \"hotel\" can sit in a line about a taxi. Work out which expense a line describes, as an accounts-payable reviewer would, and only then apply T&E-7.\n\n## The ten examples and the hidden set\n\nThe ten claims in `/app/claims/` are easy: two- or three-day trips, short card feeds, few added lines, hardly any corrections, no typing mistakes, and a narrow selection of the ways people phrase things. The hidden set adds 56 claims under the same policy, cities and currencies. Those trips last from one to six days and carry more card charges, more out-of-pocket items, more corrections and withdrawals, typos, and wording, shorthand and number formats that never appear in the ten examples. None of them is ambiguous to someone reading it against T&E-7.\n\n## Scoring\n\nThe grader imports your module as an unprivileged user and calls `audit()` once per claim, each in a new process with a hard limit of 10 seconds of CPU, no network, and the two tables at their usual paths. Nothing one call writes is left for the next. It runs all 66 claims: the 56 hidden ones and the ten examples. The reward is 1 only when every claim has `pay_employee` and `recover` equal to the cent and exactly the right set of flags, and your files in `/app/reports/` agree with the expected audits of the ten examples. Otherwise it is 0.\n\nYou have 28800 seconds to complete this task.\n", "public_examples": 10, "hidden_cases": 56, "hidden_unit": "claims", "grading": "Pay and recover to the cent, exact flag set", "answer_key": "included", "answer_key_sha256": null, "model": "claude-opus-5-5", "blind_attempts": 5, "blind_solved": 0, "closest_attempt": "14 of 56 hidden claims wrong"}
