Same model, same tasks
All seven agents ran Kimi K3, confirmed by the evaluator. Twelve fixed tasks, one frozen repo snapshot.
Seven coding agents ran the same frozen codebase under Kimi K3. All 84 reports and every code diff are published with SHA-256 hashes — re-check anything yourself.
All seven agents ran Kimi K3, confirmed by the evaluator. Twelve fixed tasks, one frozen repo snapshot.
One task, one fresh session, one output folder. Only T12’s two rounds share a session, by design.
Reports, code diffs and scoring bases ship with hashes. The manual stage scores come with their math and their limits.
Agent versions and inference settings weren’t fully logged. MiniMax ran all 12 tasks in one batch session; its run is flagged and stays out of strict comparison until a rerun. Scope & corrections →
The exact prompts each agent received, plus the public scoring basis for each task. Prompts stay in their original Chinese so the hashes hold.
The full text loads here.
Kept with their ranks at the time
These numbers are manual judgements made after the runs, not output of the frozen scoring sheet. No per-task worksheets, no blind review behind them — treat them as a snapshot, not a verdict.
| Rank then | AGENT | Function 35% | Verify 20% | Quality 15% | Protocol 15% | Evidence 15% | Stage score | Status | |
|---|---|---|---|---|---|---|---|---|---|
| Loading scores… | |||||||||
Each dimension 0–5, then 35×function/5 + 20×verification/5 + 15×quality/5 + 15×protocol/5 + 15×evidence/5, to one decimal.
The frozen rule set is acceptance 70%, quality 15%, verification 10%, safety 5% — and no composite score while safety evidence is missing. The five-way split above was a later, provisional view.
\d. The task text is ambiguous there; the next version will pin it to [0-9] or Unicode explicitly.Numbers preserved as recorded at the time. Score versions & evidence notes →
The site is a display layer. The evidence chain: frozen commit → sample checksums → report and patch hashes → whole-site SHA256SUMS. All of it works offline.
Download the published directory, then:
shasum -a 256 -c SHA256SUMS
ed2390f2…7e5ed; every agent got the same copy.The verbatim instructions every agent received:
【扫描所有文件,并开始执行任务】 + BENCH-RUN <TASK_ID> <RUN_ID> marker.BENCHMARK_RESULT.md with the commands it actually ran.Anchors for this batch:
ledger-agent-benchmark-v1, archived 2026-09-24; one R01 run per tool per task.01a0ce53-e7f4-7352-afca-2733432c3a39; scoring reply 01a0d2ad-b637-7f51-8aa7-68a16de63205. These locate the original conversation; they don’t prove it.shasum -a 256 -c SHA256SUMS re-checks every file you see.Gaps we know about:
Corrections, disputes, or a rerun request: itderry@qq.com
Desensitized copies with original file hashes preserved; code diffs expand inline. A PASS in a report is the agent’s own claim, not an independent acceptance.