PUBLIC BENCHMARK · v1.0

One model. Twelve tasks.
Seven agents, scored in the open.

Seven coding agents ran the same frozen codebase under Kimi K3. All 84 reports and every code diff are published with SHA-256 hashes — re-check anything yourself.

One model: Kimi K3Evaluator-confirmed for all seven agents
BENCHMARK SNAPSHOT 37.38N / 122.05W
AGENTK3FIELD TEST
07AGENTS
12TASKS
84REPORTS
LEDGER AGENT BENCHMARKv1.0 · KIMI K3
01 / METHOD

How the runs were set up

Scoring rules (frozen before the runs)
01

Same model, same tasks

All seven agents ran Kimi K3, confirmed by the evaluator. Twelve fixed tasks, one frozen repo snapshot.

02

Isolated sessions

One task, one fresh session, one output folder. Only T12’s two rounds share a session, by design.

03

Evidence stays public

Reports, code diffs and scoring bases ship with hashes. The manual stage scores come with their math and their limits.

Agent versions and inference settings weren’t fully logged. MiniMax ran all 12 tasks in one batch session; its run is flagged and stays out of strict comparison until a rerun. Scope & corrections →

02 / TASK SET

Read the tasks yourself

Samples v1.0 · 12 tasks

The exact prompts each agent received, plus the public scoring basis for each task. Prompts stay in their original Chinese so the hashes hold.

Loading tasks…
PUBLIC SAMPLE · V1.0

Pick a task on the left

The full text loads here.

Frozen prompts; scoring bases archived separatelySample notes ↗
03 / SCORE ARCHIVE

Stage scores, as recorded

Manual scores · not official

Kept with their ranks at the time

Read this first.

These numbers are manual judgements made after the runs, not output of the frozen scoring sheet. No per-task worksheets, no blind review behind them — treat them as a snapshot, not a verdict.

R01 · seven-agent historical scoresSub-scores 0–5 · weighted total at right
Rank thenAGENTFunction
35%
Verify
20%
Quality
15%
Protocol
15%
Evidence
15%
Stage scoreStatus
Loading scores…
MiniMax scored 84.0 at rank 7 in that round; the batch-session deviation keeps it out of strict comparison.
35FunctionTask behaviour + regression
20VerificationChecks actually run and rerun
15Code qualityReadability, scope, compatibility
15ProtocolIsolation and constraint compliance
15EvidenceCommands, output, sourcing
CALCULATION

The math used then

Each dimension 0–5, then 35×function/5 + 20×verification/5 + 15×quality/5 + 15×protocol/5 + 15×evidence/5, to one decimal.

RULE CHANGE

Not the frozen rule

The frozen rule set is acceptance 70%, quality 15%, verification 10%, safety 5% — and no composite score while safety evidence is missing. The five-way split above was a later, provisional view.

Honest limits of these numbers
  • One run per task. No repeats, so no stability claims.
  • Safety runs lack full operation logs; MiniMax’s batch session is flagged; other tools’ session isolation isn’t independently audited yet.
  • Kimi K3 is evaluator-confirmed, but tool versions, exact model IDs and inference settings weren’t archived.
  • Qoder T03 lost points under Python’s Unicode-default \d. The task text is ambiguous there; the next version will pin it to [0-9] or Unicode explicitly.

Numbers preserved as recorded at the time. Score versions & evidence notes →

04 / VERIFICATION

Don’t trust this page — check it

SHA-256 · OPEN PROTOCOL

The site is a display layer. The evidence chain: frozen commit → sample checksums → report and patch hashes → whole-site SHA256SUMS. All of it works offline.

CHAIN

Integrity chain

Download the published directory, then:

shasum -a 256 -c SHA256SUMS
  • Baseline frozen at commit ed2390f27e5ed; every agent got the same copy.
  • 84 reports, each with source and published-copy SHA-256 (report index).
  • Every code diff against the frozen commit ships with its hash (patch index).
PROTOCOL

Run protocol

The verbatim instructions every agent received:

  • Trigger phrase 【扫描所有文件,并开始执行任务】 + BENCH-RUN <TASK_ID> <RUN_ID> marker.
  • One fresh session per task; only T12’s two rounds share its own session.
  • Each run writes BENCHMARK_RESULT.md with the commands it actually ran.
Verbatim fixed prompt (规程/统一执行提示.md) ↗
PROVENANCE

Provenance

Anchors for this batch:

  • Batch ledger-agent-benchmark-v1, archived 2026-09-24; one R01 run per tool per task.
  • Session locator 01a0ce53-e7f4-7352-afca-2733432c3a39; scoring reply 01a0d2ad-b637-7f51-8aa7-68a16de63205. These locate the original conversation; they don’t prove it.
  • The site ships its own SHA256SUMS; shasum -a 256 -c SHA256SUMS re-checks every file you see.
LIMITS

What this batch can’t tell you

Gaps we know about:

  • Stability — one run per task, no variance data.
  • Official ranking — blind scoring hasn’t happened yet.
  • Cost — token usage and spend weren’t logged in R01. Same model ≠ same cost: context size and retries differ per tool. The next batch records them.
  • Prompts and reports stay in their original Chinese so hashes stay verifiable; UI translation never touches evidence files.
CONTACT

Corrections, disputes, or a rerun request: itderry@qq.com

05 / RUN RECORDS

84 reports, each with a hash

84 R01 reports

Desensitized copies with original file hashes preserved; code diffs expand inline. A PASS in a report is the agent’s own claim, not an independent acceptance.

84 reports
Loading report index…