Event Summary
Epoch AI and METR published MirrorCode results on long-horizon coding tasks that require models to reimplement programs without source-code or web access. One Claude Opus 4.7 run completed a 16,000-line target in 14 hours at a reported $251 inference cost.
Context & Narrative
MirrorCode uses end-to-end tests over 25 target programs. The results show that models can complete some substantial reimplementation tasks under the benchmark’s conditions. Epoch also reports a 56% headline score for Opus 4.7 and warns that target programs may have appeared in training data.
Key Findings
-
Fact Grade B
MirrorCode evaluates end-to-end program reimplementation without source-code or web access across 25 target programs.
-
Impact Grade B
Epoch reported that one Claude Opus 4.7 run reimplemented a 16,000-line target in 14 hours at a $251 inference cost.
Sources [1] -
Limitation Grade B
Epoch reports that pretraining exposure to open-source targets may inflate benchmark performance.
Sources [1]
Impact Assessment
-
Capability Leap +2 · Medium-term
MirrorCode provides evidence that current models can complete some long-horizon program-reimplementation tasks under controlled conditions.
Affected Groups: software engineers, AI-agent researchers, developer-tool builders
Consensus & Sources
-
1
Reference Evidence Citation logged Live source
-
2
Reference Evidence Citation logged Live source
-
3
Reference Evidence Citation logged Live source