Back to Timeline

Event Summary

Epoch AI and METR published MirrorCode results on long-horizon coding tasks that require models to reimplement programs without source-code or web access. One Claude Opus 4.7 run completed a 16,000-line target in 14 hours at a reported $251 inference cost.

Context & Narrative

MirrorCode uses end-to-end tests over 25 target programs. The results show that models can complete some substantial reimplementation tasks under the benchmark’s conditions. Epoch also reports a 56% headline score for Opus 4.7 and warns that target programs may have appeared in training data.

Key Findings

  • Fact Grade B

    MirrorCode evaluates end-to-end program reimplementation without source-code or web access across 25 target programs.

    Sources [1][2]
  • Impact Grade B

    Epoch reported that one Claude Opus 4.7 run reimplemented a 16,000-line target in 14 hours at a $251 inference cost.

    Sources [1]
  • Limitation Grade B

    Epoch reports that pretraining exposure to open-source targets may inflate benchmark performance.

    Sources [1]

Impact Assessment

  • Capability Leap +2 · Medium-term

    MirrorCode provides evidence that current models can complete some long-horizon program-reimplementation tasks under controlled conditions.

    Affected Groups: software engineers, AI-agent researchers, developer-tool builders

Consensus & Sources

Significance L1
Category Capability Breakthrough
Consensus Emerging Consensus
Impact Index 4/10