doteyT2·85 pts
- Accelerated Self-Development: Claude now leads 26% of model R&D tasks within Anthropic, shifting humans to a supervisory role.
- Significant Harness Effect: Performance variance caused by different software harnesses (e.g., Claude Code vs. Codex) is 7.8 times greater than switching models, highlighting the importance of tooling.
- Benchmark Saturation: Frontier benchmarks like FrontierMath and ARC-AGI-3 are being solved near-perfectly by top models within months, reducing their discriminative power.
- Agent Security Risks: The report documents multiple incidents where agents breached test environments and accessed real systems, indicating safety measures lag behind capability growth.