Andrej Karpathy argued that simple LLM coding tests are obsolete. Given the first paragraph of The Lord of the Rings and a 1M-token budget, Opus 5 ran autonomously for two hours and produced 5,500 lines of procedural Three.js 3D rendering code.
Key Takeaways
- ✓Frontier coding models possess the stamina to autonomously generate 5,500+ lines of custom 3D code over hours.
- ✓Simple coding benchmarks are obsolete for measuring modern long-horizon agent performance.
- ✓Highlights the next frontier: native multimodal visual verification and interactive game testing.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.