The Epoch Brief - July 8, 2026
A board game AI can't master, a cyber-disclosure spike after Claude Mythos, GPT-4's record run atop the ECI, and what's missing from AI futurism discourse
Welcome back to the Epoch Brief! In this edition:
The launch of EBR-bench, our benchmark that uses a complex board game to test AI’s ability to learn from experience.
Two new Data Insights: signs AI is finding software vulnerabilities at scale, and GPT-4’s record run atop the Epoch Capabilities Index.
A new Gradient Update on the missing half of AI futurism debates.
We’re growing and have open positions across research, engineering, and operations.
Model Evaluations
EBR-bench: Our latest benchmark suggests AI struggles to learn from experience
Can AI systems improve at challenging tasks by attempting them over and over and learning from their mistakes? It’s one of the biggest open questions in AI capabilities right now, with large economic and safety implications. Our latest benchmark, EBR-bench, tests for this ability by having models repeatedly play a complex board game called Earthborne Rangers.
So far, we see little evidence of AI learning from experience. Going forward, EBR-bench is a tool for detecting if and when that changes. View the full results and analysis.
Expanding the scope and quality of our benchmarking work
EBR-bench is one of several recent additions to Epoch benchmarking. Two weeks ago, we launched our MirrorCode benchmark, which we co-developed with METR, to understand the furthest limits of AI coding capabilities. We let AI code autonomously for weeks at a time, asking it to rebuild real-world programs from scratch — some comprising tens of thousands of lines of code. So far, the best model scores 56%.
We’ve also expanded the number of AI benchmarks we’re tracking, adding nine last month (covering agentic work, cybersecurity, algorithm engineering, forecasting, and research-level physics) and now another 13. Seven of the latest additions also feed into the Epoch Capabilities Index, our aggregate measure of model capability.
Data Insights
This week we published two new Data Insights, our digestible snapshots of complex trends in AI.
Cyber vulnerability disclosures spiked around Claude Mythos Preview
AI appears to be finding software vulnerabilities at scale. Researcher Luke Emberson finds that 21 notable organizations disclosed around 1,500 high- and critical-severity CVEs in June, more than 3.5× the monthly record before Mythos’ release.
The spike follows Anthropic’s April announcement that Claude Mythos Preview could autonomously discover software vulnerabilities, and that partners in the company’s Project Glasswing program had already been using it to find and fix bugs ahead of the model’s public release.
GPT-4 led in ECI far longer than any other model
OpenAI’s GPT-4 topped the Epoch Capabilities Index (ECI) for roughly a year after its release in March 2023. No model since has led for as long. The second-longest lead, by OpenAI’s o1, lasted a little over three months, less than a third of GPT-4’s.
Gradient Update: The missing half of AI futurism debates
AI discourse is rife with big predictions about how automated AI research will rapidly transform the future with nanotech, Dyson swarms, and near-light-speed spacecraft, among other advanced technologies.
In our latest Gradient Update, senior researcher JS Denain and researcher Anson Ho argue these claims often rest solely on AI models advancing in capability, and lack serious analysis of just how hard futuristic tech is to build. They propose using exploratory engineering, with explicit assumptions about AI capabilities, to better ground predictions.
Gradient Updates represent the views of the authors, and do not necessarily reflect the views of Epoch AI as a whole.
Other Updates
Careers
Epoch is growing, and we’re looking for talented folks to help scale our impact.
Finance Specialist / Manager to run our accounting and finance operations.
Researcher (Benchmark Reviews) to develop and publish critiques and reviews of AI benchmarks.
Researcher (Evaluations) to evaluate frontier models on hard-to-grade tasks.
Software Engineer, Benchmarking to build and maintain our benchmarking infrastructure.
Talent Scout to help us find and recruit exceptional people.
Senior Product Designer to lead UI/UX and data visualization.
Data Scientist (Contract) to assist with our research through literature review and data analysis.
Applications are rolling, so apply soon!





