Models that are good at math benchmarks tend to be good at coding and reasoning benchmarks too, pointing to a common factor driving AI capabilities.
Please forgive the question, but is than a news update compared to https://epochai.substack.com/p/benchmark-scores-general-capability? (Have you added new data since then?)
Please forgive the question, but is than a news update compared to https://epochai.substack.com/p/benchmark-scores-general-capability? (Have you added new data since then?)