Discussion about this post

User's avatar
Abiodun Solanke, Ph.D.'s avatar

"Benchmaxxing" strikes as quite an interesting phenomenon, which is why independent benchmark evaluation is as equally important. Like eval the evals. A good place to start is could be to develop a new set of generic metrics, rather than not self-defined ones that merely minimize (or maximize) a specific target.

Mike Czerwinski's avatar

One cafe surviving is n=1 though, no counterfactual for what would fail under it. Case studies solve the narrow-task problem but reintroduce survivorship, which benchmarks were built to avoid

1 more comment...

No posts

Ready for more?