We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choic
We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choices about API settings, harness design, and prompting. If you’re an API developer trying to maximize performance, we recommend using the