- Joined
- Dec 19, 2005
- Messages
- 18,023
"The asterisk on Astra’s 98.6%
Astra was evaluated through the company’s Responses API harness, with two settings changed to better reflect how the model performs in real-world use. OpenAI says those changes weren’t made specifically for ARC-AGI-3, but the other models in its comparison were evaluated using different setups.ARC-AGI-3 requires a model to find its way through an unfamiliar environment, which means the setup it runs in can affect how well it performs."
https://thenewstack.io/astra-arc-agi-benchmark/