AC-Small Improves on APEX-Agents Dev Set
AC-Small improved significantly on held-out benchmarks after post-training on the APEX-Agents dev set, with +5.7pp on APEX, +8.0pp on Toolathalon, and +7.7pp on GDPval. This improvement showcases the potential of fine-tuning models on specific datasets to enhance their performance. The results have implications for AI model development and the pursuit of more accurate and robust models.