AC-Small improved significantly on held-out benchmarks
AC-Small improved significantly on held-out benchmarks after post-training on the APEX-Agents dev set, with +5.7pp on APEX, +8.0pp on Toolathalon, and +7.7pp on GDPval. This improvement showcases the potential of post-training on diverse datasets. The results have significant implications for AI enthusiasts and professionals, as they demonstrate the effectiveness of fine-tuning models on specific tasks. The improvement in benchmarks also highlights the importance of continuous learning and adaptation in AI models.