AC-Small Improves on APEX-Agents Dev Set
AC-Small improved significantly on held-out benchmarks after post-training on the APEX-Agents dev set, with +5.7pp on APEX, +8.0pp on Toolathalon, and +7.7pp on GDPval. This demonstrates the potential for large language models to generalize well beyond their initial training data. The improvement is notable and suggests that post-training on specific datasets can enhance model performance.