SkillForge co-evolves agent skills and policy via a fitness-driven skill lifecycle
A NeurIPS 2026-accepted arXiv paper introduces SkillForge, an agentic RL method that evolves a skill library through a fitness-driven lifecycle of trial, active, stable and retired states, so skills and the model co-evolve during training. A pre-RL phase uses the base model's own rollouts to pre-retire low-fitness skills, seeding supervised fine-tuning; RL then continues with selective retirement, stabilization and LLM-guided mutation each iteration. The authors report the highest aggregate success rate across several interactive agent benchmarks, up to 7.8% relative improvement over the strongest baseline while keeping the library compact, and release SkillFurnace, a 5k+ record annotated dataset of retirement-filtered SFT trajectories, evolved skill libraries and retirement events.