VeriFine: Scaling verification for self-improvement in embodied reasoning
A new arXiv paper introduces VeriFine, an agent harness framework that scales verification by co-evolving the policy, training curriculum and judge. Its Policy Improvement Loop uses a rubric judge to diagnose recurring failures and build an adaptive curriculum; when verification plateaus, a Judge Improvement Loop selectively queries human guidance and refines the judge through coactive calibration. The authors report continuous self-improvement in both policy and judge capability on driving and robot navigation tasks under reinforcement and supervised fine-tuning.