Anthropic adds build-eval and hillclimb tuning commands to Claude Code
On September 28 Anthropic published a tuning guide introducing two Claude Code entry points, /claude-api build-eval and /claude-api hillclimb (Claude Code v2.1.259 or later). The first turns "what counts as a good answer" into repeatable test cases and scoring rules; the second lets Claude propose one change per round — to prompts, skill files, tool descriptions or model config — keeping it only if the evaluation improves and reverting otherwise, with a hidden held-out set used to check for overfitting. Anthropic says an internal customer-support evaluation improved decision accuracy from 78.6% to 90.5% on 14 held-out tickets while cutting model call cost to roughly one fifth.