Godot
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- SWE-Game benchmark tests whether coding agents can build games
An arXiv paper introduces SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D, with five task types: brief-to-game, implementation from a design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Across six evaluated models, Opus5 scored highest overall in all five task types, yet best overall scores on the three construction tasks stayed below 60 out of 100, with Brief-to-Game at 50.38; the authors flag requirement omissions and gameplay logic errors as the predominant problems. The paper reports executable checks reaching 92.59% balanced accuracy on human-labeled behaviors from 100 agent-built games versus 78.41% for a video-based VLM judge, and rubric visual scores with a 0.829 Spearman correlation to human ratings.
Hugging Face · Papers · 🔥 0
Experience and discussion from the community
Share my Godot experienceAsk about Godot
Nobody has shared their experience with Godot yet.