UniPat AI
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- UniPat AI's PaperBenchX benchmark: GPT-6 Astra fully reproduces only 13.98% of papers
UniPat AI released PaperBenchX, which it calls the first multidisciplinary end-to-end benchmark for reproducing published paper results: 93 reproduction tasks built from 93 papers span 12 research directions and 10 domain-native scientific environments (including Ansys HFSS, Ansys Lumerical, Meep, PySCF and ABACUS), with 3,168 expert-verified scoring items. Across the 93 tasks, the strongest tested configuration, GPT-6 Astra, achieved only a 13.98% full-reproduction rate; averaged over tested configurations, modeling and execution scored about 62% each while validation reached just 42.8%. Per-task budgets run 4–24 hours (median 7 hours), and 12 tasks are open-sourced with 81 held out.
量子位(原生 RSS) · 🔥 11
Experience and discussion from the community
Share my UniPat AI experienceAsk about UniPat AI
Nobody has shared their experience with UniPat AI yet.