Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

UniPat AI

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 1

  1. UniPat AI's PaperBenchX benchmark: GPT-6 Astra fully reproduces only 13.98% of papers

    UniPat AI released PaperBenchX, which it calls the first multidisciplinary end-to-end benchmark for reproducing published paper results: 93 reproduction tasks built from 93 papers span 12 research directions and 10 domain-native scientific environments (including Ansys HFSS, Ansys Lumerical, Meep, PySCF and ABACUS), with 3,168 expert-verified scoring items. Across the 93 tasks, the strongest tested configuration, GPT-6 Astra, achieved only a 13.98% full-reproduction rate; averaged over tested configurations, modeling and execution scored about 62% each while validation reached just 42.8%. Per-task budgets run 4–24 hours (median 7 hours), and 12 tasks are open-sourced with 81 held out.

    量子位(原生 RSS) · 🔥 11

Experience and discussion from the community

Share my UniPat AI experienceAsk about UniPat AI

Nobody has shared their experience with UniPat AI yet.