Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

Opus 5

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 2

  1. Four frontier models each got $100 to build a PDF editor — nearly all shipped bugs

    Researcher nielstron gave Gemini 3.8 Flash, GPT Astra 6, Opus 5 and Fable 5 $100 each (about $400 total) and asked them to autonomously build an open-source PDF editor with an excellent user experience. According to OSChina's report, bugs could be found within a few clicks in nearly every result. The author argues coding agents cannot interact with software the way humans do, which is a key reason their flaws surface so quickly.

    开源中国 · 🔥 14
  2. PaperBenchX: best AI agent fully reproduces only 14% of papers

    UniPat AI's PaperBenchX benchmark gives 10 frontier agent configurations a published paper and asks them to re-run the key experiments in real scientific software, across 93 tasks covering 12 research directions. According to the team's published results, 70 of the 93 tasks were not fully reproduced by any configuration, and the top performer, GPT-6 Astra, fully cleared only 14% (13 tasks), while Chinese models clustered around 7% — roughly half that. Scoring separates modeling, execution and scientific verification across 3,168 scorable criteria, with a 4–24 hour per-task budget and outputs wiped and re-run in an isolated environment before grading.

    品玩 实时要闻 · 🔥 13

Experience and discussion from the community

Share my Opus 5 experienceAsk about Opus 5

Nobody has shared their experience with Opus 5 yet.