Four frontier models each got $100 to build a PDF editor — nearly all shipped bugs
Researcher nielstron gave Gemini 3.8 Flash, GPT Astra 6, Opus 5 and Fable 5 $100 each (about $400 total) and asked them to autonomously build an open-source PDF editor with an excellent user experience. According to OSChina's report, bugs could be found within a few clicks in nearly every result. The author argues coding agents cannot interact with software the way humans do, which is a key reason their flaws surface so quickly.
Why it matters: Small comparative experiments like this suggest coding agents' weak point may lie as much in interacting with and verifying software as in generating code.
一个研究 LLM for Code 的研究者,拿 400 美元做了个实验:给 Gemini 3.8 Flash、GPT Astra 6、Opus 5、Fable 5 四个前沿模型各 $100 预算,让它们自主开发一款「用户体验极好」的开源 PDF 编辑器——结果是「点几下就能在几乎每个里找到 bug」。作者 nielstron 据此提出:coding agent 不能像人一样和软件交互,所以它们...
This is the outlet's own summary. Read the full story on the original site.
Read the original →How we got here
- Innolight's triple bind as optical modules become a policy flashpoint虎嗅 AI · Google
- OpenAI APAC public policy lead said to depart after six months36氪快讯 · OpenAI
- Hands-on with GPT-6's Intelligent UI: interactive cards, mini-games, mixed results36氪 人工智能 · OpenAI
- Trump announces 'AI Force' and an AI czar虎嗅 AI · Anthropic
- Terence Tao warns OpenAI's bulk AI math proofs cause 'proof indigestion'AIbase AI新闻 · OpenAI
- OpenAI Asia-Pacific policy head to leave after six monthsBloomberg Technology · OpenAI