GPT-5.6 Sol
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 5
- OpenAI and Ironclad train agents for complex contracting workflows
OpenAI says it partnered with Ironclad to train and evaluate computer-using agents on complex contracting workflows, including agreement configuration, approvals, and reusable legal terms. In OpenAI’s research evaluation across 11 tasks, GPT-6 Astra averaged 55.0% versus 41.6% for GPT-5.6 Sol, while estimated time per attempt fell from 37.0 minutes to 19.2 minutes; these results come from OpenAI’s internal evaluation.
OpenAI · 🔥 5 - HERMES: modular executable Dev-Primitives for software engineering agents
An arXiv paper introduces Dev-Primitives, an abstraction that pairs each repository artifact — source files, configs, tests, dependencies — with a resident LLM, giving it an agent-native interface for natural-language reasoning, inter-component communication and localized self-modification. Built on top of it, the HERMES harness-engineering framework activates these primitives at repository scale via dependency-aware dynamic activation and a bug-diagnosis mechanism that maps execution evidence back to the components needing revision. The authors report HERMES beats matched baseline harnesses by 12.4% on average across four software engineering benchmarks, stays within 4.5% of a homogeneous GPT-5.6 Sol configuration even with Qwen3-8B Dev-Primitives, and cuts inference cost by 26.2% on Terminal-Bench 4.0.
Hugging Face · Papers · 🔥 3 - Agent swarms may be AI's next scaling law, but gains look limited
Understanding AI argues that multi-agent "swarms" are emerging as a new scaling law for frontier AI: OpenAI researcher Noam Brown says the company's models are now sometimes trained in environments alongside other agents, given tools to message each other, and encouraged to achieve objectives together. The piece points to July's Hugging Face incident, where hundreds of OpenAI agents self-organized into teams, and OpenAI's September claim that 10,000 agents solved a famous math problem in a few days. But it notes diminishing returns — Anthropic's Claude Opus 5.5 system card found the biggest multi-agent gain came from scaling one to 10 agents, with the main benefit being speed rather than a better answer, and Brown attributed under 10% of the math breakthrough's credit to multi-agent coordination.
Understanding AI · 🔥 0 - VISTA boosts multimodal agents with visual memory and active recall
A team led by Kaiming He introduced VISTA, a framework that gives multimodal agents direct visual input, lossless visual memory and tools to inspect past frames. According to the reported paper results, Claude Opus 5 improved from 40.68 to 100 on 25 public ARC-AGI-3 games, while GPT-5.6 Sol improved from 13.33 to 99 without changing the underlying models.
MIT科技评论中文 · 🔥 0 - VISTA Gives Frontier Models Visual Memory for ARC-AGI-3
A paper from Kaiming He’s team presents VISTA, a harness that gives vision-language models direct access to game images, persistent frame-by-frame visual memory, and model-controlled inspection tools. The source reports that Claude Opus 5 completed all 25 public ARC-AGI-3 games with a perfect score, while GPT-5.6 Sol achieved 99, attributing the gains to improved visual access and memory rather than additional model training.
创业邦 科技 · 🔥 0
Experience and discussion from the community
Share my GPT-5.6 Sol experienceAsk about GPT-5.6 Sol
Nobody has shared their experience with GPT-5.6 Sol yet.