OpenAI releases 719 AI math solutions, drawing controversy
DoNews reports that OpenAI has published 719 AI math solutions, a move that has sparked controversy. The report does not specify which model or project produced the solutions, nor the exact nature of the dispute.
ByteDance Seed ties DeepSeek-V4 long-context swings to 'phase sensitivity'
In late September 2026, ByteDance's Seed team published an arXiv paper arguing that "phase sensitivity" introduced by chunked KV-cache compression is the main cause of long-context retrieval swings in the DeepSeek-V4 family. The work covers open models including DeepSeek-V4-Flash, V4-Pro and V4.1-Flash, and reports that token phase coordinates created by the compression window can make retrieval accuracy for the same information differ by as much as 40 percentage points depending on phase, producing periodic degradation that averaged benchmark scores hide. The team calls for better cache-compression designs to improve long-context stability.
PaperBenchX: best AI agent fully reproduces only 14% of papers
UniPat AI's PaperBenchX benchmark gives 10 frontier agent configurations a published paper and asks them to re-run the key experiments in real scientific software, across 93 tasks covering 12 research directions. According to the team's published results, 70 of the 93 tasks were not fully reproduced by any configuration, and the top performer, GPT-6 Astra, fully cleared only 14% (13 tasks), while Chinese models clustered around 7% — roughly half that. Scoring separates modeling, execution and scientific verification across 3,168 scorable criteria, with a 4–24 hour per-task budget and outputs wiped and re-run in an isolated environment before grading.
MIT Tech Review roundtable: AI-designed viruses with Samuel King
MIT Technology Review will host a subscriber-only online roundtable on October 16, 2026, with AI reporter James O'Donnell interviewing Samuel King, a Stanford bioengineering PhD candidate and Innovators Under 35 honoree. The listing recalls that in 2025 King used a generative AI model to propose genetic blueprints for microscopic viruses, noting this is not yet AI-generated life but "that could be next." The conversation will cover his work and "new ways of seeing biology."
ByteDance Seed ties long-context swings to chunked KV cache 'phase sensitivity'
ByteDance's Seed team says in a new research paper that large language models' performance swings on very long inputs stem mainly from "phase sensitivity" introduced by chunked KV cache compression. To cut memory use in long-context inference, the method compresses consecutive token windows into fewer entries at a fixed stride, creating a new positional coordinate: a token's "phase" relative to the compression window boundary. Experiments show the same information becomes much harder or easier to retrieve at different phases, with long-context retrieval accuracy gaps of up to 40 percentage points in some large open-source models.