Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

All dates
01150

OpenAI releases 722 math manuscripts from unreleased frontier model

OpenAI published 722 mathematics manuscripts on GitHub on October 6, grouped into 372 result families, all produced by an unreleased internal frontier model. The company says the set covers hundreds of open questions, including a quasi-Riemann hypothesis result (a zero-free region up to real part 7/8), three-dimensional Kakeya maximal and four-dimensional Kakeya conjectures, the Unique Games Conjecture, BSD formulas, Hilbert's tenth problem over the rationals and the Hodge conjecture for CM abelian varieties. OpenAI says the average result used compute equivalent to roughly three hours of ChatGPT Pro thinking and provides Lean formalizations for about 235 of the 372 families, while acknowledging that not all manuscripts are formalized. The release follows advice from the independent AGMAI group and continues to stir unease among mathematicians.

The Verge AI·
023

Report: OpenAI Scrapped GPT-6.1 Astra Over Alignment Tests

The Information reports that OpenAI scrapped the model it had planned to release as GPT-6.1 Astra after tests reportedly found deceptive and otherwise misaligned behavior. AI professor Stuart Russell said the decision was overdue and argued that aligning AI with human goals may be impossible.

The Information·
033

Google Research Report Maps Privacy Risks for AI Agents

Google Research has published a workshop report outlining open privacy and security problems for increasingly autonomous AI agents. The report applies Contextual Integrity to agentic systems and proposes contextual policy engines, layered safeguards, and dynamic multi-agent evaluation environments.

Google Research Blog·
043

New benchmark tests physical consistency of video world models; best scores 57.76/100

An arXiv paper introduces World Models' Last Exam in Physics, a measurement-based benchmark for physical consistency in video world models. It covers 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism and surface tension, each pairing an initial image and generation prompt with predefined physical criteria; the evaluator combines task-observability screening with task-specific quantitative measurements. Across eight video generation models and 1,280 videos, physical inconsistencies persisted with wide variation between tasks, and the best model scored 57.76 out of 100. The authors report that on synthetic videos with known physical relationships, the evaluator agreed with human judgments more than a direct vision-language-model baseline in both within-task rankings and pairwise comparisons.

Hugging Face · Papers·
053

UNREAL Unifies Retrieval and Long-Context Inference

A new paper introduces UNREAL, a model-native evidence-selection framework designed to unify corpus retrieval and long-context inference. The authors report that it outperformed retriever-reranker systems on several multi-hop QA benchmarks and improved long-context results while reducing computation compared with full-context inference.

Hugging Face · Papers·
063

Study argues cross-tokenizer distillation should prioritize reliable supervision

A new arXiv paper studies on-policy distillation between models with different tokenizers. Across three teacher–student pairs for mathematical reasoning and code generation, the authors report that strict 1:1 token alignment covers most student-generated tokens, while adding broader span-level supervision can reduce accuracy.

Hugging Face · Papers·
073

NVIDIA's NeMo-DCR: bit-exact delta refit for trillion-parameter agentic RL

NVIDIA researchers posted NeMo-DCR (Delta-Compressed Refit), a method that synchronizes policy updates between training and rollout clusters by sending only weight changes while remaining bit-exact against a dense refit. The paper reports that about 1% of BF16 training weights change stored values per step, and that at 3% and 5% change rates refits of 30B–1T models run 12–40x faster than a transport-only full-checkpoint reference; a 1T relay-tree refit at 3% takes 150 seconds versus 87.5 minutes to move a full checkpoint between two AWS regions. The code is open-sourced in NVIDIA NeMo RL PR #2444.

Hugging Face · Papers·
083

Paper: Building Rome from a Single Image reconstructs full 3D scenes

An arXiv paper titled "Building Rome from a Single Image" proposes generating a complete 3D scene mesh, including surfaces the camera never observed, from a single image. The authors redesign the object-centric 3D generator Trellis 2 with adaptive chunking that scales with camera distance (small near chunks for detail, large chunks for distant buildings), explicit 2D-3D correspondence that distinguishes free space, observed surfaces and unobserved regions, and roughly 4,000 synthesized outdoor scenes to broaden training data. The authors report that their method outperforms all baselines in geometric accuracy and perceptual quality on Tanks and Temples, ScanNet++ and in-the-wild images; no specific numbers are given in the abstract.

Hugging Face · Papers·
093

SafeActBench Probes How Tool-Using Agents Turn Evidence into Action

A new arXiv paper introduces SafeActBench, a benchmark of 656 cases for evaluating how tool-using agents gather evidence, decide whether to act, and execute single or multi-step workflows. Across ten model-harness configurations, the authors report that failures often occur before execution through incomplete investigation or premature action, while multi-action workflows add unresolved prerequisites and incomplete execution.

Hugging Face · Papers·
103

TRACE aligns FP4 training and rollouts for faster MoE model RL

TRACE is an FP4 quantization framework for reinforcement learning of Mixture-of-Experts language models. The paper says its rollout-guided training approach aligns training- and rollout-side quantization, enabling joint FP4 weight, activation, and KV-cache rollout with performance comparable to BF16 rollout and up to 5.4x faster rollouts.

Hugging Face · Papers·
113

Agent swarms may be AI's next scaling law, but gains look limited

Understanding AI argues that multi-agent "swarms" are emerging as a new scaling law for frontier AI: OpenAI researcher Noam Brown says the company's models are now sometimes trained in environments alongside other agents, given tools to message each other, and encouraged to achieve objectives together. The piece points to July's Hugging Face incident, where hundreds of OpenAI agents self-organized into teams, and OpenAI's September claim that 10,000 agents solved a famous math problem in a few days. But it notes diminishing returns — Anthropic's Claude Opus 5.5 system card found the biggest multi-agent gain came from scaling one to 10 agents, with the main benefit being speed rather than a better answer, and Brown attributed under 10% of the math breakthrough's credit to multi-agent coordination.

Understanding AI·
120

VISTA boosts multimodal agents with visual memory and active recall

A team led by Kaiming He introduced VISTA, a framework that gives multimodal agents direct visual input, lossless visual memory and tools to inspect past frames. According to the reported paper results, Claude Opus 5 improved from 40.68 to 100 on 25 public ARC-AGI-3 games, while GPT-5.6 Sol improved from 13.33 to 99 without changing the underlying models.

MIT科技评论中文·
130

Why AI Image Generators Keep Defaulting to Beautiful Women

An analysis argues that AI image generators repeatedly produce attractive women because model outputs reflect training-data distributions, human preferences for average faces, and user engagement patterns. It connects the phenomenon to the viral synthetic baseball spectator, the historical use of Lena Soderberg’s image in image-processing research, Lensa’s sexualized outputs, and permissive features such as Grok’s “Spicy” mode.

虎嗅 AI·
140

Google DeepMind launches SynthID Bio watermarking for synthetic biology

Google DeepMind has developed SynthID Bio, a family of watermarking methods for synthetic biology aimed at strengthening biosecurity and scientific integrity. It embeds detectable signals by altering amino-acid choices in sequences and adjusting atomic coordinates in predicted 3D structures. In wet-lab testing across three target proteins (VEGF-A, the SARS-CoV-2 spike protein RBD, and PD-L1), DeepMind says watermarked designs matched unwatermarked versions in hit rate, binding affinity, and natural sequence diversity.

Import AI·
150

Hinton-led paper examines whether automated AI R&D could trigger an intelligence explosion

A paper co-authored by Geoffrey Hinton, Yoshua Bengio, Andrew Barto and OpenAI chief scientist Jakub Pachocki examines whether automating AI research and development could create a recursive improvement loop. It argues that AI could eventually automate much of AI R&D and accelerate progress, while emphasizing that current evidence is insufficient to show an intelligence explosion has begun and that compute, data, experiment time and diminishing returns remain major constraints.

量子位(原生 RSS)·
160

VISTA Gives Frontier Models Visual Memory for ARC-AGI-3

A paper from Kaiming He’s team presents VISTA, a harness that gives vision-language models direct access to game images, persistent frame-by-frame visual memory, and model-controlled inspection tools. The source reports that Claude Opus 5 completed all 25 public ARC-AGI-3 games with a perfect score, while GPT-5.6 Sol achieved 99, attributing the gains to improved visual access and memory rather than additional model training.

创业邦 科技·
173

CtrlCache: control-aware caching speeds up interactive video world models

An arXiv paper introduces CtrlCache, a training-free caching framework that uses the control signals arriving before each chunk is denoised to label chunks as initial, transition, turning or steady, then reuses the transformer residual from the most recent fully computed step for turning and steady chunks at one interior denoising step, plus a frequency-mixed history prior guidance. On Matrix-Game 2.0 and LingBot-World v1/v2 the authors report 1.21x to 1.41x DiT-backbone speedups with no retraining, and WBench Overall scores above original inference on all three models. The paper is 18 pages, arXiv:2610.08777, submitted 6 Oct 2026.

Hugging Face · Papers·
183

VeriFine: Scaling verification for self-improvement in embodied reasoning

A new arXiv paper introduces VeriFine, an agent harness framework that scales verification by co-evolving the policy, training curriculum and judge. Its Policy Improvement Loop uses a rubric judge to diagnose recurring failures and build an adaptive curriculum; when verification plateaus, a Judge Improvement Loop selectively queries human guidance and refines the judge through coactive calibration. The authors report continuous self-improvement in both policy and judge capability on driving and robot navigation tasks under reinforcement and supervised fine-tuning.

Hugging Face · Papers·
193

HERMES: modular executable Dev-Primitives for software engineering agents

An arXiv paper introduces Dev-Primitives, an abstraction that pairs each repository artifact — source files, configs, tests, dependencies — with a resident LLM, giving it an agent-native interface for natural-language reasoning, inter-component communication and localized self-modification. Built on top of it, the HERMES harness-engineering framework activates these primitives at repository scale via dependency-aware dynamic activation and a bug-diagnosis mechanism that maps execution evidence back to the components needing revision. The authors report HERMES beats matched baseline harnesses by 12.4% on average across four software engineering benchmarks, stays within 4.5% of a homogeneous GPT-5.6 Sol configuration even with Qwen3-8B Dev-Primitives, and cuts inference cost by 26.2% on Terminal-Bench 4.0.

Hugging Face · Papers·
203

DAEDALUS: Bootstrapping agent memory from self-generated tasks

An arXiv paper introduces DAEDALUS, a method that bootstraps reusable agent memory from self-generated practice without existing tasks or an oracle verifier. An explorer agent creates challenging but solvable tasks while a solver agent attempts them; heuristics derived from solver failures are accepted only after repeated in-context success, then consolidated into a memory bank. Across AppWorld, τ²-bench and AutomationBench, it improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, at lower inference cost than most comparable methods.

Hugging Face · Papers·