Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

AI News

All dates
019

Stanford's Fei-Fei Li team open-sources OpenWAM world-action framework

A Stanford team including Fei-Fei Li, Jiajun Wu and Ehsan Adeli published OpenWAM, an open framework for composable world-action models. It starts from Alibaba's open Wan2.2-5B video model, continues pretraining on roughly 3.34 million robot and human-interaction videos with a causal, future-blind setup, then joins a 5B video expert and a 2B action expert via a shared Mixture-of-Transformers architecture with four configurable video-action interaction orders (VTA, ATV, Joint, Decoupled). The team also built LIBERO-Long-CF, a counterfactual dataset of about 32,000 clips and 4.1 million control records, and trained transferable local inverse- and forward-dynamics components; they report a frozen local IDM reaching 84.0% average success on four new LIBERO-90 tasks versus 47.0% for a full-context IDM and 21.5% for demonstration-only local IDM, plus 92.1% and 91.9% average success for VTA and Joint on real dual-arm Franka FR3 tasks. Training, evaluation and composition code and pretrained video-model weights have been open-sourced; the piece also notes Black Forest Labs' FLUX 3 Action world-action model.

36氪 人工智能·
0214

One prompt hijacked every AgentCore agent in an AWS account, Zenity says

Security firm Zenity Labs published research on a vulnerability chain it calls "AgentCorruption": chat access to a single public Amazon Bedrock AgentCore agent was enough to trick it, in plain language, into querying the instance metadata service at 169.254.169.254 and sending its own temporary AWS credentials to an external server. Using AgentCore's overly broad default execution role, the researchers then listed and invoked every agent in the same account and region, downloaded their source code and container images, read private user-agent conversations and secrets from AWS Secrets Manager, and poisoned long-term memory to make the hijack persistent. Zenity says it reported the findings to AWS on 25 December 2025; AWS has since made IMDSv2 the default for new AgentCore deployments and narrowed the default role.

The Decoder·
0332

OpenAI withdraws three math papers over a sign error

In its math repository history, OpenAI said a sign error in “Algebraicity of Weil classes on split abelian eightfolds” invalidates a stabilization-trace cancellation argument and the construction used by two dependent papers, leading it to withdraw three manuscripts. Fourteen other manuscripts were revised with proof repairs, corrected statements and clearer hypotheses, and 13 more were updated to cite the revised editions. Six additional formalizations and five other additions bring the total share of top-line results formalized to 300/719, about 42%.

Hacker News · AI(100+ 分)·
047

OpenAI releases 722 math research papers with AI-generated solutions

OpenAI released 722 research papers on Tuesday containing AI-created solutions to problems across a range of math subjects, according to The Information. The report attributes the success to the verifiability of math solutions (much like code), the overlap between mathematics and AI research work, and abundant compute—one advising mathematician described it as "brute force and stamina," even though the AI still errs on simple arithmetic.

The Information·
0510

Anthropic: Claude Science helps produce first complete UV sky map

Astrophysicist Brice Ménard worked with Anthropic's Claude Science to create the first complete ultraviolet map of the sky. Complete sky maps exist from radio through gamma rays, but large regions had never been observed in UV; Ménard guided Claude to find and combine existing datasets and fill gaps with statistical inference. Anthropic says the work would have taken humans weeks but took a few days with Claude, while Ménard worked on other projects.

Anthropic (X)·
069

Weizmann's Brain-IT AI reconstructs viewed images from fMRI scans

Researchers at the Weizmann Institute of Science, led by professor Michal Irani, built Brain-IT, a model that uses fMRI scans to reconstruct images a person is viewing — and also works in reverse to predict brain activity for a given image. The team says it outperforms earlier methods on image content and details such as composition and color, and needs only about one hour of fMRI data from a new subject to match results other approaches reach with 40 hours of recordings. The work remains confined to lab settings; the researchers hope it could eventually help people who cannot communicate and, more speculatively, decode dreams.

CNET·
078

OpenAI's math proof release falls short of new field standards

OpenAI this week published hundreds of claimed solutions to hard math problems (719 manuscripts, per the article), saying it consulted an advisory group of elite mathematicians to avoid the controversy its earlier result sparked. But only ten of the manuscripts included the model's chain of thought, and by the article's account roughly 42% of the proofs had not gone through formalization in Lean; the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton's Institute for Advanced Study, said it is ultimately up to the mathematical community to judge whether its recommendations were followed — its first request being to stop testing advanced problems on proprietary models. A paper from Cambridge and King's College London mathematicians also documents at least two discrepancies between OpenAI's natural-language proof and the Lean code for a problem derived from the Navier-Stokes equations, concluding such autoformalized proofs should not be trusted without peer review.

TechCrunch AI·
088

Periodic Labs founders on 'synthesis superintelligence' and autonomous labs

In a Latent Space podcast interview, Periodic Labs co-founders Liam Fedus and Ekin Dogus Cubuk laid out their "synthesis superintelligence" thesis: rather than only training on internet data, the lab grounds reinforcement learning environments in real physical experiments so models can reason over noisy, incomplete and scarce measurements while predicting, synthesizing and characterizing materials. They argue failed experiments and negative results may be the most valuable training data, and that giving every lab instrument "a 140 IQ" could compress decades of scientific trial-and-error into months. They also say even future frontier models will still need to run real experiments.

Latent Space·
0910

GPT-6 Astra finds exact plasma equilibria, overturning 59-year-old Grad conjecture

University of Maryland plasma physicist Matt Landreman prompted GPT-6 Astra Pro to design an asymmetric three-dimensional plasma equilibrium; the model returned a family of exact analytic solutions in 20 minutes 34 seconds and, the next day, after a failed first attempt plus a proof that the first construction could not yield a non-integer rotational transform, produced a second family with magnetic shear in 33 minutes 37 seconds. Two arXiv papers published a day apart in late September overturn Harold Grad's 1967 conjecture: Landreman's paper credits GPT-6 Astra Pro in its acknowledgements, says parts were drafted by it, and publishes the prompts and verification scripts. A separate 147-page paper, "Counterexamples to the Grad conjecture," from Brown, Oxford and Bar-Ilan researchers used GPT-5.6 Sol, Claude Fable 5 and Claude Opus 5 for technical detail and computation, and ships a Lean 4 formalisation.

36氪 人工智能·
109

Google DeepMind Institute essay: AI as an 'invention of a method of invention'

Google DeepMind Institute published an essay, "Bending the Curve of Discovery," by Alex Imas and James Manyika on AI and science. It argues that today's LLMs and specialized models like AlphaFold act as economic complements, with most handoffs between them still run by the scientist; in a possible future, LLMs would orchestrate those handoffs automatically — prompting specialized models, auditing outputs and looping until a question is answered or a non-automated stage such as wet-lab testing is reached — shifting the scientist's role to designing that loop. The authors also raise epistemic questions about scientific understanding, theory-building, training the next generation of scientists and motivation, and note that as hypothesis generation gets cheap the bottleneck moves downstream to verification and choosing which questions matter, making this an organizational and institutional challenge as much as a technical one.

Google DeepMind (X)·
119

UniPat AI's PaperBenchX benchmark: GPT-6 Astra fully reproduces only 13.98% of papers

UniPat AI released PaperBenchX, which it calls the first multidisciplinary end-to-end benchmark for reproducing published paper results: 93 reproduction tasks built from 93 papers span 12 research directions and 10 domain-native scientific environments (including Ansys HFSS, Ansys Lumerical, Meep, PySCF and ABACUS), with 3,168 expert-verified scoring items. Across the 93 tasks, the strongest tested configuration, GPT-6 Astra, achieved only a 13.98% full-reproduction rate; averaged over tested configurations, modeling and execution scored about 62% each while validation reached just 42.8%. Per-task budgets run 4–24 hours (median 7 hours), and 12 tasks are open-sourced with 81 held out.

量子位(原生 RSS)·
127

OpenAI publishes 377 AI math results at once

DoNews reports that OpenAI released 377 AI-related mathematics results in a single batch. The report does not specify the form of the results, the research areas involved, or where they were published.

DoNews·
137

Biohub leads AI drug-discovery data infrastructure, scale grows to $1.8B

DoNews reports that Biohub is leading an effort to build data infrastructure for AI-driven drug discovery, with the scale of the initiative growing to $1.8 billion. The report does not detail participating organizations, the datasets covered, or how the funding will be used.

DoNews·
149

OpenAI's math advances spark a reckoning for academia

A Bloomberg Technology newsletter reports that OpenAI's advances in mathematics are prompting a reckoning in academia, with one mathematician reflecting on how AI is changing his profession. The item offers no specifics on models, benchmarks or results.

Bloomberg Technology·
1510

DeepMind stresses math collaboration and tooling

Google DeepMind researcher Pushmeet Kohli said on X that advancing mathematics is about empowering mathematicians, and that DeepMind has worked closely with the mathematical community since its early days to build new tools that push the boundaries of what is possible. The post names no specific tool, model or launch date.

Google DeepMind (X)·
166

Anthropic: AI-enabled attack risk can't be judged by technique count alone

Anthropic's threat intelligence team analyzed 832 accounts banned for malicious cyber activity between March 2025 and March 2026, mapping 13,873 actions across 482 techniques and all 14 MITRE ATT&CK tactics. A state-sponsored espionage campaign it disrupted in November 2025 used 30 techniques across 13 tactics — comparable to many medium-risk actors — yet scored the maximum 100 on Anthropic's own risk methodology. The data also shows AI use shifting toward post-compromise activity, with account discovery up 8.9% and AI-assisted phishing down 8.6%, while actors rated medium risk or higher rose from 33% to 56%.

TechRadar·
175

Inside OpenAI's 722 math manuscripts: the headline claims, formalized or not

OpenAI published 722 math manuscripts produced by an unreleased internal model to a GitHub repository, and a 36Kr/机器之心 review picks out the most famous results among them. The manuscripts are grouped into 372 "result families" spanning 17 areas such as number theory, geometry, theoretical computer science and mathematical physics; about 50 families claim falsifications or counterexamples, and roughly 60% include Lean formalization. Claims listed include the quasi-Riemann hypothesis, the BSD formula and Goldfeld's conjecture, Hilbert's tenth problem over the rationals, the Unique Games Conjecture, a Hadwiger conjecture counterexample and the free group factor isomorphism problem — all described as OpenAI's own claims and mostly not yet peer reviewed.

36氪 人工智能·
186

AI labs said to probe whether models can break cryptographic protocols

Theoretical computer scientist Omer Reingold, citing sources, says several AI labs have quietly begun testing whether their models can break important cryptographic protocols; Scott Aaronson relayed the claim on his Shtetl-Optimized blog with Reingold's permission. No labs are named, no protocols or results are specified, and the claim remains secondhand and unverified.

Techmeme·
195

Anthropic: Glasswing verified 129,000 vulnerabilities, joins Cyber Verification Program

Anthropic said on October 6 that its standalone AI-assisted vulnerability research effort, Project Glasswing, is being folded into an expanded Cyber Verification Program (CVP). It reported that partners verified at least 129,000 software vulnerabilities between April and July, while its own Claude-based scanning of open-source projects uncovered another 5,500 from April to October; more than 33,000 of the total were rated critical or high severity.

iThome 台湾·
205

Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

Apple Machine Learning Research published Normalizing Trajectory Models (NTM) at NeurIPS, which models each reverse diffusion step as an expressive conditional normalizing flow with exact likelihood training, pairing shallow invertible blocks per step with a deep parallel predictor across the trajectory and trainable from scratch or from pretrained flow-matching models. Its exact trajectory likelihood also enables self-distillation: a lightweight denoiser trained on the score function induced by the model itself produces high-quality samples in four steps. On text-to-image benchmarks, NTM matches or outperforms strong image generation baselines in just four sampling steps while uniquely retaining exact likelihood over the generative trajectory.

Apple Machine Learning Research·