Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

All dates
0148

OpenAI posts 722 AI-generated math manuscripts on GitHub, reigniting ethics row

OpenAI published 722 math manuscripts, grouped into 372 result families and produced by an unreleased internal model, in a public GitHub repository rather than peer-reviewed journals. The work spans a claimed "quasi-Riemann hypothesis," the BSD conjecture and Hilbert's tenth problem over the rationals, drawn from roughly 4,000 research problems at an average of about three hours of ChatGPT Pro thinking compute per result, with Lean formalizations for many. Some mathematicians criticized the company for bypassing academic publication norms and for a credit dispute, while the same internal model's claimed Navier–Stokes Millennium Prize solution remains under formal review.

WIRED AI·
025

Meta says Muse Spark helped address five open math questions

Meta says researchers used Muse Spark 1.1 and 1.2 in Thinking Mode through the regular Meta AI chat interface while working on six mathematics papers. Five papers addressed previously open questions, but mathematicians selected the problems, guided the work, and independently checked, corrected, and refined the model’s contributions.

TechRepublic·
035

OpenAI and Ironclad train agents for complex contracting workflows

OpenAI says it partnered with Ironclad to train and evaluate computer-using agents on complex contracting workflows, including agreement configuration, approvals, and reusable legal terms. In OpenAI’s research evaluation across 11 tasks, GPT-6 Astra averaged 55.0% versus 41.6% for GPT-5.6 Sol, while estimated time per attempt fell from 37.0 minutes to 19.2 minutes; these results come from OpenAI’s internal evaluation.

OpenAI·
045

Google Earth AI model shows promise for public-health forecasting

Google Research reports that its Population Dynamics Foundation Model (PDFM), part of Google Earth AI, can provide plug-in geospatial embeddings for public-health and epidemiological workflows. Partner evaluations covered vaccination, cardiovascular disease, dengue, postpartum depression and cholera, with reported improvements or comparable performance against conventional inputs in several settings.

Google Research Blog·
055

AdvSim2Real trains web agents against adaptive prompt injection

The AdvSim2Real paper presents a simulated training setup that co-evolves web tasks, adaptive prompt-injection attacks, and an agent in a frozen web world model. The authors report that training a 4B agent this way improved completion with and without attacks, transferred to a real browser, and increased completion under an unseen frontier-model adversary by 33.6% relative to the base agent on 150 web tasks.

Hugging Face · Papers·
065

Erdosproblems.com freezes proof claims amid AI-generated math submissions

Erdosproblems.com founder Thomas Bloom says the site will freeze new problem comments and proof claims following a wave of AI-generated submissions, many without explanations. The site will also remove problem statuses and credit-oriented language, while emphasizing high-quality expositions and formalizations.

Hacker News · AI(100+ 分)·
074

Francis Halzen Reportedly Wins the 2026 Nobel Prize in Physics

The supplied report says Francis Halzen of the University of Wisconsin–Madison won the 2026 Nobel Prize in Physics alone for decisive contributions to the IceCube Neutrino Observatory and the discovery of high-energy astrophysical neutrinos. IceCube uses 5,160 optical sensors embedded 1,450–2,450 meters beneath Antarctic ice to detect faint light from neutrino interactions.

量子位(原生 RSS)·
085

Fleming Initiative launches AI evaluation programme for antimicrobial resistance

The Fleming Initiative announced a new three-year programme supported by Google DeepMind to develop methods and standards for evaluating AI systems used in antimicrobial resistance. The programme aims to assess whether such systems are accurate, reliable and ready for use.

Google DeepMind (X)·
094

Terence Tao Warns of “Proof Indigestion” as AI Mathematics Accelerates

A report says Fields Medalist Terence Tao has urged AI companies to slow down their pursuit of mathematical breakthroughs, arguing that machine-checked proofs are advancing faster than human interpretation, peer review and textbook integration. It also reports that OpenAI has formed an independent Mathematics and AI Advisory Group at the Institute for Advanced Study, while stating that the group will not advise on OpenAI’s internal mathematical progress.

量子位(原生 RSS)·
105

Microsoft Research podcast: what AI evaluation gets wrong

In a Microsoft Research podcast episode, host Chad Atalla speaks with Jennifer Neville, who leads the AI Interaction and Learning team at Microsoft Research and is a professor at Purdue, about how evaluation pushes the performance boundaries of today's AI systems and about the “surprising failures” that appear when models are tested beyond traditional benchmarks. Neville argues that the benchmarks commonly used in ML/AI are fairly simple relative to real-world use, so her team designs evaluations around multiturn behavior, collaborative settings and long-horizon tasks to expose performance gaps and then drive algorithmic and model improvements. She also offers practical guidance for working with current AI systems and stresses examining the data closely when results defy expectations.

Microsoft Research Blog·
113

Apple's RISED uses rubrics for multi-environment agent training

Apple Machine Learning Research published RISED, a method that repurposes rubrics to guide online data selection and policy supervision when training a single LLM agent jointly across diverse interactive environments. An LLM judge tags each rollout with a rubric vocabulary shared across environments; the resulting profiles select data that matches the mixed-environment batch's behavioural composition while limiting overlap, positive rubrics provide privileged context for an on-policy self-distillation teacher's token-level supervision, and negative rubrics steer later rollouts away from recurring failure modes. The paper reports that across model backbones RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment.

Apple Machine Learning Research·

That's everything.