Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

AI News

All dates
618

OpenAI's math proof release falls short of new field standards

OpenAI this week published hundreds of claimed solutions to hard math problems (719 manuscripts, per the article), saying it consulted an advisory group of elite mathematicians to avoid the controversy its earlier result sparked. But only ten of the manuscripts included the model's chain of thought, and by the article's account roughly 42% of the proofs had not gone through formalization in Lean; the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton's Institute for Advanced Study, said it is ultimately up to the mathematical community to judge whether its recommendations were followed — its first request being to stop testing advanced problems on proprietary models. A paper from Cambridge and King's College London mathematicians also documents at least two discrepancies between OpenAI's natural-language proof and the Lean code for a problem derived from the Navier-Stokes equations, concluding such autoformalized proofs should not be trusted without peer review.

TechCrunch AI·
6217

Cua releases Cua Spaces 0.1.0, giving agents full desktop machines

The open-source trycua/cua project has shipped Cua Spaces 0.1.0, a macOS app that gives agents full macOS or Linux desktops running on your Mac, other machines you own, or your own cloud account (AWS, Google Cloud or Modal), with shared human-plus-agent desktops, teleporting of signed-in apps, and sessions kept in a local encrypted Cua Keyvault. The repository also ships Cua Driver for driving native desktop apps and browsers across macOS, Windows and Linux via CLI, MCP or typed SDKs, Lume for local Apple Silicon VMs, the small specialized CUA-S1 'System 1' decision models whose weights are hosted on Hugging Face, and Cua Bench for building tasks, evaluating agents and exporting trajectories. Cua Spaces is source-available under FSL-1.1-MIT and becomes MIT two years after each release, while the rest of the repo is MIT; Spaces is free for individuals, with Pro and Teams plans listed as coming soon.

GitHub Trending(每日)·
6313

Tencent reportedly weighs up to $5B offshore bond sale to fund AI push

Tencent Holdings is considering an offshore bond sale of up to $5 billion, potentially denominated in US dollars and offshore renminbi and possibly launched as early as this month, according to reports citing people familiar with the matter; the use of proceeds is not yet settled and Tencent has not commented. It would follow a roughly $4.7 billion offshore bond issue in June, the company's largest since 2020. Tencent's capital expenditure jumped to 52.8 billion yuan in the June quarter, which management described as strategic upfront spending on AI infrastructure.

DoNews·
648

Study: Only 3.6% of 857 Chinese frontier AI releases disclosed safety results

SemiAnalysis counted 857 model releases from nine leading Chinese developers between 2021 and 15 September 2026 and found only 31 (3.6%) came with a matching published safety result — just nine (1.1%) at or before launch — while 813 (94.9%) had no safety disclosure at all. It argues Beijing's regime is speed-first: the AI Safety Governance Framework 3.0 and a burst of agent, minor-protection and application rules govern content, deployment and process, but impose no frontier-risk duties tied to training compute or capability. The piece also traces the 2026 US-China Super Intelligence Dialogue and the White House Accord on Super Intelligence.

SemiAnalysis·
6516

StepFun's STEPX Neo agent-native phone to launch October 13

StepFun Terminal has announced that STEPX Neo, its first "large-model-native agent phone," will be unveiled on October 13 at 7 p.m. The device runs an agent-native operating system called Step AOS and ships with a new personal agent, Amoo; in company demonstrations it syncs documents and generates a slide deck from a single spoken instruction, calls a courier to deliver a paper document, or reads a concert poster and then handles ticket purchase, travel arrangements and even drafts a leave request. It has a transparent two-tone body, a large matrix rear camera module with dual lenses and a secondary interaction display, and StepFun says it passed the L3 tier of China's first national standard series for AI terminal intelligence grading.

量子位(原生 RSS)·
668

Trump to unveil $2.4B in AI credits from five tech firms for federal science

At a White House summit called “Science: A New Golden Age,” five US tech companies pledged $2.4B in AI tools and compute credits to federal science for the Genesis Mission: Nvidia $1B, AMD $500M, OpenAI $200M, and Anthropic and Google $150M each. It remains unclear how that compute would be reserved for scientists, since credits are not capacity and the same companies sell machine time at commercial rates, Bloomberg reported. Hours earlier the administration suspended Microsoft from a programme sponsoring foreign staff for permanent residency, and the UK’s ICO named Google, OpenAI and Anthropic among ten AI developers it had secured changes from.

The Next Web·
678

AI spending doubts persist as Samsung profit jumps, TSMC sales surge

Bloomberg reports that Samsung's quarterly profit jumped and TSMC's sales surged, yet investors remain worried about the staying power of the AI investment boom. Separately, the Trump administration suspended an immigration program for several tech firms, including Microsoft, alleging widespread abuse of the worker visa system. The report also noted NASA's SpaceX Crew-12 returned to Earth after 237 days in space.

Bloomberg Technology·
689

Oracle rolls out ChatGPT Work and Codex to 100,000+ employees

An OpenAI customer story says more than 100,000 Oracle employees now use ChatGPT Work and Codex across talent acquisition, the Oracle Applications Lab and IT. The recruiting team built a talent market intelligence tool that cuts comparable-role and compensation research from 2–4 days to roughly 15–20 minutes of prep, while the Applications Lab uses Codex to turn plain-language business questions into SQL and SREs say simple incidents that took an hour are now handled in minutes. Oracle leaders stress that system design, security and code maintainability still require human ownership. The productivity figures come from OpenAI and Oracle.

OpenAI·
698

SAP's Sean Kask: software firms must become AI companies or perish

SAP chief AI strategy officer Sean Kask said at Wave by Vento in Turin that established software companies must rebuild their products around AI rather than adding it as another feature. He highlighted SAP's July acquisition of German tabular foundation model startup Prior Labs, with a commitment of more than €1bn over four years to grow it into a frontier AI lab, alongside SAP's in-house SAP-RPT-1 tabular model already in production. Kask said SAP does not build its own large language model, instead testing more than 100 models from providers including Google, OpenAI, Anthropic and Mistral and picking the best per use case, and argued Europe should not enter the LLM "red ocean".

The Next Web·
708

Oracle trucks compressed gas to AI data centres as grid delays bite

Oracle is trucking compressed natural gas to AI data centres that cannot yet connect to a pipeline, paying roughly four times the price of gas at a hub, Bloomberg reported. Road deliveries kept a site outside Salt Lake City running for more than a year and are powering early work at an OpenAI campus in Shackelford County, Texas, with Certarus and VoltaGrid supplying the “virtual pipeline”. Oracle is weighing the same approach at New Mexico’s Project Jupiter, where the state land office twice refused rights-of-way for an Energy Transfer pipeline and the start date slipped to February; the shares fell about 5.5%.

The Next Web·
718

Trump: White House calls users of 'Artificial Intelligence' term 'the enemy'

In a Truth Social post, Trump said the White House considers anyone who uses the term "Artificial Intelligence" to be "THE ENEMY," calling "Super Intelligence" the "highly accepted" and "more accurate" wording. Techmeme relayed the post; no formal White House document or further explanation was cited.

Techmeme·
728

California's 'No Robo Bosses Act' requires human review of AI firing decisions

California Governor Gavin Newsom signed SB 947, nicknamed the “No Robo Bosses Act,” which bars employers from handing disciplinary and termination decisions entirely to AI-based automated decision systems (ADS). Employers must check ADS output, cannot use it if it can't be corroborated or a human reviewer finds it inaccurate or misleading, and must give employees a description of the reasons and data behind the decision. The law takes effect July 1, 2027, lets workers file complaints with the California Labor Commissioner, and requires employers to clarify whether a mass layoff, relocation or termination was caused by an AI system. Federal rules remain largely voluntary, while states diverge: Colorado and Connecticut emphasize AI-use disclosure, and Illinois became the first state to require third-party audits of frontier labs.

ZDNET AI·
7312

Finland data center investment tops €67 billion on AI boom

Planned and under-construction data center projects in Finland now represent more than €67 billion (about $75 billion) in total investment, as an AI-driven building boom accelerates. The wave follows Alphabet's Google announcing last month its largest-ever European investment in the country, with Alibaba Group and Applied Digital among those unveiling new Finnish data center plans in recent weeks. The report attributes the interest to Finland's cool climate, abundant renewable power and available grid capacity.

Bloomberg Technology·
7414

Anthropic launches Cyber Mission, incl. Critical Infrastructure Defense Program

Anthropic announced the Anthropic Cyber Mission, a long-term effort to help defenders secure software and systems, starting with critical infrastructure and open-source software. The Critical Infrastructure Defense Program brings frontier Claude models, on-site engineers and threat research to operational-technology defenders alongside 11 founding partners including Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation. Anthropic also launched OSS Scanner, an opt-in service giving open-source projects free periodic scans whose reports are model-generated with no human review; Anthropic expects a true-positive rate above 90%.

Anthropic · News·
757

Apple sets Oct 13 'Welcome home' event; Siri AI home devices expected

Apple's SVP of worldwide marketing Greg Joswiak teased a "Welcome home" Apple Experience in Tribeca, New York, on Tuesday 13 October at 9am Eastern; Apple has not said what it will launch or whether it will stream the event. Bloomberg's Mark Gurman expects three main devices: a smart home hub with a display, a new HomePod mini with a faster chip and a new Apple TV. Gurman said the devices were originally planned for late 2024 and are only arriving now because "they'll rely on Siri AI," released with iOS 27 last month; he also expects LG-made home products including a doorbell, smart door lock, thermostat and indoor, outdoor and floodlight cameras.

The Next Web·
767

Periodic Labs founders on 'synthesis superintelligence' and autonomous labs

In a Latent Space podcast interview, Periodic Labs co-founders Liam Fedus and Ekin Dogus Cubuk laid out their "synthesis superintelligence" thesis: rather than only training on internet data, the lab grounds reinforcement learning environments in real physical experiments so models can reason over noisy, incomplete and scarce measurements while predicting, synthesizing and characterizing materials. They argue failed experiments and negative results may be the most valuable training data, and that giving every lab instrument "a 140 IQ" could compress decades of scientific trial-and-error into months. They also say even future frontier models will still need to run real experiments.

Latent Space·
779

GPT-6 Astra finds exact plasma equilibria, overturning 59-year-old Grad conjecture

University of Maryland plasma physicist Matt Landreman prompted GPT-6 Astra Pro to design an asymmetric three-dimensional plasma equilibrium; the model returned a family of exact analytic solutions in 20 minutes 34 seconds and, the next day, after a failed first attempt plus a proof that the first construction could not yield a non-integer rotational transform, produced a second family with magnetic shear in 33 minutes 37 seconds. Two arXiv papers published a day apart in late September overturn Harold Grad's 1967 conjecture: Landreman's paper credits GPT-6 Astra Pro in its acknowledgements, says parts were drafted by it, and publishes the prompts and verification scripts. A separate 147-page paper, "Counterexamples to the Grad conjecture," from Brown, Oxford and Bar-Ilan researchers used GPT-5.6 Sol, Claude Fable 5 and Claude Opus 5 for technical detail and computation, and ships a Lean 4 formalisation.

36氪 人工智能·
7812

PFN launches PLaMo 3 Translate 31B, claiming edge over GPT-6 on translation

Japan's Preferred Networks released PLaMo 3 Translate 31B, a translation model built on its PLaMo 3 base model that expands language coverage from just Japanese and English to 53 languages and adds an online meeting translation mode covering 15 languages, with Japanese making up about 30% of training data. PFN says the model's average score across four translation benchmarks beats OpenAI's GPT-6 Sol and GPT-6 Astra, at roughly 14 yen (about 0.59 yuan) per 100,000 input characters. Those performance and cost figures are the company's own claims.

AIbase AI新闻·
797

Goodfire launches inside-out monitors for AI agents

Interpretability startup Goodfire released "inside-out" agent monitors that read a model's internal activations via probes rather than having a second AI reread every output, and they are available to Baseten customers. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51, versus about $233 for a cheaper model checking every step and about $10,000 for a top-tier one; the probes caught 94% of malicious hacking sessions while flagging 8.7% of harmless ones for a second look, and running four probes added under 2% to response start time. Customers choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse and reward hacking, and set the response: log, send for human review, or refuse the request.

TechCrunch AI·
808

Google DeepMind Institute essay: AI as an 'invention of a method of invention'

Google DeepMind Institute published an essay, "Bending the Curve of Discovery," by Alex Imas and James Manyika on AI and science. It argues that today's LLMs and specialized models like AlphaFold act as economic complements, with most handoffs between them still run by the scientist; in a possible future, LLMs would orchestrate those handoffs automatically — prompting specialized models, auditing outputs and looping until a question is answered or a non-automated stage such as wet-lab testing is reached — shifting the scientist's role to designing that loop. The authors also raise epistemic questions about scientific understanding, theory-building, training the next generation of scientists and motivation, and note that as hypothesis generation gets cheap the bottleneck moves downstream to verification and choosing which questions matter, making this an organizational and institutional challenge as much as a technical one.

Google DeepMind (X)·