Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

All dates
415

Claude Opus 5.5 composes retro game music in Scrimshaw Jukebox test

On his blog, Simon Willison tested whether Claude Opus 5.5 could compose music, asking it to first design a simple text-based music format, then build an artifact that plays it aloud with example tracks, aiming for the quality of the original Secret of Monkey Island. The result was "Scrimshaw Jukebox," a pixel-art browser player containing six original adventure-game tracks (Moonlit Harbor, The Ghost Galleon and others), written as plain text and performed by an in-browser synthesizer with a piano-roll score view and editing. Willison called the output "surprisingly good" while noting it leaned harder into the Monkey Island theme than he intended.

Simon Willison's Weblog·
425

Google Earth AI model shows promise for public-health forecasting

Google Research reports that its Population Dynamics Foundation Model (PDFM), part of Google Earth AI, can provide plug-in geospatial embeddings for public-health and epidemiological workflows. Partner evaluations covered vaccination, cardiovascular disease, dengue, postpartum depression and cholera, with reported improvements or comparable performance against conventional inputs in several settings.

Google Research Blog·
435

Interconnects: open-weight cyber-risk discourse is broken

An Interconnects opinion piece argues the public debate over open-weight models' cyber risk has gone off the rails: Anthropic's report on GLM-5.3 as an offensive cyber tool is largely reasonable technically, but ducks cross-cutting questions such as what happens if open models are banned and why Chinese labs consider their models safe to release. The author notes that documented cyberattacks so far mostly involve closed models, so "open dangerous, closed safe" may really be "open unsafe, closed unsafe," and says Chinese companies must register major model releases with the government along with evaluations, though it is unclear whether those cover cyber or bio risks. He adds that GLM-5.3, not Claude Mythos, is the model that crossed the capability threshold, and that more than a month after its weights were released there is little public evidence of a step change in harm.

Interconnects·
447

Christian Szegedy:AI将重塑数学研究

新智元解读了 xAI 联合创始人 Christian Szegedy 关于数学未来的一篇长文。Szegedy 认为,随着 AI 在定理证明和自动形式化方面进展,数学可能从少数专家的手艺转变为科学技术的基础设施;文章同时引用了 OpenAI 和 Anthropic 关于解决数学问题的说法,但未独立核实这些案例。

新智元·
454

vLLM v0.31.0 expands serving, speculation, and model support

vLLM v0.31.0 adds broad serving and model-support updates, including DeepSeek-V4.1-Flash optimizations, Model Runner V2 speculative decoding, larger-scale expert parallelism, and improved multimodal support. It also introduces the `vllm preload` weight-cache daemon for faster engine restarts and experimental CRIU-based initialized-engine snapshots, alongside several security and breaking changes.

vLLM Releases·
464

Pinterest turns beauty Pins into AI-powered salon action plans

Pinterest introduced Beauty Guides, an AI feature that turns saved beauty Pins into actionable salon plans. Powered by Pinterest Intelligence, the guides translate visual inspiration into stylist terminology and provide estimated process times, price ranges, and maintenance requirements for hairstyles, hair color, nail shapes, and finishes.

TechCrunch AI·
474

Whistleblowers Say AI Researchers Are Checking Work Less Often

Three former AI lab employees told a New York City Council hearing that researchers are increasingly letting AI write code and conduct research, while checking its work less frequently. The labs have emphasized that AI has accelerated research, and Anthropic said in a recent report that AI was leading more than a quarter of its model R&D as of August.

The Information·
484

TechCrunch Disrupt 2026 publishes AI-focused breakout agenda

TechCrunch has published the breakout-session agenda for Disrupt 2026, scheduled for October 13–15 in San Francisco. The AI-focused sessions will cover AI agents, physical AI, inference economics, trustworthy AI, space applications, and growth strategies, alongside fundraising and company-building topics.

TechCrunch AI·
494

AdvSim2Real trains web agents against adaptive prompt injection

The AdvSim2Real paper presents a simulated training setup that co-evolves web tasks, adaptive prompt-injection attacks, and an agent in a frozen web world model. The authors report that training a 4B agent this way improved completion with and without attacks, transferred to a real browser, and increased completion under an unseen frontier-model adversary by 33.6% relative to the base agent on 150 web tasks.

Hugging Face · Papers·
504

Claude Threat Detection Reportedly Leads to Florida Arrest

A Florida woman was arrested after messages allegedly sent to Claude described plans to attack the Lee County Sheriff’s Office and mentioned obtaining a gun. According to local reports cited by TechRepublic, Anthropic’s safety systems flagged the conversation, human reviewers examined it, and the company notified law enforcement under its emergency disclosure policy.

TechRepublic·
515

Jump Trading Uses GPT-6 Astra for Agentic Quant Research

OpenAI says Jump Trading is using GPT-6 Astra to expand agentic workflows for coding, quantitative studies, and hypothesis validation. The trading firm describes long-running agents that can evaluate findings, redirect analysis, and combine improvements, while human review and controlled environments remain part of the process.

OpenAI·
5212

Utopai X reportedly ranks second in Artificial Analysis Video Arena

A Machine Intelligence report says Utopai X ranked second globally, with an Elo score of 1150, in Artificial Analysis's Video Arena in late September. The article attributes the result to custom post-training built on the MiniMax H3 architecture and integration with Utopai's Production Intelligence platform, while noting that the claims come from the reported evaluation and company materials.

机器之心·
534

Instinct’s AI assistant bet puts trust ahead of model benchmarks

Instinct, a 14-person AI assistant startup, reportedly raised $1 billion at a $10 billion valuation while offering an agent that users can contact by phone, text, email, WhatsApp, or iMessage. The article compares it with Meta Muse and OpenAI Dots, while highlighting unresolved issues around transaction claims, user trust, authorization, privacy, and costly agent errors.

虎嗅 AI·
544

AI Agents Challenge Workday App Usage and Growth

Workday customers are using AI agents from Anthropic, Microsoft and other providers to retrieve and analyze data without directly visiting the Workday app. Interviews with five consultants and partners serving more than 2,000 customers also suggest that many companies paying for Workday’s own AI tools do not use them regularly.

The Information·
554

Gamma 5 launches rebuilt presentation platform to strip the 'AI smell' from slides

Gamma introduced Gamma 5 on Oct. 6, a rebuilt presentation platform with a redesigned editor, a more capable AI agent and thousands of visual templates, aimed at complaints that AI-generated decks look repetitive and generic. The platform combines more than 20 AI models, including ones from Anthropic, OpenAI and Google, plus roughly 20 connectors to outside sources such as Slack, Granola, Fireflies and Fathom. Gamma says an internal benchmark for the visual accuracy of imported files improved from about 30% early in development to over 97%.

SiliconANGLE AI·
564

Visual Studio 2026 Adds More AI Assistance for Developers

TechRepublic describes Microsoft Visual Studio Professional 2026 as a 64-bit, cross-platform IDE with expanded AI assistance through IntelliCode, including context-aware suggestions, refactoring, and code completion. The article also promotes a limited-time $29.97 offer against a stated $499.99 MSRP, with prices subject to change and affiliate disclosures.

TechRepublic·
574

NVIDIA promotes open models for telecom AI and announces Nemotron 3 LTM

NVIDIA says telecom operators are increasingly using open models to gain more control, customization and deployment flexibility across network operations and customer care. It also announced the 30-billion-parameter Nemotron 3 Large Telco Model, fine-tuned by AdaptKey on open telecom datasets, and released a NeMo-based recipe for adapting open models to operator-specific data.

NVIDIA Blog·
584

Antseed launches decentralized marketplace for AI inference

Antseed launched a peer-to-peer inference marketplace that routes requests to hundreds of independent and mainstream model providers, bypassing centralized gateways, with a self-hosted local router that lets providers set their own prices. Its controlling entity, the Antseed Foundation, closed a $2.4 million token funding round led by Spark Capital, with participation from Collider, DCG, North Island Ventures, Reciprocal Ventures, Venice.ai and Relay Capital. The company claims users can reach frontier models for up to 97% less than official API prices, and says it has processed nearly 150 billion tokens across 202 active providers — figures it reports itself.

SiliconANGLE AI·
594

Anaconda launches agent swarms and AI red-teaming tools for enterprise platform

Anaconda is expanding its enterprise AI platform beyond Python packaging with agent swarms, agentic security testing and production deployment capabilities built on its acquisitions of Kilo Code, Enkrypt AI and Outerbounds. The new Kilo Desktop brings swarms into Visual Studio Code with access to more than 500 models, local model execution and automatic model routing; on security, it offers autonomous red-teaming across more than 300 attack categories, runtime controls and a public Agent Incident Registry. Anaconda cites its own research that 63% of respondents are moving toward swarms, and Enkrypt research finding vulnerabilities in 73% of agent tools examined across over 25,000 MCP servers, while saying 95% of the Fortune 500 use its software.

SiliconANGLE AI·
604

Erdosproblems.com freezes proof claims amid AI-generated math submissions

Erdosproblems.com founder Thomas Bloom says the site will freeze new problem comments and proof claims following a wave of AI-generated submissions, many without explanations. The site will also remove problem statuses and credit-oriented language, while emphasizing high-quality expositions and formalizations.

Hacker News · AI(100+ 分)·