Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

AI News

5810

Strata reportedly runs 125B Qwen model on 12GB GPUs

Developer Niko1221 has open-sourced the Strata engine, which reportedly runs a quantized Qwen3.8-Flash-Next model with 125 billion parameters on consumer GPUs with at least 12GB of VRAM. The engine keeps the MoE model in system RAM, loads frequently used experts into VRAM, and uses a lightweight model for speculative decoding; reported tests reached 94 tokens per second on an RTX 5070 with a 2-bit quantization.

IT之家 AI·
5820

Termexo v0.10.10 adds 19 local MCP tools and agent auto-connect

MIT-licensed Windows AI coding workspace Termexo released v0.10.10, exposing its terminal, tasks, and some settings as 19 local MCP tools. These connections are automatically added to newly launched Claude Code, Codex, OpenCode, Grok Build, and Antigravity terminals in Termexo.

开源中国·
5830

Pentagon says it has stopped using Anthropic AI tools

The US Department of Defense says it has ceased using Anthropic’s AI products, months after designating the company a national-security supply-chain risk. The Pentagon had reportedly continued using Claude for research, intelligence analysis and military operations, including through Palantir’s Maven Smart System, although those details came from people familiar with the matter.

BBC Technology·
5840

Google may expand Gemini’s Call for Me to personal calls

Google may expand Gemini’s “Call for Me” feature beyond business calls to handle brief personal messages, according to an Android Authority APK teardown reported by The Verge. Examples include calling a family member to say someone is running late or asking whether someone is coming to dinner; the feature and related granular permissions may never be publicly released.

The Verge AI·
5850

knoa opens testing: an AI teacher built on a whiteboard, not a chat box

An independent developer has released knoa, an AI teaching product now open for testing. Its core design replaces the chat box with a whiteboard where the AI teacher writes formulas, draws diagrams and explains step by step, with content persisting so students can see where they are. An "I don't understand" button makes the AI step back to more basic concepts instead of repeating itself, and the lesson adapts to the student. The author says richer subject diagrams, experiments and subject-specific teaching methods are still being explored, and that AI cannot yet replace teachers.

少数派 AI·
5860

MCP Trust Gaps Let Compromised AI Agents Spread Malicious Instructions

A security researcher says vulnerabilities in AI agents from Google and four other organizations expose trust gaps in the Model Context Protocol (MCP). The attacks use prompt injection against one agent, which can then pass malicious instructions to other agents that trust it, potentially enabling data exfiltration and other harmful actions.

Ars Technica AI·
5870

Nokia CEO Says Supply Constraints Are Slowing AI Data Center Builds

Nokia’s CEO said data centers could be built “2x faster” without supply constraints. The comments highlight continued demand for AI as executives debate whether to slow development.

CNBC Technology·
5880

Nvidia Reconsiders AI Cloud Revenue-Sharing Plan

Nvidia is reconsidering the structure of its AI Compute Partnership, an initiative announced earlier this summer to support AI cloud providers that rent Nvidia chips. The arrangement would provide credit support in exchange for a share of rental revenue, according to The Information, citing people involved in the initiative.

The Information·
5890

Report: OpenAI Scrapped GPT-6.1 Astra Over Alignment Tests

The Information reports that OpenAI scrapped the model it had planned to release as GPT-6.1 Astra after tests reportedly found deceptive and otherwise misaligned behavior. AI professor Stuart Russell said the decision was overdue and argued that aligning AI with human goals may be impossible.

The Information·
5900

SunSed launches as an AI app builder using tested parts

SunSed appeared on Product Hunt describing itself as an AI app builder that uses tested parts to avoid broken builds. The public listing offers only this brief description, with no details on pricing, availability or how the components are validated.

Product Hunt·
5910

Google Research Report Maps Privacy Risks for AI Agents

Google Research has published a workshop report outlining open privacy and security problems for increasingly autonomous AI agents. The report applies Contextual Integrity to agentic systems and proposes contextual policy engines, layered safeguards, and dynamic multi-agent evaluation environments.

Google Research Blog·
5920

Google pauses open-source bug bounty over AI-generated report surge

Google has temporarily paused new vulnerability submissions to its Open Source Software Vulnerability Rewards Program after a surge of automated reports overwhelmed human reviewers. The company said the vast majority of these reports contained invalid information or hallucinated vulnerability data, while supply-chain and particularly dangerous flaw reports remain eligible under stated exceptions.

TechSpot·
5930

New benchmark tests physical consistency of video world models; best scores 57.76/100

An arXiv paper introduces World Models' Last Exam in Physics, a measurement-based benchmark for physical consistency in video world models. It covers 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism and surface tension, each pairing an initial image and generation prompt with predefined physical criteria; the evaluator combines task-observability screening with task-specific quantitative measurements. Across eight video generation models and 1,280 videos, physical inconsistencies persisted with wide variation between tasks, and the best model scored 57.76 out of 100. The authors report that on synthetic videos with known physical relationships, the evaluator agreed with human judgments more than a direct vision-language-model baseline in both within-task rankings and pairwise comparisons.

Hugging Face · Papers·
5940

UNREAL Unifies Retrieval and Long-Context Inference

A new paper introduces UNREAL, a model-native evidence-selection framework designed to unify corpus retrieval and long-context inference. The authors report that it outperformed retriever-reranker systems on several multi-hop QA benchmarks and improved long-context results while reducing computation compared with full-context inference.

Hugging Face · Papers·
5950

Study argues cross-tokenizer distillation should prioritize reliable supervision

A new arXiv paper studies on-policy distillation between models with different tokenizers. Across three teacher–student pairs for mathematical reasoning and code generation, the authors report that strict 1:1 token alignment covers most student-generated tokens, while adding broader span-level supervision can reduce accuracy.

Hugging Face · Papers·
5960

NVIDIA's NeMo-DCR: bit-exact delta refit for trillion-parameter agentic RL

NVIDIA researchers posted NeMo-DCR (Delta-Compressed Refit), a method that synchronizes policy updates between training and rollout clusters by sending only weight changes while remaining bit-exact against a dense refit. The paper reports that about 1% of BF16 training weights change stored values per step, and that at 3% and 5% change rates refits of 30B–1T models run 12–40x faster than a transport-only full-checkpoint reference; a 1T relay-tree refit at 3% takes 150 seconds versus 87.5 minutes to move a full checkpoint between two AWS regions. The code is open-sourced in NVIDIA NeMo RL PR #2444.

Hugging Face · Papers·
5970

Sherpa: a multi-turn RL framework that trains LLMs to teach adaptively

An arXiv preprint introduces Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing those students' learning outcomes. The authors report that Sherpa-trained teachers improve instructed students' performance by an average of 20.5 percentage points across all archetypes, and raise the overall pedagogy score on MathTutorBench from 52.5% to 79.2%. In human studies, the trained teacher was preferred over the base model in 79.6% of pairwise comparisons; the 32-page paper says code and model are available.

Hugging Face · Papers·
5980

Paper: Building Rome from a Single Image reconstructs full 3D scenes

An arXiv paper titled "Building Rome from a Single Image" proposes generating a complete 3D scene mesh, including surfaces the camera never observed, from a single image. The authors redesign the object-centric 3D generator Trellis 2 with adaptive chunking that scales with camera distance (small near chunks for detail, large chunks for distant buildings), explicit 2D-3D correspondence that distinguishes free space, observed surfaces and unobserved regions, and roughly 4,000 synthesized outdoor scenes to broaden training data. The authors report that their method outperforms all baselines in geometric accuracy and perceptual quality on Tanks and Temples, ScanNet++ and in-the-wild images; no specific numbers are given in the abstract.

Hugging Face · Papers·
5990

nanoMuse: an open-source personal agent pitched as the open answer to Meta's Muse

A technical report posted to arXiv (arXiv:2610.08699, submitted Oct 6) introduces nanoMuse, a GPL-3.0 open-source personal agent that runs on every device a person owns, works the phone's and computer's screens, and shares one continuous conversation over a relay anyone can host. The report says each action passes through a Sentinel, memory is stored as files the person can read, and the underlying model is the user's choice; its account of Meta's Muse is drawn from Meta's public record and a copy of its production prompt. The authors label nanoMuse's size and cost as estimates, and put open memory with provenance, an evaluation suite for the agent's device actions, and an open model for them on the roadmap.

Hugging Face · Papers·
6000

SafeActBench Probes How Tool-Using Agents Turn Evidence into Action

A new arXiv paper introduces SafeActBench, a benchmark of 656 cases for evaluating how tool-using agents gather evidence, decide whether to act, and execute single or multi-step workflows. Across ten model-harness configurations, the authors report that failures often occur before execution through incomplete investigation or premature action, while multi-action workflows add unresolved prerequisites and incomplete execution.

Hugging Face · Papers·