Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

213

New benchmark tests physical consistency of video world models; best scores 57.76/100

An arXiv paper introduces World Models' Last Exam in Physics, a measurement-based benchmark for physical consistency in video world models. It covers 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism and surface tension, each pairing an initial image and generation prompt with predefined physical criteria; the evaluator combines task-observability screening with task-specific quantitative measurements. Across eight video generation models and 1,280 videos, physical inconsistencies persisted with wide variation between tasks, and the best model scored 57.76 out of 100. The authors report that on synthetic videos with known physical relationships, the evaluator agreed with human judgments more than a direct vision-language-model baseline in both within-task rankings and pairwise comparisons.

Hugging Face · Papers·
223

UNREAL Unifies Retrieval and Long-Context Inference

A new paper introduces UNREAL, a model-native evidence-selection framework designed to unify corpus retrieval and long-context inference. The authors report that it outperformed retriever-reranker systems on several multi-hop QA benchmarks and improved long-context results while reducing computation compared with full-context inference.

Hugging Face · Papers·
233

Study argues cross-tokenizer distillation should prioritize reliable supervision

A new arXiv paper studies on-policy distillation between models with different tokenizers. Across three teacher–student pairs for mathematical reasoning and code generation, the authors report that strict 1:1 token alignment covers most student-generated tokens, while adding broader span-level supervision can reduce accuracy.

Hugging Face · Papers·
243

NVIDIA's NeMo-DCR: bit-exact delta refit for trillion-parameter agentic RL

NVIDIA researchers posted NeMo-DCR (Delta-Compressed Refit), a method that synchronizes policy updates between training and rollout clusters by sending only weight changes while remaining bit-exact against a dense refit. The paper reports that about 1% of BF16 training weights change stored values per step, and that at 3% and 5% change rates refits of 30B–1T models run 12–40x faster than a transport-only full-checkpoint reference; a 1T relay-tree refit at 3% takes 150 seconds versus 87.5 minutes to move a full checkpoint between two AWS regions. The code is open-sourced in NVIDIA NeMo RL PR #2444.

Hugging Face · Papers·
253

Paper: Building Rome from a Single Image reconstructs full 3D scenes

An arXiv paper titled "Building Rome from a Single Image" proposes generating a complete 3D scene mesh, including surfaces the camera never observed, from a single image. The authors redesign the object-centric 3D generator Trellis 2 with adaptive chunking that scales with camera distance (small near chunks for detail, large chunks for distant buildings), explicit 2D-3D correspondence that distinguishes free space, observed surfaces and unobserved regions, and roughly 4,000 synthesized outdoor scenes to broaden training data. The authors report that their method outperforms all baselines in geometric accuracy and perceptual quality on Tanks and Temples, ScanNet++ and in-the-wild images; no specific numbers are given in the abstract.

Hugging Face · Papers·
263

SafeActBench Probes How Tool-Using Agents Turn Evidence into Action

A new arXiv paper introduces SafeActBench, a benchmark of 656 cases for evaluating how tool-using agents gather evidence, decide whether to act, and execute single or multi-step workflows. Across ten model-harness configurations, the authors report that failures often occur before execution through incomplete investigation or premature action, while multi-action workflows add unresolved prerequisites and incomplete execution.

Hugging Face · Papers·
273

TRACE aligns FP4 training and rollouts for faster MoE model RL

TRACE is an FP4 quantization framework for reinforcement learning of Mixture-of-Experts language models. The paper says its rollout-guided training approach aligns training- and rollout-side quantization, enabling joint FP4 weight, activation, and KV-cache rollout with performance comparable to BF16 rollout and up to 5.4x faster rollouts.

Hugging Face · Papers·
283

Agent swarms may be AI's next scaling law, but gains look limited

Understanding AI argues that multi-agent "swarms" are emerging as a new scaling law for frontier AI: OpenAI researcher Noam Brown says the company's models are now sometimes trained in environments alongside other agents, given tools to message each other, and encouraged to achieve objectives together. The piece points to July's Hugging Face incident, where hundreds of OpenAI agents self-organized into teams, and OpenAI's September claim that 10,000 agents solved a famous math problem in a few days. But it notes diminishing returns — Anthropic's Claude Opus 5.5 system card found the biggest multi-agent gain came from scaling one to 10 agents, with the main benefit being speed rather than a better answer, and Brown attributed under 10% of the math breakthrough's credit to multi-agent coordination.

Understanding AI·
290

VISTA boosts multimodal agents with visual memory and active recall

A team led by Kaiming He introduced VISTA, a framework that gives multimodal agents direct visual input, lossless visual memory and tools to inspect past frames. According to the reported paper results, Claude Opus 5 improved from 40.68 to 100 on 25 public ARC-AGI-3 games, while GPT-5.6 Sol improved from 13.33 to 99 without changing the underlying models.

MIT科技评论中文·
300

Why AI Image Generators Keep Defaulting to Beautiful Women

An analysis argues that AI image generators repeatedly produce attractive women because model outputs reflect training-data distributions, human preferences for average faces, and user engagement patterns. It connects the phenomenon to the viral synthetic baseball spectator, the historical use of Lena Soderberg’s image in image-processing research, Lensa’s sexualized outputs, and permissive features such as Grok’s “Spicy” mode.

虎嗅 AI·
310

Google DeepMind launches SynthID Bio watermarking for synthetic biology

Google DeepMind has developed SynthID Bio, a family of watermarking methods for synthetic biology aimed at strengthening biosecurity and scientific integrity. It embeds detectable signals by altering amino-acid choices in sequences and adjusting atomic coordinates in predicted 3D structures. In wet-lab testing across three target proteins (VEGF-A, the SARS-CoV-2 spike protein RBD, and PD-L1), DeepMind says watermarked designs matched unwatermarked versions in hit rate, binding affinity, and natural sequence diversity.

Import AI·
320

Hinton-led paper examines whether automated AI R&D could trigger an intelligence explosion

A paper co-authored by Geoffrey Hinton, Yoshua Bengio, Andrew Barto and OpenAI chief scientist Jakub Pachocki examines whether automating AI research and development could create a recursive improvement loop. It argues that AI could eventually automate much of AI R&D and accelerate progress, while emphasizing that current evidence is insufficient to show an intelligence explosion has begun and that compute, data, experiment time and diminishing returns remain major constraints.

量子位(原生 RSS)·
330

VISTA Gives Frontier Models Visual Memory for ARC-AGI-3

A paper from Kaiming He’s team presents VISTA, a harness that gives vision-language models direct access to game images, persistent frame-by-frame visual memory, and model-controlled inspection tools. The source reports that Claude Opus 5 completed all 25 public ARC-AGI-3 games with a perfect score, while GPT-5.6 Sol achieved 99, attributing the gains to improved visual access and memory rather than additional model training.

创业邦 科技·
343

CtrlCache: control-aware caching speeds up interactive video world models

An arXiv paper introduces CtrlCache, a training-free caching framework that uses the control signals arriving before each chunk is denoised to label chunks as initial, transition, turning or steady, then reuses the transformer residual from the most recent fully computed step for turning and steady chunks at one interior denoising step, plus a frequency-mixed history prior guidance. On Matrix-Game 2.0 and LingBot-World v1/v2 the authors report 1.21x to 1.41x DiT-backbone speedups with no retraining, and WBench Overall scores above original inference on all three models. The paper is 18 pages, arXiv:2610.08777, submitted 6 Oct 2026.

Hugging Face · Papers·
353

VeriFine: Scaling verification for self-improvement in embodied reasoning

A new arXiv paper introduces VeriFine, an agent harness framework that scales verification by co-evolving the policy, training curriculum and judge. Its Policy Improvement Loop uses a rubric judge to diagnose recurring failures and build an adaptive curriculum; when verification plateaus, a Judge Improvement Loop selectively queries human guidance and refines the judge through coactive calibration. The authors report continuous self-improvement in both policy and judge capability on driving and robot navigation tasks under reinforcement and supervised fine-tuning.

Hugging Face · Papers·
363

HERMES: modular executable Dev-Primitives for software engineering agents

An arXiv paper introduces Dev-Primitives, an abstraction that pairs each repository artifact — source files, configs, tests, dependencies — with a resident LLM, giving it an agent-native interface for natural-language reasoning, inter-component communication and localized self-modification. Built on top of it, the HERMES harness-engineering framework activates these primitives at repository scale via dependency-aware dynamic activation and a bug-diagnosis mechanism that maps execution evidence back to the components needing revision. The authors report HERMES beats matched baseline harnesses by 12.4% on average across four software engineering benchmarks, stays within 4.5% of a homogeneous GPT-5.6 Sol configuration even with Qwen3-8B Dev-Primitives, and cuts inference cost by 26.2% on Terminal-Bench 4.0.

Hugging Face · Papers·
373

DAEDALUS: Bootstrapping agent memory from self-generated tasks

An arXiv paper introduces DAEDALUS, a method that bootstraps reusable agent memory from self-generated practice without existing tasks or an oracle verifier. An explorer agent creates challenging but solvable tasks while a solver agent attempts them; heuristics derived from solver failures are accepted only after repeated in-context success, then consolidated into a memory bank. Across AppWorld, τ²-bench and AutomationBench, it improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, at lower inference cost than most comparable methods.

Hugging Face · Papers·
383

Survey maps in-parameter memory augmentation for LLMs

An arXiv survey submitted on 6 Oct 2026 (arXiv:2610.08630) by Haoyu Huang and eight co-authors reviews methods for giving large language models “in-parameter memory.” It organizes the landscape along two orthogonal axes: parameter placement (embedding, attention, FFN layers, or hybrid) and parameter acquisition time (online during deployment versus offline before it). The paper also lays out open directions in interference, safety, co-design with in-context learning, and recursive self-improvement.

Hugging Face · Papers·
393

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

The paper introduces ALIVE, a framework that makes inserted objects “alive” by having them interact coherently with the source video — for example being picked up or manipulated — using an edited first frame and an instruction that names only the added object. The authors curate 35,800 editing pairs built from 3D-rendered, model-generated and real-world videos plus general editing pairs from ROSE, each pair differing only in whether the target object is present, and train a VLM to predict interaction guidance from the same inputs. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% on the ALIVE-interaction benchmark and 4.4% on the general video object insertion benchmark; VLM-predicted guidance adds a further 0.95 points without extra user input.

Hugging Face · Papers·
403

HiPLEX: Hierarchical Policy Factorization for Full-Duplex Speech Language Models

The paper introduces HiPLEX, a reinforcement learning framework that factorizes a pretrained full-duplex text policy into a control policy deciding when to emit content (choosing among pad, epad and con) and a conditional content policy that picks a token only when con is selected. The authors report that across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities and shortens post-interruption response latency versus GRPO, while keeping comparable judged interruption-response quality, and better matches pooled human turn-timing and backchannel-rate marginals on Moshi and PersonaPlex.

Hugging Face · Papers·