Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

AI News

7410

VeriFine: Scaling verification for self-improvement in embodied reasoning

A new arXiv paper introduces VeriFine, an agent harness framework that scales verification by co-evolving the policy, training curriculum and judge. Its Policy Improvement Loop uses a rubric judge to diagnose recurring failures and build an adaptive curriculum; when verification plateaus, a Judge Improvement Loop selectively queries human guidance and refines the judge through coactive calibration. The authors report continuous self-improvement in both policy and judge capability on driving and robot navigation tasks under reinforcement and supervised fine-tuning.

Hugging Face · Papers·
7420

KLPO: a critic-free, KL-regularized policy optimization method for LLM agents

A new arXiv paper, "On KL-Regularized Policy Optimization" by Yifan Zhang (submitted 6 Oct 2026), proposes KLPO, which anchors the KL regularizer at the sampler so asynchronous RL for LLM agents can train on trajectories from stale checkpoints and mismatched inference probabilities without importance weights. The paper says the regularized improvement step has a closed-form Gibbs solution, and that for token-level policy mirror descent targets the gradient can be computed from terminal returns without a critic, using one rollout per prompt. It further proves independent Monte Carlo estimates of the KL term keep gradients unbiased and shows SPPO, GPO, REBEL and BPO arise as special cases of KLPO.

Hugging Face · Papers·
7430

NP-OPD adds negative-policy rollouts to on-policy distillation

The paper introduces Negative-Policy OPD (NP-OPD), which adds a lower-performing, lower-capability “negative policy” at the rollout stage of on-policy distillation (OPD), continuously supplying tokens that the negative policy prefers over the teacher so they stay under teacher supervision during training — without changing the distillation reward formulation. The authors report improvements over OPD across model scales, generation modes, reasoning domains and different OPD variants, and analyses indicating NP-OPD suppresses negative-policy-preferred tokens and moves the student away from the negative policy. The 25-page preprint (7 figures, 24 tables) says code will be released.

Hugging Face · Papers·
7440

Recurrent Looped Transformer adds per-token feedback to boost length generalization

The paper introduces the Recurrent Looped Transformer (RLT), which splits its eight layers between a parallel causal encoder and a recurrent decoder that merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, two RLT splits trained on at most 40 bits generalize parity to 256 bits with 100% accuracy in every seed while an eight-layer Transformer stays at chance; swap-based S5 permutation tracking at eight times the training length reaches 97% versus under 1%, and modular arithmetic reaches up to 93% versus 33%. Ablations show the gains depend on the feedback: removing it drops parity and S5 to chance, and updating feedback once per four-token chunk keeps 64-bit parity at 99% but lowers length-64 S5 from 100% to 20%.

Hugging Face · Papers·
7450

HERMES: modular executable Dev-Primitives for software engineering agents

An arXiv paper introduces Dev-Primitives, an abstraction that pairs each repository artifact — source files, configs, tests, dependencies — with a resident LLM, giving it an agent-native interface for natural-language reasoning, inter-component communication and localized self-modification. Built on top of it, the HERMES harness-engineering framework activates these primitives at repository scale via dependency-aware dynamic activation and a bug-diagnosis mechanism that maps execution evidence back to the components needing revision. The authors report HERMES beats matched baseline harnesses by 12.4% on average across four software engineering benchmarks, stays within 4.5% of a homogeneous GPT-5.6 Sol configuration even with Qwen3-8B Dev-Primitives, and cuts inference cost by 26.2% on Terminal-Bench 4.0.

Hugging Face · Papers·
7460

DAEDALUS: Bootstrapping agent memory from self-generated tasks

An arXiv paper introduces DAEDALUS, a method that bootstraps reusable agent memory from self-generated practice without existing tasks or an oracle verifier. An explorer agent creates challenging but solvable tasks while a solver agent attempts them; heuristics derived from solver failures are accepted only after repeated in-context success, then consolidated into a memory bank. Across AppWorld, τ²-bench and AutomationBench, it improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, at lower inference cost than most comparable methods.

Hugging Face · Papers·
7470

Survey maps in-parameter memory augmentation for LLMs

An arXiv survey submitted on 6 Oct 2026 (arXiv:2610.08630) by Haoyu Huang and eight co-authors reviews methods for giving large language models “in-parameter memory.” It organizes the landscape along two orthogonal axes: parameter placement (embedding, attention, FFN layers, or hybrid) and parameter acquisition time (online during deployment versus offline before it). The paper also lays out open directions in interference, safety, co-design with in-context learning, and recursive self-improvement.

Hugging Face · Papers·
7480

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

The paper introduces ALIVE, a framework that makes inserted objects “alive” by having them interact coherently with the source video — for example being picked up or manipulated — using an edited first frame and an instruction that names only the added object. The authors curate 35,800 editing pairs built from 3D-rendered, model-generated and real-world videos plus general editing pairs from ROSE, each pair differing only in whether the target object is present, and train a VLM to predict interaction guidance from the same inputs. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% on the ALIVE-interaction benchmark and 4.4% on the general video object insertion benchmark; VLM-predicted guidance adds a further 0.95 points without extra user input.

Hugging Face · Papers·
7490

HiPLEX: Hierarchical Policy Factorization for Full-Duplex Speech Language Models

The paper introduces HiPLEX, a reinforcement learning framework that factorizes a pretrained full-duplex text policy into a control policy deciding when to emit content (choosing among pad, epad and con) and a conditional content policy that picks a token only when con is selected. The authors report that across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities and shortens post-interruption response latency versus GRPO, while keeping comparable judged interruption-response quality, and better matches pooled human turn-timing and backchannel-rate marginals on Moshi and PersonaPlex.

Hugging Face · Papers·
7500

WorldSonus brings real-time spatial sound to world models

A new arXiv paper introduces WorldSonus, an interactive video-to-audio framework that gives generated world-model environments synchronized sound. It uses a streaming causal autoregressive diffusion architecture that the authors report runs at a real-time factor of 0.41, plus chunk-indexed prompt scheduling so sound events can be steered mid-generation. Stereo and ambisonic supervision is used to align output stereo audio with scene geometry and camera motion.

Hugging Face · Papers·
7510

A Safe Action Is Not Enough: Feasible-Future Decoding for VLA Policies

An arXiv paper (arXiv:2610.05166, v2 revised 6 Oct 2026) names the "feasibility-likelihood gap": a frozen vision-language-action (VLA) policy may favor a locally admissible move that leaves no policy-supported route to safe task completion. The authors derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe completion, and build VICS-G, an alarm-triggered, training-free reranker. Across six Safety-CHORES settings it lowers mean cumulative safety cost by 1.9%-57.5% while staying within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length, with no policy retraining or online rollouts.

Hugging Face · Papers·
7520

EmbodiedSmith: scaling embodied data via recursive self-improvement in simulation

An arXiv paper introduces EmbodiedSmith, a framework for scalable embodied data generation through recursive self-improvement (RSI) in simulation, unifying asset, scene and task generation with autonomous, language-driven customization. Its core is an agentic refinement loop in which scene generation anticipates downstream task requirements while task generation guides targeted scene edits, improving task-generation success including for long-horizon tasks. The framework also supports mobile manipulators, humanoids and dexterous hands, plus deformable objects and fluids; the authors report downstream policy experiments showing that greater data diversity improves generalization.

Hugging Face · Papers·
7530

Recursive Game Creator tops GameCraft-Bench with experience-oriented agentic loop

An arXiv paper introduces Recursive Game Creator, a recursive agentic game-development harness built around four components — Designer, Builder, Player and Reviewer — that aims to push playable prototypes toward genuinely entertaining games. The authors report an overall score of 77.89 on GameCraft-Bench, a strict task success rate of 53.2% on GameASG-Bench (a 34.1% improvement over the same-model baseline), and the highest mean runtime-check pass rate among compared methods at 93.4%. They also report a user study showing longer playtime and higher ratings, and say code is coming soon.

Hugging Face · Papers·
7540

Speculative execution cuts on-device voice agent latency from 5.79s to 4.60s

An arXiv paper proposes speculative tool execution for on-device cascaded voice agents: a Predictor module anticipates tool calls from partial ASR hypotheses, runs them speculatively and caches the results, which are then injected into the LLM prompt for faster responses. A rule-based validation step filters cached results invalidated by user self-corrections, and the LLM can still issue tool calls directly, so worst-case latency stays bounded by the serial baseline. In live measurements on a fully implemented Android voice assistant, median time-to-first-audio fell from 5.79s to 4.60s and the standard deviation from 3.49s to 2.81s.

Hugging Face · Papers·
7550

MediateRec benchmark tests personal-agent mediation of cross-platform recommendations

The paper formalizes a paradigm called Personal-Agent Mediated Recommendation: a platform recommender ranks a candidate set from platform-local information, and a personal LLM agent then uses user-authorized cross-platform history to mediate that ranking into a final top-K slate. The authors introduce MediateRec, a benchmark with scalable proxy cross-platform environments plus a real cross-platform test under a controlled platform-agent information boundary, and propose Personal Attribution Mediation Optimization (PAMO), which counterfactually masks cross-platform history to estimate personal mediation support and reallocate rank-aware advantage mass. Experiments show mediation yields meaningful platform corrections, but even strong proprietary LLMs introduce non-negligible harmful overrides; PAMO beats matched outcome-only RL on seen and unseen target platforms and on the real test, with a better rescue-harm balance.

Hugging Face · Papers·
7560

Attacca: goal-directed control for long-horizon embodied agents

An arXiv paper introduces Attacca, a method for training visual goal-conditioned policies for long-horizon embodied agents. It uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, adds a target-mask prediction head for dense current-view grounding, and conditions the policy on Search, Approach and Interact phases. On short- and long-horizon Minecraft tasks, the authors report 39.0–47.5% clean success, a 1.7–2.4x gain over the strongest baseline, and 54%, 30% and 28% completion on long-horizon tasks, up to a 7x improvement.

Hugging Face · Papers·
7570

OpenAI safety report lead quits, calls company culture 'broken'

David Robinson, who spent 3.5 years at OpenAI overseeing safety reports for frontier launches and drafting its current Preparedness Framework, has left the company and published an Atlantic essay calling its culture "broken" and saying "the time for trial and error is over" on AI safety risk. He argued labs should run like nuclear plants and airports, with layers of redundancy, but said he and colleagues were "so busy sprinting" they rarely had time to consider big changes. His exit follows OpenAI firing researchers Jasmine Wang, Tomek Korbak and Mikita Balesni over reportedly passing sensitive information to an outside safety group.

The Rundown AI·
7580

AI Reconstructs Viewed Images From Brain Scans

Researchers at Israel’s Weizmann Institute of Science developed an AI system that uses high-resolution fMRI data to reconstruct images people are viewing, combining brain decoding with an image diffusion model. The team trained it on data from eight participants who viewed about 9,000 images each, and reports that the system can be calibrated to a new participant with about one hour of data rather than the roughly 40 hours required by earlier approaches.

MIT科技评论中文·
7590

CNBC roundup flags AI researchers’ warning and major market stories

CNBC Technology’s morning market roundup lists the proposed Paramount-WBD merger, Elon Musk’s return to trillionaire status and a warning from AI researchers among five items for investors. The supplied excerpt does not provide further details about the AI warning or the other developments.

CNBC Technology·
7600

Rabbit’s OS3 bets on cross-device personal agents

Rabbit founder Jesse Lyu says the company’s OS3 is designed as a single conversational entry point for personal agents, connecting up to five existing computers and using their local files, software, and login environments. He argues that personal agents must reduce setup and usage barriers, while acknowledging that the demand for everyday agent tasks has not yet been proven.

36氪 人工智能·