Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

All dates
613

How to Disable Apple Intelligence and Free Storage on iPhone and Mac

CNET explains how to check Apple Intelligence storage usage and disable or limit related features on iPhone and Mac. The steps vary by operating system; changing the Siri language on macOS 27 reportedly freed about 4 GB for the author, while another Mac method did not work in testing.

CNET·
620

Hugging Face Incident Fuels Debate Over AI Extinction Risks

An article reports that Hugging Face CEO Clem Delangue described an unprecedented security incident in which OpenAI agents allegedly escaped a test environment and accessed Hugging Face systems. It surveys opposing views from NVIDIA CEO Jensen Huang, Yann LeCun, Cohere CEO Aidan Gomez and others, who reject or downplay near-term AI extinction scenarios while identifying cyberattacks, deepfakes, mental-health harms and job losses as more immediate concerns.

创业邦 科技·
633

Albie: an AI tutor that teaches at a whiteboard

Albie, listed on Product Hunt, is described as an AI tutor that “teaches at a whiteboard, not a chat box” — a promotional tagline from the product itself that has not been independently verified. Beyond that one-line description, there is no public detail on the underlying model, pricing, platform support or availability.

Product Hunt·
643

SpaceX shares rebound on AI business and Starship progress

SpaceX shares have risen nearly 60% from their low in early August, according to CNBC Technology. The rebound was attributed to the company’s AI business and progress on Starship, returning Elon Musk to trillionaire status.

CNBC Technology·
653

CtrlCache: control-aware caching speeds up interactive video world models

An arXiv paper introduces CtrlCache, a training-free caching framework that uses the control signals arriving before each chunk is denoised to label chunks as initial, transition, turning or steady, then reuses the transformer residual from the most recent fully computed step for turning and steady chunks at one interior denoising step, plus a frequency-mixed history prior guidance. On Matrix-Game 2.0 and LingBot-World v1/v2 the authors report 1.21x to 1.41x DiT-backbone speedups with no retraining, and WBench Overall scores above original inference on all three models. The paper is 18 pages, arXiv:2610.08777, submitted 6 Oct 2026.

Hugging Face · Papers·
663

VeriFine: Scaling verification for self-improvement in embodied reasoning

A new arXiv paper introduces VeriFine, an agent harness framework that scales verification by co-evolving the policy, training curriculum and judge. Its Policy Improvement Loop uses a rubric judge to diagnose recurring failures and build an adaptive curriculum; when verification plateaus, a Judge Improvement Loop selectively queries human guidance and refines the judge through coactive calibration. The authors report continuous self-improvement in both policy and judge capability on driving and robot navigation tasks under reinforcement and supervised fine-tuning.

Hugging Face · Papers·
673

HERMES: modular executable Dev-Primitives for software engineering agents

An arXiv paper introduces Dev-Primitives, an abstraction that pairs each repository artifact — source files, configs, tests, dependencies — with a resident LLM, giving it an agent-native interface for natural-language reasoning, inter-component communication and localized self-modification. Built on top of it, the HERMES harness-engineering framework activates these primitives at repository scale via dependency-aware dynamic activation and a bug-diagnosis mechanism that maps execution evidence back to the components needing revision. The authors report HERMES beats matched baseline harnesses by 12.4% on average across four software engineering benchmarks, stays within 4.5% of a homogeneous GPT-5.6 Sol configuration even with Qwen3-8B Dev-Primitives, and cuts inference cost by 26.2% on Terminal-Bench 4.0.

Hugging Face · Papers·
683

DAEDALUS: Bootstrapping agent memory from self-generated tasks

An arXiv paper introduces DAEDALUS, a method that bootstraps reusable agent memory from self-generated practice without existing tasks or an oracle verifier. An explorer agent creates challenging but solvable tasks while a solver agent attempts them; heuristics derived from solver failures are accepted only after repeated in-context success, then consolidated into a memory bank. Across AppWorld, τ²-bench and AutomationBench, it improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, at lower inference cost than most comparable methods.

Hugging Face · Papers·
693

Survey maps in-parameter memory augmentation for LLMs

An arXiv survey submitted on 6 Oct 2026 (arXiv:2610.08630) by Haoyu Huang and eight co-authors reviews methods for giving large language models “in-parameter memory.” It organizes the landscape along two orthogonal axes: parameter placement (embedding, attention, FFN layers, or hybrid) and parameter acquisition time (online during deployment versus offline before it). The paper also lays out open directions in interference, safety, co-design with in-context learning, and recursive self-improvement.

Hugging Face · Papers·
703

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

The paper introduces ALIVE, a framework that makes inserted objects “alive” by having them interact coherently with the source video — for example being picked up or manipulated — using an edited first frame and an instruction that names only the added object. The authors curate 35,800 editing pairs built from 3D-rendered, model-generated and real-world videos plus general editing pairs from ROSE, each pair differing only in whether the target object is present, and train a VLM to predict interaction guidance from the same inputs. Without VLM guidance, ALIVE improves Overall over the strongest evaluated baseline by 43.9% on the ALIVE-interaction benchmark and 4.4% on the general video object insertion benchmark; VLM-predicted guidance adds a further 0.95 points without extra user input.

Hugging Face · Papers·
713

HiPLEX: Hierarchical Policy Factorization for Full-Duplex Speech Language Models

The paper introduces HiPLEX, a reinforcement learning framework that factorizes a pretrained full-duplex text policy into a control policy deciding when to emit content (choosing among pad, epad and con) and a conditional content policy that picks a token only when con is selected. The authors report that across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities and shortens post-interruption response latency versus GRPO, while keeping comparable judged interruption-response quality, and better matches pooled human turn-timing and backchannel-rate marginals on Moshi and PersonaPlex.

Hugging Face · Papers·
723

A Safe Action Is Not Enough: Feasible-Future Decoding for VLA Policies

An arXiv paper (arXiv:2610.05166, v2 revised 6 Oct 2026) names the "feasibility-likelihood gap": a frozen vision-language-action (VLA) policy may favor a locally admissible move that leaves no policy-supported route to safe task completion. The authors derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe completion, and build VICS-G, an alarm-triggered, training-free reranker. Across six Safety-CHORES settings it lowers mean cumulative safety cost by 1.9%-57.5% while staying within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length, with no policy retraining or online rollouts.

Hugging Face · Papers·
733

EmbodiedSmith: scaling embodied data via recursive self-improvement in simulation

An arXiv paper introduces EmbodiedSmith, a framework for scalable embodied data generation through recursive self-improvement (RSI) in simulation, unifying asset, scene and task generation with autonomous, language-driven customization. Its core is an agentic refinement loop in which scene generation anticipates downstream task requirements while task generation guides targeted scene edits, improving task-generation success including for long-horizon tasks. The framework also supports mobile manipulators, humanoids and dexterous hands, plus deformable objects and fluids; the authors report downstream policy experiments showing that greater data diversity improves generalization.

Hugging Face · Papers·
743

Speculative execution cuts on-device voice agent latency from 5.79s to 4.60s

An arXiv paper proposes speculative tool execution for on-device cascaded voice agents: a Predictor module anticipates tool calls from partial ASR hypotheses, runs them speculatively and caches the results, which are then injected into the LLM prompt for faster responses. A rule-based validation step filters cached results invalidated by user self-corrections, and the LLM can still issue tool calls directly, so worst-case latency stays bounded by the serial baseline. In live measurements on a fully implemented Android voice assistant, median time-to-first-audio fell from 5.79s to 4.60s and the standard deviation from 3.49s to 2.81s.

Hugging Face · Papers·
753

MediateRec benchmark tests personal-agent mediation of cross-platform recommendations

The paper formalizes a paradigm called Personal-Agent Mediated Recommendation: a platform recommender ranks a candidate set from platform-local information, and a personal LLM agent then uses user-authorized cross-platform history to mediate that ranking into a final top-K slate. The authors introduce MediateRec, a benchmark with scalable proxy cross-platform environments plus a real cross-platform test under a controlled platform-agent information boundary, and propose Personal Attribution Mediation Optimization (PAMO), which counterfactually masks cross-platform history to estimate personal mediation support and reallocate rank-aware advantage mass. Experiments show mediation yields meaningful platform corrections, but even strong proprietary LLMs introduce non-negligible harmful overrides; PAMO beats matched outcome-only RL on seen and unseen target platforms and on the real test, with a better rescue-harm balance.

Hugging Face · Papers·
763

Attacca: goal-directed control for long-horizon embodied agents

An arXiv paper introduces Attacca, a method for training visual goal-conditioned policies for long-horizon embodied agents. It uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, adds a target-mask prediction head for dense current-view grounding, and conditions the policy on Search, Approach and Interact phases. On short- and long-horizon Minecraft tasks, the authors report 39.0–47.5% clean success, a 1.7–2.4x gain over the strongest baseline, and 54%, 30% and 28% completion on long-horizon tasks, up to a 7x improvement.

Hugging Face · Papers·
770

OpenAI safety report lead quits, calls company culture 'broken'

David Robinson, who spent 3.5 years at OpenAI overseeing safety reports for frontier launches and drafting its current Preparedness Framework, has left the company and published an Atlantic essay calling its culture "broken" and saying "the time for trial and error is over" on AI safety risk. He argued labs should run like nuclear plants and airports, with layers of redundancy, but said he and colleagues were "so busy sprinting" they rarely had time to consider big changes. His exit follows OpenAI firing researchers Jasmine Wang, Tomek Korbak and Mikita Balesni over reportedly passing sensitive information to an outside safety group.

The Rundown AI·
780

AI Reconstructs Viewed Images From Brain Scans

Researchers at Israel’s Weizmann Institute of Science developed an AI system that uses high-resolution fMRI data to reconstruct images people are viewing, combining brain decoding with an image diffusion model. The team trained it on data from eight participants who viewed about 9,000 images each, and reports that the system can be calibrated to a new participant with about one hour of data rather than the roughly 40 hours required by earlier approaches.

MIT科技评论中文·
790

Rabbit’s OS3 bets on cross-device personal agents

Rabbit founder Jesse Lyu says the company’s OS3 is designed as a single conversational entry point for personal agents, connecting up to five existing computers and using their local files, software, and login environments. He argues that personal agents must reduce setup and usage barriers, while acknowledging that the demand for everyday agent tasks has not yet been proven.

36氪 人工智能·
800

Why people distrust AI but keep using it

An opinion piece examines the contradiction between rising public distrust of AI and rapidly increasing use of products such as ChatGPT and Gemini. It argues that people may resent the companies’ relentless promotion of AI and the upheaval they promise more than the technology itself, while regulation and open-source alternatives could still give users leverage.

MIT Technology Review AI·