Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

4413

PersonTTS: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

An arXiv preprint introduces PersonTTS, an amortized agentic policy-discovery framework that reframes test-time scaling (TTS) as finding controllers that maximize joint satisfaction of a user's accuracy, latency and inference-cost requirements, rather than optimizing one resource dimension at a time. It reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while still evaluating each candidate against the target profile. On AIME and HMMT, the authors report that PersonTTS substantially outperforms strong TTS baselines on joint requirement satisfaction for unseen user profiles and held-out problems, and that cross-user experience reuse improves policy quality while cutting discovery-agent time and cost under the same evaluation budget.

Hugging Face · Papers·
4423

RunningTab adds environment-side tabs to agent workspace interaction

A new arXiv paper introduces RunningTab, a framework that equips direct workspace interaction (DWI) with an environment-side tab. In DWI, an LLM agent searches and reads any file in a workspace from a terminal without indexing, but the context window keeps no trace of what the task asked for, what was read, or what was listed but never opened; RunningTab has the agent add its requirements while the environment records each read file as an excerpt with provenance and each listed-but-unopened file as a candidate. The agent can then match each requirement against its best excerpts and top candidates, resolve it or set it aside with a reason, and receive a finish check if it tries to complete with requirements still open. The authors report that across three benchmarks and three LLMs, RunningTab consistently outperforms plain DWI and baselines that keep the record inside the model.

Hugging Face · Papers·
4433

Study probes hybrid attention mechanics, proposes Sliding-Window Linear Attention

An arXiv paper submitted on 7 Oct 2026, "Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position" by Xiaoran Liu, Ziwei He and Xipeng Qiu, compares hybrids of full attention with sliding-window attention (SWA) or gated linear attention variants (GLA, GDN). It reports a "Seesaw Effect": LA hybrids benefit more from long-context continual pretraining, while SWA hybrids do better under length extrapolation, and describes a Short-Context Learning Trap, Short-Window Weariness and Long-Window Laziness in SWA hybrids. The authors propose Sliding-Window Linear Attention, claiming 16x training-free length extrapolation while maintaining 100% accuracy on NIAH-SK1 at 64k context; the 60-page paper is under review.

Hugging Face · Papers·
4443

Tetris3D generates physically coherent 3D scenes from a single image

An arXiv paper introduces Tetris3D, a generative framework for single-image 3D scene reconstruction that explicitly conditions each object's generation on the geometry and physical relationships of surrounding objects, keeping shape and pose geometrically and physically plausible within the scene. The authors report that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art results in generation quality and physical stability. The work also releases ComOb, a physics-simulation-based dataset of 1.2M scenes with per-object meshes and pairwise physical relation annotations.

Hugging Face · Papers·
4453

AdSpark: A 300K-sample dataset and benchmark for product ad video generation

The paper introduces AdSpark, a large-scale dataset and benchmark for product-centric advertisement video generation. AdSpark-300K contains roughly 300K reference image–prompt–video triplets drawn from a major e-commerce platform, split into real and synthetic subsets, with structured ad annotations such as product identity, selling-point descriptions, creative plans and aligned audio scripts. The authors also propose AdSpark-Bench, which scores generated ads on six dimensions including visual quality, product fidelity, instruction adherence, temporal coherence, audio alignment and advertising effectiveness; evaluations reveal remaining challenges in product preservation, multi-shot storytelling and selling-point visualization, and the dataset is to be released upon acceptance.

Hugging Face · Papers·
4463

SkillForge co-evolves agent skills and policy via a fitness-driven skill lifecycle

A NeurIPS 2026-accepted arXiv paper introduces SkillForge, an agentic RL method that evolves a skill library through a fitness-driven lifecycle of trial, active, stable and retired states, so skills and the model co-evolve during training. A pre-RL phase uses the base model's own rollouts to pre-retire low-fitness skills, seeding supervised fine-tuning; RL then continues with selective retirement, stabilization and LLM-guided mutation each iteration. The authors report the highest aggregate success rate across several interactive agent benchmarks, up to 7.8% relative improvement over the strongest baseline while keeping the library compact, and release SkillFurnace, a 5k+ record annotated dataset of retirement-filtered SFT trajectories, evolved skill libraries and retirement events.

Hugging Face · Papers·
4473

QuadTok: quadtree visual tokenizer cuts tokens ~10% for autoregressive image generation

The paper introduces QuadTok, a hierarchical quadtree framework for visual tokenization that bridges 2D spatial binding and 1D sequence flexibility, allocating representational capacity to visually intricate areas while leaving homogeneous regions at coarse resolution. Its ImageNet-trained tokenizer reportedly saves about 10% of tokens on ImageNet versus a fixed 256-token grid, and about 9% when transferred zero-shot to COCO, while maintaining comparable reconstruction fidelity. Conditioned on a quadtree topology supplied before generation, a 947M GPT-style generative model reaches 2.08 gFID on ImageNet 256×256, and the preserved spatial correlation also enables zero-shot spatially controlled image generation.

Hugging Face · Papers·
4483

WebFovea: harness fixes lift web agent score from 31 to 57

WebFovea, a vision-based web agent, placed second in the WebRetriever Challenge 2026 with an official hidden-set score of 57.0 out of 100 on Protocol III of the WebRetriever benchmark. Its technical report argues that many failures on live sites occurred in the harness between the model and the page rather than in the model's reasoning: a coordinate-space mismatch placed every click at 3/4 of its intended coordinates, actions on native dropdowns, inside iframes and in text boxes failed silently, and self-generated chat-template tokens contaminated 4.9% of task episodes. Using the same model across all four submissions, the author says the score rose from 31.0 to 57.0 through harness changes, alongside run-to-run variance on live sites; code is available.

Hugging Face · Papers·
4493

Supermicro summit: three takeaways on AI storage strategy

SiliconANGLE rounded up three insights from theCUBE's interview series at the Supermicro Open Storage Summit, where the theme was enterprises moving from successful AI experiments to systems that must deliver dependable business results. The coverage argues that how organizations store, serve and manage data increasingly shapes the performance, cost and practicality of AI deployments, and that the challenge goes beyond buying faster hardware to rethinking architecture as workloads evolve.

SiliconANGLE AI·
4506

ChatPlayground AI lifetime deal bundles 25+ models for $59.97

An affiliate-linked TechRepublic article promotes ChatPlayground AI's "Unlimited" lifetime plan, saying a one-time $59.97 payment (listed at a regular $619) gives access to more than 25 models — including GPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro — through a single web interface. The post says the platform offers side-by-side comparison of up to six models, an in-browser AI sidebar, and a document editor that can query PDFs, YouTube videos and page elements. These are promotional claims, and StackSocial notes prices are subject to change.

TechRepublic·
4513

Building a context-aware AI assistant with OpenClaw on AWS AgentCore

An AWS Machine Learning Blog tutorial shows how to run the open-source agentic system OpenClaw on Amazon Bedrock AgentCore runtime and give it continuity with AgentCore memory, using a gardening assistant called Sprout as the example; the whole system ships in a single CloudFormation template that deploys with one command. On each message the agent retrieves long-term memory records from a per-user namespace, ranks explicit preferences ahead of inferred facts, and injects them into the system prompt, degrading gracefully to a memory-free answer if retrieval times out within a 3-second budget or errors. Tasks are routed to different models — Claude Haiku 4.5 for text and Claude Sonnet 4.5 for image diagnosis — with Telegram as the front door. The author estimates roughly $1–2 per month for light personal use versus about $35 per month for an always-on EC2 instance.

AWS Machine Learning Blog·
4523

MasterClass bets on AI teaching agents to cut tutoring costs

MasterClass Chief Product Officer Mandar Bapaye said at CoreWeave's Fully Connected event that MasterClass Executive, the company's AI-native business program, uses a multi-agent system to plan lessons around each learner, watching for signals such as cognitive overload and fading motivation and then changing its approach. Roughly 10 agents run behind every learner interaction, he said, with inputs, outputs, tool calls and inter-agent communication tracked. MasterClass has selected CoreWeave's W&B Weave to trace and monitor those teaching agents, and built its own agent that reviews traces nightly to flag issues and likely root causes. He said the first cohort drew 30,000 applications for about 500 spots, with the second nearing 50,000 applications.

SiliconANGLE AI·
4533

WIRED Reviews ‘Artificial,’ a Dark Comedy About OpenAI and AI Risk

WIRED reviews Luca Guadagnino’s film *Artificial*, a black comedy-drama about OpenAI, Ilya Sutskever, Sam Altman and the AI race. The review says the movie portrays the industry’s leaders as vain and reckless while dramatizing conflicts over AI safety, rapid growth and control of the technology.

WIRED AI·
4540

Terence Tao Warns of “Proof Indigestion” as AI Mathematics Accelerates

A report says Fields Medalist Terence Tao has urged AI companies to slow down their pursuit of mathematical breakthroughs, arguing that machine-checked proofs are advancing faster than human interpretation, peer review and textbook integration. It also reports that OpenAI has formed an independent Mathematics and AI Advisory Group at the Institute for Advanced Study, while stating that the group will not advise on OpenAI’s internal mathematical progress.

量子位(原生 RSS)·
4553

Claude Code’s Suggested Messages May Serve the Model First

A blog post discusses Claude Code’s suggested message feature and argues that the model, rather than the human user, may be its real customer. The item was surfaced through comments on Hacker News.

Hacker News · AI(100+ 分)·
4560

OpenAI’s 28-Day Push Starts With GPT-6 Speed Claims and User Skepticism

OpenAI reportedly began a 28-day Codex and Work improvement push by increasing the default inference speed of GPT-6 Astra and GPT-6.1 Sol by about 50%, according to Tibo and coverage of the announcement. The report says the rollout drew skepticism because user tests allegedly fell short of the claimed 50 TPS and coincided with reports of ChatGPT visual ads, EU text watermarking, and changes to subscription value.

量子位(原生 RSS)·
4570

Sony Music Seeks Removal of 260,000 AI Impersonation Tracks

Sony Music Entertainment reportedly asked streaming platforms to remove more than 260,000 AI-generated tracks impersonating its artists by the end of September, nearly double the 135,000 requests made by the end of March. The company said the unauthorized voice and likeness imitations harm artists and mislead fans; Deezer separately reported that AI-generated songs now account for more than half of new uploads on its platform.

IT之家 AI·
4580

South Korea Plans 4.7T Won Frontier AI Model Initiative

South Korea plans to launch a 4.7 trillion won initiative to develop frontier AI models from March 2027, pending approval of the 2027 budget by the National Assembly. The government plans to combine public equity investment with private capital and select a lead developer through competitive bidding.

IT之家 AI·
4593

Vast pitches tiered storage to ease AI agent memory pressure

In an interview with theCUBE at CoreWeave's Fully Connected 2026 event, Vast Data co-founder and CTO Alon Horev said agent memory differs from ordinary inference: it includes both in-session context and long-term memory that lets an agent review past interactions. Vast's approach tiers storage — GPU memory first, then CPU memory on the same machine, then persistent media holding petabytes of KV cache — with Nvidia's Dynamo software orchestrating the process. Horev said a 500,000-token session can occupy one-tenth to one-twentieth of a GPU's memory, and offloading such sessions to storage avoids repeat recalculation; enterprises also need to record and retain everything their agents do, data that can feed fine-tuning or purpose-built models.

SiliconANGLE AI·
4603

AgentGuard scans AI agent Skills for risks before you install them

A Product Hunt listing introduces AgentGuard, described as a tool that scans AI agent Skills for risks before installation. Beyond that one-line pitch, no technical details, supported platforms or pricing are given.

Product Hunt·