Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

AI News

4613

SkillForge co-evolves agent skills and policy via a fitness-driven skill lifecycle

A NeurIPS 2026-accepted arXiv paper introduces SkillForge, an agentic RL method that evolves a skill library through a fitness-driven lifecycle of trial, active, stable and retired states, so skills and the model co-evolve during training. A pre-RL phase uses the base model's own rollouts to pre-retire low-fitness skills, seeding supervised fine-tuning; RL then continues with selective retirement, stabilization and LLM-guided mutation each iteration. The authors report the highest aggregate success rate across several interactive agent benchmarks, up to 7.8% relative improvement over the strongest baseline while keeping the library compact, and release SkillFurnace, a 5k+ record annotated dataset of retirement-filtered SFT trajectories, evolved skill libraries and retirement events.

Hugging Face · Papers·
4623

QuadTok: quadtree visual tokenizer cuts tokens ~10% for autoregressive image generation

The paper introduces QuadTok, a hierarchical quadtree framework for visual tokenization that bridges 2D spatial binding and 1D sequence flexibility, allocating representational capacity to visually intricate areas while leaving homogeneous regions at coarse resolution. Its ImageNet-trained tokenizer reportedly saves about 10% of tokens on ImageNet versus a fixed 256-token grid, and about 9% when transferred zero-shot to COCO, while maintaining comparable reconstruction fidelity. Conditioned on a quadtree topology supplied before generation, a 947M GPT-style generative model reaches 2.08 gFID on ImageNet 256×256, and the preserved spatial correlation also enables zero-shot spatially controlled image generation.

Hugging Face · Papers·
4633

WebFovea: harness fixes lift web agent score from 31 to 57

WebFovea, a vision-based web agent, placed second in the WebRetriever Challenge 2026 with an official hidden-set score of 57.0 out of 100 on Protocol III of the WebRetriever benchmark. Its technical report argues that many failures on live sites occurred in the harness between the model and the page rather than in the model's reasoning: a coordinate-space mismatch placed every click at 3/4 of its intended coordinates, actions on native dropdowns, inside iframes and in text boxes failed silently, and self-generated chat-template tokens contaminated 4.9% of task episodes. Using the same model across all four submissions, the author says the score rose from 31.0 to 57.0 through harness changes, alongside run-to-run variance on live sites; code is available.

Hugging Face · Papers·
4643

Supermicro summit: three takeaways on AI storage strategy

SiliconANGLE rounded up three insights from theCUBE's interview series at the Supermicro Open Storage Summit, where the theme was enterprises moving from successful AI experiments to systems that must deliver dependable business results. The coverage argues that how organizations store, serve and manage data increasingly shapes the performance, cost and practicality of AI deployments, and that the challenge goes beyond buying faster hardware to rethinking architecture as workloads evolve.

SiliconANGLE AI·
4655

ChatPlayground AI lifetime deal bundles 25+ models for $59.97

An affiliate-linked TechRepublic article promotes ChatPlayground AI's "Unlimited" lifetime plan, saying a one-time $59.97 payment (listed at a regular $619) gives access to more than 25 models — including GPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro — through a single web interface. The post says the platform offers side-by-side comparison of up to six models, an in-browser AI sidebar, and a document editor that can query PDFs, YouTube videos and page elements. These are promotional claims, and StackSocial notes prices are subject to change.

TechRepublic·
4663

Building a context-aware AI assistant with OpenClaw on AWS AgentCore

An AWS Machine Learning Blog tutorial shows how to run the open-source agentic system OpenClaw on Amazon Bedrock AgentCore runtime and give it continuity with AgentCore memory, using a gardening assistant called Sprout as the example; the whole system ships in a single CloudFormation template that deploys with one command. On each message the agent retrieves long-term memory records from a per-user namespace, ranks explicit preferences ahead of inferred facts, and injects them into the system prompt, degrading gracefully to a memory-free answer if retrieval times out within a 3-second budget or errors. Tasks are routed to different models — Claude Haiku 4.5 for text and Claude Sonnet 4.5 for image diagnosis — with Telegram as the front door. The author estimates roughly $1–2 per month for light personal use versus about $35 per month for an always-on EC2 instance.

AWS Machine Learning Blog·
4673

MasterClass bets on AI teaching agents to cut tutoring costs

MasterClass Chief Product Officer Mandar Bapaye said at CoreWeave's Fully Connected event that MasterClass Executive, the company's AI-native business program, uses a multi-agent system to plan lessons around each learner, watching for signals such as cognitive overload and fading motivation and then changing its approach. Roughly 10 agents run behind every learner interaction, he said, with inputs, outputs, tool calls and inter-agent communication tracked. MasterClass has selected CoreWeave's W&B Weave to trace and monitor those teaching agents, and built its own agent that reviews traces nightly to flag issues and likely root causes. He said the first cohort drew 30,000 applications for about 500 spots, with the second nearing 50,000 applications.

SiliconANGLE AI·
4683

WIRED Reviews ‘Artificial,’ a Dark Comedy About OpenAI and AI Risk

WIRED reviews Luca Guadagnino’s film *Artificial*, a black comedy-drama about OpenAI, Ilya Sutskever, Sam Altman and the AI race. The review says the movie portrays the industry’s leaders as vain and reckless while dramatizing conflicts over AI safety, rapid growth and control of the technology.

WIRED AI·
4690

Terence Tao Warns of “Proof Indigestion” as AI Mathematics Accelerates

A report says Fields Medalist Terence Tao has urged AI companies to slow down their pursuit of mathematical breakthroughs, arguing that machine-checked proofs are advancing faster than human interpretation, peer review and textbook integration. It also reports that OpenAI has formed an independent Mathematics and AI Advisory Group at the Institute for Advanced Study, while stating that the group will not advise on OpenAI’s internal mathematical progress.

量子位(原生 RSS)·
4703

Claude Code’s Suggested Messages May Serve the Model First

A blog post discusses Claude Code’s suggested message feature and argues that the model, rather than the human user, may be its real customer. The item was surfaced through comments on Hacker News.

Hacker News · AI(100+ 分)·
4710

OpenAI’s 28-Day Push Starts With GPT-6 Speed Claims and User Skepticism

OpenAI reportedly began a 28-day Codex and Work improvement push by increasing the default inference speed of GPT-6 Astra and GPT-6.1 Sol by about 50%, according to Tibo and coverage of the announcement. The report says the rollout drew skepticism because user tests allegedly fell short of the claimed 50 TPS and coincided with reports of ChatGPT visual ads, EU text watermarking, and changes to subscription value.

量子位(原生 RSS)·
4720

Sony Music Seeks Removal of 260,000 AI Impersonation Tracks

Sony Music Entertainment reportedly asked streaming platforms to remove more than 260,000 AI-generated tracks impersonating its artists by the end of September, nearly double the 135,000 requests made by the end of March. The company said the unauthorized voice and likeness imitations harm artists and mislead fans; Deezer separately reported that AI-generated songs now account for more than half of new uploads on its platform.

IT之家 AI·
4730

South Korea Plans 4.7T Won Frontier AI Model Initiative

South Korea plans to launch a 4.7 trillion won initiative to develop frontier AI models from March 2027, pending approval of the 2027 budget by the National Assembly. The government plans to combine public equity investment with private capital and select a lead developer through competitive bidding.

IT之家 AI·
4743

Vast pitches tiered storage to ease AI agent memory pressure

In an interview with theCUBE at CoreWeave's Fully Connected 2026 event, Vast Data co-founder and CTO Alon Horev said agent memory differs from ordinary inference: it includes both in-session context and long-term memory that lets an agent review past interactions. Vast's approach tiers storage — GPU memory first, then CPU memory on the same machine, then persistent media holding petabytes of KV cache — with Nvidia's Dynamo software orchestrating the process. Horev said a 500,000-token session can occupy one-tenth to one-twentieth of a GPU's memory, and offloading such sessions to storage avoids repeat recalculation; enterprises also need to record and retain everything their agents do, data that can feed fine-tuning or purpose-built models.

SiliconANGLE AI·
4753

AgentGuard scans AI agent Skills for risks before you install them

A Product Hunt listing introduces AgentGuard, described as a tool that scans AI agent Skills for risks before installation. Beyond that one-line pitch, no technical details, supported platforms or pricing are given.

Product Hunt·
4763

Giorgia Meloni Seeks Voice Trademark to Fight AI Deepfakes

Italian Prime Minister Giorgia Meloni has applied to register her voice as a trademark with the European Union Intellectual Property Office, citing the need to combat AI-generated deepfakes. The application includes a four-second Italian recording and remains under review; a trademark would add legal obstacles but would not fully prevent voice cloning.

IT之家 AI·
4773

Mirror Particle pitches a world model for changing human behavior

Mirror Particle is developing a foundation model intended to simulate how human behavior changes over time, using visual, social and other data alongside language. The San Francisco startup says its system focuses on revealed behavior and is initially targeting market research and brand strategy, while reporting that it has raised an angel round and is close to closing its first venture round.

TechCrunch AI·
4780

Why US communities are resisting AI data centers

The article examines growing US community opposition to AI data centers, citing concerns over electricity prices, water use, noise, limited long-term employment and strained local infrastructure. It argues that fragmented power markets, lengthy permitting processes and the rush to expand AI capacity are deepening tensions between technology companies, governments and residents.

36氪 人工智能·
4793

Microsoft Research podcast: what AI evaluation gets wrong

In a Microsoft Research podcast episode, host Chad Atalla speaks with Jennifer Neville, who leads the AI Interaction and Learning team at Microsoft Research and is a professor at Purdue, about how evaluation pushes the performance boundaries of today's AI systems and about the “surprising failures” that appear when models are tested beyond traditional benchmarks. Neville argues that the benchmarks commonly used in ML/AI are fairly simple relative to real-world use, so her team designs evaluations around multiturn behavior, collaborative settings and long-horizon tasks to expose performance gaps and then drive algorithmic and model improvements. She also offers practical guidance for working with current AI systems and stresses examining the data closely when results defy expectations.

Microsoft Research Blog·
4803

AWS explains ISO/IEC 42005:2025 AI impact assessment guidance

AWS published a post on how organizations can use the ISO/IEC 42005:2025 standard to run AI system impact assessments and fold them into existing enterprise risk, privacy and security reviews. It covers the standard's guidance on when to assess, lightweight triage, required documentation and reassessment triggers, plus Annex D for integrating existing assessments without duplication and Annex E for a standalone template. AWS also points to its Well-Architected Responsible AI Lens and its ISO/IEC 42001 implementation guide, and says Amazon Bedrock, Amazon Q Business, Amazon Transcribe and Amazon Textract hold ISO/IEC 42001 certification.

AWS Machine Learning Blog·