Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

AI News

All dates
0118

Cua releases Cua Spaces 0.1.0, giving agents full desktop machines

The open-source trycua/cua project has shipped Cua Spaces 0.1.0, a macOS app that gives agents full macOS or Linux desktops running on your Mac, other machines you own, or your own cloud account (AWS, Google Cloud or Modal), with shared human-plus-agent desktops, teleporting of signed-in apps, and sessions kept in a local encrypted Cua Keyvault. The repository also ships Cua Driver for driving native desktop apps and browsers across macOS, Windows and Linux via CLI, MCP or typed SDKs, Lume for local Apple Silicon VMs, the small specialized CUA-S1 'System 1' decision models whose weights are hosted on Hugging Face, and Cua Bench for building tasks, evaluating agents and exporting trajectories. Cua Spaces is source-available under FSL-1.1-MIT and becomes MIT two years after each release, while the rest of the repo is MIT; Spaces is free for individuals, with Pro and Teams plans listed as coming soon.

GitHub Trending(每日)·
0216

Microsoft MarkItDown trends on GitHub as an LLM document converter

Microsoft's open-source Python utility MarkItDown converts PDF, PowerPoint, Word, Excel, images (EXIF metadata and OCR), audio, HTML, CSV/JSON/XML, ZIP archives, YouTube URLs and EPUB files into Markdown for LLM and text-analysis pipelines. It is positioned as a lightweight converter focused on preserving document structure such as headings, lists, tables and links, with third-party plugins and optional higher-quality conversion via Azure Document Intelligence or Azure Content Understanding. The documentation warns that it performs I/O with the current process's privileges, so inputs should be sanitized in untrusted environments.

GitHub Trending · Python·
038

Addy Osmani open-sources agent-skills, a 25-skill pack for AI coding agents

Developer Addy Osmani has published agent-skills on GitHub, an open-source pack that encodes senior-engineering workflows, quality gates and best practices as 25 skills (24 lifecycle skills plus a meta-skill) mapped to six phases: DEFINE, PLAN, BUILD, VERIFY, REVIEW and SHIP, with nine slash commands such as /spec, /plan, /build, /test, /review and /ship. The repo says the pack installs via the npx skills CLI into 70+ agents including Claude Code, Cursor, Codex and GitHub Copilot, with native plugin and local install paths; its README also notes that installing a single skill skips the repo-level references/ checklists, a portability gap tracked in issue #361.

GitHub Trending(每日)·
047

Ai2 releases olmOCR v0.4.0, lifting benchmark score ~4 points with RL

The Allen Institute for AI released olmOCR v0.4.0 on October 21, 2025, a new model trained with synthetic data and reinforcement learning that raises the project's olmOCR-Bench score by roughly 4 points. The toolkit converts PDF, PNG and JPEG documents into clean Markdown, handling equations, tables, handwriting, multi-column layouts and insets while stripping headers and footers and preserving natural reading order. Ai2 says it runs on a 7B vision-language model requiring a GPU, at under $200 per million pages, and reports an olmOCR-Bench score of 82.4±1.1 for v0.4.0 — just below Chandra OCR 0.1.0 (83.1) on its own leaderboard.

GitHub Trending · Python·
056

Uber open-sources ADR agentic AI detection and response system

Uber has open-sourced ADR (Agentic AI Detection and Response), an enterprise security system for AI agents that is deployed in production at the company. The repository includes ADR Discovery, ADR Sensor, ADR-Bench and ADR Detector, which inventory AI apps, CLI agents, IDE extensions, local model runtimes and MCP servers on endpoints, and capture agent intent, tool use and execution traces from 7+ coding tools including Claude Code, Cursor, Codex, GitHub Copilot CLI, DeepSeek Harness, opencode, Claude Desktop and Gemini CLI. The project says ADR-Bench spans 304 tasks and 134 MCP servers across all 17 agent attack techniques; the ADR Prevention component and the ADR Explorer engine are not part of the open-source release, and an accompanying paper was accepted to MLSys 2026.

GitHub Trending · Python·
065

Google open-sources ML Drift, a cross-platform GPU inference engine for on-device AI

Google's AI Edge team released ML Drift under Apache 2.0, a cross-platform on-device GPU compute engine for AI/ML inference that abstracts OpenGL ES, OpenCL, Metal and WebGPU backends and serves as the GPU acceleration engine inside LiteRT, while also being usable standalone. The release adds 5D tensor support, an extensible custom-op framework with an agent-oriented SKILL.md guide, and stage-aware optimizations for on-device LLM prefill and decode; a WebGPU backend built on Dawn extends it to Windows and Linux as a desktop preview. Google says ML Drift already runs on millions of devices powering Chrome, YouTube Shorts, Photos and Meet, with YouTube Shorts seeing up to 40% lower average frame latency, and partners including Adobe and Snap reporting up to 30% faster on-device performance — all vendor-reported figures.

Google Developers Blog·
076

LMCache multiprocess mode hit by 9.8-rated RCE flaw

According to iThome Taiwan, LMCache — a platform for large language model inference acceleration and cache management — has a critical remote code execution (RCE) vulnerability in its multiprocess mode, tracked as CVE-2026-105192 with a CVSS score of 9.8. The report does not state affected version ranges, exploitation prerequisites, whether the flaw has been exploited in the wild, or any official fix or mitigation.

iThome 台湾·
085

Tencent open-sources EVIE-4.5B visual document retriever

Tencent Hunyuan released EVIE-4.5B on Hugging Face under Apache-2.0, a visual document retrieval model built on a ColQwen3.5 late-interaction multi-vector architecture with a single-projection Prefix-MRL head that can be truncated at runtime from 2048 down to 64–2048 dimensions without separate checkpoints. The model card reports 66.02 on ViDoRe V3 (66.75 for the 8B flagship teacher) and pairs it with training-free HAC token compression that cuts roughly 750 vectors per page to 32, bringing a 1M-page index to about 3.81 GiB. Tencent says weights, training pipelines (including ARD distillation from the 8B teacher), the HAC algorithm and evaluation suites are all open-sourced, with a formal paper to follow.

腾讯混元 · Hugging Face·
099

gpt-instruct publishes Codex jailbreak prompt pack with new pre-releases

The GitHub project MDX-Tom/gpt-instruct offers Codex jailbreak prompts plus a reproducible evaluation toolchain aimed at improving first-pass execution, process continuity and artifact verification on complex tasks, with runnable rollback. It maintains three lines: the stable gpt-5.6-sol-v45 and two pre-releases, gpt-6-astra-v2-rc1 and gpt-6.1-sol-v1-rc2; the author reports B-tier non-cloud scores of 42/50 cases for Astra and 34/50 for 6.1, with C-tier tests not run for either. The project says it is for AI-safety purposes only, uses Codex's official configuration mechanism, and warns that jailbreak activity carries account risk, recommending throwaway accounts.

GitHub Trending · Python·
105

LLM-built Rust port of the TypeScript compiler beats Microsoft's Go version on speed

Developer pingdotgg has ported Microsoft's native TypeScript compiler (written in Go) to Rust as ts-rust (npm package tsc-rs). The early 0.1.0 release claims 100% compatibility on every real-world project tested, all 181,711 ported Go tests passing, and geometric-mean type-check times about 1.61× faster than tsc 7 across six open-source apps. The author says more than $400,000 of OpenAI GPT-5.6 Sol and GPT-6 Astra tokens never exceeded ~84% compatibility, while Claude Opus 5.5 rewrote the port from scratch and produced a working v0 in 10 hours for roughly $24,047 of API spend over two weeks.

Hacker News · AI(100+ 分)·
115

Google open-sources AQuA, an agent that diagnoses production AI agents

Google's developer blog introduces AQuA (Ambient Quality Agent), released as composable building blocks in the google/adk-recipes repo. It runs alongside your agent in your Google Cloud project, sampling up to 1,000 recent sessions from Cloud Trace, Cloud Logging or BigQuery on a schedule, after deployments or on demand, and pushes them through a five-stage pipeline (sample, review, cluster, verify, track) that yields structured findings; it never sits in the request path and does not edit code or open pull requests itself. In a demo sweep of 32 travel-concierge sessions, it produced 42 findings and 9 candidate clusters, of which the Gemini 3.7 Flash verifier rejected 3 false positives, leaving 6 verified issues.

Google Developers Blog·
127

vLLM climbs GitHub Trending among Python projects

vLLM is a high-throughput, memory-efficient inference and serving engine for large language models, originally developed in the Sky Computing Lab at UC Berkeley and now maintained by a community of more than 2,000 contributors. It manages attention key-value memory with PagedAttention and supports continuous batching, chunked prefill, prefix caching, speculative decoding and quantization formats including FP8, INT8/INT4, GPTQ/AWQ and GGUF, alongside an OpenAI-compatible API server. The engine runs on NVIDIA, AMD and Intel GPUs plus x86/ARM/PowerPC CPUs, and supports 200+ model architectures on Hugging Face.

GitHub Trending · Python·
136

hackingtool: AI-guided all-in-one security testing toolkit trends on GitHub

hackingtool is an open-source all-in-one toolkit for authorized security testing that the project describes as bundling 215 curated tools across 21 categories, with an AI layer on top: you describe a goal in plain English (e.g. "find subdomains of example.com") and it maps intent onto a fixed 63-tag taxonomy to surface the right tool and command. The AI layer is opt-in and bring-your-own-key, using an OpenAI-compatible endpoint or local Ollama, and degrades to offline keyword matching with no model configured; /goal drafts a step-by-step plan, asks for authorization confirmation, and never feeds tool output back to the model. It runs only on Linux/macOS, emphasizes no auto-execution, SHA-256-verified downloads, and refuses out-of-scope requests such as DoS or mass targeting.

GitHub Trending · Python·
145

WordPress 7.1.3 patches 7 flaws, 3 reported by Anthropic

WordPress released version 7.1.3 on October 6, fixing seven security vulnerabilities and four program bugs. Anthropic reported three of the seven flaws, while security firms Trail of Bits and Patchstack each reported one, three independent researchers jointly submitted one, and WordPress's own security team found the last. Anthropic's findings cover a delayed-trigger injection in the WXR exporter, a denial-of-service flaw in WP_Http::make_absolute_url(), and an author-role privilege issue allowing posts to be pinned; the report says the most serious issue was a stored XSS in the comment moderation interface found by Trail of Bits' Thomas Chauchefoin.

AIbase AI新闻·
155

Apollo 3.0 moves toward Agent First configuration management

Open-source news site OSChina reports that Apollo 3.0 sets out an "Agent First" direction for configuration management. The article notes that large models are shifting from generating content to executing tasks, with agents already taking part in coding, testing, operations and software delivery, which makes reading, changing, publishing and rolling back configurations a natural next need. It also stresses that a configuration center carries production runtime control and may hold sensitive settings, so opening configuration management to agents is not simply a matter of adding a few interfaces.

开源中国·

That's everything.