Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

or

New here?

AI News

All dates
0129

OpenAI launches GPT-6.1 Sol Ultrafast in API, Codex and ChatGPT Work

OpenAI rolled out an Ultrafast version of GPT-6.1 Sol across the API, Codex and ChatGPT Work, saying it delivers near-Astra intelligence at up to 8x the speed of Sol Standard, priced at $12 per million input tokens and $60 per million output tokens. In Codex and ChatGPT Work, access is limited to the $500/month Pro 500 plan, qualifying usage-based Enterprise and Edu users, and some developers report quotas draining quickly at high reasoning effort. OpenAI also previewed an "instant steering" capability in Codex that lets users redirect a run mid-generation.

OpenAI (X)·
029

Anthropic, Google, Mistral roll out Haiku 5.5, Gemini 4 Argon and more

Over the past week Anthropic, Google and Mistral AI each shipped new models. Anthropic released Claude Haiku 5.5 — its first major refresh in a year for its cheapest, fastest Claude — which it says costs about 75% less to run than the previous Haiku and is the first Haiku-level model with an adjustable effort setting, alongside Claude Sonnet 5.5, which posts sizable gains in agentic coding and knowledge work and alignment and cybersecurity capabilities the company compares to Opus 5. Google's Gemini 4 Argon, a frontier model with a 1 million-token output limit and stronger deep reasoning, is limited to trusted cybersecurity workers through its Fairwind program, while Mistral's flagship Mistral Large 4, nicknamed "Le Chonk," is in public preview and, per Mistral, ranks among the world's top open-weight models ahead of an Oct. 27 rollout.

CNET·
0318

GPT-6 rolls out to all ChatGPT users alongside a major output UI revamp

During OpenAI's run of 28 consecutive days of updates, GPT-6 was pushed to all ChatGPT users: previously only Pro subscribers could use the strongest model, GPT-6 Astra, while other users' thinking levels ran on GPT-5.6-series models — now all of them are GPT-6. ChatGPT also got an output UI overhaul for all users, with cleaner formatting and interactive charts available in every thinking mode. The author's hands-on test says the fastest mode produced a Chongqing itinerary with images and a city map in about 10 seconds, and that complex questions are now answered as the model reasons rather than after all reasoning finishes.

36氪 人工智能·
049

Tavus says Griffin-Lite is first video chat model to pass a 'video Turing test'

San Francisco AI video company Tavus released a new model, Griffin-Lite, which it calls the first real-time video conversation model to pass a "video Turing test" and describes as a "human interaction model." Tavus says that after one-minute video calls, 26 of 54 participants (48%) mistook it for a real person, versus 2.4% for the previous-generation system under the same test. Unlike voice assistants, Griffin-Lite can react in real time while both sides are speaking, whether agreeing or interrupting.

36氪快讯·
058

Odyssey launches Odyssey-3 world model that drives cars and controls humanoids

Palo Alto lab Odyssey released Odyssey-3, a world model it says can control robot arms and humanoids, drive a car, fly drones indoors, generate environments for AI agent training and play Grand Theft Auto V; a research preview is available now. Odyssey says Odyssey-3 Pro scored 66.1 on the video-to-video test of Physics-IQ Verified — the highest reported on that leaderboard, using the best of eight attempts per task — and ranked first in three of four WorldMark categories, a result from its own evaluation. Humanoid company Flexion has built control policies on Odyssey-3, which Odyssey says handled lighting changes that caused tested baselines to fail; the car was driven on real roads in India with a policy trained on about 20 hours of data.

The Next Web·
0612

PFN launches PLaMo 3 Translate 31B, claiming edge over GPT-6 on translation

Japan's Preferred Networks released PLaMo 3 Translate 31B, a translation model built on its PLaMo 3 base model that expands language coverage from just Japanese and English to 53 languages and adds an online meeting translation mode covering 15 languages, with Japanese making up about 30% of training data. PFN says the model's average score across four translation benchmarks beats OpenAI's GPT-6 Sol and GPT-6 Astra, at roughly 14 yen (about 0.59 yuan) per 100,000 input characters. Those performance and cost figures are the company's own claims.

AIbase AI新闻·
078

Step 5 Preview goes free for a week on OpenCode, Cline and Nous Portal

StepFun says its flagship model for agentic and professional work, Step 5 Preview, is free for one week in OpenCode, Cline and Nous Research's Nous Portal (Hermes Agent). The company touts 1M context, multi-modal input and zero data retention. Cline claims the model scores ahead of Kimi K3 and GLM-5.3 on DeepSWE, calling it one of the strongest open-weight coding models available — a partner claim rather than an official benchmark.

阶跃星辰 StepFun·
087

ShengShu opens Vidu Q4 preview, 720p video from about ¥0.6/sec

ShengShu Technology has opened a preview of Vidu Q4, its next-generation flagship video generation model, to creators. In its hands-on test, QbitAI put 720p generation as low as roughly ¥0.09 per second, while the company's MaaS pricing is cited at about ¥0.6/sec for 720p image-to-video and ¥0.75/sec for 1080p, with a two-month limited-time discount on the SaaS side during the preview. The report says the model outputs up to 4K, generates clips up to 16 seconds, accepts as many as 15 reference images and three reference audio clips, and improves facial emotion, camera moves including FPV, and effects that interact with characters and lighting.

量子位(原生 RSS)·
097

Jev: TypeSafe AI's probability-only model, explained

TypeSafe AI emerged from stealth on September 15 with Jev, a model that does not produce free-form text but instead returns estimated probabilities over a fixed answer set — yes/no, multiple choice, or a rating. It went viral: the announcement drew nearly 40 million views on Twitter, over 100,000 people joined its Discord within a week, some 2,000 GitHub projects adopted it, and signups opened on September 20 were paused two days later amid overwhelming demand. Because it is faster and cheaper and its output slots directly into code patterns like if statements, Jev has spawned similar models from the likes of Cloudflare, Amazon and OpenAI.

Understanding AI·
105

Reflection launches Beam, Mistral debuts Large 4 as Western open models push back

Reflection, the US startup billed as an "American DeepSeek", released its first open-weight model Beam: a 501B-parameter MoE with 23B activated per token for coding, reasoning and agent workloads, whose final RL run used more than 10,500 Nvidia GB300 GPUs and over 100 million rollouts. A day later Mistral launched Mistral Large 4, its largest model yet at 1.05T parameters (52B active), 1M-token context and native multimodality, trained from scratch on about 3,800 Grace Blackwell GPUs in Europe; it is in API preview with full weights due in late October. Both companies concede Chinese open models remain ahead: Beam trails GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash on DeepSWE v1.1 (44.4), Terminal Bench v2.1 (80.1) and HLE (36.2), so Reflection is pitching intelligence per unit of inference compute instead.

36氪 人工智能·
116

Vidu Q4 Preview hands-on: flagship video model from ¥0.09/second

Shengshu Technology has opened preview access to Vidu Q4 Preview, a next-generation video generation model it positions as a "high-expressiveness flagship" focused on character acting, camera work and complex VFX, starting at ¥0.09 per second. In a hands-on test, ifanr found the model strong on camera direction, subject stability and effects that interact with the scene, with support for up to 15 reference images, three reference audio clips and optional 2K/4K output, but weaker on pacing during sharp emotional turns, style control and the hallucination issues common to AI video. The company says the preview is meant to gather real-world feedback before a full release.

爱范儿·
125

Why isn't the industry freaking out about DeepSeek 4.1 Flash?

A developer blogger says a month of heavy use of DeepSeek 4.1 Flash left him unable to tell it apart from frontier models such as Claude Opus 5.5 in everyday coding, at a fraction of the cost; he sometimes uses Opus for a final code review and has DeepSeek apply the fixes. He also credits a roughly 437x smaller KV cache versus DeepSeek V1 for keeping all-day sessions under a dollar, and dismisses the dispute over Chinese labs allegedly distilling Claude's training data as irrelevant to developers chasing value.

Hacker News · AI(100+ 分)·

That's everything.