Create

Sign in to ReadmeX

Sign in to join communities, post, vote and chat.

New here?

Story

AI's simmering question: where is all the automation?

AI summary

A The Information column frames the industry debate over why automation has not arrived even as OpenAI said its unreleased model solved or made major progress on hundreds of significant math and computer science problems. Cognition founder Scott Wu argued that ever-more-capable agents plus more data centers, memory and sandboxes will yield "virtual employees," while ex-OpenAI researcher Diogo Almeida, now at TypeSafe AI, maker of Jev, said today's LLMs are optimized for assistance and fail at unattended automation. Anthropic's Claude Code product lead Cat Wu conceded much automation stalls at a semi-automated stage where Claude does 80% and a human does 20%.

Why it matters: It highlights the gap between rising model capability and real-world automation, nudging practitioners to focus on reliability and productization rather than benchmarks alone.

OpenAIAnthropicClaude Code

8
Source excerptThe Information

A new age of discovery and automation is here. Can you feel it? The leading AI companies certainly can. But you may not.

The rare slow news day gave us more time to marinate on the significance of OpenAI’s late-breaking announcement on Tuesday that its unreleased AI model had solved or made major progress on hundreds of significant problems in math and computer science. What does it mean for those of us who don’t exert much brain power on algebra every day? 

It’s difficult to separate these huge technical achievements from more commercially driven announcements that just keep the AI money train going. And it wasn’t lost on me that Tuesday’s announcement, while groundbreaking, came amid reports of OpenAI raising more cash, and a week after its buzzy developer day mostly underwhelmed. Meanwhile, the latest AI craze—consumer agents—is starting to crash into the reality that a lot of websites want to block bots. And economists don’t see signs yet, at least in major statistics, that AI is making us more productive.

One question animating this phase of the AI gold rush is: How will companies convert raw, impressive intelligence into actual, demonstrable usefulness? 

I witnessed the debate recently at the Midway, a dark auditorium near the San Francisco waterfront that often hosts EDM shows and art exhibitions. Last week, nearly 800 software developers filed inside for a conference hosted by Modal—not an EDM DJ, but a highly valued seller of computing infrastructure and software tools for developers.

Modal gave headline billing to Scott Wu, who used to be best known for his prowess in math competitions, before the coding agent startup he founded, Cognition, rocketed to a $48 billion valuation last month. 

He stood on stage in front of a slide that read: “It’s time to be ambitious.” He talked about a near future where companies are full of “virtual employees,” which would “work on outcomes instead of tasks.” With enough data centers, memory and sandboxes, the industry would get there, he said. “Agents keep getting more and more capable. That’s the entire trend.”

And what gave him confidence for his call to arms was math. “The moment for me was the AIME moment,” he said, referring to the time last year when advanced AI models began reasoning well enough to outperform humans on the kinds of high-level math competitions he used to compete in. “That’s when I knew it was all over”—meaning, humans no longer claim a monopoly on intelligence.

I heard a different tone later that afternoon at the conference from Diogo Almeida, a former OpenAI researcher who described himself as “co-author of some of OpenAI’s greatest hits.” Almeida, smiling and wearing a pink fur coat on stage, spoke with the insider zeal of someone who had defected from a problematic country or religion and could now tell you all about it. He now runs TypeSafe AI, which recently launched a popular alternative to large language models called Jev. 

“Where the fuck is all the automation?” he asked the crowd, before firing off a string of recognizable acronyms, at least to this crowd. “My TL;DR on the state of what is going on in AI is 100% of LLMs today are optimized for assistance with RLHF.” (For the uninitiated, the latter acronym is reinforcement learning from human feedback.) In other words, LLMs are trained to please a human in the loop, so they excel at assistance and fail at unattended automation.

At the same event I heard some of the most important AI leaders pushing for more automation among their software engineering ranks, even as they acknowledged it could be a tough change. They didn’t sound ready to let the machines take over, either. 

“One area we’re pushing a lot more on is [to] think, what are the repetitive tasks I’m doing all the time? Why am I still in that loop?” said Cat Wu, head of product for Claude Code at Anthropic, perhaps the most AI-pilled company in the world. “A lot of automation gets stuck in this stage of being semi-automated. Claude does 80% of the work, and I do 20%. And you’re doing the tasks repetitively.”

And some software engineering leaders are still vexed more by taming humans than AI agents. Dax Raad, a prominent open-source developer who built OpenCode, said his company’s latest problem was that human engineers aren’t actually paying enough attention to code that AI is writing. “I’m just trying to figure out how to get my team to ask the right questions and feel bad when something goes out that they didn’t pay attention to,” Raad said. “It’s the age-old problem: how do you get people to care more?”

In Other News

• Elon Musk announced on Tuesday night that SpaceX’s AI unit will now use some AI models from competitors to power Grok Bot, signaling a shift away from relying exclusively on models developed in-house.

• Microsoft on Wednesday unveiled its latest effort to run AI in PCs powered by its Windows software, rather than running the AI in the cloud, which the company said would bring down costs for customers. More here.

• Anjney Midha, a former general partner at Andreessen Horowitz, along with former executives at Google, Apple and Nvidia have launched a company that aims to make it easier and more affordable for smaller companies or startups to access such compute.

• The U.S. Treasury Department said Wednesday it had fined the parent company of startup accelerator Plug and Play Tech Center, the first penalty issued under an outbound investment restriction program targeting China’s tech sector. 

Today on The Information’s TITV

Check out today's episode of TITV in which Akash Pasricha speaks with Wedbush‘s Matt Bryson about Intel’s likely role in Musk's Terafab.

New From Our Reporters

Exclusive

Ex-Trump AI Adviser Sriram Krishnan Targets Raising a $500 Million Venture Fund

By Leo Schwartz

Exclusive

Where Nvidia’s $100 Billion Dealmaking Juggernaut Will Go Next

By Valida Pau and Phoebe Liu

Recommended Newsletter

Start your day with Applied AI, the newsletter from The Information that uncovers how leading businesses are leveraging AI to automate tasks across the board. Subscribe now for free to get it delivered straight to your inbox twice a week.

This is the outlet's own summary. Read the full story on the original site.

Read the original →

How we got here

  1. GPT-6 Astra Solves 59-Year-Old Plasma Fusion Conjecture in Two Days36氪 人工智能 · OpenAI
  2. GitHub tops 225M users as non-coders flood in36氪 人工智能 · Microsoft
  3. CITIC Securities: AI industry focus shifting to inference and monetization36氪快讯 · Anthropic
  4. Inside OpenAI's 722 math manuscripts: the headline claims, formalized or not36氪 人工智能 · OpenAI
  5. Fired OpenAI researchers urge labs to halt work that impairs AI monitoringTechmeme · OpenAI
  6. AI boom has dashed US reindustrialization, essay argues虎嗅 AI · Microsoft

Comments

I've used this: share my experience What I think: share my view
How important is this story?No ratings yet

No comments yet. Start the conversation.