Strategic Initiatives
12505 stories
·
45 followers

The AI Price War Is Heating Up—and OpenAI Is Gaining Ground on Anthropic - WSJ

1 Share

LLM (google/gemini-3.8-flash) summary:

  • Market Share Shift: spending among corporate users became evenly split between anthropic and openai by september
  • Product Strategy: openai released cost efficient models including the gpt 5 6 lineup to lower operational expenses
  • Price Reductions: openai reduced prices for models such as luna by 80 percent and terra by 20 percent
  • Infrastructure Strain: high demand for claude code caused anthropic to experience frequent outages and computing bottlenecks
  • Data Retention Concerns: certain clients avoided fable 5 because anthropic required a 30 day data retention period
  • Public Offerings: both companies are preparing for initial public offerings to justify high valuations
  • Vendor Diversification: businesses are shifting workloads among different providers to avoid relying on a single model developer
  • Efficiency Countermove: anthropic introduced the claude 5 5 family in september emphasizing reduced cost and improved efficiency

BPC > Try for full article text (no need to report issue for external site) | no article content found! | : | archive.today | archive.vn
Dario Amodei, the CEO of Anthropic, and Sam Altman, the CEO of OpenAI.Dario Amodei, the CEO of Anthropic, and Sam Altman, the CEO of OpenAI. David Paul Morris/Bloomberg News, Jeff Chiu/AP

Anthropic edged to the front of the AI race earlier this year with cutting-edge models and a coding tool that corporate America loved. By some measures, OpenAI is now hot on its heels.

Demand for Anthropic’s Claude Code was so strong in the spring that the company faced frequent outages and a computing crunch. 

But businesses grappling with mounting costs as more employees experimented with artificial intelligence realized they didn’t always need the most advanced options for most tasks. Companies, which now use a variety of models, increasingly turned to low-cost, open-weight models from China for some tasks. And some Claude users balked at guardrails Anthropic put on its powerful Fable 5 model when it was released in early June.

OpenAI stoked the price war this summer with the release of its GPT-5.6 lineup—Sol, Terra and Luna—giving customers access to models with varying capabilities, including some that are less expensive to run. Its Codex coding product and other business tools are now gaining ground, and it has released additional iterations of its cost-efficient models.

An analysis from OpenRouter, a startup that allows developers to access different models, found that among some 120,000 companies that use Anthropic and OpenAI tools, the share of spending was roughly even between the two AI giants in September. At the beginning of the year, Anthropic commanded three-quarters of that total. OpenRouter said its data largely represents spending by AI-native startups as well as some slightly older tech companies and large enterprises.

Share of business spending on OpenAI and Anthropic models, weekly

100

%

75

Anthropic

50

25

OpenAI

OpenAI’s GPT-5.6 models

0

Jan.

Feb.

March

April

May

June

July

Aug.

Sept.

100

%

75

Anthropic

50

25

OpenAI

OpenAI’s

GPT-5.6 models

0

Jan.

Feb.

March

April

May

June

July

Aug.

Sept.

100

%

75

Anthropic

50

25

OpenAI

OpenAI’s

GPT-5.6 models

0

Jan.

Feb.

March

April

May

June

July

Aug.

Sept.

100

%

Anthropic

75

50

25

OpenAI

OpenAI’s

GPT-5.6 models

0

Jan.

Feb.

March

April

May

June

July

Aug.

Sept.

100

%

Anthropic

75

50

25

OpenAI

OpenAI’s

GPT-5.6 models

0

Jan.

Feb.

March

April

May

June

July

Aug.

Sept.

Note: Spending by 120,000 companies on Anthropic and OpenAI models. Some companies in data set also use tools from other providers. OpenRouter says data largely represents spending by AI-native startups as well as some older tech companies and large enterprises. Chart data begins with the first Monday in January.

Source: OpenRouter

OpenAI, which said 2.5 million businesses now use its products, is narrowing the gap in the battle for business customers at a crucial time, as both model makers are planning initial public offerings. Anthropic is working toward an IPO as soon as November while OpenAI is likely to go public next year. 

Both AI companies have skyrocketing capital expenditures, and the coming months are critical in showing Wall Street they have sustainable businesses and revenue streams to back up trillion-plus dollar valuations.  

David Zhu, co-founder and chief executive officer of AI sales-platform startup Reevo, said his company’s AI use is shifting away from Anthropic to OpenAI models for two reasons: cost and a desire to be less reliant on any one model maker. “The honeymoon phase of being tied to one model maker is gone,” said Zhu. He said his preference could shift again. “Things change so quickly,” said Zhu.

On Wednesday morning, a sea of people snaked around a downtown warehouse in San Francisco, waiting to get into the hottest event in town: the Claude Founder House, hosted by Anthropic. Founders and developers started lining up an hour before the event’s start time.

An event staffer at one point shouted, “You need to have a ticket. If you are on the wait list, we will not be approving you.” A day earlier, some people waited for three hours, only to be turned away because the networking event reached capacity, with the large crowd drawing comparisons on social media to lines outside the Coachella music festival. 

Michael Szklarski, co-founder of videogaming startup ReadyM, who waited in line for Claude Founder House on Wednesday, said his company often needs the most cutting-edge models and prefers Anthropic’s Fable 5.1 model. But he hoped to meet with Anthropic engineers to discuss, among other things, how to bring down the cost.

Ara Kharazian, an economist at finance startup Ramp, said he looked at data from his employer’s 70,000 customers in mid-September and observed that businesses were spending more on OpenAI’s models than Anthropic’s for the first time since December. But the lead was short-lived. By the end of the week, Anthropic was back on top by a slim margin. Ramp said its data set skews toward high-growth, tech-forward companies of varying sizes, with some Fortune 500 companies included. Its data omits spending by individual developers.

Many founders and industry analysts point to OpenAI’s June release of the GPT-5.6 models—designed with cost efficiency in mind—as a turning point.

Shortly after the models launched, OpenAI lowered the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%.

“The GPT-5.6 opened up this new lane for OpenAI in a way that the Anthropic family of models doesn’t really have,” said Peter Walker, head of insights at OpenRouter, which is owned by fintech company Stripe. 

David Hsu, founder and CEO of software-development platform Retool, said that at the start of the year, his company moved back and forth between models from OpenAI and Anthropic. But the release of GPT-5.6 models was “the main catalyst” for why Retool mostly uses OpenAI models now, he said. “We want to find the cheapest models because our revenue goes up the cheaper the models are,” said Hsu.

Today, he estimates that he spends about 20% less using OpenAI’s models than models from Anthropic.

Another factor in Hsu’s shift was Anthropic’s announcement that with Fable 5, the company would keep data from users for 30 days for what it said were trust and safety purposes. That was a problem for companies working in industries that handled sensitive data—including Retool. Anthropic has since tried to address the issue with some customers by giving them control of the stored data.

“Pretty much all the contracts we’ve signed, the default is all data is deleted,” said Hsu. “But we could not attest to that if we had used Fable, so we just never used it.” 

But any lead in AI can change. In late September, Anthropic started releasing its new Claude 5.5 family of models. The benefits it touted: lower cost and better efficiency.

Copyright ©2026 Dow Jones & Company, Inc. All Rights Reserved. 87990cbe856818d5eddac44c7b1cdeb8

Appeared in the October 8, 2026, print edition as 'OpenAI Is Close on Anthropic’s Heels Amid Price War'.

Angel Au-Yeung is a finance and technology reporter for The Wall Street Journal in San Francisco. She covers business leaders, startups and Silicon Valley culture. She has won several national awards for her work, including investigations into a Russian billionaire's ownership of dating app Bumble, the final months of the late former CEO of Zappos Tony Hsieh and the downfall of crypto-trading firm FTX.

She is the co-author of "Wonder Boy: Tony Hsieh, Zappos and the Myth of Happiness in Silicon Valley," which was named one of the best business books of 2023 by the Financial Times and described by the New Yorker as "mandatory reading for anyone who is interested in big tech."


Up Next


Videos

Read the whole story
bogorad
3 hours ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete

Moët Hennessy Directs AI to Sniff Out a Stubborn Problem in Winemaking | The Morning Download for Oct. 7 - WSJ

1 Share

LLM (google/gemini-3.8-flash) summary:

  • Aroma Detection: analog devices and moet hennessy created an artificial intelligence sensor system that detects grape crop defects with high accuracy
  • Sensor Approach: the device analyzes overall smell patterns rather than measuring individual chemical compounds to identify bad samples
  • Agricultural Applications: researchers plan to test the technology on wildfire smoke damage and plant viruses affecting grape crops
  • Model Alignment: appian chief executive matt calkins advocated for strict testing standards to prevent misaligned artificial intelligence models
  • Cloud Funding: lambda is raising four billion dollars ahead of an initial public offering reaching a valuation of fourteen point five billion dollars
  • Model Release: mistral announced plans to release weights for its mistral large 4 model on october 27
  • Nuclear Power: google and constellation energy signed a twenty year agreement to supply nuclear energy for data centers
  • Cybersecurity Risks: jamie dimon stated that the release of the anthropic mythos model increased global cyber vulnerabilities tenfold

Moët Hennessy’s Robert-Jean de Vogüé Research Center in Oiry, FranceMoët Hennessy's Robert-Jean de Vogüé Research Center in Oiry, France Moët Hennessy

Good morning. Sensing technology—like cameras and microphones—is good at capturing the sights and sounds around us and feeding them into AI models for analysis. But what about the things we smell? 

Sensing these invisible molecules is much less straightforward, yet it holds a lot of business promise—especially for a company like Moët Hennessy.

Manuel Reman, member of Moët Hennessy Executive Committee.Moët Hennessy’s Manuel Reman Moët Hennessy

The luxury wine unit at LVMH said a pesky defect known as Fresh Mushroom Aroma has started to appear across its grape crop and the whole Champagne region in recent years. The defect is undetectable to humans until after the grapes are fermented into wine, when it then spoils the yield. It doesn’t occur every year (most recently manifesting in the 2023 crop). But when it does, the results are disastrous and can cost the company millions of euros in wastage, said Manuel Reman, member of the Moët Hennessy Executive Committee, overseeing R&D.

Finding a way to detect the defect in the grape juice before the lengthy fermentation process was the focus of a recent research collaboration between the winemaker, the semiconductor and software firm Analog Devices and the University of California, Davis.


Analog Devices built a machine that used a new type of chemical sensor developed in-house to capture olfactory data from the grape juice. Then together with Moët Hennessy and UC Davis, it trained an AI model to determine the likelihood that a it was infected with Fresh Mushroom Aroma. Early testing showed 99% accuracy.


Newsletter Sign-up

WSJ | CIO Journal

The Morning Download delivers daily insights and news on business technology from the CIO Journal team.

Subscribe

“I’m not telling you everybody was crying in the room, but nearly. I still have goosebumps. It’s so huge,” Reman said. 

So how did they do it? 

Today’s sensing technology struggles with smell partly because it tries to identify and measure individual chemicals, a challenging task, said Max Shulaker, chief of Health Solutions at Analog Devices. Shulaker said he took a different approach, building miniaturized sensors that capture the overall aroma fingerprint of a sample. The AI learns to recognize patterns associated with outcomes of interest. 

“Rather than asking ‘What chemicals are present?,’ we ask ‘Does this smell like a good sample or a bad sample?,’ For many real-world applications, that’s the more relevant question,” Shulaker said. 

Analog Device’s system in Moët Hennessy’s Robert-Jean de Vogüé Research Center in Oiry, France.The Analog Devices system in Moët Hennessy's research center. Moët Hennessy

Moët Hennessy provided data from previous spoiled crops to help train the AI algorithm on what to recognize. 

Moët isn’t ready to start using the machine in production yet. The amount of time it takes to read a sample has come down significantly, but it’s still around two hours—too high given that a small Champagne house like Moët Hennessy’s Krug would need to run around 300 samples (and a bigger one like Moët & Chandon might need thousands), Reman said. 

Analog Devices’s Shulaker said he’s confident about bringing that time down as the AI model continues learning from the samples it takes in, getting smarter about exactly what it’s looking for. 

Reman said he hopes to put some of the devices in limited production next year. 

But beyond detecting Fresh Mushroom Aroma, the possibilities are vast for this technology, according to Ben Montpetit, professor and chair, Department of Viticulture and Enology at UC Davis. For example, smoke from fires is another problem impacting wine production that is often undetectable early in the production processes. 

“In 2020 this cost the wine industry close to $4 billion in losses here in California,” Montpetit said “So we’re also exploring that use case as well as plant viruses that are emerging.”


The Quest to ‘Align’ AI

Appian Chief Executive Officer Matt CalkinsAppian Chief Executive Officer Matt Calkins Appian

Matt Calkins, the chief executive of AI automation software company Appian, is taking a firm stance on AI safety: The U.S. needs to institute a strict model “alignment” test, which would ultimately slow the pace of AI development, he told a small group of reporters on Tuesday night.

Model alignment is a term commonly used by researchers to describe AI that acts in ways that match human intentions. And model misalignment refers to AI that acts in ways that ignore or conflict with human intentions.

The idea of model alignment has become more well-known recently as debate rages over whether AI could one day wipe out humans. When taken to the extreme, misaligned AI models could see killing humans simply as a necessary step toward accomplishing their goals.

Avoiding cataclysm. Calkins, who said he doesn’t like to use “cataclysmic language,” still sees misaligned AI as a serious threat: “You come up with a technology as powerful as AI, with the capabilities that it has, and it’s just inevitable that 10 years from now, it’s either massively empowering us or substantially oppressing us,” he said.

His proposed solution is a model alignment test designed by top AI researchers and thinkers, such as Nobel laureate Geoffrey Hinton and Alphabet chief scientist Demis Hassabis. “They would be willing to set the precedent for an alignment test that would be applied to the U.S. and would therefore be copied elsewhere,” he said.

Calkins’s thinking isn’t too different from that of Anthropic and OpenAI leaders, who’ve said they’re researching how they will make sure superintelligent AI models remain aligned. However, both say they don’t yet have a reliable way to do so.

The problem is, Calkins doesn’t think the AI labs will go far enough on their own. “What they’re doing right now is testing alignment gently, realizing that their models are not aligned, and letting them loose anyway,” he said.

Appian itself isn’t strictly an AI company—the firm builds automation software that it sells to governments and enterprises. But to Calkins, talking about AI safety isn’t a matter of selling more software. “I was on an investor tour, and my CFO said, ‘Stop talking about alignment. This isn’t getting you any more investors,’” he said. “But I feel a duty to say this.”

—Belle Lin


On Our Radar

Lambda CEO Michel Combes
Lambda CEO Michel Combes Marlene Awaad/Bloomberg News
  • Neocloud company Lambda is raising up to $4 billion in a final round of fundraising before the company’s planned IPO. The new funding will give the company a valuation of $14.5 billion, excluding the amount of the money being raised, WSJ reports.
  • France’s Mistral said it would soon release Mistral Large 4, dubbed Le Chonk, a new open-weight AI model that it claimed would rank among the best globally. It plans to release the weights, or the trained parameters of the model, on Oct. 27, WSJ reports.
  • A month after its Millennium Prize solution, OpenAI released findings on more than 300 problems—and tried to win back the world of math, WSJ reports.
  • Google and Constellation Energy agreed to a 20-year nuclear-power deal that would boost output at 11 existing reactors, the latest tie-up between the tech and energy industries to power new data centers. The WSJ reports that the upgrades will provide additional power to the grid that is roughly equivalent to building a new large reactor.
  • JPMorgan Chase CEO Jamie Dimon in an interview with Bloomberg said the release of Anthropic’s Mythos model raised the stakes of cybersecurity risks across the globe. Risks “went up 10-fold after Mythos,” Dimon said. “AI created vulnerabilities that we didn’t know about.”

The WSJ Technology Council

The WSJ Tech Council brings together CIOs, CTOs and CISOs advancing innovation and shaping the future. Join this trusted community where tech executives connect with peers to explore emerging trends and gain the perspective they need to stay ahead of disruption.

Request Information


About Us

Follow Isabelle Bousquette on LinkedIn, Instagram, X, and TikTok for more behind the scenes on her tech and AI coverage, and lately, her contributions to the WSJ Leadership Institute’s new Executive Resilience series, where she’s profiling America’s top execs about their fitness and wellness habits.

Follow Belle Lin on LinkedIn and X for her latest reporting on enterprise technology and AI.

Steven Rosenbush is chief of the enterprise technology bureau at the WSJ Leadership Institute. He also has a column. You can follow him on LinkedIn.

Tom Loftus is the editor of The Morning Download. He suggests following Isabelle, Belle and Steve on their various social channels. But if you insist, here’s his LinkedIn.

Read the whole story
bogorad
6 hours ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete

Stanford-Led Study Finds Cheaper AI Models Cost More in 32% of Comparisons

1 Share
Stanford-Led Study Finds Cheaper AI Models Cost More in 32% of Comparisons

In April 2026, Uber chief technology officer Praveen Neppalli Naga sat down to demonstrate the company’s AI coding tools. Over the next two hours, he used $1,200 worth of tokens, the units providers charge for when their models process and generate text.

By then, Uber had already exhausted its entire 2026 AI budget, about four months into the year.

AI work is billed by the token, and a model’s rate card is a poor guide to its bill because token consumption varies between models and even between runs of the same model. Gemini 3 Flash was listed 80% cheaper than GPT-5.4 at May 1, 2026 prices, but cost 38% more across the study’s tasks. Only 11% of 396 enterprises surveyed in April and May 2026 could forecast AI costs within 10%, down from 15% in 2025, in a report released July 29 by Benchmarkit and Mavvrik, which sells AI cost-management software; the figures are self-reported.

“I’m back to the drawing board, because the budget I thought I would need is blown away already,” Naga said.

The Breakdown

  • Researchers from Stanford, Carnegie Mellon, UC Berkeley and Microsoft Research found that in 106 of 336 pairwise comparisons (32%), the model with the lower listed price cost more in total.
  • Gemini 3 Flash was listed 80% cheaper than GPT-5.4 at May 1, 2026 prices but cost 38% more across the study's tasks.
  • Repeated runs of the same query on the same model varied by up to 9.7 times in cost.
  • Only 11% of 396 enterprises surveyed in April and May 2026 could forecast AI costs within 10%, down from 15% in 2025.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

The study

Lingjiao Chen, a researcher at Stanford University and Microsoft Research, tested whether listed API prices predict what a model actually costs to run, with co-authors from Carnegie Mellon, UC Berkeley and Microsoft Research. Their paper, “The Price Reversal Phenomenon,” first appeared on March 25, 2026, and was revised May 28.

The revised study tested eight frontier reasoning models across 12 tasks. Of 336 pairwise cost comparisons, 106, or 32%, showed the model with the lower listed price costing more in total.

No model in the study was consistently the cheapest or the most expensive across all of its benchmarks; the ranking changed from task to task. Listed price was defined as input plus output rates.

Thinking tokens are the hidden reasoning a model writes before answering. Their volume helped explain the reversals. On one MMLU-Pro problem in the study, Gemini 3 Flash consumed more than 60,000 thinking tokens; GPT-5.4 solved the same problem with 25.

For tasks requiring a model to interact repeatedly with tools or a computer environment, the number of turns also drove costs. Each turn can include earlier conversation history as input, adding another charge as the model continues working.

“The practical takeaway is clear,” Chen said. “Price alone should not be used to infer which model is actually cheaper.”

On one prompt in the researchers’ data, the cheaper model cost 14 times as much and still failed. Gemini 3.1 Pro finished in 85 steps for about $1. Gemini 3 Flash went through nearly 1,000 steps, accumulated $14 in token charges and failed.

The researchers published their per-run cost data and code so companies could repeat the comparison on their own workloads.

Same prompt, different bill

Choosing a model that used fewer tokens in one test did not guarantee a repeatable bill. In the May revision, repeated runs of an identical query on the same model varied by up to 9.7 times between the cheapest and most expensive run.

The paper describes an irreducible noise floor, a baseline of randomness that no forecaster can get below: models can follow different reasoning paths even when their inputs stay fixed. That makes predicting the cost of an individual query difficult. The variation cannot be removed by re-prompting.

A follow-up analysis of the researchers’ data on two selected programming prompts showed variation across models from Anthropic, Google and OpenAI. Anthropic models were added through additional data collection after the original study, which had not included them for this measure. Each model received the same prompt five times, with a different cost on each run. Unsuccessful runs consumed tokens and incurred charges too.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

That uncertainty has reached customers building their own software. Mazda Marvasti, co-founder and chief executive of Amberd.ai, said some customers abandoned internally built automation tools because they could not forecast or justify the costs. Amberd.ai builds on private, open-source models run on bare-metal servers using QumulusAI hardware; his remarks come from a Futurum report sponsored by QumulusAI.

“When they start deploying it throughout the organization, the cost starts skyrocketing because it’s a useful tool that somebody built, but it’s now priced on a variable basis,” he said.

The other side

“Some prompt-level fluctuation is inherent to AI, and our testing shows this averages out across a high volume of real-world, diverse workloads,” a Google spokeswoman said.

“Total costs depend on many factors for a given task,” she said, “which can make it hard to forecast new and evolving technology with precision.” Google offers spending caps and flexible pricing. Google, Anthropic and OpenAI have also released newer models that perform better on industry benchmarks since the models in the study were tested.

The paper uses a single pricing snapshot from May 1, 2026, runs each model at one reasoning setting and measures cost separately from answer quality. The cost comparison does not account for whether the answer was right, so a cheap model that fails and an expensive one that succeeds are compared on cost alone.

Budget limits

By June 2026, Uber had capped agentic coding tools at $1,500 per employee per month, per tool.

Inside Uber, chief operating officer Andrew Macdonald described the difficulty of connecting usage measures to what customers receive. On the Rapid Response podcast, he said: “It's very hard to draw a line between one of those stats and 'OK, now we're actually producing like 25% more useful consumer features.'”

Frequently Asked Questions

What is the price reversal phenomenon?

It is the finding that a reasoning model with a lower listed API price can cost more in total to run than a pricier model. In the study by Lingjiao Chen and co-authors, this happened in 106 of 336 pairwise comparisons, or 32%, across eight frontier models and 12 tasks.

Why can a cheaper AI model end up costing more?

Models consume very different numbers of tokens on the same work. On one MMLU-Pro problem, Gemini 3 Flash used more than 60,000 thinking tokens while GPT-5.4 used 25. In tasks with tools, the number of interaction turns also drove costs.

Does the same model cost the same every time?

No. Repeated runs of an identical query on the same model varied by up to 9.7 times in cost. The paper says this variation cannot be removed by re-prompting, which makes the cost of an individual query difficult to predict.

What did Google say about the findings?

A Google spokeswoman said some prompt-level fluctuation is inherent to AI and that Google's testing shows it averages out across a high volume of diverse workloads. She said Google offers spending caps and flexible pricing.

What are the study's limits?

It uses a single pricing snapshot from May 1, 2026, runs each model at one reasoning setting, and measures cost separately from answer quality, so a cheap model that fails and an expensive one that succeeds are compared on cost alone.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

OpenAI’s Sol costs half as much as Opus 5.5; Trump renames AI super intelligence
IMPLICATOR .ai Morning Briefing · From San Francisco   Wednesday, September 23, 2026 10 stops From San Francisco 1 The Editorial   Morning, humans. Today’s theme: w
Palo Alto Networks CEO Says AI Token Costs Must Fall Up to 90%
Palo Alto Networks CEO Nikesh Arora said on CNBC on Thursday that AI token costs need to fall as much as 90% to support large-scale enterprise adoption. He called OpenAI CEO Sam Altman’s claim that th
Anthropic shifts enterprise billing to per-token pricing. The flat-fee era is over.
Anthropic has restructured its enterprise plan to bill Claude, Claude Code, and Cowork usage separately from seat fees, moving its largest business customers to per-token pricing at standard API rates
Read the whole story
bogorad
2 days ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete

DeepSeek Narrows US AI Benchmark Lead to 3% After September Release

1 Share
DeepSeek Narrows US AI Benchmark Lead to 3% After September Release

DeepSeek’s September release narrowed the US lead over China to roughly 3% in a LiveBench score comparison. The V4.1 Flash model moved close to Anthropic’s leading system on LiveBench’s overall tests, which include reasoning and coding. Bloomberg Intelligence senior analyst Robert Lea expects the improved performance to bring Chinese developers further market share gains.

In the leaderboard captured on October 4, DeepSeek V4.1 Flash Max Effort scored 81.1 against 83.4 for Anthropic’s Claude Fable 5.1 Max Effort. That is a difference of 2.3 score points, or about 2.8% of Anthropic’s score. Lea’s comparison put the gap at about 9% in May and 15% earlier in 2026.

The scores differ by task. In the October 4 snapshot, DeepSeek scored 77.3 on agentic coding, compared with 66.1 for Claude Fable 5.1 Max Effort.

The US still holds the lead in this comparison. These benchmark results do not establish national AI leadership or show whether US chip export restrictions are effective.

What Changed

  • DeepSeek narrowed China’s gap in a LiveBench comparison to about 3%, from roughly 9% in May.
  • The October 4 leaderboard showed DeepSeek at 81.1 overall, against Anthropic’s 83.4.
  • OpenRouter token traffic does not measure the whole AI market and underrepresents enterprise workloads.
  • Analyst Robert Lea forecasts that China’s AI industry could remain unprofitable until 2030.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

DeepSeek’s efficiency changes

DeepSeek’s September 10 release notes describe a design intended to cut the computing and memory needed to serve answers. Only a small portion of the model’s stored parameters, the values learned during training, is active when it reads input or generates output. The company says it also shrank the cache that holds information used during a conversation.

Those efficiency claims come from DeepSeek. The release supports native visual understanding and continues to offer off-peak API rates at half the peak price, under the new pricing that took effect on September 10. The benchmark gain does not by itself establish how reliably the model will perform inside a company or whether it can be deployed securely.

OpenRouter usage and enterprise demand

Chinese models processed more tokens than US models on OpenRouter in the trailing week to September 7. But US models handled slightly more requests in that period, the usage window examined in JPMorgan Asset Management’s September 9 analysis. OpenRouter is a service developers use to send requests to different AI systems.

Token share is not market share. Automated agents can produce large volumes of tokens while working through a task, so heavy token consumption can make a provider look more widely used than a count of requests does.

OpenRouter also underrepresents enterprise activity. Most of those workloads run through major cloud platforms or directly with developers such as OpenAI and Anthropic. Large US companies still predominantly use US models, while large Chinese companies tend to use domestic ones.

Lower token prices do not always mean cheaper completed tasks. Models consume different amounts of text before reaching an answer, and downloadable models still incur serving costs. A price comparison based on tokens alone can therefore miss the cost of completing the work.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Enterprise procurement also depends on regulation and data governance. Downloading model weights permits private hosting, but does not reveal the training data or resolve intellectual property questions.

Profitability remains uncertain

Lea forecasts that China’s AI industry could remain unprofitable until 2030 despite the benchmark gains. In his assessment, low-margin token supply and a domestic price war may prevent developers from building a lasting commercial advantage.

China’s market is now crowded with more than 1,100 large language models.

Lea identifies ByteDance’s Doubao as the frontrunner in AI app monetization, while rival chatbots from DeepSeek and Tencent remain free.

“Putting China’s AI sector on a sustainable profit footing will require a cooling of competitive pressures, an industry shakeout, and a more rational approach to pricing,” he said.

Frequently Asked Questions

How close is DeepSeek to the US leader on LiveBench?

The October 4 snapshot showed DeepSeek V4.1 Flash Max Effort at 81.1 overall and Anthropic’s Claude Fable 5.1 Max Effort at 83.4. The 2.3-point difference is about 2.8% of Anthropic’s score, rounded to roughly 3% in the comparison.

Does DeepSeek lead on every task?

No. DeepSeek scored 77.3 on agentic coding against 66.1 for Claude Fable 5.1 Max Effort in the October 4 snapshot. Anthropic held the higher overall score.

What changed in DeepSeek’s September release?

DeepSeek described a design intended to reduce the computing and memory needed to serve answers, using only a small part of its stored parameters at a time and a smaller conversation cache. These efficiency claims come from the company.

Does Chinese token traffic prove market leadership?

No. JPMorgan’s September 9 analysis found more Chinese-origin token traffic on OpenRouter but slightly more US-origin requests. The service underrepresents enterprise workloads, and agents can consume many tokens per task.

When could China’s AI industry become profitable?

Robert Lea forecasts that the industry could remain unprofitable until 2030. That is an analyst forecast, tied to competitive pressure and pricing, rather than a demonstrated outcome.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Stanford AI Index 2026 pegs US-China AI gap at 2.7%
Stanford's 2026 AI Index closes the US-China performance gap to 2.7%. But the Chinese labs that closed it are pivoting to closed source, and DeepSeek V4 has gone silent on Huawei silicon. China caught up the week Chinese labs quit the game that got them there.
Z.ai Delays GLM-5.3 Weights After CyberGym Score Tops Mythos
GLM-5.3 scored 84.5% on CyberGym, edging Anthropic's restricted Mythos 5, and Z.ai responded by holding its downloadable weights until around August 28. The lead vanishes on exploitation benchmarks, and every figure came from Z.ai's own harness.
Read the whole story
bogorad
3 days ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete

Why Aliança Catalana Is Not Running in Spain’s General Election: "If They Ask Me Whether I Hate Spain..."

1 Comment

Orriols’s party has been growing nonstop for two years

  • Decision not to run: Aliança Catalana will not contest Spain’s general elections because its political project is focused exclusively on Catalonia and aims ultimately at independence.
  • Rapid growth: Led by Sílvia Orriols, the party went from being largely unknown outside Ripoll to winning six of 17 council seats there in 2023 and making Orriols mayor.
  • Parliamentary breakthrough: In the May 2024 Catalan elections, Aliança Catalana won two seats and 118,035 votes, or 3.79% of the total, gaining representation from Girona and Lleida.
  • Rejection of Spanish institutions: The party views participation in the Spanish Congress as inconsistent with its objective of creating an independent Catalan state and does not want to become another party within Spain’s political system.
  • Contrast with other separatists: Unlike ERC and Junts, which use their congressional seats to negotiate with Spain’s governments, Aliança rejects seeking additional powers, funding, or agreements from Madrid in favor of breaking with Spain.
  • Immigration-focused platform: The party advocates highly restrictive policies toward illegal immigration and links immigration to security, Catalan identity, demographic change, and opposition to what it calls “Islamization.”
  • Electoral consequences: By staying out of the general election, Aliança gives up the chance to win seats in Congress while concentrating its campaign on Catalonia, leaving Junts, ERC, PSC, PP, Vox, and Sumar to compete for the region’s vote.
Read the whole story
bogorad
3 days ago
reply
Interesting, a different approach
Barcelona, Catalonia, Spain
Share this story
Delete

The End of College as We Know It - WSJ

1 Share

LLM (google/gemini-3.8-flash) summary:

  • Enrollment Declines: undergraduate numbers fell nationwide leading to budget reductions and deficits across major institutions
  • Employer Skepticism: companies increasingly avoid ivy league candidates due to graduates lacking problem solving skills
  • Admissions Criteria: holistic reviews and athletic preferences contribute to underqualified students entering the workforce
  • Research Funding: institutions face scrutiny after charging high indirect overhead rates on federal grants
  • Escalating Costs: tuition and expenses exceed one hundred thousand dollars annually at sixteen institutions
  • Artificial Intelligence: machine systems outperform traditional instruction methods and pass licensing exams
  • Experiential Learning: virtual apprenticeships and yearlong team projects offer hands on alternatives to standard lectures
  • Declining Perception: the share of americans viewing higher education as very important fell to 31 percent


Andy Kessler

Oct. 4, 2026 2:52 pm ET

image Getty Images

What’s the matter with universities? Syracuse missed its 2026 enrollment target and may run a $30 million deficit. Minnesota is cutting its budget by $225 million over two years. Tulsa is slashing tuition by more than half. Colleges enroll 4.2% fewer undergraduates than last year. Cornell, Yale, Vanderbilt and Washington University in St. Louis put out plans—not radical enough, if you ask me—to reimagine the university. Wossamotta U?

Colleges are known for left-leaning professors, speech suppression, bizarre courses—Occidental College taught “The Unbearable Whiteness of Barbie”—and easy grading more than critical thinking, judgment and problem solving. It’s easy to blame artificial intelligence, but problems run deeper:

• Aptitude. In Griggs v. Duke Power (1971), the Supreme Court effectively banned corporations from giving intelligence tests, so companies rely on colleges to do their sorting via aptitude tests like the SAT. That led to branding and credentialism, but that game is over. A Cornell report noted employers’ “jarring” skepticism of college graduates because they lack job skills such as “the ability to handle uncertainty and solve problems that do not have clear answers.” A Forbes survey found almost half of corporate executives are either less likely to hire Ivy League graduates than they were five years ago or would never hire them.

• Admissions. “Holistic reviews” that ignore aptitude and grades put a thumb on the scale and harm college brands. In combination with athletic scholarships and easy majors, they bring in underqualified students who are churned out as subpar employees.

• Accounting. In 1945, Vannevar Bush suggested the U.S. government fund “basic research in the colleges, universities and research institutes,” which became a big business. But universities got greedy and abused the system by overcharging the government for overhead.

That business model is now suspect. The Trump administration instituted a 15% cap on “indirect cost rates.” A 2025 study found that at more than 350 institutions funded by the National Institutes of Health, the average negotiated indirect cost rate was 58% and the effective actual overhead rate averaged 42%. The Trump administration’s 15% cap was struck down in most cases and the administration has backed off for now, but we’ve seen research funds cut off. Watch for “let’s hit up alumni for more” as research overhead may no longer be an overstuffed cow to pay for campus excesses.

That includes professor salaries. Many eyebrows lifted in 2024 when Claudine Gay was fired as Harvard president and went back to being a professor, at a $900,000 salary. The average full professor salary at Harvard in 2026 is $293,619, up 6.4% from last year. I’d bet many of them hate capitalism but are happy to enjoy its fruits.

• Affordability. Student loans lead to constant tuition hikes. Same for high professor salaries and administrative bloat. Sixteen colleges and universities cost more than $100,000 a year in tuition and expenses. Where’s the value? Students could buy a decent GPU rack and start their own AI company for less.

• Artificial Intelligence. Professors and students seem to be going through the motions of learning, so there’s a huge structural problem. Do ChatGPT-written essays get graded by Claude? Why bother?

Three years ago, OpenAI’s GPT-4 scored 1,410 on the SAT, maybe enough to get into Michigan. Since then, AI has passed bar exams and medical licensing exams, which says more about these outdated tests than the inadequacy of humans, perhaps a sign of an antiquated education system. Even worse is a 2025 Harvard study, “AI tutoring outperforms in-class active learning.” Are teachers and professors becoming obsolete?

Universities can take advantage of AI rather than ban it. The sage-on-a-stage format is over. Learning is experiential. As corporate America rejects college branding, graduates would benefit from more hands-on experience. Universities may be uniquely positioned to implement what I call “virtual apprenticeships,” real job experience while still at school.

I suggest all students spend their entire junior year on team projects. Assignments might include building nuclear autonomous buses or robots that cook. Cross-discipline self-organizing teams will form. You’ll need engineers and coders, but also prelaw and political-science majors to clear the regulatory path, philosophers and marketing majors to write press releases, designers, nutritionists, etc.

Self-paced AI-tutored instruction can help students learn what they need as problems arise. A smart entrepreneur might even create a student transfer portal to recruit talented students from other universities.

According to Gallup, in 2013, 70% of Americans said college was very important. Today it’s only 31%. If universities don’t change, more than a few will implode, while others will wither away over decades. AI and alternative forms of learning are already rising. Something radical is needed to stop the decay of our once great universities. Who will step up?

Write to <a href="mailto:kessler@wsj.com">kessler@wsj.com</a>.

Copyright ©2026 Dow Jones & Company, Inc. All Rights Reserved. 87990cbe856818d5eddac44c7b1cdeb8

Appeared in the October 5, 2026, print edition as 'The End of College as We Know It'.

Andy Kessler is the author of Inside View, a column he writes for The Wall Street Journal on technology and markets and where they intersect with culture. He won the 2019 Gerald Loeb Award for commentary. He is the author of several books including Wall Street Meat and Eat People. He used to design chips at Bell Labs before working on Wall Street for PaineWebber and Morgan Stanley and then as a founder of the hedge fund Velocity Capital.


Videos

Read the whole story
bogorad
3 days ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete
Next Page of Stories