Strategic Initiatives
12503 stories
·
45 followers

Stanford-Led Study Finds Cheaper AI Models Cost More in 32% of Comparisons

1 Share
Stanford-Led Study Finds Cheaper AI Models Cost More in 32% of Comparisons

In April 2026, Uber chief technology officer Praveen Neppalli Naga sat down to demonstrate the company’s AI coding tools. Over the next two hours, he used $1,200 worth of tokens, the units providers charge for when their models process and generate text.

By then, Uber had already exhausted its entire 2026 AI budget, about four months into the year.

AI work is billed by the token, and a model’s rate card is a poor guide to its bill because token consumption varies between models and even between runs of the same model. Gemini 3 Flash was listed 80% cheaper than GPT-5.4 at May 1, 2026 prices, but cost 38% more across the study’s tasks. Only 11% of 396 enterprises surveyed in April and May 2026 could forecast AI costs within 10%, down from 15% in 2025, in a report released July 29 by Benchmarkit and Mavvrik, which sells AI cost-management software; the figures are self-reported.

“I’m back to the drawing board, because the budget I thought I would need is blown away already,” Naga said.

The Breakdown

  • Researchers from Stanford, Carnegie Mellon, UC Berkeley and Microsoft Research found that in 106 of 336 pairwise comparisons (32%), the model with the lower listed price cost more in total.
  • Gemini 3 Flash was listed 80% cheaper than GPT-5.4 at May 1, 2026 prices but cost 38% more across the study's tasks.
  • Repeated runs of the same query on the same model varied by up to 9.7 times in cost.
  • Only 11% of 396 enterprises surveyed in April and May 2026 could forecast AI costs within 10%, down from 15% in 2025.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

The study

Lingjiao Chen, a researcher at Stanford University and Microsoft Research, tested whether listed API prices predict what a model actually costs to run, with co-authors from Carnegie Mellon, UC Berkeley and Microsoft Research. Their paper, “The Price Reversal Phenomenon,” first appeared on March 25, 2026, and was revised May 28.

The revised study tested eight frontier reasoning models across 12 tasks. Of 336 pairwise cost comparisons, 106, or 32%, showed the model with the lower listed price costing more in total.

No model in the study was consistently the cheapest or the most expensive across all of its benchmarks; the ranking changed from task to task. Listed price was defined as input plus output rates.

Thinking tokens are the hidden reasoning a model writes before answering. Their volume helped explain the reversals. On one MMLU-Pro problem in the study, Gemini 3 Flash consumed more than 60,000 thinking tokens; GPT-5.4 solved the same problem with 25.

For tasks requiring a model to interact repeatedly with tools or a computer environment, the number of turns also drove costs. Each turn can include earlier conversation history as input, adding another charge as the model continues working.

“The practical takeaway is clear,” Chen said. “Price alone should not be used to infer which model is actually cheaper.”

On one prompt in the researchers’ data, the cheaper model cost 14 times as much and still failed. Gemini 3.1 Pro finished in 85 steps for about $1. Gemini 3 Flash went through nearly 1,000 steps, accumulated $14 in token charges and failed.

The researchers published their per-run cost data and code so companies could repeat the comparison on their own workloads.

Same prompt, different bill

Choosing a model that used fewer tokens in one test did not guarantee a repeatable bill. In the May revision, repeated runs of an identical query on the same model varied by up to 9.7 times between the cheapest and most expensive run.

The paper describes an irreducible noise floor, a baseline of randomness that no forecaster can get below: models can follow different reasoning paths even when their inputs stay fixed. That makes predicting the cost of an individual query difficult. The variation cannot be removed by re-prompting.

A follow-up analysis of the researchers’ data on two selected programming prompts showed variation across models from Anthropic, Google and OpenAI. Anthropic models were added through additional data collection after the original study, which had not included them for this measure. Each model received the same prompt five times, with a different cost on each run. Unsuccessful runs consumed tokens and incurred charges too.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

That uncertainty has reached customers building their own software. Mazda Marvasti, co-founder and chief executive of Amberd.ai, said some customers abandoned internally built automation tools because they could not forecast or justify the costs. Amberd.ai builds on private, open-source models run on bare-metal servers using QumulusAI hardware; his remarks come from a Futurum report sponsored by QumulusAI.

“When they start deploying it throughout the organization, the cost starts skyrocketing because it’s a useful tool that somebody built, but it’s now priced on a variable basis,” he said.

The other side

“Some prompt-level fluctuation is inherent to AI, and our testing shows this averages out across a high volume of real-world, diverse workloads,” a Google spokeswoman said.

“Total costs depend on many factors for a given task,” she said, “which can make it hard to forecast new and evolving technology with precision.” Google offers spending caps and flexible pricing. Google, Anthropic and OpenAI have also released newer models that perform better on industry benchmarks since the models in the study were tested.

The paper uses a single pricing snapshot from May 1, 2026, runs each model at one reasoning setting and measures cost separately from answer quality. The cost comparison does not account for whether the answer was right, so a cheap model that fails and an expensive one that succeeds are compared on cost alone.

Budget limits

By June 2026, Uber had capped agentic coding tools at $1,500 per employee per month, per tool.

Inside Uber, chief operating officer Andrew Macdonald described the difficulty of connecting usage measures to what customers receive. On the Rapid Response podcast, he said: “It's very hard to draw a line between one of those stats and 'OK, now we're actually producing like 25% more useful consumer features.'”

Frequently Asked Questions

What is the price reversal phenomenon?

It is the finding that a reasoning model with a lower listed API price can cost more in total to run than a pricier model. In the study by Lingjiao Chen and co-authors, this happened in 106 of 336 pairwise comparisons, or 32%, across eight frontier models and 12 tasks.

Why can a cheaper AI model end up costing more?

Models consume very different numbers of tokens on the same work. On one MMLU-Pro problem, Gemini 3 Flash used more than 60,000 thinking tokens while GPT-5.4 used 25. In tasks with tools, the number of interaction turns also drove costs.

Does the same model cost the same every time?

No. Repeated runs of an identical query on the same model varied by up to 9.7 times in cost. The paper says this variation cannot be removed by re-prompting, which makes the cost of an individual query difficult to predict.

What did Google say about the findings?

A Google spokeswoman said some prompt-level fluctuation is inherent to AI and that Google's testing shows it averages out across a high volume of diverse workloads. She said Google offers spending caps and flexible pricing.

What are the study's limits?

It uses a single pricing snapshot from May 1, 2026, runs each model at one reasoning setting, and measures cost separately from answer quality, so a cheap model that fails and an expensive one that succeeds are compared on cost alone.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

OpenAI’s Sol costs half as much as Opus 5.5; Trump renames AI super intelligence
IMPLICATOR .ai Morning Briefing · From San Francisco   Wednesday, September 23, 2026 10 stops From San Francisco 1 The Editorial   Morning, humans. Today’s theme: w
Palo Alto Networks CEO Says AI Token Costs Must Fall Up to 90%
Palo Alto Networks CEO Nikesh Arora said on CNBC on Thursday that AI token costs need to fall as much as 90% to support large-scale enterprise adoption. He called OpenAI CEO Sam Altman’s claim that th
Anthropic shifts enterprise billing to per-token pricing. The flat-fee era is over.
Anthropic has restructured its enterprise plan to bill Claude, Claude Code, and Cowork usage separately from seat fees, moving its largest business customers to per-token pricing at standard API rates
Read the whole story
bogorad
1 hour ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete

DeepSeek Narrows US AI Benchmark Lead to 3% After September Release

1 Share
DeepSeek Narrows US AI Benchmark Lead to 3% After September Release

DeepSeek’s September release narrowed the US lead over China to roughly 3% in a LiveBench score comparison. The V4.1 Flash model moved close to Anthropic’s leading system on LiveBench’s overall tests, which include reasoning and coding. Bloomberg Intelligence senior analyst Robert Lea expects the improved performance to bring Chinese developers further market share gains.

In the leaderboard captured on October 4, DeepSeek V4.1 Flash Max Effort scored 81.1 against 83.4 for Anthropic’s Claude Fable 5.1 Max Effort. That is a difference of 2.3 score points, or about 2.8% of Anthropic’s score. Lea’s comparison put the gap at about 9% in May and 15% earlier in 2026.

The scores differ by task. In the October 4 snapshot, DeepSeek scored 77.3 on agentic coding, compared with 66.1 for Claude Fable 5.1 Max Effort.

The US still holds the lead in this comparison. These benchmark results do not establish national AI leadership or show whether US chip export restrictions are effective.

What Changed

  • DeepSeek narrowed China’s gap in a LiveBench comparison to about 3%, from roughly 9% in May.
  • The October 4 leaderboard showed DeepSeek at 81.1 overall, against Anthropic’s 83.4.
  • OpenRouter token traffic does not measure the whole AI market and underrepresents enterprise workloads.
  • Analyst Robert Lea forecasts that China’s AI industry could remain unprofitable until 2030.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

DeepSeek’s efficiency changes

DeepSeek’s September 10 release notes describe a design intended to cut the computing and memory needed to serve answers. Only a small portion of the model’s stored parameters, the values learned during training, is active when it reads input or generates output. The company says it also shrank the cache that holds information used during a conversation.

Those efficiency claims come from DeepSeek. The release supports native visual understanding and continues to offer off-peak API rates at half the peak price, under the new pricing that took effect on September 10. The benchmark gain does not by itself establish how reliably the model will perform inside a company or whether it can be deployed securely.

OpenRouter usage and enterprise demand

Chinese models processed more tokens than US models on OpenRouter in the trailing week to September 7. But US models handled slightly more requests in that period, the usage window examined in JPMorgan Asset Management’s September 9 analysis. OpenRouter is a service developers use to send requests to different AI systems.

Token share is not market share. Automated agents can produce large volumes of tokens while working through a task, so heavy token consumption can make a provider look more widely used than a count of requests does.

OpenRouter also underrepresents enterprise activity. Most of those workloads run through major cloud platforms or directly with developers such as OpenAI and Anthropic. Large US companies still predominantly use US models, while large Chinese companies tend to use domestic ones.

Lower token prices do not always mean cheaper completed tasks. Models consume different amounts of text before reaching an answer, and downloadable models still incur serving costs. A price comparison based on tokens alone can therefore miss the cost of completing the work.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Enterprise procurement also depends on regulation and data governance. Downloading model weights permits private hosting, but does not reveal the training data or resolve intellectual property questions.

Profitability remains uncertain

Lea forecasts that China’s AI industry could remain unprofitable until 2030 despite the benchmark gains. In his assessment, low-margin token supply and a domestic price war may prevent developers from building a lasting commercial advantage.

China’s market is now crowded with more than 1,100 large language models.

Lea identifies ByteDance’s Doubao as the frontrunner in AI app monetization, while rival chatbots from DeepSeek and Tencent remain free.

“Putting China’s AI sector on a sustainable profit footing will require a cooling of competitive pressures, an industry shakeout, and a more rational approach to pricing,” he said.

Frequently Asked Questions

How close is DeepSeek to the US leader on LiveBench?

The October 4 snapshot showed DeepSeek V4.1 Flash Max Effort at 81.1 overall and Anthropic’s Claude Fable 5.1 Max Effort at 83.4. The 2.3-point difference is about 2.8% of Anthropic’s score, rounded to roughly 3% in the comparison.

Does DeepSeek lead on every task?

No. DeepSeek scored 77.3 on agentic coding against 66.1 for Claude Fable 5.1 Max Effort in the October 4 snapshot. Anthropic held the higher overall score.

What changed in DeepSeek’s September release?

DeepSeek described a design intended to reduce the computing and memory needed to serve answers, using only a small part of its stored parameters at a time and a smaller conversation cache. These efficiency claims come from the company.

Does Chinese token traffic prove market leadership?

No. JPMorgan’s September 9 analysis found more Chinese-origin token traffic on OpenRouter but slightly more US-origin requests. The service underrepresents enterprise workloads, and agents can consume many tokens per task.

When could China’s AI industry become profitable?

Robert Lea forecasts that the industry could remain unprofitable until 2030. That is an analyst forecast, tied to competitive pressure and pricing, rather than a demonstrated outcome.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Stanford AI Index 2026 pegs US-China AI gap at 2.7%
Stanford's 2026 AI Index closes the US-China performance gap to 2.7%. But the Chinese labs that closed it are pivoting to closed source, and DeepSeek V4 has gone silent on Huawei silicon. China caught up the week Chinese labs quit the game that got them there.
Z.ai Delays GLM-5.3 Weights After CyberGym Score Tops Mythos
GLM-5.3 scored 84.5% on CyberGym, edging Anthropic's restricted Mythos 5, and Z.ai responded by holding its downloadable weights until around August 28. The lead vanishes on exploitation benchmarks, and every figure came from Z.ai's own harness.
Read the whole story
bogorad
22 hours ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete

Why Aliança Catalana Is Not Running in Spain’s General Election: "If They Ask Me Whether I Hate Spain..."

1 Comment

Orriols’s party has been growing nonstop for two years

  • Decision not to run: Aliança Catalana will not contest Spain’s general elections because its political project is focused exclusively on Catalonia and aims ultimately at independence.
  • Rapid growth: Led by Sílvia Orriols, the party went from being largely unknown outside Ripoll to winning six of 17 council seats there in 2023 and making Orriols mayor.
  • Parliamentary breakthrough: In the May 2024 Catalan elections, Aliança Catalana won two seats and 118,035 votes, or 3.79% of the total, gaining representation from Girona and Lleida.
  • Rejection of Spanish institutions: The party views participation in the Spanish Congress as inconsistent with its objective of creating an independent Catalan state and does not want to become another party within Spain’s political system.
  • Contrast with other separatists: Unlike ERC and Junts, which use their congressional seats to negotiate with Spain’s governments, Aliança rejects seeking additional powers, funding, or agreements from Madrid in favor of breaking with Spain.
  • Immigration-focused platform: The party advocates highly restrictive policies toward illegal immigration and links immigration to security, Catalan identity, demographic change, and opposition to what it calls “Islamization.”
  • Electoral consequences: By staying out of the general election, Aliança gives up the chance to win seats in Congress while concentrating its campaign on Catalonia, leaving Junts, ERC, PSC, PP, Vox, and Sumar to compete for the region’s vote.
Read the whole story
bogorad
22 hours ago
reply
Interesting, a different approach
Barcelona, Catalonia, Spain
Share this story
Delete

The End of College as We Know It - WSJ

1 Share

LLM (google/gemini-3.8-flash) summary:

  • Enrollment Declines: undergraduate numbers fell nationwide leading to budget reductions and deficits across major institutions
  • Employer Skepticism: companies increasingly avoid ivy league candidates due to graduates lacking problem solving skills
  • Admissions Criteria: holistic reviews and athletic preferences contribute to underqualified students entering the workforce
  • Research Funding: institutions face scrutiny after charging high indirect overhead rates on federal grants
  • Escalating Costs: tuition and expenses exceed one hundred thousand dollars annually at sixteen institutions
  • Artificial Intelligence: machine systems outperform traditional instruction methods and pass licensing exams
  • Experiential Learning: virtual apprenticeships and yearlong team projects offer hands on alternatives to standard lectures
  • Declining Perception: the share of americans viewing higher education as very important fell to 31 percent


Andy Kessler

Oct. 4, 2026 2:52 pm ET

image Getty Images

What’s the matter with universities? Syracuse missed its 2026 enrollment target and may run a $30 million deficit. Minnesota is cutting its budget by $225 million over two years. Tulsa is slashing tuition by more than half. Colleges enroll 4.2% fewer undergraduates than last year. Cornell, Yale, Vanderbilt and Washington University in St. Louis put out plans—not radical enough, if you ask me—to reimagine the university. Wossamotta U?

Colleges are known for left-leaning professors, speech suppression, bizarre courses—Occidental College taught “The Unbearable Whiteness of Barbie”—and easy grading more than critical thinking, judgment and problem solving. It’s easy to blame artificial intelligence, but problems run deeper:

• Aptitude. In Griggs v. Duke Power (1971), the Supreme Court effectively banned corporations from giving intelligence tests, so companies rely on colleges to do their sorting via aptitude tests like the SAT. That led to branding and credentialism, but that game is over. A Cornell report noted employers’ “jarring” skepticism of college graduates because they lack job skills such as “the ability to handle uncertainty and solve problems that do not have clear answers.” A Forbes survey found almost half of corporate executives are either less likely to hire Ivy League graduates than they were five years ago or would never hire them.

• Admissions. “Holistic reviews” that ignore aptitude and grades put a thumb on the scale and harm college brands. In combination with athletic scholarships and easy majors, they bring in underqualified students who are churned out as subpar employees.

• Accounting. In 1945, Vannevar Bush suggested the U.S. government fund “basic research in the colleges, universities and research institutes,” which became a big business. But universities got greedy and abused the system by overcharging the government for overhead.

That business model is now suspect. The Trump administration instituted a 15% cap on “indirect cost rates.” A 2025 study found that at more than 350 institutions funded by the National Institutes of Health, the average negotiated indirect cost rate was 58% and the effective actual overhead rate averaged 42%. The Trump administration’s 15% cap was struck down in most cases and the administration has backed off for now, but we’ve seen research funds cut off. Watch for “let’s hit up alumni for more” as research overhead may no longer be an overstuffed cow to pay for campus excesses.

That includes professor salaries. Many eyebrows lifted in 2024 when Claudine Gay was fired as Harvard president and went back to being a professor, at a $900,000 salary. The average full professor salary at Harvard in 2026 is $293,619, up 6.4% from last year. I’d bet many of them hate capitalism but are happy to enjoy its fruits.

• Affordability. Student loans lead to constant tuition hikes. Same for high professor salaries and administrative bloat. Sixteen colleges and universities cost more than $100,000 a year in tuition and expenses. Where’s the value? Students could buy a decent GPU rack and start their own AI company for less.

• Artificial Intelligence. Professors and students seem to be going through the motions of learning, so there’s a huge structural problem. Do ChatGPT-written essays get graded by Claude? Why bother?

Three years ago, OpenAI’s GPT-4 scored 1,410 on the SAT, maybe enough to get into Michigan. Since then, AI has passed bar exams and medical licensing exams, which says more about these outdated tests than the inadequacy of humans, perhaps a sign of an antiquated education system. Even worse is a 2025 Harvard study, “AI tutoring outperforms in-class active learning.” Are teachers and professors becoming obsolete?

Universities can take advantage of AI rather than ban it. The sage-on-a-stage format is over. Learning is experiential. As corporate America rejects college branding, graduates would benefit from more hands-on experience. Universities may be uniquely positioned to implement what I call “virtual apprenticeships,” real job experience while still at school.

I suggest all students spend their entire junior year on team projects. Assignments might include building nuclear autonomous buses or robots that cook. Cross-discipline self-organizing teams will form. You’ll need engineers and coders, but also prelaw and political-science majors to clear the regulatory path, philosophers and marketing majors to write press releases, designers, nutritionists, etc.

Self-paced AI-tutored instruction can help students learn what they need as problems arise. A smart entrepreneur might even create a student transfer portal to recruit talented students from other universities.

According to Gallup, in 2013, 70% of Americans said college was very important. Today it’s only 31%. If universities don’t change, more than a few will implode, while others will wither away over decades. AI and alternative forms of learning are already rising. Something radical is needed to stop the decay of our once great universities. Who will step up?

Write to <a href="mailto:kessler@wsj.com">kessler@wsj.com</a>.

Copyright ©2026 Dow Jones & Company, Inc. All Rights Reserved. 87990cbe856818d5eddac44c7b1cdeb8

Appeared in the October 5, 2026, print edition as 'The End of College as We Know It'.

Andy Kessler is the author of Inside View, a column he writes for The Wall Street Journal on technology and markets and where they intersect with culture. He won the 2019 Gerald Loeb Award for commentary. He is the author of several books including Wall Street Meat and Eat People. He used to design chips at Bell Labs before working on Wall Street for PaineWebber and Morgan Stanley and then as a founder of the hedge fund Velocity Capital.


Videos

Read the whole story
bogorad
1 day ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete

Dyson CameraJet

1 Share

So when I saw Dyson had a $500 toothbrush, I was excited. Finally, advertising that targets me! I love brushing my teeth, and I have more money than I know how to spend. Not because I’m particularly rich, but because most stuff doesn’t really appeal to me.

Like if I owned a helicopter it would just be a headache, because like imagine one day I get a call from the hangar saying the hangar is flooding and the water is rising and you need to move your helicopter. I’m thousands of miles away and need a helicopter pilot in the next 30 minutes, a new place to store it, was the maintenance even done will we even be able to take off on short notice and really I just am upset with myself because I made the poor decision to purchase a helicopter, and once I come back to reality I feel relieved that I don’t own a helicopter and this scenario will never happen to me.


I do however, by means of my birthday, own a Dyson CameraJet (pictured above). It broke within 30 seconds of the first brushing. None of the LEDs turn on anymore. I spent an hour investigating, finally opening the user removable battery compartment to find the Spearmint Dyson Low-foaming mouth rinse had leaked inside. And by how the toothbrush is designed, it’s clear the entire electronics compartment was flooded with the stuff. Here’s the top comment on Reddit about this toothbrush:


Apparently this is happening to everyone, “a potential for water seepage” they say. Dyson wants me to find the receipt and return it through some obtuse process that probably doesn’t work, dude it was a gift I just want my $500 toothbrush to work.

They claim they worked on it for 6 years, but it’s clear their QA Process doesn’t include putting any liquid in the device. It clearly should, ideally for all devices but at least for spot checks on some. It’s sad to see this. At comma, we put every comma four in a highly stressful environment for 16 hours, a superset of the state it’s in driving, while testing all peripherals: the cameras, IMU, GPS, screen, etc… We have gotten the failure rate super low by doing this, and for the few that do fail it’s usually after a while.

There’s no excuse for a mature consumer electronics company to not design a procedure to fully test the functionality of each device before shipping. This shows some serious dysfunction at the company, and they should take this as a wake up call to fix their processes and issue a recall for the toothbrush.



Dyson, if you see this post, e-mail me when I can drop by the Dyson store in ifc mall Hong Kong and swap it for a new one. I don’t want a stupid process, I want a real technical explanation of the issue and a working fancy toothbrush.

Read the whole story
bogorad
1 day ago
reply
Barcelona, Catalonia, Spain
Share this story
Delete

Will the rain derail the housing encampment? The Tenants’ Union asks people to store the tents in Barcelona

1 Comment

Heavy rain is forecast in Catalonia

  • Weather precautions: The Tenants’ Union recommended that participants in the housing encampment at Barcelona’s Plaça de Catalunya seek shelter until tomorrow and protect their tents and canopies.
  • Severe rain forecast: Torrential rain is expected to affect Catalonia’s coastal and pre-coastal areas this weekend.
  • Emergency measures: Mobility restrictions will take effect in 20 Catalan counties, and the Inuncat emergency flood plan has been activated.
  • Decisions remain local: The union said it is part of the encampment but does not decide for all participants, including groups that arrived several days earlier.
  • Participant discretion: Each group will decide how best to respond to the rain alerts and protect people and equipment.
  • Related alerts: Civil Protection issued or planned ES-Alerts covering torrential rain and a possible Mediterranean “mini-hurricane.”
Read the whole story
bogorad
2 days ago
reply
hehehe
Barcelona, Catalonia, Spain
Share this story
Delete
Next Page of Stories