The $100 Frontier Compute Problem

Community Article
Published September 29, 2026

What happens when AI stops answering questions and starts working for hours?

I have been thinking about the economics of frontier coding models.

My own usage is a useful example.

I use GPT-5.6 Sol Medium heavily for software development. A normal month for me is roughly eight hours a day across 25 working days: about 200 hours of access.

I pay around $100 per month.

Now compare that with trying to reproduce something in roughly the same capability class using open weights and rented GPUs.

The comparison is imperfect. A ChatGPT subscription is a shared, quota-limited service. Renting GPUs gives me dedicated hardware. And GLM-5.3 is not GPT-5.6 Sol.

But the numbers expose something interesting about frontier AI economics.

How close is the open-weight alternative?

GLM-5.3 is one of the strongest open-weight coding models available today. It has 753 billion total parameters, with around 40 billion active parameters.

On Z.ai's published evaluations, GLM-5.3 scored 88.2% on Terminal-Bench 2.1, compared with 88.8% for GPT-5.6 Sol.

On the harder DeepSWE v1.1 benchmark, GLM-5.3 scored 66.9% versus 72.7% for Sol.

On Terminal-Bench 3.0 the difference was 28.3% versus 34.6%.

So we are not talking about a tiny local coding model compared with a frontier proprietary system. On some agentic coding workloads, GLM-5.3 is genuinely in the same conversation.

But the weights alone are roughly 756 GB in the released FP8 model.

That changes the economics quickly.

The self-hosting bill

An NVIDIA H200 has 141 GB of HBM3e memory.

Six H200s give 846 GB of VRAM. Once runtime overhead, KV cache, context and other memory requirements are considered, six to seven H200s is a reasonable order-of-magnitude configuration for serving a model this large.

RunPod currently lists multi-GPU H200 cluster capacity at about $4.31 per GPU-hour.

For my approximately 200 working hours each month:

6 × $4.31 × 200 = $5,172/month

Seven GPUs:

7 × $4.31 × 200 = $6,034/month

Compare that with my approximately $100/month ChatGPT subscription.

That is a difference of roughly 50–60×.

It obviously does not mean OpenAI spends $5,000 per month serving me.

I am comparing shared infrastructure with dedicated infrastructure.

That distinction is precisely what makes the economics interesting.

Squeeze the model and the gap gets smaller

Suppose instead that I use GLM-5.3 Flash.

Flash has roughly 320 billion total parameters but activates only around 18 billion for each token. Quantize the weights aggressively to around four bits and it may be possible to squeeze the model onto something like a 192 GB AMD MI300X, depending on implementation, context size and runtime overhead.

A RunPod MI300X currently costs roughly $2.39/hour.

For 200 hours:

$2.39 × 200 = $478/month

Now the gap is around 5×, rather than 50×.

But I have achieved that by accepting compromises: quantization, tighter memory margins, potentially lower quality, my own inference stack, deployment work and hardware-specific operational issues.

Meanwhile, the $100 subscription gives me access to the frontier model without any of that infrastructure.

That is an unusually good deal for a heavy user.

There is another important number: how much work the model does

Parameter count is only part of the story.

Artificial Analysis currently measures GLM-5.3 Max as producing roughly 71,000 output tokens per task on its Intelligence Index, compared with roughly 8,000 for GPT-5.6 Sol Medium.

Its estimated cost per task is about $2.01 for GLM-5.3 Max versus $0.50 for Sol Medium, even though GLM's raw token prices are much lower.

This matters.

The economics of an agent are determined by more than:

model size × token price

They increasingly look like:

reasoning tokens
× agent turns
× tool calls
× context processing
× retries
× parallel agents
× task duration

A model that solves the same problem in fewer iterations can be more economically useful even if each individual token is more expensive.

That is why agent efficiency is becoming as important as benchmark performance.

Agents change the unit of consumption

Traditional ChatGPT interaction looks something like this:

one user → one prompt → one response

Agentic computing looks more like this:

one user → one objective → dozens or hundreds of model actions

A coding agent may inspect 40 files, search a repository, modify three files, launch tests, inspect the failure, search documentation, change the implementation, rerun the tests and then ask another agent to review the result.

ChatGPT Work extends that pattern beyond programmers.

An agent can search files, operate applications, analyse information, generate documents and keep a project running for hours or even days.

The model isn't merely answering.

It is consuming compute while doing work.

And OpenAI's own numbers show how quickly this behaviour is developing.

The usage curve is already moving

In June, OpenAI reported more than 5 million weekly Codex users, up more than 6× since the desktop app launched in February. Around 20% were already non-developers.

By May 2026, more than 70% of individual Codex users had submitted at least one task estimated to represent more than an hour of human work.

Enterprise adoption is moving even faster outside engineering.

Between February and August, weekly active enterprise Codex users grew:

108× in legal
41× in sales
41× in recruiting
26× in marketing
5× in engineering

By June, Codex alone accounted for 64% of combined ChatGPT and Codex output tokens among OpenAI enterprise customers.

And the distribution is extremely uneven.

OpenAI says its top 10% of enterprise customers by AI usage now generate 8.3× as many output tokens per active user as typical firms. In January that gap was only 2.6×.

That is almost exactly the pattern I would expect from agents.

The average user matters less than the power user running several expensive workflows simultaneously.

OpenAI's own researchers show where this could go

Perhaps the most interesting numbers come from inside OpenAI itself.

By mid-August 2026, OpenAI said the median researcher was consuming the equivalent of more than $600 per day at API prices through coding agents.

The 90th-percentile researcher was consuming more than $7,000 per day.

Across the research organisation, agent runtime had reached the equivalent of 3.1 agent working days for every human working day.

That is a glimpse of what happens when agents become genuinely useful.

People don't necessarily use less compute because the model gets better.

They discover more work worth delegating.

Better agents can therefore create their own demand.

GPT-5.6 Sol already demonstrated the problem

Tibo Sottiaux's July update about GPT-5.6 Sol was particularly interesting.

Users were burning through Codex allowances faster than expected.

OpenAI said it had not deliberately reduced plan limits. Sol was simply more willing to keep working: longer runs, more tool calls and more coordination between tools and subagents.

After optimisations, OpenAI expected normal Sol allowances to last around 18% longer.

That is a small example of a much larger issue.

The thing that makes a model better as an agent—its willingness to continue working—can also make it substantially more expensive to serve.

And OpenAI itself now describes the world as "compute-constrained", saying demand is growing faster than capacity and that routing, scheduling, kernels, caching and model implementation have become central to serving GPT-5.6 efficiently.

But economics are only half of the problem

It is tempting to assume that sufficiently good economics will solve everything.

GPUs become cheaper. Models become more efficient. Datacentres become larger. OpenAI raises another round of capital.

But underneath the economics there is physics.

A GPU consumes electricity.

Electricity becomes heat.

That heat has to go somewhere.

Power must reach the site.

Transformers and switchgear have to exist.

Transmission capacity has to exist.

Cooling systems need electricity and, depending on their design, water.

Suitable land needs to be acquired and permitted.

None of these scale at software speed.

The power numbers are getting large

The International Energy Agency expects worldwide datacentre electricity consumption to more than double to around 945 TWh by 2030.

That is slightly more electricity than Japan currently consumes.

AI-optimised datacentres are expected to be the fastest-growing part of that load. The IEA expects electricity use from accelerated servers to grow around 30% annually in its base case.

The United States illustrates the local problem even more clearly.

U.S. datacentres consumed about 176 TWh in 2023, or 4.4% of U.S. electricity.

Berkeley Lab now estimates that datacentres could reach around 11.8% of total U.S. electricity consumption by 2030, with scenarios ranging from 9.5% to 15.3%.

This isn't an abstract future resource problem.

Google has said grid connection delays in some U.S. locations can reach 12 years.

High-voltage transformer lead times have reached as much as 160 weeks.

You cannot solve a 160-week transformer lead time with a software update.

And then there is water

Cooling design can reduce direct water use substantially, and not every datacentre relies heavily on evaporative cooling.

But at system level the numbers are still significant.

Berkeley Lab estimated direct U.S. datacentre water consumption at roughly 65 billion litres in 2023, with its 2028 scenarios rising substantially depending on datacentre growth and cooling technology.

We are also seeing the trade-offs directly.

OpenAI and NEXTDC's planned 612 MW Sydney facility dropped a recycled-water cooling plan after the required water infrastructure could not be approved and built, instead moving toward a waterless design that uses more electricity.

The constraint did not disappear.

It moved from water to power.

People also have to live next to these things

Datacentres are physical industrial facilities.

They generate noise, require transmission infrastructure, alter local development and can compete for electricity and water.

That creates another constraint that spreadsheets tend to ignore: social licence.

In September, Goodman withdrew plans for a A$1.2 billion, 90 MW datacentre in Lane Cove, Sydney, after substantial local opposition. The proposed facility would have been around 20 metres from some homes. Concerns included noise, power and water consumption and local environmental effects.

In the U.S., Reuters reported coordinated protests against datacentre construction across 42 states in July.

Capital can buy GPUs.

It cannot instantly manufacture grid capacity, permits, transmission lines or community acceptance.

Scaling is also producing diminishing economic returns

There is another uncomfortable issue.

The original neural scaling laws did not say performance grows linearly with compute.

They found that training loss follows a power law as compute, data and model size increase.

That is an important distinction.

It means the claim that "cost increases geometrically while intelligence increases logarithmically" is not technically correct.

But the underlying intuition is useful: pushing farther along an existing scaling curve can require large increases in compute for progressively smaller reductions in loss.

And frontier training compute has been growing extraordinarily quickly. Stanford's 2025 AI Index reported that training compute for notable AI systems had been doubling roughly every five months.

There is an opposing trend.

The same report found that ML hardware performance had been improving around 43% annually, price-performance around 30% annually, and energy efficiency around 40% annually.

So this is not a simple story of compute becoming infinitely expensive.

We have two curves racing each other:

demand for compute is exploding

while

hardware and algorithmic efficiency are improving

The question is which curve dominates.

Agentic workloads make that question much harder because they create new demand as quickly as efficiency frees capacity.

Hardware itself is a wasting asset

There is also a financial problem hidden underneath all this infrastructure.

AI servers do not have the economics of a hydroelectric dam.

Hyperscalers generally depreciate server and networking equipment over roughly five or six years.

Amazon actually shortened the estimated useful life of part of its server fleet from six years back to five years in 2025, explicitly citing the faster pace of AI and machine-learning technology development.

Meta uses approximately 5.5 years for most server and networking assets.

Microsoft reports server and network equipment lives ranging from two to six years.

Accounting life and technical usefulness are different things.

An H100 does not suddenly become useless when a B200 arrives.

Older accelerators can continue serving smaller models and less latency-sensitive inference.

But their relative economic value can deteriorate quickly when a new generation produces significantly more tokens per watt, more memory capacity or better interconnect performance.

That makes AI infrastructure unusual.

Companies are spending tens or hundreds of billions of dollars on assets whose competitive value can fall significantly before the building containing them has finished depreciating.

Datacentres themselves create another mismatch

A chip can be replaced.

A partially completed 500 MW datacentre cannot easily become something else.

Land, transmission agreements, substations, cooling infrastructure and construction capital are location-specific and illiquid.

And delays are becoming financially meaningful.

Oracle recently invoked a force-majeure provision relating to potential power delays at the $18 billion-financed Project Jupiter datacentre development associated with Stargate in New Mexico.

This is where the AI infrastructure story begins to resemble energy and heavy industry more than software.

Vertical integration matters—but the picture is complicated

There is also a structural difference between AI labs.

Google has spent years building its own TPU accelerator architecture and integrating hardware, networking, datacentres, software and models.

But even Google does not manufacture leading-edge silicon itself. Semiconductor fabrication remains concentrated in specialist foundries.

xAI similarly relies heavily on NVIDIA hardware rather than owning semiconductor manufacturing.

OpenAI is moving further down the stack. In June it announced its first custom inference accelerator, developed with Broadcom and manufactured by TSMC, with deployment planned for the end of 2026.

Anthropic remains much more dependent on external infrastructure.

It works closely with Amazon on Trainium, and says it already uses more than one million Trainium2 chips. In April it announced an agreement with Amazon covering up to 5 GW of compute and more than $100 billion of AWS technology commitments over ten years.

That dependence has now become one of the central financial risks in Anthropic's IPO filing.

Anthropic's prospectus puts numbers on the problem

The numbers disclosed this week are remarkable.

According to the confidential IPO prospectus reviewed by Reuters, Anthropic generated approximately $4.6 billion in revenue in 2025.

Its operating loss was approximately $8.06 billion.

It spent $7.33 billion on compute and infrastructure during that year alone, around three times its 2024 spending.

Its reported net loss was almost $42 billion, although roughly $34 billion of that was an accounting charge associated with financing instruments, rather than operating cash burn.

Anthropic is reportedly targeting a valuation above $2 trillion.

More strikingly, the company has disclosed around $518 billion of planned cloud, computing and infrastructure commitments over the coming decade.

Approximately 80% of those commitments are non-cancelable or require payment regardless of actual usage.

Among them are commitments of roughly:

$111.1B to Google
$110B to Amazon
$31.4B to Microsoft
$161.2B in Broadcom-related equipment leases

Anthropic itself told prospective investors that future AI growth could be constrained "principally by the availability of compute."

Whether a $2 trillion valuation ultimately makes sense is a decision for investors.

But the prospectus makes the economic bet unusually clear.

This is no longer a conventional software company cost structure.

It is a software company making infrastructure commitments on the scale of heavy industry.

There is no escape from the physical layer

This is why I don't think the future can be analysed purely through token prices.

Maybe inference becomes ten times more efficient.

History suggests enormous efficiency improvements are possible. The cost of reaching roughly GPT-3.5-level MMLU performance fell from around $20 per million tokens in November 2022 to $0.07 by October 2024, according to Stanford's AI Index—a more than 280× reduction.

That is extraordinary.

But if agent usage grows 100× at the same time, the physical problem remains.

And OpenAI's current enterprise numbers—108× growth in legal Codex users, 41× in sales, increasingly parallel agents and hours-long tasks—show how that can happen.

Efficiency does not automatically reduce resource consumption when cheaper compute causes people to consume much more of it.

This is the real $100 question

So I don't think the important question is:

Will OpenAI run out of GPUs?

That is too simplistic.

OpenAI can build more infrastructure.

It can deploy custom silicon.

Models will become more efficient.

Inference software will improve.

Smaller models will handle easier work.

Caching and batching will improve.

Older hardware will move down the workload stack.

But there are limits to how quickly physical infrastructure can follow exponential software demand.

A power station is not an API call.

A transmission line cannot be deployed overnight.

A semiconductor fab takes years and tens of billions of dollars.

A transformer with a 160-week lead time cannot be parallelised by launching another agent.

And a neighbourhood that doesn't want a 90 MW datacentre next door is not a CUDA optimisation problem.

That makes the $100 subscription question increasingly interesting.

Today a serious developer can obtain hundreds of hours of access to frontier intelligence for roughly the price of a conventional software subscription.

The equivalent dedicated open-weight infrastructure can cost several times more even under aggressive quantization, and tens of times more when running a genuinely frontier-scale model.

That gap exists because shared infrastructure is extraordinarily efficient.

The risk is that agents change the workload faster than infrastructure economics can compensate.

We are moving from:

AI answering questions

to:

AI working for minutes

to:

AI working for hours

to:

multiple AIs working continuously in parallel.

At that point, intelligence stops behaving like ordinary SaaS.

The marginal cost matters again.

Electricity matters.

Memory bandwidth matters.

Datacentre construction matters.

Capital cost matters.

Hardware depreciation matters.

And physics matters.

The next frontier-model race therefore may not be decided simply by who produces the highest benchmark score.

GPT-5.6 Sol already scores 88.8% on Terminal-Bench 2.1. GLM-5.3 reaches 88.2%. We are reaching a point where multiple models can perform extremely well on benchmarks that looked difficult only recently.

Stanford's 2026 AI Index found that four major labs were already clustered within only 25 Arena Elo points, while the gap between the leading closed and open models was only a few percentage points.

The scarce resource may increasingly be something else:

how much useful intelligence can you deliver per watt, per dollar and per unit of scarce datacentre capacity?

And for heavy users like me, there is an even simpler version of that question:

For how long can $100 continue buying hundreds of hours of frontier-level agent work?

Community

Sign up or log in to comment