Skip to the stories

An H100 now rents for $4.50 an hour. Idle cards lose to the API.

Nebius raised H100 rent to $4.50 an hour on October 1. Serving one request at a time, a self-hosted 70B model now costs more than OpenAI Sol in batch.

The Editors · AI money · 7 min read · Oct 2, 2026

Black Gigabyte graphics card

By the numbers

$4.50Nebius on-demand H100 per GPU-hour from October 1, 2026
+53%H100 rent since early May, when it was $2.95
$2 / $10GPT-6.1 Sol per million input / output tokens
120 vs 2,400output tokens per second on one H100, 1 request vs 100 at once

Nebius raised its on-demand H100 rent to $4.50 an hour on October 1, 2026, up from $3.85. Two days earlier, OpenAI launched GPT-6.1 Sol at $2 per million input tokens and $10 per million output, a fifth of what its flagship charges. Rent went up while tokens got cheaper. If you run a small AI product on a rented card, which side you land on depends on one number: how many requests your GPU serves at the same time.

Our math, from published prices and a public benchmark: a 70B open model on one H100, answering one request at a time, now costs about $2.67 per 1,000 typical requests. Sol costs $3.58 at standard rates and $1.79 through its batch tier. Serve ten requests at once on the same card and your cost falls to $0.49. The rent hike hurts the operator whose GPU sits idle. The busy one barely notices.

What Nebius changed on October 1

Nebius lists four new on-demand rates on its pricing page, effective October 1, 2026:

  • H100: $3.85 to $4.50 per GPU-hour
  • H200: $4.50 to $5.40
  • B200: $7.15 to $8.50
  • B300: $7.85 to $9.50

That is the second rise since spring. The Motley Fool reported on September 27 that the same pricing page showed the H100 at $2.95 in early May. From $2.95 to $4.50 is a 53% rise in five months.

The company says demand pulls it there. Its CFO told investors "we sold out of capacity because, as fast as we bring capacity online, we can sell it", and the Fool counts $37.5 billion of remaining performance obligations as of June against $582 million of quarterly revenue. Investors liked the price card: Nebius stock rose 10% to $229.67 on the news.

Nebius is not the most expensive option either. A comparison published by Spheron, a GPU marketplace that competes with Nebius, lists CoreWeave's H100 at $6.16 an hour and AWS Capacity Blocks for the B200 at $12.355. If cards get scarcer, $4.50 may look cheap by spring.

What the API side did in the same week

OpenAI released GPT-6.1 Sol on September 29 at DevDay. Two separate write-ups, DataCamp's and eesel's, give the same price card:

ModelInput / 1M tokensOutput / 1M tokens
GPT-6.1 Sol$2.00$10.00
GPT-6.1 Sol, batch$1.00$5.00
GPT-6 Astra$10.00$50.00

Cached input on Sol costs $0.10 per million. DataCamp reports Sol matching Astra on the DeepSWE coding benchmark and trailing it on security and lab-science tests. For routine business work, the frontier price just dropped 80%.

The cost per request, worked out

A GPU-hour and a token are different units. To compare them you need a throughput figure. We used a 2026 benchmark from Spheron: Llama 3.3 70B in FP8 on one H100 with vLLM 0.18, prompts averaging 512 input tokens and 256 output tokens.

Requests served at onceOutput tokens per second
1120
10650
501,850
1002,400

At one request at a time, each answer takes about 2.1 seconds, so the card serves roughly 1,690 requests an hour. At $4.50 an hour that's $2.67 per 1,000 requests. The same 1,000 requests on Sol, at 512 tokens in and 256 out, cost $1.02 of input plus $2.56 of output: $3.58. Batch halves it to $1.79.

$3.58$2.67$1.79$0.49$0.17
View as table
Cost per 1,000 requests (512 tokens in, 256 out)
labelCost per 1,000 requests (512 tokens in, 256 out)
GPT-6.1 Sol, standard$3.58
H100 at $4.50, 1 request at a time$2.67
GPT-6.1 Sol, batch$1.79
H100 at $4.50, 10 at a time$0.49
H100 at $4.50, 50 at a time$0.17
Source: Market Anarchy calculation from Nebius and OpenAI price cards and Spheron's H100 benchmark, October 2026

Three things come out of that table.

One stream at a time now loses to batch. In May, at $2.95 an hour, the single-stream card cost $1.75 per 1,000 requests, level with what Sol's batch tier charges today. At $4.50 it costs 49% more than batch. If your workload can wait for a batch job (overnight summaries, tagging, back-office drafts), the API wins this month on price and gives you a stronger model.

Idle hours are the real bill. The $2.67 assumes the card works every second you pay for it. Nebius bills by the hour whether requests arrive or not. A single-stream card busy 75% of the time costs about $3.56 per 1,000 requests, the same as Sol at full price. Busy a quarter of the time, it costs $10.67. Most side projects and small agencies sit closer to the second number than the first.

Concurrency wipes out the hike. Ten requests at once cut the cost to $0.49 per 1,000, seven times cheaper than Sol standard even after the 53% rise. At fifty, it's $0.17. A product with steady traffic, like a support bot or a document pipeline that queues work, still self-hosts at a deep discount. For that operator, the October price card raises a $0.42 cost to $0.49. That is noise.

What this means if you're choosing now

Pick by traffic shape, then by model.

  • Spiky or low traffic, one user at a time: the API is cheaper per request once idle hours count, and Sol is a stronger model than a 70B open one. Pay per token.
  • Work that can wait hours: Sol batch at $1 / $5 beats a single-stream rented H100 outright.
  • Steady traffic you can queue to ten or more concurrent requests: renting still wins by a wide margin, and the hike is a rounding error. Lock in a rate if your provider offers one; Nebius has raised twice in five months.

We covered the other side of this trade in August, when OpenRouter's 5.5% fee to load credits added a cost layer on top of the token price. That fee applies here too if you buy Sol through a router instead of directly. And the scarcity behind Nebius's hike traces back to the chip supply chain we looked at in Nvidia's $279B in supply commitments.

What could make this math wrong

The throughput figures are one vendor's benchmark, and Spheron rents GPUs, so it has a reason to show cards in a good light. Your own prompts may be longer, your model smaller or larger, your serving stack slower. Llama 3.3 70B is also a weaker model than Sol: on hard coding or reasoning tasks, a cheaper wrong answer costs more than a pricier right one. And OpenAI can move its price card as fast as Nebius moved its own. Run the same arithmetic on your own logs before you move a workload: requests per hour, average tokens in and out, and the share of paid hours your card actually works.

Sources

This is not financial advice.

Share this piece

More from AI money