An H100 now rents for $4.50 an hour. Idle cards lose to the API.
Nebius raised H100 rent to $4.50 an hour on October 1. Serving one request at a time, a self-hosted 70B model now costs more than OpenAI Sol in batch.
By the numbers
Nebius raised its on-demand H100 rent to $4.50 an hour on October 1, 2026, up from $3.85. Two days earlier, OpenAI launched GPT-6.1 Sol at $2 per million input tokens and $10 per million output, a fifth of what its flagship charges. Rent went up while tokens got cheaper. If you run a small AI product on a rented card, which side you land on depends on one number: how many requests your GPU serves at the same time.
Our math, from published prices and a public benchmark: a 70B open model on one H100, answering one request at a time, now costs about $2.67 per 1,000 typical requests. Sol costs $3.58 at standard rates and $1.79 through its batch tier. Serve ten requests at once on the same card and your cost falls to $0.49. The rent hike hurts the operator whose GPU sits idle. The busy one barely notices.
What Nebius changed on October 1
Nebius lists four new on-demand rates on its pricing page, effective October 1, 2026:
- H100: $3.85 to $4.50 per GPU-hour
- H200: $4.50 to $5.40
- B200: $7.15 to $8.50
- B300: $7.85 to $9.50
That is the second rise since spring. The Motley Fool reported on September 27 that the same pricing page showed the H100 at $2.95 in early May. From $2.95 to $4.50 is a 53% rise in five months.
The company says demand pulls it there. Its CFO told investors "we sold out of capacity because, as fast as we bring capacity online, we can sell it", and the Fool counts $37.5 billion of remaining performance obligations as of June against $582 million of quarterly revenue. Investors liked the price card: Nebius stock rose 10% to $229.67 on the news.
Nebius is not the most expensive option either. A comparison published by Spheron, a GPU marketplace that competes with Nebius, lists CoreWeave's H100 at $6.16 an hour and AWS Capacity Blocks for the B200 at $12.355. If cards get scarcer, $4.50 may look cheap by spring.
What the API side did in the same week
OpenAI released GPT-6.1 Sol on September 29 at DevDay. Two separate write-ups, DataCamp's and eesel's, give the same price card:
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| GPT-6.1 Sol | $2.00 | $10.00 |
| GPT-6.1 Sol, batch | $1.00 | $5.00 |
| GPT-6 Astra | $10.00 | $50.00 |
Cached input on Sol costs $0.10 per million. DataCamp reports Sol matching Astra on the DeepSWE coding benchmark and trailing it on security and lab-science tests. For routine business work, the frontier price just dropped 80%.
The cost per request, worked out
A GPU-hour and a token are different units. To compare them you need a throughput figure. We used a 2026 benchmark from Spheron: Llama 3.3 70B in FP8 on one H100 with vLLM 0.18, prompts averaging 512 input tokens and 256 output tokens.
| Requests served at once | Output tokens per second |
|---|---|
| 1 | 120 |
| 10 | 650 |
| 50 | 1,850 |
| 100 | 2,400 |
At one request at a time, each answer takes about 2.1 seconds, so the card serves roughly 1,690 requests an hour. At $4.50 an hour that's $2.67 per 1,000 requests. The same 1,000 requests on Sol, at 512 tokens in and 256 out, cost $1.02 of input plus $2.56 of output: $3.58. Batch halves it to $1.79.
View as table
| label | Cost per 1,000 requests (512 tokens in, 256 out) |
|---|---|
| GPT-6.1 Sol, standard | $3.58 |
| H100 at $4.50, 1 request at a time | $2.67 |
| GPT-6.1 Sol, batch | $1.79 |
| H100 at $4.50, 10 at a time | $0.49 |
| H100 at $4.50, 50 at a time | $0.17 |
Three things come out of that table.
One stream at a time now loses to batch. In May, at $2.95 an hour, the single-stream card cost $1.75 per 1,000 requests, level with what Sol's batch tier charges today. At $4.50 it costs 49% more than batch. If your workload can wait for a batch job (overnight summaries, tagging, back-office drafts), the API wins this month on price and gives you a stronger model.
Idle hours are the real bill. The $2.67 assumes the card works every second you pay for it. Nebius bills by the hour whether requests arrive or not. A single-stream card busy 75% of the time costs about $3.56 per 1,000 requests, the same as Sol at full price. Busy a quarter of the time, it costs $10.67. Most side projects and small agencies sit closer to the second number than the first.
Concurrency wipes out the hike. Ten requests at once cut the cost to $0.49 per 1,000, seven times cheaper than Sol standard even after the 53% rise. At fifty, it's $0.17. A product with steady traffic, like a support bot or a document pipeline that queues work, still self-hosts at a deep discount. For that operator, the October price card raises a $0.42 cost to $0.49. That is noise.
What this means if you're choosing now
Pick by traffic shape, then by model.
- Spiky or low traffic, one user at a time: the API is cheaper per request once idle hours count, and Sol is a stronger model than a 70B open one. Pay per token.
- Work that can wait hours: Sol batch at $1 / $5 beats a single-stream rented H100 outright.
- Steady traffic you can queue to ten or more concurrent requests: renting still wins by a wide margin, and the hike is a rounding error. Lock in a rate if your provider offers one; Nebius has raised twice in five months.
We covered the other side of this trade in August, when OpenRouter's 5.5% fee to load credits added a cost layer on top of the token price. That fee applies here too if you buy Sol through a router instead of directly. And the scarcity behind Nebius's hike traces back to the chip supply chain we looked at in Nvidia's $279B in supply commitments.
What could make this math wrong
The throughput figures are one vendor's benchmark, and Spheron rents GPUs, so it has a reason to show cards in a good light. Your own prompts may be longer, your model smaller or larger, your serving stack slower. Llama 3.3 70B is also a weaker model than Sol: on hard coding or reasoning tasks, a cheaper wrong answer costs more than a pricier right one. And OpenAI can move its price card as fast as Nebius moved its own. Run the same arithmetic on your own logs before you move a workload: requests per hour, average tokens in and out, and the share of paid hours your card actually works.
Sources
- Nebius AI Cloud: compute pricing
- The Motley Fool: Nebius is raising the price of its AI compute on Oct. 1 (September 27, 2026)
- Yahoo Finance: IREN climbs 6%, Nebius jumps 10% as GPU rental rates head higher
- Spheron: Nebius pricing 2026, H100 to B300 rates rise up to 21%
- DataCamp: GPT-6.1 Sol features, benchmarks, pricing
- eesel: GPT-6.1 Sol pricing
- Spheron: vLLM vs TensorRT-LLM vs SGLang, H100 benchmarks 2026
This is not financial advice.