
Nvidia's claim is that a 5 million dollar GB200 NVL72 system generates 75 million dollars of token revenue over three years, a 15x return, using results from an independent benchmark run by SemiAnalysis. Reconstruct it and the maths works: it needs roughly 11 trillion output tokens a year, about 362,000 every second without pause, sold at the price DeepSeek listed for R1 in January 2025. The fragile part is not the hardware. It is the price of a token, which on DeepSeek's own current list ranges from 66 cents to 3.96 dollars per million depending on tier and time of day, and which the same benchmark shows can fall 21x from a single software change. The number is a ceiling presented as a forecast.
The number
If you have read anything about the economics of AI hardware in the last year, you have probably met this sentence: a 5 million dollar investment in an Nvidia GB200 NVL72 system can generate 75 million dollars in token revenue. That is a 15x return.
It comes from Nvidia's own blog, published in October 2025, and Nvidia has repeated it steadily since. The figure caption on the technical version of the post adds the detail that gets dropped in most retellings: the 75 million is over three years, and it is revenue, not profit.
What makes the claim more interesting than ordinary vendor marketing is where the underlying data came from. Nvidia did not measure this itself. The numbers come from InferenceMAX, an open-source benchmark built by the research firm SemiAnalysis, which runs the same models across hundreds of chips from different manufacturers, publishes the results daily, and lets anyone reproduce them. Nvidia gave SemiAnalysis the hardware and then quoted the score.
So this is not a made-up number. It is worth taking seriously, and it is worth taking apart, because the thing the arithmetic depends on is not the part anyone talks about.
What the machine is
First, the object. A GB200 NVL72 is not a server in the sense of a box on a shelf. It is a full rack, about the size of a large wardrobe, weighing roughly 1.4 tonnes, containing 72 Blackwell chips wired together so tightly that software can treat them as one very large processor rather than 72 small ones.
That wiring is the whole point. Big modern models do not fit on one chip, so the chips have to pass work between each other constantly. The more of them that can behave as a single unit, the less time is lost in the handoff.
It draws around 120 to 132 kilowatts, according to Nvidia's own documentation and HPE's product listing. A typical older data centre rack was built for 10 to 30 kilowatts, which is why these things force buildings to be rebuilt around them. Most of that heat is removed with liquid, not air.
On price: SemiAnalysis wrote in February 2026 that a GB200 NVL72 rack can cost 3.3 million dollars. Nvidia's figure is 5 million, which is consistent if you include the networking, storage and installation that a rack on its own does not include. Nothing suspicious there.
The arithmetic, reconstructed
Nvidia does not show its working, so here is ours.
The model in the claim is DeepSeek R1, the Chinese reasoning model that caused a global stir in January 2025. When DeepSeek launched it, the company charged 2.19 dollars per million output tokens. A token is roughly three quarters of a word, and output tokens are the ones the model writes, as opposed to the ones you feed it.
75 million dollars over three years is 25 million dollars a year. At 2.19 dollars per million tokens, that requires 11.4 trillion output tokens a year. Spread across a year of seconds, that is about 362,000 tokens every second, continuously, with no gaps. Divided across the 72 chips, roughly 5,000 tokens per second from each one, forever.
That is a demanding number, but it is not a fantasy number. It sits inside the range InferenceMAX actually measured for this rack when it is tuned for total throughput rather than for individual speed. The arithmetic is internally consistent, which is the first honest thing to say about it.
Here is the second. Nvidia's own technical post states the cost side too: on a GB200 NVL72, producing a million tokens costs just over 10 cents, against 1.56 dollars on the previous generation H200. So the picture Nvidia is painting is a machine that makes a million tokens for about 10 cents and sells them for 2.19 dollars. Almost the entire 15x lives in that gap.
362,000 tokens every second, continuously, for three years. Roughly 5,000 per second from each of the 72 chips, forever.
The bakery
Think of it as an industrial bakery oven that costs 5 million kroner and can bake a loaf for 10 øre.
The salesman's slide says: at 2 kroner a loaf, this oven earns 75 million kroner. Every number on that slide is true. What the slide assumes is that the oven runs at full load every hour of every night for three years, that every loaf finds a buyer the moment it comes out, that the price of bread stays at 2 kroner, and that in year three people still want this particular kind of bread.
None of those are properties of the oven. They are properties of the market the oven sits in. The salesman is quoting an engineering measurement and letting you read it as a business forecast.
This is the ordinary shape of capital equipment marketing and there is nothing scandalous about it. It only becomes a problem when the number gets repeated by people who have stopped noticing which half is measured and which half is assumed.
A token does not have one price
The weakest assumption is the price. Not because prices are falling in a simple way, which is the lazy version of this argument, but because a token has never had one price.
Look at DeepSeek's current published price list. It no longer sells R1. It sells two models. The cheaper one, V4 Flash, charges 66 cents per million output tokens off-peak and 1.32 dollars at peak. The larger one, V4 Pro, charges 1.98 dollars off-peak and 3.96 at peak. Peak hours are defined as specific windows on weekdays, and off-peak rates are exactly half.
Run the same three years of tokens through those prices instead. At the cheapest rate the output is worth about 23 million dollars. At the most expensive it is worth about 135 million. Nvidia's 75 million sits in the middle of a range that spans six times over, and every point in that range comes from one company's own price sheet on one day.
That is the chart at the top of this piece. The scaling is our arithmetic, not Nvidia's, and it compares different models, so treat it as an illustration of the spread rather than a forecast. The spread is the point.
What the same benchmark says now
InferenceMAX has since been renamed InferenceX and kept running. In February 2026 SemiAnalysis published its second version, covering every Nvidia data centre chip from the last four years and every AMD one from the last three, using close to a thousand GPUs for a full benchmark sweep. Two findings from it sit awkwardly next to the 15x slide.
The first is how much the cost of making a token moves without any new hardware. Take DeepSeek R1 on a GB300 rack, serving each user at 150 tokens per second. The baseline cost is about 2.35 dollars per million tokens. Turn on one optimisation called multi-token prediction, where the model guesses several words ahead and checks them in one pass, and the cost drops to about 11 cents. That is a 21-fold change from a software switch, on the same metal.
The second is how much the speed you promise your users changes the price. On a B200 running the same model, serving at 50 tokens per second per user costs about 56 cents per million output tokens. Push to 125 tokens per second and it rises to around 4 dollars. Going 2.5 times faster costs roughly 7 times more. SemiAnalysis uses this to explain why Anthropic's Opus fast mode is priced 6 to 12 times higher for about 2.5 times the speed, and points out that no new chip is involved.
So the cost of a token is not a fixed property of a rack. It is a number you choose, by deciding how fast to serve people and how much engineering to do. A three-year revenue projection built on any single point of that curve is a snapshot presented as a plan.
A 21-fold change in cost per token, from a software switch, on the same metal.
What the operators actually earn
The most useful thing in the InferenceX report is not a benchmark at all. It is a worked example of a real company's margins.
SemiAnalysis took public data showing that Crusoe, a cloud provider, serves DeepSeek at 36 tokens per second per user for 1.35 dollars per million input tokens and 5.40 dollars per million output tokens. Cross-referencing that against measured costs, and counting depreciation of the hardware as a cost, it estimates Crusoe earns up to 83 per cent gross margin on input tokens and 45 per cent on output tokens.
SemiAnalysis is careful to say those assumptions may not be exactly right and that the calculation ignores downtime and underused capacity, which are precisely the things that eat real returns. Still, this is the closest thing the industry has to a published margin, and it came from an outside firm reverse-engineering it rather than from any provider disclosing it.
One more number from the same report, for perspective on who is doing best out of all this. SemiAnalysis puts Nvidia's own gross margin at around 75 per cent, roughly a fourfold markup on the cost of making the chips. The company selling the 15x return is running better economics than anyone it sells to.
Is any of this new?
The genre is old. Mainframe vendors published cost-per-transaction slides. Telecoms companies laying fibre in the late 1990s published revenue-per-mile projections that assumed the traffic would arrive. Shale drillers published well economics at one oil price. In every case the engineering was sound and the assumption about the sale price was the part that broke.
What is genuinely different this time is the direction of the two prices. The thing the machine makes is getting cheaper fast, through both competition and software. The machine itself is getting more expensive: Fortune reported on 22 August that Nvidia is raising prices by around 15 per cent on systems shipping early next year, including its next-generation Vera Rubin and current Grace Blackwell lines, and the rental price of GPUs has been climbing rather than falling.
A 15x return calculated when your input cost is rising and your output price is under pressure is the best case, not the base case. That is a fair thing for a vendor to publish. It is not a fair thing for a buyer to plan on.
What to watch
One thing, and it is the same thing we flagged a week ago when writing about the companies that sell AI tokens wholesale: a published gross margin.
Not a benchmark score, not a return-on-investment slide, not a revenue run rate. An actual margin, disclosed by a company that buys these racks and sells what they produce, audited if possible. Nobody in the category has published one. Anthropic's stock market filing, when the public version arrives, may be the first document that forces some of this into daylight.
Until then, the honest summary of the 15x claim is this. The measurement is real and independently produced. The machine is genuinely a large step up from the one before it. And the revenue is a ceiling calculated at one price, on one model, at full utilisation, in a market where the price of the product has a six-fold spread on a single vendor's own website.
Curious about AI? Come build with us.
Oslo Vibe Coding runs free, beginner-friendly drop-ins where we build real things with AI. No one codes alone.