GPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete. Ornn Data — a firm that publishes GPU rental indices and also rents GPUs — argues the other way: open-weight demand gives older silicon a class of workloads it can still serve cheaply, and the contract-price curve is already saying so.
The chain starts with how closed models are sold. Access runs through subscription allowances the provider can reset, so the posted token rate card is the marginal price of additional usage, not a commodity price that falls on its own. Open weights remove the gate: the same checkpoint can be served by anyone. September OpenRouter snapshots Ornn cites listed eighteen to twenty-two providers for several widely served open models, with highest-to-lowest output prices spread 1.8x to 5.6x. The paper’s framing of the real distinction is that “the economic distinction is who can make the deployment decision.”
Self-hosting turns that price spread into compute-only cost — spot GPU-hour rent divided by achieved throughput, adjusted for utilization and headroom — and lands at $0.12 to $0.35 per million output tokens at full utilization. The interesting part is that the hardware ranking reverses by workload. On gpt-oss-120b, a sparse mixture-of-experts model with 5.1B active parameters, the A100 comes out cheaper than the H100 at spot and at the three- and five-year term prices. On dense Llama-2-70B, newer hardware wins.
Then the market evidence. A100 occupancy rose from 74% in March 2026 to 90% in September, with listed capacity up 13% and the spot index up 20% — supply and price rising together, which is the tell that demand moved rather than scarcity. Ornn’s term curves show the A100’s five-year mark retaining 80.2% of its one-month price for a contract ending in September 2031, when Ampere is 11.3 years old. H100 retains 59.8% over the same tenor, B200 53.8%, H200 43.7%.
Why this can hold: long-running agents, batch evaluation, and parts of reinforcement learning tolerate latency and are hardware-agnostic. Price-elastic demand routes to whatever serves the work most cheaply. Older hardware does not need to win every workload to keep a role.
Two things keep this from being the clean story the abstract implies.
First, the paper’s own limitations section is unusually thorough, and it lands on the headline. The A100-beats-H100 sparse result combines two different third-party serving setups. The dense A100 row is estimated from a vendor throughput ratio, not a matched MLPerf submission, and the paper says alternative precision assumptions can reverse its ranking. Matching models at an equal Artificial Analysis Index score does not establish equal task success. And on the central causal claim they simply decline: the paper does not establish that open-weight demand caused the A100 occupancy or rental-price behavior — device retirements, dense inference, non-LLM work, and financing effects remain live alternatives. Forward marks are analyst-produced indicators, not executable quotes, and occupancy tracks rental status across tracked providers, not the installed base.
Second, the commercial interest. Ornn publishes the paper and licenses the rental index, occupancy series, and term marks used in it, and rents GPUs through Ornn Compute. All of it is disclosed. It is also a reason to treat “old GPUs hold value” as a thesis two of their businesses want to be true.
One footnote the paper treats as validation: the abstract cites NVIDIA’s acquisition of Hugging Face on 3 September 2026 — confirmed at $12.93B, NVIDIA’s largest software acquisition — as evidence that open weights matter strategically. That reading is fair, but it cuts both ways. The distribution chokepoint for open checkpoints now sits inside the company that sells the GPUs and profits from their useful life.
Strip the causal story and what’s left is still worth keeping: the term-price curves, and the reframing that the right measure of hardware usefulness is the cost of the work it can serve and the demand for that work — “not simply the age of the chip.”