Nvidia Vera Rubin vs Blackwell: What a 10x Cut in AI Costs Means for NVDA Stock

Aadi Bihani Image

Aadi Bihani

Last updated:
17 min read
Vera Rubin vs Blackwell: Can Cheaper AI Make Nvidia More Money?
Table Of Contents
  • Nvidia Vera Rubin’s Key Numbers At a Glance
  • The Short Answer: Cheaper AI Can Make Nvidia More Money, But 10x is Not the Reason
  • What is Nvidia Vera Rubin?
  • Vera Rubin vs Blackwell: What Actually Changes?
  • What Nvidia’s “10x Cheaper Tokens” Claim Really Means
  • The Token-Rack-Wallet Equation
  • Three Ways the Rubin Story Can Play Out For Nvidia
  • The Efficiency Dividend: Who Keeps the Savings?
  • Why Power May Matter More Than GPU Count
  • Can Rubin Cannibalise Blackwell?
  • Nvidia’s Financial Starting Point is Unusually Strong
  • What Does Rubin Mean for NVDA Stock Valuation?
  • Our View: Rubin is Bullish for the Business, But the Stock Still Has to Earn It

Nvidia has built a machine that, by its own estimate, can make some AI tokens about 10 times cheaper. That sounds wonderful for customers and slightly alarming for shareholders. If every data centre can do the same job with fewer GPUs, is Vera Rubin a growth engine or the most efficient way Nvidia has found to shrink its own market?

Let’s break down what Vera Rubin actually changes, how it compares with Blackwell, where Nvidia’s 10x claim applies, and the usage growth required for cheaper AI to create more revenue for Nvidia rather than less.

Nvidia Vera Rubin’s Key Numbers At a Glance

Key metricNvidia’s claim or timeline
10xHigher inference throughput per watt
0.1xToken cost versus the stated baseline
4xFewer GPUs for a large MoE training run
2H 2026Partner availability begins

Source: Nvidia’s Vera Rubin platform announcement and Vera Rubin NVL72 product page. These are Nvidia’s projected performance claims and are workload-specific.

The Short Answer: Cheaper AI Can Make Nvidia More Money, But 10x is Not the Reason

Our view is that Rubin is strategically bullish for Nvidia. It improves the economic return from a fixed power budget, expands the size of workloads customers can afford, and lets Nvidia sell a larger part of the data-centre system. Yet the stock thesis does not work simply because Nvidia wrote “10x” on a slide.

The critical question is demand elasticity. This means how much usage rises when the price falls. If the cost of useful AI work drops 90% and usage rises only 2x, customers need materially fewer systems for that work. If usage rises 10x or 15x, total compute demand can still expand.

There is another twist. Nvidia no longer sells only the engine. With Rubin, it is trying to sell more of the car: the Vera CPU, Rubin GPUs, NVLink fabric, Spectrum-X networking, BlueField data-processing units, storage components and software. A customer may need fewer racks for a fixed job, but Nvidia could collect more dollars from each rack.

That leads to our central framework:

Nvidia’s Rubin revenue outcome = usage growth divided by efficiency growth, multiplied by Nvidia’s revenue captured per rack.

Rubin is attractive if the first and third terms grow faster than the efficiency term.

What is Nvidia Vera Rubin?

Vera Rubin is Nvidia’s next full AI computing platform after Blackwell. “Vera” is the CPU, “Rubin” is the GPU, and NVL72 is the rack-scale system that links 72 GPUs with 36 CPUs so they can behave like one very large computer.

In March 2026, Nvidia said all seven major Rubin chips were in full production. Partner availability is expected to begin in the second half of 2026 through cloud providers and AI infrastructure companies including Amazon Web Services, Microsoft Azure, Google Cloud, Oracle Cloud Infrastructure, CoreWeave and others.

On Nvidia’s May earnings call, management said production shipments would start in the third quarter, with a larger ramp in the fourth quarter and the first quarter of fiscal 2028. That is a launch schedule, not a revenue guarantee.

Blackwell, meanwhile, is not one product. The original GB200 NVL72 was followed by GB300 NVL72, usually called Blackwell Ultra. GB300 is available now and improves dense FP4 computing and attention performance over the original Blackwell GPU.

This distinction matters because Nvidia’s headline 10x Rubin comparison uses the original GB200 NVL72 as the baseline, not GB300 Blackwell Ultra.

Vera Rubin vs Blackwell: What Actually Changes?

FeatureGB200 NVL72GB300 NVL72Vera Rubin NVL72Why investors should care
GPU generationBlackwellBlackwell UltraRubinRubin starts the next annual product cycle
GPUs per rack727272Nvidia keeps the rack-scale design
CPUs per rack36 Grace36 Grace36 VeraMore CPU value can sit inside each Rubin rack
GPU memory per rack13.4 TB HBM3EAbout 20 TB HBM3E20.7 TB HBM4Larger models and context can remain close to compute
GPU memory bandwidth per rack576 TB/sUp to 576 TB/s1,580 TB/sFaster movement of model weights and data
NVLink bandwidth per rack130 TB/s130 TB/s260 TB/sGPUs can exchange data twice as quickly
Product status on Aug. 25, 2026In productionAvailablePartner launch in 2H 2026Rubin still carries ramp and execution risk

Sources: GB200 NVL72, GB300 NVL72 and Vera Rubin NVL72. Nvidia labels Rubin specifications as preliminary and subject to change.

The largest practical leap is data movement. A Rubin GPU has 288 GB of HBM4 memory and 22 TB per second of memory bandwidth. Nvidia says the bandwidth is 2.8 times Blackwell or Blackwell Ultra. NVLink 6 provides 3.6 TB per second of GPU-to-GPU bandwidth per GPU.

Why does that matter?

Modern AI is often less like a student solving one equation and more like a restaurant kitchen serving thousands of customised orders. The chefs are the computing cores. The ingredients are the model weights and the user’s context. Adding chefs does little if ingredients arrive slowly. Rubin widens the pantry doors, doubles the rack’s interconnect bandwidth and gives the kitchen more memory.

Rubin is also designed for mixture-of-experts, or MoE, models. An MoE model contains many specialist subnetworks but activates only a few for each token.

Think of a hospital that routes a patient to the relevant specialists instead of asking every doctor to attend every appointment. This can reduce computation, but it creates heavy communication between GPUs. Rubin’s faster memory and interconnect are aimed directly at that bottleneck.

What Nvidia’s “10x Cheaper Tokens” Claim Really Means

The clean headline needs a messy footnote.

Nvidia’s product page compares one Vera Rubin NVL72 rack with one original GB200 NVL72 rack on Kimi-K2-Thinking, using a 32,000-token input and an 8,000-token output.

Under that specific long-context workload, Nvidia projects:

  • Up to 10x more tokens per megawatt.
  • Token cost at one-tenth of the previous level.
  • Training a 10-trillion-parameter MoE model in one month with one-fourth as many GPUs.

These are not broad guarantees for every model, prompt length, data centre or software stack. The results are projected, subject to change, and do not compare Rubin with GB300 in the headline test.

The phrase “up to” also describes a ceiling, not an average.

How investors should read the benchmark

That does not make the result unimportant. Kimi-K2-Thinking is a demanding reasoning workload, exactly the kind of application Nvidia expects agentic AI to create.

But investors should not plug a universal 90% price cut into a spreadsheet.

A fair interpretation is: Rubin may improve the economics of some advanced inference workloads by an order of magnitude, while the realised fleet-wide gain will depend on the workload mix and utilisation.

Training deserves a separate note. Nvidia’s one-fourth-GPU claim assumes a 10-trillion-parameter MoE model, 100 trillion training tokens and a fixed one-month training time. It means Rubin could make the frontier larger or faster. It does not mean all training clusters suddenly need 75% fewer GPUs.

The Token-Rack-Wallet Equation

Most chip analysis stops at speed. Investors need to go one step further and connect performance to Nvidia’s wallet.

Let:

  • U = growth in AI usage or workload.
  • E = improvement in useful work per rack.
  • P = change in Nvidia revenue captured per rack.

Then, as a simple model:

Relative Nvidia revenue = U ÷ E × P

Suppose a Rubin system is 10 times more efficient for a given task. If usage stays flat and Nvidia earns the same revenue per rack, the customer theoretically needs one-tenth as many racks.

If Nvidia captures 1.5 times more revenue per rack by selling more CPU, networking, storage and software, workload must grow roughly 6.7 times to keep Nvidia’s revenue unchanged.

How much usage growth offsets better efficiency?

Useful work per rackIf revenue/rack is 1.0x1.25x1.5x2.0x
4x4.0x usage3.2x2.7x2.0x
6x6.0x usage4.8x4.0x3.0x
8x8.0x usage6.4x5.3x4.0x
10x10.0x usage8.0x6.7x5.0x

Source: INDmoney analysis. This is a sensitivity model, not a forecast. Nvidia does not publish Rubin system pricing or revenue per rack.

The table gives investors a better question than “Is Rubin 10x faster?”

Ask: Will usage grow by five to eight times over the relevant years, and can Nvidia capture materially more value per installation?

A simple 10x-efficiency scenario

Assume Rubin delivers 10x useful work per rack and Nvidia captures 1.5x as much revenue per rack. Here is what different levels of usage growth would imply for revenue from that workload:

Usage growthRelative Nvidia revenueChange from baseline
2x0.30x-70%
5x0.75x-25%
7x1.05x+5%
10x1.50x+50%
15x2.25x+125%

Source: INDmoney analysis using U ÷ E × P, with E = 10 and P = 1.5. The 1.5x revenue-per-rack assumption is illustrative.

This is the core of the Rubin debate. A large efficiency gain can hurt units in a low-growth market and accelerate revenue in a highly elastic one.

Three Ways the Rubin Story Can Play Out For Nvidia

Scenario 1: Efficiency wins and customers buy fewer GPUs

In the bear case, AI usage grows, but not fast enough. Enterprises discover that many agent projects are expensive demonstrations rather than reliable products. Cloud providers use Rubin to serve existing demand with fewer racks. Custom chips take the most predictable inference workloads.

Nvidia may still sell premium systems, but unit growth, pricing or both weaken.

This risk is easy to dismiss during a capacity shortage, but it is real. Nvidia’s latest Form 10-Q explicitly warns that frequent architecture launches can lead customers to defer orders, reduce channel inventory and create revenue volatility.

Its inventory rose to $25.8 billion at the end of April 2026 from $21.4 billion in January, and the company recorded $1.1 billion of inventory and purchase-commitment provisions in the quarter.

Scenario 2: Usage expansion wins

In the base-to-bull case, lower costs unlock tasks that were uneconomic before: longer reasoning, always-on copilots, autonomous software agents, personalised video, robotics simulation and AI embedded in routine business workflows.

History supports the direction, though not a precise Rubin outcome.

Stanford’s 2025 AI Index found that the cost of achieving GPT-3.5-level benchmark performance fell from about $20 per million tokens in November 2022 to $0.07 in October 2024, a decline of more than 280 times. AI adoption rose rather than disappeared.

In the 2026 Index, 88% of surveyed organisations reported using AI in at least one business function and 70% reported generative AI use. Yet agent deployment remained in the single digits across almost every function.

That gap matters. Generative AI has reached broad experimentation, while autonomous agents are still early.

If an agent reads 200 pages, calls five tools, checks its work and repeats the task every day, it can consume far more tokens than a chatbot answering one question. Nvidia says agentic systems may use up to 15 times more tokens.

Treat that as a company claim, but the underlying logic is sound: more steps per task can absorb a large efficiency gain.

Scenario 3: Full-stack value wins

Rubin is not only a faster GPU. The standard NVL72 combines Rubin GPUs, Vera CPUs and NVLink. Nvidia’s broader rack platform adds Spectrum-X networking, BlueField-4 data-processing units and storage systems. Its software layer includes CUDA and Nvidia AI Enterprise.

The early financial evidence is visible in networking.

In Q1 FY2027, Nvidia’s data-centre compute revenue grew 77% year over year to $60.4 billion, while networking revenue grew 199% to $14.8 billion. Networking represented almost one-fifth of data-centre revenue in the quarter.

Software creates another layer of potential value.

Nvidia AI Enterprise’s self-managed one-year subscription has a list price of $4,500 per GPU. At list price, 72 fully licensed GPUs would equal $324,000 a year.

That is only an illustration. Actual discounts, cloud arrangements and product bundles vary, and it should not be treated as Rubin revenue guidance. Still, recurring software can make the lifetime value of an installation larger than its initial hardware sale.

The Efficiency Dividend: Who Keeps the Savings?

When token cost falls from 100 units to 10, a 90-unit “efficiency dividend” appears. It does not automatically belong to Nvidia.

Possible recipientHow it captures valueWhat investors should watch
NvidiaHigher system price, more networking, CPUs and softwareRevenue per deployed rack and networking mix
Cloud or AI providerKeeps token prices steady and expands gross marginCloud AI pricing versus infrastructure cost
End customerReceives lower API or cloud pricesFaster adoption and new workloads
The application itselfSpends savings on longer context and more reasoning stepsTokens per task and agent activity
CompetitorsWin standardised inference with cheaper custom siliconNvidia’s accelerator and networking share

This framework reveals the most important risk: not simply fewer GPUs, but value migration.

If cheaper inference becomes a standardised utility, hyperscalers may steer stable, high-volume tasks to their own chips. Nvidia does best when models change quickly, workloads are diverse and developers value the flexibility of CUDA and a common system. Custom silicon does best when a large workload is predictable enough to optimise tightly.

Competition is already moving.

Google introduced TPU 8t for training and TPU 8i for inference in April 2026, claiming up to 2.7 times better training performance per dollar and up to 80% better inference performance per dollar than its prior systems. Google also added native PyTorch support in preview, directly attacking the software-friction advantage that has helped Nvidia.

AMD is scaling too. Meta and AMD announced a multi-year agreement covering up to 6 gigawatts of AMD Instinct GPUs, with the first 1-gigawatt deployment expected to begin shipping in the second half of 2026.

Oracle separately plans a 50,000-GPU AMD MI450 deployment starting in the third quarter.

Rubin therefore has two jobs: make Nvidia infrastructure cheaper to operate and keep that infrastructure more useful than a collection of narrower alternatives.

Why Power May Matter More Than GPU Count

The “fewer GPUs” concern assumes customers are trying to complete a fixed amount of work. Many leading data centres face a different problem: a fixed amount of electricity.

Imagine a cinema with every show sold out. A projector that can serve 10 times as many screens per unit of power does not make the owner close nine screens. It lets the owner sell more tickets within the same electricity limit.

Cloud providers care about tokens per megawatt because access to power, grid connections and cooling can limit capacity. If demand exceeds supply, Rubin’s efficiency can raise the number of billable tokens produced by an already constrained site.

Nvidia’s claim of 10x tokens per megawatt is therefore economically more important than “10x faster” on its own.

The catch is that the cloud provider, not Nvidia, initially owns those extra billable tokens. Nvidia benefits only if better returns encourage more capital spending, support premium system prices or expand Nvidia’s share of the system.

Can Rubin Cannibalise Blackwell?

Yes, and that is partly the point.

Nvidia has moved to a roughly annual architecture cadence. Customers that know a better system is close can delay some purchases. Blackwell Ultra, Rubin and future platforms can also complicate supply planning. That raises the risk of a temporary air pocket during a transition.

Nvidia’s defence is software compatibility and a broad installed base.

On the May earnings call, management said rental prices for H100 had risen 20% year to date and A100 prices had risen 15%, arguing that older GPUs remained useful beyond their accounting lives.

This is management’s market observation, not independently audited pricing, but it suggests new chips have expanded the market without immediately making old chips worthless.

The more important test comes after Rubin ships at scale. If cloud rental prices and utilisation for H100, H200 and Blackwell remain healthy while Rubin ramps, the market is absorbing more total compute.

If older-platform pricing falls sharply before Rubin volumes grow, efficiency and substitution may be winning.

Nvidia’s Financial Starting Point is Unusually Strong

Rubin is arriving while Nvidia is still growing at a scale that few companies have ever reached.

Q1 FY2027 metricResultYear-over-year change
Total revenue$81.6 billion+85%
Data-centre revenue$75.2 billion+92%
Data-centre compute$60.4 billion+77%
Data-centre networking$14.8 billion+199%
GAAP gross margin74.9%+14.2 percentage points
Q2 FY2027 revenue outlook$91.0 billionCompany guidance

Source: Nvidia Q1 FY2027 earnings release.

Data centre produced 92% of quarterly revenue. That concentration makes Rubin more important, not less. A strong launch can sustain the main engine. A weak launch, supply problem or customer shift can hit almost the whole company.

Customer concentration adds another layer.

Three direct customers accounted for 21%, 17% and 16% of first-quarter revenue, or 54% combined. Nvidia notes that direct customers can be contract manufacturers, original equipment makers and cloud providers, so these percentages do not map neatly to three end users.

Even so, a small number of purchasing channels have enormous bargaining power.

What Does Rubin Mean for NVDA Stock Valuation?

Technology can be excellent while a stock is expensive. Investors have to price both. Nvidia closed at $208.48 on August 24, 2026, with a market value of roughly $5.05 trillion. At that size, even a respectable shareholder return requires a vast increase in earnings.

Here is a simple reverse valuation. For a 10% annual return over three years, Nvidia’s market value would need to reach about $6.72 trillion, ignoring dividends, buybacks and changes in net cash.

What earnings and revenue might support that value?

Assumed P/E in three yearsNet income requiredRevenue at 45% net marginIncrease vs $364B annualised Q2 guide
25x$269B$598B+64%
30x$224B$498B+37%
35x$192B$427B+17%

Source: INDmoney analysis. Starting market value is approximately $5.05 trillion; future value is $5.05T × 1.10³. The $364 billion baseline annualises Nvidia’s $91 billion Q2 guidance. The 45% net margin and future P/E ratios are illustrative, not forecasts.

Why use a 45% margin? It is below Nvidia’s recent non-GAAP level and leaves room for product transitions, competition and a less exceptional pricing environment.

Why avoid a simple trailing GAAP P/E? First-quarter GAAP net income included $15.9 billion of gains on equity securities, which makes headline earnings a poor measure of the operating run rate.

The table is not a price target. It shows the burden of proof.

At a 30x future multiple, the company would need roughly $224 billion in annual net income and about $498 billion in revenue at our assumed margin for shareholders to earn 10% a year from the August 24 starting point.

That outcome is possible, but it needs more than a successful chip launch. It needs sustained AI infrastructure spending, high margins, durable market share and a larger Nvidia wallet inside each data centre.

Our View: Rubin is Bullish for the Business, But the Stock Still Has to Earn It

We think usage expansion and full-stack value are more likely to outweigh unit efficiency over the next few years. Three facts support that view.

  • First, agent adoption is still in the single digits across most business functions. The addressable workload can grow much faster than the number of companies saying they “use AI.”
  • Second, Nvidia’s networking revenue is already growing faster than its compute revenue, evidence that its wallet is expanding beyond GPUs.
  • Third, scarce power makes customers value output per megawatt rather than a raw GPU count.

But there is no free lunch.

In our 10x-efficiency and 1.5x-revenue-per-rack illustration, usage must rise about 6.7x merely to preserve Nvidia revenue from the same category of work. That is plausible in agentic AI, not proven.

A large efficiency number should make investors demand stronger evidence of token growth, not relax their standards.

Our conclusion is deliberately split:

  • For Nvidia’s competitive position: Rubin looks positive. It attacks the biggest customer constraints: power, memory movement and system complexity.
  • For Nvidia’s revenue: Positive if usage grows at least five to eight times for the workloads receiving the largest efficiency gains, or if Nvidia captures much more of each rack.
  • For NVDA stock: The technology strengthens the long-term story, but a roughly $5 trillion valuation already assumes extraordinary execution. Better products do not remove valuation risk.

The most underappreciated point is this: Nvidia does not need customers to buy more GPUs for every fixed task. It needs the world’s total demand for useful AI work, multiplied by Nvidia’s share of the infrastructure wallet, to grow faster than efficiency.

Rubin is designed to make that possible. Investors still need to see it happen.

Share: