AMD Says Its Chips Beat Nvidia by 30% on Cost. The Real Number Might Be Higher.
AMD has spent the last two weeks telling anyone who'll listen that its Instinct MI350 chips undercut Nvidia's B200 on price by roughly 30%. That's the headline number from its Advancing AI 2026 event. But go look at the actual independent benchmarks, and the marketing pitch turns out to be conservative — not inflated.

Photo by panumas nikhomkhai on Pexels
SemiAnalysis runs a live benchmark platform called InferenceX that tests real inference workloads across GPU architectures, not vendor slide decks. On the GLM-5.1 model at higher throughput settings, their numbers show MI355X running 40% cheaper per million tokens than B200 on SGLang. On other configurations tracked by the same platform, comparing MI355X against Nvidia's H200, the gap widened even further — in some throughput regimes the cost-per-token advantage exceeded 100%. The story isn't "AMD is a bit cheaper." It's "AMD is dramatically cheaper in specific, real workloads, and the size of that gap depends heavily on which model and which throughput target you're running."
Why the Cost Story Actually Matters More Than the Price Story
Here's the thing that gets lost when people fixate on sticker price: hyperscalers don't buy GPUs, they buy tokens per dollar per watt. A chip that's 30% cheaper to purchase but burns more power per token isn't actually cheaper to operate at scale. That's why the SemiAnalysis inference benchmarks matter more than AMD's own MSRP comparisons — they measure the thing that actually shows up on a cloud provider's P&L.
This is also why the commercial deals AMD signed this year are the more important story than any single cost claim. As I covered in AMD Says AI Chips Are a $1.4 Trillion Market. Who's Actually Paying?, the answer to "who's actually paying" now has real names attached:
- OpenAI committed to 6 gigawatts of AMD Instinct GPUs across multiple chip generations, with the first 1GW of MI450 deployment starting in the second half of 2026 — and AMD handed OpenAI a warrant for up to 160 million AMD shares, vesting in tranches as deployment milestones hit.
- Oracle will be the first hyperscaler to offer a public AI supercluster built on 50,000 MI450 GPUs, going live in Q3 2026.
- Meta signed a five-year, $60 billion supply agreement anchored on a 1GW opening MI450 deployment, with a custom variant co-engineered for Meta's own workloads.
Those aren't hypothetical TAM numbers. They're contracted revenue with real capacity commitments behind them, which is a very different thing than "AMD says its chips are cheap."

Photo by analogicus on Pixabay
The Stock Has Already Priced In a Lot of This
AMD closed near $553 during the Advancing AI event, more than double where it started 2026. Analysts didn't wait for confirmation — KeyBanc has a $725 target, UBS is at $700, and Bank of America bumped its target to $560 from $500. When a stock is up over 140% in seven months and sell-side targets are already clustering in the $600s and $700s, the "AMD is undervalued" trade is mostly over. What's left is a bet that execution continues to match or beat what's already baked into the price.
That's a meaningfully different bet than the one investors were making back when AMD was a cheap alternative nobody believed in. Now the stock needs the OpenAI, Oracle, and Meta ramps to actually hit their gigawatt targets on schedule, with margins that hold up once AMD is manufacturing at real volume instead of showcase volume.
The Part of the Bull Case That Still Has a Hole In It
Cost-per-token benchmarks are compelling, but they don't erase Nvidia's real moat, which was never really about silicon. CUDA has roughly 4 million trained developers and deep integration across PyTorch, TensorFlow, and nearly every production inference stack, while ROCm — AMD's answer to CUDA — is still catching up on tooling maturity, kernel optimization, and the long tail of framework compatibility that took Nvidia a decade to build. Nvidia's AI GPU market share has slipped from about 92% to roughly 85% over the past two years, which is real movement, but it's still a dominant position, and switching costs for teams with production code built on CUDA are not trivial.
That gap is narrowing — AMD has been pouring R&D into ROCm and landed official PyTorch packaging and a Hugging Face partnership — but "narrowing" and "closed" are different words. If a hyperscaler's internal ML teams hit friction migrating workloads to ROCm at scale, the token-cost advantage on paper doesn't matter, because engineering time is the real bottleneck, not GPU rental cost.
What Would Actually Change My Mind Here
The bull case holds if the MI450 ramp with Oracle and Meta hits its 2026 capacity targets without slipping, and if AMD keeps showing up in independent, model-by-model benchmarks like InferenceX rather than just its own comparisons. The bear case shows up if OpenAI's deployment milestones slip (which would delay the warrant vesting too — a signal worth watching), or if ROCm friction shows up in customer commentary during Q3/Q4 earnings calls. Either of those would matter a lot more than another slide about price-per-GPU.
This is not financial advice — always do your own research before making investment decisions.
For now, the more interesting number isn't AMD's 30% claim — it's that independent testing suggests the real advantage, in the workloads that matter, is often bigger than AMD itself is advertising. That's unusual for a company whose stock has already run this hard, and it's worth watching whether that gap holds up as MI450 volume ramps through the back half of 2026.
댓글
댓글 쓰기