
Trainium3 is Amazon's third home-made AI chip, built by its Annapurna Labs unit on TSMC's 3-nanometre process, with 144 GB of high-bandwidth memory and roughly double the low-precision maths of its predecessor. SemiAnalysis's deep dive, the most detailed public account of the chip, says Amazon's guiding rule is "deliver the fastest time to market at the lowest TCO", total cost of ownership, rather than the highest performance. The evidence is in the compromises: most 2026 racks are air-cooled with 64 chips, because liquid-cooled buildings are the bottleneck; the first network switches force data to take up to three hops between chips in the same rack, with faster one-hop switches to be swapped in later; every chip carries 16 spare wires so a switch tray can be replaced without stopping the rack, something Nvidia's GB200 cannot do; and Amazon earns share warrants in its switch supplier Astera Labs worth an effective discount of roughly 23 percent. On 20 April Anthropic committed more than $100 billion over ten years to AWS, for up to 5 gigawatts of capacity across Trainium2, 3 and 4, with nearly 1 GW due online by the end of 2026. Amazon's Andy Jassy calls the silicon "high performance at significantly lower cost". SemiAnalysis still expects Nvidia to stay "King of the Jungle" as long as it keeps accelerating, and notes Trainium3's next generation may use Nvidia's own NVLink connector.
The claim
Every article about an Amazon AI chip asks the same question: is it an Nvidia killer? SemiAnalysis, the research firm whose reports we lean on for the hardware side of this brief, answered it in an unusual way in its Trainium3 teardown. It did not lead with speed. It led with a sentence about accounting: "With Trainium3, AWS remains laser-focused on optimizing performance per total cost of ownership (perf per TCO). Their hardware North Star is simple: deliver the fastest time to market at the lowest TCO."
Some translation. AWS is Amazon Web Services, the cloud business that rents computing by the hour and is Amazon's main profit engine. Trainium is its own line of AI chips, designed by Annapurna Labs, an Israeli chip start-up Amazon bought in 2015, and made by TSMC in Taiwan. TCO, total cost of ownership, is what a rack of chips costs over its whole life: the chips, the network, the cooling, the building, the electricity and the people who keep it running. Performance per TCO is how much useful work you get for all of that money. It is a different scoreboard from the one Nvidia tops, which is how much work one chip can do in one second.
On that second scoreboard Trainium3 does not win, and SemiAnalysis does not pretend it does. The chip doubles its predecessor's low-precision maths (the rough-and-fast arithmetic that modern AI mostly runs on), moves to TSMC's 3-nanometre process and carries 144 GB of high-bandwidth memory with 70 percent more bandwidth than before, partly because Amazon switched memory suppliers from Samsung to SK Hynix and Micron. Higher-precision maths is unchanged, and the newest four-bit number format runs no faster than eight-bit, where Nvidia's chips double up. SemiAnalysis's estimate, in the paid part of the report, is that Trainium3 is about 30 percent better than Nvidia's GB300 rack on cost per unit of eight-bit performance and much worse on four-bit. That is a chip built to be cheap on the work its customers actually run, not to top a benchmark.
"Their hardware North Star is simple: deliver the fastest time to market at the lowest TCO."
Four choices that give the game away
First, air over water. Nvidia's flagship rack, the GB200 NVL72, packs 72 chips into a liquid-cooled cabinet, and most of the world's data centres cannot cool one. Amazon's main 2026 product is a two-rack, 64-chip, air-cooled system. SemiAnalysis calls it "the time to market SKU" and says it is "the only SKU with a switched scale-up architecture that can be deployed in datacenters that are not liquid cooled ready", which is why "the majority of the Trainium3 chips deployed in 2026" will be that version. A slightly slower rack you can install this year beats a faster one waiting for a building.
Second, a worse network now, a better one later. Inside an AI rack the chips must talk to each other constantly, and the ideal is a network where every chip is one hop from every other. Amazon's first version is not that. Because the big PCIe switches it wanted were not yet available in volume, it launched with 160-lane switches that force some traffic to take "up to three hops" between two chips in the same rack, and four hops across the two-rack version. The plan is three switch generations over the chip's life: 320-lane switches that bring every chip to one hop, then a new industry standard called UALink. SemiAnalysis's phrase is that the first generation "will be rather quickly replaced". Amazon shipped a rack it already intends to upgrade, because shipping mattered more than perfection.
Third, spare wires. Each Trainium3 has 80 wiring lanes to the rack's backplane, of which 16 are deliberately unused spares. They exist so that a switch tray can be swapped while the rack keeps working, and so a rack with a few dead lanes does not stall the thousands of others in a training run. SemiAnalysis contrasts this with Nvidia's GB200, where "operators must first drain all the workloads from the rack before swapping switch trays", and notes Nvidia's backplane "has had poor reliability". Amazon's servers are also "cableless": every signal runs through the circuit board rather than through hand-plugged cables, which costs some signal quality and needs extra signal-boosting chips, but makes the racks faster to build and less likely to be assembled wrong. Nvidia is now copying the idea for its next generation.
Fourth, the parts pay for themselves. Those signal boosters and switches come from Astera Labs, and Amazon's purchase agreement gives it warrants on Astera stock that vest as it hits volume targets, at a strike price of $20.34. SemiAnalysis worked out that the warrants vested by late September 2025 amounted to "an effective discount of roughly 23 percent" on the parts. We wrote last week about Amazon making Qualcomm pay it in shares to become a customer; here it is the same trick pointed at a supplier. The more Amazon buys, the more it is paid.
Who is buying
The customer that makes this strategy work is Anthropic. On 20 April it announced a new agreement with Amazon: "We are committing more than $100 billion over the next ten years to AWS technologies, securing up to 5GW of new capacity to train and run Claude." The deal spans Trainium2 through Trainium4, with the option to buy future generations. Anthropic said it already runs "over one million Trainium2 chips" and expects "nearly 1GW total of Trainium2 and Trainium3 capacity coming online by the end of 2026", with scaled Trainium3 "later this year". For scale, one gigawatt is roughly the electricity draw of a city of several hundred thousand homes.
Andy Jassy, Amazon's chief executive, gave the sales pitch in the same announcement: "Our custom AI silicon offers high performance at significantly lower cost for customers, which is why it's in such hot demand." Note the word order. Not the fastest; lower cost. Anthropic, which said in the same post that its revenue run-rate had passed $30 billion and that its own growth was straining its infrastructure, is the kind of buyer that cares about cost per unit of work far more than about a chip's peak speed.
The relationship is also circular, which readers of our Oracle and Qualcomm briefs will recognise. Amazon said it was investing a further $5 billion in Anthropic, with up to $20 billion more to come, on top of the $8 billion already invested. So the chip's biggest customer is partly owned by the chip's maker, and a good share of the money Anthropic raises flows back to AWS as compute bills. SemiAnalysis also reports that the next chip, Trainium4, will come in a version that connects using Nvidia's own NVLink technology, and it believes Amazon "is unlikely to be paying Nvidia's typical ~75% gross margins" for the privilege. The supposed Nvidia killer may end up carrying an Nvidia part.
Is this actually new?
Amazon has run this play before. Its Graviton processors, launched in 2018, were slower than Intel's best on many single tasks and were sold on a simple promise: the same job for less money. Graviton now runs a large share of AWS. Before that, Google built its early search infrastructure out of cheap commodity computers that failed often, and wrote software to route around the failures rather than buy the expensive, reliable machines the industry sold. Amazon's spare lanes and hot-swappable switch trays are that philosophy in hardware.
SemiAnalysis itself reaches for an older precedent, and it is a warning to Nvidia rather than praise for Amazon: "In the same way that Intel stayed complacent in the CPU while others like AMD and ARM raced ahead, if Nvidia stays complacent they will lose their pole position even more rapidly." The report's conclusion is that Nvidia stays "King of the Jungle" only "as long as they continue to keep accelerating their pace of development". What is new is not the strategy but the company running it. Amazon is the largest cloud in the world, its chip now has a ten-year anchor customer, and it is the first company outside Nvidia to ship a full switched rack, ahead of AMD by about a year.
The everyday version
Think of a budget airline. It does not fly faster than the flag carrier. It flies one type of aircraft so any crew can fly any plane, it turns them around in 25 minutes, it uses the secondary airport that has a free gate today rather than waiting for a slot at the main one, and it negotiates so hard that some airports pay it to land there. The seat is a little less comfortable and the network has a few more connections. If you only care about the cost of getting from A to B, you book it anyway.
Trainium3 is that airline. The chip is not the fastest. The rack is air-cooled because that building exists now. The network takes an extra hop until better switches arrive. The parts come with a rebate paid in shares. And its biggest customer has just booked ten years of seats, because it is buying tokens per dollar, not speed.
What to take from it
If you want one sentence: Amazon has stopped trying to beat Nvidia's chip and started trying to beat Nvidia's bill, and its biggest customer just signed a $100 billion vote that the bill is what matters.
For anyone buying AI rather than building chips, this is the number that decides your costs. The price of a million tokens (the units of text a model reads and writes) is set less by whose chip is fastest than by whose rack does the most work per dollar over five years. Every cheaper useful chip that reaches volume pushes that price down.
Three things to watch. Whether Anthropic's nearly 1 GW of Trainium capacity actually lands by December, because the whole case rests on time to market. Whether the software catches up: SemiAnalysis said in December that the chip mode most outside researchers want was not due until mid-2026, and Amazon's open-sourcing of its compiler is its attempt to dig the kind of developer moat Nvidia has. And whether Trainium4 really ships with Nvidia's connector inside, which would tell you that the two companies have decided to share the market rather than fight over it.
Curious about AI? Come build with us.
Oslo Vibe Coding runs free, beginner-friendly drop-ins where we build real things with AI. No one codes alone.