\n\n\n\n Nobody Benchmarks the Substrate - AgntBox Nobody Benchmarks the Substrate - AgntBox \n

Nobody Benchmarks the Substrate

📖 5 min read•803 words•Updated Sep 17, 2026

Remember the GPU shortage of the crypto years, when everyone learned that a graphics card was not really one product but a supply chain wearing a trench coat? People who had never thought about component sourcing suddenly knew about GDDR6 lead times and shipping container rates. The lesson faded fast once cards were back on shelves.

I think about that a lot when I read AI accelerator spec sheets in 2026. We have gotten very good at arguing about the logic die. Transistor counts, process nodes, TFLOPS at whatever precision the vendor finds most flattering. Meta’s MTIA is on its third iteration now, expected on TSMC’s N3P with likely more than 100 billion transistors and HBM attached. That number is the part everyone will quote. It is also the part that tells you the least about whether the thing ships in volume.

The part that holds everything together

The unglamorous answer is packaging, and specifically the ABF substrate underneath. Modern accelerators are not a chip. They are several large logic dies sitting next to stacks of memory, all mounted on a substrate that has to carry signal and power across that whole assembly without warping, cracking, or losing yield. As designers pack more silicon per accelerator to hit the memory bandwidth frontier models want, that substrate gets bigger and more layered, and the tolerances get meaner.

This is the other chip in every AI accelerator, in the sense that it is a manufactured component with its own capacity constraints, its own small set of suppliers, and its own failure modes. It just does not have a marketing name, so nobody puts it in the keynote.

As someone who reviews tooling for a living, I have a bias here: I care less about peak numbers than about whether a thing is available, predictable, and boring to operate. Packaging is where those three properties are decided.

Reading the 2026 market share differently

The headline story is that NVIDIA’s data center AI accelerator revenue share is estimated around 80 to 85 percent this year, down from roughly 92 percent in 2024. AMD is up to something like 5 to 7 percent from about 2 percent. Google’s TPU line continues on its own track.

The usual reading is that AMD’s MI300X+ finally got competitive on software, which is partly true. The reading I find more useful is that a few points of share moving at this scale is mostly a story about who could get manufactured product into racks. When demand outstrips supply this badly, share is a supply metric dressed up as a competitiveness metric. Intel’s Gaudi 3 underperforming is a product problem. Broadcom’s custom XPUs gaining traction is partly a design-win story and partly a story about hyperscalers buying their way into the packaging and memory queue directly instead of standing in NVIDIA’s line.

Where the money is actually going

Two funding rounds from February tell you what investors think the constraint is. Cerebras raised $1 billion in a Series H on February 3, and Positron AI raised $230 million in a Series B on February 4.

Cerebras is the most literal possible bet against conventional packaging. The Wafer Scale Engine 3 sidesteps the multi-die-on-substrate problem by not cutting the wafer up in the first place. You can call that elegant or you can call that stubborn, but it is a direct architectural response to interconnect and assembly limits rather than a faster core.

Positron describes its work as a memory-centric inference accelerator, which is the other half of the same observation. If the hard part is feeding compute rather than having compute, design around the memory path and stop optimizing the number that goes in the press release.

Neither approach is guaranteed to work out. Both are honest about where the pain is.

What this means if you are choosing hardware

For most readers of this site, you are not buying accelerators by the tray. You are picking a cloud instance type or deciding which runtime to build against. The practical version of all this:

  • Availability beats peak specs. A part you can rent today at a predictable price is better than a faster part with a six-month queue.
  • Memory bandwidth and capacity per accelerator will tell you more about your real throughput than headline FLOPS, especially for inference.
  • Vendor share numbers are supply signals right now. Do not read a few points of movement as a verdict on software maturity.
  • Architectures that route around packaging limits, wafer-scale or memory-first, are worth testing on your own workload rather than trusting benchmarks tuned for someone else’s.

The transistor count is the easy number. The substrate under it, and the memory beside it, are what decide whether the accelerator is a product or a demo. That is not a new observation in hardware. It just keeps getting rediscovered every time compute demand outruns the factories that have to assemble it.

đź•’ Published:

đź§°
Written by Jake Chen

Software reviewer and AI tool expert. Independently tests and benchmarks AI products. No sponsored reviews — ever.

Learn more →
Browse Topics: AI & Automation | Comparisons | Dev Tools | Infrastructure | Security & Monitoring
Scroll to Top