What if the thing throttling your AI budget next quarter isn’t GPU supply, model licensing, or token pricing, but a stack of memory chips glued to the side of a processor you’ll never touch?
Reuters reported on September 10 that Chinese AI chipmakers, including Huawei Technologies and Cambricon, have sharply raised prices for both current and next-generation AI processors. According to three people familiar with the matter, the cause is a high-bandwidth memory shortage. Huawei has lifted the indicated price of its Ascend 950DT accelerator card, which packages an AI processor together with memory and other components, to above 250,000 yuan, roughly $37,255, per two of those sources. Separate coverage put the increases at around 20 to 50 percent above quotes given a couple of months earlier.
I review tools for a living. I spend most of my time asking whether a product actually does what its landing page claims. So my instinct with a story like this is to ask a narrower question than the macro analysts are asking: does this change what I recommend to people building things?
Why a memory shortage shows up in your invoice
High-bandwidth memory is the part of an accelerator that feeds the compute units. Modern AI chips are rarely starved for arithmetic. They’re starved for data arriving fast enough. That’s why HBM sits physically stacked next to the processor, and why an accelerator card like the Ascend 950DT is sold as an integrated unit rather than a bare chip.
The practical consequence: when memory gets scarce, you can’t swap in a cheaper substitute and keep the same performance profile. The memory is the performance profile. That’s what makes this different from a generic component squeeze. There’s no downgrade path that leaves your throughput intact.
Reuters’ reporting notes the increase is affecting the cost of AI model development and inference. That’s the sentence worth sitting with, because it covers both sides of the lifecycle. Training gets more expensive, which is expected. Inference gets more expensive too, and inference is the recurring cost — the one that compounds every day your product is live.
What I’d actually change in how I evaluate tools
I’ve been guilty of reviewing AI tooling as though compute were a flat utility. You plug in, you pay per token or per hour, the price only ever goes down. That assumption has held up well enough for a few years that it stopped feeling like an assumption. This story is a reminder that it is one.
A few things I’m adjusting:
- Treat per-token pricing as a variable, not a constant. If you built a business model on current inference rates staying flat or falling, write down what happens if they move the other way. Hardware cost pressure eventually reaches the API layer.
- Ask vendors about hardware exposure. Not in a gotcha way. Just: what silicon are you running on, and do you own it or rent it? A tool running on hardware its vendor already bought has a different cost trajectory than one buying capacity monthly.
- Value efficiency work more than I used to. Quantization, caching, smaller task-specific models, batching. These used to read as nice-to-have optimizations. Under memory cost pressure they read as insurance.
- Re-read your exit paths. If a provider raises prices mid-contract, what’s the actual cost of moving? For most teams the honest answer is higher than they think, because the prompts, evals, and glue code are all shaped around one model’s behavior.
The part I can’t tell you
I’m not going to pretend I know how long this lasts or where prices settle. The reporting covers Chinese chipmakers specifically, and Huawei has previously said its Ascend 950 series would use two proprietary HBM technologies, HiBL 1.0 for the 950PR and HiZQ 2.0 for the 950DT, though the company hasn’t disclosed details beyond that. What that means for supply, I genuinely don’t know, and anyone who tells you confidently probably doesn’t either.
What I can say is that the shape of the constraint is clear. Memory bandwidth is the scarce resource in AI hardware right now, and scarce resources get priced accordingly. That pressure moves through the stack in a predictable order: chip vendors, then cloud providers, then API pricing, then your monthly bill.
A more boring kind of advice
The useful response here isn’t dramatic. It’s the unglamorous engineering work that good teams do anyway: know your token volume, know which calls actually need the biggest model, cache what you can, and keep a rough idea of what your stack would cost if inference pricing rose meaningfully.
That’s less exciting than a supply chain story, but it’s the part you control. The memory market isn’t waiting on your opinion. Your architecture is.
🕒 Published: