AI hardware is on track to need something like 10 gigawatts of power, a number so large it gets its own conference talks. Meanwhile, the most promising fix on the table right now is not a new transistor or a magic cooling fluid. It is rearranging the pieces you already have into a better layout.
That tension is the whole story of chiplet co-design, and it is why I have been paying attention to it despite covering mostly software tooling here. When the hardware underneath your models gets cheaper to design and cheaper to run, that eventually shows up in your inference bill. So let’s look at what is actually being claimed, and what I would want to see before treating any of it as settled.
What the research actually says
Recent work out of the University of Michigan and other institutions points to chiplet co-design frameworks cutting both energy and design costs for AI accelerators. The core idea is not exotic. Instead of designing one enormous monolithic chip and hoping it fits every workload, you compose an accelerator out of smaller dies and optimize the composition against the workload you actually run.
A paper called Fengshui puts numbers on it. For datacenter mixture-of-experts and dense LLM serving, it reports reducing prefill energy by up to 16.8%, and energy-delay product by up to 28.7%. The mechanisms named are operator-level heterogeneity and expert parallelism, which in plain terms means different chiplets handle different kinds of math, and expert routing gets mapped onto that heterogeneity rather than fighting it.
Separately, Dr. Nasrullah’s talk at Chiplet Summit 2026 frames the 10GW problem around four techniques, with node mixing and low-power approaches among them. Node mixing is the pragmatic one: not every part of an accelerator needs the newest process node, so put the parts that benefit on the expensive node and the rest on cheaper silicon. That is where design cost savings come from, not just energy.
Why I find the cost angle more interesting than the energy angle
Energy numbers get the headlines because gigawatts sound dramatic. But a 16.8% prefill energy reduction, while real, is an efficiency gain, not a change in trajectory. If demand for inference keeps growing at anything like its current rate, a 17% improvement gets consumed quickly.
The design cost side is where the structural change lives. Monolithic accelerator design is brutally expensive, which is why so few organizations attempt it. If composing chiplets genuinely lowers that barrier, you get more players designing purpose-built accelerators instead of everyone renting time on the same handful of general-purpose parts. That is a supply-side shift, and supply-side shifts tend to matter more than efficiency percentages.
The institutional signals back this up. The U.S. CHIPS National Advanced Packaging Manufacturing Program explicitly includes chiplet ecosystems and co-design in its scope. Cadence has a chiplet platform aimed at exactly these engineering and business problems. Public money and EDA vendor roadmaps are pointing the same direction, which usually means the tooling gap is the real bottleneck, not the concept.
The parts nobody is selling hard
A few things I would want clarified before getting enthusiastic:
- “Up to” is doing work in those numbers. Up to 16.8% and up to 28.7% are ceilings measured on specific workloads. The typical case is almost certainly lower, and the paper’s framing does not tell us how much lower.
- Bespoke means bespoke. The Fengshui work is explicitly about bespoke neural network accelerator co-design. Co-optimizing hardware against a workload works beautifully until the workload changes. Model architectures have not exactly been stable.
- Thermal behavior is an open question. There is active work on open benchmarks for evaluating AI thermal models in 2.5D packaging, which tells you thermal modeling for these stacked designs is still being standardized. Unsolved modeling problems have a way of eating projected gains.
- Design cost reduction assumes an ecosystem exists. Cheap composition only works if there are chiplets to compose, with interfaces that actually interoperate. That is a coordination problem, and coordination problems are slower than engineering problems.
What this means if you are not designing silicon
For most people reading a toolkit review site, the practical takeaway is patience with a side of attention. You are not going to specify chiplet layouts. But the direction of this research tells you something about where inference economics are heading: toward more specialized hardware, chosen per workload, rather than one accelerator to rule them all.
If that plays out, the skill that gains value is not knowing hardware internals. It is knowing your own workload well enough to match it to the right silicon. Profiling your inference patterns is unglamorous work that nobody puts on a conference slide, and it is probably the most useful preparation available.
The research here is solid and the incentives are aligned. I would just hold off on treating single-digit-to-high-teens percentage gains as the answer to a 10-gigawatt problem. They are a real contribution to it, which is a different claim.
đź•’ Published: