RSS Amplifier

Patrick Kennedy's Axautik Group LLC and ServeTheHome Stack · Apr 17, 2026

The Low-Hanging AI Compute Fruit: A New PCIe GPU

0
Sign in to vote or save

Patrick Kennedy · Patrick Kennedy's Axautik Group LLC and ServeTheHome Stack

Something is fairly striking in the current AI accelerator market. There is no HBM-powered current-generation GPU available for PCIe CEM (Card Electromechanical) slots. With the surge of local AI occurring alongside massive AI cluster build-outs, there are several good reasons new products are not launching. At the same time, this is one of those areas where it makes a lot of sense to have a big HBM-powered PCIe card because we are seeing the need daily with our work with local AI models on STH. Let us get into the state of the industry, along with some opportunities and challenges to making a big HBM PCIe card a reality.

The AI revolution has created an unprecedented demand for compute resources that can handle increasingly large models. While much attention has focused on data center accelerators connected via NVLink, InfiniBand, Ethernet, UltraEthernet, and so forth, there exists a significant opportunity in the traditional PCIe GPU market that remains largely unaddressed by today’s accelerators. The industry needs an HBM3E-based PCIe graphics card that delivers both high bandwidth and high memory capacity in a form factor compatible with existing server infrastructure.

Current AI workloads face a fundamental bottleneck: the trade-off between model size and available memory. The answer, if you want to run big models fast, is to buy a rack-scale GPU cluster, get a lot of power to run it, and have at it. While that is great, not everyone has a multi-million dollar budget to implement that plan, and just finding power is not easy anymore. So let us start small and work up.

At STH, we have ten NVIDIA GB10 systems and have an 8-node GB10 cluster with 200GbE RDMA networking running. It runs large models like Kimi-K2.5 but slowly because each node has limited memory bandwidth and only 128GB of capacity. Likewise, the AMD Ryzen AI Max+ 395 systems at 128GB run into memory bandwidth and capacity challenges. Our Apple Mac Studio M3 Ultra 512GB has better memory capacity and bandwidth, but we may need 1-3 more because we have seen it drop to single-digit tokens per second while processing 150K+ context tokens. It is super fast, until it goes so slowly that you can clock it with a sundial.

All of these are instructive for the need for a data center PCIe GPU because they address memory capacity, enabling large models at the expense of speed. The main alternatives to these would be getting an NVIDIA RTX Pro 6000 Blackwell with 96GB of GDDR7. That has more FLOPS and higher memory bandwidth, but a GB10 can run larger models and retain more context due to its capacity.

The NVIDIA GB300 Station deserves special mention, as it features a de-rated 252GB of HBM3E memory from the higher-end Blackwell Ultra servers. In a 1.6kW power budget for the entire system, you only get one of these GPUs in the GB300 Station platforms. That means a cluster of four GB10’s will be vastly slower, but will run larger models. As a result, NVIDIA has a Grace CPU with 496GB of memory to help expand the total memory in a system. People want more, so now you can also add a 96GB RTX Pro 6000 Blackwell.

When we look at what you can put into servers, a standard PCIe GPU that fits into a standard CEM slot, it feels like the world has forgotten this segment. Four GPUs, eight GPUs, or even just one or two in a standard server. This is a form factor and segment that are somewhat compatible with existing infrastructure and with which people are familiar.

We did a guide on this in 2025, called The 2025 PCIe GPU in Server Guide, and it was extremely popular, garnering over 170K views each on both the written and video forms. This has been a popular topic folks have been researching for the last nine months. At the same time, there is no true PCIe GPU flagship for this generation that combines leading memory capacity, ample memory bandwidth, and a lot of FLOPS. That is a challenge.

Memory bandwidth and capacity directly determine what AI models can run efficiently on a given system. HBM3E technology offers substantial advantages over traditional GDDR memory architectures. The three-dimensional stacking approach enables significantly wider memory interfaces within the same physical footprint.

For AI inference workloads, memory capacity and memory bandwidth often matter more than raw FLOPS. A model must fit entirely within GPU memory to avoid costly swaps to system RAM or storage. We are already running models in the lab that require 800GB or more of memory between the model and its context storage, and we are on the verge of a new generation of larger models that could be quantized down to that range but will use a lot more memory. That 800GB figure is important because 96GB per GPU and 8 GPUs is 768GB of memory. Unfortunately, memory capacity is one where you either have it or you do not, but you are always trying to get right up to the limit.

High bandwidth of that memory ensures that once a model is loaded, data can flow to compute units without creating bottlenecks. Transformer architectures particularly benefit from sustained memory throughput during attention operations and token generation phases. The combination of high capacity and high bandwidth enables organizations to run larger models at higher batch sizes while maintaining acceptable latency profiles. Our NVIDIA GB10 cluster has the capacity, but with LPDDR5X and then having to sync over the network in a TP=8 configuration, it is like an extreme case of “it works” but it is slow. Just like when we ran a large BF16 Deepseek-R1 model on an AMD EPYC server.

When you have tested and used the extreme memory bandwidth starved configurations to get capacity, you crave speed. On the other hand, when you have traded capacity to get that speed and lose the accuracy of the models you are running, you quickly yearn for capacity.

If you are searching for a GPU server that uses “only” 5kW for four GPUs or 10kW for and eight GPU system, your chance with a more advanced solution like the HGX NVL8 platforms stopped somewhere around NVIDIA Hopper. PCIe GPUs at around 600W each (some can do more with liquid cooling) mean you can build servers to hit those rack power limits and retrofit them into existing data centers.

That begs the question, what is out there? Let us get to that next.

Read the original on axautikgroupllc.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.