AI capability and demand are growing exponentially, driving massive data center build-out and unprecedented energy demand. The scale is staggering: hyperscalers are projecting over $1T in capex and even then, timelines stretch 5-8+ years just to secure power and bring facilities online (as referenced by our colleagues Kevin and Nick in an earlier post).
Inside the data center, the challenges are just as severe: a single pod now approaches the power consumption of a small city, and large facilities contain enough optical fiber to wrap around the world five times. As a result, energy and operational complexity have become key design constraints, instead of afterthoughts. The explosive growth of LLMs and their infrastructure has made physics the bottleneck. At Congruent, we think the solutions will come from a stack that’s energy-aware from the ground up. Historically, power was treated as a utility - you’d design the compute, then provision enough electricity. The new stack inverts this structure. Compute systems are starting to understand power constraints, and power systems are starting to understand compute workloads. That bidirectional awareness is the thread tying together what might otherwise look like scattered opportunities. That’s where we’re focused: orchestration software that coordinates across power, cooling, and workloads; tools that compress the years-long path to grid interconnection; and step-function hardware bets in memory, photonics, and power electronics. If you’re building in any of these areas, we’d love to talk.
What makes this moment unique is that none of these problems are waiting their turn. Physical limits in power distribution, memory bandwidth, and grid stability are all hitting at once. Moore’s Law economics are, to some extent, breaking down precisely when AI demand is accelerating exponentially, which means there’s never been more urgency, or frankly more opportunity, for architectural innovation. Breakthrough innovations at any single layer - particularly in memory architecture - would have sweeping effects on their own. However, because these constraints are interdependent, solutions that work across layers can compound in ways that point solutions can’t. The entire stack is evolving together, and this interdependence creates opportunities for coordinated innovation beyond what point solutions can achieve.
In this post, we’ll cover the three primary data center constraints, what the energy-aware stack looks like, and where new venture opportunities may emerge.
Three fundamental constraints are colliding simultaneously, driving innovation across the entire compute stack.
1. Model size outpacing hardware (The Memory Wall)
Model parameters are growing exponentially while memory capacity and bandwidth improve linearly. The widening gap forces tight coupling of GPUs - models require low-latency memory access, which means compute must be densely packed. Training workloads are often memory-bound, leaving expensive GPUs underutilized. The result is monolithic racks that are a nightmare to cool and operate.
New memory and compute architectures represent the largest opportunity for breakthrough innovation, though the path ahead remains uncertain.
A range of approaches are being explored - vertical 3D integration, compute-in-memory, compute-in-network architectures - all of which place memory closer to compute. Memory specialization by workload type, analogous to today’s specialized accelerators, might dramatically improve efficiency for specific tasks, though it’s early. Groq’s LPU architecture, which uses on-chip SRAM to sidestep HBM bandwidth bottlenecks for inference, is one example of this approach attracting significant capital. If these approaches work, they wouldn’t just improve performance, they’d simplify compute architecture and substantially reduce energy consumption. We’re watching this space closely, as is everyone else in the industry, because step-function improvements here would change everything downstream.
Photonic interconnects offer a clearer path forward. As bandwidth scales, they deliver the highest performance while addressing several key challenges. On the packaging front, they move more bits through fewer, more manageable cables and connectors. They also promise lower energy per bit. And because optical signals can travel longer distances with minimal loss and latency, they enable more flexible rack layouts within data centers and more flexible data center architectures across buildings or regions.
Here’s where the interdependence becomes visible. Photonic disaggregation enables smaller distributed data centers mainly for inference workloads, which are individually less demanding on local power grids and can be deployed more quickly, with greater flexibility in site selection and power procurement.
2. Surging power needs at the rack level and unsustainable current densities
Memory bandwidth requirements force dense coupling of GPUs, which has power implications that are hard to overstate. GPU power draw has increased from around 300W to over 1000W in just a few years, and NVIDIA’s upcoming Rubin platform will push entire racks to 600kW - the power consumption of hundreds of homes, concentrated in a single cabinet. At these densities, the physics of low-voltage distribution starts to break down. Resistive losses scale with the square of current, copper busbars become prohibitively massive, and you quickly find that floor space is being consumed by power infrastructure rather than compute.
The industry is responding by moving to high-voltage DC distribution (±400V, 800V, with 1500V under consideration) which dramatically changes the local power distribution architecture. A 600 kW rack at 48V requires over 12,000 amps; at 800V, just 750 amps. That difference shows up in smaller copper busbars, less space consumed by power infrastructure, and more room for actual compute. Superconducting transmission, which companies like Congruent’s portfolio company VEIR are commercializing, could push efficiency even further. That said, HVDC creates its own demands. Efficient power conversion from medium voltage down to chip-level voltages may require semiconductor technologies and circuit topologies that are still being developed, and safety standards for HVDC protection haven’t fully caught up yet. Liquid cooling is a similar story - it enables higher densities, but demands tight integration with power delivery and rack design, which is why companies like portfolio company Alloy Enterprises are focused specifically on the thermal management piece. Beyond the rack, microgrids and solid-state transformers (the likes of HeronPower, Amperesand, Blixt, and unannounced incumbent vendor efforts, among others) provide smart grid connection and bidirectional power management.
What ties all this together is that memory growth drives the need for architectural density, and power delivery is what makes that density physically possible. Neither problem is being solved in isolation, which is another reason why we find this space so interesting.
3. Grid compliance and load volatility
The challenge here operates at two levels: securing interconnection, then operating as a responsible grid participant once connected.
As we’ve discussed in previous posts, securing grid interconnection takes 5-8+ years and has become a primary bottleneck for data center development. These timeframes can be dramatically shortened to [< 1 yr] by using flexible interconnections that maximize existing grid capacity through real-time management and data center load flexibility. Software platforms like portfolio company Camus Energy are using this approach to reduce time-to-power by working with grid constraints rather than against them.
Once connected, AI workloads create unprecedented volatility. Training and inference alternate between compute-intensive phases (high power) and communication-intensive phases (lower power), with fluctuations of hundreds of megawatts on millisecond timescales. For grid operators accustomed to slow, predictable load changes, this is unfamiliar territory and not contemplated in the controls architecture in legacy grid systems. Massive, rapid load swings threaten voltage and frequency stability across entire regional grids.
Massive power fluctuations can lead to grid instability, as nodes simultaneously transition between compute-intensive (high power) and communication-intensive (low power) phases.
The operational challenges that follow are significant. First, the data center must smooth load profiles to prevent grid instability. Smart energy management addresses this: workload schedulers that understand which jobs can pause run in the background, energy storage that buffers fluctuations, and on-site generation for baseload support. Solutions like Emerald AI and Phaidra actively manage consumption patterns while maximizing compute utilization. Second, data centers must respond appropriately when grid events occur. Grid operators increasingly expect large loads to provide services (or at the least stability), not just consume power passively. At the application layer, schedulers must distinguish which training runs can pause during grid stress from which inference services require guaranteed power for latency SLAs. These signals propagate down through compute hardware to facility-level power management. Solid-state transformers function as intelligent power routers, dynamically balancing between grid, solar, and storage inputs to optimize cost and availability. Virtual power plant platforms like portfolio company Leap can rapidly deliver capacity from existing grid assets while also enabling data centers to participate - effectively turning them into distributed grid resources, which is a significant shift in how we think about what a data center represents as a grid asset.
This kind of coordination doesn’t exist today. It requires integration across what are currently vendor-balkanized layers.
The sheer scale and complexity of AI data centers requires both best-in-class performance at each layer and energy-aware coordination across the stack to deliver:
Faster deployment time via accelerated power access (more flexible siting of smaller data-centers regionally clustered and flexible interconnect)
Stable grid operation through full-stack energy management to stabilize data center loads and comply with stringent grid requirements
Higher compute utilization via smart workload orchestration that minimizes curtailment
Lower operating costs and lower carbon footprint through coordinated energy management
Individual elements provide value independently but amplify when operating as a system. HVDC distribution enables higher rack densities and provides the foundation for microgrids, energy storage, and solid-state transformers. When you combine those with workload scheduling and compute controllers, you can smooth data center load for grid stability, provide increased capacity when energy is available, and shed loads intelligently when curtailed.
Photonic interconnects provide the lowest power and fastest connections, and they simplify deployment by breaking up mega racks for easier siting. Absent algorithmic breakthroughs, training still favors dense, co-located clusters, but inference workloads are different - they can leverage optical connections to create logical mega data centers across distributed physical sites, which is where the deployment flexibility of photonics really pays off.
With this foundation in place, new optimization surfaces open up. Some opportunities are structural: they emerge because power, thermals, and grid constraints are now first-class design variables rather than their historical status as afterthoughts. Others are paradigmatic: step-function breakthroughs that can rewrite the constraints entirely. Our core thesis is that, regardless of where breakthroughs land, the stack is becoming energy-aware. Workloads will increasingly express flexibility and requirements - whether a job can pause, what latency it requires - and infrastructure will increasingly translate those intents into decisions about compute, power, thermals, and grid participation.
We see opportunities falling into two broad categories, though we’re certainly not capturing everything here.
Platform and infrastructure plays
These are companies building on the new constraints, not trying to eliminate them, but turning them into product advantages. They tend to have clearer paths to revenue and more predictable technical risk:
Platform orchestration and energy management is where we’re spending a lot of time. The facility-level platform that coordinates power distribution, energy storage, cooling, workload scheduling, and grid interaction. Someone is going to own this layer, and it’s unlikely that it is a single hardware vendor. Power components may commoditize, but the orchestration layer creates platform value with strong network effects, particularly systems providing bidirectional information exchange across applications, schedulers, hardware, facilities, and grid. We’re looking for platforms that can actuate, not just monitor, while working with existing hardware systems.
Workload-aware inference routing is the application-layer counterpart to facility-level orchestration - dynamic routing of inference requests across distributed infrastructure based on latency SLAs, energy costs, and real-time capacity. This feels inevitable to us; the question is how it integrates with the layers below and whether there is room for an independent company to own this functionality, or whether it will be insourced.
Accelerating deployment is perhaps the most immediately actionable category. Tools that turn data center load flexibility into faster grid interconnection, or that enable operators to site facilities based on real grid capacity rather than nameplate ratings. The unlock here is regulatory as much as it is technical, and is key to bringing data centers online in a reasonable timeframe.
Step-function bets
These are higher-variance opportunities where success could rewrite the constraints entirely. They’re harder to underwrite as a venture investor, but the market scale justifies taking thoughtful swings.
Novel memory architectures represent the biggest prize. High-throughput memory, application-specific memory architectures, compute-in-memory designs that eliminate the bandwidth bottleneck rather than engineering around it. A true breakthrough here simplifies everything downstream, which is why everyone is chasing it and why we remain cautious about picking winners too early. The scale, capital needs, and expertise to address the memory challenge are not for the faint of heart.
Alternative computational substrates: Thermodynamic computing, stochastic architectures, photonic compute - approaches that could achieve order-of-magnitude improvements in energy per operation for specific workloads. These are longer shots but the payoff would be enormous if any of them pan out at scale.
Photonic integration: Copackaged optics offers near-term power and latency gains by moving optical components onto the package. The material platform question - silicon photonics vs. InP vs. hybrid approaches - is still being contested;our sense is that the answer will vary by application, which may mean room for multiple winners.
Advanced cooling and power electronics: Immersion cooling, two-phase systems, and novel power conversion topologies. Current approaches have hard limits on rack density and efficiency. Innovations here could break through those limits in ways that unlock value across the rest of the stack.
We are energized by the challenge here: to maximize intelligence and outcomes delivered per watt. If you’re working on any of this, we’d love to hear from you.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.