Technology
Memory-Centric Computing Architectures to Overcome the von Neumann Bottleneck
Quick fact
In today's computers, moving data between memory and CPU can consume more than 100 times the energy of performing a single arithmetic operation, and this gap is widening every year.
Why this is interesting
You've felt it: your computer stalls when loading a huge file, even though the processor is fast. What if the memory did the computing instead of just storing data?
Read the full explanation
Understanding Memory-Centric Computing Architectures to Overcome the von Neumann Bottleneck
To appreciate memory-centric computing, first picture the classic von Neumann architecture: a central processing unit (CPU) fetches instructions and data from a separate memory, executes operations, and writes results back. This separation was revolutionary in the 1940s, but it created a fundamental constraint—every computation requires shuttling data back and forth over a limited bus. This is the von Neumann bottleneck. As processors got faster and software demands grew, moving data became the limiting factor, not the chip speed. Memory-centric computing flips this model. Instead of bringing data to the processor, it brings computation to the data. Think of it like a library: instead of running a book to a scholar each time, you make the bookshelf itself do some summarizing. In practice, memory-centric architectures integrate processing abilities directly into or very near the memory chips. Two broad flavors exist: near-memory computing (placing processor cores close to memory, e.g., stacked memory with logic) and compute-in-memory (using the memory cells themselves to perform logical or arithmetic operations). A key technology enabling compute-in-memory is the memristor, a resistor whose resistance can change based on the history of applied voltage. In a memory-centric circuit, an array of memristors can perform analog multiply-and-accumulate operations—the bread and butter of neural networks—simultaneously, using Ohm's law (current = voltage × conductance) and Kirchhoff's law (currents sum at a node). The data stays put; the answer emerges from the physics of the array.
A deeper explanation
The von Neumann bottleneck is not just a theoretical concern—it quantifies the energy and time cost of data movement. In a conventional system, every operation involves fetching instructions and data from memory, which can take hundreds of cycles and consume orders of magnitude more energy than the operation itself. This 'memory wall' means that increasing processor speed has diminishing returns because memory bandwidth and latency cannot keep pace. Memory-centric computing overcomes this by exploiting the physical properties of memory devices to perform computation in place. For example, a memristor can store a weight as its conductance. When a voltage corresponding to an input is applied, the current through it equals the product of input and weight—a multiply operation happens instantly in analog form. When a column of memristors is driven by an input vector, the currents sum along the column (due to Kirchhoff's current law), yielding a dot product without any data movement. This forms the core of matrix-vector multiplication, the dominant operation in deep learning and scientific computing. Beyond analog computing, processing-in-memory (PIM) can also be digital. For instance, logic operations can be performed by sending signals into memory arrays, or dedicated logic circuits can be embedded on the memory chip. The common goal is to avoid transferring data across the slow, energy-hungry CPU-memory interface. Memory-centric designs also include near-memory computing, where high-bandwidth 3D-stacked memory (like HBM) sits directly beneath a processor, dramatically reducing the distance data travels. The implications are enormous: for data-intensive applications like machine learning training and inference, real-time analytics, and genome sequencing, memory-centric architectures can deliver order-of-magnitude energy savings and speedups. However, they come with challenges: programming models are new, precision can be limited in analog approaches, and manufacturing requires advanced integration. Despite this, they represent a vital pathway to reclaiming performance gains as Moore's law slows.