In just a few years, AI’s explosive growth has made it a mainstay in everyday life, but also sparked major resource consumption concerns. Large language models (LLMs) such as ChatGPT and Claude have significantly larger computational and energy costs versus smaller neural networks: ChatGPT alone drew an estimated 23 gigawatt hours of electricity per month in 2024, exceeding the annual consumption of many countries. Yet, LLMs continue to be widely deployed, making efficiency improvements ever more critical.
Faced with this challenge, Tao Luo, Head of the High Performance Computing Chapter at the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), and Zhehui Wang, A*STAR IAIC Senior Scientist, are adapting novel electronic components known as memristor crossbars—originally trialled to significant success in computer vision models—to LLMs.
“Memristor crossbars are very compact memory-and-computing devices where the memory portions are arranged in crossbar-like structures, allowing other portions to physically perform mathematical operations directly where the data is stored,” said Luo. “The key advantages of these devices are their density and energy efficiency.”
This two-in-one device architecture cuts down a typically energy-hungry portion of computing: data movement. Traditional computers store data in a memory device before moving them to a separate processor device for computation; memristor crossbars remove the need for data to travel, potentially easing a major bottleneck for LLM deployment.
Adapting memristor crossbars for LLMs, however, can be challenging. Besides using volumes of data so large they push the limits of denser on-chip memory technologies, LLMs also contain attention mechanisms and rely on nonlinear operations, which traditional memristor crossbars have not been optimised for.
Together with collaborators from Nanyang Technological University, Singapore; the National University of Singapore; and the Chinese Academy of Sciences, the team developed a memristor crossbar architecture that combines a computation crossbar and a dense crossbar to overcome these hurdles.
“The dense crossbar provides large-capacity storage, while the computation crossbar provides flexible and energy-efficient computing,” Wang said. “By integrating them on the same chip or package, our architecture reduces the need to move data between separate chips or external memory.”
Through simulation experiments with several state-of-the-art LLMs including GPT-3 and LLaMa, the researchers found that their memristor crossbar architecture was as accurate as conventional GPU-based computation, while also being about 68 times more efficient overall in terms of chip space and computation time combined. Their device could also be implemented on chip areas 39 times smaller than traditional memristor crossbars while reducing energy consumption 18-fold.
“We showed a feasible path for deploying LLMs on memristor-based architectures, not just smaller or more regular neural networks,” Wang commented. “This work achieves an architectural synergy between dense storage and flexible computation, providing an important step toward more compact and sustainable AI hardware.”
Moving forward, the team aims to scale their memristor crossbar design to support larger models and move it towards practical deployment.
The A*STAR-affiliated researchers contributing to this research are from the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC).
