Highlights

In brief

A novel memristor architecture incorporating a dense crossbar and computation crossbar enables deployments of state-of-the-art LLMs on a single chip, reducing time and energy losses through off-chip communication.

Photo by Svitlana Lishchyshyna | Magnific

Greener crossbars for large language models

10 Aug 2026

Inspired by computer vision hardware, researchers adapt a two-in-one crossbar-based chip design to boost the energy efficiency of AI systems.

In just a few years, AI’s explosive growth has made it a mainstay in everyday life, but also sparked major resource consumption concerns. Large language models (LLMs) such as ChatGPT and Claude have significantly larger computational and energy costs versus smaller neural networks: ChatGPT alone drew an estimated 23 gigawatt hours of electricity per month in 2024, exceeding the annual consumption of many countries. Yet, LLMs continue to be widely deployed, making efficiency improvements ever more critical.

Faced with this challenge, Tao Luo, Head of the High Performance Computing Chapter at the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), and Zhehui Wang, A*STAR IAIC Senior Scientist, are adapting novel electronic components known as memristor crossbars—originally trialled to significant success in computer vision models—to LLMs.

“Memristor crossbars are very compact memory-and-computing devices where the memory portions are arranged in crossbar-like structures, allowing other portions to physically perform mathematical operations directly where the data is stored,” said Luo. “The key advantages of these devices are their density and energy efficiency.”

This two-in-one device architecture cuts down a typically energy-hungry portion of computing: data movement. Traditional computers store data in a memory device before moving them to a separate processor device for computation; memristor crossbars remove the need for data to travel, potentially easing a major bottleneck for LLM deployment.

Adapting memristor crossbars for LLMs, however, can be challenging. Besides using volumes of data so large they push the limits of denser on-chip memory technologies, LLMs also contain attention mechanisms and rely on nonlinear operations, which traditional memristor crossbars have not been optimised for.

Together with collaborators from Nanyang Technological University, Singapore; the National University of Singapore; and the Chinese Academy of Sciences, the team developed a memristor crossbar architecture that combines a computation crossbar and a dense crossbar to overcome these hurdles.

“The dense crossbar provides large-capacity storage, while the computation crossbar provides flexible and energy-efficient computing,” Wang said. “By integrating them on the same chip or package, our architecture reduces the need to move data between separate chips or external memory.”

Through simulation experiments with several state-of-the-art LLMs including GPT-3 and LLaMa, the researchers found that their memristor crossbar architecture was as accurate as conventional GPU-based computation, while also being about 68 times more efficient overall in terms of chip space and computation time combined. Their device could also be implemented on chip areas 39 times smaller than traditional memristor crossbars while reducing energy consumption 18-fold.

“We showed a feasible path for deploying LLMs on memristor-based architectures, not just smaller or more regular neural networks,” Wang commented. “This work achieves an architectural synergy between dense storage and flexible computation, providing an important step toward more compact and sustainable AI hardware.”

Moving forward, the team aims to scale their memristor crossbar design to support larger models and move it towards practical deployment.

The A*STAR-affiliated researchers contributing to this research are from the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC).

Want to stay up to date with breakthroughs from A*STAR? Follow us on Twitter and LinkedIn!

References

Wang, Z., Luo, T., Liu, C., Liu, W., Goh, R.S.M., et al. Enabling energy-efficient deployment of large language models on memristor crossbar: A synergy of large and small. IEEE Transactions on Pattern Analysis and Machine Intelligence 47 (2), 916–933 (2025). | article

About the Researchers

Tao Luo received his PhD degree from the School of Computer Science and Engineering, Nanyang Technological University, Singapore, in 2017. He is currently the Head of the High Performance Computing Chapter at the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC). He has led multiple research and industry projects and published extensively in leading international conferences and journals. His research interests include high-performance computing, green and efficient AI, quantum computing, hardware-software co-exploration and AI applications.
Zhehui Wang received his PhD degree in electronic and computer engineering from the Hong Kong University of Science and Technology, Hong Kong, in 2016. He is currently a Senior Scientist at the A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), where he leads the hardware-aware AI research area. His research interests include efficient AI deployment, AI on emerging computing technologies, hardware-software co-design, green AI, and high-performance computing.

This article was made for A*STAR Research by Wildtype Media Group