As artificial intelligence (AI) evolves from chatbots to reasoning and agent stages, the bottleneck in AI systems is shifting from processing units to memory. SK Hynix has identified that the core challenges of AI systems, including performance, energy efficiency, and resource utilization, are increasingly centered on memory. The company argues that 'memory-centric computing,' which reduces the distance between memory and processing units while allowing multiple accelerators to share memory, is the next frontier for AI expansion.
Joo Young-pyo, Vice President of System Architecture at SK Hynix, stated at the 18th Good Growth, Good Jobs Global Forum (2026 GGGF) held in Seoul on September 2, "It seems that memory is the issue these days. Performance, energy efficiency, and resource optimization challenges are all centered around memory." He emphasized, "The bottleneck is where computing is focused, and currently, that bottleneck is memory."
The demand for memory is being driven by changes in AI services. While traditional chatbots only needed to respond at a conversational pace, agent-based AI systems communicate with each other, generating and executing code, correcting errors, and reworking tasks. Joo explained, "When conversing with humans, around 20 tokens per second are sufficient, but when machines communicate, they need to respond at thousands of tokens per second," indicating a significant increase in system performance requirements.
In response, SK Hynix is innovating memory solutions along three axes: performance, energy efficiency, and resource utilization. Key factors include how much memory bandwidth can be increased within a given area and how to reduce the energy required to transfer a single bit of data to processing units like GPUs, NPUs, and TPUs. Enhancing system utilization to ensure expensive GPUs and memory operate continuously is also a priority.
High Bandwidth Memory (HBM) is a prime example of this transformation. Joo noted, "While the rise of GPUs has brought attention to HBM, it is also true that the existence of memory like HBM has enabled GPUs to achieve better performance." He added that there is growing concern in the industry about whether the current combination of HBM and GPUs will be sufficient in the era of agent-based AI.
One alternative is to further reduce the distance between memory and processing units. Stacking DRAM in three dimensions above or below processors can shorten data transfer distances. Processing In Memory (PIM), which performs calculations directly within or near memory, follows the same principle. Joo remarked, "It takes significantly more energy to retrieve data from memory and deliver it to processing units than to perform a calculation. As the distance increases, energy efficiency drops sharply." However, he acknowledged that 3D stacking presents technical challenges, such as heat generation and reduced storage space.
Another proposed solution is CXL, which allows multiple accelerators to share memory. Previously, GPUs had to wait for one another to finish tasks before passing data, but with shared memory, data can be stored and retrieved as needed by the next processing unit. SK Hynix is researching ways to reduce waste in data copying and synchronization processes between GPUs through this approach. Joo explained that while CXL was not specifically designed for AI, its open memory interface characteristics are expanding its application in AI systems.
Joo drew a line against the notion that optimizing AI algorithms will directly lead to reduced memory demand. He cited the 'Jevons Paradox,' where improved fuel efficiency in cars has led to increased overall oil consumption, stating, "As the cost of AI services decreases, more AI services will be utilized." This implies that even if individual services use less memory due to algorithm optimization, overall AI demand could still drive an increase in memory needs.
Joo concluded, "In the past, the CPU was the bottleneck of the system, and improving the CPU directly enhanced system performance. Now, the bottleneck is memory." He emphasized, "By enhancing memory performance and energy efficiency, AI systems can significantly improve."
* This article has been translated by AI.
Copyright ⓒ Aju Press All rights reserved.