On September 23, at the GMIF 2026 Global Memory Industry Innovation Summit, John Xavier Lionel, Head of Global Storage Business at Arm, delivered a keynote speech titled "From Data to Tokens: Reshaping Computing and Storage Architecture for Edge AI," sharing the latest insights on the evolution trends of computing and storage in the era of Generative Artificial Intelligence (AI).
AI is Reshaping Storage Demands
As AI model sizes continue to expand, context lengths grow, and the number of concurrent users increases rapidly, storage is becoming a key factor affecting AI system capabilities. During inference, the Decode phase requires continuous access to the KV Cache, meaning that memory capacity, bandwidth, and data transfer efficiency directly impact system performance and throughput. Therefore, for generative AI, system performance is no longer determined solely by computing resources but increasingly depends on whether memory, storage, and data transfer systems can consistently and efficiently provide the required data to computing resources.
Meanwhile, as generative AI scales up, Tokens have become an important unit for measuring AI operational costs. For example, if one million devices perform 20 AI interactions per day, each generating 500 Tokens, this results in approximately 10 billion Tokens generated daily. Each interaction's Tokens correspond to costs in computing, memory, storage, networking, and energy consumption. Thus, inference is transforming from a one-time technical investment into ongoing operational expenditure. For the industry, the truly important question in the future is no longer "how many Tokens can be generated," but "how efficiently and at what low cost valuable intelligence can be generated."
Competition in the AI era is also shifting from simply pursuing higher TOPS to establishing a more efficient system balance among computing, memory, storage, and data movement.
Distributed AI Enables Intelligence to Run in the Most Suitable Location
As AI continues to extend from the cloud to the edge, the industry's focus is shifting from the choice between "cloud or edge" to "how to run the right AI workloads in the most suitable location."
Different AI workloads have varying requirements for performance, latency, privacy, power consumption, and cost, so they need to run at the most appropriate computing tier. Lightweight tasks such as always-on sensing, wake-word detection, and anomaly detection are suitable for running locally on low-power endpoint devices; applications sensitive to latency and privacy, such as AI assistants and multimodal interactions, are better executed on high-performance edge devices like smartphones, PCs, XR devices, and robots; industry models and localized AI services can rely on edge infrastructure; while ultra-large models, complex reasoning, model training, and global services are better deployed in the cloud.
By reasonably distributing AI workloads across the cloud, edge infrastructure, and end devices, enterprises can achieve a better balance between capability, latency, privacy, cost, and energy efficiency. For users, this means faster response times, enhanced privacy protection, more reliable experiences, and more personalized intelligent services.
Storage Has Become a Key Component of Competitiveness in Edge AI Systems
Thanks to continuous advancements in Small Language Models (SLMs), quantization techniques, heterogeneous computing architectures, and software ecosystems, edge AI is entering a phase of rapid development. AI capabilities that could previously only run in the cloud are being deployed to endpoint devices, improving response speeds, protecting data privacy, and enabling local intelligent applications. Meanwhile, as model sizes, local knowledge bases, and user contexts continue to expand, edge AI systems face increasing demands for storage capacity, memory bandwidth, model loading speed, security, and power management, bringing new system design challenges.
For example, generative AI inference requires continuous access to model weights, context, and Key-Value Cache (KV Cache). Even with powerful AI accelerators, system performance will still be limited if data cannot flow efficiently between storage, memory, and compute units.
In this process, the role of storage systems is also undergoing profound changes. In the past, storage was mainly used to save operating systems, applications, and user data; however, as more AI capabilities migrate to the device side, AI assets such as models, vector data, local knowledge bases, and Agent memories are also beginning to reside long-term on devices. Storage now carries not just data, but the models, knowledge, and intelligent capabilities themselves.
Armv9 Computing Platform Expands the Boundaries of Edge AI Capabilities
As edge AI expands from lightweight applications such as always-on sensing, keyword detection, and anomaly detection to AI PCs, multimodal AI, robotics, and local generative AI applications, models and context are growing from megabytes to gigabytes, placing higher demands on system computing, memory, and storage capabilities.
For low-power edge devices, Arm Cortex-M CPUs can provide efficient AI capabilities within limited power budgets. Its core design principles are: minimize external data movement, keep active data as close to the compute units as possible, efficiently store models in Flash, and configure AI acceleration capabilities on demand.
As model sizes, context lengths, and local knowledge bases continue to grow, edge AI is entering the era of large models. At this stage, system challenges come not only from computing power itself but also from capacity, bandwidth, and data movement efficiency. Based on the same design philosophy, Arm has launched the Armv9 Edge AI Computing Platform featuring the Cortex-A320 CPU and Ethos-U85 NPU. Through an enhanced memory architecture, more efficient data access capabilities, and collaborative computing design between CPU and NPU, it supports running large language models with over one billion parameters, providing stronger edge-side computing capabilities for generative AI, multimodal AI, and agent applications.
The next competition in the AI era is no longer just about TOPS; it is a contest of system design that achieves optimal synergy among computing, memory, storage, and data movement. Arm is helping ecosystem partners unleash AI capabilities more efficiently through its computing platforms covering Everythinggj from low-power endpoints to high-performance edge devices, along with continuously optimized heterogeneous computing architectures, driving edge AI into the billion-parameter era.
– End –

