VAST Data and AMD say they have expanded their collaboration to combine VAST’s AI Operating System with AMD’s 6th generation EPYC CPUs and Instinct GPUs, targeting large-scale AI training, inference and agentic AI deployments for cloud providers and enterprises.
In a statement, the companies positioned the expanded partnership as a response to rising infrastructure costs and GPU underutilisation as organisations shift focus from training to production inference. The announcement included performance claims from early VAST testing using an AMD Instinct MI355X GPU, reporting a 9x improvement in time-to-first-token and a 9.7x increase in token throughput when using VAST for KV cache offloading under high-concurrency agentic AI workloads. VAST said the performance and cost claims were reviewed but not independently verified by AMD.
The collaboration includes VAST selecting 6th Gen AMD EPYC processors, previously codenamed “Venice,” to power new generations of its CBox and EBox platforms. VAST said PCIe Gen-6 support would provide increased I/O bandwidth and lower latency for data services including database, data warehousing and event streaming via VAST’s DataBase and DataEngine capabilities.
VAST, AMD and DriveNets also announced an AI infrastructure reference architecture based on AMD Helios rack-scale infrastructure, the VAST AI OS and DriveNets AI Fabric networking. The reference designs cover model training, inference, reinforcement learning and KV cache workloads, with sizing guidance aimed at simplifying deployment for enterprises and AI cloud providers.
Additional ecosystem work includes collaboration with TensorMesh and EmbeddedLLM on inference architectures intended for agentic AI use cases, as well as KV cache and inference optimisations combining AMD Instinct GPUs, AMD Infinity Context, AMD ROCm software and the VAST AI OS.
VAST said it is integrating automated KV cache lifecycle management through its data lifecycle policies, aimed at expiring and deleting cache data to help meet security, privacy and compliance requirements. The announcement also highlighted AMD Pensando Pollara 400 AI NICs for moving data between GPU memory and VAST’s NVMe-based storage using NFS over TCP and NFS over RDMA.
“The industry is discovering that inference is fundamentally a data problem. Success depends on how effectively organisations can bring data, compute, memory and intelligence together as a single system,” said John Mao, vice president of global technology alliances at VAST Data, in a release.
AMD corporate vice president Derek Dicker said the companies are targeting “infrastructure efficiency” as AI moves into production environments.
Several infrastructure and cloud providers, including Core42, Crusoe and Vultr, provided supporting statements describing the partnership in the context of production AI deployments and enterprise requirements.
VAST said it showcased the architecture and reference designs at AMD Advancing AI 2026 in San Francisco, held July 22–23.

