AMD has announced that the Instinct MI300X accelerator and Instinct MI300A accelerated processing unit (APU) are now available as part of the company’s new Instinct MI300 Series, aimed at high-performance computing, generative AI, and similar applications.
Here’s what AMD had to say about the Instinct MI300X, which is said to feature industry-leading memory bandwidth for generative AI and leadership performance for large language model (LLM) training and inferencing, as well as increased throughput versus the NVIDIA H100 HGX when running inference for certain LLMs:
AMD Instinct MI300X accelerators are powered by the new AMD CDNA 3 architecture. When compared to previous generation AMD Instinct MI250X accelerators, MI300X delivers nearly 40% more compute units, 1.5x more memory capacity, 1.7x more peak theoretical memory bandwidth3 as well as support for new math formats such as FP8 and sparsity; all geared towards AI and HPC workloads.
Today’s LLMs continue to increase in size and complexity, requiring massive amounts of memory and compute. AMD Instinct MI300X accelerators feature a best-in-class 192 GB of HBM3 memory capacity as well as 5.3 TB/s peak memory bandwidth to deliver the performance needed for increasingly demanding AI workloads. The AMD Instinct Platform is a leadership generative AI platform built on an industry standard OCP design with eight MI300X accelerators to offer an industry leading 1.5TB of HBM3 memory capacity. The AMD Instinct Platform’s industry standard design allows OEM partners to design-in MI300X accelerators into existing AI offerings and simplify deployment and accelerate adoption of AMD Instinct accelerator-based servers.
Compared to the Nvidia H100 HGX, the AMD Instinct Platform can offer a throughput increase of up to 1.6x when running inference on LLMs like BLOOM 176B and is the only option on the market capable of running inference for a 70B parameter model, like Llama, on a single MI300X accelerator; simplifying enterprise-class LLM deployments and enabling outstanding TCO.
And here’s the word on the Instinct MI300A, which AMD has noted as being one of the first products of its kind:
The AMD Instinct MI300A APUs, the world’s first data center APU for HPC and AI, leverage 3D packaging and the 4th Gen AMD Infinity Architecture to deliver leadership performance on critical workloads sitting at the convergence of HPC and AI. MI300A APUs combine high-performance AMD CDNA 3 GPU cores, the latest AMD “Zen 4” x86-based CPU cores and 128GB of next-generation HBM3 memory, to deliver ~1.9x the performance-per-watt on FP32 HPC and AI workloads, compared to previous gen AMD Instinct MI250X.
Energy efficiency is of utmost importance for the HPC and AI communities, however these workloads are extremely data- and resource-intensive. AMD Instinct MI300A APUs benefit from integrating CPU and GPU cores on a single package delivering a highly efficient platform while also providing the compute performance to accelerate training the latest AI models. AMD is setting the pace of innovation in energy efficiency with the company’s 30×25 goal, aiming to deliver a 30x energy efficiency improvement in server processors and accelerators for AI-training and HPC from 2020-2025.
The APU advantage means that AMD Instinct MI300A APUs feature unified memory and cache resources giving customers an easily programmable GPU platform, highly performant compute, fast AI training and impressive energy efficiency to power the most demanding HPC and AI workloads.
AMD Instinct MI300 Series Specifications
| AMD Instinct | Architecture | GPU CUs | CPU Cores | Memory | Memory Bandwidth (Peak theoretical) | Process Node | 3D Packaging w/ 4th Gen AMD Infinity Architecture |
| MI300A | AMD CDNA 3 | 228 | 24 “Zen 4” | 128GB HBM3 | 5.3 TB/s | 5nm / 6nm | Yes |
| MI300X | AMD CDNA 3 | 304 | N/A | 192GB HBM3 | 5.3 TB/s | 5nm / 6nm | Yes |
| Platform | AMD CDNA 3 | 2,432 | N/A | 1.5 TB HMB3 | 5.3 TB/s per OAM | 5nm / 6nm | Yes |
“AMD Instinct MI300 Series accelerators are designed with our most advanced technologies, delivering leadership performance, and will be in large scale cloud and enterprise deployments,” said Victor Peng, president, AMD. “By leveraging our leadership hardware, software and open ecosystem approach, cloud providers, OEMs and ODMs are bringing to market technologies that empower enterprises to adopt and deploy AI-powered solutions.”


Discussion (8 replies)
Join Discussion →I hope we never see the day when bullshit like this is integrated in mainstream CPU's.
Uhhh too late? Even mobile CPU's are doing AI acceleration now. It's been like that or projected to be that for over a year now.
To not expect it at this point (even if it is just a marketing blurb) is primary school yard fit throwing. It's coming, it is literally already happening in modern desktop/laptop CPU's.
Just just hope they don't or in your case do paywall it behind some subscription theme to unlock AI acceleration. (Like Intel wants to do with their next gen Xeon CPU's. )
Apple has had an AI accelerator in their SOC since A11 (iPhone X) -- they first used it to power Face ID/ biometrics.
The latest Apple Watch even added 4 AI cores to the watch SOC - they use it there to process voice commands for Siri and new gesture-based controls, among other things.
Well, to each their own, but I want absolutely no AI computation on my computer or any devices of mine. I also don't want anything derived from AI computation elsewhere to fins its way onto my computer or any other devices.
I want honest, local, static computing. That's it. Not just today, but every day, for now and for all eternity.
Lets be real all AI processing nodes are, are nodes with instruction sets for better processing of specific data types. In this case used for 'ai' like functions... that are just enhanced scripting. Just beyond what is easily processed on a normal core.
I'm fine with whatever accelerators they want to put in my stuff - so long as I have a reasonable understanding of what it's doing and who it's talking it.
Google reportedly taps AMD to design next-generation TPU — hybrid AI ASIC could integrate on-package CPU cores for reinforcement learning
News
By Anton Shilov
Published 4 hours ago
A custom TPU for agentic and RL workloads?
"Market chatter suggests is working with AMD on a TPU project in the v10 generation,"
— a SemiAnalysis note for clients cited by Sean reads.
Google hardly needs AMD to design a conventional TPU. Hence, chances that AMD will implement Google's TPU v10i for inference or V10t for training are low. Hence, Google might need something only a CPU maker like AMD could provide, including CPU IP, programmable logic, interconnects, or certain advanced packaging know-how.
SemiAnalysis claims that Google and its customers are pushing for TPUs with on-package CPU cores for reinforcement learning and potentially other CPU-heavy workloads. While conventional LLM training remains overwhelmingly accelerator-heavy, reinforcement learning for reasoning and agentic models can require considerably more general-purpose compute around accelerator operations.
This is where AMD comes into play, as it already has experience developing a data center-grade design — the Instinct MI300A — that packs both x86 and accelerator chiplets.
A hypothetical Google design could therefore combine Google-developed TPU compute chiplets with AMD CPU and HBM in a tightly integrated package built by AMD.
https://www.tomshardware.com/tech-industry/artificial-intelligence/google-reportedly-taps-amd-to-design-next-generation-tpu-hybrid-ai-asic-could-integrate-on-package-cpu-cores-for-reinforcement-learning
Pretty big and unexpected news. Weird that they chose AMD even though they have great success making their TPU accelerators. This is something they might be doing only due to time related pressure because they have enough money to deploy their own licensed ARM CPU cores if they wanted or develop something in-house. Could also be that they lost someone important from their silicon team.
Ahctually....I can see someone creating a Transmeta Crusoe like AI accelerated CPU that optimizes running code the longer it runs. That would be cool. As then developers can just focus on writing code and not spending time optimizing it.