Memory for the model you actually want to run
192 GB of local HBM3 gives serving teams room for larger models, longer context, and higher-concurrency deployments without treating memory as an afterthought.
AMD Instinct accelerator
A high-memory AMD accelerator for production inference. Start with one isolated GPU VM or scale to an eight-GPU server without changing the operational model.
At a glance
MI300X pairs a large HBM3 footprint with the bandwidth needed to serve open-source models reliably, without turning deployment into a hardware project for your team.
HBM3 memory
192 GB
Memory bandwidth
5.3 TB/s
Compute units
304
Accelerators per server
8x
Inference fit
The accelerator is only part of the offering. Hot Aisle automates the networking, PXE, operating system, ROCm, virtualization, and billing layers required to make this capacity usable.
192 GB of local HBM3 gives serving teams room for larger models, longer context, and higher-concurrency deployments without treating memory as an afterthought.
5.3 TB/s of bandwidth keeps the accelerator supplied as requests move through the model, supporting responsive token generation at production scale.
Provision a NUMA-balanced KVM virtual machine for a single GPU, or take a complete eight-GPU server when the workload calls for it.
System profile
Infrastructure is more useful when the physical system, virtualization model, and control surface are designed together.
CDNA 3 combines eight XCDs, 304 compute units, and 256 MB of AMD Infinity Cache in a dense accelerator designed for model serving.
Eight MI300X accelerators are available in each Dell PowerEdge XE9680, with a high-bandwidth fabric and full-system automation already in place.
Deploy MI300X
Create a team, add credit, and provision isolated AMD GPU compute from the terminal UI, API, or CLI. No sales handoff is required.
Quick start