Amazon SageMaker AI inference now supports G7 instances
Amazon SageMaker AI inference now supports G7 instances powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, delivering up to 4.6x AI inference performance over previous-generation G6 instances
Amazon SageMaker AI inference now supports G7 instances powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. These instances offer up to 4.6x AI inference performance compared to previous-generation G6 instances. G7 instances provide 32 GB of GPU memory per GPU with 5th Generation Tensor Cores, up to 700 Gbps of EFA-enabled networking (7x compared to G6), and up to 7.6 TB of local NVMe SSD storage for keeping large models close to compute. These capabilities make G7 instances well suited for serving models in the 7B-30B parameter range, image and video generation workloads, and multi-model inference endpoints that benefit from higher memory bandwidth and throughput. You can deploy models on G7 instances using the SageMaker AI Inference console, API, or SDK by specifying G7 instance types (such as ml.g7.xlarge through ml.g7.48xlarge) in your endpoint configuration.