SageMaker AI inference expands to G7e instances in new regions
Amazon SageMaker AI inference now offers G7e instances in Seoul, London, and Tokyo regions, supporting large LLM inference and high-memory workloads
Amazon SageMaker AI inference now offers EC2 G7e instances in Asia Pacific (Seoul), Europe (London), and Asia Pacific (Tokyo) regions. G7e instances feature up to eight NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs with 96GB of memory per GPU, 5th Generation Intel Xeon processors, and up to 1,600 Gbps of Elastic Fabric Adapter networking bandwidth. This delivers up to 2.3x inference performance compared to previous-generation G6e instances. With up to 768GB of total GPU memory on a single instance, G7e instances enable serving medium-to-large language models of up to 70 billion parameters with FP8 precision without multi-node configurations. These instances are well suited for LLM inference, image and video generation, spatial computing, and scientific computing workloads requiring high GPU memory capacity and bandwidth.