Gemma-4-31B-it-assistant and Gemma-4-31B-IT-NVFP4 models now available on Amazon SageMaker JumpStart
Amazon SageMaker JumpStart now includes Google DeepMind's Gemma-4-31B-it-assistant and NVIDIA's Gemma-4-31B-IT-NVFP4 models, expanding the portfolio of foundation models for AWS customers.
Google DeepMind's Gemma-4-31B-it-assistant and NVIDIA's Gemma-4-31B-IT-NVFP4 models are now available on Amazon SageMaker JumpStart. These models adapt Google's flagship 31B dense architecture for enterprise workloads, offering both full-precision and optimized quantized variants. The Gemma-4-31B-it-assistant specializes in multimodal inference, coding, and agent workflows, accepting text and images (with video treated as frame sequences) and outputting text. It supports a 256K token context window and over 140 languages, ranking third among open models on the Arena AI text reader leaderboard. The Gemma-4-31B-IT-NVFP4 is quantized to 4-bit FP4 precision using NVIDIA's ModelOpt framework, reducing memory usage to about 18.5 GB (68% smaller than the base model) and speeding up inference by roughly 2.5x while retaining 97-99% of the original model's quality. SageMaker JumpStart allows these models to be deployed in just a few clicks for specific AI use cases.
Why it matters
Amazon SageMaker JumpStart is a platform that allows AWS customers to easily access AI foundation models, and the addition of these new models enables a range of AI applications, including multimodal reasoning and agent development.