A field report on running Google's Gemma-4 on AWS Inferentia2: mixed attention heads, the vLLM / optimum-neuron / NxD dead-ends, and the neuronx-cc compiler limits.
Back to Blog
AI & ML 16 min read
Porting Gemma-4 (2B / 4B / 12B) to AWS Inferentia2
xbill
July 13, 2026
Originally published on Dev.to: View original article →