Google logo

Google: Gemma 4 31B IT

Largest Gemma 4 instruction model for open, self-hosted chat and reasoning

Source-linked gemma Open weights Released 2026-04-02
API record Report
InputT
OutputT
Input price$0.09/M
Output price$0.34/M
Context262.144K
Max output32.768K
Providers33
Inference availability

Providers

Provider-specific identifiers, limits, and listed prices per million tokens. Every row links back to the provider's own documentation.

Report a price
Providers offering Gemma 4 31B IT
ProviderProvider model IDContextMax outputInputOutputCache readCapabilitiesDocs
Cloudflare AI Gateway google-ai-studio/gemma-4-31b-it 262.144K 32.768K ReasoningToolsJSON Docs ↗
QVAC gemma4-31b 262.144K 32.768K ReasoningToolsJSON Docs ↗
Requesty gemma-4-31b-it 262.144K 8.192K ReasoningTools Docs ↗
Kenari gemma-4-31b-it 262.144K 32.768K ReasoningToolsJSON Docs ↗
Nvidia google/gemma-4-31b-it 256K 16.384K ReasoningTools Docs ↗
Google gemma-4-31b-it 262.144K 32.768K ReasoningToolsJSON Docs ↗
UnoRouter gemma-4-31b-it:free 262.144K 32.768K ReasoningToolsJSON Docs ↗
OpenRouter google/gemma-4-31b-it 262.144K 16.384K $0.09 $0.34 $0.05 ReasoningToolsJSON Docs ↗
Kilo Gateway google/gemma-4-31b-it 262.144K 16.384K $0.09 $0.34 $0.05 ReasoningToolsJSON Docs ↗
DevPass (LLM Gateway) gemma-4-31b-it 262.144K 32.768K $0.1 $0.25 $0.01 ReasoningToolsJSON Docs ↗
CrofAI gemma-4-31b-it 262.144K 262.144K $0.1 $0.3 $0.02 ReasoningToolsJSON Docs ↗
NanoGPT google/gemma-4-31b-it 262.144K 131.072K $0.1 $0.45 $0.05 ReasoningToolsJSON Docs ↗
Lilac google/gemma-4-31b-it 262.1K 262.1K $0.11 $0.35 ReasoningToolsJSON Docs ↗
FastRouter google/gemma-4-31b-it 262.144K 32.768K $0.13 $0.38 ReasoningToolsJSON Docs ↗
SiliconFlow google/gemma-4-31B-it 262.144K 262.144K $0.13 $0.4 ToolsJSON Docs ↗
OrcaRouter google/gemma-4-31b-it 262.144K 32.768K $0.13 $0.38 $0.02 ReasoningToolsJSON Docs ↗
NEAR AI Cloud google/gemma-4-31B-it 262.144K 32.768K $0.13 $0.4 $0.026 ReasoningToolsJSON Docs ↗
Deep Infra google/gemma-4-31B-it 262.144K 32.768K $0.13 $0.38 ReasoningToolsJSON Docs ↗
Vercel AI Gateway google/gemma-4-31b-it 262.144K 131.072K $0.14 $0.4 ToolsJSON Docs ↗
Merge Gateway google/gemma-4-31b-it 262.144K 65.536K $0.14 $0.4 Reasoning Docs ↗
Friendli google/gemma-4-31B-it 262.144K 32.768K $0.14 $0.4 ReasoningToolsJSON Docs ↗
Amazon Bedrock google.gemma-4-31b 262.144K 32.768K $0.14 $0.4 ReasoningToolsJSON Docs ↗
Abacus google/gemma-4-31b-it 262.144K 131.072K $0.14 $0.4 ReasoningToolsJSON Docs ↗
Hugging Face google/gemma-4-31B-it 262.144K 32.768K $0.14 $0.4 ReasoningToolsJSON Docs ↗
NovitaAI google/gemma-4-31b-it 262.144K 131.072K $0.14 $0.4 ReasoningToolsJSON Docs ↗
Crusoe google/gemma-4-31b-it 262.144K 32.768K $0.14 $0.4 $0.14 ReasoningToolsJSON Docs ↗
ai& google/gemma-4-31b-it 262.144K 32.768K $0.2 $0.5 ReasoningToolsJSON Docs ↗
Cortecs gemma-4-31b-it 262K 262K $0.223 $0.39 ReasoningToolsJSON Docs ↗
Infomaniak google/gemma-4-31B-it 100K 32.768K $0.25 $0.5 ReasoningToolsJSON Docs ↗
Tinfoil gemma4-31b 262.144K 32.768K $0.4 $1 ReasoningToolsJSON Docs ↗
Regolo AI gemma4-31b 100K 100K $0.46 $2.42 ReasoningToolsJSON Docs ↗
Pioneer google/gemma-4-31B-it 32.768K 32.768K $0.5 $0.5 $0.5 ReasoningToolsJSON Docs ↗
Cerebras gemma-4-31b 131.072K 40.96K $0.99 $1.49 ReasoningToolsJSON Docs ↗

Capability badges appear only where the provider catalog explicitly lists support. A blank cell means the source is silent, not that the feature is absent.

Listed rates

Price across providers

Input price per million tokens as published by each provider. Bars are drawn from listed rates only — no traffic weighting, since the catalog observes no requests.

Lowest input $0.09/M

Across 26 priced providers

Median input $0.14/M

Midpoint of listed rates

Highest input $0.99/M

11.0× the lowest listed rate

Output range $0.25 – $2.42

Per million output tokens

OpenRouter $0.09/M
Kilo Gateway $0.09/M
DevPass (LLM Gateway) $0.1/M
CrofAI $0.1/M
NanoGPT $0.1/M
Lilac $0.11/M
FastRouter $0.13/M
SiliconFlow $0.13/M
OrcaRouter $0.13/M
NEAR AI Cloud $0.13/M
Deep Infra $0.13/M
Vercel AI Gateway $0.14/M
Merge Gateway $0.14/M
Friendli $0.14/M
Amazon Bedrock $0.14/M
Abacus $0.14/M
Hugging Face $0.14/M
NovitaAI $0.14/M
Crusoe $0.14/M
ai& $0.2/M
Cortecs $0.223/M
Infomaniak $0.25/M
Tinfoil $0.4/M
Regolo AI $0.46/M
Pioneer $0.5/M
Cerebras $0.99/M
Cost calculator

Estimate a workload

$0.00
Excludes taxes, non-token charges, and tiered discounts.

2 providers list the identical $0.09 input rate, so price alone will not separate them — compare context limits, max output, and capabilities above.

Context limits also differ by provider, from 32.768K to 262.144K tokens. Compare the provider table above before choosing on price alone.

Specification

Capabilities

Recorded from the source catalog and provider listings.

Reasoning Yes
Tool calling Yes
Structured output Yes
Attachments Yes
Vision input Yes
Open weights Yes
Creator
Google
Model family
gemma
Knowledge cutoff
Not documented
License
Not documented
Release date
2026-04-02
Model ID
google/gemma-4-31b-it

Weights: Hugging Face ↗

Published evaluations

Benchmarks

Every result stays attached to its source, version, metric and harness. Scores from different versions are never merged, and the profile below plots each benchmark against its own population rather than on a shared scale.

Benchmark registry
Benchmark profile
Coding IndexCoding Index Agentic IndexAgentic Index IntelligenceIntelligence Index This model — Coding Index: 43.4 index (43.4th percentile) This model — Agentic Index: 6.7 index (31.0th percentile) This model — Intelligence Index: 15.4 index (27.2th percentile)

Each axis is this model's percentile among the 3 benchmarks it has published results for, measured against every other model with a score on that same benchmark. Percentiles are used because benchmarks do not share a scale — a 60 on one is not a 60 on another. Hover any point for the raw score.

Coding Index index Source ↗
Coding Index #115 of 204
Agentic Index index Source ↗
Agentic Index #132 of 192
Intelligence Index index Source ↗
Intelligence Index #138 of 191

Each strip shows every published score for that benchmark, with this model marked. The lighter shape behind the ticks is the density of results, and the dashed line is the median.

Cost against capability

Price and performance

Listed input price plotted against Coding Index, the benchmark with the widest published coverage that this model appears in.

0 43 86 $0.01 $0.1 $1 $10 Anthropic: Claude Fable 5.1 (batch) — $5/M, score 81.6 Claude Opus 5 — $5/M, score 78.0 Mistral: Ministral 3 8B 2512 (batch) — $0.075/M, score 9.7 Gemini 3.5 Flash Lite — $0.3/M, score 49.3 Mistral: Mistral Large 3 2512 (batch) — $0.25/M, score 20.1 OpenAI: GPT-3.5 Turbo (batch) — $0.25/M, score 10.7 OpenAI: o3 Mini High (batch) — $0.55/M, score 16.3 Qwen: Qwen3.8 Max (0902) — $2/M, score 71.8 Gemini 3.5 Flash — $1.5/M, score 70.1 Z.ai: GLM 5.3 Flash (batch) — $0.075/M, score 71.5 DeepSeek V4 Flash — $0.15/M, score 52.0 GPT-4 Turbo — $10/M, score 21.5 GPT-4.1 mini — $0.4/M, score 20.2 Claude Sonnet 5 — $2/M, score 71.5 Mistral: Ministral 3 8B 2512 — $0.15/M, score 9.7 OpenAI: o1 (batch) — $7.5/M, score 39.7 Z.ai: GLM 5.1 — $0.966/M, score 55.8 Gemini 3.7 Flash — $0.75/M, score 76.1 Z.ai: GLM 5.3 — $1.4/M, score 74.8 LongCat-2.0 — $0.3/M, score 45.3 Z.ai: GLM 5.3 (batch) — $0.7/M, score 74.8 Anthropic: Claude Opus 4.8 (batch) — $2.5/M, score 74.3 Google: Gemini 3.5 Flash (batch) — $0.75/M, score 70.1 DeepSeek V4 Pro — $0.435/M, score 59.4 OpenAI: GPT-5.5 (batch) — $2.5/M, score 74.9 GPT-6 Astra — $10/M, score 76.9 Anthropic: Claude Opus 4.7 — $5/M, score 73.6 OpenAI: GPT-5.4 Nano (batch) — $0.1/M, score 56.1 Inception: Mercury 2 — $0.25/M, score 31.1 Google: Gemini 3.1 Pro Preview (batch) — $1/M, score 68.8 Gemini 3.6 Flash — $0.75/M, score 69.2 Anthropic: Claude Sonnet 4.5 — $3/M, score 52.1 Qwen: Qwen3.5-9B (batch) — $0.17/M, score 28.7 Google: Gemini 2.5 Pro (batch) — $0.625/M, score 33.3 GPT-5.4 mini — $0.75/M, score 56.1 Gemma 4 26B A4B IT — $0.042/M, score 39.3 Claude Fable 5 — $10/M, score 76.5 Mistral: Mistral Small 4 (batch) — $0.075/M, score 26.6 Z.ai: GLM 5.3 Flash — $0.15/M, score 71.5 MiniMax: MiniMax M3 — $0.3/M, score 58.6 DeepSeek: DeepSeek V4 Pro 0813 (batch) — $0.66/M, score 68.8 Thinking Machines: Inkling Small (batch) — $0.5/M, score 52.9 Gemini 3.1 Flash Lite Preview — $0.25/M, score 34.7 Gemma 3 12B IT — $0.05/M, score 5.8 Anthropic: Claude Fable 5.1 — $10/M, score 81.6 Google: Gemma 4 31B (batch) — $0.39/M, score 43.4 Qwen: Qwen3.8 2.4T A95B (batch) — $2/M, score 71.9 Mistral: Devstral 2 2512 — $0.4/M, score 31.3 MoonshotAI: Kimi K2.7 Code (batch) — $0.95/M, score 60.8 DeepSeek V4 Flash 0731 — $0.05/M, score 69.1 DeepSeek-R1 — $0.7/M, score 24.6 Inkling — $1.87/M, score 52.1 OpenAI: gpt-oss-20b (batch) — $0.05/M, score 20.7 Muse Spark 1.1 — $1.25/M, score 71.3 MiMo-V2.5 — $0.14/M, score 56.8 Google: Gemini 3.5 Flash Lite (batch) — $0.15/M, score 49.3 Kimi K2 Thinking — $0.4/M, score 21.0 Mistral: Mistral Medium 3.5 (batch) — $0.75/M, score 46.9 GPT-5.6 Sol — $4/M, score 77.4 GPT-5.6 Luna — $0.2/M, score 71.4 Qwen: Qwen3.8 Max (0803) — $2/M, score 68.9 Mistral: Mistral Medium 3.1 (batch) — $0.2/M, score 20.5 MiMo-V2.5-Pro — $0.435/M, score 60.2 GPT-3.5-turbo — $0.5/M, score 10.7 GPT OSS 120B — $0.03/M, score 30.4 GPT-5.1 — $1.25/M, score 49.4 GPT-5 — $1.25/M, score 37.8 GPT-5.5 — $5/M, score 74.9 Qwen: Qwen3.8 27B — $0.42/M, score 68.1 Kimi K2.6 — $0.95/M, score 61.8 Gemma 4 31B IT — $0.09/M, score 43.4 Google: Gemini 3.8 Flash (batch) — $0.375/M, score 76.3 OpenAI: GPT-5 Mini (batch) — $0.125/M, score 15.6 Gemini 2.5 Pro — $1.25/M, score 33.3 MoonshotAI: Kimi K3 (batch) — $3/M, score 76.2 Inkling Small — $0.45/M, score 52.9 Anthropic: Claude Sonnet 5 (batch) — $1/M, score 71.5 DeepSeek V3.2 — $0.18/M, score 44.2 Anthropic: Claude Fable 5 (batch) — $5/M, score 76.5 Nemotron 3 Super 120B A12B — $0.2/M, score 37.7 IBM: Granite 4.2 8B — $0.06/M, score 22.4 Anthropic: Claude Opus 4.8 — $5/M, score 74.3 Thinking Machines: Inkling (batch) — $1/M, score 52.1 Gemma 3 4B IT — $0.04/M, score 2.7 SpaceXAI: Grok 4.6 — $2/M, score 76.8 DeepSeek: DeepSeek V3.1 Terminus — $0.27/M, score 43.5 GPT-5 Mini — $0.25/M, score 15.6 Claude Opus 5 (batch) — $2.5/M, score 78.0 SpaceXAI: Grok 4.3 — $1.25/M, score 42.2 OpenAI: GPT-6 Astra (batch) — $5/M, score 76.9 OpenAI: GPT-5.6 Terra (batch) — $1/M, score 76.7 OpenAI: GPT-5.4 (batch) — $1.25/M, score 71.1 Mistral: Mistral Large 3 2512 — $0.5/M, score 20.1 OpenAI: GPT-4o-mini (batch) — $0.075/M, score 11.4 DeepSeek V4 Pro 0813 — $0.442/M, score 68.8 Z.ai: GLM 5.2 (batch) — $0.7/M, score 68.8 Google: Gemini 3.6 Flash (batch) — $0.375/M, score 69.2 Nemotron 3 Nano 30B A3B — $0.05/M, score 14.4 GPT-5.6 Terra — $2/M, score 76.7 MiniMax: MiniMax M3 (batch) — $0.3/M, score 58.6 Gemini 3.8 Flash — $0.75/M, score 76.3 NVIDIA: Nemotron 3 Ultra (batch) — $0.6/M, score 49.3 MiniMax: MiniMax M2.7 — $0.3/M, score 52.6 SpaceXAI: Grok 4.5 — $2/M, score 72.4 Mistral: Mistral Medium 3.1 — $0.4/M, score 20.5 OpenAI: gpt-oss-120b (batch) — $0.15/M, score 30.4 OpenAI: GPT-4.1 Mini (batch) — $0.2/M, score 20.2 Gemma 3 27B IT — $0.08/M, score 10.1 GPT-4 — $30/M, score 13.1 Solar Pro 4 — $0.3/M, score 52.7 Gemini 3.1 Pro Preview — $2/M, score 68.8 Anthropic: Claude Opus 4.7 (batch) — $2.5/M, score 73.6 OpenAI: GPT-4.1 Nano (batch) — $0.05/M, score 11.1 OpenAI: GPT-5.1 (batch) — $0.625/M, score 49.4 Kimi K2.7 Code — $0.95/M, score 60.8 Kimi K3 — $3/M, score 76.2 GPT-5.4 — $2.5/M, score 71.1 Anthropic: Claude Sonnet 4.6 — $3/M, score 63.0 Anthropic: Claude Sonnet 4.6 (batch) — $1.5/M, score 63.0 Google: Gemini 3.7 Flash (batch) — $0.375/M, score 76.1 Muse Spark 1.2 — $1.25/M, score 72.2 DeepSeek: DeepSeek V4 Flash 0731 (batch) — $0.11/M, score 69.1 inclusionAI: Ling 3.0 Flash — $0.021/M, score 50.6 Qwen: Qwen3.7 Plus — $0.32/M, score 55.9 SpaceXAI: Grok 4.3 (batch) — $1/M, score 42.2 Nemotron 3 Ultra 550B A55B — $0.5/M, score 49.3 Anthropic: Claude Sonnet 4 — $3/M, score 37.6 OpenAI: GPT-5.6 Luna (batch) — $0.1/M, score 71.4 OpenAI: GPT-5.4 Mini (batch) — $0.375/M, score 56.1 GPT-5.4 nano — $0.2/M, score 56.1 Qwen: Qwen3.8 2.4T A95B — $2/M, score 71.9 OpenAI: GPT-5.6 Sol (batch) — $1/M, score 77.4 Qwen: Qwen3.6 Plus — $0.325/M, score 54.5 Anthropic: Claude Haiku 4.5 (batch) — $0.5/M, score 43.9 Z.ai: GLM 4.6 — $0.43/M, score 45.8 Anthropic: Claude Sonnet 4.5 (batch) — $1.5/M, score 52.1 GPT-4o (2024-05-13) — $5/M, score 24.2 OpenAI: GPT-5 (batch) — $0.625/M, score 37.8 Qwen: Qwen3 Next 80B A3B Thinking — $0.15/M, score 17.4 OpenAI: GPT-4 Turbo (batch) — $5/M, score 21.5 Nemotron 3.5 Lightning 30B A3B — $0.05/M, score 26.8 Hy3 preview — $0.066/M, score 58.8 GPT-4.1 nano — $0.1/M, score 11.1 GPT OSS 20B — $0.02/M, score 20.7 GPT-4o mini — $0.15/M, score 11.4 o1 — $15/M, score 39.7 Kimi K2.5 — $0.3/M, score 46.8 Trinity Large Thinking — $0.25/M, score 25.8 Qwen: Qwen3 30B A3B Thinking 2507 — $0.2/M, score 12.1 Z.ai: GLM 5.2 — $0.966/M, score 68.8 Qwen: Qwen3.7 Max — $1.475/M, score 66.0 SpaceXAI: Grok Build 0.1 — $1/M, score 51.5 Qwen: Qwen3.6 35B A3B — $0.1/M, score 41.9 Qwen: Qwen3.6 27B — $0.3/M, score 53.7 inclusionAI: Ling-2.6-flash — $0.01/M, score 25.3 Mistral: Mistral Small 4 — $0.15/M, score 26.6 Kwaipilot: KAT-Coder-Pro V2 — $0.3/M, score 59.5 Qwen: Qwen3.5-9B — $0.1/M, score 28.7 Qwen: Qwen3.5-35B-A3B — $0.312/M, score 37.0 Qwen: Qwen3.5-122B-A10B — $0.26/M, score 45.7 Upstage: Solar Pro 3 — $0.15/M, score 16.2 Z.ai: GLM 4.7 — $0.4/M, score 45.3 Amazon: Nova 2 Lite — $0.3/M, score 23.0 Mistral: Ministral 3 3B 2512 — $0.1/M, score 4.8 Google: Gemma 3n 4B — $0.06/M, score 3.2 Meta: Llama 4 Maverick — $0.2/M, score 16.3 OpenAI: o3 Mini High — $1.1/M, score 16.3 Meta: Llama 3.3 70B Instruct — $0.1/M, score 11.9 Meta: Llama 3.1 8B Instruct — $0.05/M, score 5.4 Nex AGI: Nex-N2-Pro — $0.25/M, score 59.1 Mistral: Mistral Medium 3.5 — $1.5/M, score 46.9 inclusionAI: Ring-2.6-1T — $0.075/M, score 42.8 IBM: Granite 4.1 8B — $0.05/M, score 9.5 Qwen: Qwen3.5 397B A17B — $0.55/M, score 48.2 Qwen: Qwen3 Coder Next — $0.12/M, score 36.2 Mistral: Ministral 3 14B 2512 — $0.2/M, score 14.4 Anthropic: Claude Haiku 4.5 — $1/M, score 43.9 Qwen: Qwen3 235B A22B Thinking 2507 — $0.23/M, score 22.1 Qwen: Qwen3 8B — $0.117/M, score 9.0 Qwen: Qwen3 14B — $0.227/M, score 13.8 Qwen: Qwen3 32B — $0.08/M, score 15.3 Meta: Llama 4 Scout — $0.1/M, score 8.2 DeepSeek: DeepSeek V3 0324 — $0.25/M, score 21.2 Cohere: Command A — $2.5/M, score 27.8 Step 3.7 Flash — $0.185/M, score 39.6 Gemma 4 31B IT Input price per million tokens (log scale) Index

The stepped line is the efficient frontier: at each price, the best score available for that money or less. A model sitting on it is not being beaten by anything cheaper. Price is log-scaled because listed rates span four orders of magnitude. Only models with both a listed price and a score on this benchmark can appear.

Catalog activity

Change log

Field-level changes detected between successful source imports.

Full change log
Max Output Tokens16384 → 32768
Price Completion0.33999999999999997 → 0.34
Max Output Tokens32768 → 16384
Price Completion0.34 → 0.33999999999999997
Max Output Tokens16384 → 32768
Price Completion0.33999999999999997 → 0.34
Max Output Tokens32768 → 16384
Price Completion0.34 → 0.33999999999999997
Max Output Tokens16384 → 32768
Price Completion0.33999999999999997 → 0.34
Provenance

Sources & verification

Every figure on this page traces back to one of these records.

Methodology
Public API

Use this record

Fetch the complete source-linked model record. No key, no account, no rate-limited tier.

API documentation
Endpoint
GET https://model.kyssta.lol/api/v1/models/google/gemma-4-31b-it
curl
curl "https://model.kyssta.lol/api/v1/models/google/gemma-4-31b-it"
Common questions

Frequently asked questions

Answered directly from the stored record — nothing here is generated beyond the catalog's own fields.

What is Gemma 4 31B IT?

Largest Gemma 4 instruction model for open, self-hosted chat and reasoning. It is published by Google and catalogued here from Models.dev.

How much does Gemma 4 31B IT cost?

Listed input pricing starts at $0.09 per million tokens from OpenRouter, rising to $0.99 across 26 listed providers.

What is the context length of Gemma 4 31B IT?

Gemma 4 31B IT accepts up to 262.144K tokens of context and returns up to 32.768K output tokens.

Does Gemma 4 31B IT support tool calling and structured output?

Provider catalogs list support for tool calling, structured output, reasoning, and image input.

Which providers serve Gemma 4 31B IT?

33 providers list this model: Cloudflare AI Gateway, QVAC, Requesty, Kenari, Nvidia, Google and 27 more.

Are the weights for Gemma 4 31B IT open?

Yes. The weights are published and downloadable from Hugging Face.

When was Gemma 4 31B IT released?

The catalog records a release date of 2026-04-02, last verified Sep 11, 2026.

More models from Google

View all →