Fast Gemini workhorse for multimodal apps where latency and price matter
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Google's proven reasoning model for coding, math, and multimodal analysis
Long-lived GPT workhorse for coding, instruction following, and production apps
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Affordable GPT-4.1 lane for fast coding help and structured extraction
Open multimodal Llama for strong reasoning with efficient everyday serving
Open Llama with long-context vision for efficient multimodal agents
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen instruction model for multilingual chat, reasoning, and tool use
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 | $2.5 | 2025-06-17 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 | $0.4 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 | $10 | 2025-06-17 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 | $8 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 | $0.4 | 2025-04-14 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 | $1.6 | 2025-04-14 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.14 | $0.59 | 2025-04-05 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 3.5M | $0.17 | $0.66 | 2025-04-05 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 | $0.2 | 2024-11-01 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 | $1.2 | 2024-01-25 |