Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...
Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
This model always redirects to the latest model in the Google Gemini Pro family.
This model always redirects to the latest model in the Google Gemini Flash family.
Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio...
The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market. It routes you based on what the OpenRouter community collectively spends on...
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Thinking Machines: Inkling (free)thinkingmachines/inkling:free | 1.04858M | Free | Free | — | |||
| Google: Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | 1.04858M | $1.25 | $10 | — | |||
| Google: Gemini 3.7 Flash (batch)google/gemini-3.7-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| OpenAI: GPT Audioopenai/gpt-audio | 128K | $2.5 | $10 | — | |||
| Auto Router (Beta)openrouter/auto-beta | 2M | — | — | — | |||
| NVIDIA: Nemotron 3 Nano Omni (free)nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 256K | Free | Free | — | |||
| Google Gemini Pro Latest~google/gemini-pro-latest | 1.04858M | $2 | $12 | — | |||
| Google Gemini Flash Latest~google/gemini-flash-latest | 1.04858M | $0.75 | $3.75 | — | |||
| Mistral: Voxtral Small 24B 2507mistralai/voxtral-small-24b-2507 | 32.768K | $0.1 | $0.3 | — | |||
| Auto Routeropenrouter/auto | 2M | — | — | — | |||
| Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1.04858M | $0.25 | $1.5 | — | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 1.04858M | $0.625 | $5 | — | |||
| Meta: Muse Spark 1.2 Contributormeta/muse-spark-1.2-contributor | 1.04858M | $0.1 | $0.2 | — | |||
| Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch | 1.04858M | $0.75 | $4.5 | — | |||
| OpenAI: GPT Audio Miniopenai/gpt-audio-mini | 128K | $0.6 | $2.4 | — | |||
| Meta: Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor | 1.04858M | $0.1 | $0.2 | — | |||
| Google: Gemini 2.5 Flash Lite (batch)google/gemini-2.5-flash-lite:batch | 1.04858M | $0.05 | $0.2 | — | |||
| Thinking Machines: Inkling Small (batch)thinkingmachines/inkling-small:batch | 524.288K | $0.5 | $1.2 | — | |||
| Google: Gemini 3.1 Pro Preview (batch)google/gemini-3.1-pro-preview:batch | 1.04858M | $1 | $6 | — | |||
| Thinking Machines: Inkling Small (free)thinkingmachines/inkling-small:free | 1.04858M | Free | Free | — | |||
| Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | 1.04858M | $0.15 | $1.25 | — | |||
| Google: Gemini 2.5 Pro Preview 05-06google/gemini-2.5-pro-preview-05-06 | 1.04858M | $1.25 | $10 | — | |||
| Thinking Machines: Inkling (batch)thinkingmachines/inkling:batch | 524.288K | $1 | $4.05 | — | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 1.04858M | $0.15 | $1.25 | — |