Reasoning Tools JSON

Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows

google/gemini-3.8-flash 2026-09-02 1.04858M context $0.75/M input $3.75/M output
17 providers
Reasoning Tools JSON

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It improves long-horizon agent collaboration, instruction following, and coding efficiency relative to Muse Spark 1.2.

meta/muse-spark-1.3 2026-09-02 1.04858M context $1.25/M input $4.25/M output
9 providers

Speech transcription model for accurate audio-to-text and captioning workflows

google/gemini-3.5-transcribe-live 2026-08-26 Not documented context Input not listed Output not listed
1 provider
Reasoning Tools JSON

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning

google/gemini-3.7-flash 2026-08-13 1.04858M context $0.75/M input $3.75/M output
23 providers
Reasoning Tools JSON

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning

google/gemini-flash-latest 2026-08-13 1.04858M context $0.75/M input $3.75/M output
7 providers
Reasoning Tools JSON

Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1 with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows.

meta/muse-spark-1.2 2026-08-05 1.04858M context $1.25/M input $4.25/M output
14 providers
Reasoning Tools JSON

Fast Gemini model balancing multimodal reasoning, tool use, and cost

google/gemini-flash-lite-latest 2026-07-21 1.04858M context $0.3/M input $2.5/M output
6 providers
Reasoning Tools JSON

Fast Gemini model balancing multimodal reasoning, tool use, and cost

google/gemini-3.5-flash-lite 2026-07-21 1.04858M context $0.3/M input $2.5/M output
23 providers
Reasoning Tools JSON

Fast Gemini model balancing multimodal reasoning, tool use, and cost

google/gemini-3.6-flash 2026-07-21 1.04858M context $0.75/M input $3.75/M output
24 providers
Reasoning

Realtime speech-to-speech model with configurable reasoning, tool use, and robust voice-agent behavior

openai/gpt-realtime-2.1 2026-07-06 128K context $4/M input $24/M output
2 providers

Low-latency audio-to-audio model for real-time speech translation across 70+ languages

google/gemini-3.5-live-translate-preview 2026-06-09 131.072K context $3.5/M input $21/M output
1 provider
Reasoning Tools JSON

Fast Gemini model balancing multimodal reasoning, tool use, and cost

google/gemini-3.5-flash 2026-05-19 1.04858M context $1.5/M input $9/M output
31 providers
Reasoning Tools JSON

Low-latency Gemini model for high-volume multimodal and agent workloads

google/gemini-3.1-flash-lite 2026-05-07 1.04858M context $0.25/M input $1.5/M output
23 providers

Streaming speech-to-text model for low-latency transcript deltas from live audio

openai/gpt-realtime-whisper 2026-05-07 Not documented context Input not listed Output not listed
1 provider

Multimodal embedding model mapping text, images, video, audio, and PDFs into a unified embedding space

google/gemini-embedding-2 2026-04-22 8.192K context $0.2/M input Output not listed
3 providers

Maximum-comprehensiveness agentic researcher for multi-step investigation, synthesis, and cited reports

google/deep-research-max-preview-04-2026 2026-04-21 1.04858M context Input not listed Output not listed
1 provider

Agentic model for autonomous multi-step research, synthesis, and cited reports

google/deep-research-preview-04-2026 2026-04-21 1.04858M context Input not listed Output not listed
1 provider

Low-latency speech generation with steerable prompts and expressive audio tags

google/gemini-3.1-flash-tts-preview 2026-04-15 8.192K context $1/M input $20/M output
2 providers
Reasoning Tools JSON

Vision-language model for embodied reasoning: spatial understanding, task planning, and physical-world agentic robotics

google/gemini-robotics-er-1.6-preview 2026-04-14 131.072K context $1/M input $5/M output
2 providers
Reasoning Tools

High-quality, low-latency Live API model for real-time dialogue and voice-first AI applications

google/gemini-3.1-flash-live-preview 2026-03-26 131.072K context $0.75/M input $4.5/M output
1 provider

Music generation model for short 30-second clips, loops, and previews from text or image prompts

google/lyria-3-clip-preview 2026-03-25 131.072K context Input not listed Output not listed
4 providers

Music generation model for full-length songs from text or images with vocals and structure

google/lyria-3-pro-preview 2026-03-25 131.072K context Input not listed Output not listed
3 providers
Reasoning Tools

MiMo omni model for text, image, video, audio, and agents

xiaomi/mimo-v2-omni 2026-03-18 262.144K context $0.14/M input $0.28/M output
4 providers
Reasoning Tools JSON

Low-latency Gemini model for high-volume multimodal and agent workloads

google/gemini-3.1-flash-lite-preview 2026-03-03 1.04858M context $0.25/M input $1.5/M output
13 providers
Reasoning Tools JSON

Advanced Gemini model for complex reasoning, coding, and multimodal analysis

google/gemini-3.1-pro-preview-customtools 2026-02-19 1.04858M context $2/M input $12/M output
11 providers
Reasoning Tools JSON

Reasoning-first Gemini preview for agentic coding and complex problem solving

google/gemini-3.1-pro-preview 2026-02-19 1.04858M context $2/M input $12/M output
31 providers
Reasoning Tools JSON

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

google/gemini-3-flash-preview 2025-12-17 1.04858M context $0.5/M input $3/M output
25 providers
Reasoning Tools JSON

Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts

google/gemini-3-pro-preview 2025-11-18 1.04858M context $0.57/M input $3.43/M output
10 providers

Speech generation model for controllable voice, narration, and audio delivery

google/gemini-2.5-flash-tts 2025-09-30 32.768K context $0.5/M input $10/M output
1 provider

Speech generation model for controllable voice, narration, and audio delivery

google/gemini-2.5-pro-tts 2025-09-30 32.768K context $1/M input $20/M output
1 provider
Reasoning Tools JSON

Fast Gemini workhorse for multimodal apps where latency and price matter

google/gemini-2.5-flash 2025-06-17 1.04858M context $0.3/M input $2.5/M output
30 providers
Reasoning Tools JSON

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents

google/gemini-2.5-flash-lite 2025-06-17 1.04858M context $0.1/M input $0.4/M output
18 providers
Reasoning Tools JSON

Google's proven reasoning model for coding, math, and multimodal analysis

google/gemini-2.5-pro 2025-06-17 1.04858M context $1.25/M input $10/M output
29 providers

Qwen omni model for text, vision, audio, and multimodal agent tasks

alibaba/qwen-omni-turbo 2025-01-19 32.768K context $0.07/M input $0.27/M output
4 providers
Tools

Earlier Gemini Flash workhorse for responsive multimodal apps and tool use

google/gemini-2.0-flash 2024-12-11 1.04858M context $0.1/M input $0.42/M output
2 providers
Tools

Low-latency Gemini model for high-volume multimodal and agent workloads

google/gemini-2.0-flash-lite 2024-12-11 1.04858M context $0.052/M input $0.21/M output
2 providers

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...

meta/muse-spark-1.2-contributor 1.04858M context $0.1/M input $0.2/M output

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling:free 1.04858M context Free input Free output

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

google/gemini-3.6-flash:batch 1.04858M context $0.375/M input $1.875/M output

The Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market. It routes you based on what the OpenRouter community collectively spends on...

openrouter/auto 2M context Input not listed Output not listed

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio...

mistralai/voxtral-small-24b-2507 32.768K context $0.1/M input $0.3/M output

This model always redirects to the latest model in the Google Gemini Flash family.

~google/gemini-flash-latest 1.04858M context $0.75/M input $3.75/M output

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

google/gemini-3.5-flash:batch 1.04858M context $0.75/M input $4.5/M output

This model always redirects to the latest model in the Google Gemini Pro family.

~google/gemini-pro-latest 1.04858M context $2/M input $12/M output

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...

openai/gpt-audio-mini 128K context $0.6/M input $2.4/M output

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free 256K context Free input Free output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro-preview 1.04858M context $1.25/M input $10/M output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro:batch 1.04858M context $0.625/M input $5/M output

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

google/gemini-2.5-flash-lite:batch 1.04858M context $0.05/M input $0.2/M output

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...

openai/gpt-audio 128K context $2.5/M input $10/M output