4,213 models

No provider description is available for this model yet.

azure_ai/grok-4-fast-non-reasoning 131.072K context $0.2/M input $0.5/M output

No provider description is available for this model yet.

azure_ai/grok-4-fast-reasoning 131.072K context $0.2/M input $0.5/M output

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

google/gemini-3-flash-preview:batch 1.04858M context $0.25/M input $1.5/M output

No provider description is available for this model yet.

nvidia/MiniMax-M2.7-DFlash Not documented context Input not listed Output not listed

No provider description is available for this model yet.

azure_ai/grok-code-fast-1 131.072K context $0.2/M input $1.5/M output

No provider description is available for this model yet.

azure_ai/jais-30b-chat 8.192K context $3200/M input $9710/M output

No provider description is available for this model yet.

azure_ai/jamba-instruct 70K context $0.5/M input $0.7/M output

No provider description is available for this model yet.

azure_ai/kimi-k2.6 262.144K context $0.95/M input $4/M output

No provider description is available for this model yet.

azure_ai/ministral-3b 128K context $0.04/M input $0.04/M output

No provider description is available for this model yet.

azure_ai/mistral-large 32K context $4/M input $12/M output

No provider description is available for this model yet.

bedrock/us.twelvelabs.pegasus-1-2-v1:0 Not documented context Input not listed $7.5/M output

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...

openai/gpt-chat-latest 400K context $5/M input $30/M output

No provider description is available for this model yet.

azure_ai/mistral-large-latest 128K context $2/M input $6/M output

No provider description is available for this model yet.

azure_ai/mistral-large-3 256K context $0.5/M input $1.5/M output

No provider description is available for this model yet.

azure_ai/mistral-medium-2505 131.072K context $0.4/M input $2/M output

No provider description is available for this model yet.

azure_ai/mistral-nemo 131.072K context $0.15/M input $0.15/M output

No provider description is available for this model yet.

azure_ai/mistral-small 32K context $1/M input $3/M output

No provider description is available for this model yet.

nvidia/Kimi-K2.7-Code-DFlash Not documented context Input not listed Output not listed

No provider description is available for this model yet.

text-completion-openai/babbage-002 16.384K context $0.4/M input $0.4/M output

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5 1M context $3/M input $15/M output

No provider description is available for this model yet.

bedrock/twelvelabs.pegasus-1-2-v1:0 Not documented context Input not listed $7.5/M output

No provider description is available for this model yet.

nvidia/Nemotron-Labs-Audex-2B Not documented context Input not listed Output not listed

No provider description is available for this model yet.

azure/global-standard/gpt-4o-mini 128K context $0.15/M input $0.6/M output

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

qwen/qwen3-235b-a22b 131.072K context $0.455/M input $1.82/M output

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

inference-net/schematron-v2-turbo 128K context $0.03/M input $0.15/M output

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

inclusionai/ling-3.0-flash-fin:free 262.144K context Free input Free output

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

openai/gpt-4-turbo:batch 128K context $5/M input $15/M output

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...

moonshotai/kimi-k2 131.072K context $0.57/M input $2.3/M output

No provider description is available for this model yet.

bedrock/moonshotai.kimi-k2-thinking 262.144K context $0.73/M input $3.03/M output

No provider description is available for this model yet.

bedrock/moonshotai.kimi-k2.5 262.144K context $0.6/M input $3.03/M output