4,213 models

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

google/gemini-3.5-flash-lite:batch 1.04858M context $0.15/M input $1.25/M output

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...

openai/gpt-5.4-pro:batch 1.05M context $15/M input $90/M output

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

deepseek/deepseek-v4-flash-0731:batch 1.04858M context $0.11/M input $0.33/M output

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

nvidia/nemotron-3-ultra-550b-a55b:free 1M context Free input Free output

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

z-ai/glm-5.3-flash 1.04858M context $0.15/M input $0.5/M output

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

qwen/qwen3-235b-a22b-thinking-2507 131.072K context $0.23/M input $2.3/M output

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

openai/gpt-5:batch 400K context $0.625/M input $5/M output

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

anthropic/claude-sonnet-4 200K context $3/M input $15/M output

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,...

openai/gpt-5.2-pro:batch 400K context $10.5/M input $84/M output

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

google/gemini-3.5-flash:batch 1.04858M context $0.75/M input $4.5/M output

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

qwen/qwen3-coder-plus 1M context $0.65/M input $3.25/M output
Open weights

No provider description is available for this model yet.

Qwen/Qwen-Drive-1.0-4B Not documented context Input not listed Output not listed

Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o, Claude, and others into your files at...

relace/relace-apply-3 256K context $0.85/M input $1.25/M output

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...

openai/gpt-4o:batch 128K context $1.25/M input $5/M output
Open weights

No provider description is available for this model yet.

deepseek-ai/DeepSeek-V4.1-Flash Not documented context Input not listed Output not listed

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

inclusionai/ling-3.0-flash 262.144K context $0.021/M input $0.063/M output

No provider description is available for this model yet.

novita/moonshotai/kimi-k3 1.04858M context $3/M input $15/M output

No provider description is available for this model yet.

azure_ai/fw-glm-5.2-fast 1.04858M context $2.1/M input $6.6/M output

No provider description is available for this model yet.

azure_ai/fw-inkling 1.04858M context $1/M input $4.05/M output

No provider description is available for this model yet.

fireworks_ai/deepseek-v4-flash-0731 1.04858M context $0.14/M input $0.28/M output

No provider description is available for this model yet.

fireworks_ai/glm-5p2-fast 1.04858M context $2.1/M input $6.6/M output

No provider description is available for this model yet.

friendliai/zai-org/glm-5.1 202.752K context $1.4/M input $4.4/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k2.5 262.144K context $0.66/M input $3.3/M output

No provider description is available for this model yet.

fireworks_ai/glm-5p2-fast-us 1.04858M context $2.1/M input $6.6/M output

No provider description is available for this model yet.

fireworks_ai/kimi-k3 1.04858M context $3/M input $15/M output

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

cognitivecomputations/dolphin-mistral-24b-venice-edition:free 32.768K context Free input Free output

No provider description is available for this model yet.

fireworks_ai/kimi-k3-fast 1.04858M context $4.5/M input $22.5/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k2.6 262.144K context $1.045/M input $4.4/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k2.7-code 262.144K context $1.05/M input $4.4/M output

No provider description is available for this model yet.

fireworks_ai/kimi-k3-us 1.04858M context $3.3/M input $16.5/M output

No provider description is available for this model yet.

fireworks_ai/qwen3p8-max 262.144K context $2/M input $6/M output

No provider description is available for this model yet.

fireworks_ai/muse-glimmer-30b 131.072K context $0.35/M input $1.5/M output

No provider description is available for this model yet.

azure_ai/fw-kimi-k3 1.04858M context $3.3/M input $16.5/M output

No provider description is available for this model yet.

azure_ai/fw-minimax-m2.5 1M context $0.33/M input $1.32/M output

No provider description is available for this model yet.

azure_ai/fw-minimax-m3 512K context $0.33/M input $1.32/M output

No provider description is available for this model yet.

azure_ai/fw-nemotron-3-ultra-nvfp4 262.144K context $0.6/M input $2.4/M output