3,229 models

No provider description is available for this model yet.

anthropic/claude-4-opus-20250514 200K context $15/M input $75/M output

No provider description is available for this model yet.

bedrock/ap-south-1/deepseek.v3.2 163.84K context $0.74/M input $2.22/M output

No provider description is available for this model yet.

bedrock/moonshotai.kimi-k2.5 262.144K context $0.6/M input $3.03/M output

No provider description is available for this model yet.

azure_ai/gpt-5.4-mini-2026-03-17 272K context $0.75/M input $4.5/M output

No provider description is available for this model yet.

bedrock/moonshotai.kimi-k2-thinking 262.144K context $0.73/M input $3.03/M output

No provider description is available for this model yet.

azure_ai/gpt-5.4-mini 272K context $0.75/M input $4.5/M output

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...

bytedance-seed/seed-2-1-turbo 262.144K context $0.5/M input $2.5/M output

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

thinkingmachines/inkling:batch 524.288K context $1/M input $4.05/M output

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

google/gemini-3.6-flash:batch 1.04858M context $0.375/M input $1.875/M output

No provider description is available for this model yet.

azure_ai/gpt-5.4-2026-03-05 1.05M context $2.5/M input $15/M output

No provider description is available for this model yet.

cerebras/zai-glm-4.7 128K context $2.25/M input $2.75/M output

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...

meta-llama/llama-3.2-3b-instruct:free 131.072K context Free input Free output

No provider description is available for this model yet.

gemini/gemini-3.7-flash 1.04858M context $0.75/M input $3.75/M output

No provider description is available for this model yet.

vertex_ai/gemini-3.7-flash 1.04858M context $0.75/M input $3.75/M output

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

inference-net/schematron-v2-turbo 128K context $0.03/M input $0.15/M output

No provider description is available for this model yet.

azure_ai/mistral-small-2503 128K context $0.1/M input $0.3/M output

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

qwen/qwen3-coder 262.144K context $0.3/M input $1/M output