3,424 models

No provider description is available for this model yet.

azure_ai/gpt-5.5 1.05M context $5/M input $30/M output

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

moonshotai/kimi-k2.7-code:batch 262.144K context $0.95/M input $4/M output

No provider description is available for this model yet.

azure_ai/gpt-5.4 1.05M context $2.5/M input $15/M output

No provider description is available for this model yet.

xai/grok-build-latest 500K context $2/M input $6/M output

No provider description is available for this model yet.

scaleway/glm-5.2 256K context $1.8/M input $5.5/M output

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses more compute to think harder and provide...

openai/o1-pro:batch 200K context $75/M input $300/M output

No provider description is available for this model yet.

azure_ai/kimi-k2.7-code 262.144K context $0.95/M input $4/M output

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

openai/gpt-3.5-turbo:batch 16.385K context $0.25/M input $0.75/M output

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

qwen/qwen3-coder-plus 1M context $0.65/M input $3.25/M output

No provider description is available for this model yet.

azure_ai/fw-nemotron-3-ultra-nvfp4 262.144K context $0.6/M input $2.4/M output

No provider description is available for this model yet.

azure_ai/gpt-oss-120b 131.072K context $0.15/M input $0.6/M output

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...

meta-llama/llama-3.2-3b-instruct:free 131.072K context Free input Free output

No provider description is available for this model yet.

azure_ai/gpt-5.4-2026-03-05 1.05M context $2.5/M input $15/M output

No provider description is available for this model yet.

vertex_ai/gemini-3.8-flash 1.04858M context $0.75/M input $3.75/M output

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...

bytedance-seed/seed-2-1-turbo 262.144K context $0.5/M input $2.5/M output

No provider description is available for this model yet.

azure_ai/gpt-5.4-mini 272K context $0.75/M input $4.5/M output

No provider description is available for this model yet.

azure/computer-use-preview 8.192K context $3/M input $12/M output

No provider description is available for this model yet.

azure_ai/gpt-5.4-mini-2026-03-17 272K context $0.75/M input $4.5/M output

No provider description is available for this model yet.

azure_ai/gpt-5.4-nano 272K context $0.2/M input $1.25/M output

No provider description is available for this model yet.

azure_ai/gpt-5.4-nano-2026-03-17 272K context $0.2/M input $1.25/M output

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

qwen/qwen3.8-2.4t-a95b:batch 1.01M context $2/M input $6/M output

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

google/gemini-3.1-flash-lite:batch 1.04858M context $0.125/M input $0.75/M output