Compact GPT model for low-latency assistance and high-volume workloads
Models
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Compact GPT model for low-latency assistance and high-volume workloads
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of Arena-Hard-Auto, where LLMs evaluate each other’s responses. It is a fine-tuned base model...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...
No provider description is available for this model yet.
No provider description is available for this model yet.
GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...
No provider description is available for this model yet.
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
This model always redirects to the latest GLM model from Z.ai.
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
No provider description is available for this model yet.
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
No provider description is available for this model yet.
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 | $30 | 2023-11-06 | |||
| GPT-4openai/gpt-4 | 8.192K | $30 | $60 | 2023-11-06 | |||
| GPT-3.5-turboopenai/gpt-3.5-turbo | 16.385K | $0.5 | $1.5 | 2023-03-01 | |||
| glm-5.3-flashzai/glm-5.3-flash | 1.04858M | $0.15 | $0.5 | — | |||
| meta-models/Muse-Glimmer-30Bdeepinfra/meta-models/muse-glimmer-30b | 131.072K | $0.3 | $1.2 | — | |||
| zai-org/glm-4.7-hnovita/zai-org/glm-4.7-h | 204.8K | $0.6 | $2.2 | — | |||
| zai-org/GLM-5.3together_ai/zai-org/glm-5.3 | 1.04858M | $1.4 | $4.4 | — | |||
| thinkingmachines/Inkling-Smalldeepinfra/thinkingmachines/inkling-small | 524.288K | $0.45 | $1.2 | — | |||
| moonshotai/kimi-k2.5novita/moonshotai/kimi-k2.5 | 262.144K | $0.6 | $3 | — | |||
| qwen/qwen3.7-maxnovita/qwen/qwen3.7-max | 1M | $1.25 | $3.75 | — | |||
| kimi-k2.7-codemoonshot/kimi-k2.7-code | 262.144K | $0.95 | $4 | — | |||
| gemma-4-31b-itgemini/gemma-4-31b-it | 262.144K | — | — | — | |||
| databricks-gemini-3-flashdatabricks/databricks-gemini-3-flash | 1.04858M | $0.625 | $3.75 | — | |||
| gemma-4-26b-a4b-itgemini/gemma-4-26b-a4b-it | 262.144K | — | — | — | |||
| databricks-glm-5-3-flashdatabricks/databricks-glm-5-3-flash | 1.04858M | — | — | — | |||
| databricks-gemini-3-1-prodatabricks/databricks-gemini-3-1-pro | 1.04858M | $2.5 | $15 | — | |||
| xiaomimimo/mimo-v2.5novita/xiaomimimo/mimo-v2.5 | 1.04858M | $0.168 | $0.336 | — | |||
| AionLabs: Aion-RP 1.0 (8B)aion-labs/aion-rp-llama-3.1-8b | 32.768K | $0.8 | $1.6 | — | |||
| google/gemma-4-31B-it-turbodeepinfra/google/gemma-4-31b-it-turbo | 262.144K | $0.09 | $0.34 | — | |||
| Qwen/Qwen3-Maxdeepinfra/qwen/qwen3-max | 256K | $1.2 | $6 | — | |||
| deepseek/deepseek-ocr-2novita/deepseek/deepseek-ocr-2 | 8.192K | $0.03 | $0.03 | — | |||
| XiaomiMiMo/MiMo-V2.5deepinfra/xiaomimimo/mimo-v2.5 | 262.144K | $0.4 | $2 | — | |||
| qwen-coderqwencloud/qwen-coder | 1M | $0.3 | $1.5 | — | |||
| Qwen: Qwen3 235B A22Bqwen/qwen3-235b-a22b | 131.072K | $0.455 | $1.82 | — | |||
| qwen/qwen3-coder-nextnovita/qwen/qwen3-coder-next | 262.144K | $0.2 | $1.5 | — | |||
| kimi-k2.7-codeqwencloud/kimi-k2.7-code | 229.376K | $0.95 | $4 | — | |||
| OpenAI: GPT-5.1 Chatopenai/gpt-5.1-chat | 128K | $1.25 | $10 | — | |||
| glm-5.2qwencloud/glm-5.2 | 1.04858M | $1.4 | $4.4 | — | |||
| Meta: Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor | 1.04858M | $0.1 | $0.2 | — | |||
| google/gemini-3.5-flashdeepinfra/google/gemini-3.5-flash | 1M | $1.5 | $9 | — | |||
| zai-org/glm-5novita/zai-org/glm-5 | 202.8K | $1 | $3.2 | — | |||
| qwen-flashqwencloud/qwen-flash | 997.952K | — | — | — | |||
| Z.ai: GLM Latest~z-ai/glm-latest | 1.04858M | $0.873 | $3.36 | — | |||
| Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash | 131.072K | $0.061 | $0.4 | — | |||
| qwen-flash-2025-07-28qwencloud/qwen-flash-2025-07-28 | 997.952K | — | — | — | |||
| qwen-maxqwencloud/qwen-max | 30.72K | $1.6 | $6.4 | — | |||
| qwen-plusqwencloud/qwen-plus | 129.024K | $0.4 | $1.2 | — | |||
| qwen-plus-2025-01-25qwencloud/qwen-plus-2025-01-25 | 129.024K | $0.4 | $1.2 | — | |||
| qwen-plus-2025-04-28qwencloud/qwen-plus-2025-04-28 | 129.024K | $0.4 | $1.2 | — | |||
| qwen-plus-2025-07-14qwencloud/qwen-plus-2025-07-14 | 129.024K | $0.4 | $1.2 | — | |||
| qwen-plus-2025-07-28qwencloud/qwen-plus-2025-07-28 | 997.952K | — | — | — | |||
| qwen-plus-2025-09-11qwencloud/qwen-plus-2025-09-11 | 997.952K | — | — | — | |||
| qwen-plus-latestqwencloud/qwen-plus-latest | 997.952K | — | — | — | |||
| qwen-turboqwencloud/qwen-turbo | 129.024K | $0.05 | $0.2 | — | |||
| qwen-turbo-2024-11-01qwencloud/qwen-turbo-2024-11-01 | 1M | $0.05 | $0.2 | — | |||
| qwen-turbo-2025-04-28qwencloud/qwen-turbo-2025-04-28 | 1M | $0.05 | $0.2 | — | |||
| qwen-turbo-latestqwencloud/qwen-turbo-latest | 1M | $0.05 | $0.2 | — | |||
| qwen3-30b-a3bqwencloud/qwen3-30b-a3b | 129.024K | — | — | — | |||
| Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch | 1.04858M | $0.375 | $1.875 | — | |||
| glm-5.1qwencloud/glm-5.1 | 202.745K | $1.4 | $4.4 | — |