LFM2.5-1.2B-Instruct is a compact, high-performance instruction-tuned model built for fast on-device AI. It delivers strong chat quality in a 1.2B parameter footprint, with efficient edge inference and broad runtime support.

liquid/lfm-2.5-1.2b-instruct:free 32.768K context Free input Free output

No provider description is available for this model yet.

snowflake/llama3.2-1b 128K context Input not listed Output not listed

No provider description is available for this model yet.

cerebras/zai-glm-4.6 128K context $2.25/M input $2.75/M output

No provider description is available for this model yet.

novita/kwaipilot/kat-coder-pro 256K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

lambda_ai/llama3.1-8b-instruct 131.072K context $0.025/M input $0.04/M output

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

openai/o3-mini-high:batch 200K context $0.55/M input $2.2/M output

No provider description is available for this model yet.

azure_ai/mistral-large-latest 128K context $2/M input $6/M output

No provider description is available for this model yet.

novita/zai-org/glm-4.6 204.8K context $0.55/M input $2.2/M output

No provider description is available for this model yet.

novita/zai-org/glm-4.6v 131.072K context $0.3/M input $0.9/M output

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

qwen/qwen3-next-80b-a3b-instruct:free 262.144K context Free input Free output

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

inclusionai/ling-3.0-flash 262.144K context $0.021/M input $0.063/M output

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

openai/gpt-5.6-sol:batch 1.05M context $1/M input $5/M output

No provider description is available for this model yet.

lambda_ai/llama3.1-70b-instruct-fp8 131.072K context $0.12/M input $0.3/M output

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

google/gemma-4-31b-it:batch 262.144K context $0.39/M input $0.97/M output

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

x-ai/grok-4.20-multi-agent 2M context $1.25/M input $2.5/M output

No provider description is available for this model yet.

ollama/llama2:70b 4.096K context Input not listed Output not listed

No provider description is available for this model yet.

bedrock/sa-east-1/deepseek.v3.2 163.84K context $0.74/M input $2.22/M output

No provider description is available for this model yet.

ollama/llama2:7b 4.096K context Input not listed Output not listed

No provider description is available for this model yet.

ollama/llama3 8.192K context Input not listed Output not listed

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,...

openai/gpt-5.2-pro:batch 400K context $10.5/M input $84/M output

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

anthropic/claude-haiku-4.5:batch 200K context $0.5/M input $2.5/M output

No provider description is available for this model yet.

friendliai/google/gemma-4-31b-it 262.144K context $0.14/M input $0.4/M output

No provider description is available for this model yet.

friendliai/zai-org/glm-5.2 1.04858M context $1.4/M input $4.4/M output

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5:batch 1M context $1.5/M input $7.5/M output

Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with...

stealth/ox-alpha 1.04858M context Input not listed Output not listed

No provider description is available for this model yet.

novita/deepseek/deepseek-v3.2-exp 163.84K context $0.27/M input $0.41/M output