4,213 models

LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model optimized for agentic tasks, data extraction, and RAG—while still running comfortably on edge devices. It supports long context (up to 32K tokens) and is...

liquid/lfm-2.5-1.2b-thinking:free 32.768K context Free input Free output

No provider description is available for this model yet.

novita/qwen/qwen2.5-7b-instruct 32K context $0.07/M input $0.07/M output

No provider description is available for this model yet.

cloudflare/@cf/qwen/qwq-32b 24K context $0.66/M input $1/M output

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

inception/mercury-2.5 260K context $0.04/M input $0.15/M output

LFM2.5-1.2B-Instruct is a compact, high-performance instruction-tuned model built for fast on-device AI. It delivers strong chat quality in a 1.2B parameter footprint, with efficient edge inference and broad runtime support.

liquid/lfm-2.5-1.2b-instruct:free 32.768K context Free input Free output

No provider description is available for this model yet.

snowflake/llama3.2-1b 128K context Input not listed Output not listed

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5:batch 1M context $1.5/M input $7.5/M output

No provider description is available for this model yet.

cerebras/zai-glm-4.6 128K context $2.25/M input $2.75/M output

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...

qwen/qwen3-max 262.144K context $0.78/M input $3.9/M output

No provider description is available for this model yet.

novita/qwen/qwen3-4b-fp8 128K context $0.03/M input $0.03/M output

No provider description is available for this model yet.

novita/qwen/qwen3-8b-fp8 128K context $0.035/M input $0.138/M output

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

openai/gpt-5.6-luna:batch 1.05M context $0.1/M input $0.6/M output

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

qwen/qwen3-next-80b-a3b-instruct:free 262.144K context Free input Free output

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

mistralai/codestral-2508 256K context $0.3/M input $0.9/M output

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

qwen/qwen3-coder 262.144K context $0.3/M input $1/M output

No provider description is available for this model yet.

mistral/magistral-medium-2509 40K context $2/M input $5/M output

No provider description is available for this model yet.

novita/baidu/ernie-4.5-21b-a3b 120K context $0.07/M input $0.28/M output

No provider description is available for this model yet.

novita/baidu/ernie-4.5-vl-28b-a3b 30K context $0.14/M input $0.56/M output

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

mistralai/mistral-large-2512:batch 262.144K context $0.25/M input $0.75/M output

No provider description is available for this model yet.

ollama/llama2:70b 4.096K context Input not listed Output not listed

No provider description is available for this model yet.

bedrock/sa-east-1/deepseek.v3.2 163.84K context $0.74/M input $2.22/M output

No provider description is available for this model yet.

ollama/llama2:7b 4.096K context Input not listed Output not listed

No provider description is available for this model yet.

ollama/llama3 8.192K context Input not listed Output not listed

No provider description is available for this model yet.

mistral/magistral-medium-2506 40K context $2/M input $5/M output

No provider description is available for this model yet.

friendliai/google/gemma-4-31b-it 262.144K context $0.14/M input $0.4/M output

No provider description is available for this model yet.

friendliai/zai-org/glm-5.2 1.04858M context $1.4/M input $4.4/M output