4,213 models

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

z-ai/glm-5v-turbo 202.752K context $1.2/M input $4/M output

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

z-ai/glm-4.6 198K context $0.43/M input $1.75/M output

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

anthropic/claude-sonnet-4.5:batch 1M context $1.5/M input $7.5/M output

No provider description is available for this model yet.

bedrock/eu-north-1/deepseek.v3.2 163.84K context $0.74/M input $2.22/M output

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

mistralai/mistral-medium-3.1:batch 131.072K context $0.2/M input $1/M output

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

google/gemini-2.5-flash:batch 1.04858M context $0.15/M input $1.25/M output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro-preview-05-06 1.04858M context $1.25/M input $10/M output

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

google/gemini-3-flash-preview:batch 1.04858M context $0.25/M input $1.5/M output

No provider description is available for this model yet.

novita/qwen/qwen3.5-122b-a10b 262.144K context $0.4/M input $3.2/M output

No provider description is available for this model yet.

novita/qwen/qwen3.5-27b 262.144K context $0.3/M input $2.4/M output

No provider description is available for this model yet.

cloudflare/@cf/zai-org/glm-5.2 262.144K context $1.4/M input $4.4/M output

No provider description is available for this model yet.

openrouter/openai/gpt-3.5-turbo 16.385K context $1.5/M input $2/M output

LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model optimized for agentic tasks, data extraction, and RAG—while still running comfortably on edge devices. It supports long context (up to 32K tokens) and is...

liquid/lfm-2.5-1.2b-thinking:free 32.768K context Free input Free output

No provider description is available for this model yet.

novita/minimax/minimax-m2.7 204.8K context $0.3/M input $1.2/M output

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

inception/mercury-2.5 260K context $0.04/M input $0.15/M output

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

z-ai/glm-5.3 1.04858M context $1.4/M input $4.4/M output

LFM2.5-1.2B-Instruct is a compact, high-performance instruction-tuned model built for fast on-device AI. It delivers strong chat quality in a 1.2B parameter footprint, with efficient edge inference and broad runtime support.

liquid/lfm-2.5-1.2b-instruct:free 32.768K context Free input Free output

No provider description is available for this model yet.

openrouter/moonshotai/kimi-k2.5 262.144K context $0.6/M input $3/M output