4,182 models

No provider description is available for this model yet.

baseten/moonshotai/kimi-k2-thinking Not documented context $0.6/M input $2.5/M output

No provider description is available for this model yet.

novita/minimax/minimax-m2.7 204.8K context $0.3/M input $1.2/M output
Open weights

No provider description is available for this model yet.

meta-llama/CodeLlama-7b-Instruct-hf Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

meta-llama/CodeLlama-7b-Python-hf Not documented context Input not listed Output not listed
Open weights

No provider description is available for this model yet.

meta-llama/CodeLlama-7b-hf Not documented context Input not listed Output not listed

No provider description is available for this model yet.

baseten/openai/gpt-oss-120b Not documented context $0.1/M input $0.5/M output

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

google/gemini-2.5-pro:batch 1.04858M context $0.625/M input $5/M output

No provider description is available for this model yet.

openai/chatgpt-4o-latest 128K context $5/M input $15/M output

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

openai/gpt-5.6-luna:batch 1.05M context $0.1/M input $0.6/M output
Open weights

No provider description is available for this model yet.

Qwen/Qwen3.8-27B Not documented context Input not listed Output not listed

No provider description is available for this model yet.

novita/xiaomimimo/mimo-v2.5 1.04858M context $0.168/M input $0.336/M output

No provider description is available for this model yet.

nebius/nousresearch/hermes-4-70b 131.072K context $0.13/M input $0.4/M output

No provider description is available for this model yet.

cerebras/zai-glm-4.6 128K context $2.25/M input $2.75/M output

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

anthropic/claude-sonnet-4 200K context $3/M input $15/M output

No provider description is available for this model yet.

nebius/nousresearch/hermes-4-405b 131.072K context $1/M input $3/M output

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

meta/muse-spark-1.3-contributor 1.04858M context $0.1/M input $0.2/M output

No provider description is available for this model yet.

baseten/deepseek-ai/deepseek-v3.1 Not documented context $0.5/M input $1.5/M output

No provider description is available for this model yet.

baseten/deepseek-ai/deepseek-v3-0324 Not documented context $0.77/M input $0.77/M output

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...

openai/gpt-audio-mini 128K context $0.6/M input $2.4/M output

No provider description is available for this model yet.

gmi/zai-org/glm-4.7-fp8 202.752K context $0.4/M input $2/M output

No provider description is available for this model yet.

azure/gpt-audio-mini 128K context $0.6/M input $2.4/M output

No provider description is available for this model yet.

nebius/moonshotai/kimi-k3 1.024M context $3/M input $15/M output
Open weights

No provider description is available for this model yet.

google/gemma-4-12B Not documented context Input not listed Output not listed

No provider description is available for this model yet.

snowflake/llama3.2-1b 128K context Input not listed Output not listed

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

inclusionai/ling-3.0-flash-fin:free 262.144K context Free input Free output

No provider description is available for this model yet.

nebius/moonshotai/kimi-k2.7-code 262.144K context $0.95/M input $4/M output

No provider description is available for this model yet.

qwencloud/qwen-flash 997.952K context Input not listed Output not listed