Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...

inclusionai/ling-3.0-flash-vl 131.072K context $0.06/M input $0.18/M output

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...

meta/muse-spark-1.2-contributor 1.04858M context $0.1/M input $0.2/M output

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

qwen/qwen3-235b-a22b 131.072K context $0.455/M input $1.82/M output

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...

moonshotai/kimi-k2 131.072K context $0.57/M input $2.3/M output

No provider description is available for this model yet.

cerebras/zai-glm-4.7 128K context $2.25/M input $2.75/M output

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and...

meta-llama/llama-3.2-11b-vision-instruct 131.072K context $0.345/M input $0.345/M output

No provider description is available for this model yet.

nlp_cloud/chatdolphin 16.384K context $0.5/M input $0.5/M output

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

nousresearch/hermes-3-llama-3.1-405b:free 131.072K context Free input Free output

No provider description is available for this model yet.

gmi/deepseek-ai/deepseek-v3-0324 163.84K context $0.28/M input $0.88/M output

No provider description is available for this model yet.

gmi/google/gemini-3-pro-preview 1.04858M context $2/M input $12/M output

No provider description is available for this model yet.

gmi/google/gemini-3-flash-preview 1.04858M context $0.5/M input $3/M output

No provider description is available for this model yet.

gmi/moonshotai/kimi-k2-thinking 262.144K context $0.8/M input $1.2/M output

No provider description is available for this model yet.

gmi/minimaxai/minimax-m2.1 196.608K context $0.3/M input $1.2/M output

No provider description is available for this model yet.

baseten/minimaxai/minimax-m2.5 Not documented context $0.3/M input $1.2/M output

No provider description is available for this model yet.

baseten/nvidia/nemotron-120b-a12b Not documented context $0.3/M input $0.75/M output

No provider description is available for this model yet.

baseten/zai-org/glm-5 Not documented context $0.95/M input $3.15/M output

No provider description is available for this model yet.

baseten/zai-org/glm-4.7 Not documented context $0.6/M input $2.2/M output

No provider description is available for this model yet.

baseten/zai-org/glm-4.6 Not documented context $0.6/M input $2.2/M output

No provider description is available for this model yet.

baseten/moonshotai/kimi-k2.5 Not documented context $0.6/M input $3/M output

No provider description is available for this model yet.

baseten/moonshotai/kimi-k2-thinking Not documented context $0.6/M input $2.5/M output

Coder‑Large is a 32 B‑parameter offspring of Qwen 2.5‑Instruct that has been further trained on permissively‑licensed GitHub, CodeSearchNet and synthetic bug‑fix corpora. It supports a 32k context window, enabling multi‑file...

arcee-ai/coder-large 32.768K context $0.5/M input $0.8/M output

No provider description is available for this model yet.

baseten/openai/gpt-oss-120b Not documented context $0.1/M input $0.5/M output

No provider description is available for this model yet.

openai/chatgpt-4o-latest 128K context $5/M input $15/M output