4 models
Ranked by NL2Repo
Reasoning Tools 55.9

Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows

alibaba/qwen3.8-max-preview 2026-07-19 1M context $2/M input $6/M output
6 providers
Open weights 48.2

Large coding-reasoning model for agentic software tasks and RL search

deepreinforce/ornith-1.0-397b 2026-06-25 262.144K context Input not listed Output not listed
Open weights 34.6

Large coding-reasoning model for agentic software tasks and RL search

deepreinforce/ornith-1.0-35b 2026-06-25 262.144K context Input not listed Output not listed
Open weights 27.2

Open coding-reasoning model for repository tasks and self-improving agents

deepreinforce/ornith-1.0-9b 2026-06-25 262.144K context Input not listed Output not listed