Open multimodal Llama model for image understanding, captioning, and visual QA
3 providers
Every model in the catalog with source-linked pricing, context limits, provider availability, and published benchmark results.
Open multimodal Llama model for image understanding, captioning, and visual QA
| Model | Creator | Inputs | Context | Input | Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Llama-3.2-11B-Vision-Instructmeta/llama-3.2-11b-vision-instruct | 128K | $0.055 | $0.055 | 2024-09-25 |