Model comparisonSo sánh model模型对比
meta/llama-3.2-11b-vision-instruct Details | |
|---|---|
| Overview | |
| Description | The Llama 3.2-Vision instruction-tuned models are optimized for visual recognition, image reasoning, captioning, and answering general questions about an image. |
| Category | Text Generation |
| Context length | 0 |
| Providers | 1 |
| Pricing (per 1M tokens) | |
| Input | $0.05 |
| Output | $0.68 |
| Capabilities | |
| Reasoning | – |
| Tool calling | |
| Vision | |
| Streaming | |
| Token activity (30 days) | |
| Usage |