Data as of Aug 25, 2026 · Based on 3,181,687 AI responses across 10,525 prompts · See how Parse measures this
InternVL is a large-scale vision-language foundation model that scales the vision encoder to 6 billion parameters and aligns it with large language models for generic visual-linguistic tasks. The InternVL 2.5 series, ranging from 1B to 78B parameters, achieves over 70% on the MMMU benchmark, matching the performance of leading closed-source commercial models like GPT-4o.
Sources
kdnuggets.com shapes more of what AI says about InternVL than any other source, at 67% of its citations.
laonpeople.com
The market map
AI Creative Design Tools →Where AI ranks InternVL