Best Vision-language radar.
Automated radar for multimodal, image understanding, document vision, OCR, and VLM developer options. Rankings use existing source-backed radar data and do not invent prices, benchmark scores, release dates, or capabilities.
NVIDIA NIM Model Catalog
Highest automated fit score using source-backed task evidence and radar signals.
NVIDIA NIM Model Catalog
Strong candidate when the source metadata indicates API or hosted-service availability.
NVIDIA NIM Model Catalog
Strong candidate when open weights, GitHub, permissive license, or self-hosting signals are visible.
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
Strong candidate with fresh GitHub, Hub, Product Hunt, or radar momentum.
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
Strong paper or research-linked candidate when a source-backed research signal exists.
| # | Candidate | Type | Access | Source | Fit | Updated | Why it matched | Evidence |
|---|---|---|---|---|---|---|---|---|
| 01 | NVIDIA NIM Model CatalogNVIDIA inference platform | service | Free endpoint | Official inference catalog | RDR82 | Jun 26, 2026 | Matched vision-language, vision language, multimodal; 3 source links; official inference catalog signal; access model: Free endpoint | build.nvidia.com |
| 02 | Hugging Face Inference ProvidersHugging Face model hub | service | Paid API | Official inference catalog | RDR79 | Jun 26, 2026 | Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid API | huggingface.co |
| 03 | Fireworks AI Serverless ModelsFireworks AI inference platform | service | Paid API | Official inference catalog | RDR78 | Jun 26, 2026 | Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid API | docs.fireworks.ai |
| 04 | Together AI Serverless ModelsTogether AI inference platform | service | Paid API | Official inference catalog | RDR78 | Jun 26, 2026 | Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid API | docs.together.ai |
| 05 | ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMshttp://arxiv.org/abs/2608.04010v1 | paper | Research-only | arxiv-ai | RDR77 | Aug 4, 2026 | Matched vision-language, vision language, multimodal; 2 source links; access model: Research-only; freshly updated | arxiv.org |
| 06 | ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagramshttp://arxiv.org/abs/2607.24707v1 | paper | Research-only | arxiv-ai | RDR76 | Jul 27, 2026 | Matched vision-language, vision language, multimodal; 2 source links; access model: Research-only | arxiv.org |
| 07 | ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understandinghttp://arxiv.org/abs/2607.24743v1 | paper | Research-only | arxiv-ai | RDR75 | Jul 27, 2026 | Matched multimodal, image understanding, visual question answering; 1 source link; access model: Research-only | arxiv.org |
| 08 | Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMshttp://arxiv.org/abs/2607.27122v1 | paper | Research-only | arxiv-ai | RDR72 | Jul 29, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Research-only | arxiv.org |
| 09 | Qualcomm AI Hub ModelsQualcomm model zoo | service | Downloadable pretrained | Official model zoo | RDR72 | Jun 26, 2026 | Matched vision-language, vision language, multimodal; 1 source link; official model zoo signal; access model: Downloadable pretrained | aihub.qualcomm.com |
| 10 | Antfly: A Distributed Search Engine for Multimodal AI DataRAG & Search | article | Pricing not verified | AI on Radar article | RDR70 | Jun 5, 2026 | Matched vision-language, vision language, multimodal; 1 source link; access model: Pricing not verified | github.com |
| 11 | OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Modelshttp://arxiv.org/abs/2607.28609v1 | paper | Research-only | arxiv-ai | RDR70 | Jul 30, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Research-only | arxiv.org |
| 12 | ReToken: One Token to Improve Vision-Language Models for Visual Retrievalhttp://arxiv.org/abs/2607.28627v1 | paper | Research-only | arxiv-ai | RDR69 | Jul 30, 2026 | Matched vision-language, vision language; 2 source links; access model: Research-only | arxiv.org |
| 13 | FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Samplinghttp://arxiv.org/abs/2607.29596v1 | paper | Research-only | arxiv-ai | RDR69 | Jul 31, 2026 | Matched vision-language, vision language, multimodal; 1 source link; access model: Research-only | arxiv.org |
| 14 | VLMs May Not Globally Enhance Human Alignment over LLMs During Natural ReadingResearch Papers | article | Pricing not verified | AI on Radar article | RDR69 | May 28, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verified | arxiv.org |
| 15 | HumanCLAW: Can Vision-Language Models Act Through a Body?http://arxiv.org/abs/2607.27180v1 | paper | Research-only | arxiv-ai | RDR68 | Jul 29, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Research-only | arxiv.org |
| 16 | KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainabilityhttp://arxiv.org/abs/2607.24730v1 | paper | Research-only | arxiv-ai | RDR68 | Jul 27, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Research-only | arxiv.org |
| 17 | ContextRL: Reinforcement Learning for Improved LLM Reasoning and Multimodal PerformanceResearch Papers | article | Pricing not verified | AI on Radar article | RDR67 | Jun 16, 2026 | Matched multimodal, image understanding, visual question answering; 1 source link; access model: Pricing not verified | arxiv.org |
| 18 | TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAMhttp://arxiv.org/abs/2607.27205v1 | paper | Research-only | arxiv-ai | RDR67 | Jul 29, 2026 | Matched vision-language, vision language; 2 source links; access model: Research-only | arxiv.org |
| 19 | NEO-ov: A Native One-Vision Foundation Model for End-to-End Spatiotemporal ModelingResearch Papers | article | Pricing not verified | AI on Radar article | RDR67 | May 28, 2026 | Matched vision-language, vision language, multimodal; 1 source link; access model: Pricing not verified | arxiv.org |
| 20 | Data Pyramid for Embodied Manipulationhttp://arxiv.org/abs/2607.24744v1 | paper | Research-only | arxiv-ai | RDR66 | Jul 27, 2026 | Matched vision-language, vision language, multimodal; 1 source link; access model: Research-only | arxiv.org |
| 21 | Evidence Attribution in Visual Document Understanding without Coordinates or Region Labelshttp://arxiv.org/abs/2607.24651v1 | paper | Research-only | arxiv-ai | RDR66 | Jul 27, 2026 | Matched vision-language, vision language, multimodal; 1 source link; access model: Research-only | arxiv.org |
| 22 | VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screeninghttp://arxiv.org/abs/2607.26042v1 | paper | Research-only | arxiv-ai | RDR65 | Jul 28, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Research-only | arxiv.org |
| 23 | PAR3D: A Unified 3D-MLLM for Part-Aware Scene UnderstandingResearch Papers | article | Pricing not verified | AI on Radar article | RDR65 | Jun 5, 2026 | Matched vision-language, vision language, multimodal; 1 source link; access model: Pricing not verified | arxiv.org |
| 24 | ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programshttp://arxiv.org/abs/2607.28538v1 | paper | Research-only | arxiv-ai | RDR65 | Jul 30, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Research-only | arxiv.org |
| 25 | ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Enginehttp://arxiv.org/abs/2607.28625v1 | paper | Research-only | arxiv-ai | RDR65 | Jul 30, 2026 | Matched vision-language, vision language; 1 source link; access model: Research-only | arxiv.org |
| 26 | MemoryVLA++: Enhancing Vision-Language-Action Models with Temporal Memory and Imagination for RoboticsRobotics | article | Pricing not verified | AI on Radar article | RDR64 | Jun 9, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verified | arxiv.org |
| 27 | P-ai: A Self-Growing Desktop AI Assistant for Automation and Long-Running TasksAI Tools | article | Pricing not verified | AI on Radar article | RDR64 | May 26, 2026 | Matched image-to-text, image to text; 1 source link; access model: Pricing not verified | github.com |
| 28 | InSight: Self-Guided Skill Acquisition via Steerable VLAsRobotics | article | Pricing not verified | AI on Radar article | RDR63 | Jun 24, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verified | arxiv.org |
| 29 | PerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionBenchmarks | article | Pricing not verified | AI on Radar article | RDR63 | Jun 29, 2026 | Matched multimodal, visual question answering; 1 source link; access model: Pricing not verified | arxiv.org |
| 30 | WCM: A World Critic Model for Vision-Language-Action Reinforcement Learninghttp://arxiv.org/abs/2607.29613v1 | paper | Research-only | arxiv-ai | RDR63 | Jul 31, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Research-only | arxiv.org |
| 31 | MonoIR-RS: A New Benchmark for Infrared Remote Sensing Vision-Language UnderstandingResearch Papers | article | Pricing not verified | AI on Radar article | RDR63 | Jul 8, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verified | arxiv.org |
| 32 | LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box DecodingResearch Papers | article | Pricing not verified | AI on Radar article | RDR63 | May 27, 2026 | Matched vision-language, vision language; 1 source link; access model: Pricing not verified | arxiv.org |
| 33 | MAPS: A Novel Framework for Joint Vision-Language Geo-LocalizationResearch Papers | article | Pricing not verified | AI on Radar article | RDR63 | Jun 23, 2026 | Matched vision-language, vision language, multimodal; 1 source link; access model: Pricing not verified | arxiv.org |
| 34 | NewtPhys: A New Benchmark for Newtonian Physics Understanding in Foundation ModelsResearch Papers | article | Pricing not verified | AI on Radar article | RDR62 | Jun 3, 2026 | Matched vision-language, vision language; 1 source link; access model: Pricing not verified | arxiv.org |
| 35 | SOCO: A New Benchmark for Semantic Object Correspondence in Vision Foundation ModelsBenchmarks | article | Pricing not verified | AI on Radar article | RDR61 | Jun 1, 2026 | Matched vision-language, vision language; 1 source link; access model: Pricing not verified | arxiv.org |
| 36 | NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language ModelsResearch Papers | article | Pricing not verified | AI on Radar article | RDR61 | Jun 23, 2026 | Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verified | arxiv.org |
| 37 | $π\mathbf{R}^2$: Reactive Real-time Flow Policieshttp://arxiv.org/abs/2607.26055v1 | paper | Research-only | arxiv-ai | RDR60 | Jul 28, 2026 | Matched vision-language, vision language; 1 source link; access model: Research-only | arxiv.org |
| 38 | Anatomy Contextualized Adaption of CT Foundation Modelshttp://arxiv.org/abs/2607.27154v1 | paper | Research-only | arxiv-ai | RDR60 | Jul 29, 2026 | Matched vision-language, vision language; 1 source link; access model: Research-only | arxiv.org |
| 39 | DLAM: Distributional Latent Actions with Temporal Constraintshttp://arxiv.org/abs/2607.27138v1 | paper | Research-only | arxiv-ai | RDR60 | Jul 29, 2026 | Matched vision-language, vision language; 1 source link; access model: Research-only | arxiv.org |
| 40 | cero2k6/gideal-rag-v1cero2k6 model | model | Open weights | Hugging Face Models | RDR59 | Aug 6, 2026 | Matched image-to-text, image to text; 2 source links; access model: Open weights; freshly updated | huggingface.co |
| 41 | HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answeringhttp://arxiv.org/abs/2607.29638v1 | paper | Research-only | arxiv-ai | RDR59 | Jul 31, 2026 | Matched ocr, visual question answering; 1 source link; access model: Research-only | arxiv.org |
| 42 | a3124371940/radeonvla_reflex_evaluation_videosa3124371940 model | model | Open weights | Hugging Face Models | RDR58 | Aug 5, 2026 | Matched vision-language, vision language; 2 source links; access model: Open weights; freshly updated | huggingface.co |
| 43 | jwidmer/trocr-hanse-test-inferencejwidmer model | model | Open weights | Hugging Face Models | RDR58 | Aug 6, 2026 | Matched image-to-text, image to text; 2 source links; access model: Open weights; freshly updated | huggingface.co |
| 44 | RoboTTT: Scaling Robot Policy Context to 8K TimestepsRobotics | article | Pricing not verified | AI on Radar article | RDR58 | Jul 17, 2026 | Matched vision-language, vision language; 1 source link; access model: Pricing not verified | arxiv.org |
| 45 | PolicyTrim: Enhancing Vision-Language-Action Model EfficiencyResearch Papers | article | Pricing not verified | AI on Radar article | RDR58 | Jun 23, 2026 | Matched vision-language, vision language; 1 source link; access model: Pricing not verified | arxiv.org |
| 46 | Staged Executable Inverse Graphics (SEIG) with Vision-Language Models in BlenderResearch Papers | article | Pricing not verified | AI on Radar article | RDR57 | Jun 2, 2026 | Matched vision-language, vision language; 1 source link; access model: Pricing not verified | arxiv.org |
| 47 | CARA: Concept-Aware Risk Attention for Interpretable Collision Anticipationhttp://arxiv.org/abs/2607.22494v1 | paper | Research-only | arxiv-ai | RDR57 | Jul 24, 2026 | Matched vision-language, vision language; 1 source link; access model: Research-only | arxiv.org |
| 48 | IT-Help-San-Diego/calibration-scopeRust repository | repo | Open source | GitHub | RDR57 | Jul 29, 2026 | Matched vision-language, vision language; 1 source link; access model: Open source; open weights signal | github.com |
Track Vision-language changes
Get private alerts when source-backed vision-language candidates, comparisons, or access signals change. No manual content, no invented claims.
API and bulk access