LIVE-Last scan updating-53 sources active-693 signals today-OTHERCowAgent: Open-Source AI Assistant and Agent Harness
Automated best-of radar

Best Vision-language radar.

Automated radar for multimodal, image understanding, document vision, OCR, and VLM developer options. Rankings use existing source-backed radar data and do not invent prices, benchmark scores, release dates, or capabilities.

Candidates
48
matched automatically
Top fit
82
NVIDIA NIM Model Catalog
Sources
57
canonical links retained
Access
All
all access paths
Updated
Aug 6, 2026
listed in sitemap
All48Free/open6Paid API3Open source6
Best overall radar pick

NVIDIA NIM Model Catalog

Highest automated fit score using source-backed task evidence and radar signals.

RDR82serviceNVIDIA
Best API or hosted option

NVIDIA NIM Model Catalog

Strong candidate when the source metadata indicates API or hosted-service availability.

RDR82serviceNVIDIA
Best open-source option

NVIDIA NIM Model Catalog

Strong candidate when open weights, GitHub, permissive license, or self-hosting signals are visible.

RDR82serviceNVIDIA
Best developer momentum

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Strong candidate with fresh GitHub, Hub, Product Hunt, or radar momentum.

RDR77paperarxiv-ai
Best research signal

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Strong paper or research-linked candidate when a source-backed research signal exists.

RDR77paperarxiv-ai
#CandidateTypeAccessSourceFitUpdatedWhy it matchedEvidence
01NVIDIA NIM Model CatalogNVIDIA inference platformserviceFree endpointOfficial inference catalogRDR82Jun 26, 2026Matched vision-language, vision language, multimodal; 3 source links; official inference catalog signal; access model: Free endpointbuild.nvidia.com
02Hugging Face Inference ProvidersHugging Face model hubservicePaid APIOfficial inference catalogRDR79Jun 26, 2026Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid APIhuggingface.co
03Fireworks AI Serverless ModelsFireworks AI inference platformservicePaid APIOfficial inference catalogRDR78Jun 26, 2026Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid APIdocs.fireworks.ai
04Together AI Serverless ModelsTogether AI inference platformservicePaid APIOfficial inference catalogRDR78Jun 26, 2026Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid APIdocs.together.ai
05ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMshttp://arxiv.org/abs/2608.04010v1paperResearch-onlyarxiv-aiRDR77Aug 4, 2026Matched vision-language, vision language, multimodal; 2 source links; access model: Research-only; freshly updatedarxiv.org
06ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagramshttp://arxiv.org/abs/2607.24707v1paperResearch-onlyarxiv-aiRDR76Jul 27, 2026Matched vision-language, vision language, multimodal; 2 source links; access model: Research-onlyarxiv.org
07ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understandinghttp://arxiv.org/abs/2607.24743v1paperResearch-onlyarxiv-aiRDR75Jul 27, 2026Matched multimodal, image understanding, visual question answering; 1 source link; access model: Research-onlyarxiv.org
08Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMshttp://arxiv.org/abs/2607.27122v1paperResearch-onlyarxiv-aiRDR72Jul 29, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Research-onlyarxiv.org
09Qualcomm AI Hub ModelsQualcomm model zooserviceDownloadable pretrainedOfficial model zooRDR72Jun 26, 2026Matched vision-language, vision language, multimodal; 1 source link; official model zoo signal; access model: Downloadable pretrainedaihub.qualcomm.com
10Antfly: A Distributed Search Engine for Multimodal AI DataRAG & SearcharticlePricing not verifiedAI on Radar articleRDR70Jun 5, 2026Matched vision-language, vision language, multimodal; 1 source link; access model: Pricing not verifiedgithub.com
11OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Modelshttp://arxiv.org/abs/2607.28609v1paperResearch-onlyarxiv-aiRDR70Jul 30, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Research-onlyarxiv.org
12ReToken: One Token to Improve Vision-Language Models for Visual Retrievalhttp://arxiv.org/abs/2607.28627v1paperResearch-onlyarxiv-aiRDR69Jul 30, 2026Matched vision-language, vision language; 2 source links; access model: Research-onlyarxiv.org
13FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Samplinghttp://arxiv.org/abs/2607.29596v1paperResearch-onlyarxiv-aiRDR69Jul 31, 2026Matched vision-language, vision language, multimodal; 1 source link; access model: Research-onlyarxiv.org
14VLMs May Not Globally Enhance Human Alignment over LLMs During Natural ReadingResearch PapersarticlePricing not verifiedAI on Radar articleRDR69May 28, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verifiedarxiv.org
15HumanCLAW: Can Vision-Language Models Act Through a Body?http://arxiv.org/abs/2607.27180v1paperResearch-onlyarxiv-aiRDR68Jul 29, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Research-onlyarxiv.org
16KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainabilityhttp://arxiv.org/abs/2607.24730v1paperResearch-onlyarxiv-aiRDR68Jul 27, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Research-onlyarxiv.org
17ContextRL: Reinforcement Learning for Improved LLM Reasoning and Multimodal PerformanceResearch PapersarticlePricing not verifiedAI on Radar articleRDR67Jun 16, 2026Matched multimodal, image understanding, visual question answering; 1 source link; access model: Pricing not verifiedarxiv.org
18TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAMhttp://arxiv.org/abs/2607.27205v1paperResearch-onlyarxiv-aiRDR67Jul 29, 2026Matched vision-language, vision language; 2 source links; access model: Research-onlyarxiv.org
19NEO-ov: A Native One-Vision Foundation Model for End-to-End Spatiotemporal ModelingResearch PapersarticlePricing not verifiedAI on Radar articleRDR67May 28, 2026Matched vision-language, vision language, multimodal; 1 source link; access model: Pricing not verifiedarxiv.org
20Data Pyramid for Embodied Manipulationhttp://arxiv.org/abs/2607.24744v1paperResearch-onlyarxiv-aiRDR66Jul 27, 2026Matched vision-language, vision language, multimodal; 1 source link; access model: Research-onlyarxiv.org
21Evidence Attribution in Visual Document Understanding without Coordinates or Region Labelshttp://arxiv.org/abs/2607.24651v1paperResearch-onlyarxiv-aiRDR66Jul 27, 2026Matched vision-language, vision language, multimodal; 1 source link; access model: Research-onlyarxiv.org
22VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screeninghttp://arxiv.org/abs/2607.26042v1paperResearch-onlyarxiv-aiRDR65Jul 28, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Research-onlyarxiv.org
23PAR3D: A Unified 3D-MLLM for Part-Aware Scene UnderstandingResearch PapersarticlePricing not verifiedAI on Radar articleRDR65Jun 5, 2026Matched vision-language, vision language, multimodal; 1 source link; access model: Pricing not verifiedarxiv.org
24ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programshttp://arxiv.org/abs/2607.28538v1paperResearch-onlyarxiv-aiRDR65Jul 30, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Research-onlyarxiv.org
25ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Enginehttp://arxiv.org/abs/2607.28625v1paperResearch-onlyarxiv-aiRDR65Jul 30, 2026Matched vision-language, vision language; 1 source link; access model: Research-onlyarxiv.org
26MemoryVLA++: Enhancing Vision-Language-Action Models with Temporal Memory and Imagination for RoboticsRoboticsarticlePricing not verifiedAI on Radar articleRDR64Jun 9, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verifiedarxiv.org
27P-ai: A Self-Growing Desktop AI Assistant for Automation and Long-Running TasksAI ToolsarticlePricing not verifiedAI on Radar articleRDR64May 26, 2026Matched image-to-text, image to text; 1 source link; access model: Pricing not verifiedgithub.com
28InSight: Self-Guided Skill Acquisition via Steerable VLAsRoboticsarticlePricing not verifiedAI on Radar articleRDR63Jun 24, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verifiedarxiv.org
29PerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionBenchmarksarticlePricing not verifiedAI on Radar articleRDR63Jun 29, 2026Matched multimodal, visual question answering; 1 source link; access model: Pricing not verifiedarxiv.org
30WCM: A World Critic Model for Vision-Language-Action Reinforcement Learninghttp://arxiv.org/abs/2607.29613v1paperResearch-onlyarxiv-aiRDR63Jul 31, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Research-onlyarxiv.org
31MonoIR-RS: A New Benchmark for Infrared Remote Sensing Vision-Language UnderstandingResearch PapersarticlePricing not verifiedAI on Radar articleRDR63Jul 8, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verifiedarxiv.org
32LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box DecodingResearch PapersarticlePricing not verifiedAI on Radar articleRDR63May 27, 2026Matched vision-language, vision language; 1 source link; access model: Pricing not verifiedarxiv.org
33MAPS: A Novel Framework for Joint Vision-Language Geo-LocalizationResearch PapersarticlePricing not verifiedAI on Radar articleRDR63Jun 23, 2026Matched vision-language, vision language, multimodal; 1 source link; access model: Pricing not verifiedarxiv.org
34NewtPhys: A New Benchmark for Newtonian Physics Understanding in Foundation ModelsResearch PapersarticlePricing not verifiedAI on Radar articleRDR62Jun 3, 2026Matched vision-language, vision language; 1 source link; access model: Pricing not verifiedarxiv.org
35SOCO: A New Benchmark for Semantic Object Correspondence in Vision Foundation ModelsBenchmarksarticlePricing not verifiedAI on Radar articleRDR61Jun 1, 2026Matched vision-language, vision language; 1 source link; access model: Pricing not verifiedarxiv.org
36NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language ModelsResearch PapersarticlePricing not verifiedAI on Radar articleRDR61Jun 23, 2026Matched vision-language, vision language, vlm; 1 source link; access model: Pricing not verifiedarxiv.org
37$π\mathbf{R}^2$: Reactive Real-time Flow Policieshttp://arxiv.org/abs/2607.26055v1paperResearch-onlyarxiv-aiRDR60Jul 28, 2026Matched vision-language, vision language; 1 source link; access model: Research-onlyarxiv.org
38Anatomy Contextualized Adaption of CT Foundation Modelshttp://arxiv.org/abs/2607.27154v1paperResearch-onlyarxiv-aiRDR60Jul 29, 2026Matched vision-language, vision language; 1 source link; access model: Research-onlyarxiv.org
39DLAM: Distributional Latent Actions with Temporal Constraintshttp://arxiv.org/abs/2607.27138v1paperResearch-onlyarxiv-aiRDR60Jul 29, 2026Matched vision-language, vision language; 1 source link; access model: Research-onlyarxiv.org
40cero2k6/gideal-rag-v1cero2k6 modelmodelOpen weightsHugging Face ModelsRDR59Aug 6, 2026Matched image-to-text, image to text; 2 source links; access model: Open weights; freshly updatedhuggingface.co
41HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answeringhttp://arxiv.org/abs/2607.29638v1paperResearch-onlyarxiv-aiRDR59Jul 31, 2026Matched ocr, visual question answering; 1 source link; access model: Research-onlyarxiv.org
42a3124371940/radeonvla_reflex_evaluation_videosa3124371940 modelmodelOpen weightsHugging Face ModelsRDR58Aug 5, 2026Matched vision-language, vision language; 2 source links; access model: Open weights; freshly updatedhuggingface.co
43jwidmer/trocr-hanse-test-inferencejwidmer modelmodelOpen weightsHugging Face ModelsRDR58Aug 6, 2026Matched image-to-text, image to text; 2 source links; access model: Open weights; freshly updatedhuggingface.co
44RoboTTT: Scaling Robot Policy Context to 8K TimestepsRoboticsarticlePricing not verifiedAI on Radar articleRDR58Jul 17, 2026Matched vision-language, vision language; 1 source link; access model: Pricing not verifiedarxiv.org
45PolicyTrim: Enhancing Vision-Language-Action Model EfficiencyResearch PapersarticlePricing not verifiedAI on Radar articleRDR58Jun 23, 2026Matched vision-language, vision language; 1 source link; access model: Pricing not verifiedarxiv.org
46Staged Executable Inverse Graphics (SEIG) with Vision-Language Models in BlenderResearch PapersarticlePricing not verifiedAI on Radar articleRDR57Jun 2, 2026Matched vision-language, vision language; 1 source link; access model: Pricing not verifiedarxiv.org
47CARA: Concept-Aware Risk Attention for Interpretable Collision Anticipationhttp://arxiv.org/abs/2607.22494v1paperResearch-onlyarxiv-aiRDR57Jul 24, 2026Matched vision-language, vision language; 1 source link; access model: Research-onlyarxiv.org
48IT-Help-San-Diego/calibration-scopeRust repositoryrepoOpen sourceGitHubRDR57Jul 29, 2026Matched vision-language, vision language; 1 source link; access model: Open source; open weights signalgithub.com
Custom alerts

Track Vision-language changes

Get private alerts when source-backed vision-language candidates, comparisons, or access signals change. No manual content, no invented claims.

API and bulk access
Topics
Choose segments and get a private RSS feed plus preference link.