Best Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs alternatives.
Live source-backed alternatives to Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs for Vision-language. Alternatives are selected from the same task category and update whenever the best-of index rebuilds.
Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions. We propose a simple yet effective domain-generalized training framework built on two practical strategies. First, we introduce a balanced minibatch sampling scheme that strategically samples tampered and real images in each minibatch, preventing biased optimization toward either manipulated artifacts or clean-image priors and avoiding training collapse, ensuring that each optimization step receives proper sampled gradient signals. Second, we adopt a simple late-injection strategy, where the detector is first trained on large-scale base data until stable convergence, and then exposed to a small amount of newly selected supporting data from emerging VLM distributions, improving adaptability without overfitting to limited new domains. Together, these components provide a simple yet strong recipe for improving pixel-level tampering localization and OOD robustness across modern VLMs. Despite the conceptual simplicity, our framework outperforms the prior state-of-the-art PIXAR by a large margin of 26.1% and 26.8% relative improvement in average gIoU and cIoU, respectively, across OOD VLMs of GPT-Images-2.0, Gemini-3.1, FLUX.2, and Seedream 4.5. Our code is available at https://github.com/VILA-Lab/PIXAR-DG cs.CV Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions. We propose a simple yet effective domain-generalized training framework built on two practical strategies. First, we introduce a balanced minibatch sampling scheme that strategically samples tampered and real images in each minibatch, preventing biased optimization toward either manipulated artifacts or clean-image priors and avoiding training collapse, ensuring that each optimization step receives proper sampled gradient signals. Second, we adopt a simple late-injection strategy, where the detector is first trained on large-scale base data until stable convergence, and then exposed to a small amount of newly selected supporting data from emerging VLM distributions, improving adaptability without overfitting to limited new domains. Together, these components provide a simple yet strong recipe for improving pixel-level tampering localization and OOD robustness across modern VLMs. Despite the conceptual simplicity, our framework outperforms the prior state-of-the-art PIXAR by a large margin of 26.1% and 26.8% relative improvement in average gIoU and cIoU, respectively, across OOD VLMs of GPT-Images-2.0, Gemini-3.1, FLUX.2, and Seedream 4.5. Our code is available at https://github.com/VILA-Lab/PIXAR-DG Research signal collected from arXiv metadata; Gemini enrichment can add a clearer summary. cs.CV cs.AI
NVIDIA NIM Model Catalog
Matched vision-language, vision language, multimodal; 3 source links; official inference catalog signal; access model: Free endpoint
Hugging Face Inference Providers
Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid API
| # | Alternative | Kind | Access | Fit | Why it appears | Source |
|---|---|---|---|---|---|---|
| 01 | NVIDIA NIM Model Catalog | service | Free endpoint | RDR83 | Matched vision-language, vision language, multimodal; 3 source links; official inference catalog signal; access model: Free endpoint | build.nvidia.com |
| 02 | Hugging Face Inference Providers | service | Paid API | RDR80 | Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid API | huggingface.co |
| 03 | Fireworks AI Serverless Models | service | Paid API | RDR79 | Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid API | docs.fireworks.ai |
| 04 | Together AI Serverless Models | service | Paid API | RDR79 | Matched vision-language, vision language, multimodal; 2 source links; official inference catalog signal; access model: Paid API | docs.together.ai |
| 05 | AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding | paper | Research-only | RDR76 | Matched vision-language, vision language, multimodal; 1 source link; access model: Research-only | arxiv.org |
| 06 | Symbal: Detecting Systematic Misalignments in Model-Generated Captions | paper | Research-only | RDR76 | Matched vision-language, vision language, multimodal; 2 source links; access model: Research-only; freshly updated | arxiv.org |
| 07 | Qualcomm AI Hub Models | service | Downloadable pretrained | RDR73 | Matched vision-language, vision language, multimodal; 1 source link; official model zoo signal; access model: Downloadable pretrained | aihub.qualcomm.com |
Track Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs alternatives
Get private alerts when source-backed vision-language alternatives, access signals, or comparison evidence change.
API and bulk access