LIVE-Last scan updating-53 sources active-864 signals today-BENCHMARKSPiloting the world's first double-blind AI evaluations
Source-linked AI news and decision tools

Track the AI changes that matter to builders.

AI on Radar monitors official releases, model and API changes, GitHub momentum, papers, and developer tools, then publishes concise reporting with visible sources and practical calculators.

01
LLM API selectorFilter source-backed model routes by workload, context, capabilities, and budget.
02
API cost calculatorEstimate monthly input, output, cache-read, and per-request charges.
03
GPU / VRAM calculatorEstimate model-weight and optional architecture-aware KV-cache memory.

Latest on the radar

5 signalsView all articles
FeaturedBenchmarksRDR81

Piloting the world's first double-blind AI evaluations

Google DeepMind is pioneering the first double-blind evaluations for AI systems. This approach aims to mitigate biases in AI assessment by ensuring neither the evaluators nor the developers know which AI is being tested.

2 min - 5h agoAI evaluation
Research PapersRDR83

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Researchers have introduced VBVR-Pro, a closed-loop testbed designed to advance native visual reasoning. This suite aims to make visual reasoning trainable, verifiable, optimizable, and controllable by providing a scalable task space and reliable reward mechanisms.

2 min - 5h agovisual reasoning
BenchmarksRDR82

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

A new benchmark, SWE Refactor Bench, has been introduced to evaluate the ability of coding agents to perform complex, whole-repository software migrations. Existing benchmarks are insufficient as they do not verify if the migration actually occurred, allowing agents to pass tests by copying original code. This benchmark addresses that gap by assessing both migration completeness and behavioral correctness.

2 min - 2d agocoding agents

Evergreen decision library

6 compare / 6 alternatives / 6 use cases / 4 changesOpen compare hub

Signal surfaces

Live across tracked layers

Topic intelligence

Graph

Cross-source topics rebuilt from clusters, entities, source mix, and article decisions.

  • New Gemma-2-9b-it Datasets for Math, Physics, and Chemistry on Hugging FaceResearch Papers - 8 linked signalsRDR61
  • penfever/nemotron-gym-competitive-coding-minimax-m27-131k-traces_chunk1 Dataset signal on Hugging FaceModels - 24 linked signalsRDR59
  • ​ Fine-tuning preview CLI commandSignals - 3 linked signalsRDR58
  • ​ Model revision validation statusSignals - 3 linked signalsRDR58
12 active topicsOpen

GitHub momentum

Top 25

Open-source attention measured by star velocity, freshness, and topic relevance, not raw stars.

  • santifer/career-opsJavaScript - 67,113 stars - +2779 7dRDR92
  • holaboss-ai/holaOSTypeScript - 9,611 stars - +2202 7dRDR87
  • opensandbox-group/OpenSandboxPython - 14,063 stars - +1570 7dRDR89
  • stablyai/orcaTypeScript - 50,371 stars - +2588 7dRDR89
8,847 repos trackedOpen

Model watch

Sourced

Releases and capability changes from Hugging Face, OpenRouter, Artificial Analysis, and source APIs.

  • orcarouter/spoken-multihop-ragorcarouter - context not listed - unknownRDR57
  • AtesiT/ru-llm-judge-datasetAtesiT - context not listed - unknownRDR60
  • aziz9788/saudi-llmaziz9788 - context not listed - unknownRDR58
  • rmems/agentic-coding-trajectoriesrmems - context not listed - unknownRDR60
6,087 models trackedOpen

Research queue

arXiv

Preprints summarized through a developer-impact lens with source links and conservative claims.

  • CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modeshttp://arxiv.org/abs/2608.27455v1 - Yufan Wu, Yinghui HeRDR83
  • Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090http://arxiv.org/abs/2608.27370v1 - Kairong Luo, Jiarui CuiRDR87
  • RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolutionhttp://arxiv.org/abs/2608.27439v1 - Junjie Zhang, Hui LiuRDR82
  • Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Informationhttp://arxiv.org/abs/2608.27417v1 - Chanho Park, Daehyeon ChoiRDR84
2,287 papers trackedOpen

Tool launches

Gated

Product and developer-tool launches with compliance-aware sourcing.

  • VMs won't contain cyber-capable agentshacker-news-ai - Community discussionRDR64
  • Z.ai confirms Ox Alpha is a new GLM-series model and will release its weightshacker-news-ai - Community discussionRDR64
  • CEO fired developers to make room for AI. Developers create open source AI CEOhacker-news-ai - Community discussionRDR64
  • Serve Markdown to AI Agents with Accept Headershacker-news-ai - Community discussionRDR62
120 launches 7dOpen
Get the AI builder radar

Custom alerts and RSS for the segments you actually track.

Pick agents, AI coding, models, research, GitHub radar, providers, and a minimum score. Feeds use published, source-linked radar items only.

Topics
Choose segments and get a private RSS feed plus preference link.
How the radar works

Transparent scoring. No invented numbers.

53 active sources

Official blogs, GitHub, arXiv, Hugging Face, OpenRouter, RSS feeds, and public APIs with source links kept visible.

367 scans in 24h

Scan counts and signal totals come from the live radar, not from static homepage counters.

Radar score, 0-100

Scores combine reliability, freshness, novelty, technical importance, developer relevance, ecosystem signal, and confidence.

993 raw signals in 24h

Low-confidence items stay off public pages until their source support is strong enough.