Research radar
arXiv preprints, filtered for developer impact.
Concise summaries of new AI papers, ranked for how likely they are to change how production systems are built. arXiv collection remains rate-limited.
Papers tracked
2,287
matching records
Shown
10
current page
Top radar
86
http://arxiv.org/abs/2608.26086v1
Code links
2
on this page
| Paper | Authors | arXiv ID | Categories | Published | Code | Radar | Summary |
|---|---|---|---|---|---|---|---|
| TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development | Jiarui Yan, Weiwei Sun, Sijie Li | http://arxiv.org/abs/2608.26086v1 | cs.LG, cs.AI | Aug 26, 2026 | Detected | RDR86 | Large language models write correct code for isolated problems but remain far weaker at auton... |
| WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution | Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng | http://arxiv.org/abs/2608.27454v1 | cs.AI, cs.CL | Aug 27, 2026 | None | RDR75 | Agent skills package specialized knowledge and workflows into reusable resources that extend... |
| SWE-Prime: Fewer Trajectories, Better Performance | Dewu Zheng, Ruizhe Ye, Yanlin Wang | http://arxiv.org/abs/2608.27449v1 | cs.SE, cs.AI | Aug 27, 2026 | None | RDR81 | To improve large language models' ability to resolve real-world software issues, prior work h... |
| From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench | Dewu Zheng, Yanlin Wang, Xiwen Wang | http://arxiv.org/abs/2608.27442v1 | cs.SE, cs.AI | Aug 27, 2026 | None | RDR81 | In real-world software development, code review typically involves iterative interactions bet... |
| Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners | Qianlong Lan, Vinothini Pandurangan, Anuj Kaul | http://arxiv.org/abs/2608.27424v1 | cs.CR, cs.AI | Aug 27, 2026 | None | RDR82 | Static scanners are increasingly used to identify executable or otherwise unsafe content in m... |
| Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms | Siye Wu, Kai Yang, Yuchen Cai | http://arxiv.org/abs/2608.27409v1 | cs.CL | Aug 27, 2026 | None | RDR80 | Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large... |
| CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases | Sil Hamilton, Albert Yu Sun, Oscar J. Romero | http://arxiv.org/abs/2608.27391v1 | cs.AI, cs.CL | Aug 27, 2026 | None | RDR87 | LLMs are increasingly able to answer complex questions about enterprise-scale document collec... |
| D2C-Routing: Dimension-to-Composition Evidence Routing for Mixed-Origin AI-Generated Text Detection | Xin Chen, Fuwei Zhang, Yiqi Tong | http://arxiv.org/abs/2608.27380v1 | cs.CL | Aug 27, 2026 | Detected | RDR86 | AI-generated text detection is commonly framed as a binary document-level judgment about whet... |
| VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning | Junxiang Xu, Ruisi Wang, Fanyi Pu | http://arxiv.org/abs/2608.26105v1 | cs.CV, cs.AI | Aug 26, 2026 | None | RDR86 | Native visual reasoning treats visual generation as the medium of reasoning itself: visual st... |
| TTPO: Test-Time Policy Optimization | Aozhe Wang, Zhengxi Lu, Jianze Wang | http://arxiv.org/abs/2608.27448v1 | cs.CL | Aug 27, 2026 | None | RDR74 | Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Sel... |