Research radar
arXiv preprints, filtered for developer impact.
Concise summaries of new AI papers, ranked for how likely they are to change how production systems are built. arXiv collection remains rate-limited.
Papers tracked
2,287
matching records
Shown
10
current page
Top radar
76
http://arxiv.org/abs/2608.24848v1
Code links
0
on this page
| Paper | Authors | arXiv ID | Categories | Published | Code | Radar | Summary |
|---|---|---|---|---|---|---|---|
| BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes | Fei Tang, Huawen Shen, Zhiqiong Lu | http://arxiv.org/abs/2608.24848v1 | cs.CL | Aug 25, 2026 | None | RDR76 | Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading... |
| StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments | Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav | http://arxiv.org/abs/2608.24804v1 | cs.AI, cs.SE | Aug 25, 2026 | None | RDR77 | We present StarHarness, a framework for evolving environment-specific agent harnesses while k... |
| CAFE: Self-Improving Search Agents Need Co-Evolving Feedback | Boyang Liu, Senjie Jin, Peixin Wang | http://arxiv.org/abs/2608.24794v1 | cs.AI | Aug 25, 2026 | None | RDR78 | Outcome-supervised search agents learn when and how to retrieve evidence, but terminal reward... |
| Image Difference Quantification Using Autoencoder-Based Latent Representations | Manish Sharma, Timothy Yim, Clifton Forlines | http://arxiv.org/abs/2608.24782v1 | cs.CV | Aug 25, 2026 | None | RDR84 | Traditional image similarity metrics such as Mean Squared Error (MSE), Peak Signal-to-Noise R... |
| Stochastic Estimation of Transduced Language Models | Vésteinn Snæbjarnarson, Samuel Kiegeland, Manuel de Prada Corral | http://arxiv.org/abs/2608.27428v1 | cs.CL | Aug 27, 2026 | None | RDR73 | Transduced language models (TLMs) compose a pretrained \emph{source} language model with a fu... |
| ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing | Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer | http://arxiv.org/abs/2608.26083v1 | cs.LG, cs.AI | Aug 26, 2026 | None | RDR80 | Deep neural networks often exploit spurious associations in their training data, a failure kn... |
| Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study | Leonardo Duart, Tiago Fonseca, Thiago Chacón | http://arxiv.org/abs/2608.26060v1 | cs.CL, stat.ML | Aug 26, 2026 | None | RDR79 | Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recen... |
| Beyond Local Surprise: Grounded Dialogue as Selective Belief Revision under Referential Uncertainty | Ziming Liu, Bhanu Chaitanya Jasti, Ziyang Xu | http://arxiv.org/abs/2608.26035v1 | cs.CL | Aug 26, 2026 | None | RDR74 | When a speaker refers to a scene that the listener cannot directly see, the listener must dec... |
| Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning | Sixiang Chen, Jiaming Liu, Jixian Wu | http://arxiv.org/abs/2608.24885v1 | cs.RO, cs.CV | Aug 25, 2026 | None | RDR79 | Action-conditioned world models are increasingly used as learned simulators for policy evalua... |
| What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation | Hao Chen | http://arxiv.org/abs/2608.24881v1 | stat.ML, cs.LG | Aug 25, 2026 | None | RDR83 | Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inceptio... |