- Research PapersScore84
Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen
A new study questions the reliability of compile rate as a metric for evaluating LLM-based C/C++ vulnerability repair. Experiments show compile rate is heavily influenced by evaluation artifacts and can be misleading, even rewarding non-repairs. The research proposes a change-aware metric, diff_F1, as a more suitable preliminary screen.
- BenchmarksScore81
OpenAI Introduces MentalHealthBench for Evaluating AI in Mental Health Conversations
OpenAI has introduced MentalHealthBench, a new benchmark designed to assess the safety and helpfulness of AI responses in simulated mental health dialogues. This benchmark was developed with input from experts in the field.
- Developer ToolsScore82
Share an App Window on Windows
OpenAI has updated its developer tools to allow sharing an application window on Windows. This feature aims to enhance the capabilities of applications built with OpenAI's technologies.
AI on Radar Digest - Sep 24, 2026
3 AI signals selected from today's radar.