1. Research PapersScore84

    Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

    A new study questions the reliability of compile rate as a metric for evaluating LLM-based C/C++ vulnerability repair. Experiments show compile rate is heavily influenced by evaluation artifacts and can be misleading, even rewarding non-repairs. The research proposes a change-aware metric, diff_F1, as a more suitable preliminary screen.

    The findings are based on an empirical study published on arXiv by Om Nepal, Sushant Aryal, Oluseyi Olukola, and Nick Rahimi, titled "Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen". Full analysis
  2. BenchmarksScore81

    OpenAI Introduces MentalHealthBench for Evaluating AI in Mental Health Conversations

    OpenAI has introduced MentalHealthBench, a new benchmark designed to assess the safety and helpfulness of AI responses in simulated mental health dialogues. This benchmark was developed with input from experts in the field.

    Source: OpenAI Full analysis
  3. Developer ToolsScore82

    Share an App Window on Windows

    OpenAI has updated its developer tools to allow sharing an application window on Windows. This feature aims to enhance the capabilities of applications built with OpenAI's technologies.

    Source: OpenAI Full analysis