What changed
This research critically examines the effectiveness of compile rate as a primary metric for evaluating Large Language Models (LLMs) in automatically repairing C/C++ security vulnerabilities. The study, conducted by Om Nepal and colleagues, argues that compile rate—the percentage of generated patches that successfully compile—is scientifically unreliable for single-function vulnerability repair. Through five controlled experiments involving 203 vulnerable functions from the Big-Vul dataset, three open-source code LLMs (ranging from 350M to 6.7B parameters), and three prompting strategies, the researchers identified several significant issues.
Firstly, compile rate shows minimal responsiveness to interventions that genuinely improve the generated code. Secondly, it is heavily dominated by artifacts from the evaluation harness and dataset rather than reflecting true model quality. Approximately 64% of compile failures were found to be unrelated to the model's output, a proportion that remained consistent across different model sizes. Thirdly, the compile rate for identical patches can fluctuate significantly—by 1.8 to 2.7 times—simply due to changes in compiler standard flags, without any regressions in the code itself. Fourthly, compile rate can rank models in an order opposite to that of reference-similarity metrics. Finally, when used as an optimization target in a compiler-feedback loop, compile rate can inadvertently reward non-repairs, such as deletions or placeholder insertions, as similarity to the human-written fix decreases.
The study also found that the common fallback metric, whole-function CodeBLEU, is also problematic. An unmodified copy of the vulnerable input code scored higher than any of the tested models, indicating that it prioritizes unchanged context over actual repair. As an alternative, the researchers examined diff_F1, a change-aware screen that focuses solely on the edited region. While not a definitive repair-quality metric, diff_F1 assigns zero credit to no-operation changes and very little credit to some deletion-based patches that can game the compilation metric. However, it does credit genuine partial edits and may serve as a useful preliminary screen before more in-depth, execution-based analysis.
Why it matters for builders
Builders developing or utilizing LLMs for code vulnerability repair must be aware of the limitations of current evaluation methodologies. The findings suggest that relying solely on compile rate can lead to a false sense of progress and misdirect development efforts. Understanding that a significant portion of compilation failures are due to environmental factors or dataset artifacts, rather than model deficiencies, is crucial for accurate benchmarking. Furthermore, the susceptibility of compile rate and CodeBLEU to manipulation highlights the need for more robust and meaningful evaluation metrics that truly reflect the quality and security impact of LLM-generated code fixes.
Practical impact
Developers working with LLM-based code repair should reconsider their primary evaluation metrics. Instead of solely relying on compile rate, it is recommended to incorporate change-aware metrics like diff_F1 for an initial screening of generated patches. For a more comprehensive assessment, builders should prioritize execution-based analysis and vulnerability-specific testing where possible. The research also points to the importance of standardizing evaluation environments and compiler flags to ensure reproducible and comparable results across different models and experiments. The authors provide code and data for their experiments, encouraging further investigation into these metrics.
Caveats and source limits
This study focuses specifically on single-function vulnerability repair in C/C++ code and uses the Big-Vul dataset. The proposed diff_F1 metric is presented as a screen, not a definitive measure of repair quality, and the paper acknowledges its limitations, noting that it may not penalize all forms of gaming patches sufficiently. The research did not include pricing information for the LLMs used, nor did it provide specific release dates for the models themselves. The study's findings are based on empirical experiments and may not generalize to all LLM architectures or vulnerability types without further investigation. The source does not provide independent benchmark results for diff_F1 against other proposed metrics.
Sources
Claim check: 5/5 supported claims - 5 evidence links - 93% avg confidence
- Compile rate is a scientifically unreliable metric for single-function vulnerability repair.supported - arxiv.org
- Approximately 64% of compile failures in LLM-based code vulnerability repair are not attributable to the model.supported - arxiv.org
- Compile rate can shift by 1.8 to 2.7 times on identical patches under a single compiler-standard flag, with zero regressions.supported - arxiv.org
- Whole-function CodeBLEU scores an unchanged copy of vulnerable input higher than every tested model.supported - arxiv.org
- The `diff_F1` metric gives zero credit to a no-op and near-zero credit to some deletion-based gaming patches.supported - arxiv.org
Caveats
- Based on empirical study findings.
- This share is nearly invariant across models tested.
- Indicates sensitivity to toolchain configuration.
- Suggests CodeBLEU rewards preserved context over repair.
- It is proposed as a cheap screen before deeper analysis, not a repair-quality metric.
- Single-source caution: verify critical details at the linked source.
Radar score 81/100 - how it was calculated
- Reliability 80: Research metadata source
- Freshness 50: Fresh research date
- Novelty 87: Research implementation signal
- Technical 90: Research technical evidence
- Developer 74: Research developer relevance
- Ecosystem 68: Research implementation signal
- Confidence 98: Claims have reliable evidence