Evidence · DiggingBeagle record
BaxBench: Can LLMs Generate Secure and Correct Backends?
BaxBench evaluates 392 security-critical backend tasks with functional tests and expert-designed exploits. Its project page reports that 62% of solutions from the best tested model were either incorrect or vulnerable and that roughly half of correct solutions were insecure on average.
- Published
- 2025 (year precision)
- Source role
- primary disclosure
Evidence record
BaxBench evaluates 392 security-critical backend tasks with functional tests and expert-designed exploits. Its project page reports that 62% of solutions from the best tested model were either incorrect or vulnerable and that roughly half of correct solutions were insecure on average.
Claim-level citations (2)
- contextFive coding agents produced 69 findings across 15 matched application builds: The evidence does not support the blanket claim that generated security failures occur only when the requirement was absent from the prompt: Tenzai reports authorization failures despite detailed guidance, and SUSVIBES reports that preliminary security prompting did not eliminate the secure-coding gap.
Key Takeaways and prompt-setting descriptions
- supportsFive coding agents produced 69 findings across 15 matched application builds: Independent benchmarks support the distinction between functional completion and secure completion but do not reproduce Tenzai's 15-app result. BaxBench reports that end-to-end exploits succeeded against around half of functionally correct generated backends on average, while the final ICML 2026 SUSVIBES paper reports 57% functional correctness but only 11.8% secure solutions for SWE-Agent with Claude 4 Sonnet.
PMLR abstract and key findings: 392 backend-generation tasks evaluated with functional tests and end-to-end exploits; security exploits succeeded against around half of the functionally correct programs generated by each LLM on average.