Evidence · DiggingBeagle record

BaxBench: Can LLMs Generate Secure and Correct Backends?

BaxBench evaluates 392 security-critical backend tasks with functional tests and expert-designed exploits. Its project page reports that 62% of solutions from the best tested model were either incorrect or vulnerable and that roughly half of correct solutions were insecure on average.

Published
2025 (year precision)
Source role
primary disclosure

Evidence record

BaxBench evaluates 392 security-critical backend tasks with functional tests and expert-designed exploits. Its project page reports that 62% of solutions from the best tested model were either incorrect or vulnerable and that roughly half of correct solutions were insecure on average.

Read the original source ↗

Claim-level citations (2)