We turned deepsec.sh into a cyber security eval ⛨
I think the benchmark is really helpful because the AI is solving a real world task that takes real time and drives substantial cost. Running deepsec on larger codebases scales linearly in cost.
Respectively, the cost data and the quality tradeoffs are directly applicable to real world workloads.
Link
GitHub - vercel-labs/deepsec: Deepsec is a security harness for finding vulnerabilities in your...
Deepsec is a security harness for finding vulnerabilities in your codebase powered by coding agents - vercel-labs/deepsec
github.com
Vercel (@vercel)
DeepsecBench evaluates model accuracy, cost, and speed in finding cybersecurity vulnerabilities.
Latest results:
▪️ GPT-5.6 Sol scores highest
▪️ Kimi K3 gets half the top score at 1/5 the cost
▪️ Grok 4.5 wins best score/cost ratio in top 10
vercel.com/blog/deepsecbench…
DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities
Today we're releasing DeepsecBench, a benchmark that measures how well different models find security vulnerabilities in application code.
vercel.com
— https://nitter.net/vercel/status/2081846100173177313#m