Twitter/X

On 2026-07-27 @cramforce converted deepsec.sh into DeepsecBench, a cybersecurity…

Brief

cramforce turned deepsec.sh into DeepsecBench (published 2026-07-27), a benchmark that evaluates AI models on finding security vulnerabilities in real application code by measuring accuracy, cost, and speed. The benchmark highlights linear cost scaling with codebase size and shows GPT-5.6 Sol top-scoring, Kimi K3 at half the score for 1/5 the cost, and Grok 4.5 best score/cost among top-10 models.

Why it matters

On 2026-07-27 @cramforce converted deepsec.sh into DeepsecBench, a cybersecurity evaluation that runs AI agents against real code to find vulnerabilities, because the task reflects real-world time and monetary costs.

Key details

  • DeepsecBench explicitly measures model accuracy, cost, and speed when finding application-code vulnerabilities; running deepsec on larger codebases scales linearly in cost, making cost/quality tradeoffs directly applicable to production workloads.
  • Vercel's published DeepsecBench results show GPT-5.6 Sol scored highest; Kimi K3 achieved roughly half the top score while incurring about one-fifth the cost; Grok 4.5 delivered the best score-to-cost ratio among the top 10 models.
Cleaned source text

We turned deepsec.sh into a cyber security eval ⛨

I think the benchmark is really helpful because the AI is solving a real world task that takes real time and drives substantial cost. Running deepsec on larger codebases scales linearly in cost.

Respectively, the cost data and the quality tradeoffs are directly applicable to real world workloads.

Link

GitHub - vercel-labs/deepsec: Deepsec is a security harness for finding vulnerabilities in your...

Deepsec is a security harness for finding vulnerabilities in your codebase powered by coding agents - vercel-labs/deepsec

github.com

Vercel (@vercel)

DeepsecBench evaluates model accuracy, cost, and speed in finding cybersecurity vulnerabilities.

Latest results:

▪️ GPT-5.6 Sol scores highest

▪️ Kimi K3 gets half the top score at 1/5 the cost

▪️ Grok 4.5 wins best score/cost ratio in top 10

vercel.com/blog/deepsecbench…

DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities

Today we're releasing DeepsecBench, a benchmark that measures how well different models find security vulnerabilities in application code.

vercel.com

— https://nitter.net/vercel/status/2081846100173177313#m