Twitter/X

DeepsecBench runs the unmodified deepsec security analysis; GPT‑5.6 Sol costs $56…

Brief

DeepsecBench runs the unmodified deepsec security scan so reported costs scale linearly: GPT‑5.6 Sol costs $56 to analyze DeepsecBench's 50 entry files (implying $1,120 for 1,000 files), while Grok 4.5 costs $5.60 (10%) — ~$112 for 1,000 files — but finds less than half the issues. Vercel's results put GPT‑5.6 Sol highest, Kimi K3 at half score for 1/5 the cost, and Grok 4.5 with the best score/cost ratio, highlighting a clear accuracy vs. cost trade-off.

Why it matters

DeepsecBench runs the unmodified deepsec security analysis; GPT‑5.6 Sol costs $56 to analyze the benchmark's 50 entry files, which extrapolates to ~$1,120 per run for 1,000 files and ~$11,200 for 10,000 files (post dated 2026-07-28).

Key details

  • Grok 4.5 costs $5.60 to run the same 50-file set (10% of GPT‑5.6 cost), extrapolating to ~$112 for 1,000 files, but Grok 4.5 finds less than half the issues; multiple runs can partly recover coverage but still leave many findings missed.
  • Vercel's DeepsecBench ranks models by accuracy, cost, and speed: GPT‑5.6 Sol scores highest; Kimi K3 achieves ~50% of the top score at ~1/5 the cost; Grok 4.5 has the best score/cost ratio among the top 10—illustrating a quality vs. cost trade-off.
Source evidence

The beauty of DeepsecBench is that it straight-up just runs the deepsec security analysis unmodified, which means you can extrapolate from the cost reported in the benchmark.

The leader GPT5.6 Sol costs $56 to analyze the 50 entry files that the benchmark is based on. Imagine your code base has 1000 files that deepsec considers interesting enough to start a scan from: This would cost you $1,120 per run. And $11,200 for 10K files and so forth

Running deepsec on the same set with Grok 4.5 costs $5.60 (coincidentally exactly 10%). And so at the 1K files is only $112.

Now, that same Grok run find less than half the issues. At 90% off you can make up some of the coverage through multiple runs, but you will find less stuff to fix.

It's a real trade-off between quality and cost, but I'm glad we have the options.

Vercel (@vercel)

DeepsecBench evaluates model accuracy, cost, and speed in finding cybersecurity vulnerabilities.

Latest results:
▪️ GPT-5.6 Sol scores highest
▪️ Kimi K3 gets half the top score at 1/5 the cost
▪️ Grok 4.5 wins best score/cost ratio in top 10
vercel.com/blog/deepsecbench…

Link

DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities

Today we're releasing DeepsecBench, a benchmark that measures how well different models find security vulnerabilities in application code.
vercel.com

— https://nitter.net/vercel/status/2081846100173177313#m