Twitter/X

LFM2.5-2.6B matched DeepSeek-V4-Flash on agentic tool-calling (both achieved…

Brief

LFM2.5–2.6B, per a 2026-08-06 post, matched DeepSeek‑V4‑Flash’s tool-calling correctness (35/35 calls) across three agentic tasks while running ~3.7× faster on a single box with 4×RTX 5090 GPUs (366 tok/s, 19s vs 77 tok/s, 70s). The author highlights that LFM was trained for agents and is small enough to run on phones.

Why it matters

LFM2.5-2.6B matched DeepSeek-V4-Flash on agentic tool-calling (both achieved 35/35 tool calls) in a 3-job benchmark run reported 2026-08-06.

Key details

  • Measured on one box with 4×RTX 5090 GPUs, LFM2.5-2.6B ran ~3.7× faster: 366 tok/s and 19s total vs DeepSeek-V4-Flash at 77 tok/s and 70s; both models fired 12 tool calls in a single turn.
  • Benchmark jobs: (1) weather and local time for six cities, (2) convert one budget into six currencies, (3) check and book four hotels for one date; author emphasizes LFM was trained for agents and can run on a phone.
Source evidence

LFM vs DeepSeek Flash on agentic tool calls

atomic.chat (@atomicchathq)

New LFM2.5-2.6B hits DeepSeek-V4 level on tool calling and runs 3.7x faster!

We ran @liquidai 's new LFM2.5-2.6B against DeepSeek-V4-Flash on one box with 4x RTX 5090. Both got the same three jobs, and each one only completes if the model fires every tool call

Topics:
-weather and local time in six cities
-one budget into six currencies
-four hotels checked and booked for one date

Outputs:
LFM2.5-2.6B: 35/35 tool-calls, 366 tok/s, 19s
DeepSeek-V4-Flash: 35/35 tool-calls, 77 tok/s, 70s

Both models made all 35 calls, and both fired twelve of them in one turn. LFM was trained for agents, and tool calling is where that shows. Strong result for a model you can run right on your phone

Video

— https://nitter.net/atomicchathq/status/2085405031474343963#m