Twitter/X

LFM2.5-230M released by @liquidai

Brief

LFM2.5-230M is a 230M-parameter model from @liquidai, pre-trained on 19T tokens with a 32K context and distilled from LFM2.5-350M, designed to run on CPUs, NPUs, and GPUs for on-device agentic workloads. LiquidAI reports 213 tok/s on a Galaxy S25 Ultra CPU and 42 tok/s on a Raspberry Pi 5, claiming it matches or outperforms models over twice its size on instruction following, data extraction, and tool use.

Why it matters

LFM2.5-230M released by @liquidai: a 230M-parameter model built to run on CPUs, NPUs, and GPUs for agentic tasks on phones, robots, and automation devices (announcement published 2026-06-25).

Key details

  • Training and architecture details: built on the LFM2 architecture, pre-trained on 19T tokens with a 32K context extension, and post-trained via distillation from LFM2.5-350M.
  • Performance claims: 213 tok/s decoding on a Galaxy S25 Ultra (CPU) and 42 tok/s on a Raspberry Pi 5 (CPU); reported to compete with or beat models over twice its size on instruction following, data extraction, and tool use.
Source evidence

Introducing LFM2.5-230M: our smallest model yet, built to run fast anywhere (CPUs, NPUs, and GPUs) to enable agentic tasks on phones, robots, home and network automation devices.

> 230M parameters, built on the LFM2 architecture
> Pre-trained on 19T tokens, with a 32K context extension
> Post-trained with distillation from LFM2.5-350M
> 213 tok/s decode speed on Galaxy S25 Ultra (CPU)
> 42 tok/s on a Raspberry Pi 5 (CPU)
> Competes with and often beats models more than twice its size on instruction following, data extraction, and tool use.
> use it for large-scale data extraction pipelines or lightweight on-device agentic workloads.

🧵