Twitter/X

Inkling-small is a 276B-parameter model with 12B active weights (~1/4 the size of…

Brief

Inkling-small is a 276B total / 12B active model the author says matches a 4× larger sibling and sometimes outperforms it on benchmarks. With SGLang+DSpark achieving 648 tok/s (288 tok/s without) on 8× B200 and a Miles-verified multimodal RL workflow, the release is positioned to enable accessible full-parameter RL and production-ready inference/training.

Why it matters

Inkling-small is a 276B-parameter model with 12B active weights (~1/4 the size of its sibling); the post asserts it matches a 4× larger model in capability and even wins on some benchmarks.

Key details

  • SGLang+DSpark yields 648 tok/s decode (simulated acc len=4) and 288 tok/s without DSpark on 8× NVIDIA AI B200 (TP 8, NVFP4, bs=1); Miles verified a multimodal RL workflow and the release is presented as ready for full-parameter RL and day‑0 inference/training, with LoRA and full-parameter training within reach for normal teams.
Source evidence

a model that matches its 4x larger sibling in capability while fitting full-parameter RL on accessible hardware is the release that actually moves the ecosystem. big models make headlines, trainable models make impact. 276B/12B active hits the size where a normal team can take their own multimodal data and turn it into real capability gains through RL - not just prompt engineering around a frozen model. Miles verified for exactly this workflow, 648 tok/s decode with DSpark on the serving side. inference and training both ready day 0. this is what full-stack open infrastructure looks like @lmsysorg @thinkymachines @radixark

LMSYS Org (@lmsysorg)

Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x @NVIDIAAI B200, TP 8, NVFP4, bs=1).

What makes this model different is the size. 276B total with 12B active is a sweet spot for RL, and both LoRA and full-parameter training become well within reach. Miles is ready and verified for multimodal RL on Inkling-small, so you can turn your multimodal data into real capability gains.

At ~1/4 the size, Inkling-small matches the bigger version in capability and even wins on some benchmarks.

Run Inkling-small with SGLang, and customize it with Miles.

Video

— https://nitter.net/lmsysorg/status/2082890993179955322#m