Twitter/X

Thinking Machines released Inkling-Small on 2026-07-30

Brief

Inkling-Small, released 2026-07-30 by Thinking Machines, is a 276B-parameter MoE model with 12B active parameters that the team says matches Inkling at one quarter the size. Apache-2.0 licensed and multimodal (text/image/audio) with a 1M-token context, full GGUF weights and docs are available for local 128GB RAM use and Tinker fine-tuning.

Why it matters

Thinking Machines released Inkling-Small on 2026-07-30: a 276B-parameter Mixture-of-Experts model with 12B active parameters that the company claims matches Inkling's performance at one quarter the size.

Key details

  • Inkling-Small is Apache-2.0 licensed, described by @UnslothAI as 'the strongest open model for its size', and supports multimodal inputs (text, image, audio) plus a 1,000,000-token (1M) context window and controllable reasoning effort.
  • Full weights and GGUF are available (docs and download links published); the model can run locally on 128GB RAM, can be fine-tuned on Tinker, and used for chat in the Tinker Playground.
Source evidence

You can now run Inkling-Small, a new 276B model by Thinking Machines.

Inkling-Small is the strongest open model for its size and runs local on 128GB RAM.

Apache-2.0 Licensed, it has image, audio + 1M context support.

Guide: unsloth.ai/docs/models/inkli…
GGUF: huggingface.co/unsloth/Inkli…

Thinking Machines (@thinkymachines)

Today, we are releasing Inkling-Small.

Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.

thinkingmachines.ai/news/ink…

Fine-tune it on Tinker today, or chat with it in text, image, and audio on Tinker Playground.

Link

Introducing Inkling-Small

An open-weights model that matches Inkling at a quarter of the size: multimodal, Mixture-of-Experts, with controllable reasoning effort. Fine-tune it on Tinker.
thinkingmachines.ai

— https://nitter.net/thinkymachines/status/2082885869426631032#m