Twitter/X

AutoRouter now integrates with AutoTune to detect when to train a custom SLM for…

Brief

AutoRouter integrates with AutoTune and today launches Inference AutoTune (Sam Hogan, 2026-07-21), enabling users to distill frontier models into 1–30B parameter task-specific SLMs with about 25 lines of code. The system promises ~2-hour, <$250 training runs, automatic request routing that reduces cost and latency by over 90%, user-owned weights, and availability in private beta.

Why it matters

AutoRouter now integrates with AutoTune to detect when to train a custom SLM for cost-, quality-, and latency-sensitive use cases (announcement by Sam Hogan, published 2026-07-21).

Key details

  • Inference AutoTune can distill any frontier model into a 1–30B parameter task-specific SLM with ~25 lines of code, trains in ~2 hours for under $250, and claims automatic routing that cuts cost and latency by >90%; weights remain user-owned and the feature is in private beta.
Source evidence

AutoRouter has first-class integration with AutoTune, our automatic fine-tuning and distillation engine, to detect when a custom SLM should be trained for use cases where cost, quality, and latency are paramount.

Sam Hogan 🇺🇸 (@samhogan)

We're releasing Inference AutoTune

Distill any frontier model into a 1-30B parameter task-specific SLM with only 25 lines of code

automatically route requests to reduce cost and latency by >90%

~2 hours and <$250 to train. You own the weights

Available in private beta today

— https://nitter.net/samhogan/status/2076044602554159240#m