Twitter/X

Author claims Unsloth is "one of the most underrated teams in AI" and built…

Brief

Unsloth has developed custom Triton kernels and math optimizations that enable local LLM training on very low-end hardware (3 GB VRAM) with claimed 2× speedups, 70% less memory and no accuracy loss. On 2026-07-20 they announced AMD support for 500+ models across Radeon/Instinct/Ryzen (Windows/WSL/Linux) and an open-source local UI for training, inference and deployment.

Why it matters

Author claims Unsloth is "one of the most underrated teams in AI" and built custom Triton kernels and math algorithms that allow fine-tuning LLMs on 3 GB VRAM with up to 2× faster runtime and 70% less memory use, reportedly with no accuracy loss.

Key details

  • On 2026-07-20 Unsloth announced an AMD collaboration enabling training and inference for 500+ models on Radeon, Instinct, Ryzen and data-center GPUs across Windows/WSL/Linux; supports ROCm-optimized GGUF & Safetensors, can train Qwen and Gemma families, and ships an open-source local UI with tool-call healing, code execution, remote APIs and HTTPS deployment.
Source evidence

unsloth is one of the most underrated teams in ai and it's not close. while the timeline fights about the frontier, they quietly write the kernels that let you fine-tune on 3gb of vram, 2x faster with 70% less memory, no accuracy lost.

and now they've brought real training support to amd hardware that was effectively cuda only until this week. radeon, ryzen, instinct, a massive install base that couldn't touch local training just got a first class path in.

the labs making the models get the headlines, the people making the models run on your actual hardware get a quote tweet. that ratio is broken. go give unsloth their due.

Unsloth AI (@UnslothAI)

Introducing Unsloth for AMD 🚀
You can now train & run LLMs on your AMD hardware

• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
• Works on Windows, WSL, Linux
• Train Qwen, Gemma on 3GB VRAM

GitHub: github.com/unslothai/unsloth

Works on Radeon, Instinct, Ryzen and data center GPUs with up to 2× faster with 70% less VRAM and no accuracy loss via our custom Triton kernels and math algorithms. We also support optimized ROCm builds for GGUF & Safetensors inference.

Unsloth is an open-source local UI for faster LLM training and inference, with tool-call healing, code execution, secure web search, remote APIs, and HTTPS deployment. Connect local models to Claude Code, Codex agents and run the latest Kimi, GLM, DeepSeek, Qwen3.6, and Gemma 4 models.

🔗Blog + Guide: unsloth.ai/docs/basics/amd

— https://nitter.net/UnslothAI/status/2079207457788952944#m