unsloth is one of the most underrated teams in ai and it's not close. while the timeline fights about the frontier, they quietly write the kernels that let you fine-tune on 3gb of vram, 2x faster with 70% less memory, no accuracy lost.
and now they've brought real training support to amd hardware that was effectively cuda only until this week. radeon, ryzen, instinct, a massive install base that couldn't touch local training just got a first class path in.
the labs making the models get the headlines, the people making the models run on your actual hardware get a quote tweet. that ratio is broken. go give unsloth their due.
Unsloth AI (@UnslothAI)
Introducing Unsloth for AMD 🚀
You can now train & run LLMs on your AMD hardware
• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
• Works on Windows, WSL, Linux
• Train Qwen, Gemma on 3GB VRAM
GitHub: github.com/unslothai/unsloth
Works on Radeon, Instinct, Ryzen and data center GPUs with up to 2× faster with 70% less VRAM and no accuracy loss via our custom Triton kernels and math algorithms. We also support optimized ROCm builds for GGUF & Safetensors inference.
Unsloth is an open-source local UI for faster LLM training and inference, with tool-call healing, code execution, secure web search, remote APIs, and HTTPS deployment. Connect local models to Claude Code, Codex agents and run the latest Kimi, GLM, DeepSeek, Qwen3.6, and Gemma 4 models.
🔗Blog + Guide: unsloth.ai/docs/basics/amd
— https://nitter.net/UnslothAI/status/2079207457788952944#m