Twitter/X

Author @0xSero applied a method to Qwen3.6-35B (repo…

Brief

0xSero shares a Hugging Face release (Qwen3.6-35B-Hyrbid-3.25bpw) claiming a method that buys more context when the model is run on 24 GB and 32 GB GPUs. The thread also highlights a community project that got GLM-5.2 running at near-lossless quality on a roughly $15,000 home rig, using insights from extensive RTX Pro 6000 message analysis.

Why it matters

Author @0xSero applied a method to Qwen3.6-35B (repo: 0xSero/Qwen3.6-35B-Hyrbid-3.25bpw on Hugging Face) that reportedly increases usable context when running on 24 GB and 32 GB GPUs.

Key details

  • The same account reports a community effort to run GLM-5.2 at near-lossless quality on a $15,000 home setup, based on analysis of tens of thousands of messages using an RTX Pro 6000.
Cleaned source text

Applied this method to Qwen3.6-35B buys you more context on 24gb/32gb

huggingface.co/0xSero/Qwen3.…

Link

0xSero/Qwen3.6-35B-Hyrbid-3.25bpw · Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co

0xSero (@0xSero)

Article

GLM-5.2 at home for 15,000$

Here's how a community of tinkerers got GLM-5.2 running at near lossless quality on a 15,000$ budget. This article was put together using tens of thousands of analyzed messages in the RTX Pro 6000

— https://nitter.net/0xSero/status/2079230064840106173#m