Hey everyone!
I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion.
I fused weights from @liquidai's LFM2.5-2.6B & @Alibaba_Qwen's Qwen3.6-35B-A3B.
It's capable of near Qwen3.6-35B-A3B performances, while being around a fifth of the size.
This is the first model of a new series, which has been over 6 weeks of work so far fully dedicated on this.
I have two models coming soon with this architecture :
- a mini version of minimax m3 ( already built btw )
- a mini version of deepseek v4 flash ( already built too ).
26B and 12B.
Would deffinitely love to have more compute though in order benchmark those and run more experiments and make them even better. This is I believe the fastest way for us to achieve frontier intelligence locally.
@0xSero you have a lot of compute, maybe you could help me out finish my work in order to release these models opensource for everyone to use.
Then I can move on to the big boys ( glm 5.2, kimi k3 and soon qwen 3.8 hehehe ).
Link to the model : huggingface.co/Akahsizrr/fus…
Still working on a lot of quantizations, and a checkpoint with better sparse activation.
Link
Akahsizrr/fuse-1-Lite · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co