It’s crazy how much better our model training techniques have gotten.
The model thinks a lot more than swe 1.6 but also understands implicit correctness requirements much better. With fast serving I think this is a pretty good tradeoff.
Cognition (@cognition)
Introducing SWE-1.7, the most capable model we’ve trained yet.
It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s.
RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale
— https://nitter.net/cognition/status/2074882968770728416#m