gpt 5.6 will run on Cerebras chips at 750 tokens per second. can you imagine how fast (and cheap) this will be once their custom jalapeño chip launches end of year?
openai’s bet on owning the ai hardware stack is going to kill. cheaper intelligence, lightning speed AND saves on compute which they can redistribute to (you guessed it): building a better model
v cool
Adam.GPT (@TheRealAdamG)
openai.com/index/previewing-…
"GPT‑5.6 is priced per 1M tokens across three model sizes: Sol is $5 input / $30 output; Terra is $2.50 input / $15 output; and Luna is $1 input / $6 output. GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints and a 30-minute minimum cache life. For GPT‑5.6 and later models, cache writes are billed at 1.25x the model’s uncached input rate, while cache reads continue to receive the 90% cached-input discount.
We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity."
— https://nitter.net/TheRealAdamG/status/2070573862224580821#m