Twitter/X

GPT‑5.6 will run on Cerebras hardware at up to 750 tokens per second; OpenAI says…

Brief

GPT‑5.6 will run on Cerebras hardware at up to 750 tokens/sec (OpenAI plans Sol on Cerebras in July 2026) while introducing per‑1M‑token pricing for Sol/Terra/Luna and tightened prompt‑caching rules. The poster argues OpenAI owning the hardware stack—and Cerebras' upcoming 'jalapeño' chip later in 2026—will drive much faster, cheaper inference and accelerate model development.

Why it matters

GPT‑5.6 will run on Cerebras hardware at up to 750 tokens per second; OpenAI says GPT‑5.6 Sol on Cerebras will be available in July 2026 with initial access limited to select customers.

Key details

  • OpenAI priced GPT‑5.6 per 1M tokens: Sol $5 input / $30 output; Terra $2.50 input / $15 output; Luna $1 input / $6 output; new caching rules include explicit cache breakpoints, a 30‑minute minimum cache life, cache writes billed at 1.25× the uncached input rate, and cache reads receiving a 90% cached‑input discount.
  • The author claims OpenAI’s strategy to own the AI hardware stack (Cerebras now, and a custom 'jalapeño' chip expected end of year 2026) will make intelligence much cheaper and faster and free compute to 'build a better model.'
Source evidence

gpt 5.6 will run on Cerebras chips at 750 tokens per second. can you imagine how fast (and cheap) this will be once their custom jalapeño chip launches end of year?

openai’s bet on owning the ai hardware stack is going to kill. cheaper intelligence, lightning speed AND saves on compute which they can redistribute to (you guessed it): building a better model

v cool

Adam.GPT (@TheRealAdamG)

openai.com/index/previewing-…

"GPT‑5.6 is priced per 1M tokens across three model sizes: Sol is $5 input / $30 output; Terra is $2.50 input / $15 output; and Luna is $1 input / $6 output. GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints and a 30-minute minimum cache life. For GPT‑5.6 and later models, cache writes are billed at 1.25x the model’s uncached input rate, while cache reads continue to receive the 90% cached-input discount.

We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity."

— https://nitter.net/TheRealAdamG/status/2070573862224580821#m