Twitter/X

Brad Flora asserts there is more AI compute than ever but it's "absurdly hard" to…

Brief

OpenRelay claims to solve access friction by unifying accelerators across clouds into one endpoint: the company reports 100B tokens/week across 22 locations on four continents, supports eight accelerator SKUs (NVIDIA, TPU, Trainium, AMD), and markets an Inference Delivery Network/CDN at up to 20% lower cost than renting hardware.

Why it matters

Brad Flora asserts there is more AI compute than ever but it's "absurdly hard" to access and points to OpenRelay, claiming it makes every chip across every cloud look like a single endpoint with 100B tokens/week already running through it.

Key details

  • OpenRelay says it runs across 22 locations on 4 continents, offers 8 accelerator SKUs (NVIDIA, TPU, Trainium, AMD), positions itself as an Inference Delivery Network/CDN, and promises up to 20% cost savings versus renting hardware directly.
Source evidence

There is more AI compute in the world than ever and it is still absurdly hard to actually get any of it. OpenRelay makes every chip across every cloud look like one endpoint. 100B tokens a week already running through it. If you need compute; they've got it.

OpenRelayInc (@OpenRelayInc)

Launching @openrelay: One Endpoint to Rule Them All

  • 100B tokens/week across 22 locations on 4 continents
  • 8 accelerator SKUs: NVIDIA, TPU, Trainium, AMD
  • Up to 20% cheaper than renting the hardware yourself

Every chip on earth. None of them your problem.

Run inference on the world’s compute
piped.video/7uNhqcjBXJg?si=rsBj…

Link

OpenRelay Launch

OpenRelay is building an Inference Delivery Network on our CDN of G...
youtube.com

— https://nitter.net/OpenRelayInc/status/2085815430250426709#m