Twitter/X

@godofprompt argues that the cost problem in AI apps stems from sending every…

Brief

@godofprompt argues that the cost problem in AI apps stems from sending every prompt to the priciest model and that the solution is routing: use cheap models for trivial tasks, stronger models selectively, and fallbacks for outages. OrcaRouter (@OrcaRouter) launches, claiming it doesn’t take a cut while other LLM routers charge about a 5% per-token fee.

Source evidence

Most AI apps are still burning money by sending every prompt to the most expensive model.

That won’t last.

The real unlock is routing: cheap models for simple jobs, stronger models when it actually matters, fallback when providers break.

This is the infra layer everyone eventually needs.

OrcaRouter (@OrcaRouter)

1/ OrcaRouter goes live today!

The other LLM routers tax you ~5% on every token.
OrcaRouter doesn't take a cut.

Here's how that works, and why it matters for what you're building ↓

Video

— https://nitter.net/OrcaRouter/status/2052780131459149918#m