We spend a ton of time with customers struggling to write evals and figure out the best model for each task.
This CLI wraps that experience into a single tool, allowing you to leverage all 500+ models on OpenRouter, their benchmarks, and real-world usage data.
It's brand new. Give us feedback!
OpenRouter (@OpenRouter)
Introducing Ori Eval: the easiest way to write your first eval.
There's no definitive best model, only the best model for each task. Ori Eval leverages OpenRouter's APIs for each task in your codebase, and then evaluates the results.
curl -fsSL openrouter.ai/skills/spawn-o…
— https://nitter.net/OpenRouter/status/2084301100078027143#m