Want to try Kimi K3 in production but worried how it might change your product?
Don’t worry, we got you:
- Install Inference Gateway (docs.inference.net)
- Keep sending traffic to your current provider
- Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours.
- Gateway starts mirroring live traffic to Kimi K3 to run evals. Traffic is mirrored - you’re still using your old model in prod.
- Once evals look healthy, you get a Slack notification letting you know it’s safe to switch.
- Switch model id in your code to “kimi-k3”
Congrats, you just saved 50% on your monthly token bill, and you own your LLM stack end to end.
Link
Catalyst by Inference.net - Inference.net Documentation
docs.inference.net