Twitter/X

Install Inference Gateway (docs.inference.net) and continue sending traffic to…

Brief

samhogan presents a six-step flow to trial Kimi K3 in production via Inference.net's Inference Gateway: install the Gateway, keep your current provider, mirror live traffic to Kimi K3 while an RLM runs evals for ~24 hours, receive a Slack approval, then change your model id to "kimi-k3"—claiming a 50% token-cost reduction and full stack ownership.

Why it matters

Install Inference Gateway (docs.inference.net) and continue sending traffic to your current provider while the Gateway mirrors live traffic to Kimi K3 for evaluation.

Key details

  • The Gateway uses an RLM to generate evals for your app over ~24 hours and sends a Slack notification when evals look healthy so you can safely switch.
  • After switching your code's model id to "kimi-k3", the author claims you save 50% on your monthly token bill and 'own your LLM stack end to end.'
Source evidence

Want to try Kimi K3 in production but worried how it might change your product?

Don’t worry, we got you:

  1. Install Inference Gateway (docs.inference.net)
  2. Keep sending traffic to your current provider
  3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours.
  4. Gateway starts mirroring live traffic to Kimi K3 to run evals. Traffic is mirrored - you’re still using your old model in prod.
  5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch.
  6. Switch model id in your code to “kimi-k3”

Congrats, you just saved 50% on your monthly token bill, and you own your LLM stack end to end.

Link

Catalyst by Inference.net - Inference.net Documentation

docs.inference.net