Twitter/X

cryptopunk7213: Over the next 6 months, the most valuable capabilities will be…

Brief

cryptopunk7213 and Brian Armstrong argue the near-term value lies in model selection per task, multi-model routing/harnesses, caching, and AI-cloud memory/DB management as aggregation becomes the bottleneck. Armstrong reports defaulting to GLM 5.2 and Kimi 2.7, 91% of employees never hit caps, a LibreChat cache hit improvement from 5%→60%, and nearly halved AI spend while tokens grow.

Why it matters

cryptopunk7213: Over the next 6 months, the most valuable capabilities will be per-task model choice, multi-model routing/harnesses, caching to reduce token spend, and AI-cloud management of memory/databases; the bottleneck is shifting from LLM capacity to aggregation.

Key details

  • Brian Armstrong: Coinbase is defaulting to open-weight models (GLM 5.2 and Kimi 2.7) via an LLM gateway to control cost rather than hard caps; 91% of employees never hit usage caps so cheaper defaults steer behavior without friction.
  • Armstrong operational wins: preprocess routing and cache-awareness raised LibreChat cache hit rate from 5% to 60%; combined better defaults, routing, caching, lean context, and visibility cut AI spend nearly in half while token usage keeps growing.
Source evidence

few things that’ll become increasingly more valuable over the next 6 months:

  • which ai models you use for what tasks
  • how you route across multiple models per task (harness)
  • caching to reduce token spend
  • ai cloud: efficiently managing memory and data bases

the bottleneck is shifting from LLMs to aggregation

Brian Armstrong (@brian_armstrong)

How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching.

Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work.

Better Routing – In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task.

Better Caching – Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so we’re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% → 60% in LibreChat once properly implemented.

Keep Context Lean – Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted.

Better Visibility – Our engineers can use as many tokens as they want, from whatever model they want, but we’ve made usage visible – and the more you spend on AI, the more impact we expect.

The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable.

Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.

— https://nitter.net/brian_armstrong/status/2070670644577280109#m