Twitter/X

At a PE-backed medical billing company serving hundreds of healthcare facilities…

Brief

A PE-backed medical billing company encoded two decades of a co-founder’s SQL-based expertise into an AI agent in six weeks: one senior engineer plus a fractional delivery lead (1.5 FTE), $200–$300/month compute. An evaluation harness of 60 golden test cases produced 59/60 correct, denial analyses reconciled to the penny, 15/15 automated grading passes, and claimed nine out of ten hours saved per week; author argues AI raises the premium on domain experts.

Why it matters

At a PE-backed medical billing company serving hundreds of healthcare facilities, one senior AI engineer plus a fractional AI delivery lead (total 1.5 FTE) encoded 20 years of a co-founder’s SQL-based domain expertise into an agent in six weeks, using $200–$300/month in compute.

Key details

  • An evaluation harness of 60 golden test cases (verified by the co-founder’s SQL) produced 59/60 correct on first deploy; the flagship denial analysis reconciled to the penny, 15/15 denial cases passed automated grading, and the agent self-corrected when the CPTO challenged an answer.
  • The team reported 'Nine out of ten hours saved per week' for the co-founder; the author argues AI raises the premium on domain expertise because knowledge capture is the critical condition for scaling—an 'AI Velocity Pod' lets the expert compound impact by reviewing and directing the agent.
Source evidence

“AI raised the premium on domain expertise.”

Agreed

Mark Ajzenstadt (@mardehaym)

At a PE-backed medical billing company processing claims for hundreds of healthcare facilities, one AI engineer encoded 20 years of domain expertise in six weeks.

Twenty years of data, a working rules engine, and one person who could answer any analytical question about the system: the co-founder, hand-writing SQL against a legacy schema. She was the only person who understood the data.

Inside the pod:
1. One senior AI engineer building the agent and encoding domain methodology
2. A fractional AI delivery lead on architecture, review, and quality gates

Total: 1.5 FTE. $200 to $300 a month in compute.

The old way: the co-founder ran denial analyses, financial health assessments, and payer mix questions by hand. Full working days on single research requests. She was the only path to the data.

The engineer spent the first three weeks writing zero code. Each session went to encoding how the co-founder reasons through queries. Denial definitions became scoped lookup rules. Remittance logic became a sequenced pipeline. Evaluation order became a decision tree the agent follows before output.

The eval harness came next: 60 golden test cases from the co-founder's actual workload, verified by her own SQL against production. Automated judges grade responses on correctness and grounding. Score of 1: the agent invented a number. Score of 5: the claim traces to the data. The team doesn't ship any release that drops accuracy.

After six weeks on the same data:

  1. 59 of 60 golden questions correct on first deploy
  2. Flagship denial analysis reconciles to the penny against the source database
  3. 15 of 15 denial cases pass automated grading
  4. CPTO challenged one answer; the agent self-corrected and matched
  5. Nine out of ten hours saved per week. $200 to $300 a month in compute.

The co-founder had been spending full working days on those same requests. Nobody else could navigate the schema.

One condition makes this work or fail. The knowledge capture can't be skipped.

Buyers describe the same bottleneck: their most valuable domain expert buried under manual queries nobody else can run.

AI raised the premium on domain expertise.

With an AI Velocity Pod, that person compounds: reviewing and directing while the agent handles volume, verified to the penny.

— https://nitter.net/mardehaym/status/2085632086048752126#m