Twitter/X

A PE-backed medical billing company enabled one senior AI engineer plus a…

Brief

A PE-backed medical billing company used an AI Velocity Pod (one senior engineer + fractional AI delivery lead, 1.5 FTE) to capture 20 years of a co-founder’s SQL-based domain expertise in six weeks. They formalized denial rules, remittance pipelines, and a decision-tree evaluation, built a 60-case golden test harness with automated judges, and achieved 59/60 accuracy, penny-perfect reconciliation, and major time savings while running on $200–$300/month compute.

Why it matters

A PE-backed medical billing company enabled one senior AI engineer plus a fractional AI delivery lead (total 1.5 FTE) to encode 20 years of domain expertise into an agent in six weeks, running on $200–$300/month in compute.

Key details

  • The team built an eval harness of 60 golden test cases (verified by the co-founder’s SQL) with automated judges scoring 1–5; results on first deploy: 59/60 correct, flagship denial analysis reconciled to the penny, 15/15 denial cases passed, and the agent self-corrected after a CPTO challenge.
  • The single critical condition is knowledge capture: the engineer spent three weeks eliciting the co-founder's reasoning (denial definitions as lookup rules, remittance logic as a sequenced pipeline, evaluation order as a decision tree); buyers report the same bottleneck of one indispensable domain expert.
Source evidence

At a PE-backed medical billing company processing claims for hundreds of healthcare facilities, one AI engineer encoded 20 years of domain expertise in six weeks.

Twenty years of data, a working rules engine, and one person who could answer any analytical question about the system: the co-founder, hand-writing SQL against a legacy schema. She was the only person who understood the data.

Inside the pod:
1. One senior AI engineer building the agent and encoding domain methodology
2. A fractional AI delivery lead on architecture, review, and quality gates

Total: 1.5 FTE. $200 to $300 a month in compute.

The old way: the co-founder ran denial analyses, financial health assessments, and payer mix questions by hand. Full working days on single research requests. She was the only path to the data.

The engineer spent the first three weeks writing zero code. Each session went to encoding how the co-founder reasons through queries. Denial definitions became scoped lookup rules. Remittance logic became a sequenced pipeline. Evaluation order became a decision tree the agent follows before output.

The eval harness came next: 60 golden test cases from the co-founder's actual workload, verified by her own SQL against production. Automated judges grade responses on correctness and grounding. Score of 1: the agent invented a number. Score of 5: the claim traces to the data. The team doesn't ship any release that drops accuracy.

After six weeks on the same data:

  1. 59 of 60 golden questions correct on first deploy
  2. Flagship denial analysis reconciles to the penny against the source database
  3. 15 of 15 denial cases pass automated grading
  4. CPTO challenged one answer; the agent self-corrected and matched
  5. Nine out of ten hours saved per week. $200 to $300 a month in compute.

The co-founder had been spending full working days on those same requests. Nobody else could navigate the schema.

One condition makes this work or fail. The knowledge capture can't be skipped.

Buyers describe the same bottleneck: their most valuable domain expert buried under manual queries nobody else can run.

AI raised the premium on domain expertise.

With an AI Velocity Pod, that person compounds: reviewing and directing while the agent handles volume, verified to the penny.