Skip to content

Operative coding · scoped engine validation

In development

Highest supportable surgical coding — two separate proofs.

AI operative coding for orthopedic surgery. CodeIQ OR prepares CPT from the operative note. Historical agreement with surgeon-coded cases is one proof class. Adjudication concordance on billed lines is another. They are not the same claim.

The moment it solves

The case is closed. The coding still has to be right.

Surgeons should not be reconstructing CPT from memory at the end of a long list. CodeIQ OR proposes the highest supportable line and waits. Most of the adjudication set was retrospective — concordance with what was paid is not a claim that the engine caused the payment.

How it works

Prepared work. Physician still decides.

  1. 01

    Read the operative note

    Primary procedure, laterality, implants, and bundling constraints.

  2. 02

    Propose the supportable line

    Highest supportable coding, with ambiguity routed to the surgeon instead of guessed.

  3. 03

    Surgeon confirms

    The comparison target for the historical benchmark is the operating surgeon's billing-record coding — agreement, not clinical ground truth.

Evidence of value

Classed evidence. Not every number is cash.

MEASURED PERFORMANCE

96.1%

Adjudication concordance on billed + fully adjudicated engine-suggested lines

Of billed + fully adjudicated engine-suggested lines, 96.1% were paid. n=1,436 lines. Most cases were retrospective.

How this was measured

OpAgent Benchmark: paid rate among billed and fully adjudicated lines the engine suggested. Concordance with what was paid, not a causal claim that CodeIQ caused the payment.

Sample: n=1,436 lines

Do not imply CodeIQ caused the paid outcome. Most cases were retrospective.

Source: OpAgent Benchmark

MEASURED PERFORMANCE

1.2%

Strict coding/bundling false-positive rate

Strict coding/bundling false positives in the benchmark adjudication set.

How this was measured

OpAgent Benchmark: strict coding/bundling false-positive rate.

A false-positive rate is not an accuracy percentage and not an ROI figure.

Source: OpAgent Benchmark

MODELED OPPORTUNITY

$275,682

Tier A modeled opportunity

Modeled. Not recovered cash. Not caused-by-CodeIQ revenue.

How this was measured

OpAgent Benchmark: Tier A modeled opportunity on the adjudicated surgical set.

Modeled opportunity, not realized recovery.

Source: OpAgent Benchmark

MODELED OPPORTUNITY

$104,348

Tier B expected value

Modeled expected value. Not recovered cash.

How this was measured

OpAgent Benchmark: Tier B expected value on the adjudicated surgical set.

Modeled, not realized recovery.

Source: OpAgent Benchmark

MODELED OPPORTUNITY

~$509

Tier A modeled opportunity per adjudicated case

Tier A modeled opportunity divided across adjudicated cases.

How this was measured

OpAgent Benchmark: ~$509 Tier A per adjudicated case.

Modeled per-case opportunity, not recovered cash per case.

Source: OpAgent Benchmark

MEASURED PERFORMANCE

97.5%

Agreement with surgeon-coded historical cases

Weighted combined agreement across 359 held-out surgical cases spanning 5 procedure families, scored against the operating surgeon's own billing-record coding — a historical agreement benchmark, not clinical ground truth.

How this was measured

The validation dataset is 3,951 de-identified operative cases (2018–2026) from a single high-volume arthroplasty and trauma practice, held as a Limited Data Set. Cases were split with a fixed seed into few-shot, evaluation, and locked-test partitions; the locked test set was physically quarantined and never viewed during prompt iteration. The comparison target is the operating surgeon's billing-record CPT assignments. Prompt changes were accepted only when the evaluation set improved. 4.2% of cases self-flagged as ambiguous and routed to human review.

Sample: n=359 locked test set

Period: April 2026

Historical surgical coding agreement, not clinical ground truth, not adjudication, and not implied across other OpAgent agents.

Source: April 2026

Methodology / caveat

Adjudication concordance and strict coding/bundling false-positive rate are one proof class. Economic figures are modeled opportunity, not recovered cash. The 97.5% figure is historical surgical coding agreement against surgeon billing-record CPT assignments — not clinical ground truth, and not a claim that CodeIQ caused paid outcomes. Most cases were retrospective.

Full historical-agreement methodology →

Where it connects

The rest of OpAgent is already in the room.

  • Wrap

    Operative documentation that feeds the coding pass.

  • RadiologiQ

    Imaging that often sits in the same episode.

  • AdminQ

    The authorization packet that should have matched this case.

  • Proof

    Locked-test methodology for historical agreement.