AI for Medical Billers: Draft a Denial Appeal in Minutes

A medical biller's guide to using ChatGPT or Claude for denial appeals and code checks — PHI stripped first, every claim verified before you file.

If you work in medical billing or coding, you already know the feeling: another denial lands in your queue, the reason code is barely useful, and you’re about to spend the next forty minutes hunting down the payer’s policy language, matching it to the chart, and writing the same kind of appeal letter you’ve written a hundred times before. Search “AI for medical billing” right now and you’ll find a wall of software vendors selling six-figure platforms and a couple of academic papers — but not one plain answer to “can I just use the ChatGPT or Claude I already have?”

You can. Not blindly, and not without a few rules that matter more here than almost anywhere else AI gets used at work — because the thing you’re pasting is somebody’s protected health information. This is the walkthrough nobody’s written yet: how to actually use a consumer AI tool to draft a denial appeal and sanity-check a code, what the real accuracy numbers say (not the marketing numbers), and the one rule that comes before all of it.

What Just Changed: Payers Have to Tell You Why

Starting January 1, 2026, the rules around denials themselves changed in a way that makes this whole workflow more useful than it would have been a year ago. CMS finalized the Interoperability and Prior Authorization final rule (CMS-0057-F) back in January 2024, and its operational provisions — the ones that actually bite — took effect this year. Under the rule, Medicare Advantage organizations, state Medicaid/CHIP programs, and Medicaid/CHIP managed care plans must now issue prior-authorization decisions within 72 hours for expedited requests and 7 calendar days for standard requests, and, critically, they must give a specific reason for every denial — not a generic boilerplate code.

That last part matters enormously for what follows. CMS itself says the point is to stop providers from having to “reverse-engineer” a vague denial before they can even start an appeal. In plain terms: payers can no longer hide behind a one-line code and make you guess. You have more to work with than billers did a year ago — which is exactly what an AI tool needs to draft something useful.

CMS.gov fact sheet page for the Interoperability and Prior Authorization Final Rule CMS-0057-F, dated January 17, 2024 Source: CMS.gov — CMS-0057-F fact sheet

(Worth knowing the limits: this rule doesn’t apply to traditional Medicare fee-for-service or fully commercial payers outside the ACA exchanges, and a companion proposal extending similar rules to prescription-drug prior authorizations isn’t finalized yet. Check whether your payer is actually covered before assuming the new reason-code requirement applies.)

The Rule That Comes Before Everything Else: Strip PHI First

Before a single workflow step below, here’s the one rule that isn’t optional. Free and consumer-tier ChatGPT and Claude accounts are not HIPAA-eligible, and pasting real patient information into them is a compliance violation the moment you hit send — regardless of how secure the vendor’s infrastructure is.

Here’s exactly where each vendor draws the line, as of this writing:

  • OpenAI will sign a Business Associate Agreement (BAA) for ChatGPT Enterprise and the OpenAI API with zero-data-retention configured — but explicitly will not offer a BAA for ChatGPT Free, Plus, Team, or the newer consumer “ChatGPT Health” product.
  • Anthropic offers BAAs for Claude Enterprise (once your account’s Primary Owner formally activates HIPAA compliance in account settings) and the first-party API, plus HIPAA-eligible configurations through AWS Bedrock and Google Vertex AI — but explicitly not for Claude.ai’s free or Pro consumer tiers.

If your workplace hasn’t set up an Enterprise or API deployment with a signed BAA, the safe path is full de-identification before anything touches the AI tool. Under the HIPAA Privacy Rule’s “Safe Harbor” method (45 CFR §164.514(b)), that means removing all 18 of these identifier categories, per HHS’s own official guidance:

HHS.gov official guidance page: “Methods for De-identification of Protected Health Information” under the HIPAA Privacy Rule Source: HHS.gov — De-identification of PHI guidance

  1. Names
  2. Geographic subdivisions smaller than a state (street address, county, ZIP — except the first three digits if the area has 20,000+ people)
  3. All dates except year (birth, admission, discharge dates; ages over 89 aggregated to “90 or older”)
  4. Telephone numbers
  5. Fax numbers
  6. Email addresses
  7. Social Security numbers
  8. Medical record numbers
  9. Health plan beneficiary numbers
  10. Account numbers
  11. Certificate/license numbers
  12. Vehicle identifiers, including license plates
  13. Device identifiers and serial numbers
  14. Web URLs
  15. IP addresses
  16. Biometric identifiers
  17. Full-face photographs and comparable images
  18. Any other unique identifying number or code — including internal MRNs and claim numbers, which count as “unique identifying codes” under this rule

That last point trips people up constantly: your internal claim number and MRN count as identifiers too. A note with the name removed but the original claim number intact is not de-identified. Replace real identifiers with a placeholder (“Patient A,” “Claim #001”) before you paste anything.

The Walkthrough: From Denial to Filed Appeal

Here’s the actual workflow, built around how experienced billers already structure an appeal letter — because AI doesn’t replace that structure, it just fills it in faster once you give it the right pieces.

Step 1 — Pull the denial details and de-identify. From the Explanation of Benefits (EOB), grab the exact Claim Adjustment Reason Code (CARC) and Remittance Advice Remark Code (RARC) — for example, CO-50 (“not deemed a medical necessity”), CO-197 (missing precertification), or CO-16 (claim lacks information). Replace every identifier from the list above with a placeholder.

Step 2 — Give the AI the full skeleton, not just the denial code. Experienced billers build appeal letters around a consistent structure, and your prompt should hand the AI that same structure rather than asking it to freelance:

  • Identifiers block (placeholder patient/claim info + the CPT/HCPCS or ICD-10-CM code at issue)
  • The denial reason, quoted verbatim from the EOB (the CARC/RARC code and text)
  • An opening statement naming the procedure, date, denial reason, and the action requested (reprocessing or reconsideration)
  • The clinical/medical-necessity argument — diagnosis, presentation, prior conservative treatments tried, why this service was the appropriate next step
  • The payer’s own coverage policy — its Local Coverage Determination (LCD), National Coverage Determination (NCD), or internal medical policy number — matched criterion by criterion against the chart
  • Regulatory-rights language where relevant (ERISA’s right to full and fair review under 29 U.S.C. §1133, or Medicare appeal rights under 42 U.S.C. §1395ff)
  • A documentation/enclosures list
  • A closing request for reconsideration, with an offer of peer-to-peer physician review for medical-necessity denials

Expected result: a structured draft letter hitting every section above, in the payer’s expected format — not a finished, filed document.

Step 3 — Verify every cited policy line against the actual chart. This is the step that cannot be skipped, and the research explains exactly why in the next section.

Step 4 — Have a certified coder or biller review before it goes out. Treat the AI draft the way you’d treat a draft from a new hire: useful starting point, not a final answer.

A Worked Example

A claim for an MRI comes back denied CO-50, “not deemed a medical necessity.” The chart shows six weeks of physical therapy and an NSAID trial that didn’t resolve the patient’s symptoms before the MRI was ordered — exactly the kind of conservative-treatment history payers want to see cited.

You paste (de-identified): the CARC code and text, the procedure and diagnosis codes, a summary of the treatment history, and the payer’s LCD number if you have it. The AI returns a structured draft: an opening paragraph naming the denial and requesting reconsideration, a medical-necessity paragraph citing the six weeks of PT and the failed NSAID trial, a policy-citation paragraph referencing the LCD criteria, and a closing paragraph requesting peer-to-peer review.

You then do the step that actually determines whether this appeal wins: pull the real LCD text and confirm the cited criteria match what’s actually written in the policy, not what the AI assumed was in it. If it matches, the letter goes out on letterhead with the real identifiers restored. If it doesn’t, you fix the citation before anything gets filed.

A Second Worked Example: Coding as a Copilot, Not an Autopilot

Denial appeals aren’t the only place this workflow applies. Coders are also experimenting with AI as a shortlist tool for ICD-10-CM and CPT assignment — and the right mental model here is copilot, not autopilot, given the accuracy numbers in the next section.

Say a de-identified progress note describes a patient with type 2 diabetes, a foot ulcer, and documented peripheral neuropathy. Instead of asking the AI to “give me the code,” the more useful prompt asks it to shortlist 2-3 candidate ICD-10-CM codes with its reasoning for each — for instance, distinguishing between a diabetic foot ulcer code that requires documentation of the specific site and laterality versus a general neuropathy code. The AI’s job here is to surface the candidates and flag what documentation would confirm each one; your job as the coder is to verify against the actual encounter note and pick the code that’s actually supported by the documentation in front of you — not the one that sounds most complete.

This shortlist-then-verify pattern is specifically what keeps a coder in the loop where the accuracy data says they need to be. It also directly answers the anxiety question threaded through this whole topic: the workflow makes you the coder who got faster at spotting good and bad candidates, not the coder AI made obsolete.

What the Accuracy Data Actually Says (Not the Marketing)

This is where a lot of AI-in-billing content gets dishonest, so here’s the real picture, split by source type.

Appeal-drafting success rates (vendor-reported, treat with appropriate skepticism): RapidClaims’ RapidRecovery product reports a 55% overall appeal success rate (68% overturned within 30 days). Vellix’s DenialiQ, built on Claude, reports an 87% appeal win rate. These are vendor numbers from purpose-built RCM platforms with structured integrations — not what you’ll get pasting into a consumer chat window, and neither publishes an independently audited methodology.

Coding accuracy — the peer-reviewed numbers are much lower than vendor claims. A study published in NEJM AI tested general-purpose LLMs on medical code assignment and found GPT-4, the best performer, hit only 33.9% exact-match accuracy for ICD-10-CM codes, 45.9% for ICD-9-CM, and 49.8% for CPT codes — with models “frequently” producing imprecise or outright fabricated codes. A related benchmark found fabricated-code rates as high as 45.9% for Llama2-70b and 37.4% for Gemini Pro. A 2025 study in npj Health Systems showed domain fine-tuning could lift lab accuracy to 97.48% — but that number collapsed to 69.20% exact match (87.16% category match) once tested on real-world clinical notes, leading the authors to state plainly that “manual expert review remains essential.”

Compare that to the “96–98% AI coding accuracy” figures floating around some RCM vendor marketing — those numbers appear with no disclosed methodology. The honest summary: off-the-shelf consumer AI is not reliable enough for autonomous code assignment. Procedure codes (CPT) are markedly worse than diagnosis codes in every benchmark — roughly four times higher error rates in the NEJM AI study — which tracks with how much more clinical judgment procedure coding requires.

Where the newer CMS rule actually helps: with payers now required to give specific denial reasons instead of generic codes, appeals should genuinely get easier to draft — you’re no longer reverse-engineering the denial before you can even start.

Comparison: Manual vs. AI-Assisted vs. Purpose-Built RCM Software

ApproachCostSpeedAccuracy ceilingPHI risk
Fully manualStaff time onlySlowest — 30-40+ min per appealDepends entirely on biller experienceNone beyond normal workflow
Consumer ChatGPT/Claude, de-identifiedFree–$20/moFast draft, verification still required34–50% exact-match on codes unverified; appeal drafts need policy-citation checksRequires strict de-identification discipline every time
Enterprise/API tier with BAATeam/org pricingFast, PHI can be entered directlySame underlying model accuracy — verification still requiredCompliant, but still needs human review
Purpose-built RCM platform (vendor-integrated)Often five to six figures annuallyFastest — integrated with claims dataVendor-reported 55–87% appeal win rates (unaudited)Built for compliance, but locks you into the vendor

The honest takeaway from this table: the accuracy ceiling is set by the underlying model, not by which product wraps it — a consumer tool with careful verification and a six-figure platform both need a human checking the output. The real differences are speed, integration, and whether PHI can legally flow through the tool at all.

What This Means for You

If you’re a solo biller at a small practice: Start with a free Claude or ChatGPT account, strictly de-identified, for drafting appeal letters. The time savings on the drafting step alone — going from a blank page to a structured draft — is real, even with the verification step added back in.

If you’re a certified coder worried this replaces you: The data doesn’t support that fear. BLS projects positive job growth, not decline, for coding-adjacent roles through 2034 — see the job-market section below. Your value shifts toward verification and judgment, which the accuracy numbers above show AI genuinely cannot replace yet.

If you manage a revenue-cycle team: Push for an Enterprise or API-tier deployment with a signed BAA before your team starts using AI on real claims. The compliance gap between “free tier” and “BAA-covered tier” is not a technicality — it’s the line between compliant and non-compliant.

If you’re new to billing and coding: Learn the manual appeal-letter structure first. You need to be able to spot when an AI draft cites a policy wrong, and that requires knowing what right looks like.

If you’re evaluating a purpose-built RCM AI platform: Ask for the actual accuracy methodology, not the headline percentage. The gap between vendor-reported and peer-reviewed numbers in this space is large enough that “96% accurate” without a disclosed test set should be treated as a marketing claim, not a fact.

If your practice handles a high volume of Medicare Advantage denials: Worth knowing — industry analysis of federal MA data shows only about 11.5% of denials get appealed, but 80.7% of those appeals win. That gap between low appeal rates and high win rates suggests a lot of winnable appeals are simply never filed, which is exactly the bottleneck a faster drafting workflow can help close.

If you’re deciding whether to buy software or use what you have: Purpose-built platforms integrate with claims systems and typically cost significantly more. A consumer AI workflow costs nothing extra if you already have ChatGPT or Claude, but it demands more manual discipline — de-identification, verification, and no shortcuts on either.

Edge Cases and Troubleshooting

The AI cites a policy criterion that doesn’t actually appear in the LCD/NCD. This is the single most common and most dangerous failure mode. Always pull the real policy text and check the citation word-for-word — never trust a policy quote you haven’t verified against the source document.

The AI suggests a CPT code that seems plausible but isn’t supported by the note. This is the “excess/extraneous code” failure pattern documented in the research — a code with no grounding in the clinical documentation. If left unreviewed and billed, this is an upcoding risk with real audit exposure, not a harmless mistake.

You’re not sure whether your payer is covered by the new CMS-0057-F reason-code requirement. The rule covers Medicare Advantage, state Medicaid/CHIP fee-for-service and managed care, and ACA marketplace plans (with some exceptions) — it does not cover traditional Medicare fee-for-service or fully commercial payers outside the exchanges. Check which category your denial falls into before assuming you’re entitled to a specific reason.

A coworker wants to paste a real chart note into ChatGPT to save time. This is the moment to have the BAA conversation. Even a well-intentioned shortcut on a free-tier account is a compliance exposure the moment PHI is entered — there’s no “just this once” exception under the Privacy Rule.

The appeal keeps getting denied on resubmission. Check whether the medical-necessity argument is actually addressing the payer’s specific policy criteria, not just restating the clinical facts. A common weak point in AI-drafted appeals is a strong clinical narrative paired with a policy citation that doesn’t precisely match the payer’s actual coverage language.

You’re auditing AI-assisted claims after the fact and finding coding errors. Given that even fine-tuned, specialized models drop from 97% lab accuracy to 69% real-world accuracy, a documented human audit trail — not a trust-the-percentage approach — is the actual safeguard compliance analysts recommend.

Your organization is deciding whether AI-assisted claims need extra documentation. Given the False Claims Act exposure a fabricated code can create once it flows into a real claim, treat every AI-suggested code the same way you’d treat a code from a brand-new coder: it gets a second set of eyes before it’s billed.

What This Can’t Do

It can’t replace certified-coder verification. Every accuracy benchmark in this piece — 33.9% to 69.2% exact-match depending on conditions — makes the same point from different angles: unverified AI code assignment is not production-ready.

It can’t legally touch PHI on a consumer account. No amount of careful prompting changes the BAA status of ChatGPT Free/Plus/Team or Claude.ai free/Pro. This is a hard legal line, not a best practice suggestion.

It doesn’t know your specific payer’s unpublished internal criteria. LCDs and NCDs are public, but payers sometimes apply additional internal review standards that don’t appear in any document an AI tool can access. A clean AI-drafted appeal can still lose against criteria it never saw.

It won’t catch a fabricated code you don’t already know to look for. If you don’t have enough coding knowledge to recognize when a suggested code is wrong, you also won’t catch it when the AI is wrong — which is exactly why this is a copilot workflow for people who already know the field, not a replacement for learning it.

It doesn’t reduce your compliance exposure — it can increase it if used carelessly. A fabricated code that makes it into a filed claim is what compliance analysts call a “billing-integrity event,” with real False Claims Act and audit exposure. The speed AI adds only helps if the verification step keeps pace with it.

The Training Landscape Is Catching Up

Worth knowing if you’re weighing whether this is a fringe habit or a mainstream shift: AAPC — the field’s principal credentialing and continuing-education body — has moved past treating AI as an optional curiosity. Its published guidance through 2024–2025 consistently frames AI as a productivity tool that requires human oversight, explicitly encouraging coders to “embrace AI with the right attitude” rather than resist it. For 2026, that framing turned into structured continuing education: a dedicated “AI-Assisted Medical Coding and Billing: Applied Practice” course now sits in AAPC’s CE catalog alongside a foundational “Intro to AI” course.

AAPC’s “Intro to AI Concepts for Medical Billing and Coding” continuing-education course page Source: AAPC — Intro to AI Concepts for Medical Billing and Coding

That’s a meaningful signal. Continuing-education catalogs at a credentialing body don’t add a topic casually — it means AAPC now treats baseline AI literacy as part of what it takes to maintain certification, not a nice-to-have. If your employer hasn’t formalized an AI-use policy yet, pointing to AAPC’s own CE catalog is a straightforward way to make the case that this is now a standard professional competency, not a workaround.

Frequently Asked Questions

Can I use free ChatGPT or Claude for this at all? Yes, but only with patient information fully de-identified per the HIPAA Safe Harbor standard — all 18 identifier categories removed, including internal MRNs and claim numbers. Never paste real, identifiable PHI into a free or consumer-tier account.

Is AI actually taking medical billing and coding jobs? The data doesn’t support that. BLS projects medical records specialists to grow 7.1% (about 13,800 jobs) from 2024–2034, and names AI as a factor that moderates the growth rate — not one that causes job losses. Health information technologists and medical registrars are projected to grow 15% over the same period, “much faster than average.”

What does AAPC officially say about using AI in coding? AAPC, the field’s main credentialing body, treats AI as a productivity tool requiring human oversight, not a coder replacement, and has added a dedicated “AI-Assisted Medical Coding and Billing: Applied Practice” course to its 2026 continuing-education catalog — signaling AI literacy is now treated as a core competency for certification maintenance.

How accurate is AI at suggesting ICD-10 or CPT codes? Independently published, peer-reviewed benchmarks — not vendor marketing — show general-purpose LLMs achieving roughly 34–50% exact-match accuracy depending on the code type, with CPT/procedure codes performing notably worse than diagnosis codes. Domain-specific fine-tuned models do better but still drop sharply on real-world (vs. lab) documentation.

What’s the difference between Claude/ChatGPT’s free tier and their HIPAA-eligible tiers? Free and consumer-paid tiers (ChatGPT Plus/Team, Claude Pro) are not covered by a Business Associate Agreement from either vendor. Enterprise tiers and the first-party APIs (with the right configuration) are BAA-eligible. Only the BAA-covered tiers may legally receive real, identifiable PHI.

Do I need special software, or can I really just use what I already have? For drafting and code-checking with proper de-identification, a free or standard ChatGPT/Claude account works. For directly processing real PHI at scale, you need either an Enterprise/API deployment with a signed BAA or a purpose-built, compliance-designed RCM platform.

How has the CMS prior-authorization rule changed appeals in 2026? As of January 1, 2026, covered payers must issue prior-auth decisions within 72 hours (expedited) or 7 calendar days (standard) and must give a specific denial reason instead of a generic code — intended to stop providers from having to guess why a claim was denied before appealing.

What’s the single biggest mistake to avoid? Trusting an AI-cited policy quote without checking it against the real LCD/NCD text. This is the failure mode most likely to sink an otherwise well-argued appeal, and it’s fully preventable with one verification step.

Should I disclose to my employer that I’m using AI for this? Yes. Given the compliance stakes around PHI and the audit exposure around fabricated codes, this should be a documented, employer-sanctioned workflow with clear rules about de-identification and verification — not something done informally on the side.

Can AI help me figure out whether an appeal is even worth filing? Yes, and this may be the highest-leverage use of all. Industry analysis of Medicare Advantage data shows only about 11.5% of denials get appealed, yet 80.7% of those appeals succeed — a gap that suggests many winnable appeals simply never get filed because drafting one feels too time-consuming. A faster first-draft step can shift that math, especially on medical-necessity denials where the clinical argument is strong but nobody had time to write it up.

What if my organization uses a purpose-built AI coding or RCM platform instead of ChatGPT/Claude directly? The same verification discipline applies regardless of which product wraps the underlying model. Ask your vendor for their actual tested accuracy methodology and sample size — not just a headline percentage — since the gap between vendor-reported figures (55–87% appeal win rates) and independently published, peer-reviewed benchmarks (34–50% coding exact-match) in this space is large enough to matter.

The Bottom Line

The gap here is real and it’s wide open: search “AI for medical billing” today and you’ll find enterprise software vendors and academic papers, but nothing written in plain language for a working biller who wants to know whether ChatGPT or Claude can actually help — and how to do it without risking PHI or filing an appeal built on a citation that doesn’t exist. It can help, genuinely, on the drafting side. The CMS rule change means you have more to work with than a year ago. But the coding-accuracy numbers are lower than any vendor will tell you upfront, and the PHI rule is not negotiable.

The honest workflow: strip identifiers first, use AI to build the first draft against the real appeal-letter skeleton, then verify every policy citation and every suggested code against the source before anything goes out the door. That’s not a limitation of the tool — it’s the same discipline experienced billers already apply to their own first drafts. If you want to build this into a repeatable, PHI-safe routine for your whole team, FindSkill’s course library has AI Fundamentals for the tool basics and AI Fluency for Healthcare Workers for the profession-specific layer on top.

Sources

Build Real AI Skills

Step-by-step courses with quizzes and certificates for your resume