Pharmacists: What ChatGPT Is Safe For (and What It Isn't)

New 2026 studies show ChatGPT misses most dangerous drug interactions. What pharmacists can safely use it for — and the one thing to never trust it with.

A patient in Lagos watched her pharmacist quietly type her lab results into ChatGPT mid-consult, trying to talk her out of a diagnosis her doctor had already confirmed. A family member almost got an extra, nonexistent daily dose of a post-surgical medication because an AI summary “sounded more confident than the prescription leaflet.” And in a nursing home, a relative’s life was likely saved when someone ran two prescriptions through ChatGPT and caught a severely dangerous combination that two doctors — working from records that hadn’t synced — never saw coming. All three of these happened in 2026. All three are true. And they point to the exact same conclusion pharmacy’s own research now backs up with hard numbers: ChatGPT is genuinely useful for some parts of your job, and genuinely dangerous for one specific part of it, and the two categories don’t overlap the way most people assume.

This isn’t a “ban AI” post or a “trust AI” post. It’s the line — backed by five separate 2025–2026 peer-reviewed studies — between where ChatGPT earns its keep in a pharmacy workflow and where it needs to stay far away from the final call.

What just changed: the evidence caught up with the anecdotes

Pharmacists have been quietly using ChatGPT for over a year now — drafting counseling scripts, summarizing medication therapy management (MTM) visits, translating patient leaflets. What’s new in 2026 is that the research finally caught up with specific numbers, and the numbers are worse than most people using the tool casually would guess.

The headline study: researchers led by Erin Ready, a clinical associate professor in pharmaceutical sciences at the University of British Columbia, tested ChatGPT against real antiretroviral (HIV) drug-interaction pairs, checked against the HIV/HCV Medication Guide and the University of Liverpool’s HIV Drug Interaction Checker — two of the field’s gold-standard references. Published ahead of print in the journal AIDS on April 13, 2026, the result: ChatGPT correctly classified only 40.4% of the interaction pairs. Sensitivity was 46.0%. Specificity was 29.0%. And the negative predictive value — how much you can trust a “no interaction” answer — was just 20.9%. Worse, 60.7% of the model’s incorrect answers were false negatives: cases where a real, dangerous interaction existed and ChatGPT said everything was fine.

That last number is the one that should stop you. A false positive wastes your time double-checking something that was actually fine. A false negative is the one that reaches a patient.

The American Society of Health-System Pharmacists (ASHP) homepage, the professional body whose August 2025 AI statement sets the current standard cited throughout this research Source: ASHP — American Society of Health-System Pharmacists

And the HIV study isn’t an outlier. It’s the pattern:

  • A 311-patient ICU cohort study (published January 2026 in the Journal of the American Pharmacists Association) found ChatGPT-4.1 and ChatGPT-5 identified only 8 and 10 potential drug-drug interactions per patient respectively, against a reference-tool average of 22. Worse, when the same patient’s medication list was re-queried, ChatGPT-5 only gave a consistent answer 68% of the time.
  • A separate 101-patient ICU study (Frontiers in Pharmacology, accepted August 2026) found ChatGPT-5 had 69.8% sensitivity and just 36.7% negative predictive value against UpToDate’s Drug Interaction Checker — while also generating six alerts per patient on average versus UpToDate’s median of zero, meaning it both missed real signals and buried pharmacists in false alarms.
  • At ASHP’s 2025 Midyear meeting, a study testing ChatGPT against complex simulated pharmacy cases found it was error-free on only 48% of them — sometimes missing interactions that were genuinely there, sometimes flagging interactions that didn’t exist at all.

Five studies. Five different patient populations, prompts, and reference tools. All pointing the same direction: general-purpose ChatGPT is not a reliable standalone drug-interaction checker, in any clinical context tested so far.

The one exception that matters, and why it’s misleading

Here’s the twist that makes this genuinely confusing if you’ve had a good experience with the tool: ChatGPT can and does catch real errors sometimes — and when it does, it’s memorable enough to go viral.

In late August 2026, a patient posted that ChatGPT flagged a prescription for hydralazine (a blood-pressure medication) when the intended drug was hydroxyzine (an antihistamine) — two look-alike, sound-alike names for completely different drugs. The clinic confirmed the error and corrected it. A different case, also documented on X that same week, involved a nursing-home relative on two severely incompatible medications that two different doctors had prescribed without knowing about each other’s orders, due to a records-sync failure — caught before administration, by someone running the list through ChatGPT.

Both of these are real, and both matter. But notice what they have in common: they’re look-alike/sound-alike name confusion and records-sync gaps — pattern-matching problems, where a model that’s good at spotting “these two words are suspiciously similar” or “this combination looks unusual” adds genuine value as a second set of eyes. That’s a fundamentally different task from correctly classifying the clinical severity, mechanism, and management of a known drug-drug interaction — which is exactly the task the five studies above show it performing at 40–70% accuracy, well below what any pharmacy would accept from a human colleague.

The honest read: ChatGPT is a decent gut-check for “does this look wrong,” and an unreliable authority for “is this actually safe.” Confusing the two is how a useful habit turns into a dangerous one.

U.S. Pharmacist coverage of an ASHP Midyear study finding ChatGPT’s drug-related answers were frequently incomplete or wrong, establishing the pattern later 2025-2026 studies confirmed Source: U.S. Pharmacist — “Pharmacists Should Be Wary of ChatGPT Information on Drugs”

The Ontario audit: the same failure mode, one layer up the chain

If the pharmacy-specific studies feel like an isolated concern, a parallel 2026 finding from outside pharmacy makes the pattern impossible to dismiss as a fluke. A government auditor in Ontario, Canada tested AI medical scribes — tools used to turn doctor-patient conversations into structured clinical notes — across the systems approved for use in the province. The results, reported in May 2026: the AI agents routinely hallucinated treatments that were never discussed, and in roughly 60% of tested cases, recorded a different drug entirely than the one actually prescribed. Almost half the tested notes contained fabricated information of some kind.

Notice the mechanism: this isn’t a scribe tool failing at drug-interaction math. It’s failing at the more basic task of accurately transcribing what a drug even is. That’s directly relevant to pharmacy, because a drug-interaction check is only as good as the medication list feeding it — and if an upstream AI tool has already swapped a drug name during documentation, no amount of careful interaction-checking downstream catches an error that was baked in before the pharmacist ever saw the chart. The lesson generalizes past this one audit: wherever AI touches a medication name, dose, or drug identity — scribing, summarizing, interaction-checking — the same name-confusion and fabrication failure modes show up. Treat any AI-touched medication data as unverified at every step it passes through, not just the last one.

How pharmacists worldwide are actually using it (and how nervous they are about it)

Survey data from three different countries in 2025-2026 shows a strikingly consistent split: real enthusiasm for AI’s potential, paired with real caution about handing it anything patient-specific.

In Egypt, a cross-sectional study of 428 pharmacists (214 of them in community practice) found 73.6% recognized ChatGPT’s anticipated benefits for pharmacy practice — but 65.9% disagreed or stayed neutral on using it to analyze patient medical information and give individualized advice, 67.1% worried specifically about erroneous answers, and 61.7% raised confidentiality concerns. When asked about actual use, nearly 30% said they’d used it to check drug-drug interactions specifically — a sizable minority already doing the exact thing the accuracy studies above say not to do, unsupervised.

In Jordan, two 2023 surveys of community pharmacists found comparable ambivalence: in a 221-pharmacist sample, 48.4% were willing to integrate ChatGPT into practice and 47.5% rated its perceived benefit highly — but more than 70% said it lacked the human and ethical judgment pharmacy work requires. A second Jordanian survey found strong support (69.9%) for AI-generated educational material specifically, while flagging the same accuracy, privacy, and legal concerns.

In the United States, a 2025 survey of pharmacy-practice preceptors across Indiana, Illinois, and Michigan found only 30.4% had used an AI chatbot at all, though 51.5% planned to start or continue. Where they had used it, the top reported tasks were summarizing information (46.4%), writing recommendation letters (32.9%), and looking up general disease-state information (32.5%) — notably, none of the top three uses were interaction-checking. The study’s own conclusion: most respondents had not used a chatbot and were unlikely to base patient-care decisions on one.

Read together, the international picture is reassuring in one specific way and concerning in another. Reassuring: the pharmacists most engaged with the research literature are already gravitating toward exactly the safe uses this article recommends — drafting, summarizing, education — and staying skeptical of autonomous clinical judgment. Concerning: a meaningful minority (that 30% figure from Egypt) are already using general ChatGPT for interaction-checking specifically, which is precisely the use case the accuracy data says is least reliable. If you recognize that pattern in your own pharmacy or your team’s habits, this is the moment to redirect it — not because the tool has no value, but because it has real value sitting one task-category away from where people are currently pointing it.

The walkthrough: the workflow pharmacists are actually using safely

Here’s what the community-practice evidence — from a 2025 Netherlands community-pharmacy survey, a McKesson ideaShare 2026 case study, and Singapore’s NUH MedBot hospital counseling program — shows pharmacists doing right now, without incident:

Step 1 — Strip every patient identifier before anything gets typed. Name, MRN, date of birth, room number — gone. This isn’t optional; it’s the first rule in every safe workflow the research surfaced.

Step 2 — Give the model the actual label or guideline text, not just a drug name. “Summarize this in plain language” applied to an official package insert produces a far more reliable draft than asking ChatGPT to recall drug information from memory. Feed it source material; don’t ask it to be the source.

Step 3 — Ask for a draft at a specific reading level or note format. “Write this at a 6th-grade reading level” or “format this as a SOAP note with suggested CPT codes” is the actual prompt pattern behind the most common real-world use: a pharmacist at McKesson’s 2026 ideaShare conference described pasting a de-identified post-MTM-visit summary into an LLM and getting back a structured SOAP note draft in place of 15–20 minutes of manual documentation.

Step 4 — Check every interaction and dose claim in a real, licensed interaction database. Lexicomp, Micromedex, UpToDate’s Drug Interaction Checker, or your pharmacy’s own clinical decision-support tool — not ChatGPT’s memory. This step is non-negotiable and it’s the one every study above says can’t be skipped.

Step 5 — Personalize and counsel face-to-face. The AI-drafted script is a starting point, not the final word to the patient. Singapore’s NUH MedBot program — which runs AI-assisted counseling for a locked list of 66 pre-approved medications and reports saving roughly 28 hours of pharmacist time per month — still routes every single script through pharmacist review before it reaches a patient. Speed, not autonomy, is the win.

A worked example, end to end. A community pharmacist finishes a 40-minute MTM consult with a patient on eight medications, three of which were recently adjusted. Instead of spending the next 20 minutes typing a comprehensive medication review from scratch, they paste a de-identified summary — medication list, adherence notes, patient-reported symptoms, no names or dates — into ChatGPT with the prompt: “Draft a SOAP-format MTM note from this summary, including a personal medication record and a medication-related action plan. Flag anything that needs my review before I finalize it.” They get a structured draft back in under a minute. They then open Lexicomp, verify every interaction flag the draft implicitly touches, correct two details the AI got slightly wrong about timing, add their own clinical judgment on a borderline renal-dosing question, and sign the note. Total time: roughly 8 minutes instead of 20 — with the interaction-checking step exactly as rigorous as it would have been without AI in the loop at all.

A second worked example, for the counter-level catch scenario. A patient hands over two new prescriptions from two different prescribers who didn’t know about each other’s orders — the exact scenario behind the nursing-home near-miss referenced earlier. A pharmacist, uneasy about the combination but not immediately placing why, types both drug names and doses into ChatGPT as a same-shift sanity check: “Do these two medications, at these doses, raise any obvious safety concerns together?” The response flags a serious interaction. The pharmacist does not treat that flag as confirmation — they immediately pull up Lexicomp and verify independently, because the same tool that just caught something real is, per the studies above, wrong in the reassuring direction roughly 1 time in 5. The Lexicomp check confirms the interaction is real and serious. The pharmacist holds the second prescription and calls both prescribers. Total elapsed time from “something feels off” to “verified and escalated”: under five minutes — with the AI functioning as a prompt to look closer, not as the authority that decided the outcome.

Comparison: what to trust ChatGPT for, and what to route around it

Task2026 evidence-backed verdictWhy
Patient counseling scripts, plain-language rewrites✅ Safe with pharmacist reviewASHP identifies this as a supported use; Netherlands survey found writing assistance “valuable and time-saving”
MTM note / SOAP note drafting✅ Safe with pharmacist reviewMcKesson ideaShare 2026 case study; structure and formatting, not clinical decisions
Medication leaflet translation⚠️ Safe only with bilingual pharmacist reviewReported as a current use case, but no published translation-accuracy data exists yet — treat every translation as unverified until checked
Spotting look-alike/sound-alike drug-name confusion⚠️ Useful as a second check, not a first lineReal catches documented (hydralazine/hydroxyzine), but this is pattern-matching, not interaction classification
Drug-drug interaction classification (severity, mechanism, management)❌ Do not trust standalone40.4% correct classification (HIV study); 48% error-free on complex cases (ASHP Midyear 2025); consistently underperforms Lexicomp/Micromedex/UpToDate across five 2025-26 studies
Individualized dosing recommendations (renal/hepatic impairment, narrow-therapeutic-index drugs)❌ Do not trust standaloneRequires patient-specific clinical judgment the studies show current models don’t reliably capture

What this means for you

If you’re a community pharmacist handling high volume: The counseling-script and leaflet-drafting workflow above is your highest-value, lowest-risk starting point. Start there before touching anything interaction-related.

If you’re in hospital or health-system pharmacy doing MTM or clinical documentation: The SOAP-note drafting pattern from McKesson’s 2026 case study is worth piloting on a small scale — de-identify first, draft with AI, verify every clinical claim against your existing interaction database exactly as you already do, sign as usual.

If you’re working in ICU, transplant, oncology, or any high-acuity specialty: The two ICU studies above are your specific evidence — ChatGPT under-detected interactions by roughly half to two-thirds compared to reference tools in these exact settings. Treat this as confirmation to keep your existing verification protocol fully intact, not as a reason to add an AI shortcut.

If you’re managing HIV/antiretroviral therapy specifically: The Ready et al. study is a direct warning about this drug class. ART interactions depend on pharmacokinetic detail (boosters, enzyme induction, therapeutic drug monitoring) that a fluent-sounding AI answer can mask rather than reveal. Check Liverpool’s HIV Drug Interaction Checker first, every time.

If you’re a pharmacy student or new practitioner: Learning to spot the difference between “this looks fine” and “this is verified fine” is now a core professional skill, the same way learning to distrust an unfamiliar drug name on sight already is. Build the habit early: AI drafts, you verify, always in that order.

If you’re a pharmacy manager or owner writing an AI-use policy: ASHP’s August 2025 statement — the field’s current formal position — is explicit that generative AI should augment, not replace, pharmacist judgment, and that any AI tool touching medication safety needs validation, transparency, and ongoing surveillance before deployment. Use the five-step workflow above as the skeleton of your written policy, and make the “always check a licensed interaction database” step mandatory, not optional guidance.

If you’re a pharmacy technician working alongside licensed pharmacists: The AI-drafting workflow isn’t reserved for licensed staff only — a technician can just as safely use ChatGPT to draft a counseling handout or format a discharge medication schedule. What matters is that the interaction-checking and final sign-off step stays with the pharmacist, exactly as your existing scope-of-practice rules already require regardless of AI involvement.

If you’re a telehealth or mail-order pharmacist without in-person patient contact: The “counsel face-to-face” step in the workflow above becomes a phone or video call instead, but it still has to happen — an AI-drafted counseling script read verbatim over a chat window, with no verification conversation, loses the one safeguard (a real-time chance to catch confusion or ask a clarifying question) that makes the drafting step safe in the first place.

If you’re a patient reading this because you’ve used ChatGPT to double-check your own prescriptions: The lookalike-name catches are real and worth doing as a gut-check. But treat a “no interaction found” answer from ChatGPT as exactly zero reassurance — the studies above show that’s precisely where it fails most often. Call your pharmacist with any real concern; don’t let a confident-sounding AI answer substitute for the conversation.

Edge cases and troubleshooting

“ChatGPT gave me a confident, detailed-sounding answer about a drug interaction — should I trust the detail level?” No. The research is explicit that fluency doesn’t correlate with accuracy here. The Ready et al. study’s incorrect answers weren’t vague hedges — they were confidently wrong classifications 60.7% of the time in the false-negative direction specifically.

“I asked ChatGPT to check interactions on a full medication list, and it only flagged one or two things — is that reassuring?” The opposite, actually. The 311-patient ICU study found ChatGPT-4.1 and -5 identified 8 and 10 interactions per patient on average against a reference-tool count of 22 — meaning a short list from ChatGPT is more likely evidence of under-detection than a genuinely clean chart.

“A colleague swears by using ChatGPT for quick interaction checks at the counter and hasn’t had a problem yet.” Absence of an observed error isn’t the same as absence of risk — the false-negative failure mode by definition doesn’t announce itself until something goes wrong. The five studies above exist precisely because “seems fine so far” isn’t a safety standard.

“What about the newer, more specialized medical AI tools — are they different from general ChatGPT?” The studies above tested general-purpose ChatGPT-4.1 and ChatGPT-5 specifically, not purpose-built clinical decision-support systems. Domain-specific tools validated against clinical outcomes are a different category with their own (also worth-scrutinizing) evidence base — don’t assume the findings transfer, but don’t assume they don’t, either, without checking that tool’s own published validation.

“I used ChatGPT to translate a medication leaflet into Spanish for a patient — how do I know it’s accurate?” There’s currently no published translation-accuracy data for this specific task. Treat every AI-translated leaflet as a first draft requiring review by a bilingual pharmacist or qualified medical translator before it reaches a patient — the same rule that applies to any other AI-drafted clinical content.

“My pharmacy doesn’t have a formal AI-use policy yet — where do I start?” ASHP’s August 2025 statement is the current field-standard reference. At minimum, your policy needs: mandatory de-identification before any AI tool sees patient information, a hard rule that AI-drafted interaction or dosing claims get checked against a licensed database before acting on them, and a named pharmacist responsible for periodic review of how the tool is actually being used in practice.

“Is it ever appropriate to skip the licensed-database check?” Not according to any of the evidence reviewed here. Even in the perioperative study that showed ChatGPT’s best documented performance — 95% sensitivity on 80 known interactions in a controlled, anesthesiologist-reviewed vignette set — the authors were explicit that this doesn’t generalize to unscripted, real-world polypharmacy. Controlled-test performance and real-patient reliability are not the same claim.

What it can’t do

It can’t replace a licensed drug-interaction database. Every study reviewed here, across five separate research groups and patient populations, reaches the same conclusion: general ChatGPT underperforms Lexicomp, Micromedex, and UpToDate at the specific task of interaction classification.

It can’t give you a trustworthy “no interaction” answer. The negative predictive value numbers (20.9% in the HIV study, 36.7% in one ICU study) mean a clean-sounding AI summary is specifically the least reliable output type it produces.

It can’t account for the individual patient in front of you. Renal function, hepatic status, pregnancy, drug-food interactions, adherence history, and the dozens of other patient-specific factors that inform real clinical judgment aren’t reliably captured by a general-purpose model working from a text prompt.

It can’t substitute for institutional AI governance. ASHP’s own position is explicit: safe deployment requires validation, transparency, and ongoing surveillance — none of which a pharmacist using consumer ChatGPT on their own initiative can provide for their organization.

It can’t be trusted to repeat itself. The 68% reproducibility rate found in the ICU cohort study means the same medication list, queried twice, can return different interaction assessments — a property no clinical safety tool should have.

It can’t tell you which of its own answers to double-check. This is the failure mode underneath all the others: the model doesn’t flag its own uncertainty in a way that reliably tracks its actual accuracy. A confidently-worded false negative reads exactly like a confidently-worded correct answer, which is why the workflow above treats every interaction claim as needing verification rather than trying to guess which ones are the risky 40-60%.

FAQ

Is ChatGPT ever appropriate for checking drug interactions at all? As a secondary, gut-check layer for spotting obviously wrong things like look-alike drug names — yes, cautiously. As a primary or standalone interaction-safety check — no, not according to any of the 2025-26 studies reviewed here.

Why do some pharmacists report good experiences with it, then? Selection bias in what gets shared publicly (a dramatic catch is memorable; a silent miss usually isn’t discovered), plus real value in the tasks ChatGPT genuinely is good at — drafting, formatting, plain-language rewriting — that get conflated with the interaction-checking task it’s bad at.

Does ChatGPT-5 perform better than earlier versions on drug interactions? The ICU cohort study tested both ChatGPT-4.1 and ChatGPT-5 and found neither outperformed reference tools, though ChatGPT-5 generated more alerts overall — meaning better recall of possible issues but also more false alarms to sort through, not a clean accuracy win.

What does ASHP actually say pharmacists should do? Its August 2025 statement calls for the pharmacy workforce to lead AI validation and implementation, evaluate any tool for accuracy and transparency before use, and confine fully automated AI to tasks where performance genuinely matches a human counterpart — which, per the evidence above, interaction classification does not yet meet.

Is there a safe way to use AI for medication questions if I’m a patient, not a pharmacist? Use it as a prompt to ask your pharmacist a more specific question, never as a replacement for asking them. If an AI tool flags something concerning, call and confirm — don’t act on the AI answer alone in either direction (neither “it’s fine” nor “this is dangerous”).

What’s the single most important habit to build from this research? Treat every AI-generated interaction or dosing claim as unverified until checked against a licensed database — the same instinctive skepticism you’d apply to an unfamiliar claim from any non-authoritative source, applied consistently, every time, with no exceptions for time pressure.

Do the accuracy problems apply equally to drug-food and drug-supplement interactions, or just drug-drug? All the studies cited above tested drug-drug interactions specifically. Supplement and food interactions (like the Vitamin A over-dosing case referenced earlier) involve even less standardized reference data across AI training sources, so if anything, treat AI-generated answers in that category with at least as much caution — there’s no published accuracy study yet establishing even a baseline to compare against.

How does this compare to using AI for non-medication clinical questions, like general disease-state information? The US preceptor survey found disease-state look-ups were among the more common and less controversial uses, likely because a wrong answer there is more easily cross-checked against general knowledge than a wrong interaction-severity classification is. The same core rule still applies — verify against an authoritative source before it informs a patient conversation — but the stakes and error-detection difficulty differ.

Where can I read the actual studies instead of a summary? The Ready et al. HIV study, the ICU cohort studies, and ASHP’s full statement are linked in the Sources section below — worth reading directly if you’re building a department policy, since the exact methodology matters more than any summary, including this one.

The bottom line

The 2026 evidence draws a clean line through what looked, a year ago, like a single undifferentiated question of “should pharmacists use ChatGPT.” The answer splits cleanly in two: for drafting, formatting, translation-with-review, and plain-language rewriting, the tool is a genuine time-saver that working pharmacists are already using safely, at scale, with documented time savings. For classifying whether two drugs are dangerous together, it fails at a rate no pharmacy would accept from a human colleague — 40 to 70% accuracy depending on the study, with the failures skewing toward the false negatives that matter most. Know which task you’re doing before you open the chat window, and never let the fluency of an AI-drafted answer stand in for the licensed database check your training already tells you to run. If you want a structured way to build exactly this workflow — the safe drafting uses, the mandatory verification step, and where the line sits — FindSkill’s AI for Pharmacists course walks through it lesson by lesson, with a certificate at the end.

Sources

Build Real AI Skills

Step-by-step courses with quizzes and certificates for your resume