LLMs Recognize TDM Problems but Reason Overconfidently on NTI Drugs
This benchmark shows current large language models interpret phenytoin and digoxin levels well but falter on pharmacokinetic reasoning and uncertainty acknowledgment, reinforcing that pharmacist oversight remains mandatory.
phenytoindigoxintherapeutic drug monitoringTDMnarrow therapeutic index drugslarge language modelsartificial intelligenceclinical decision supportChatGPTClaudeGeminiclinical pharmacybenchmark studyPharmacokinetics / TDM
Background Information
LLMs show promise for drug monitoring, yet caution remains due to potential reasoning errors.
Large language models (LLMs) are being considered as tools for therapeutic drug monitoring (TDM) but their reliability in clinical settings remains uncertain. Claude Sonnet 4.6 had the highest performance (94.9% ± 8.3%), outperforming ChatGPT 5.5 and Gemini 3.1 Pro, particularly in reasoning-dependent tasks. LLMs can identify TDM issues but struggle with reasoning tasks, highlighting the importance of pharmacist oversight to ensure patient safety. Pharmacist supervision is crucial when using LLMs for therapeutic drug monitoring to prevent safety concerns from overconfidence errors.
📊 Deep Analysis Available
Subscribers get the full breakdown: study design, outcomes, pharmacist implications, clinical pearls, GRADE evidence grade, and a downloadable PDF report.
Citation:Azmakan H, Azamakan T Performance of large language models on narrow therapeutic index drug monitoring: Implications for clinical pharmacy practice. Journal of the American Pharmacists Association : JAPhA. 2026:103524. doi:10.1016/j.japh.2026.103524
Subscribers receive access to a complete clinical analysis, available online and as a downloadable PDF.
Executive Summary
Study Design
Patient Population
Primary Outcomes
Secondary Outcomes & Safety Profile
Pharmacist Implications
Clinical Pearls
Limitations
Controversies & Evidence Gaps
Cost & Logistics
Want the full analysis?
Subscribe to PharmD Signal for unrestricted access to every deep analysis, plus PDF reports, GRADE evidence grades, and pharmacist implications for every article we publish.