Large Language Models in Interventional Cardiology: Current Evidence and Future Directions
© 2026 HMP Global. All Rights Reserved.
Any views and opinions expressed are those of the author(s) and/or participants and do not necessarily reflect the views, policy, or position of the Journal of Invasive Cardiology or HMP Global, their employees, and affiliates.
J INVASIVE CARDIOL 2026. doi:10.25270/jic/26.00176. Epub August 5, 2026.
Key Clinical Summary
- Contemporary large language models (LLM) show meaningful concordance with Heart Team decisions and may approach early-career operator performance; standardized, detailed prompts improve reliability more than model choice alone.
- Current LLMs cannot directly interpret angiographic, intracoronary imaging, or hemodynamic data, and clinically consequential hallucinations remain a major limitation.
- LLMs should remain adjunctive, with expert oversight and strict data confidentiality; current evidence best supports documentation, education, and patient communication, pending prospective workflow-integrated validation.
Abstract
Large language models (LLMs) have rapidly emerged as a transformative class of artificial intelligence systems capable of understanding and generating human-like text from vast corpora of clinical and scientific literature. While their use in general cardiology has been extensively reviewed, their specific role in interventional cardiology and within the catheterization laboratory (cath lab) environment remains less well characterized. This narrative review synthesizes the current evidence on LLM applications across the interventional cardiology workflow, including coronary revascularization decision-making, multidisciplinary Heart Team support, structural heart intervention planning, periprocedural communication, and acute cath lab decision support. Recent studies suggest that contemporary LLMs such as ChatGPT-4 (OpenAI), Claude (Anthropic), and Gemini (Alphabet, Inc.) can achieve clinically meaningful concordance with expert Heart Team recommendations for percutaneous coronary intervention vs coronary artery bypass grafting, and that their outputs may approach or even match those of early-career interventional cardiologists in simulated emergency scenarios. However, important limitations persist, including variable accuracy across prompt formats, susceptibility to hallucinations, lack of multimodal integration with angiographic and intravascular imaging, opacity of training sources, and unresolved medicolegal issues. The authors discuss the current evidence base, highlight promising avenues such as multimodal LLMs and retrieval-augmented generation tools that leverage current guidelines, and outline regulatory and ethical considerations that must accompany clinical adoption. LLMs are unlikely to replace interventional cardiologists in the foreseeable future, but they may meaningfully support decision-making, education, and patient communication in a domain where speed, complexity, and risk converge.
Introduction
Interventional cardiology has long been a fertile ground for technological innovation, from the introduction of balloon angioplasty in 1977 to the more recent expansion of transcatheter aortic valve replacement (TAVR) and transcatheter edge-to-edge repair (TEER) for mitral and tricuspid valve disease. In parallel, artificial intelligence (AI) has progressively entered the catheterization laboratory (cath lab) through deep learning applications in coronary angiography, intravascular imaging, computed tomography (CT)-derived fractional flow reserve, and patient-specific digital twins for procedural planning.1,2 Within this broader AI ecosystem, large language models (LLMs) represent a relatively recent but particularly disruptive development.
LLMs are transformer-based neural networks trained on massive corpora of text and capable of generating contextually relevant responses to natural language prompts. The public release of ChatGPT (OpenAI) in late 2022 catalyzed widespread interest in their potential medical applications, and the cardiology community has produced an expanding body of literature evaluating LLM performance in tasks ranging from board-style examinations to patient education and clinical documentation.3,4 Yet most existing systematic and narrative reviews have focused on general cardiology, preventive cardiology, or cardiovascular medicine broadly, with limited dedicated synthesis of LLM use in interventional cardiology.5,6
The cath lab represents one of the most demanding environments for clinical AI because decisions frequently occur under time pressure, uncertainty, and substantial procedural risk. Decisions are often time-sensitive, anatomy is complex and patient-specific, multiple guidelines may apply with subtle differences between European and North American recommendations, and procedural outcomes are highly operator-dependent. These features make the cath lab both a high-stakes test bed for LLMs and a domain in which the limitations of these models—hallucination, lack of multimodality, and opaque reasoning—are particularly consequential.7 The aim of this narrative review is to summarize the current evidence on LLM applications in interventional cardiology (Figure 1), to identify the main strengths and weaknesses of LLMs across the procedural workflow, and to outline future research priorities.
Methods
This narrative review was informed by a non-systematic PubMed/MEDLINE (National Institutes of Health) search performed through May 2026, using combinations of the terms "large language models," "artificial intelligence," "ChatGPT," "interventional cardiology," "percutaneous coronary intervention," and "Heart Team." English-language articles published between 2020 and 2026 were considered, supplemented by manual screening of reference lists and relevant editorials.
LLMs in Heart Team Decisions for Coronary Revascularization
One of the most extensively studied applications of LLMs in interventional cardiology is the support of multidisciplinary Heart Team (MDHT) decisions for coronary revascularization. Current guidelines from the European Society of Cardiology (ESC) and the American College of Cardiology (ACC)/American Heart Association (AHA)/Society for Cardiovascular Angiography and Interventions advocate for MDHT discussion in patients with complex coronary artery disease eligible for either percutaneous coronary intervention (PCI) or coronary artery bypass grafting (CABG).8 However, implementation of formal MDHT meetings varies substantially across institutions and geographic regions, creating an opportunity for LLM-based decision support.
In an early evaluation, Sudri et al analyzed 86 consecutive coronary angiography cases using ChatGPT-3.5 and ChatGPT-4 with case presentations including demographics, medical background, angiographic findings, and SYNTAX I and II scores.9 ChatGPT-4 demonstrated high concordance with MDHT decisions (accuracy 0.82, sensitivity 0.80, specificity 0.83, Cohen κ 0.59), whereas ChatGPT-3.5 performed substantially worse (accuracy 0.67, κ 0.12) (Figure 2A). Concordance was highest when the most detailed prompt format was used, and subgroup analysis showed the best agreement in patients with high SYNTAX scores in whom both the MDHT and ChatGPT-4 recommended CABG. An accompanying editorial highlighted that, for intermediate SYNTAX scores, the model tended to over-recommend CABG, illustrating the persistent challenge of integrating nonanatomic factors such as frailty, surgical risk, and patient preference.10
More recent work has extended these findings using newer models. In a 128-patient single-center analysis, ChatGPT-1 and ChatGPT-4 were compared with MDHT decisions across patients undergoing coronary angiography in 2024.11 ChatGPT-1 demonstrated superior sensitivity for CABG (82%, F1 score 82.4%), while ChatGPT-4 showed higher sensitivity for PCI (68.7%). Notably, neither model recommended optimal medical therapy alone in any case, illustrating an overreliance on procedural answers when anatomy is provided. Independent investigations using GPT-4 and Google’s Bard for PCI vs surgical aortic valve replacement decisions in TAVR-eligible patients have shown comparable patterns: LLMs perform reasonably when prompted as part of a simulated MDHT but their answers can vary substantially depending on user framing.12
Beyond binary PCI vs CABG decisions, more sophisticated frameworks are emerging. An ensemble framework combining 15 LLM versions was recently proposed for PCI decision-making in patients with moderate-to-severe coronary stenosis on coronary CT angiography, using hierarchical prompts and nested cross-validation.13 This study illustrates 2 important trends: first, that LLM performance is highly sensitive to data completeness and prompt structure; second, that ensemble strategies and dedicated evaluation pipelines may be necessary to translate LLMs from anecdotal performance to reliable clinical tools. Key representative studies are summarized in Table 1.
Overall, these studies share notable methodological strengths and weaknesses. Strengths include the use of real clinical cases rather than purely hypothetical vignettes and, in more recent work, ensemble modeling and cross-validated designs. However, most series are single-center, retrospective, and modest in size (n = 86-128), few include external validation, and the reference standard itself (MDHT decisions) is subject to inter-operator variability that may inflate or deflate apparent LLM concordance.
LLMs in Acute and Emergency Cath Lab Scenarios
The transition from chronic, planned decisions to acute cath lab scenarios represents a distinct challenge. Primary PCI for ST-elevation myocardial infarction (STEMI), management of cardiogenic shock, and handling of intraprocedural complications such as coronary perforation or no-reflow require synthesis of incomplete information under time pressure. Until recently, LLMs had rarely been tested in such settings.
A first head-to-head evaluation compared 7 LLMs (ChatGPT, Gemini [Alphabet, Inc.], LLAMA [Meta Platforms, Inc.], Qwen [Alibaba Group], Bing [Microsoft], Claude [Anthropic], DeepSeek [High-Flyer]) with 5 early-career interventional cardiologists across 12 challenging inferior myocardial infarction scenarios, with responses scored by 30 experienced interventional cardiologists.14 Physicians achieved an average reference score of 80.7 (95% confidence interval [CI], 76.3-85.0). Among the LLMs, ChatGPT ranked highest (87.4, 95% CI, 82.5-92.3), followed by Claude (80.8) and other models (Figure 2B). While these results suggest that contemporary LLMs may approach or even exceed early-career performance in highly structured vignettes, the authors emphasized that the scenarios were text-based and did not include angiographic images, hemodynamic tracings, or real-time deterioration—elements that fundamentally shape decision-making in the cath lab.
Several caveats apply. First, scoring by experienced operators is subjective, and inter-rater agreement on optimal management of complex STEMI cases is often imperfect. Second, simulated emergencies do not capture the cognitive load and team dynamics of an actual procedure. Third, LLMs have a known tendency to provide confident answers in the absence of certainty, which is particularly hazardous in time-sensitive situations.7 Nonetheless, these data lay the groundwork for future trials evaluating LLMs as bedside or in-lab cognitive aids, potentially via voice interfaces or integration with electronic health record dashboards.
LLMs in Structural Heart Interventions
Structural heart interventions represent one of the fastest-growing areas of interventional cardiology, with TAVR now considered first-line therapy for most patients older than 75 years with severe aortic stenosis, and transcatheter mitral and tricuspid procedures expanding into routine practice.15 Decision-making in this space is inherently multidisciplinary, balancing anatomy, surgical risk, life expectancy, and patient preference.
A pilot study published in JACC: Cardiovascular Interventions evaluated commercially available LLMs (ChatGPT-3.5, ChatGPT-4, and Bard) on their ability to classify patients into surgical vs transcatheter pathways using clinical vignettes constructed around ESC guideline criteria.12 The models performed reasonably well but exhibited limitations including unknown information sources and inconsistent handling of conflicting European vs North American recommendations. A subsequent study demonstrated that Heart Team simulation through prompt engineering substantially improved the accuracy of LLM recommendations for aortic stenosis, supporting the hypothesis that prompt design rather than raw model capability is often the limiting factor.16
Beyond decision-making, LLMs are beginning to be explored for procedural planning support by parsing preprocedural cardiac CT reports, prior imaging summaries, and clinical notes to highlight risk features such as low coronary heights, severely calcified annuli, or unfavorable mitral annular geometry.17 When combined with patient-specific computational simulations and digital twins, LLMs may eventually act as an interpretive layer that translates engineering outputs into clinically actionable language for the Heart Team.2
Beyond procedural planning, future structural heart applications may increasingly rely on multimodal systems capable of integrating text with imaging and physiological information. Structural interventions frequently require simultaneous interpretation of CT, echocardiography, fluoroscopy, and clinical variables rather than isolated data sources. Multimodal foundation models may therefore evolve into integrative platforms capable of synthesizing patient-specific anatomical and procedural information to support individualized treatment strategies and Heart Team discussions. A summary of LLM applications mapped across the interventional cardiology workflow is provided in Table 2.
Documentation, Education, and Patient Communication
Several non-procedural applications of LLMs are particularly relevant to interventional cardiology. First, LLMs can assist with structured data extraction from procedural reports. A recent study using both prompt-based and fine-tuned approaches on invasive coronary angiography and echocardiography reports demonstrated that cloud-based models achieved accuracies of 0.87 for culprit vessel identification and 1.0 for left ventricular function classification, with smaller locally hosted models achieving acceptable performance for less complex tasks.17 Such pipelines could substantially accelerate registry construction, quality improvement, and research.
Second, LLMs offer scalable support for patient education before and after invasive procedures. Existing reviews suggest that ChatGPT-3.5 can answer 91% of heart failure questions accurately, although readability often requires college-level comprehension.5 Patient-facing materials before PCI or TAVR could benefit from LLM-assisted generation tailored to literacy level and language, although safety evaluation remains essential.
Third, LLMs are increasingly explored as educational tools for trainees, including fellows in interventional cardiology. Their ability to generate case-based teaching scenarios, summarize landmark trials, and explain complex anatomic concepts in plain language has clear pedagogical value, though their tendency to hallucinate citations remains a concern in an academic environment.4 This concern extends beyond isolated citation errors: didactic inaccuracies conveyed to trainees carry safety implications analogous to those described elsewhere in this review, since misinformation absorbed during training may later be applied in clinical practice or propagated to future trainees, an indirect but clinically meaningful pathway for LLM-related harm.
Beyond isolated applications, near-term implementation within the catheterization laboratory will likely involve passive cognitive support systems integrated into existing clinical workflows rather than autonomous decision-making platforms. Potential applications may include summarization of prior procedural history, automated generation of structured reports, retrieval of guideline-based recommendations, and rapid access to relevant evidence during multidisciplinary discussions. Such approaches may reduce administrative burden and improve information accessibility while preserving physician oversight as the central component of clinical decision-making.
Limitations and Risks
It should be emphasized that LLM development is extremely rapid, and models such as ChatGPT-3.5 or early ChatGPT-4 versions that were evaluated in many studies discussed here are already outdated. The accuracy and concordance figures reported should therefore be read as historical benchmarks tied to specific model versions and time points, not as fixed estimates of current LLM capability, which has likely already improved substantially.
Despite encouraging early results, several limitations must be acknowledged before LLMs can be considered for routine clinical use in interventional cardiology. Hallucinations remain a fundamental failure mode and can include fabricated trial data, incorrect drug doses, and invented citations.18,19 In a high-acuity environment such as the cath lab, even a single confidently incorrect answer could contribute to patient harm.
A second concern is multimodality. Current general-purpose LLMs handle text well but integrate angiographic images, intravascular ultrasound, optical coherence tomography, and hemodynamic waveforms imperfectly. Several vendors are developing multimodal foundation models capable of jointly processing imaging and text, which may close this gap, but clinical-grade validation in interventional cardiology is still preliminary.1,17
Third, prompt sensitivity remains underappreciated. Across multiple studies, accuracy varies considerably with how clinical information is structured, the order in which findings are presented, and whether the model is prompted to participate in a simulated Heart Team.9,10,16 This variability complicates real-world deployment and demands rigorous prompt engineering and standardized input templates.
Benchmarking and explainability represent additional barriers to adoption. Unlike diagnostic imaging AI, which is often evaluated against standardized public benchmarks, LLM performance in interventional cardiology has been assessed using heterogeneous, largely single-institution vignette sets, limiting cross-study comparability. Standardized, guideline-anchored benchmark datasets would enable more rigorous comparison across models and versions. Explainability is a related concern: most LLMs act as black boxes, providing a recommendation without a transparent, verifiable reasoning chain tied to guideline criteria or patient anatomy. Chain-of-thought prompting, retrieval-augmented generation with source citation, and attention-based visualization may improve interpretability, but none has been validated specifically in this field.
Finally, ethical, legal, and regulatory considerations are non-trivial. LLMs are typically trained on data with unclear provenance and may reflect biases in published literature. Guideline conflicts between European and North American societies pose particular challenges for global deployment. Most LLMs currently used in published studies are not regulated as medical devices, raising questions about liability when their recommendations contribute to a clinical decision.19 Governance frameworks for LLMs in cardiovascular medicine have begun to emerge and should be central to future implementation efforts.20,21 Confidentiality and cybersecurity are equally important. Many widely used LLMs run on cloud platforms, and entering patient-identifiable clinical or angiographic details into public-facing interfaces risks breaching confidentiality and violating regulations such as the Health Insurance Portability and Accountability Act or General Data Protection Regulation. Institutions should prioritize locally hosted or enterprise-grade, encrypted LLM environments with audited data flows, alongside clear policies restricting public LLM use for identifiable patient data. Cybersecurity risks such as prompt injection and adversarial manipulation of clinical inputs remain largely unstudied and warrant dedicated investigation before LLMs are embedded into cath lab workflows. Key limitations and proposed mitigation strategies are summarized in Table 3.
Future Directions
Several directions are likely to shape the next phase of LLM integration into interventional cardiology (Figure 3). Retrieval-augmented generation may be particularly attractive for interventional cardiology because recommendations could be linked directly to contemporary ESC and ACC/AHA guidelines, landmark randomized trials, institutional protocols, and continuously updated evidence repositories. By grounding outputs in verifiable sources, these architectures may reduce hallucinations and improve transparency and reproducibility of generated recommendations. Domain-adapted models fine-tuned on interventional cardiology corpora (including procedural reports, multidisciplinary discussions, and high-quality educational material) may further improve performance. Multimodal LLMs integrating angiographic images, intracoronary imaging, and clinical text are likely to be transformative if validated rigorously, and could naturally connect with digital twin platforms used for personalized PCI and structural intervention planning.20,21
Beyond retrieval-augmented and multimodal architectures, several emerging paradigms merit attention. Agentic LLM systems that are capable of autonomously planning and executing multistep tasks such as retrieving prior imaging, cross-checking guideline criteria, and drafting a structured Heart Team summary, may extend LLM utility from passive question-answering to active workflow support. Federated learning, in which models are trained collaboratively across institutions without centralizing patient-level data, could enable domain-adapted models while addressing the data governance concerns above. Reinforcement learning, including reinforcement learning from human feedback and outcome-based reward modeling, may further align LLM recommendations with expert preference and, eventually, measured clinical outcomes rather than expert concordance alone. These directions remain largely unexplored in interventional cardiology.
Crucially, the field must move from cross-sectional vignette studies to prospective evaluations of clinical workflow integration. Trials comparing LLM-augmented Heart Team meetings with conventional meetings, randomized assessments of LLM-supported patient education, and pragmatic studies of in-lab decision support would generate evidence of real-world utility. Standardized reporting frameworks for LLM studies in cardiovascular medicine, analogous to TRIPOD-AI (Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis–Artificial Intelligence) for predictive models, will be essential to ensure transparency and comparability.
Conclusions
LLMs have rapidly entered the interventional cardiology landscape and demonstrate meaningful potential as cognitive aids for Heart Team revascularization decisions, structural heart intervention planning, documentation, education, and patient communication. Contemporary models can achieve substantial concordance with expert Heart Team recommendations under favorable prompt conditions and may approach early-career operator performance in simulated cath lab emergencies. However, the technology remains immature for autonomous clinical use. Hallucinations, prompt sensitivity, limited multimodal integration, and unresolved regulatory and ethical questions currently constrain deployment. In the medium term, LLMs are most likely to function as collaborative decision-support tools integrated within broader multimodal ecosystems combining imaging AI, digital twins, and evidence-based guidance. Their successful adoption, however, will likely depend less on larger models and more on safer systems, validated workflows, and human-centered oversight.
Affiliations and Disclosures
Dr Skalidis and Dr Simioni contributed equally.
Disclosures: The authors report no financial relationships or conflicts of interest regarding the content herein.
Data availability statement: Data available upon request to the corresponding author.
Address for Correspondence: Ioannis Skalidis, MD, PhD, HFR – Fribourg Cantonal Hospital and University, Fribourg, Switzerland. Email: Skalidis7@gmail.com
References
1. Samant S, Bakhos JJ, Wu W, et al. Artificial intelligence, computational simulations, and extended reality in cardiovascular interventions. JACC Cardiovasc Interv. 2023;16(20):2479-2497. doi:10.1016/j.jcin.2023.07.022
2. Skalidis I, Stalikas N, Collet C, et al. Digital twins and simulations in transcatheter coronary and structural heart interventions. Eur Heart J Digit Health. 2025;7(2):ztaf129. doi:10.1093/ehjdh/ztaf129
3. Wehbe RM. Charting the future of cardiology with large language model artificial intelligence. Nat Rev Cardiol. 2025;22(3):143-144. doi:10.1038/s41569-024-01105-y
4. Skalidis I, Cagnina A, Fournier S. Use of large language models for evidence-based cardiovascular medicine. Eur Heart J Digit Health. 2023;4(5):368-369. doi:10.1093/ehjdh/ztad041
5. Gendler M, N Nadkarni G, Sudri K, et al. Large language models in cardiology: systematic review. JMIR Cardio. 2026;10:e76734. doi:10.2196/76734
6. Ferreira Santos J, Dores H. Large language models in cardiovascular prevention: a narrative review and governance framework. Diagnostics (Basel). 2026;16(3):390. doi:10.3390/diagnostics16030390
7. Sengupta PP, Dey D, Davies RH, Duchateau N, Yanamala N. Challenges for augmenting intelligence in cardiac imaging. Lancet Digit Health. 2024;6(10):e739-e748. doi:10.1016/S2589-7500(24)00142-0
8. Lawton JS, Tamis-Holland JE, Bangalore S, et al. 2021 ACC/AHA/SCAI guideline for coronary artery revascularization: a report of the American College of Cardiology/American Heart Association joint committee on clinical practice guidelines. Circulation. 2022;145(3):e18-e114. doi:10.1161/CIR.0000000000001038
9. Sudri K, Motro-Feingold I, Ramon-Gonen R, et al. Enhancing coronary revascularization decisions: the promising role of large language models as a decision-support tool for multidisciplinary heart team. Circ Cardiovasc Interv. 2024;17(11):e014201. doi:10.1161/CIRCINTERVENTIONS.124.014201
10. Anyanwu EC, Fanaroff AC, Maddox TM. Large language models and revascularization decisions: the newest member of your multidisciplinary heart team? Circ Cardiovasc Interv. 2024;17(11):e014775. doi:10.1161/CIRCINTERVENTIONS.124.014775
11. Mola S, Yıldırım A, Gül EB. Artificial intelligence in cardiac treatment decision-making: an evaluation of the performance of ChatGPT versus the Heart Team in coronary revascularization. Rev Cardiovasc Med. 2025;26(8):38705. doi:10.31083/RCM38705
12. Salihu A, Meier D, Noirclerc N, et al. A study of ChatGPT in facilitating Heart Team decisions on severe aortic stenosis. EuroIntervention. 2024;20(8):e496-e503. doi:10.4244/EIJ-D-23-00643
13. Lin C, Lan Y, Zeng Z, et al. Evaluation of large language models in percutaneous coronary intervention decision-making. Front Cardiovasc Med. 2026;13:1690716. doi:10.3389/fcvm.2026.1690716
14. Cicek V, Zhao L, Tur Y, et al. AI in the hot seat: head-to-head comparison of large language models and cardiologists in emergency scenarios. Med Sci (Basel). 2026;14(1):33. doi:10.3390/medsci14010033
15. Leon MB, Smith CR, Mack MJ, et al; PARTNER 2 Investigators. Transcatheter or surgical aortic-valve replacement in intermediate-risk patients. N Engl J Med. 2016;374(17):1609-1620. doi:10.1056/NEJMoa1514616
16. Garin D, Cook S, Ferry C, et al. Improving large language models accuracy for aortic stenosis treatment via Heart Team simulation: a prompt design analysis. Eur Heart J Digit Health. 2025;6(4):665-674. doi:10.1093/ehjdh/ztaf068
17. van der Loo W, van der Valk V, van den Broek T, Atsma D, Staring M, Scherptong R. Large language models for structured cardiovascular data extraction: a foundation for scalable research and clinical applications. Eur Heart J Digit Health. 2025;7(2):ztaf127. doi:10.1093/ehjdh/ztaf127
18. Aminorroaya A, Biswas D, Pedroso AF, Khera R. Harnessing artificial intelligence for innovation in interventional cardiovascular care. J Soc Cardiovasc Angiogr Interv. 2025;4(3Part B):102562. doi:10.1016/j.jscai.2025.102562
19. Moosavi A, Huang S, Vahabi M, et al. Prospective human validation of artificial intelligence interventions in cardiology: a scoping review. JACC Adv. 2024;3(9):101202. doi:10.1016/j.jacadv.2024.101202
20. Skalidis I, Salihu A, Kachrimanidis I, et al. Meta-CathLab: a paradigm shift in interventional cardiology within the metaverse. Can J Cardiol. 2023;39(11):1549-1552. doi:10.1016/j.cjca.2023.08.030
21. Chen J, Liang Y, Ge J. Artificial intelligence large language models in cardiology. Rev Cardiovasc Med. 2025;26(7):39452. doi:10.31083/RCM39452


