International Journal of Artificial Intelligence in Medicine and Healthcare
OPEN ACCESS | Volume 1 - Issue 1 - 2026
ISSN No: - | Journal DOI: 10.61148/IJAIMH
Fenella Chadwick
Harvard University, Department of public health, Massachusetts Hall, Cambridge, United States.
Corresponding author: Prof. Dr. Fenella Chadwick, Harvard University, Department of public health, Massachusetts Hall, Cambridge, United States.
Received: September 10, 2026 | Accepted: September 26, 2026 | Published: September 28, 2026
Citation: Chadwick F., (2026) “Artificial Intelligence in Healthcare: Applications, Evidence, and Responsible Implementation-A Narrative Review” International Journal of Artificial Intelligence in Medicine and Healthcare, 1(2); DOI: 10.61148/IJAIMH/007.
Copyright: ©2026. Fenella Chadwick. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Artificial intelligence (AI) is increasingly used to analyze medical images, estimate risks, support clinical decisions, draft documentation, and assist research. Yet a model that performs well on a test set does not automatically improve care in a clinic. This narrative review surveys major health applications, the strength and limits of clinical evidence, and practical requirements for safe implementation. A 2025 review of radiology practice found encouraging but context-dependent diagnostic findings and relatively few studies conducted in live pathways. A 2025 meta-analysis of generative AI diagnosis found pooled accuracy of 52.1% across 83 studies, with substantial bias concerns; this figure must not be generalized to all clinical AI. WHO emphasizes autonomy, safety, transparency, accountability, equity, and sustainability. Taken together, the literature favors clearly scoped tasks, appropriate human oversight, local validation, and ongoing monitoring rather than indiscriminate automation. [1–42].
Artificial intelligence; healthcare; machine learning; clinical decision support; generative AI; patient safety; health equity
Healthcare generates images, clinical notes, laboratory results, vital signs, and operational records. AI methods may find patterns in these data and produce classifications, predictions, summaries, or recommendations. Applications vary greatly: an image-analysis device trained for a narrow task is not equivalent to a general-purpose language model. Likewise, a model’s ability to answer questions is different from its ability to change outcomes for actual patients. These distinctions matter for procurement, evaluation, and public trust. [43,75]
This is a selective narrative review, not an original experiment or systematic review. It synthesizes institutional guidance and peer-reviewed reviews available through September 2026. Sources were selected to represent imaging, generative AI, early clinical evaluation, regulatory oversight, and ethics. Because the field changes rapidly, findings refer to the studied technologies and settings rather than guaranteeing performance after later updates.
2. Technologies and Clinical Uses
Machine-learning models learn associations from data for specified outputs, such as the probability of disease. Deep-learning systems are often applied to imaging, signals, and text. Generative models create new text or other content from prompts; large multimodal models can process more than one input type. These categories overlap but have different failure modes. A useful first question is not whether a system “uses AI,” but what it is intended to do, for whom, with what inputs, and under whose responsibility. [1,2,5]
In radiology, AI may highlight suspected abnormalities, prioritize images for review, or serve as a second reader. Other use cases include retinal screening and image enhancement. The US Food and Drug Administration (FDA) describes authorized AI-enabled medical devices across multiple functions, including diabetic-retinopathy detection and image processing. Authorization concerns a particular device and intended use; it should not be interpreted as proof that every deployment improves outcomes in every hospital. Clinical validation must match local patient populations, scanners, workflow, and comparator practice. [76-90]
Models can estimate deterioration risk or identify patients for follow-up using electronic health record data. The output is a probability or alert rather than a diagnosis by itself. Even an accurate prediction is useful only if it arrives early enough, clinicians understand its meaning, and an appropriate response is available. Alert thresholds create trade-offs: increasing sensitivity can increase false alerts and workload. Evaluation must therefore examine both model discrimination and the actions the alert actually triggers. [4,6]
3. Generative AI and Operations
Generative AI may draft visit summaries, discharge instructions, patient messages, or translations. WHO identifies clerical work, clinical care, education, patient-guided use, and research among broad application areas for large multimodal models. Drafting can be attractive when clinicians face repetitive documentation, but outputs can omit facts, invent details, or misstate uncertainty. A clinician should verify the underlying record, medication details, and language before a draft becomes part of care. Patient-facing content also needs a route to a human professional when symptoms are urgent or the answer is unclear. [91-110]
Synthetic cases and question-answering tools may help trainees practice reasoning and feedback. Researchers may use computational tools to identify patterns or generate candidate hypotheses. These uses are not equivalent to validated patient care: an educational answer can be persuasive while wrong, and a proposed research lead requires independent verification. Institutional policies should distinguish low-risk drafting from tasks that directly determine triage, treatment, or diagnosis. [1,5]
Operational analytics can help forecast demand, review appointment patterns, or identify bottlenecks. However, managers should define a concrete decision before selecting an algorithm: for example, whether bed demand is forecast for staffing, discharge coordination, or emergency capacity. Relevant outcomes include waiting time, resource use, staff workload, and unintended inequities. The non-AI alternative such as a transparent scheduling rule provides an essential comparator. Not every problem requires a complex model[111-132].
4. What the Evidence Shows
A 2025 rapid systematic scoping review included 140 studies on AI in radiology practice. Of 53 quantitative effectiveness studies, only 23 evaluated AI in a live pathway; seven of those measured diagnostic accuracy. The review found potential improvements in sensitivity and specificity in some settings, with mixed results for time and cost. Some gains were more apparent among less experienced readers, while false positives and workflow disruptions remained concerns. These results support context-specific evaluation, not a blanket claim that AI always improves diagnostic care. [3]
A 2025 systematic review and meta-analysis of 83 studies estimated pooled diagnostic accuracy of 52.1% (95% confidence interval 47.0–57.1%) for the studied generative models and tasks. Models performed significantly worse than expert physicians overall in the reported comparison. The authors rated 63 of 83 studies as high risk of bias. The pooled number combines diverse tasks, models, and study designs; it is neither a universal accuracy rate for AI nor a prediction of performance at the bedside. [5]
Offline accuracy measures how a model performs against a reference label; clinical utility concerns whether its use improves decisions or outcomes safely. The DECIDE-AI guideline focuses on early, live evaluations of AI-supported decisions and asks investigators to report clinical utility, safety, human factors, and context. For randomized evaluation, SPIRIT-AI and CONSORT-AI address clearer reporting of intended use, inputs and outputs, human-AI interaction, and error cases. Reporting standards do not by themselves establish that a study is well designed or an intervention effective. [4,6,7]
5. Ethical, Legal, and Safety Issues
Training data may underrepresent certain age groups, disease severities, institutions, or languages. A model can appear accurate on average while making more errors in a subgroup. Appropriate testing requires meaningful subgroup analyses where sample sizes permit, review of missingness and access barriers, and clear plans when performance is inadequate. WHO’s guidance highlights inclusion and equity alongside human autonomy and patient safety. Equal access to an inaccurate tool is not an equity solution. [133-145]
AI development and deployment may involve sensitive records, third-party services, or cloud processing. A health organization should identify data flows, lawful uses, retention periods, access controls, and the contractual boundaries of any vendor. De-identification lowers some risks but may not remove all possibilities of linkage. Generative systems raise the additional risk that confidential details are entered into an unapproved service. Security review should cover the full workflow, including prompts, logs, interfaces, and incident response. Local privacy law and professional rules determine specific obligations. [146-166]
Patients and clinicians need to understand the system’s purpose and limitations. A probability score is not an explanation of cause, and a fluent generated answer does not confirm that sources were checked. Organizations should designate who verifies outputs, who can override a recommendation, how errors are reported, and when the tool is withdrawn. WHO recommends governance and post-release assessment for large-scale uses of multimodal models. Oversight cannot be reduced to a notice that “AI was used.” [1]
6. Regulation and Implementation
Regulatory rules depend on jurisdiction and intended function. In the United States, the FDA regulates medical devices, including qualifying AI-enabled devices, rather than “AI” in the abstract. The agency reported more than 1,600 AI-enabled medical devices authorized for US marketing as of September 2026. That count shows breadth of authorized products, not how many improve patient outcomes. A hospital must check the exact product, authorization, indications, version, and required user role before purchasing or expanding use. [2]
First define the clinical problem, patient population, user, comparator, and measurable outcome. Next assess data quality and the evidence for the specific product. Then test on local retrospective data and in a supervised “silent” phase where outputs do not affect care. A limited live pilot should document clinician interactions, failures, and patient-safety signals. Expansion should depend on predefined criteria, governance approval, training, and a clear rollback plan. DECIDE-AI is particularly relevant to reporting small-scale live evaluation. [4]
Performance can change when patient populations, clinical practice, devices, coding, or software versions change. Record version history and data dependencies; track error rates, false alerts, missing outputs, subgroup performance, user overrides, and patient outcomes. Investigate incidents rather than relying on an aggregate accuracy score. A designated clinical owner and technical owner should review results regularly. Cost analysis should include integration, support, staff time, rework from false positives, and replacement costs, not just purchase price. [2–4]
Consider a proposed AI tool that flags urgent chest images. A pilot could measure time to review for urgent and non-urgent cases, missed urgent findings, false-positive burden, and effects on radiologist workload. The hospital should decide in advance who responds to the flag and what happens if the system fails. This is an implementation example, not a claim that a particular product has achieved these results.
7. Discussion, Conclusion
The promise of AI in healthcare is task-specific. Imaging reviews show useful signals alongside limited live-pathway evidence and mixed operational effects. Studies of generative diagnosis underscore that compelling language should not be confused with expert clinical reliability. Standards for evaluation and governance converge on the same practical point: the intervention is the model plus its data, users, workflow, and monitoring process. [1,3–7]
This review was selective and did not follow a registered search protocol. It combines sources that study different technologies, patient groups, and healthcare systems. It does not calculate a pooled effect across AI applications, and US device regulation does not define requirements in other countries. Rapid model changes mean local assessment must be repeated after substantial updates.
AI can support healthcare when it solves a defined problem and is evaluated as part of a real care pathway. The safest route is to compare it with current practice, validate it locally, preserve meaningful human responsibility, protect data, assess equity, and monitor actual clinical and operational consequences after deployment.