Can Generative Artificial Intelligence Revolutionize Medicine?
Generative artificial intelligence is gradually establishing itself as a major tool in the healthcare field. It enables the analysis of vast amounts of medical data, assists in clinical decision-making, and improves the quality of care. A recent analysis reviewed 24 studies to assess its use, performance, and environmental impact.
Generative artificial intelligence models fall into two main categories: those that process only text and those that combine multiple types of data, such as images and text. The former, often pre-trained on large general databases, are then fine-tuned for specific medical tasks. This fine-tuning step allows models to be adapted to precise needs, such as disease diagnosis or patient record analysis, while reducing the time and resources required compared to full training.
Multimodal models, though less common, prove particularly promising. They can simultaneously process medical images and texts, making them suitable for complex tasks such as interpreting scans while considering patient history. For example, some models have achieved accuracy rates of up to 99.9% in detecting ophthalmological diseases or classifying medical images.
Studies show that fine-tuned models often outperform pre-trained models in terms of accuracy, with results ranging from 70.6% to 99.9%. They are used for various tasks: disease classification, answering medical questions, or generating reports. However, their training remains highly resource-intensive. Few studies have measured the environmental impact of these technologies, but one revealed that training a model emitted 486 kilograms of CO₂, highlighting the often-overlooked carbon footprint of these tools.
Another challenge lies in the quality of the data used. Models often rely on public databases or electronic medical records, but there is a risk of data contamination. This can skew results, as the model might learn from information already present in test sets, overestimating its actual performance.
Conversational applications, such as assistants capable of answering medical questions or summarizing records, are the most widespread. They play a key role in interacting with patients and healthcare professionals. Yet, despite their potential, multimodal models remain underutilized, even though they could better leverage the diversity of medical data, such as images, texts, and recordings.
Performance evaluation often relies on standardized benchmarks, such as datasets designed to test models’ ability to answer medical questions or analyze images. These benchmarks allow for objective comparison of different systems. However, there are still no specific benchmarks for certain specialties, such as dentistry, which limits the evaluation of models in these areas.
Finally, the integration of these technologies into clinical practice raises ethical and privacy concerns. As medical data is sensitive, its use must be regulated to avoid any risk of patient identification. Although most studied models are still in the internal validation phase, their large-scale deployment will require rigorous testing to ensure their reliability and safety.
Documentary Sources / Document Base
Reference Report
DOI: https://doi.org/10.1186/s44398-026-00029-6
Title: Use of generative artificial intelligence models in healthcare – a scoping review
Journal: BMC Artificial Intelligence
Publisher: Springer Science and Business Media LLC
Authors: Ayesha Nooruddin; Abhishek Lal; Syed Muhammad Faizan Ahmed; Muhammad Huzaifa Ghori; Niha Adnan; Fahad Umer