Thesis etd-06102025-173843 |
Link copiato negli appunti
Thesis type
Tesi di laurea magistrale LM6
URN
etd-06102025-173843
Thesis title
Benchmarking Large Language Models with and without RAG on expert clinical knowledge: the case of the Italian Association of Sleep Medicine expert examination
Department
RICERCA TRASLAZIONALE E DELLE NUOVE TECNOLOGIE IN MEDICINA E CHIRURGIA
Course of study
MEDICINA E CHIRURGIA
Supervisors
.
relatore Prof. Faraguna, Ugo
Keywords
- IA and medicine
- LLM and medicine
- Sleep medicine and artificial intelligence
Graduation session start date
15/07/2025
Availability
Withheld
Release date
15/07/2095
Abstract (Inglese)
Abstract (Italiano)
The application of Large Language Models (LLMs) in medical contexts is rapidly expanding. However, their reliability in addressing highly specialized clinical topics remains a key issue. This study evaluates the accuracy of various LLMs in answering 50 multiple-choice questions from the official test of the Italian Association of Sleep Medicine (AIMS), a mandatory step to obtain the national qualification of “Sleep Disorder Specialist” in Italy.
Each of the 50 questions with the corresponding 4 options were asked by the experimenter to each LLM. The experimenter noted the correctness of the question in an Excel sheet. A score of 1 was assigned to the correct questions and 0 to each incorrect question. For each LLM, to verify the consistency, the same question was asked 5 times.
The models tested included Llama 3.2 3B, Llama 3.3 70B, Llama 3.2 3B that was perfectionated by Retrieval-Augmented Generation (RAG), Gemini 2.0 Flash, and NotebookLLM (which also utilizes RAG). RAG allows you to read documents such as txt or PDFs file so that LLM can check them before giving the answer. The uploaded documents are books and papers recommended by AIMS to pass their test.
Each of the 50 questions with the corresponding 4 options were asked by the experimenter to each LLM. The experimenter noted the correctness of the question in an Excel sheet. A score of 1 was assigned to the correct questions and 0 to each incorrect question. For each LLM, to verify the consistency, the same question was asked 5 times.
The models tested included Llama 3.2 3B, Llama 3.3 70B, Llama 3.2 3B that was perfectionated by Retrieval-Augmented Generation (RAG), Gemini 2.0 Flash, and NotebookLLM (which also utilizes RAG). RAG allows you to read documents such as txt or PDFs file so that LLM can check them before giving the answer. The uploaded documents are books and papers recommended by AIMS to pass their test.
File
| Nome file | Dimensione |
|---|---|
The thesis is not available. |
|