logo SBA

ETD

Archivio digitale delle tesi discusse presso l’Università di Pisa

Tesi etd-06052026-162156


Tipo di tesi
Tesi di laurea magistrale LM6
URN
etd-06052026-162156
Titolo
A comparative study of algorithms for sleep spindle detection
Dipartimento
RICERCA TRASLAZIONALE E DELLE NUOVE TECNOLOGIE IN MEDICINA E CHIRURGIA
Corso di studi
MEDICINA E CHIRURGIA
Relatori
.
relatore Prof. Faraguna, Ugo
Parole chiave
  • A7
  • ConceFT
  • MASS
  • spindles
  • SUMO
  • YASA
Data inizio appello
23/06/2026
Consultabilità
Non consultabile
Data di rilascio
23/06/2096
Riassunto (Inglese)
The automated detection of sleep spindles constitutes a central methodological and computational challenge in the study of sleep and its electroencephalographic (EEG) correlates. From the perspective of the research workflow, the precise extraction of these transient microstructures represents a preliminary, foundational processing step. Nevertheless, it establishes the necessary paradigm for properly evaluating spindles from a clinical and neurophysiological standpoint, enabling researchers to investigate their temporal dynamics, individual traits, and functional role.
Formally, a sleep spindle is a transient burst of neural oscillatory activity generated within the thalamic and corticothalamic networks. It occurs within the sigma frequency band, typically oscillating between 11 and 16 Hz, with a minimum continuous duration of 0.5 seconds and rarely exceeding 2 to 3 seconds. Serving as a defining electrophysiological feature of non-rapid eye movement (NREM) sleep, sleep spindles are predominantly observed during stage N2, but continue to occur at a reduced frequency throughout deep slow-wave sleep (N3).
To date, the clinical gold standard for the detection and characterization of sleep spindles strictly relies on visual scoring by expert neurophysiologists—a manual paradigm burdened by severe methodological limitations. This manual procedure is extremely time-consuming and largely inadequate for processing modern large-scale polysomnographic datasets. Furthermore, it is intrinsically affected by high intra- and inter-scorer variability. This fundamental subjectivity, leading to poor reproducibility of results, has served as the primary drive for the extensive development and proliferation of numerous automated detection algorithms.
In this thesis, four specific automated detection algorithms are analyzed, structurally implemented, and operationalized: A7, YASA, ConceFT, and SUMO. The first three belong to the family of parametric and deterministic methods. Specifically, A7 and YASA operate using moving threshold logics based on relative sigma power and correlation in the time domain, while ConceFT employs an advanced synchrosqueezing time-frequency transform to reallocate spectral energy. The fourth algorithm, SUMO, represents a machine learning topology based on deep convolutional neural networks. These frameworks were meticulously adapted, optimized, and tested on electroencephalographic data from the SS2 cohort of the Montreal Archive of Sleep Studies (MASS) dataset, and their detection performances were statistically evaluated against human annotations.
A central focus of this research was the in-depth investigation of the specific factors causing structural failures or highly unstable performances in these automated frameworks. The comparative analysis revealed that algorithmic efficacy is not uniquely determined by underlying mathematical complexity. Instead, detection success heavily depends on the intrinsic physical properties of the biological signal, particularly the signal-to-noise ratio (SNR) and phase non-stationarity driven by macro-artifacts. Moreover, the human ground-truth itself emerged as a critical variable: algorithmic performance metrics collapse or peak drastically depending on whether the reference expert adopts a highly restrictive or inclusive scoring topology, or exhibits annotation fatigue.
To overcome these physical and methodological limitations, an original pre-analytical "triage" pipeline for dynamic algorithm routing was formalized. By computing a deterministic Signal Quality Index (SQI) directly within the N2 sleep epochs, the pipeline maps the physical properties of the data, routing stable, high-SNR signals to the precision-oriented deep learning filters of SUMO, while diverting artifact-heavy EEGs toward the intrinsic spectral robustness of ConceFT. However, the introduction of a second human expert into the evaluation matrix revealed an unexpected paradigm shift. Ultimately, the findings of this thesis challenge the pursuit of a single "flawless" algorithm, demonstrating that the choice between a dynamic hybrid routing system and a robust universal spectral model fundamentally depends on which human cognitive style the machine is tasked to replicate.
Riassunto (Italiano)
La detezione automatica degli sleep spindles (fusi del sonno) costituisce un problema metodologico e computazionale centrale nello studio del sonno e dei suoi correlati elettroencefalografici (EEG). Nell'ambito delle pipeline di ricerca, l'estrazione precisa di queste microstrutture rappresenta una fase di pre-processing strutturale preliminare; tuttavia, essa costituisce la base ineludibile per poterne indagare le dinamiche temporali, i tratti individuali e il ruolo neurofisiologico in ambito clinico.
Da un punto di vista formale, uno sleep spindle è un burst transitorio di attività oscillatoria neurale, generato all'interno delle reti talamiche e cortico-talamiche. Si manifesta tipicamente all'interno della banda di frequenza sigma, con oscillazioni comprese tra gli 11 e i 16 Hz, presentando una durata minima continua di 0.5 secondi e superando raramente i 2 o 3 secondi. Costituendo una caratteristica elettrofisiologica distintiva del sonno NREM, gli spindles si osservano prevalentemente durante lo stadio N2, pur continuando a verificarsi, con frequenza ridotta, per tutta la durata del sonno ad onde lente (N3).
Ad oggi, il gold-standard clinico per la detezione e la caratterizzazione dei fusi del sonno si basa sull'individuazione visiva da parte di esperti neurofisiologi, un approccio manuale gravato da severi limiti metodologici. Oltre ad essere un processo estremamente time-consuming e inadatto all'analisi di dataset polisonnografici su larga scala, l'ispezione visiva è intrinsecamente affetta da un'altissima variabilità intra- e inter-scorer. Questa fondamentale soggettività, che porta a una scarsa riproducibilità dei risultati, è stata il drive primario per il fiorire di numerosi algoritmi di detezione automatica.
In questo lavoro di tesi sono stati analizzati, implementati e resi operativi quattro specifici algoritmi: A7, YASA, ConceFT e SUMO. I primi tre rientrano nella famiglia dei metodi parametrici e deterministici operanti nel dominio del tempo o tempo-frequenza, mentre il quarto (SUMO) rappresenta un'architettura di machine learning basata su reti neurali convoluzionali profonde. I framework sono stati ottimizzati e testati sui dati della coorte SS2 del dataset pubblico MASS (Montreal Archive of Sleep Studies), e le loro performance sono state statisticamente valutate rispetto alle annotazioni umane.
Un focus centrale della ricerca è stato lo studio approfondito dei fattori che determinano il crollo o l'instabilità prestazionale di questi sistemi automatizzati. L'analisi comparativa ha rivelato che il successo della detezione non è legato unicamente alla complessità matematica del modello, ma dipende in modo critico dal rapporto segnale-rumore (SNR) e dalla stazionarietà del tracciato biologico. Inoltre, è emerso come il ground-truth umano rappresenti esso stesso una variabile critica: le metriche algoritmiche subiscono fluttuazioni drastiche a seconda che l'esperto di riferimento adotti uno stile di marcatura inclusivo, altamente selettivo, o sia affetto da un calo di attenzione (annotation fatigue).
Per superare tali limitazioni fisiche e metodologiche, si è infine strutturata un'originale pipeline di "triage" pre-analitico. Calcolando un Indice di Qualità del Segnale (SQI) direttamente sulle epoche N2, il sistema mira a instradare dinamicamente i tracciati stabili verso la precisione delle reti neurali (SUMO), deviando i segnali ricchi di artefatti verso la robustezza spettrale di ConceFT. Tuttavia, l'introduzione di un secondo esperto clinico nella matrice di validazione ha rivelato un inaspettato ribaltamento di paradigma. I risultati di questa tesi sfidano la ricerca del singolo "algoritmo perfetto", dimostrando come la scelta tra una pipeline ibrida di routing dinamico e un modello matematico universale dipenda, in ultima analisi, da quale specifica mente umana si desidera che la macchina emuli.
File