logo SBA

ETD

Archivio digitale delle tesi discusse presso l’Università di Pisa

Tesi etd-07022026-202132


Tipo di tesi
Tesi di laurea magistrale
URN
etd-07022026-202132
Titolo
Entropy-Based Identification of Significant Patterns in Natural and Artificial Reconstruction Systems
Dipartimento
FISICA
Corso di studi
FISICA
Relatori
.
relatore Prof. Punzi, Giovanni
correlatore Dott.ssa Castellotti, Serena
Parole chiave
  • Entropy
  • HEP
  • Natural vision
  • topK
Data inizio appello
20/07/2026
Consultabilità
Non consultabile
Data di rilascio
20/07/2096
Riassunto (Inglese)
A problem occurring in several data processing systems is that of an input flow far too large to keep, that must be reduced, in real time and under tight resource limits, to the fraction that is actually worth retaining. Two examples where this occurs with severity are the real-time event selection in high-energy physics, and the early processing of visual stimuli in the brain. In high-intensity physics experiments this is the role of the trigger, which selects the events of physical interest by extracting just enough information from each to decide whether to keep it. The typical approach to this is based on exploiting the fact that the target is known in advance: one knows from knowledge of the physics and the structure of the detectors which parts of the information matter. The harder and more general version of the problem arises when it is not clear beforehand what is relevant, when the data must be reduced without a predefined notion of which configurations are worth keeping. Early vision processing in the brain faces exactly this version of the problem: the visual system needs to compress an enormous input stream with no externally given criterion of relevance, yet it is far from indifferent to what it keeps, and it is capable of adapting to changing conditions. What determines this selectivity and its adaptation to input is a question that can be examined by formulating computational models of what counts as a “significant pattern” worth preserving in the input data.One conceivable strategy is to preserve the set of patterns that carry the most information at the output, under constraints on total bandwidth, and on the number of patterns the system can memorize (Del Viva et al., 2013). This is a well-defined optimization problem with an essentially unique solution that can be worked out in detail, leading to a pattern set determined by the probability distribution of the patterns in the input. This strategy has already been experimentally tested in the past, on 3×3 pixel spatial patterns in natural images: reducing the bandwidth and the number of retained patterns, human discrimination of images filtered with the optimally selected patterns has been compared against images filtered with patterns of equal information content but not selected for efficient use of bandwidth and number. The significant difference in discrimination between the two suggests that this kind of frequency-based selection reproduces at least some simple features of vision perception data. However, extending it to more realistic higher-dimensional patterns (such as adding the temporal dimension in vision, or applying the approach to the data streams of particle detectors) requires measuring the probability of occurrence of every possible configuration in much larger spaces, a number that quickly becomes intractable as the patterns grow more complex. A further difficulty is that these data arrive in real time and their statistics can drift, so the relevant probabilities are not those over the whole history, but those within a moving time window. This is the computational challenge this thesis aims to address.
Before any selection criterion can be applied, the probabilities themselves need to be estimated from the stream, but avoiding the burden of enumerating every possible configuration a-priori. To this purpose, I have adapted the Floating Top-k algorithm of Song et al. (2019), a streaming method often utilized in social-media and network-traffic analysis, that maintains a compact, continuously updated estimate of the most frequent patterns over a sliding window, with bounded memory. The selection is then carried out on the patterns it retains.
The algorithm is analysed in detail and modified in two respects: its frequency estimator is replaced by more accurate alternatives derived from extreme-value theory and maximum-likelihood estimation, and an aggregation scheme is added to stabilise the ranking where the separation between retained and discarded patterns is most fragile.
In natural vision, the maximum-entropy model is extended from static spatial patterns to dynamic spatio-temporal ones; the selected features, now including motion structure that the static case cannot represent, are used to build sketches of natural videos, and psychophysical experiments shows that observers recognise the originals from these sketches well above chance, and more reliably than from sketches of equal information content built on non-selected patterns, consistently with the model.
The developed framework is then repurposed to real-time data processing in high-energy physics, by applying it to real data from the LHCb silicon pixel vertex detector (VELO) aiming at identifying significant patterns without any prior information (“unsupervised”). At the level of single hits, detecting which positions recur most often exposes over-active (noisy) channels and follows their drift over time, providing a new approach to automated detector monitoring and correction. At the multi-cluster level, a translation-invariant encoding lets geometrically coherent combinations across layers stand out above the combinatorial background. A geometric analysis then shows that the configurations detected by the constrained entropy principle are compatible with genuine particle trajectories within the detector. This indicates that this first-principles approach is capable of retrieving necessary information about the alignment and condition of a detector under realistic conditions in a challenging data-processing environment.
The same first-principles logic underlies both applications: relevance is inferred from the entropy of the input rather than imposed from outside, and in each it is enough to recover the patterns that an independent criterion then confirms as significant.
Riassunto (Italiano)
File