logo SBA

ETD

Archivio digitale delle tesi discusse presso l’Università di Pisa

Tesi etd-08172026-130338


Tipo di tesi
Tesi di laurea magistrale
Autore
CERISE, SARA
URN
etd-08172026-130338
Titolo
Balancing Algorithmic Safety and Fundamental Rights: General-Purpose AI Guardrails under EU Law and the Case of Llama Guard
Dipartimento
GIURISPRUDENZA
Corso di studi
DIRITTO DELL'INNOVAZIONE PER L'IMPRESA E LE ISTITUZIONI
Relatori
.
relatore Giannini, Francesco
co-supervisore Favaro, Tamara
Parole chiave
  • AI guardrails
  • EU Law
  • Fundamental Rights
  • General-purpose AI
  • Llama Guard
Data inizio appello
14/09/2026
Consultabilità
Completa
Riassunto (Inglese)
The rapid proliferation of General-Purpose Artificial Intelligence (GPAI) models has prompted supranational regulatory interventions, most notably the European Union’s (EU) 2024 Artificial Intelligence Act (AI Act). Under Chapter V of the AI Act, providers of systemic-risk GPAI models must conduct model evaluations, adversarial testing, and risk mitigation. To demonstrate compliance with these stringent requirements, under threat of severe administrative fines, the AI industry has widely adopted automated, runtime safety classifiers known as "guardrails" to filter both user inputs and model outputs.
This thesis investigates the constitutional and technical tensions arising from the deployment of these automated guardrails. While designed to mitigate risks such as toxicity, illegal activities, and privacy violations, their implementation creates a structural asymmetry. To avoid liability, providers are incentivised to optimise classifiers with a highly conservative threshold calibration. This configuration causes systematic "over-blocking" (or over-refusal) of legitimate, benign queries touching on sensitive subjects, directly interfering with the right to freedom to receive information under Article 11 of the Charter of Fundamental Rights of the European Union (CFREU). Furthermore, the lack of dedicated, binding redress mechanisms in the AI Act for users affected by automated over-blocking creates a critical due process gap under the right to an effective remedy, established by Article 47 CFREU.
To contextualise this tension, this work conducts an interdisciplinary review of regulatory and technical literature. It evaluates the limitations of existing European frameworks alongside a technical deep dive into the state-of-the-art of AI safety classifiers. More specifically, the Meta Llama Guard family of open-weight safety models will be presented as a vertical case study, offering an analysis of the architecture, taxonomy, and vulnerabilities. The analysis concludes by identifying the intersection points between the technical and the legal sides in the balance between algorithmic safety and fundamental rights protection within the EU borders, outlining key open challenges for the future of the discipline.
Riassunto (Italiano)
File