Seminar on Building and Testing AI Guardrails: Catching Problematic, Sensitive, and Harmful LLM Prompts (HWS 2026)
In this seminar, we study different approaches for building and testing LLMs guardrails. LLM guardrails that detect and correctly handle problematic, sensitive, and harmful prompts are essential for many application areas of LLMs, such as chatbots for the general public or protected groups. Accordingly the seminar is concerned with the topic of how we can test how well LLMs handle such prompts and how we can guardrail the LLMs.
Organization
- Organized by Lea Cohausz and Thilo Dieing
- Available for master and bachelor students
Goals
In this seminar, you will
- Explore relevant research topics/
papers in the realm of LLM guardrailing - Give an overview of your topic to your peers in a presentation
- Practically implement the guardrailing method / a testing method you presented
- Reflect on the implementation and your results
- Write a final research report
Schedule
tbd (first meeting will be scheduled via Doodle)
Requirements
- Good reading and presentation skills in English
- Proficiency in LaTeX
- Some programming experience
Evaluation
The grade consists of three parts:
- an individual presentation of a topic/
method - an individual presentation where you reflect on your implementation and results
- a final research report
