Begin: Can start immediately
First Supervisor: Prof. Dr. Heisenberg
Second Supervisor: Natasha Randall MSc.
Level: Master's thesis
Labels are critical for both training and evaluating supervised deep learning segmentation models, but labels are often inconsistent, noisy, and ambiguous. Many approaches have been developed to support training models on weak labels, but few currently exist to facilitate evaluating models on unreliable labels. This is a big problem, because even small amounts of errors in test labels can cause incorrect model evaluations, or wrong model selection decisions.
The image below shows a binary label (in blue) that segments flood waters from satellite images. However, the annotator has inconsistently included and excluded forest from the label in (a), incorrectly missed flood waters where the satellite image was obscured by clouds in (b), and defined ambiguous and noisy class boundaries of the flood in (c).
In order to solve this problem, methods need to be developed to support evaluating models even on unreliable segmentation labels. One new approach is "Adaptive Resolution Label Aggregation", or "ARLA", which dynamically adapts the resolution of both the label and the model prediction to the precision of the label or the level of label noise. Other examples of methods include uncertainty-based evaluation, tolerance border buffering, or soft labels.
The objective of this thesis is to perform a comprehensive study to validate the effectiveness and appropriateness of approaches to evaluate segmentation models on unreliable labels. The approaches should exploit the information encapsulated by a label while minimising the label error, extracting from the noise a clearer signal of a model's true performance and true error, without obscuring either the model's strengths or weaknesses.
- Research state-of-the-art methods for evaluating segmentation models on unreliable labels.
- Apply ARLA and other state-of-the-art methods within a large-scale test, using multiple datasets,
controlled noise experiments, and deep learning models.
- Two remote sensing (satellite) datasets with labels segmenting flooded regions and segmenting coffee & forest classes are currently available, but additional datasets may be sourced, including for non remote sensing data.
- Explore how best to adapt the resolution of ARLA (and the parameters of other methods) to the size of the label error, even when no reliable label is accessible.
- Analyse the effectiveness and appropriateness of the evaluation approaches, including their advantages and limitations.
- Optional: Develop your own method or extension of an existing method.
- Experience in coding with PyTorch and working with deep learning models.
- Willingness to learn more about model evaluation and new state-of-the-art approaches.
- The thesis will be carried out in English.
If you are interested, please get in contact via email.
A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA)
https://doi.org/10.48550/arXiv.2607.11214
