Benchmarking and assessment
Benchmarking and evaluations are essential components in the assessment of quantitative aspects of Trustworthy Artificial Intelligence. These activities serve as a stepping stone towards responsible deployment of AI applications. Given the fast-paced development of methods and the wide variety of implementations in the modern AI landscape, research at Fraunhofer HHI focuses both on specific application-level evaluation and generically applicable evaluation frameworks. Our activities prioritize the requirements of the conformity assessments prescribed for high-risk applications as defined by the European AI Act, especially in the healthcare domain.
Human-Centered AI
To successfully implement AI applications, we need more than just good algorithms. It is also important to consider the needs of the people who use them. This is where Human-Centered AI comes in: Here we examine user requirements and adapt AI systems to make them understandable, reliable, and trustworthy. Together with partners from industry, healthcare, and the public sector, we develop user-friendly, interactive AI solutions that meet the demands of real-world practice.
Uncertainty Quantification Disentanglement
For the trustworthy deployment of AI, it is essential to make uncertainties in predictions transparent. Our Applied Machine Learning Group is successfully researching methods for quantifying and disentangling different sources of uncertainty, such as aleatoric (data-related) and epistemic (model-related) uncertainty. This distinction enables better risk assessment and helps make AI systems more reliable and practical. In this way, we contribute to ensuring that AI applications are not only powerful but also transparent and robust when used in real-world scenarios.
Publications
- Labarta, T., Hoang, N., Weitz, K., Samek, W., Lapuschkin, S., & Weber, L. (2025). See What I Mean? CUE: A Cognitive Model of Understanding Explanations. arXiv preprint arXiv:2506.14775. https://arxiv.org/abs/2506.14775
- Klein, S., Prajod, P., Weitz, K., Lavit Nicora, M., Tsovaltzi, D., & André, E. (2025). Communicating Through Avatars in Industry 5.0: A Focus Group Study on Human-Robot Collaboration. In Adjunct Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work (pp. 1-8). https://doi.org/10.1145/3707640.3731923
- Ma, J., Weicken, E., Pahde, F., Weitz, K., Lapuschkin, S., Samek, W., & Wiegand, T. (2025). Künstliche Intelligenz auf dem Prüfstand: Anforderungen, Qualitätskriterien und Prüfwerkzeuge für medizinische Anwendungen. Bundesgesundheitsblatt-Gesundheitsforschung-Gesundheitsschutz, 1-9. https://dx.doi.org/10.1007/s00103-025-04101-w
- Baur, S., Samek, W., & Ma, J. (2025, September). Benchmarking Uncertainty and its Disentanglement in multi-label Chest X-Ray Classification. In International Workshop on Uncertainty for Safe Utilization of Machine Learning in Medical Imaging (pp. 193-203). Cham: Springer Nature Switzerland. https://www.arxiv.org/abs/2508.04457
- Mertes, S., Huber, T., Karle, C., Weitz, K., Schlagowski, R., Conati, C., & André, E. (2024). Relevant irrelevance: generating alterfactual explanations for image classifiers. arXiv preprint arXiv:2405.05295. https://arxiv.org/abs/2405.05295