Seeing and Modelling Humans
3D Consistent and Disentangled Gaussian Splatting GANs
We explore 3D Gaussian Splatting as an explicit representation for 3D-aware generative models. Starting from our Gaussian Splatting Decoder, which brings existing 3D-aware GANs into the 3DGS domain, CGS-GAN enables native, highly 3D-consistent Gaussian generation, while COSY takes a further step towards controllable generation and semantic editing.
Human Modelling and Animation
Our work on animatable virtual humans combines volumetric capture with geometry-, video-, and learning-based representations to reproduce highly realistic human appearance and motion. Starting from our hybrid approaches for making captured volumetric video directly animatable (read more here), we have moved towards learned pose-dependent representations of geometry and appearance for real-time animation and rendering (read more here).
Neural Head Modelling
We present a framework for the automatic creation of animatable human face models from calibrated multi-view data. Based on captured multi-view video footage, we learn a compact latent representation of facial expressions by training a variational auto-encoder on textured mesh sequences. Read more
Speech- and Video-driven Face Animation
We develop neural facial animation methods for expressive and controllable digital humans, driven by speech, text, or video. Our work spans speech-driven animation and video-based expression transfer, with 3D-Aware Latent-Space Reenactment taking a further step towards controllable animation by combining expression transfer with semantic editing.
Deep Fake Detection
We conduct research on both the generation as well as the detection of Deepfake content. In case of the former, we focus our research on the development of novel model architectures, aiming for an increase in percieved quality and photorealism of the fakes. Moreover, we study the models generalization to arbitrary identities. Read more.
Scenes, Structure and Motion
Robust Keypoints
We introduce RIPE and RIPE++ two innovative weakly-supervised training frameworks based on Reinforcement Learning. RIPE demonstrates that keypoint detection and description can be learned using only image pairs while RIPE++ shows that keypoint detection, description and matching can be learned from positive image pairs only: no depth, no camera pose, and no negative pairs.
3D Gaussian Splatting Compression
Our research on 3D Gaussian Splatting compression introduced Self-Organizing Gaussians (SOG), organizing Gaussian attributes for highly efficient image-based compression. Building on this concept, KISS-GS advances SOG into a modular compression pipeline, achieving state-of-the-art compression performance while enabling simple and efficient deployment.For a comprehensive overview of the field, see our survey 3DGS.zip.
6D Camera Localization in Unseen Environments
We present SPVLoc, a global indoor localization method that accurately determines the six-dimensional (6D) camera pose of a query image and requires minimal scene-specific prior knowledge and no scene-specific training. Read more
AI-Based Building Digitalization
We leverage AI and deep learning to automate and streamline building digitalization, reducing the need for labor-intensive manual processes. Our methods integrate and analyze information from diverse sources to support efficient monitoring and automation in the construction industry. Read more
Deep 6 DoF Object Detection and Tracking
For robust AR assistance in assembly tasks, CASAPose provides fast and accurate 6-DoF pose estimation of multiple objects in a single pass. Combining global pose estimation with precise real-time refinement, it achieves reliable registration while requiring only synthetic training data, avoiding costly manual data acquisition and annotation. Read more.
Computational Imaging and Video
Video-Based Bloodflow Analysis
The extraction of heart rate and other vital parameters from video recordings of a person has attracted much attention over the last years. In our research we examine the time differences between distinct spatial regions using remote photoplethysmography (rPPG) in order to extract the blood flow path through human skin tissue in the neck and face. Our generated blood flow path visualization corresponds to the physiologically defined path in the human body. Read more
Hyperspectral Imaging
Hyperspectral imaging records images of very many, closely spaced wavelengths, ranging from wavelengths in the ultraviolet to the long-wave infrared with applications in medicine, industrial imaging or agriculture. Reflectance characteristics can be used to derive information in the different wavelengths can be used to derive information on materials in industrial imaging, vegetation and plant health status in agriculture, or tissue characteristics in medicine. We develop capturing, calibration and data analysis techniques for HSI imaging. Read more
Learning and Inference
Anomaly Detection and Analysis
We develop deep learning based methods for the automatic detection and analysis of anomalies in images with a limited amount of training data. For example, we are analyzing and defining different types of damages that can be present in big structures and using methods based on deep learning to detect and localize them in images taken by unmanned vehicles. Some of the problems arising in this task are the unclear definition of what constitutes a damage or its exact extension, the difficulty of obtaining quality data and labeling it, the consequent lack of abundant data for training and the great variability of appearance of the targets to be detected, including the underrepresentation of some particular types. Read more
Deep Detection of Face Morphing Attacks
Facial recognition systems can easily be tricked such that they authenticate two different individuals with the same tampered reference image. We develop methods for fully automatic generation of this kind of tampered face images (face morphs) as well as methods to detected face morphs. Our face morph detection methods are based on semantic image content like highlights in the eyes or the shape and appearance of facial features. Read more
Augmented and Mixed Reality
Event-based Structured Light for Spatial AR
We present a method to estimate depth with a stereo system of a small laser projector and an event camera. We achieve real-time performance on a laptop CPU for spatial AR. Read more
Tracking for Projector-Camera Systems
We enable dynamic projection mapping on 3d objects to augment envrinments with interactive additional information. Our method establishes a distortion free projection by first analyzing and then correcting the distortion of the projection in a closed loop. For this purpose, an optical flow-based model is extended to the geometry of a projector-camera unit. Adaptive edge images arer used in order to reach a high invariande to illumination changes. Read more
Older Research Projects
For a list of previous research projects, see here.