Computer Vision and Pattern Recognition (CVPR)

The CVPR research group focuses on developing the state-of-the-art models, theories and core technologies for facial recognition, biometric system security medical informatics, medical image processing and video surveillance. The research group not only aims to develop enabling technologies of all these applications, but also to address the public concerns of security and privacy issues.


Faculty Involved:



Funded Research and Consultancy Projects in the Past Few Years:


Systems and Algorithms for 3D Digital Media Content Creation with Applications in Journalism and Art Tech
Staff Prof. CHEN, Jie
Objectives
  • To develop a high-precision light field photogrammetry system for multi-view, multi-light-source 3D data capture.
  • To design calibration and reconstruction algorithms for accurate static and dynamic 3D digital asset generation.
  • To develop high-fidelity human head avatar reconstruction and animation methods from video and audio inputs. More
  • To build generative AI algorithms for studio-quality digital human, garment, and prop asset creation. - To develop robust watermarking and IP protection techniques for AI-generated 3D digital assets.
Grant Guangdong and Hong Kong Universities “1+1+1”Joint Research Collaboration Scheme

Automatic Ultrasound Video Summarization for Improved Diagnosis with Simplified Scanning Protocols
Staff Prof. GUO, Xiaoqing
Objectives
  • First, we propose a standard plane recognition and generation model that reconstructs high-quality standard planes from keyframes identified near anatomical landmarks. This will enable clinicians to extract critical diagnostic information from low-quality sweeps.
  • Second, we introduce a 3D reconstruction model that synthesizes a volumetric representation of the anatomy by estimating the 3D pose of video frames. This comprehensive 3D view will simplify diagnosis and enhance interpretability.
  • Third, we propose a memory-efficient video report generation model that dynamically compresses keyframes into a memory bank for long-term video analysis, providing clinicians with concise, informative textual summaries. More
  • Lastly, we introduce a human-interactive video question localization model that allows clinicians to quickly locate and review diagnostically relevant frames by integrating human input into the AI workflow.
Grant Early Career Scheme (ECS)

Bridging the Performance Gap in USV Operations: Confronting Water Surface Reflection Challenges
Staff Prof. WAN, Renjie
Objectives
  • Develop a base prior with 3D particle structures capable of modelling water surface reflections across varying spatiotemporal positions in 3D, forming a firm foundation for further analysis.
  • Conduct the In-Prior handling to reduce the visual ambiguities and occlusion effects caused by the water surface reflections.
  • Create a mechanism to seamlessly integrate the 3D information in base prior into existing USV systems, ensuring improved functionality of USVs. More
  • Evaluate the performance of the proposed system using curated datasets, testing its effectiveness in USVs’ systems to validate its robustness in real-world environments.
Grant Early Career Scheme (ECS)

Synthesising Lesion Images with Pixel-level Labels Without Manual Annotations for Medical Image Analysis Model Learning
Staff Prof. YUEN, Pong Chi
Objectives
  • To investigate and develop an automatic lesion image annotation framework using unlabelled medical images and corresponding medical reports. The proposed framework will require neither any manually annotated lesion images (segmentation maps) nor manual intervention.
  • To investigate and develop an annotation-free pseudo-label-guided lesion image synthesis framework from unlabelled medical images and corresponding medical reports. The framework should be able to generate realistic lesion images with pixel-level labels for model training. The synthesised lesion images will not contain patients’ health information. More
  • To investigate and develop an annotation-free lesion image synthesis framework and method that learns the diverse variations of lesion-region masks and texture. Based on the currently available medical lesion knowledge, the proposed method will be able to generate diverse lesion images with different lesion attributes, such as locations, sizes, shapes and textures. The synthesiser will also be able to generate a balanced lesion image dataset for deep model learning.
Grant General Research Fund (GRF)

A Chain-of-Thought Prompting Framework for Responsible Visual Decision-Making
Staff Prof. ZHOU, Kaiyang
Objectives
  • Develop efficient prompt learning methods, which can turn a large language model into a multimodal agent capable of handling both image and text modalities.
  • Develop alignment strategies, which aim to align the language model's output with the visual input such that the generated text describes key features/attributes of input images.
  • Develop RL methods to train the aligned vision-language model, allowing the model to perform chain-of-thought reasoning and produce executable predictions.
Grant Early Career Scheme (ECS)