Facial Reenactment in a Monocular Video Stream for Speaker Verification and Its Applications (Yiu-ming Cheung et al.)

Let us consider a scenario: Alice takes Bob to court for a criminal offence. In court, Alice presents the judge with a video in which Bob admits his guilt. Can the judge accept this apparently incriminating evidence and immediately convict Bob? Before answering this question, let us introduce an emerging technique: facial reenactment (FR). The use of FR in a monocular video stream, e.g. Face2face (Thies et al. 2019) which is a quintessential algorithm for FR, has received increasing attention in the domain of computer vision and pattern recognition. The goal of FR is to reenact a monocular target video stream based on the mouth movements of a source actor who was also recorded in a monocular video. This technique animates the mouth movements of the target speaker and re-renders the manipulated output video in a photo-realistic fashion. This technique has many applications, such as in animation and entertainment programs. However, this technique could also result in serious security problems, as it allows a speaker in a video to be reenacted and thus faked. Therefore, without an effective approach for FR detection, all videos are untrustworthy. Consequently, regarding the above question, the judge would be unable to render a judgment based only on the video. To the best of our knowledge, an effective method of FR detection for speaker verification (i.e., a method for detecting whether the target speaker in a monocular video has been reenacted) has yet to be established. Therefore, in this proposed project, we will develop such a method for FR detection. As malicious interference or presentation attacks (e.g., impersonation and speech synthesis) are possible, analyzing the speaker’s voice in a target video would not be sufficient to determine whether the video has been reenacted. Therefore, this proposed project will develop an FR detection method based only on visual speech. The FR process destroys some of the intrinsic properties of a facial region of interest (fROI), including the intrinsic characteristics of lip motion, spontaneous subtle (SS) facial expressions, and underlying traces of facial forgery and blur (FFB). Our preliminary studies have found that FR can be detected by analyzing this visual information. Accordingly, this proposed project will develop a visual-based approach to FR detection that relies on the visual information extracted from an fROI. To this end, we will investigate and analyze the characteristics of a speaker’s lip motion and the dynamics of a speaker’s SS facial expressions, and perform FFB estimation, based on which we will design and develop corresponding models and algorithms. Subsequently, we will build a fusion model and establish an integrated learning framework for FR detection. Besides, we will explore two new applications of this novel FR detection technique: (1) fake video identification, and (2) visual speech-independent speaker verification. That is, the identity verification of a speaker who can say different content in two separate videos (one video is a genuine one, and the other one is to be verified) is made based on the visual information only. To the best of our knowledge, such applications have yet to be explored. Overall, this proposed project will develop new models and algorithms for FR detection, in addition to providing an in-depth understanding of underlying FR mechanisms. The resulting techniques, theoretical findings, and empirical discoveries will serve as promising countermeasures to malicious use of FR techniques. Furthermore, prototype systems for the two applications mentioned above will be developed, thereby providing promising new tools for identifying fake videos and verifying speaker identities.


Grant Support:

This project is supported by the Research Grants Council (RGC), Hong Kong SAR, China [Project: SRFS2324-2S02].

For further information on this research topic, please contact Prof. Yiu-ming Cheung.