Achievement Award

Pioneering research and social implementation to prevent the spread of fake media

Isao ECHIZEN
Isao ECHIZEN
Junichi YAMAGISHI
Junichi YAMAGISHI
Noboru BABAGUCHI
Noboru BABAGUCHI

When AI is allowed to learn from information originating from humans, such as faces, bodies, voices, and natural language text, it is possible for it to generate synthetic media that is almost indistinguishable from real human output. Synthetic media is used in various fields, including communication and entertainment. There is, however, a negative aspect to this capability whereby criminal organizations and terrorist organizations are able to use AI to generate and spread fake media, such as fake videos, fake audio, and fake documents, for the purpose of fraud, thought guidance, and public opinion manipulation, and this has become a global threat.

The awardees started a project to prevent the spread and generation of Media Clones (a coined term for media that has been processed and generated to look like the real thing using AI) in 2016, before deepfakes became a social issue, with funding from a Grant-in-Aid for Scientific Research (S) (PI: Babaguchi, Co-PI: Echizen, Collaborative Researcher: Yamagishi). In 2018, they proposed the world's first method for detecting fake facial videos using AI [1] (cited over 1,600 times). Furthermore, in 2019, they proposed a detection method that is robust even against unseen fake facial videos [2] (cited more than 700 times) and a method that simultaneously performs fake facial video detection and segmentation of tampered (modified) regions [3] (cited more than 500 times, awarded the IEEE BTAS/5-Year Highest Impact Award), creating a new research field called ‘deepfake detection’. In addition, a paper on a privacy protection method [4] that prevents the creation of Media Clones was accepted by a top-tier journal, and achieved remarkable academic results. The paper [5] that summarized these results of the Grant-in-Aid for Scientific Research (S) was awarded the IEICE-ISS Excellent Paper Award.

Utilizing the findings from these research results, two pioneering JST CREST projects (CREST VoicePersonae (FY2018-2023, PI: Yamagishi, Research Participant: Echizen), CREST FakeMedia (FY2020-2025, PI: Echizen, Co-PI: Babaguchi, Research Participant: Yamagishi)) were launched. CREST VoicePersonae, by pursuing the conflicting goals of utilizing and protecting voice identity, achieved the advancement of both technologies, and produced groundbreaking academic results such as a voice privacy protection method that is robust against re-identification attacks [6] and a large-scale dataset for deepfake speech detection [7] (800,000 downloads), and was selected for additional support under the AIP Accelerated Program (FY2024-2026, PI: Yamagishi) in 2023. CREST FakeMedia was anticipating harm to human society from fake media even before the threat of generative AI had become apparent, and pioneered groundbreaking academic achievements, such as a method for recovering original faces from fake ones called CyberVaccine [8], a large-scale dataset for fake facial video detection [9], and a method for detoxifying adversarial samples [10] (awarded the IEICE Best Paper Award). The team has also developed an automatic detection program for fake facial videos in collaboration with CREST VoicePersonae, and made a significant contribution to the creation of trust in social cyber implementations. In 2024, these achievements were adopted as the research and development basis for the JST K Program SYNTHETIQ X (FY2024-2029, PI: Echizen).

In light of the importance of the two CREST projects mentioned above, the Global Research Center for Synthetic Media was established at the National Institute of Informatics in 2021 to promote international joint research and social implementation. Furthermore, in 2021, that center developed an automatic detection program for fake facial videos called “SYNTHETIQ VISION”, applied for a basic patent (registered in April 2023), and have been providing paid licenses to multiple domestic companies. In 2023, CyberAgent adopted SYNTHETIQ VISION, which became the first practical example of fake facial image detection in Japan. In addition, the audio privacy protection method was put to practical use by NHK in 2024.

As described above, the awardees have not only achieved pioneering academic results in preventing the spread and creation of fake media, but have also pioneered innovations such as industrial applications of fake facial video detection and audio privacy protection, which have had an enormous academic, social, and industrial impact. The achievements of the awardees are extremely significant in contributing to the realization of a trusted human-centered cyber society and are worthy of the Achievement Award.

References

  1. D. Afchar, V. Nozick, J. Yamagishi and I. Echizen, "MesoNet: a Compact Facial Video Forgery Detection Network," 2018 IEEE International Workshop on Information Forensics and Security (WIFS), pp. 1-7, 2018
  2. H. H. Nguyen, J. Yamagishi and I. Echizen, "Capsule-forensics: Using Capsule Networks to Detect Forged Images and Videos," 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2307-2311, 2019
  3. H. H. Nguyen, F. Fang, J. Yamagishi and I. Echizen, "Multi-task Learning for Detecting and Segmenting Manipulated Facial Images and Videos," 2019 IEEE 10th International Conference on Biometrics Theory, Applications and Systems (BTAS), pp. 1-8, 2019
  4. Y. Hirose, K. Nakamura, N. Nitta and N. Babaguchi, "Anonymization of Human Gait in Video Based on Silhouette Deformation and Texture Transfer," in IEEE Transactions on Information Forensics and Security, vol. 17, pp. 3375-3390, 2022
  5. N. Babaguchi, I. Echizen, J. Yamagishi, et al., “Preventing Fake Information Generation Against Media Clone Attacks,” IEICE Trans. Info. & Sys., vol. E104, no. 1, pp. 2-11, 2021
  6. F. Fang, X. Wang, J. Yamagishi, I. Echizen, M. Todisco, N. Evans, and J.-F. Bonastre, “Speaker Anonymization Using X-vector and Neural Waveform Models,” Proc. 10th ISCA Speech Synthesis Workshop, pp. 155-160, 2019
  7. X. Wang, J. Yamagishi, et al., “ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech”, Computer Speech & Language, Volume 64, 2020
  8. C. -C. Chang, H. H. Nguyen, J. Yamagishi and I. Echizen, "Cyber Vaccine for Deepfake Immunity," in IEEE Access, vol. 11, pp. 105027-105039, 2023
  9. T. -N. Le, H. H. Nguyen, J. Yamagishi and I. Echizen, "OpenForensics: Large-Scale Challenging Dataset For Multi-Face Forgery Detection And Segmentation In-The-Wild," 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 10097-10107, 2021
  10. H. H. Nguyen, M. Kuribayashi, J. Yamagishi, and I. Echizen, “Effects of Image Processing Operations on Adversarial Noise and Their Use in Detecting and Correcting Adversarial Images,” IEICE Trans. Info. & Sys., vol. E105, no. 1, pp. 65-77, 2022
Fig.1
Fig. 1:Overview of pioneering research and social implementation results from large-scale projects