Deep learning based face-swap videos, widely known as deepfakes, have drawn wide attention due to their threat to information credibility. Recent works mainly focus on the problem of deepfake detection that aims to reliably tell deepfakes apart from real ones, in an objective way. On the other hand, the subjective perception of deepfakes, especially its computational modeling, imitation, is also a significant problem but lacks adequate study. In this paper, we focus on the photorealism assessment of deepfakes, which is defined as the automatic assessment of deepfake photorealism that approximates human perception of deepfakes. It is important for evaluating the quality, deceptiveness of deepfakes which can be used for predicting the influence of deepfakes on Internet, it also has potentials in improving the deepfake generation process by serving as a critic. This paper promotes this new direction by presenting a comprehensive benchmark called DREAM, which stands for Deepfake photoREalism AssessMent. It is comprised of a deepfake video dataset of diverse quality, a large scale annotation that includes 140, 000 photorealism scores, textual descriptions obtained from 3, 500 human annotators, a comprehensive evaluation, analysis of 18 representative photorealism assessment methods, including recent large vision language model based methods, a newly proposed description-aligned CLIP method. The benchmark, insights included in this study can lay the foundation for future research in this direction, other related areas.
