Social Science Research Council Research AMP Just Tech
Citation

Generative Deepfake Videos in the Foundation-Model Era: A Timeline of Eroding Trust in Visual Evidence

Author:
Raza, Shaina; Ho, Jessee; Raza, Mahveen; Emmanouilidis, Christos
Year:
2026

Synthetic video quality has advanced to the point where contemporary diffusion transformers generate minute-long, high-resolution clips with coherent motion and complex scenes, systematically erasing the blending boundaries, frequency peaks, and temporal flicker on which early detectors relied. This review article charts the resulting post-artifact era along two axes. First, we introduce a forensic-assumption framework (A1–A5) covering physiological integrity, temporal coherence, geometric consistency, semantic consistency, and provenance signals, and trace how each assumption erodes across the GAN, diffusion, and video foundation model eras. Second, we compare five detection paradigms: spatial–frequency, temporal, multimodal, vision–language, and agentic , against the assumptions they target, their cross-generator generalization, and their deployment readiness. We then audit thirteen benchmarks against a four-pillar evaluation protocol (generalization, robustness, trustworthiness, deployment) and show that trustworthiness and deployment readiness remain systematically undermeasured. We close with a research roadmap spanning assumption-aware detection, provenance–detection co-design, self-play adversarial training, agentic orchestration, and holistic evaluation.