The increasing realism of synthetic facial videos has reduced the reliability of detectors that depend only on visible artifacts in isolated frames. This paper presents BlinkNet, an explainable deepfake-detection framework that examines the spatial appearance and temporal kinematics of eye blinks. The system detects a face, localizes 68 facial landmarks, extracts normalized ocular crops, and computes the Eye Aspect Ratio (EAR) for each frame. Overlapping sequences of 20 frames are processed by a dual-stream Temporal-Spatial Physiological Blink Anomaly Network (TPBAN). A lightweight MobileNetV2 encoder models local visual inconsistencies, while a bidirectional gated recurrent unit models the forward and backward dynamics of eyelid motion. Temporal attention assigns a relevance weight to every frame and supports frame-level anomaly visualization. Training uses a multi-task objective for authenticity classification and blink-phase recognition, together with a class-weighted binary cross-entropy term to address the imbalance between genuine and manipulated sequences. On the FaceForensics++ c23 test partition, BlinkNet obtained 83.16% accuracy, 87.31% ROC-AUC, 96.02% average precision, and a 21.08% equal error rate. The manipulated class achieved 0.90 precision and 0.88 recall. The implementation processed video at approximately 52 frames per second on a consumer laptop GPU and was integrated into a Flask-based forensic dashboard. These results indicate that ocular dynamics can complement spatial evidence while improving efficiency and interpretability.
Introduction
The rapid advancement of deepfake technology has enabled the creation of highly realistic synthetic videos, creating challenges in areas such as digital identity protection, forensic verification, financial security, and media authenticity. Traditional deepfake detection techniques mainly analyze individual video frames to identify visual artifacts like face blending errors, unnatural textures, and lighting inconsistencies. However, these methods become less effective when generated videos are visually realistic or compressed.
The proposed BlinkNet system introduces a physiological-based approach for deepfake detection by analyzing eye-blink patterns. Unlike static facial features, eye blinking is a natural temporal process involving eye opening, closing, and reopening sequences. Deepfake videos may reproduce realistic-looking frames but often fail to maintain natural blink frequency, duration, synchronization, and motion patterns. BlinkNet detects these inconsistencies by combining eye image analysis with eye movement information.
The system follows four main design principles:
Eye-focused analysis – It uses cropped eye regions as the primary input to reduce background noise and focus on relevant facial information.
Spatial-temporal learning – A lightweight convolutional network extracts visual features, while a Bidirectional Gated Recurrent Unit (BiGRU) learns temporal blink patterns.
Temporal attention mechanism – The model identifies important video frames contributing to the final prediction and provides visual explanations.
Imbalance-aware training – Weighted loss functions improve detection performance when genuine and manipulated video samples are unevenly distributed.
The literature review highlights existing approaches such as CNN-based frame analysis, FaceForensics++ benchmark models, physiological blink analysis, and lightweight deep learning architectures. While conventional CNN methods detect visual artifacts effectively, they often ignore whether facial movements are biologically realistic. BlinkNet improves upon these approaches by combining spatial features, physiological signals, and temporal modeling.
The proposed methodology begins with video input processing using OpenCV. Face detection is performed using HOG-based detection, followed by facial landmark extraction to locate eye regions. The system calculates the Eye Aspect Ratio (EAR) to measure eye openness and capture blink patterns. Eye images and EAR values are combined to form a dual input representation.
A MobileNet-based feature extractor processes eye images, while the EAR signal provides additional physiological information. These features are passed through a Bidirectional GRU network to learn temporal relationships. A temporal attention layer assigns importance weights to different frames, and a classifier generates a deepfake probability score.
To improve reliability, BlinkNet uses multi-task learning with two objectives:
Deepfake authenticity classification.
Blink-phase prediction (pre-blink, closing, closed, opening, and post-blink stages).
The model is trained and evaluated using the FaceForensics++ dataset, which contains real and manipulated videos. Videos are processed into overlapping 20-frame sequences for training, validation, and testing. The system is implemented using Python, PyTorch, OpenCV, and related machine learning libraries.
During inference, BlinkNet analyzes continuous video frames, predicts whether the content is authentic or manipulated, and generates attention-based visual explanations. The attention heatmap highlights frames where the model detects abnormal blink behavior, helping human reviewers understand the prediction.
Conclusion
BlinkNet demonstrates that eye-blink dynamics provide useful forensic information beyond static facial appearance. The proposed TPBAN architecture combines normalized ocular imagery, EAR-based physiology, MobileNetV2 feature extraction, bidirectional GRU sequence modelling, and temporal attention. On the FaceForensics++ c23 test partition, the system achieved 83.16% accuracy, 87.31% ROC-AUC, 96.02% average precision, and 21.08% EER, with 0.90 precision and 0.88 recall for manipulated sequences. Weighted training reduced majority-class bias, while the attention mechanism exposed the frames most responsible for a decision. Together with real-time performance on consumer hardware and a Flask interface, the framework provides a practical foundation for efficient and explainable deepfake screening. Broader cross-dataset validation and multi-physiological fusion remain essential before high-stakes deployment.
References
[1] Y. Li, M.-C. Chang, and S. Lyu, “In Ictu Oculi: Exposing AI Created Fake Videos by Detecting Eye Blinking,” in Proc. IEEE Int. Workshop on Information Forensics and Security (WIFS), 2018, pp. 1–7, doi: 10.1109/WIFS.2018.8630763.
[2] A. Rössler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nießner, “FaceForensics++: Learning to Detect Manipulated Facial Images,” in Proc. IEEE/CVF Int. Conf. Computer Vision (ICCV), 2019, pp. 1–11, doi: 10.1109/ICCV.2019.00009.
[3] D. Güera and E. J. Delp, “Deepfake Video Detection Using Recurrent Neural Networks,” in Proc. IEEE Int. Conf. Advanced Video and Signal Based Surveillance (AVSS), 2018, pp. 1–6, doi: 10.1109/AVSS.2018.8639163.
[4] T. Soukupová and J. ?ech, “Real-Time Eye Blink Detection Using Facial Landmarks,” in Proc. Computer Vision Winter Workshop, 2016, pp. 1–8.
[5] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2018, pp. 4510–4520, doi: 10.1109/CVPR.2018.00474.
[6] K. Cho et al., “Learning Phrase Representations Using RNN Encoder–Decoder for Statistical Machine Translation,” in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP), 2014, pp. 1724–1734, doi: 10.3115/v1/D14-1179.
[7] J. S. Chung, A. Jamaludin, and A. Zisserman, “You Said That?: Synthesising Talking Faces from Audio,” in Proc. British Machine Vision Conf. (BMVC), 2017, doi: 10.5244/C.31.54.
[8] I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” in Proc. Int. Conf. Learning Representations (ICLR), 2019.
[9] F. Chollet, “Xception: Deep Learning with Depthwise Separable Convolutions,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1251–1258, doi: 10.1109/CVPR.2017.195.