Home › Polygraph Research › Comparing OpenFace and Deep Learning Models for Deception…
Catalogue entry · Neuroimaging & New Technologies
Comparing OpenFace and Deep Learning Models for Deception Detection in Video Calls
Laslo Dinges, Marc-André Fiedler, Ayoub Al-Hamadi, Dmitri Bershadskyy, Joachim Weimann — 2025 14th International Symposium on Image and Signal Processing and Analysis (ISPA),
CNN-based AU detection outperforms OpenFace on challenging Real-life Trial data with lower video quality and unstable head pose, while OpenFace performs better in controlled settings due to broader AU coverage. Multimodal approaches slightly outperform unimodal features in both scenarios.
Abstract
In remote deception detection scenarios, where data are scarce, modalities are restricted to audio and video, and facial cues are subtle and highly individual, modeling expressions via Action Units (AUs) offers a focused and interpretable approach. This paper investigates why many methods still rely on OpenFace for AU prediction and compares the performance of CNN-based AU prediction across different scenarios. For this study, we acquired a new dataset collected in a controlled setting where participants were instructed to deceive during mock online sales meetings. For comparison, we also use the Real-life Trial dataset. We first demonstrate that our CNN-based approach outperforms OpenFace on the EmotioNet AU dataset. When used as a feature set for deception detection, the CNN approach also performs better on the challenging, in-the-wild data of the Real-life Trial. However, OpenFace, that covers also AUs not included within EmotioNet, performs better in the semi-controlled mock sales meeting scenario.
Methodology
Comparative study using two datasets: a new 500-sample Buyer-Seller-Deception-Game dataset with mock online sales meetings, and the high-stakes Real-life Trial courtroom dataset. Researchers compared OpenFace-based AU features against CNN models trained on EmotioNet, evaluating performance for deception detection across visual and audio modalities.
Comprehensive study analysis
An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.
Background & Context
Automated deception detection from video remains one of the most challenging problems in computer vision and affective computing. In remote deception detection scenarios, data are scarce, modalities are restricted to audio and video, and facial cues are subtle and highly individual, making reliable analysis particularly difficult. Traditional approaches have relied heavily on OpenFace, an open-source facial behavior analysis toolkit, for extracting Action Units (AUs)—the fundamental building blocks of facial expressions defined by the Facial Action Coding System.
Despite rapid advances in deep learning, many deception detection systems continue to use OpenFace for AU prediction rather than more modern CNN-based approaches. This study addresses a critical question: when does the traditional OpenFace approach outperform deep learning models, and vice versa? This paper investigates why many methods still rely on OpenFace for AU prediction and compares the performance of CNN-based AU prediction across different scenarios. The answer has significant implications for developing practical deception detection tools for video conferencing and online interactions.
The research is particularly timely given the surge in remote communication platforms where deception detection could serve applications ranging from online fraud prevention to law enforcement interviews conducted via video call. Deceptive behavior is highly context dependent, shaped by factors such as the specific scenario, cultural background, and the associated stakes, necessitating robust methods that can generalize across diverse recording conditions.
Research Design & Methodology
For this study, the researchers acquired a new dataset collected in a controlled setting where participants were instructed to deceive during mock online sales meetings. This "Buyer-Seller-Deception-Game" (GAME) dataset represents a novel contribution to deception detection research. It currently includes 500 annotated video samples, with a planned extension to around 1000 before public release, where participants engage in incentivized online interactions, in which sellers attempt to persuade buyers to choose a specific card, creating naturally motivated deceptive and truthful behavior.
The experimental design compared two distinct AU detection approaches across two contrasting datasets:
- OpenFace-based features: Traditional AU extraction using the established OpenFace 2.0 toolkit
- CNN-based approach: Deep learning models trained for AU prediction on the EmotioNet dataset
- GAME dataset: Clean, controlled recordings from mock sales meetings (low-stakes, high-quality video)
- Real-life Trial dataset: High-stakes courtroom testimony videos with challenging in-the-wild conditions
The researchers first demonstrated that their CNN-based approach outperforms OpenFace on the EmotioNet AU dataset, establishing the technical superiority of the deep learning model under controlled benchmark conditions. They evaluated a variety of visual and audio based feature sets, such as gaze, head pose, facial Action-Units (AUs), and prosodic features, on their dataset as well as on a high-stakes in-the-wild deception dataset.
Results & Key Findings
The study revealed a nuanced, context-dependent performance pattern that challenges conventional assumptions about deep learning superiority. The choice between OpenFace and CNN-based AU detection depends critically on recording conditions.
When used as a feature set for deception detection, the CNN approach performs better on the challenging, in-the-wild data of the Real-life Trial. More specifically, CNN based AU predictors outperform others in the more challenging dataset with lower video quality and unstable head pose. This suggests deep learning models exhibit greater robustness to the degraded visual conditions typical of real-world courtroom footage.
However, the opposite pattern emerged in controlled settings: OpenFace based AU features perform best on clean and controlled recordings. The study's abstract notes that in the semi-controlled mock sales meeting scenario, OpenFace performs better, likely because OpenFace covers AUs not included within EmotioNet, providing more comprehensive facial analysis when video quality permits detailed feature extraction.
Additional findings include:
- Multimodal approaches slightly outperform the best unimodal features in both cases
- CNN models trained on EmotioNet showed superior AU detection accuracy on that benchmark compared to OpenFace
- The performance trade-off relates to both AU coverage and robustness to video degradation
Discussion & Significance
This research provides crucial practical guidance for deception detection system development. The key insight is that tool selection should be driven by recording conditions rather than assuming deep learning superiority. For applications involving controlled environments like structured interviews or high-quality video conferences, traditional OpenFace may offer advantages through broader AU coverage. Conversely, for real-world applications with variable lighting, camera angles, and video quality—such as archived footage or surveillance scenarios—CNN-based approaches demonstrate superior resilience.
The findings help explain why OpenFace remains prevalent in deception detection research despite the deep learning revolution. Its comprehensive AU coverage and reliable performance in controlled settings make it a pragmatic choice for many laboratory studies and structured data collection protocols. The researchers present a new high-quality low-stakes dataset specifically tailored for exploring the potential of future deception detection tools in online interactions, such as virtual sales meetings, addressing a critical gap in available training data.
The study also highlights fundamental challenges in deception detection: the scarcity of labeled training data, the context-dependency of deceptive behaviors, and the need for methods that generalize across recording conditions. By systematically comparing approaches across different data types, this work provides empirical evidence for informed technology selection in operational systems.
Limitations & Considerations
Several methodological constraints warrant consideration. The GAME dataset, while valuable, represents relatively low-stakes deception in an economic game setting, which may not fully capture the psychological dynamics of high-consequence lying. The mock sales meeting scenario, though ecologically valid for online commerce contexts, differs substantially from criminal investigations or security screening applications.
The CNN models were trained specifically on EmotioNet, which contains a subset of possible AUs. This limited AU coverage disadvantages the deep learning approach in controlled settings where comprehensive facial analysis is feasible. Different CNN architectures or training datasets might yield different performance trade-offs. Additionally, the study does not report specific accuracy metrics or statistical significance testing, making it difficult to quantify the magnitude of performance differences between approaches.
The planned expansion of the GAME dataset to approximately 1,000 samples was not yet complete at publication, potentially limiting the generalizability of findings from that portion of the study. Cross-cultural validity also remains unexplored, as facial expression norms and deceptive behaviors vary across populations.
Practical Applications
These findings offer actionable recommendations for practitioners developing automated deception detection systems. For video conferencing platforms, online interview systems, or remote sales verification, where video quality can be controlled, OpenFace represents a robust, well-validated solution with comprehensive AU detection. Its open-source nature and extensive documentation facilitate implementation.
Conversely, for forensic applications, law enforcement interrogations using archived footage, or security screening with variable recording conditions, investing in CNN-based AU detection models trained for robustness will likely yield superior performance. The enhanced tolerance for head pose variation and video degradation makes deep learning approaches better suited for uncontrolled environments. Researchers and developers should consider hybrid approaches that dynamically select methods based on real-time video quality assessment, leveraging the complementary strengths of both techniques.
The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.
Related research
Other studies in this category that may be of interest.
Modes of visual framing in neuropsychological deception detection research
[002] (2026)Enhanced visualisation of concealed target objects by infrared thermography and machine learning
[003] (2026)Buyer–Seller-Deception-Game Dataset: A new comprehensive dataset for facial expression based deception detection…
[004] (2026)A novel method based on variational mode decomposition for lie detection
[005] (2026)Influence of Stimulus Layout and Social Presence on Deception-Related Eye Movements and…
[006] (2026)Neural Signatures of Deception: An Explainable Machine LearningApproach Using EEG Signals
Join Our Examiner Network
APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.
Keep reading the ledger.
Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.