Home › Polygraph Research › A Comprehensive Meta-Analysis of the Comparison Question Polygraph…
Catalogue entry · Meta-Analyses & Systematic Reviews
A Comprehensive Meta-Analysis of the Comparison Question Polygraph Test
Charles Robert Honts; Mark Handler; Pamela K. Shaw; Michael C. Gougler — Applied Cognitive Psychology,
The largest meta-analysis ever conducted on the CQT, analyzing 138 datasets. Found a meta-analytic effect size of 0.69 including inconclusives. Motivation level showed a positive linear relationship with accuracy. No publication bias was detected. Concluded that CQT can be accurate and experimental studies are generalizable to field settings.
Abstract
This study presents the largest meta-analysis ever conducted on the Comparison Question Test (CQT), encompassing 138 datasets from both laboratory and field studies. The authors examined overall accuracy, the relationship between motivation and accuracy, publication bias, and the generalizability of experimental findings to field conditions. Using random-effects meta-analytic models, the study calculated effect sizes and examined moderating variables across the full body of CQT research.
Methodology
Random-effects meta-analysis of 138 datasets. Studies were identified through comprehensive literature searches of published and unpublished research. Effect sizes were calculated using standard meta-analytic procedures. Moderator analyses examined the influence of study type (laboratory vs. field), motivation level, scoring method, and examiner experience. Publication bias was assessed using funnel plots and Egger's regression test.
Detailed summary
Honts, Handler, Shaw, and Gougler conducted the most comprehensive meta-analysis of CQT research to date, analyzing 138 datasets spanning decades of laboratory and field research. The meta-analytic effect size of 0.69 (including inconclusive outcomes) demonstrated substantial discrimination between truthful and deceptive examinees. A key finding was the positive linear relationship between examinee motivation and test accuracy — as the stakes of the examination increased, so did physiological differentiation between truthful and deceptive responses. This finding has profound implications for real-world applications, where the consequences of detection are inherently high. The analysis found no evidence of publication bias, addressing a common criticism of polygraph research. Perhaps most significantly, the study concluded that experimental (laboratory) studies are generalizable to field settings, bridging the long-standing gap between controlled research and operational practice. This directly countered arguments that laboratory accuracy rates overestimate real-world performance.
Implications for polygraph practice
This meta-analysis provides the strongest quantitative evidence to date supporting CQT validity. The motivation-accuracy relationship suggests that real-world polygraph examinations — where stakes are inherently higher than laboratory settings — may actually perform better than lab studies indicate. The absence of publication bias strengthens confidence in the overall accuracy estimates. For practitioners, this supports the continued use of validated CQT formats, particularly in high-stakes forensic and criminal investigation contexts.
Comprehensive study analysis
An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.
Background & Context
The Comparison Question Test has long stood at the center of heated scientific debate. For decades, critics have argued that laboratory studies overestimate real-world accuracy, that lack of motivation undermines experimental validity, and that publication bias inflates reported accuracy rates. Previous reviews generally produced overall accuracy estimates of over 85%, yet skeptics dismissed these findings as methodologically flawed or biased toward positive outcomes.
This 2021 meta-analysis by Honts, Handler, Shaw, and Gougler represents the most ambitious quantitative synthesis of CQT research ever undertaken. Data and potential moderators were coded from 138 datasets, making it substantially larger than the only previous CQT meta-analysis conducted by Kircher and colleagues in 1988, which analyzed only 14 studies. The researchers deliberately used broad inclusion criteria to capture the full range of CQT research quality and address criticisms that previous reviews cherry-picked favorable studies.
The timing was critical. Following the 2003 National Research Council report that questioned polygraph validity, the field needed rigorous quantitative evidence to either validate or refute longstanding claims about the technique. This study aimed to test specific assertions made by critics, including whether experimental studies generalize to field settings and whether motivation level affects accuracy—questions with profound implications for the continued use of polygraph testing in criminal justice and national security contexts.
Research Design & Methodology
The researchers conducted a comprehensive literature search spanning published and unpublished CQT research, closing their data collection on July 1, 2018. They employed random-effects meta-analytic models to account for variability across diverse study designs. The first and third author independently coded the first 97 data sets obtained in their analysis representing 70% of their retained data sets, and agreement was high for all the moderators, ensuring reliability in how study characteristics were classified.
The researchers adopted a unique effect size metric called rdec (the product-moment correlation between ground truth and test outcome) that includes inconclusive results in accuracy calculations. rdec is sensitive to the impact of inconclusive outcomes where their occurrence reduces the value of rdec, but not by as much as an error. This approach more realistically represents operational polygraph testing, where inconclusives are routine outcomes rather than errors.
Key moderator variables examined included:
- Motivation level — categorized as No Explicit, Some, or Real-World motivation based on consequences attached to test outcomes
- Study type — laboratory experiments versus field studies with confirmed ground truth
- Subject sampling — student, community, work, or forensic populations
- Peer-review status — published versus unpublished research
A majority of the data sets (86, 63.7%) were from experiments, with the remainder coming from confirmed field cases. The researchers also calculated Information Gain statistics to compare CQT performance against unaided human deception detection, providing a practical benchmark for evaluating the technique's diagnostic value.
Results & Key Findings
The meta-analytic effect size including inconclusive outcomes was 0.69 [0.66, 0.79], representing substantial discrimination between truthful and deceptive examinees. When translated to more familiar metrics, the analysis yielded Cohen's d = 1.92, and an AUC = 0.91, although the AUC estimate was reduced by the inclusion of inconclusive outcomes while the NRC estimate of AUC did not consider inconclusive outcomes. CQT decisions in the Median Sample were 86% correct.
The most theoretically significant finding was the positive linear relationship between motivation level and accuracy. Level of motivation had a positive linear relationship with outcome measures. For the three levels of Motivation (No Explicit, Some, and Real World) IG for Truthful outcomes peaked at .35, .45, and .46, respectively. This directly contradicts critics' assertions that low-stakes laboratory studies inflate accuracy estimates—the data show the opposite pattern.
Information Gain analysis demonstrated substantial practical utility:
- A CQT truthful outcome is approximately 7 times more informative than a layperson's decision of truth teller
- A CQT deceptive outcome is approximately 6 times more informative than a layperson's decision that a person is a liar
- Information Gain analysis of CQT outcomes representing the median accuracy showed a significant information increase over interpersonal deception detection across almost the complete range of base rates
Publication bias analysis using funnel plots and regression techniques revealed no publication bias was detected, addressing a persistent criticism that only favorable studies reach publication. The generalizability finding was equally important: experimental studies are generalizable to field settings, bridging the long-standing divide between controlled research and operational practice.
Discussion & Significance
This meta-analysis fundamentally challenges the narrative that polygraph validity remains scientifically unproven. The motivation-accuracy relationship has particularly profound implications: it suggests that real-world polygraph examinations—where stakes inherently exceed laboratory conditions—may actually perform better than experimental studies indicate. This inverts the traditional criticism that lab studies provide overly optimistic accuracy estimates.
The Information Gain analysis provides crucial context often missing from polygraph debates. While 86% overall accuracy might sound modest compared to forensic techniques like DNA analysis, the comparison to unaided human deception detection (approximately 54% accuracy) demonstrates the CQT's substantial diagnostic value. The finding that CQT decisions are 6-7 times more informative than layperson judgments quantifies this advantage in a practically meaningful way.
However, the authors emphasize an important caveat. Results should not be interpreted as indicating that all CQT polygraph tests have high accuracy. There was a large range of results that does not appear to be due to the tested moderators and thus there are likely other factors at work. They specifically note concerns about certain agency practices. Analysis based upon U.S. Government generated polygraph data indicated that under the FBI's policy only 17% of the actually innocent people given FBI polygraph examinations will avoid interrogation, highlighting how misuse of validated techniques can produce unacceptable false positive rates.
Limitations & Considerations
The researchers acknowledge significant limitations in the available research base, particularly concerning field studies. A number of potential moderator variables were not included in the meta-analysis because there were an insufficient number of studies reporting the data (e.g., age, years of education, specific type of crime in field studies, and years of experience as an examiner). This prevented analysis of potentially important factors like examiner expertise that likely influence operational accuracy.
The authors discussed the limitations of the field research literature and problems within polygraph profession that lower field accuracy. Field study quality varies dramatically, with inconsistent confirmation methods for ground truth and wide variation in testing protocols. The large unexplained variance in outcomes suggests that unmeasured factors—possibly including examiner skill, agency protocols, and subject populations—play substantial roles in determining accuracy.
Additionally, the broad inclusion criteria adopted to maximize coverage meant including studies with methodological weaknesses that previous reviewers had excluded. While this approach provides a more comprehensive picture of CQT research, it may underestimate the accuracy achievable under optimal conditions with well-trained examiners using validated protocols.
Practical Applications
For polygraph practitioners, this meta-analysis provides the strongest empirical foundation to date for continued use of properly administered CQT formats in forensic and investigative contexts. The motivation-accuracy relationship suggests that the technique is particularly well-suited for high-stakes applications—precisely the contexts where polygraph testing is most commonly deployed. The absence of publication bias strengthens confidence in reported accuracy rates across the research literature.
However, the findings also carry an urgent warning: the CQT's demonstrated validity applies only when the technique is properly implemented. The stark differences between research-based accuracy and certain agency practices (such as treating inconclusives as deceptive) underscore that validated techniques can produce unacceptable error rates when misapplied. For consumers and policymakers, this research supports polygraph use in appropriate contexts while demanding accountability for protocols that deviate from evidence-based practices. The study reinforces that polygraph testing is not inherently accurate or inaccurate—outcomes depend critically on examiner training, adherence to validated protocols, and institutional policies governing decision rules and inconclusive classifications.
The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.
References in our database [2]
Studies cited by this paper that are available in our research database.
Related research
Other studies in this category that may be of interest.
Decoding Deception with the P300: A Meta-Analysis of the Concealed Information Test
[002] (2025)Neurophysiological Approaches to Lie Detection: A Systematic Review
[003] (2025)Memory Distortions, Confabulation, and Their Impact on Polygraph Examinations
[004] (2025)A Comprehensive Review and Future Prospects of Lie Detection Using Machine Learning
[005] (2025)Neuroscientific and Biometric Methodologies in Forensic Contexts: A Critical Review
[006] (2024)Psychology and the Lie Detector Industry: A Fifty-Year Perspective
Join Our Examiner Network
APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.
Keep reading the ledger.
Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.