Home › Polygraph Research › Confidence, Accuracy, and Utility of Polygraph Decisions
Catalogue entry · Field Validity Studies
Confidence, Accuracy, and Utility of Polygraph Decisions
Benjamin Kleinmuntz; Julian J. Szucko — Journal of Applied Psychology,
Examined the relationship between examiner confidence and decision accuracy. Field study included in the OTA analysis. Highlighted that structural improvements to testing protocols can reduce error rates.
Abstract
Examination of the relationship between examiner confidence in their diagnostic decisions and the actual accuracy of those decisions in CQT examinations.
Methodology
Field study examining confidence-accuracy calibration in polygraph decisions. Examiner confidence ratings compared against confirmed outcomes.
Detailed summary
Kleinmuntz and Szucko examined whether polygraph examiners' subjective confidence in their decisions correlated with actual accuracy. The study highlighted that structural improvements to testing protocols — such as standardized scoring methods and decision rules — could reduce error rates beyond what subjective judgment alone achieves. This finding anticipated the later movement toward computerized scoring and objective decision criteria.
Implications for polygraph practice
The confidence-accuracy analysis supports the transition from subjective to objective scoring methods. Examiners may be over- or under-confident in specific cases, making standardized decision rules valuable for consistency and accuracy.
Comprehensive study analysis
An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.
Background & Context
The early 1980s marked a critical period of scientific scrutiny for polygraph testing in the United States. As polygraph use expanded in both law enforcement and private sector employment screening, Benjamin Kleinmuntz and Julian Szucko's 1984 field study emerged as a landmark investigation into the actual accuracy of polygraph interpretations in real-world settings. Published in the prestigious journal Nature, this research addressed a fundamental question that had eluded previous studies: how well do polygraph examiners actually perform when evaluating genuine criminal cases?
The study gained immediate prominence by extending their earlier laboratory mock theft study to the field, revealing high false-positive and false-negative rates among polygraph interpreters. This work was particularly significant because it employed blind evaluation procedures with verified ground truth—a methodological strength that distinguished it from many contemporaneous polygraph validity studies. The study used files from a polygraph testing firm, with all cases involving theft, providing a focused examination of examiner performance in a specific category of criminal investigation.
Research Design & Methodology
The main purpose was to extend their earlier mock theft lie detection study from the laboratory to the field. The researchers employed a retrospective blind evaluation design using archival polygraph case files that had been verified through subsequent confessions. Files came from a polygraph testing firm (identity not disclosed), with all cases involving theft crimes, creating a homogeneous sample for analysis.
The study employed multiple independent polygraph interpreters who were provided with chart data but remained blind to the actual outcomes. Six polygraph raters evaluated the charts, with estimates indicating judges detected on average 71 percent of guilty subjects, but also called half of the innocent subjects deceptive. Interpreters in the study used global evaluation procedures rather than standardized numerical scoring systems.
Key methodological features included:
- Blind chart evaluation by multiple independent examiners
- Confession-verified ground truth for determining actual guilt or innocence
- Analysis using lens model equations from human judgment research
- Comparison of clinical interpretation versus statistical approaches
Results & Key Findings
The findings showed high false-positive and false-negative rates among polygraph interpreters, with data showing unacceptably high error rates. The Kleinmuntz and Szucko study had among the lowest accuracy rates compared to other field studies analyzed in the 1983 Office of Technology Assessment review.
Performance metrics revealed concerning patterns:
- Average detection rate of 71% for guilty subjects
- 50% false-positive rate—half of innocent subjects incorrectly identified as deceptive
- Five of six raters detected deception significantly better than chance, but four had fairly high false-positive rates
- 80% correct classification achievable using purely statistical analysis
The study found validity coefficients of polygraph examiners' judgments markedly inferior to a purely statistical analysis of the charts. This suggested that human clinical judgment added little value—and potentially introduced bias—compared to algorithmic approaches. The finding that theft cases might be particularly difficult to assess was noteworthy, as crimes against property may be less detectable than crimes against persons, possibly because crimes against persons have greater affect associated with them and are thus more physiologically detectable.
Discussion & Significance
The Kleinmuntz and Szucko study fundamentally challenged the prevailing confidence in clinical polygraph interpretation. The data, with their unacceptably high error rates, agreed with conclusions reached in the US Congress Office of Technology Assessment report, leading the authors to conclude that the validity and reliability of polygraphic interrogation had yet to be established. This represented a significant departure from industry claims about polygraph accuracy.
The research illuminated critical differences between subjective interpretation and objective scoring approaches. The superior performance of statistical methods over clinical judgment suggested that standardizing polygraph decision-making could reduce errors. This finding presaged the movement toward computerized scoring algorithms and numerical evaluation systems that would emerge in subsequent decades. The study also contributed to understanding of examiner confidence-accuracy relationships, demonstrating that examiner certitude did not necessarily correlate with actual correctness.
The study became frequently cited in scientific literature supporting the position that CQT theory's conclusions about bias against innocent individuals are supported by published research findings in the best social science journals. Its publication in Nature gave the findings particular credibility within the broader scientific community beyond polygraph specialists.
Limitations & Considerations
Despite generally anomalous results compared to other field studies, the exclusive focus on theft cases represented a potential limitation. The restriction to property crimes may have reduced the generalizability of findings to other offense types. Additionally, it remained unclear how comparable their lens model method of measuring validity was to the usual method of using accuracy rates, raising questions about comparability with other validation research.
The use of confession-verified cases introduced potential selection bias. Cases that result in confessions may differ systematically from those that do not, potentially affecting the representativeness of the sample. The study also relied on examiners using global evaluation procedures rather than the numerical scoring methods that were beginning to gain acceptance, which may have influenced the relatively poor performance observed.
Practical Applications
The findings provided compelling evidence for moving away from purely subjective chart interpretation toward more standardized, objective scoring protocols. The demonstration that statistical approaches outperformed clinical judgment suggested that polygraph accuracy could be improved through algorithm-based decision rules rather than relying solely on examiner intuition. This has direct implications for modern polygraph practice, where computerized scoring systems have become increasingly prevalent.
For consumers of polygraph services and legal decision-makers, the study highlighted the critical importance of understanding error rates—particularly false-positive rates that can harm innocent individuals. The 50% false-positive rate observed in this field sample underscored that polygraph results should never be the sole basis for consequential decisions about employment, security clearance, or criminal investigation. The research reinforced the need for corroborating evidence and cautious interpretation of polygraph outcomes in real-world applications.
The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.
Related research
Other studies in this category that may be of interest.
Features of information reliability assessment by polygraph method in criminal analysis
[002] (2024)How Reliable are Polygraph Examinations in Criminal Investigations? An Empirical Assessment
[003] (2022)Research with the use of a polygraph in the investigation of environmental…
[004] (2021)Directed Lie – The Correct or the Easy Way?
[005] (2020)Iacono and Ben-Shakhar's response to Ginton (2020): Validity estimates from paired testing
[006] (2019)Validity of the Control Question Test in Two Levels of the Severity…
Join Our Examiner Network
APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.
Keep reading the ledger.
Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.