Home › Polygraph Research › Multivariate classifiers perform as well as experts in…
Catalogue entry · Scoring & Test Data Analysis
Multivariate classifiers perform as well as experts in the detection of deception
Statistical multivariate classifiers — including discriminant analysis and a bootstrap approach — achieved accuracy statistically equivalent to expert human evaluators on polygraph CQT data, with the bootstrap method performing best in absolute terms, supporting the scientific viability of computerized polygraph scoring.
Abstract
This 1993 study by Honts and Devitt compared the deception detection accuracy of expert human numerical evaluators, discriminant analysis, and a novel bootstrap classification approach using 200 subjects (100 innocent, 100 guilty) from the Utah Cooperative Working Group Database. The central finding was that multivariate statistical classifiers performed at statistically equivalent levels to experienced expert human scorers on the Comparison Question Test, with the bootstrap method achieving the highest absolute accuracy.
Methodology
Using 200 mock crime subjects from the Utah Cooperative Working Group Database, the study directly compared three CQT scoring approaches — expert human numerical evaluation, discriminant analysis, and a bootstrap resampling classifier — on the same physiological data to assess relative accuracy.
Detailed summary
This early 1990s study addressed whether algorithmic approaches could rival trained human polygraph specialists by comparing three classification methods using the same physiological data. Drawing from 200 balanced laboratory CQT cases, researchers tested human numerical evaluation against discriminant analysis and innovative bootstrap classification. The bootstrap method was particularly novel for avoiding the parametric assumptions and small sample sensitivities of traditional discriminant analysis. Results showed all three approaches achieved statistically equivalent performance, providing evidence that computer-based scoring could match expert human decision-making.
Implications for polygraph practice
This research provided crucial early evidence supporting the development of automated polygraph scoring systems, suggesting that objective computer algorithms could potentially reduce examiner bias and inconsistency while maintaining accuracy levels comparable to expert human evaluation.
Comprehensive study analysis
An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.
Background & Context
By the early 1990s, polygraph science was at a critical crossroads. Human examiners had long been the sole arbiters of deception decisions, relying on numerical scoring systems that — while structured — remained subject to individual training, fatigue, bias, and inconsistency. At the same time, computational tools were becoming increasingly capable of processing multivariate physiological datasets with far greater speed and objectivity. The central question hanging over the field was whether such algorithmic approaches could rival, or even surpass, the evaluative power of trained human specialists.
The research by M. K. Devitt and Charles R. Honts, conducted under a U.S. Navy grant and completed in the early 1990s, addressed this question head-on. Working within the productive research tradition established at the University of Utah — which had already produced the influential Computer-Assisted Polygraph System (CAPS) through the earlier work of Kircher and Raskin — Devitt and Honts sought to rigorously compare the accuracy of statistical multivariate classifiers against expert human evaluators using the same polygraph data.
This work emerged against a backdrop of growing concern about the reliability of standard polygraph scoring. The Department of Defense Polygraph Institute's semi-objective criteria had been criticized for lacking robust empirical support, and discriminant analysis — while powerful — came with methodological constraints that limited its generalizability. The study therefore filled a genuine scientific gap: testing whether newer, more flexible statistical techniques could hold their own against expert numerical scoring.
Research Design & Methodology
The study drew upon the Utah Cooperative Working Group Database, a rigorously maintained archive of polygraph records derived from multiple mock crime laboratory experiments. The final analytical dataset consisted of 200 subjects total — an equal split of 100 innocent and 100 guilty participants — all of whom had undergone the Comparison Question Test (CQT) under controlled laboratory conditions. This balanced design ensured that base-rate assumptions would not artificially inflate the performance of any single classification approach.
Three distinct decision-making approaches were placed in direct competition:
- Human numerical evaluation — Scores assigned by experienced polygraph examiners drawn from the Utah research group, representing a high-caliber (though non-representative) expert standard
- Discriminant analysis — The established multivariate technique underpinning the CAPS system, using physiological difference scores weighted by a derived linear function to separate guilty from innocent subjects
- Bootstrap classification — A then-novel resampling approach that avoids the parametric assumptions of discriminant analysis, iteratively constructing decision rules from repeated samples of the data to produce more robust and generalizable classifications
Physiological data — including respiration, electrodermal activity, and cardiovascular measures — were digitized, scored into difference scores comparing relevant to comparison question responses, and then submitted to each classification method. The bootstrap method was particularly innovative for its time, using computational resampling to sidestep the three major weaknesses of discriminant analysis: small sample sensitivity, restrictive distributional assumptions, and limited generalizability from laboratory to field settings.
Results & Key Findings
The headline finding was clear: all three classification methods produced statistically equivalent performance. Multivariate statistical classifiers matched the accuracy of expert human evaluators — a result with profound implications for the field. Neither the discriminant analysis approach nor the bootstrap classifier was significantly outperformed by the experienced human scorers.
- Statistical equivalence was established across all three approaches — human numerical evaluation, discriminant analysis, and bootstrap classification
- The bootstrap method achieved the highest absolute accuracy among the three, outperforming both discriminant analysis and human scoring in absolute terms, though differences did not reach statistical significance
- Human evaluators used in the study were senior researchers and expert scorers, not average field examiners — meaning the human baseline was artificially high
- The statistical classifiers operated on the same raw physiological data as the human evaluators, confirming that the information content of the polygraph record — not subjective interpretation — was driving accuracy
- The dataset of 200 subjects (100 guilty, 100 innocent) from mock crime paradigms provided a well-controlled, balanced test of each method's discriminative power
The practical implication of the equivalence finding is amplified by a key caveat explicitly noted by the authors: because the human evaluators were elite expert scorers rather than typical field examiners, their accuracy almost certainly overestimated what an average practitioner would achieve. The statistical classifiers, by contrast, are fully standardized and would perform identically regardless of examiner skill, training level, or fatigue.
Importantly, the bootstrap approach demonstrated clear advantages in theoretical robustness. By avoiding the parametric assumptions inherent to discriminant analysis — such as equality of covariance matrices — the bootstrap method was able to operate validly even when those assumptions were violated, which prior research had shown to be a real problem in polygraph datasets.
Discussion & Significance
The study's core finding — that statistical classifiers perform as well as expert human evaluators — carries significant weight for polygraph science. It challenges the notion that the skilled human examiner is irreplaceable in the scoring phase of deception detection, and it opens the door to standardized, objective scoring systems as a legitimate and defensible alternative. In a field often criticized for its reliance on subjective human judgment, this is a meaningful step toward greater scientific rigor.
Viewed alongside the earlier CAPS research of Kircher and Raskin, and subsequent work on computer-assisted scoring systems, this study contributes to a coherent body of evidence supporting the validity of algorithmic polygraph interpretation. The bootstrap approach was particularly forward-looking: its resampling framework anticipated the later explosion of machine learning methodologies that would dominate computational deception detection research in the 2000s and beyond. The study essentially demonstrated that the information necessary to classify deception is objectively present in the physiological record — it does not require human interpretive judgment to extract.
The research also implicitly highlights an important equity argument: if computer scoring performs equivalently to elite experts but far more consistently than average examiners, widespread adoption of automated scoring could reduce examiner-to-examiner variability — one of the longstanding practical criticisms of field polygraph practice. This point about standardization would become increasingly central to polygraph research and policy discussions in subsequent decades.
Limitations & Considerations
Several important limitations temper the strength of the study's conclusions and are worth considering carefully:
- Laboratory analog paradigm: All data came from mock crime experiments, in which subjects enacted a staged theft. The motivational stakes, legal jeopardy, and emotional arousal in real criminal investigations may differ substantially, potentially affecting how well the laboratory-derived classifier equations generalize to the field.
- Unrepresentative human baseline: The expert evaluators were elite researchers — not average practicing polygraph examiners. The study itself acknowledged that human accuracy was likely substantially overestimated relative to typical field performance, which paradoxically strengthens the case for automated scoring but limits direct comparison conclusions.
- Fixed dataset: All analyses drew from the same Utah Cooperative Working Group Database, meaning the classifiers were developed and tested on related datasets. Even with cross-validation steps, full independence of training and test data cannot be assumed, raising overfitting concerns for the discriminant analysis approach in particular.
- CQT-specific findings: The study used exclusively CQT data. Whether multivariate classifiers would perform equally well compared to expert evaluators on other polygraph formats — such as the Concealed Information Test (CIT) — was not addressed.
- Sample characteristics: As with most laboratory polygraph research of the era, subjects were likely predominantly college students or recruited volunteers, potentially limiting generalizability to more diverse populations.
Practical Applications
For practicing polygraph examiners and the agencies that employ them, this research provides early empirical grounding for a message that has only grown stronger over time: objective, computerized scoring systems are a scientifically defensible complement — and potentially a superior alternative — to unaided human scoring. When human scoring accuracy depends heavily on the individual examiner's training, experience, and attentiveness, standardized algorithmic classifiers offer a path to greater consistency and accountability. Tools like the later Objective Scoring System (OSS and OSS2) would build directly on this intellectual foundation.
For consumers of polygraph services and legal professionals evaluating polygraph evidence, the study reinforces an important point: the accuracy of a polygraph outcome is not solely a function of the instrument or the examiner's reputation, but also of the scoring methodology applied to the data. Examinations that incorporate validated computer-assisted scoring systems are operating within a research tradition that has demonstrated empirical parity with expert human judgment — and in absolute performance terms, may even exceed it.
The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.
Related research
Other studies in this category that may be of interest.
Possible Sources of the Difference in the Severity of the Prior Information…
[002] (2025)'Believe your chart but don't ignore your nose'
[003] (2024)Improvement of discrimination method utilizing effect sizes to identify differential physiological responses…
[004] (2024 • COMPUTER SCORING FIELD)Comparison of OSS-3 and PolyScore in Multi-Issue and Single-Issue Field Exams
[005] (2023)Hierarchical Bayesian Modeling of Autonomic Responses during the Concealed Information Test
[006] (2023)Optimization of Signal Processing Parameters in Psychophysiological Studies on the Example of…
Join Our Examiner Network
APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.
Keep reading the ledger.
Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.