Home › Polygraph Research › A Comparison of Computer Programs Designed to Evaluate…
Catalogue entry · Scoring & Test Data Analysis
A Comparison of Computer Programs Designed to Evaluate Psychophysiological Detection of Deception Examinations: Bakeoff 1
Donnie W. Dutton — Department of Defense Polygraph Institute Report,
All five computer scoring programs achieved statistically equivalent accuracy rates of 88–91% on definitive decisions, with no single program significantly outperforming the others; four of five programs exhibited a consistent false positive bias, misclassifying more nondeceptive examinees as deceptive than the reverse.
Abstract
This 1999 Department of Defense Polygraph Institute study — known as 'Bakeoff 1' — conducted the first systematic head-to-head comparison of five commercial computer programs designed to score psychophysiological detection of deception (PDD) examinations, using a blind evaluation of 97 field examinations with verified ground truth. The study found no statistically significant differences in accuracy among the five programs, with definitive-decision accuracy ranging from 88% to 91% across all systems.
Methodology
A blind comparative evaluation in which psychophysiological data from 97 field PDD examinations (56 deceptive, 41 nondeceptive) with investigator-verified ground truth were submitted to five software vendors, who returned deception indicated, no deception indicated, or inconclusive decisions for each case.
Detailed summary
Conducted at the Department of Defense Polygraph Institute, this pioneering study compared five commercial computer scoring systems using field examination data with established ground truth through confessions or corroborating evidence. The blind evaluation included 44 MGQT and 53 ZCT examinations from suspects, witnesses, and victims across diverse demographics. Results showed equivalent performance across all systems, providing critical evidence-based guidance for federal agencies adopting automated scoring technologies. The study established important precedent for systematic software evaluation in operational polygraph settings.
Implications for polygraph practice
This research provided essential evidence-based guidance for federal law enforcement and defense agencies making procurement decisions about computer scoring systems, while establishing methodology for future comparative evaluations of polygraph technologies.
Comprehensive study analysis
An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.
Background & Context
By the late 1990s, the polygraph field was undergoing a significant technological transition. Computerized scoring systems were emerging as potential replacements for — or supplements to — traditional human numerical scoring, with the promise of greater objectivity, standardization, and consistency. Yet despite the proliferation of competing commercial software packages, no systematic, controlled head-to-head comparison of these programs had ever been conducted under controlled conditions.
This preliminary study was completed to determine if accuracy differences exist among computer programs designed to render decisions using data collected during psychophysiological detection of deception (PDD) examinations. Conducted at the Department of Defense Polygraph Institute (DoDPI) — the U.S. government's premier polygraph research institution — this work addressed a critical gap in applied polygraph science at a pivotal moment in the technology's evolution.
The study's colloquial name, "Bakeoff 1," signals its role as a competitive evaluation — akin to industry technology benchmarking — and its designation as the first in an anticipated series. The research laid important groundwork for evidence-based procurement and deployment decisions across federal law enforcement and defense agencies that were increasingly integrating automated scoring tools into operational polygraph workflows.
Research Design & Methodology
The data from 97 PDD examinations (56 deceptive and 41 nondeceptive), for which examinee veracity was known only to the investigators, were sent to software vendors with the request that a decision of deception indicated, no deception indicated, or no opinion (e.g., inconclusive) be returned for each examination using the vendors' software. This blind evaluation design effectively neutralized any potential for vendor bias in scoring.
Key design elements of the study included:
- Five software systems evaluated: AXCON version 1.2, Chart Analysis version 5.1-090-17-097, Computerized Polygraph System version 2.2, Identifi version 1.43, and PolyScore version 4.0.
- Examination formats: The reported sample was composed of 44 MGQT examinations (27 deceptive and 17 nondeceptive) and 53 ZCT examinations (29 deceptive and 24 nondeceptive).
- Examinee roles: Eighty-five examinations were of suspects under investigation, eight were of witnesses, and four were of victims.
- Demographics: Twenty-eight of the examinees were female and 69 were male. Of the 70 examinees for which race was available, 25 were African American, 39 were Caucasian, four were Hispanic, one was a Native American, and one was of South Pacific heritage.
- Age range: The examinee age range, for the 94 examinees for which age data were available, was 14 to 70 years with a mean of 32.6.
Each examination consisted of at least three charts, as verified by a certified government examiner. The government examiner also classified the examinations as either Modified General Question Test (MGQT) or Zone Comparison Test (ZCT) format. Ground truth was established rigorously: the examinations were classified as confirmed deceptive if the examinee signed a confession which included a statement describing the details of the crime, or if there was irrefutable corroborating evidence — a standard appropriate for field-verified research.
Results & Key Findings
The headline result was one of equivalence rather than differentiation: across all five programs, accuracy rates were remarkably similar, and no statistically significant winner emerged.
- Overall correct decision rates (including inconclusives): The proportion of correct, erroneous, and no opinion decisions rendered by the five evaluated computer programs ranged from .71 to .77, .07 to .10, and .13 to .21, respectively.
- Definitive decision accuracy (excluding inconclusives): When no opinion decisions were excluded, the proportion of correct decisions ranged from .88 to .91.
- No significant differences found: There were no statistically significant differences among the frequency of correct decisions, erroneous decisions, and no opinion decisions rendered by the five computer programs.
- No universal inconclusives: There was no single examination in the data set that all of the computer programs classified as no opinion.
- False positive bias: All of the computer programs, except the Computerized Polygraph System, erroneously classified more nondeceptive examinees as deceptive (i.e., false positive) than deceptive examinees as nondeceptive (i.e., false negative).
The 88%–91% accuracy range on definitive decisions is a noteworthy benchmark, suggesting that by 1999, automated computer scoring had achieved meaningful accuracy levels. The consistent false positive bias across most programs — misidentifying innocent individuals as deceptive more often than the reverse — is a critically important operational finding with real-world consequences for examinees.
The Department of Defense Polygraph Institute compared the decision accuracy of five different computer programs. The psychophysiological data from 97 PDD examinations for which examinee veracity was known were sent to the four independent computer program developers with the instructions to provide the final decisions as determined by their software. The vendors were requested to limit their decision to deception indicated (DI), no deception indicated (NDI), or no opinion (inconclusive).
Discussion & Significance
The finding of statistical equivalence among all five programs is significant in multiple dimensions. From a procurement standpoint, it suggested that federal agencies could not justify selecting one software platform over another based solely on accuracy claims — factors like cost, integration, and vendor support might legitimately drive decisions. From a scientific standpoint, the convergence in accuracy levels pointed toward a potential ceiling in what rule-based automated scoring of physiological data could achieve at that time.
The consistent false positive bias observed across four of the five programs deserves particular attention. In operational settings, a tendency to flag innocent examinees as deceptive carries significant consequences — including wrongful suspicion in criminal investigations and potential career consequences in security screening contexts. The fact that only the Computerized Polygraph System deviated from this pattern suggests meaningful algorithmic differences in how that program weighted physiological response thresholds.
The study also appeared as a published article in the journal Polygraph, titled "A comparison of computer programs designed to evaluate psychophysiological detection of deception examinations: Bakeoff 1," published in Polygraph 29(3):237–257. This dual publication — as both a government technical report and a peer-reviewed journal article — reflects the study's importance to both applied and academic polygraph communities. It also established a benchmark dataset and methodology that subsequent "Bakeoff" evaluations could build upon.
Limitations & Considerations
Several methodological considerations temper the interpretation of these findings:
- Sample size constraints: With only 97 examinations total, the study had limited statistical power to detect moderate differences between programs, which may have contributed to the null finding of significant differences.
- Format variability: No attempt was made to delineate among the format variations (e.g., single issue, multi-issue, and multi-facet; US Air Force MGQT vs. US Army MGQT; Backster ZCT vs. DoDPI ZCT; etc.). Mixed test formats may have introduced variability that individual programs handled differently.
- Version-specific results: Each software system was tested at a specific version; findings may not generalize to later releases of the same programs.
- Unequal deceptive/nondeceptive ratio: The sample included 56 deceptive versus 41 nondeceptive cases, a modest imbalance that could influence comparative false positive/negative rates.
- "Preliminary" designation: The authors themselves characterized this as a preliminary study, indicating they recognized its exploratory rather than definitive nature.
It is also worth noting that the study used field examination data, which introduces real-world complexity but also means examiners were not standardized — individual examiner technique may have influenced physiological data quality in ways that differentially affected software performance.
Practical Applications
For polygraph practitioners and the agencies that employ them, this research delivers a clear message: no single commercially available computer scoring program in 1999 demonstrated superior accuracy over its competitors. The 88%–91% definitively-scored accuracy range demonstrated that automated systems were already performing at levels consistent with trained human examiners, supporting their adoption as objective, consistent decision-support tools. The false positive bias finding also serves as an important reminder for operational users to interpret DI outputs with appropriate caution, particularly for nondeceptive examinees.
For consumers and legal professionals evaluating polygraph evidence, the study provides context for understanding automated scoring as an evolving science. The "Bakeoff" series model — blind, comparative, field-data-based evaluations — represents a gold standard approach to technology validation that the field would benefit from applying to emerging scoring algorithms and AI-based analysis tools. The study reinforced the importance of independent, government-sponsored validation research as a counterweight to vendor marketing claims.
The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.
Cited by [1]
Other studies in our database that reference this paper.
Related research
Other studies in this category that may be of interest.
Possible Sources of the Difference in the Severity of the Prior Information…
[002] (2025)'Believe your chart but don't ignore your nose'
[003] (2024)Improvement of discrimination method utilizing effect sizes to identify differential physiological responses…
[004] (2024 • COMPUTER SCORING FIELD)Comparison of OSS-3 and PolyScore in Multi-Issue and Single-Issue Field Exams
[005] (2023)Hierarchical Bayesian Modeling of Autonomic Responses during the Concealed Information Test
[006] (2023)Optimization of Signal Processing Parameters in Psychophysiological Studies on the Example of…
Join Our Examiner Network
APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.
Keep reading the ledger.
Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.