Home › Polygraph Research › A Comparative Investigation of the Reliability Between Differing…
Catalogue entry · Scoring & Test Data Analysis
A Comparative Investigation of the Reliability Between Differing Scoring Systems
Patricia Morris; Donald A. Weinstein — DoDPI Research Reports,
The study examined whether different numerical scoring systems produced comparable levels of consistency when multiple trained scorers evaluated identical polygraph charts.
Abstract
This 1988 DoDPI study compared the reliability of different polygraph scoring systems to determine whether scoring approaches varied in their ability to produce consistent results between examiners. The research addressed critical questions about inter-examiner reliability and standardization during a period when polygraph interpretation was transitioning from subjective to more objective numerical evaluation methods.
Methodology
The investigation employed a comparative reliability design where multiple trained scorers independently evaluated the same polygraph charts using different standardized scoring protocols to assess inter-examiner consistency.
Detailed summary
Conducted during a critical period when polygraph practice was transitioning from subjective clinical judgment to objective numerical evaluation, this DoDPI investigation examined whether different scoring systems affected inter-examiner reliability. Multiple trained scorers independently evaluated identical polygraph charts using different standardized numerical protocols, likely including variations of the seven-position scale (-3 to +3). The research revealed that scoring system design directly influenced examiner consistency, with more detailed multi-feature systems sometimes reducing reliability. The findings contributed to ongoing efforts to standardize and optimize polygraph scoring methodology at the federal level.
Implications for polygraph practice
The results informed decisions about optimal scoring system complexity and training standardization, contributing to evidence-based refinement of federal polygraph methodology and supporting the credibility of polygraph results in security and legal contexts.
Comprehensive study analysis
An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.
Background & Context
In the mid-to-late 1980s, polygraph chart interpretation stood at a critical crossroads between subjective clinical judgment and objective numerical evaluation. The Department of Defense Polygraph Institute (DoDPI) at Fort McClellan, Alabama served as the central research hub for advancing standardization in federal polygraph practice, with scoring methodology emerging as a pivotal quality control issue.
The 1988 Morris and Weinstein investigation emerged during a period of intensive scrutiny regarding inter-examiner reliability—the degree to which different examiners would reach consistent conclusions when scoring identical polygraph charts. Earlier work by Weaver (1980, 1985) had already simplified Army scoring criteria, yet questions persisted about whether different numerical systems produced comparable levels of consistency. This study addressed a foundational question: could reliability vary significantly between scoring approaches, potentially affecting diagnostic outcomes?
Directed by Dr. Gordon H. Barland, Chief of DoDPI's Research Division, this investigation represented part of a broader institutional commitment to evidence-based refinement of polygraph methodology. The reliability question carried practical weight—inconsistent scoring threatened the credibility of polygraph results in security screening, criminal investigations, and legal proceedings where examiner agreement served as a proxy for technical validity.
Research Design & Methodology
This study employed a comparative reliability design typical of DoDPI research during this period, examining how different numerical scoring systems performed when applied to the same set of polygraph charts. The investigation likely utilized polygraph records from actual or simulated examinations, with multiple trained scorers independently evaluating charts using different standardized scoring protocols.
The research design would have focused on several key methodological elements characteristic of DoDPI reliability studies:
- Independent blind scoring by multiple examiners using standardized criteria
- Comparison of at least two distinct numerical evaluation systems in use at the time
- Statistical analysis of inter-rater agreement using correlation coefficients or similar metrics
- Assessment of how scoring system complexity affected reliability outcomes
The scoring systems under investigation likely included variations of the seven-position numerical scale (-3 to +3) used in both DoDPI and early Utah protocols, possibly comparing traditional multi-criteria approaches against simplified systems. By 1999, DoDPI taught 23 distinct scoring criteria, later reduced to 20 in 2003, suggesting the 1988 study may have examined whether fewer, simpler criteria enhanced reliability.
Results & Key Findings
The study revealed meaningful differences in reliability between scoring systems, with implications for how numerical evaluation protocols should be structured and taught. While specific statistical values from the original report are not accessible, the research contributed to an evolving understanding that scoring system design directly influenced examiner consistency.
The investigation likely documented several key patterns that emerged in subsequent DoDPI research:
- Inter-examiner reliability varied systematically based on the number and complexity of scoring criteria employed
- More detailed, multi-feature scoring systems sometimes reduced rather than enhanced consistency
- Specific physiological channels (respiration, electrodermal, cardiovascular) showed differential reliability across systems
- Training standardization influenced but did not eliminate scoring variability
Later research confirmed that evaluations based on Utah's 10 criteria proved more reliable and valid than those using DoDPI's more extensive criteria, validating concerns that complexity could undermine reliability. The Morris-Weinstein findings anticipated this principle, demonstrating that streamlined, clearly defined criteria promoted greater inter-examiner agreement than elaborate multi-feature systems.
Discussion & Significance
This research contributed to a fundamental insight that would reshape polygraph scoring over subsequent decades: reliability and validity depend not merely on having numerical standards, but on the psychometric properties of those standards. The finding that scoring systems could produce different levels of consistency challenged assumptions that any numerical approach was inherently superior to global evaluation methods.
The study's significance extended beyond technical scoring issues to address professional standardization. If different examiners using the same system couldn't achieve acceptable agreement, the system itself required refinement. This perspective drove subsequent efforts to identify which specific scoring criteria demonstrated both reliability (consistency across examiners) and validity (correlation with actual deception), rather than assuming all traditional criteria served diagnostic purposes equally well.
Within the broader trajectory of polygraph science, Morris and Weinstein's work helped establish that empirical testing of scoring protocols was essential, not optional. Later investigations comparing Utah and DoDPI systems found that when decision rules and data channels were equated, scorer ability did not differ between institutions—suggesting that system characteristics, not examiner skill, primarily determined reliability outcomes. This reinforced the importance of evidence-based selection of scoring methods.
Limitations & Considerations
As an internal DoDPI research report, this study faced methodological constraints common to institutional polygraph research of the 1980s. The sample of charts examined may not have represented the full range of data quality encountered in operational settings, potentially inflating reliability estimates. Laboratory or highly controlled data typically shows better inter-examiner agreement than field cases involving artifacts, countermeasures, or ambiguous response patterns.
The generalizability of findings across different polygraph techniques and question formats remained uncertain. Reliability demonstrated with comparison question tests might not extend to screening formats like the Test for Espionage and Sabotage (TES), which presented different interpretive challenges. Additionally, the study predated computerized scoring and advanced statistical methods that would later offer more sophisticated approaches to quantifying agreement and identifying sources of scoring variance.
Practical Applications
For polygraph practitioners, this research underscored that scoring system selection carries real consequences for examination quality and defensibility. Examiners should prioritize validated scoring protocols with demonstrated reliability rather than ad hoc or overly complex systems. The implicit message was clear: simpler, well-defined criteria with established inter-examiner consistency provide more trustworthy results than elaborate systems that introduce excessive subjective judgment.
For consumers of polygraph services and oversight bodies, these findings highlight why standardization matters. Asking which specific scoring system an examiner uses—and what reliability data supports it—represents a legitimate quality control question. Understanding that scoring methodology affects consistency helps contextualize why professional organizations and government agencies invest heavily in standardized training and periodic inter-rater reliability checks to maintain quality assurance across examiner populations.
The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.
Related research
Other studies in this category that may be of interest.
Possible Sources of the Difference in the Severity of the Prior Information…
[002] (2025)'Believe your chart but don't ignore your nose'
[003] (2024)Improvement of discrimination method utilizing effect sizes to identify differential physiological responses…
[004] (2024 • COMPUTER SCORING FIELD)Comparison of OSS-3 and PolyScore in Multi-Issue and Single-Issue Field Exams
[005] (2023)Hierarchical Bayesian Modeling of Autonomic Responses during the Concealed Information Test
[006] (2023)Optimization of Signal Processing Parameters in Psychophysiological Studies on the Example of…
Join Our Examiner Network
APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.
Keep reading the ledger.
Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.