Home › Polygraph Research › Criterion Validity of the Empirical Scoring System with…
Catalogue entry · Scoring & Test Data Analysis
Criterion Validity of the Empirical Scoring System with experienced examiners: Comparison with the seven-position evidentiary model Using the Federal Zone Comparison Technique
Nelson, R., Krapohl, D. — Polygraph,
No statistically significant differences were found between the ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy when scored by experienced examiners. More than 50 percent of subtotal scores among confirmed truthful cases were non-positive values.
Abstract
This study compared the Empirical Scoring System (ESS) to the traditional seven-position evidentiary model using 60 confirmed Federal Zone Comparison Test examinations scored by experienced federally trained examiners. The research aimed to determine whether the simpler ESS could achieve criterion accuracy equivalent to the established seven-position scoring method.
Methodology
Within-subjects design using a cohort of experienced federal examiners who scored 60 confirmed Federal ZCT cases from the DoD archive using both seven-position and ESS scoring systems with two-stage decision rules.
Detailed summary
Nelson and Krapohl (2011) investigated whether the simpler Empirical Scoring System could match the accuracy of the traditional seven-position evidentiary model in the hands of skilled examiners analyzing Federal ZCT cases. Using confirmed cases from a DoD archive and blind scoring by experienced federal examiners, the study applied standardized two-stage decision rules to both methods and evaluated performance across 13 accuracy dimensions. Results demonstrated statistical equivalence between the two scoring systems, with no significant differences in sensitivity, specificity, overall accuracy, or inconclusive rates. The finding that over 50% of subtotal scores were non-positive in truthful cases highlighted important considerations for scoring policy. These results suggest that examiner training and adherence to standardized procedures may be more critical to accuracy than the specific complexity of the scoring system itself.
Implications for polygraph practice
The equivalence of ESS and seven-position scoring methods when used by trained examiners supports the viability of simplified scoring systems in federal polygraph practice and suggests that training programs should prioritize examiner skill development and procedural standardization. The observation regarding non-positive subtotal scores in truthful cases has direct implications for refining decision rules and scoring policies in Federal ZCT protocols.
Comprehensive study analysis
An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.
Background & Context
The development of standardized scoring systems for polygraph test data analysis has been a critical evolution in the field's scientific foundation. By 2011, the polygraph profession had largely transitioned from authority-based scoring practices to evidence-based methodologies, yet questions remained about whether newer scoring approaches could match the performance of traditional methods when applied by skilled examiners.
The Empirical Scoring System (ESS) emerged as a simpler, manual alternative to complex computer algorithms and the widely-used seven-position numerical scale. This study addressed a crucial validation question: could the ESS match the criterion accuracy of the seven-position evidentiary model when both were used by highly trained federal examiners? Understanding whether simplified scoring methods maintain accuracy with expert scorers has direct implications for training requirements, field practice, and the reliability of polygraph results in legal and security contexts.
Research Design & Methodology
The study utilized a cohort of experienced federally trained examiners who evaluated a confirmed case sample (N=60) of Federal ZCT examinations. This within-subjects design allowed direct comparison of scoring methods while controlling for examiner skill and case characteristics. The Federal Zone Comparison Technique is a specialized format commonly used in government polygraph programs, featuring three relevant questions targeting a single event or allegation.
Key methodological features included:
- Confirmed cases from the Department of Defense archive with verified ground truth
- Blind scoring by experienced federal examiners
- Two-stage decision rules applied to both seven-position and ESS scores
- 13-dimensional criterion accuracy profile for comprehensive performance assessment
Seven-position scores were compared to Empirical Scoring System (ESS) scores using two-stage decision rules, providing a rigorous standardized framework for evaluating both methods. The research team assessed multiple accuracy dimensions including sensitivity, specificity, overall accuracy, and inconclusive rates.
Results & Key Findings
The study found no statistically significant differences between the ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy. This equivalence demonstration was the headline finding, indicating that experienced examiners could achieve comparable results with either scoring system.
An important empirical observation emerged regarding score distributions:
- More than 50 percent of subtotal (spot) scores among the confirmed truthful cases were non-positive values
- There were no statistically significant differences between criterion accuracy of the ESS and the seven-position evidentiary model
- Both scoring systems demonstrated robust performance when applied by skilled federal examiners
The finding that truthful examinees frequently produced non-positive subtotal scores has implications for decision rules and scoring policies, particularly regarding the use of spot scores versus grand totals in Federal ZCT examinations.
Discussion & Significance
These findings provided important validation for the ESS as a viable alternative to traditional seven-position scoring in federal polygraph practice. The study demonstrated that scoring system simplicity need not compromise accuracy when examiners are properly trained—a finding with practical implications for examiner training programs and field operations.
The equivalence between scoring methods when used by experienced examiners suggests that examiner skill and adherence to standardized procedures may be more important than the specific numerical scale employed. This research contributed to the growing body of evidence supporting the ESS and helped establish normative data for federal polygraph techniques using evidence-based scoring approaches.
The study also addressed concerns about whether simplified scoring systems could maintain the diagnostic precision achieved with more complex methods. By demonstrating statistical equivalence across multiple accuracy dimensions, the research supported broader adoption of the ESS in government polygraph programs.
Limitations & Considerations
The sample size of 60 examinations, while sufficient for statistical analysis, represents a relatively modest dataset for establishing comprehensive normative standards. The study focused exclusively on experienced federally trained examiners, limiting generalizability to less experienced scorers or examiners trained in different polygraph schools.
The Federal ZCT format examined in this study has specific structural characteristics (three relevant questions, single-issue focus) that may not fully represent the range of polygraph techniques used in field practice. Additionally, the study relied on archival confirmed cases, which may not perfectly represent the full spectrum of real-world examination difficulty and examinee behavior encountered in operational settings.
Practical Applications
This research provided empirical support for federal agencies and polygraph programs considering adoption of the ESS as an alternative or complement to traditional seven-position scoring. The findings suggest that properly trained examiners can achieve comparable accuracy with either system, allowing programs to select scoring methods based on practical considerations such as ease of training, inter-rater reliability, and computational requirements.
For polygraph consumers and legal stakeholders, the study reinforces that examiner training and adherence to validated procedures are critical factors in polygraph accuracy—potentially more important than the specific numerical scoring scale used. The research also contributes to the scientific foundation supporting polygraph admissibility in legal proceedings by demonstrating that evidence-based scoring systems can achieve accuracy levels comparable to traditional methods when properly implemented.
The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.
Related research
Other studies in this category that may be of interest.
Possible Sources of the Difference in the Severity of the Prior Information…
[002] (2025)'Believe your chart but don't ignore your nose'
[003] (2024)Improvement of discrimination method utilizing effect sizes to identify differential physiological responses…
[004] (2024 • COMPUTER SCORING FIELD)Comparison of OSS-3 and PolyScore in Multi-Issue and Single-Issue Field Exams
[005] (2023)Hierarchical Bayesian Modeling of Autonomic Responses during the Concealed Information Test
[006] (2023)Optimization of Signal Processing Parameters in Psychophysiological Studies on the Example of…
Join Our Examiner Network
APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.
Keep reading the ledger.
Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.