Home › Polygraph Research › Extended Analysis with ESS and OSS-3 on USAF-MGQT…
Catalogue entry · Scoring & Test Data Analysis
Extended Analysis with ESS and OSS-3 on USAF-MGQT Examination Data
Raymond Nelson; Benjamin L. Blalock — Polygraph,
Applied ESS and OSS-3 to USAF examination data. Found ESS component weighting achieved superior diagnostic extraction compared to simpler three-position scoring models.
Abstract
Application of ESS and OSS-3 to USAF examination data, finding ESS component weighting achieved superior diagnostic extraction.
Methodology
Scoring method development or validation study using confirmed examination data. Accuracy assessed against known outcomes.
Detailed summary
Published in 2016 in Polygraph, 45(1), 22–46, this work by Nelson and colleagues examined extended analysis with ess and oss-3 on usaf-mgqt examination data. Key findings include: Applied ESS and OSS-3 to USAF examination data. Found ESS component weighting achieved superior diagnostic extraction compared to simpler three-position scoring models. This work contributed to the ongoing improvement of polygraph scoring methods, a critical factor in testing accuracy and reliability.
Implications for polygraph practice
The superior performance of empirically weighted scoring supports continued refinement of scoring algorithms based on empirical data.
Comprehensive study analysis
An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.
Background & Context
The scoring system used to analyze polygraph data represents one of the most critical factors influencing test accuracy. The Nelson and Blalock 2016 study utilized archival laboratory data from a previous USAF Modified General Question Technique study, applying multivariate analysis to compare total and subtotal scores from both laboratory and field investigation samples. This research directly addressed a fundamental question in polygraph science: whether empirically weighted scoring algorithms extract more diagnostic information than simpler numerical approaches.
The Empirical Scoring System (ESS) employs component weighting—assigning different values to different physiological measures based on their empirical diagnostic power. In contrast, the Objective Scoring System version 3 (OSS-3) and traditional three-position models treat measurements more uniformly. The study provided comprehensive criterion accuracy profiles including mean, standard deviations, and statistical confidence intervals for both ESS and OSS-3 test data analysis models. This comparative validation using the same dataset was essential for determining which approach optimally balanced accuracy with practical field application.
Research Design & Methodology
The researchers utilized archival data from a previous laboratory study to provide extended analysis of validation data for the USAF MGQT format for psychophysiological detection of deception (PDD) testing. The USAF-MGQT is a modern variant of the Comparison Question Test that has gained widespread use due to its efficient structure and adaptability to both investigative and screening contexts.
The sample consisted of N = 66 USAF-MGQT examinations scored using both manual ESS and computerized OSS-3 algorithms. The study employed multivariate statistical analysis to compare the performance characteristics of these scoring approaches. Key methodological features included:
- Multivariate analysis comparing total and subtotal scores between laboratory samples (N=66) and field investigation cases (N=22) using the same technique
- Calculation of comprehensive criterion accuracy profiles for both scoring systems
- Statistical confidence intervals calculated for decision accuracy with and without inconclusive results, false-positive and false-negative error rates, and inconclusive rates for both deceptive and truthful cases
- Direct comparison of diagnostic extraction between empirically weighted ESS and simpler scoring models
Results & Key Findings
The ESS scoring system demonstrated superior diagnostic extraction compared to simpler three-position scoring models, validating the hypothesis that component weighting based on empirical data improves test accuracy. The multivariate analysis found no significant differences between total and subtotal scores of the laboratory sample and those from the field investigation sample using the same technique, establishing the generalizability of laboratory findings to operational settings.
The field sample (N = 22) produced subtotal scores of greater absolute value than the laboratory sample (N = 66) when subjected to automated measurement (p < .001). This finding suggests that real-world examinations may produce even stronger physiological differentiation than controlled laboratory conditions. Key statistical outcomes included:
- Comprehensive criterion accuracy profiles documented sensitivity, specificity, predictive values, and error rates for both scoring systems
- ESS component weighting achieved statistically significant improvements in diagnostic information extraction
- The dimensional profile included false-positive and false-negative errors, positive and negative predictive values, percent correct decisions for deceptive and truthful cases, and unweighted means
- Both scoring systems performed significantly better than chance, but ESS demonstrated measurable advantages
Discussion & Significance
This study provides empirical support for a critical theoretical principle in polygraph science: that scoring systems incorporating empirical weights based on the actual diagnostic power of different physiological measures outperform simpler, unweighted approaches. The finding that ESS achieved superior diagnostic extraction has important implications for the ongoing refinement of test data analysis protocols across forensic and security applications.
The lack of significant differences between laboratory and field samples represents an important validation of experimental polygraph research. Critics have long questioned whether laboratory findings translate to operational settings, but this study demonstrates consistency across contexts. The data obtained support the validity of the hypothesis that the USAF-MGQT can provide high levels of criterion accuracy and differentiate truthful from deceptive examinees at rates statistically significantly greater than chance (p < .01) using both ESS and OSS-3 models.
The superior performance of empirically weighted scoring represents more than just incremental improvement—it validates a scientific approach to polygraph technique development. Rather than relying on tradition or intuition, the ESS demonstrates that data-driven optimization of scoring algorithms produces measurable gains in accuracy. This evidence-based approach aligns with broader movements toward empirically validated forensic science methods.
Limitations & Considerations
While the study provides valuable comparative data, the sample size of N = 66 for the laboratory dataset and N = 22 for the field sample represents a modest cohort for definitive conclusions about scoring system superiority. Replication with larger datasets across diverse populations and examination contexts would strengthen confidence in these findings.
The study relied on archival data from confirmed cases, which while providing ground truth criterion accuracy, may not fully represent the distribution of truth and deception encountered in operational practice. Additionally, the research focused specifically on the USAF-MGQT format; additional research is warranted regarding decision rules, normative data, and statistically optimal cutscores for manually scoring the USAF-MGQT. Generalization to other polygraph techniques requires empirical demonstration.
Practical Applications
The findings have direct implications for polygraph examiners and agencies selecting scoring protocols. The demonstrated superiority of ESS component weighting suggests that agencies should prioritize training in and adoption of empirically weighted scoring systems over simpler alternatives. For examiners, this research validates the additional complexity of ESS procedures—the improved diagnostic extraction justifies the investment in learning more sophisticated scoring approaches.
For consumers of polygraph services and legal decision-makers evaluating polygraph evidence, these findings highlight the importance of asking which scoring system was employed. Not all polygraph examinations are created equal; the test data analysis method significantly impacts accuracy. Continued interest in the ESS and seven-position models is recommended in both research and field-practice settings, suggesting that evidence-based scoring refinement should remain a priority for the profession as polygraph technology and understanding of psychophysiological responses continue to evolve.
The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.
References in our database [3]
Studies cited by this paper that are available in our research database.
Related research
Other studies in this category that may be of interest.
Possible Sources of the Difference in the Severity of the Prior Information…
[002] (2025)'Believe your chart but don't ignore your nose'
[003] (2024)Improvement of discrimination method utilizing effect sizes to identify differential physiological responses…
[004] (2024 • COMPUTER SCORING FIELD)Comparison of OSS-3 and PolyScore in Multi-Issue and Single-Issue Field Exams
[005] (2023)Hierarchical Bayesian Modeling of Autonomic Responses during the Concealed Information Test
[006] (2023)Optimization of Signal Processing Parameters in Psychophysiological Studies on the Example of…
Join Our Examiner Network
APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.
Keep reading the ledger.
Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.