Professional Examiners Trained to APA Standards
140+ Professional Testing Locations Across the U.S. & Canada
Trusted by 10,000+ Clients, Attorneys & Organizations
LieDetectorTest.com Private & Confidential Polygraph Provider
Research ledger
An Empirically Based Normative System for Test Data Analysis

HomePolygraph Research › An Empirically Based Normative System for Test Data…

Catalogue entry · Scoring & Test Data Analysis

An Empirically Based Normative System for Test Data Analysis

Raymond Nelson; Donald J. Krapohl; Mark Handler — Polygraph,

2008Published
732 confirmed examinations, 140 examiners, 16 cohortsSample size
Key findings

Introduced the Empirical Scoring System based on normative physiological data. Provided validation across 732 confirmed examinations scored by 140 examiners in 16 cohorts.

Abstract

Introduction of the Empirical Scoring System based on normative physiological data, validated across 732 confirmed examinations scored by 140 examiners.

Methodology

Development of normative scoring weights from large dataset of confirmed examinations. Validation across 16 cohorts with 140 examiners.

Detailed summary

Published in 2008 in Polygraph, 37(1), 1–28, this work by Nelson and colleagues examined an empirically based normative system for test data analysis. Key findings include: Introduced the Empirical Scoring System based on normative physiological data. Provided validation across 732 confirmed examinations scored by 140 examiners in 16 cohorts. This work contributed to the ongoing improvement of polygraph scoring methods, a critical factor in testing accuracy and reliability.

Implications for polygraph practice

The scale of the validation (732 examinations, 140 examiners, 16 cohorts) provides exceptional confidence in ESS reliability and generalizability.

Comprehensive study analysis

An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.

Background & Context

The 2008 introduction of the Empirical Scoring System by Nelson, Krapohl, and Handler represented a paradigm shift in polygraph scoring methodology. For decades, polygraph scoring systems evolved through incremental modifications of legacy techniques developed by practitioners with commercial interests, often lacking rigorous scientific foundations. Unlike most polygraph scoring systems in common practice, ESS was not simply a mutation of an earlier system but began as a zero-based review of the psychophysiological and decision theory literature, with developers assembling rather than inventing the system based on a significant body of scientific findings.

This scoring system, first described by Krapohl, Nelson, and Handler in 2008, allowed for the first time in the development of a polygraph hand-scored technique the application of p-values and normative data. This innovation addressed a critical weakness in polygraph science: the absence of statistically-grounded cutoff scores derived from empirical distributions of actual examination data rather than arbitrary thresholds.

The ESS development filled a significant gap by providing examiners with cut scores based on normative data that allow for the calculation of the probability of an erroneous test result, thus allowing for the selection of statistically optimal cut scores based on operational needs. This represented a fundamental advancement from subjective scoring traditions toward evidence-based practice in forensic psychophysiology.

Research Design & Methodology

The validation study examined 5,192 scored results from 732 confirmed Federal and Utah three-question Zone Comparison Test (ZCT) examinations, with data from six samples scored by 140 experienced and inexperienced scorers in 16 cohorts of study participants. This extraordinarily large dataset provided robust statistical power for establishing normative distributions and validating decision accuracy across diverse examiner populations.

The ESS methodology built upon established psychophysiological principles by identifying which scoring features produced the most diagnostic information, then simplifying the system based on the principle of parsimony, before building a database of normative data on which to establish cutoff scores. The system employed a three-position scoring scale (+1, 0, -1) using the "bigger-is-better" rule, focusing on primary physiological features identified in prior research as most diagnostic of deception.

Key methodological elements included:

  • Use of confirmed cases with established ground truth for criterion validation
  • Application of signal detection theory and normative reference distributions
  • Implementation of two-stage decision rules combining grand total and spot score analyses
  • Electrodermal activity weighting to reflect the superior diagnostic value of this channel
  • Development of p-value tables allowing examiners to select error rates matching operational requirements

Results & Key Findings

Validation data for 13 dimensions of criterion accuracy were shown for the 5,192 scored results, providing one of the most comprehensive accuracy profiles in polygraph validation literature. The massive sample size across 732 confirmed examinations evaluated by 140 examiners in 16 cohorts demonstrated exceptional generalizability across examiner experience levels and testing contexts.

Subsequent replication studies validated the ESS findings with consistent results:

  • Bootstrap analysis in a cross-cultural replication showed a mean accuracy rate of 90.1% (95% CI = 83.8% to 95.8%) with only 3.3% inconclusives (95% CI = 1.0% to 7.0%)
  • Trainee examiners achieved 88% average decision accuracy with 13.1% inconclusives, with sensitivity and specificity levels that did not differ significantly, suggesting balanced accuracy characteristics
  • Inexperienced examiners demonstrated an average rate of agreement of 85% (95% CI = 65-97%)

A critical finding was that 61% (95% CI = 51% to 70%) of the subtotal scores of truthful cases produced a non-positive score, validating the two-stage decision rule approach. The ESS's use of normative cutoffs (Grand Total ≥ +2 for No Deception Indicated; ≤ -4 for Deception Indicated) provided statistically optimal decision thresholds with known error probabilities.

Discussion & Significance

The ESS represented a conceptual breakthrough by introducing actuarial, norm-referenced scoring to polygraph practice. Based on associated p-value tables regarding specificity, sensitivity, and inconclusive rate, the decision maker or policy setter can compare the probability of error and choose the error rate that best fits into his schema for risk aversion. This flexibility allows organizations to calibrate testing protocols to match the consequences of false positive versus false negative errors—a capability unprecedented in manual polygraph scoring.

The system's validation across both experienced and novice examiners addressed longstanding concerns about scorer reliability in polygraph testing. The finding that trainees could achieve accuracy levels meeting evidentiary standards suggested the ESS's simplified, evidence-based approach reduced subjective interpretation compared to traditional seven-position scoring systems. In the time since its introduction in 2008, the ESS/ESS-M has become one of the most widely used methods for polygraph test data analysis throughout the U.S. and other countries.

The research contributed significantly to polygraph's scientific maturation by demonstrating that systematic application of psychometric principles—feature selection based on diagnostic validity, parsimony in scoring rules, and norm-referenced decision thresholds—could produce a scoring system with demonstrable and reproducible accuracy characteristics. This work influenced subsequent developments in evidence-based polygraph practice and provided a model for future scoring system development.

Limitations & Considerations

While the study's scale provided exceptional statistical power, certain methodological considerations merit attention. The validation relied on confirmed cases from Federal and Utah archives, which may not fully represent the base rates and examinee populations encountered in all polygraph contexts. The confirmed case methodology, while providing ground truth, inherently selects for cases where confessions or independent evidence became available, potentially excluding cases that remain ambiguous in real-world practice.

The ESS was specifically designed and validated for single-issue Zone Comparison Test formats with three relevant questions. Generalizability to other test formats, multiple-issue screening examinations, or significantly different questioning techniques required separate validation. Additionally, the system's reliance on manual scoring of physiological features means that inter-rater reliability in feature identification remains a potential source of variance, though the three-position scale likely reduces this compared to seven-position systems.

Practical Applications

For polygraph examiners, the ESS offers a scientifically-grounded scoring approach with known error rates and simplified decision rules. The system's transparency—with published normative tables and explicit decision thresholds—facilitates quality control, training standardization, and defensibility in legal or administrative proceedings. Examiners can select cutoff scores matching their operational context: more conservative thresholds for screening applications where false negatives carry high costs, or balanced thresholds for investigative testing where inconclusive results can be managed through additional investigation.

For consumers and policymakers, the ESS provides a benchmark for evaluating polygraph accuracy claims. The system's validation across 732 examinations and 140 examiners offers empirically-grounded confidence intervals rather than aspirational accuracy estimates. Organizations implementing polygraph testing can reference ESS validation data when establishing quality assurance standards, training requirements, and decision policies. The availability of p-value tables enables evidence-based policy decisions about acceptable error rates, moving polygraph practice from tradition-based to risk-informed decision frameworks.

Read the original study

The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.

Related research

Other studies in this category that may be of interest.

Join Our Examiner Network

APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.

Apply now →

Keep reading the ledger.

Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.

Need to book now? Our online booking system is open 24/7. Speak directly with our team about your test or booking.