Professional Examiners Trained to APA Standards
140+ Professional Testing Locations Across the U.S. & Canada
Trusted by 10,000+ Clients, Attorneys & Organizations
LieDetectorTest.com Private & Confidential Polygraph Provider
Research ledger
Reliability of the Empirical Scoring System with expert examiners

HomePolygraph Research › Reliability of the Empirical Scoring System with expert…

Catalogue entry · Scoring & Test Data Analysis

Reliability of the Empirical Scoring System with expert examiners

Benjamin L. Blalock — Polygraph,

2011Published
25 experienced examiners scoring 10 confirmed examinationsSample size
1Cited by
Key findings

Fleiss kappa of .61 (95% CI: .54–.68) indicated substantial inter-rater agreement on numerical scores, while mean decision agreement reached 95.4% excluding inconclusive results, significantly exceeding chance.

Abstract

This 2011 study by Nelson, Blalock, Oelrich, and Cushman examined the inter-rater reliability and accuracy of the Empirical Scoring System (ESS) when applied by 25 experienced polygraph examiners to 10 confirmed Federal Zone Comparison Technique examinations. Using Monte Carlo and Bootstrap statistical methods, the study found substantial inter-rater agreement and a high rate of diagnostic decision consistency among expert scorers.

Methodology

Twenty-five experienced examiners independently scored 10 confirmed Federal ZCT polygraph examinations using the ESS; Bootstrap Monte Carlo methods were used to calculate accuracy profiles and confidence intervals for both reliability and accuracy.

Detailed summary

This study examined whether the Empirical Scoring System (ESS) could produce consistent results among experienced polygraph examiners, addressing fundamental concerns about scoring reliability in polygraph science. Twenty-five seasoned professionals independently scored 10 confirmed Federal Zone Comparison Technique cases using the ESS without knowing ground-truth outcomes. The research employed rigorous statistical methods including Monte Carlo analysis and Fleiss kappa calculations. Results showed remarkable consistency with a 95.4% decision agreement rate among examiners. The findings suggest that structured, evidence-based scoring systems like the ESS can effectively standardize what has traditionally been a subjective process.

Implications for polygraph practice

The high reliability demonstrated supports the use of structured scoring systems like ESS in operational polygraph practice and suggests that standardized approaches can reduce examiner subjectivity while maintaining expert-level decision accuracy.

Comprehensive study analysis

An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.

Background & Context

The reliability of chart-scoring methods has long been considered a cornerstone concern in polygraph science. If different examiners cannot produce consistent scores from the same physiological data, the diagnostic value of any polygraph examination is fundamentally undermined. This study investigated the reliability and accuracy of the Empirical Scoring System (ESS) with experienced examiners — a question that sits at the heart of whether structured scoring systems can standardize an otherwise subjective process.

The Empirical Scoring System emerged as a response to longstanding criticism that traditional polygraph scoring relied too heavily on examiner intuition. The ESS is described as an evidence-based system used for manual test data analysis in comparison question test formats used in polygraph testing. Its development built upon growing momentum within the field to create scoring approaches grounded in normative statistical data rather than informal heuristics.

By 2011, a body of preliminary ESS validation literature had already begun to accumulate. Earlier work had replicated and validated the hand-scoring system for comparison question polygraph examinations originally proposed by Nelson, Krapohl, and Handler (2008). The present study extended this line of inquiry by specifically isolating the variable of examiner expertise, asking whether seasoned professionals applying the ESS to confirmed Federal ZCT cases would demonstrate measurable inter-rater agreement and decision consistency.

Research Design & Methodology

The study was authored by Raymond Nelson, Benjamin Blalock, Marty Oelrich, and Barry Cushman, and published in Polygraph, Volume 40, Issue 3, 2011, pages 131–139. The research employed a blind-scoring paradigm in which experienced polygraph examiners independently evaluated a set of confirmed polygraph examinations without knowledge of ground-truth outcomes.

Key design elements included:

  • Participants: A cohort of 25 experienced examiners served as scorers, constituting a professionally seasoned sample representative of operational field practice.
  • Stimulus Materials: A small sample of 10 confirmed psychophysiological detection of deception (PDD, polygraph) exams that were conducted using the Federal ZCT format was used as the scoring corpus.
  • Scoring System: All participants applied the Empirical Scoring System (ESS), an evidence-based, numerically anchored manual scoring approach.
  • Statistical Methods: Monte Carlo statistical methods were used to calculate confidence intervals for reliability and accuracy, and Bootstrap Monte Carlo methods were used to calculate the accuracy profile and statistical confidence intervals of the ESS scores.
  • Reliability Metric: Fleiss kappa was used to quantify inter-rater agreement across all 25 examiners simultaneously.

The use of confirmed cases — where ground truth (deceptive or truthful) was independently established — allowed the researchers to evaluate not only whether examiners agreed with each other, but also whether their collective decisions were accurate. The Federal Zone Comparison Technique (ZCT) format was chosen because it represents a widely used, standardized examination structure in federal law enforcement polygraph programs.

Results & Key Findings

The headline result was a mean decision agreement rate of 95.4% among the 25 experienced examiners, a figure that substantially exceeded chance-level performance. This finding indicates that when expert practitioners apply the ESS, they converge on the same diagnostic conclusions at a very high rate.

Key statistical outcomes included:

  • Inter-rater agreement (Fleiss kappa): Fleiss kappa showed a substantial agreement between the numerical scores of the study participants (k = .61, 95% CI: .54 to .68) — a level classified by convention as "substantial" on the Landis-Koch scale.
  • Decision agreement: Decision agreement was significantly better than chance, with a mean rate of decision agreement of 95.4%, excluding inconclusive results.
  • Accuracy profiling: Bootstrap Monte Carlo methods were used to calculate the accuracy profile and statistical confidence intervals of the ESS scores from the experienced examiners.
  • Examiner qualification: The cohort of 25 experienced examiners represents a real-world operational population, lending ecological validity to the findings.

The kappa value of k = .61 is particularly meaningful in context: it reflects agreement at the level of raw numerical scores, which carry more granular variability than simple three-category decisions (DI/NDI/Inconclusive). The much higher decision agreement rate of 95.4% demonstrates that even when scorers assign slightly different numerical values, their categorical conclusions remain highly consistent — a practically important distinction for operational polygraph use.

Discussion & Significance

These findings carry significant weight in the ongoing scientific debate about the standardization of polygraph scoring. One of the most persistent critiques of traditional polygraph practice has been that chart interpretation is subjective and therefore unreliable across examiners. The present study provides quantitative evidence that a structured, empirically derived scoring system like the ESS can produce high inter-rater reliability when applied by experienced practitioners.

The results align with and extend earlier ESS validation work. Bootstrap analysis in a related cross-cultural study resulted in a mean accuracy rate of 90.1% (95% CI = 83.8% to 95.8%), excluding inconclusives, demonstrating that the ESS performs robustly across different examiner populations and cultural contexts. The 95.4% decision agreement found here with expert examiners suggests that experience may further optimize ESS application. The authors recommend continued interest in the ESS as an evidence-based model for manual test data analysis in field polygraph settings and future research.

The study also contributes to a broader understanding of how scoring system design interacts with examiner expertise. By demonstrating that a relatively small set of confirmed cases (n = 10 exams) produced stable and significant reliability estimates through Monte Carlo simulation, the authors demonstrated that probabilistic bootstrapping methods can extract meaningful psychometric information even from modestly sized samples — a methodological contribution with implications beyond polygraph research.

Limitations & Considerations

Several methodological constraints should be considered when interpreting these findings:

  • Small stimulus sample: Only 10 confirmed examinations were used as scoring materials, limiting statistical power and the generalizability of accuracy estimates to the full population of Federal ZCT cases.
  • Expert-only cohort: The sample consisted exclusively of experienced examiners, meaning results may not generalize to novice or trainee scorers. Related research suggests trainee agreement rates are meaningfully lower.
  • Format restriction: All examinations were conducted in the Federal ZCT format; reliability of the ESS in other comparison question formats (e.g., Utah ZCT, MGQT) would require separate investigation.
  • Confirmed case availability: As with most polygraph field research, ground-truth confirmation methods (confession, corroborating evidence) introduce their own biases and may not be representative of all operational cases.
  • Publication venue: The study was published in the Polygraph journal, the official publication of the American Polygraph Association — a point critics sometimes raise regarding potential field-advocacy bias, though the study employed rigorous statistical methods.

The bootstrap and Monte Carlo approach, while appropriate for small-sample psychometric estimation, also introduces assumptions about distributional properties of scores. Future research with larger confirmed case archives would strengthen confidence in the reliability estimates reported here.

Practical Applications

For working polygraph examiners, this study offers meaningful reassurance that the ESS can function as a standardized, reliable scoring tool in field settings. A decision agreement rate exceeding 95% among expert scorers means that ESS-trained examiners are highly likely to reach the same diagnostic conclusion from the same physiological data — a prerequisite for defensible, legally and administratively credible polygraph outcomes. The ESS provides a procedural description for all aspects of the scoring model, including physiological features, mathematical transformations, decision rules, and cutscores based on normative data, giving examiners a transparent and replicable framework.

For polygraph consumers, administrators, and legal professionals evaluating the evidentiary weight of polygraph findings, this research supports the argument that structured, empirically derived scoring systems produce results that are not merely the idiosyncratic judgment of a single examiner. When properly trained examiners apply the ESS to Federal ZCT data, the resulting decisions reflect a systematic, reproducible analytical process — one that compares favorably to reliability benchmarks in many other forensic assessment domains.

Read the original study

The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.

Cited by [1]

Other studies in our database that reference this paper.

Related research

Other studies in this category that may be of interest.

Join Our Examiner Network

APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.

Apply now →

Keep reading the ledger.

Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.

Need to book now? Our online booking system is open 24/7. Speak directly with our team about your test or booking.