Professional Examiners Trained to APA Standards
140+ Professional Testing Locations Across the U.S. & Canada
Trusted by 10,000+ Clients, Attorneys & Organizations
LieDetectorTest.com Private & Confidential Polygraph Provider

Deceptive Reaction Zone in Polygraph Scoring: Complete Guide

Learn how the deceptive reaction zone threshold works in polygraph scoring, including technique-specific cutoffs, 3- and 7-position scales, and automated algorithms.

Published March 27, 2026 Updated July 24, 2026 38 min read All articles

The deceptive reaction zone is a scoring concept examiners use to weigh chart responses, and understanding it clarifies how the numbers behind a lie detector test come together at LieDetectorTest.com.

The deceptive reaction zone is the specific numerical score range where a polygraph examination result crosses from inconclusive territory into a Deception Indicated (DI) determination. This comprehensive guide explains how different validated techniques define this threshold, why it matters for examinees and examiners, and how automated scoring algorithms interact with these critical boundaries.

87–89%Validated Technique Accuracy
3–5Charts Per Exam
7-PositionFull Scale Range
APAGoverning Standards Body
3-PositionMinimum Scale Width

TL;DR — The Short Version

  • The deceptive reaction zone is the specific numerical score range where a polygraph result shifts from inconclusive to Deception Indicated (DI).
  • Technique-specific thresholds — the Federal ZCT, Utah ZCT, Backster ZCT, IZCT, and other validated formats each define different negative score cutoffs for entering the deceptive zone.
  • Numerical scoring assigns values of -3 to +3 (7-position scale) or -1 to +1 (3-position scale) to each relevant-comparison question pair, totaling scores across all charts.
  • Automated algorithms such as PolyScore, OSS-3, and the Empirical Scoring System (ESS) calculate probability values and apply validated cutoffs corresponding to the deceptive reaction zone.
  • The APA's 2011 meta-analysis found an aggregated decision accuracy of 89% for single-issue diagnostic techniques and 87% across all validated techniques combined.
  • Screening vs. diagnostic exam protocols use different cutoff thresholds, balancing sensitivity and specificity for the testing context.

Who This Guide Is For

  • Examinees who received a DI result and want to understand how scores are calculated and what the deceptive reaction zone means for their case
  • Attorneys evaluating polygraph evidence who need to understand scoring thresholds and how different techniques produce different outcomes
  • Polygraph examiner students learning numerical scoring systems and validated cutoff criteria in chart analysis
  • Working examiners seeking a refresher on how the deceptive reaction zone is defined across multiple APA-validated techniques
  • Human resources professionals involved in pre-employment screening who need to understand what DI findings represent
  • Therapists and treatment providers working with PCSOT programs who encounter polygraph scores in client reports

What Is the Deceptive Reaction Zone?

Defining the Core Concept

The deceptive reaction zone is a term used in polygraph chart analysis to describe the specific numerical score range at which a polygraph examination result crosses from an inconclusive outcome into a Deception Indicated (DI) determination. It is not a single number but rather a threshold boundary — a point on the scoring continuum beyond which the accumulated physiological data indicates that the examinee was significantly more responsive to the relevant (crime-related) questions than to the comparison (control) questions.

Understanding this concept requires familiarity with how polygraph scoring works. During a polygraph examination, the examiner collects multiple channels of physiological data — including respiratory patterns, electrodermal activity (EDA), and cardiovascular changes — across multiple chart runs. Each chart is scored by comparing the examinee's physiological reactions to relevant questions against their reactions to comparison questions.

When those scores are totaled across all chart runs and all physiological channels, the resulting grand total determines the examination outcome. If the grand total is sufficiently positive, the result is No Deception Indicated (NDI). If the grand total is sufficiently negative, the result falls within the deceptive reaction zone and the outcome is Deception Indicated (DI). If the score falls within a narrow band near zero, it may be classified as inconclusive. The deceptive reaction zone is the range of scores that falls on the negative (deceptive) side of the inconclusive band.

Why This Threshold Matters

For examinees, the deceptive reaction zone represents the difference between an inconclusive finding — which may allow for retesting or further discussion — and a definitive Deception Indicated result. In high-stakes contexts such as law enforcement pre-employment screening, security clearance vetting, or post-conviction sex offender testing (PCSOT), crossing into the deceptive reaction zone can have significant consequences — from disqualification from a job to changes in supervision conditions.

For examiners and attorneys, understanding where the deceptive reaction zone begins and how it is defined by different validated techniques is essential for proper test administration, quality control, and legal review. Not all scoring systems use the same cutoffs, and the APA's 2011 meta-analytic survey confirmed that the combination of all validated PDD techniques produced a decision accuracy of 87% (confidence interval 80–94%) with an inconclusive rate of 13% [1]Verified Meta-Analytic Survey of Criterion Accuracy of Validated Polygraph Techniques
Confirms 89% accuracy for single-issue diagnostic techniques (CI: 83-95%) and 87% overall accuracy (CI: 80-94%) across all validated polygraph techniques
. Single-issue diagnostic techniques achieved 89% accuracy (CI: 83–95%) [1]Verified Meta-Analytic Survey of Criterion Accuracy of Validated Polygraph Techniques
Confirms 89% accuracy for single-issue diagnostic techniques (CI: 83-95%) and 87% overall accuracy (CI: 80-94%) across all validated polygraph techniques
. A comprehensive 2021 meta-analysis of 138 datasets by Honts and Reavy similarly found that reviews of the Comparison Question Test generally produced overall accuracy estimates of over 85% [2]Verified A Comprehensive Meta-Analysis of the Comparison Question Polygraph Test
Confirms that reviews of the CQT generally produce overall accuracy estimates of over 85% across 138 datasets
.

How Numerical Scoring Creates the Zone

Spot Analysis and Score Assignment

The deceptive reaction zone only makes sense within the context of numerical scoring — the systematic process by which examiners convert raw physiological data into quantifiable values. Modern polygraph scoring begins with what is called spot analysis: the examiner identifies and marks physiological responses at specific locations on the polygraph chart corresponding to relevant questions and comparison questions within the polygraph reaction window.

For each channel of physiological data (typically pneumograph upper, pneumograph lower, EDA, and cardio), the examiner compares the magnitude and duration of the examinee's response to each relevant question against the response to the nearest comparison question. Based on this comparison, a numerical score is assigned to that spot. This process was formalized through the development of objective and empirical scoring methods by Nelson, Handler, and Krapohl, whose work resulted in more objective and stable analysis of psychophysiological outcomes [3]Verified Development of the Objective Scoring System (OSS) and Empirical Scoring System (ESS)
Confirms development and validation of OSS and ESS computerized scoring methods that improved inter-scorer reliability and objectivity
.

The Scoring Scales: How Numbers Map to Reactions

The two most commonly used scoring scales in polygraph practice are the 3-position scale and the 7-position scale [4]Verified PDD Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms definitions of 3-position and 7-position scoring scales and the Backster Either-Or Rule
. The 3-position scale is an abbreviated form of the 7-position scale, where the range of values for each comparison is from -1 to +1 [4]Verified PDD Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms definitions of 3-position and 7-position scoring scales and the Backster Either-Or Rule
.

On the 3-position scale, each comparison yields a score of -1, 0, or +1:

+1 — The comparison question produced a larger response (favorable for the examinee) 0 — No discernible difference between relevant and comparison responses -1 — The relevant question produced a larger response (unfavorable for the examinee)

On the 7-position scale, the range extends from -3 through 0 to +3, allowing finer-grained distinctions in response magnitude [4]Verified PDD Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms definitions of 3-position and 7-position scoring scales and the Backster Either-Or Rule
. By convention, negative values represent greater responding to relevant questions, while positive values indicate greater responses to comparison questions [4]Verified PDD Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms definitions of 3-position and 7-position scoring scales and the Backster Either-Or Rule
. A zero usually indicates equal or no reactions to the relevant and comparison questions.

These individual spot scores are then summed across all physiological channels and all chart runs (typically 3–5 charts are collected per exam) to produce a grand total score for each relevant question and for the exam as a whole.

From Grand Total to Outcome Classification

Once the grand total is calculated, the examiner applies the cutoff criteria specified by the validated technique being used. Every APA-validated technique publishes its own set of decision rules that define three zones on the scoring continuum:

No Deception Indicated (NDI) Zone — Grand total is sufficiently positive (at or above the NDI cutoff), producing a non-deceptive response pattern Inconclusive Zone — Grand total falls between the NDI and DI cutoffs Deception Indicated (DI) Zone — Grand total is sufficiently negative (at or below the DI cutoff)

The deceptive reaction zone is this third band — the range of scores that meet or exceed the negative threshold defined by the technique's validated cutoff rules. It is the point at which the physiological data indicates a pattern of responding statistically associated with deception. Krapohl (2005) published detailed decision rules for evidentiary and paired testing applications, providing precise mathematical boundaries for outcome classification [5]Verified Polygraph Decision Rules for Evidentiary and Paired Testing (Marin Protocol) Applications / A Comparison of 3- and 7-Position Scoring Scales with Laboratory Data
Confirms Krapohl's research on scoring scale comparisons and decision rules for evidentiary testing
.

How Different Techniques Define the Threshold

Technique-Specific Cutoffs: Not One-Size-Fits-All

One of the most important aspects of the deceptive reaction zone is that it is not universally defined. Different APA-validated polygraph techniques use different cutoff scores to determine where the deceptive reaction zone begins. This means the same raw physiological data, scored on the same scale, could potentially produce different outcomes depending on which validated technique the examiner is using.

This is not a flaw in the system — it reflects the reality that each technique was developed, validated, and calibrated against its own set of laboratory and field research data. The cutoffs are empirically derived to optimize the technique's balance between sensitivity (correctly identifying deceptive examinees) and specificity (correctly identifying truthful examinees). The APA's meta-analytic survey confirmed that multiple-issue techniques produced an aggregated decision accuracy of 85% (CI: 77–93%) with an inconclusive rate of 13% [1]Verified Meta-Analytic Survey of Criterion Accuracy of Validated Polygraph Techniques
Confirms 89% accuracy for single-issue diagnostic techniques (CI: 83-95%) and 87% overall accuracy (CI: 80-94%) across all validated polygraph techniques
.

Federal Zone Comparison Test (FZCT)

The Federal ZCT uses the 7-position scoring scale with specific per-spot and grand total cutoffs. It is widely used in U.S. federal agencies and has its own published decision rules for the DI threshold. The deceptive reaction zone is defined by a cumulative negative grand total that exceeds the technique's DI cutoff. Nelson and Krapohl (2011) found no statistically significant differences between the Empirical Scoring System and the seven-position evidentiary model across 13 dimensions of criterion accuracy when applied to Federal ZCT examinations scored by experienced examiners [6]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners: Comparison with the Seven-Position Evidentiary Model Using the Federal Zone Comparison Technique
Confirms no statistically significant differences between ESS and seven-position model across 13 dimensions of criterion accuracy
.

Utah Zone Comparison Test (UZCT)

Developed by the University of Utah research team, the Utah Numerical Scoring System is described as "a highly reliable and valid method for scoring specific-incident, comparison-question tests" [7]Verified The Utah Numerical Scoring System
Confirms the Utah method as a highly reliable and valid method for scoring specific-incident comparison-question tests using the 7-position scale
. For respiration, electrodermal activity, relative blood pressure, and peripheral vasomotor activity, a score from +3 to -3 is assigned for each presentation of a relevant question, compared to a nearby comparison question [7]Verified The Utah Numerical Scoring System
Confirms the Utah method as a highly reliable and valid method for scoring specific-incident comparison-question tests using the 7-position scale
. The Utah approach places significant emphasis on both per-spot and total-score analysis, with the deceptive zone beginning at a defined negative threshold.

Backster Zone Comparison Test

Cleve Backster's technique uses a systematic scoring approach with its own unique cutoff criteria, including the distinctive "Either-Or Rule" (EOR). According to the official APA Terminology Reference, the Either-Or Rule is a "PDD scoring rule forwarded by Cleve Backster and used exclusively in the Backster Zone Comparison Technique" [8]Verified An Assessment of the Backster 'Either-Or' Rule in Polygraph Scoring
Confirms the Either-Or Rule is a real scoring rule used in the Backster technique and that both EOR and Federal methods showed higher accuracy than Backster without EOR
. Under this rule, if a relevant question does not evoke a physiological reaction, it is scored against the adjacent comparison question with the larger reaction [8]Verified An Assessment of the Backster 'Either-Or' Rule in Polygraph Scoring
Confirms the Either-Or Rule is a real scoring rule used in the Backster technique and that both EOR and Federal methods showed higher accuracy than Backster without EOR
.

Research by Meiron, Krapohl, and Ashkenazi (2008) investigated the EOR using 100 Backster field cases blind-scored by field examiners. They found that both the Backster EOR method and the Federal method had higher accuracy levels than the Backster method without the EOR for deceptive cases [9]Verified A Field Study of the Backster Zone Comparison Technique's Either-Or Rule and Scoring System Versus Two Other Scoring Systems When Relevant Question Elicits Strong Response
Confirms comparative analysis of Backster Either-Or Rule versus federal scoring and Matte Quadri-Track system
. Matte (2010) conducted a further field study comparing the Backster Either-Or Rule and scoring system versus two other scoring systems when a relevant question elicits a strong response [10]Verified Setting the Record Concerning the Differences Between the Matte Quadri-Track and the Backster Zone Comparison Techniques
Confirms that the Quadri-Track ZCT became a unique technique with significant changes from the Backster ZCT
. Backster also introduced to the profession the concepts of "psychological set," zones, spots, anticlimax dampening, symptomatic questions, exclusionary comparison questions, and 7-position scoring for use in chart analysis [8]Verified An Assessment of the Backster 'Either-Or' Rule in Polygraph Scoring
Confirms the Either-Or Rule is a real scoring rule used in the Backster technique and that both EOR and Federal methods showed higher accuracy than Backster without EOR
.

Matte Quadri-Track ZCT (MQTZCT)

James Allan Matte's technique adds a fourth "track" (the Inside Track question) and has its own scoring matrix and cutoff values. The Quadri-Track ZCT, while emanating from the Backster Zone Comparison Technique, became a unique polygraph technique with significant changes and improvements [11]Verified The Integrated Zone Comparison Technique: A Field Utility Study in a Deceptive Population
Confirms IZCT achieved 95% overall accuracy including inconclusives and 98% excluding inconclusives with zero false negatives
. The deceptive reaction zone boundaries were validated specifically for this four-question format.

Integrated Zone Comparison Technique (IZCT)

The IZCT, developed by Nathan J. Gordon, has demonstrated strong accuracy across multiple studies. In a field utility study by Patton (2013), the IZCT achieved 95% overall accuracy including inconclusives and 98% accuracy excluding inconclusives, with zero false negatives [12]Verified Integrated Zone Comparison Polygraph Technique Accuracy with Scoring Algorithms
Confirms 100% accuracy for IZCT with all three algorithms when inconclusives excluded; 90% with ASIT Poly Suite and 72% with PolyScore/OSS when inconclusives counted as errors
. Gordon et al. (2006) found that when inconclusives were excluded, accuracy for the IZCT with all three tested algorithms (ASIT Poly Suite, PolyScore 5.5, and OSS) was 100% [13]Verified Effectiveness of the Integrated Zone Comparison Technique (IZCT) with Various Scoring Systems in a Mock Crime Experiment by Students
Confirms systematic comparison of three scoring approaches for the IZCT in controlled conditions where ground truth was known
. When inconclusives were counted as errors, ASIT Poly Suite achieved 90% accuracy while PolyScore and OSS both achieved 72% [13]Verified Effectiveness of the Integrated Zone Comparison Technique (IZCT) with Various Scoring Systems in a Mock Crime Experiment by Students
Confirms systematic comparison of three scoring approaches for the IZCT in controlled conditions where ground truth was known
. A separate mock crime experiment by Gordon and Fleisher (2012) systematically compared three scoring approaches for the IZCT where ground truth was known [14]Verified The Comparison Question Polygraph Test: A Contrast of Methods and Scoring
Confirms substantial main effects of guilt in both OSS2 computer and human scoring for probable-lie and directed-lie CQT variants with 250 participants
.

Other Validated Formats

The Directed Lie Screening Test (DLST) substitutes directed-lie comparison questions for probable-lie comparisons. Its deceptive reaction zone cutoffs were calibrated for the different psychological dynamics of directed-lie formats. Honts and Reavy (2015) conducted a mock crime experiment with 250 paid participants contrasting probable-lie and directed-lie CQT variants, finding substantial main effects of guilt in both OSS2 computer and human scoring [15]Verified A Comparison of 3- and 7-Position Scoring Scales with Field Examinations
Confirms field comparison of 3-position and 7-position scoring scales
.

The Air Force Modified General Question Test (AFMGQT) is a multi-issue screening format that uses its own scoring methodology with cutoffs designed for the broader scope of questions typical of pre-employment or counterintelligence screening examinations.

The key takeaway is that the deceptive reaction zone is technique-dependent, and examiners must apply the correct cutoff criteria for the specific validated technique they are using. This is why APA Standards of Practice require examiners to document which validated technique was employed in each examination.

3-Position vs. 7-Position Scoring Scales

Why the Scale Width Matters for the DRZ

The choice between the 3-position and 7-position scoring scale has direct implications for the deceptive reaction zone. On the 3-position scale (-1, 0, +1), the mathematical range of possible grand totals is narrower, meaning the inconclusive band and the deceptive zone are compressed into a smaller range of integers. On the 7-position scale (-3 through +3), the finer granularity allows for a wider spread of possible scores and, consequently, more precisely defined zone boundaries.

Donald Krapohl's 1998 study, "A comparison of 3- and 7-position scoring scales with laboratory data," directly addressed these differences [5]Verified Polygraph Decision Rules for Evidentiary and Paired Testing (Marin Protocol) Applications / A Comparison of 3- and 7-Position Scoring Scales with Laboratory Data
Confirms Krapohl's research on scoring scale comparisons and decision rules for evidentiary testing
. Harwell (2000) also compared 3- and 7-position scoring scales with field examinations [16]Verified Empirical Scoring System: A Cross-Cultural Replication and Extension Study
Confirms international validation with 19 trainees achieving mean accuracy of 90.1% (95% CI: 83.8–95.8%) across cultures
. The official APA Terminology Reference describes the 3-position scale as an "abbreviated form of the 7-position scale" based on the principle that physiological reactions of greater magnitude are caused by stimuli that are more salient to the examinee [4]Verified PDD Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms definitions of 3-position and 7-position scoring scales and the Backster Either-Or Rule
.

Practical Implications and Research Findings

Research has shown that the 7-position scale generally provides better inter-scorer reliability and improved classification accuracy compared to the 3-position scale. Nelson and Krapohl (2011) found that more than 50 percent of subtotal scores were identical across the ESS and seven-position evidentiary model [6]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners: Comparison with the Seven-Position Evidentiary Model Using the Federal Zone Comparison Technique
Confirms no statistically significant differences between ESS and seven-position model across 13 dimensions of criterion accuracy
, and no statistically significant differences in false-positive or false-negative errors were found between the ESS, 7-position, 3-position, or unweighted ESS models [6]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners: Comparison with the Seven-Position Evidentiary Model Using the Federal Zone Comparison Technique
Confirms no statistically significant differences between ESS and seven-position model across 13 dimensions of criterion accuracy
.

However, some validated techniques — particularly older formats — were originally calibrated using the 3-position scale, and their cutoff criteria were validated at that resolution. Using a 7-position scale with cutoffs designed for 3-position scoring (or vice versa) would be a serious methodological error. The deceptive reaction zone cutoffs are inseparable from the scoring scale used in their validation.

Modern APA-validated techniques generally favor the 7-position scale, and most current automated scoring algorithms are designed to work with 7-position manual scores or produce their own probability-based outputs. The cross-cultural ESS validation by Handler et al. (2011) — involving 19 international trainees scoring 100 confirmed cases — achieved a mean accuracy of 90.1% (95% CI: 83.8–95.8%), demonstrating that evidence-based scoring systems work reliably across different examiner populations [17]Verified The Polygraph and Lie Detection (Appendix F: Computerized Scoring of Polygraph Data)
Confirms PolyScore developed at Johns Hopkins APL and CPS developed at University of Utah as the two most prominent computerized scoring systems
.

The Inconclusive-to-Deceptive Boundary Explained

The Nature of the Inconclusive Zone

The inconclusive zone exists because physiological data is inherently variable. Not every response difference between relevant and comparison questions is meaningful — some differences may be due to random physiological fluctuations, minor artifacts, or marginal attentional shifts related to the fight-or-flight response. The inconclusive zone is essentially a buffer where the data does not point clearly in either direction.

The boundary between inconclusive and the deceptive reaction zone is typically the most consequential line on the scoring continuum. Scores just above this line result in an inconclusive finding — the examinee may have the opportunity for retesting. Scores just below this line enter the deceptive reaction zone and result in a DI determination. The APA meta-analysis reported inconclusive rates of 11% for single-issue techniques and 13% for multi-issue techniques [1]Verified Meta-Analytic Survey of Criterion Accuracy of Validated Polygraph Techniques
Confirms 89% accuracy for single-issue diagnostic techniques (CI: 83-95%) and 87% overall accuracy (CI: 80-94%) across all validated polygraph techniques
, indicating that most examinations produce definitive results.

Borderline Scores and Cumulative Patterns

Scores at or very near the inconclusive-to-DI boundary are the most challenging to interpret with confidence. A grand total of, say, -5 when the DI cutoff is -6 would remain inconclusive — the examiner cannot call it deceptive even though the data trends negative. But a grand total of -6 or -7 crosses the line and enters the deceptive reaction zone.

This is precisely why validated technique protocols specify exact cutoff values rather than leaving the boundary to examiner judgment. The cutoffs were derived from research that optimized the probability of correct classification at that boundary point. Altering the cutoffs would change the technique's balance between false positive and false negative error rates.

When we say a grand total must be "sufficiently negative" to enter the deceptive reaction zone, we mean that the accumulated weight of evidence across all physiological channels and all chart runs consistently points toward greater reactivity to relevant questions. A single negative spot score on one channel in one chart would not be enough. The deceptive reaction zone requires a pattern of negative scoring that persists across the examination data set. This cumulative requirement is one of the key safeguards against false positive results. Only consistent, systematic differences in reactivity to relevant versus comparison questions will accumulate to the point of crossing the DI threshold.

Automated Scoring Algorithms and the Deceptive Reaction Zone

How Algorithms Define Their Own Thresholds

Modern polygraph practice increasingly relies on automated scoring algorithms — software systems that analyze the raw physiological data and produce classification outputs without (or in addition to) human manual scoring. Research has shown that some automated data analysis algorithms can meet or exceed human experts in polygraph decision making [3]Verified Development of the Objective Scoring System (OSS) and Empirical Scoring System (ESS)
Confirms development and validation of OSS and ESS computerized scoring methods that improved inter-scorer reliability and objectivity
. A key advantage of automated data analysis is reliability — the reproducibility of analytic results [3]Verified Development of the Objective Scoring System (OSS) and Empirical Scoring System (ESS)
Confirms development and validation of OSS and ESS computerized scoring methods that improved inter-scorer reliability and objectivity
.

These algorithms do not directly replicate the manual numerical scoring process. Instead, they apply statistical models — often derived from logistic regression, discriminant function analysis, or other pattern-recognition techniques — to extract features from the physiological waveforms and calculate a probability of deception. This probability value is then compared against the algorithm's own validated cutoff thresholds to classify the result as NDI, Inconclusive, or DI.

PolyScore (Johns Hopkins APL)

PolyScore was developed by the Johns Hopkins University Applied Physics Laboratory [18]Verified Development of a Deep-Learning-Based Computerized Scoring Algorithm
Confirms PolyScore uses linear logistic regression and CPS uses linear discriminant analysis for deception detection
. As described by the National Research Council (2003), PolyScore is one of the two most prominent computerized scoring systems alongside the CPS system from the University of Utah [18]Verified Development of a Deep-Learning-Based Computerized Scoring Algorithm
Confirms PolyScore uses linear logistic regression and CPS uses linear discriminant analysis for deception detection
. PolyScore uses linear logistic regression for deception detection [19]Verified Objective Scoring System, Version 3 (OSS-3): Development and Validation
Confirms OSS-3 algorithm description and validation using brute-force methods, demonstrating balanced sensitivity/specificity and outperforming 10 human scorers
and applies a probability threshold to classify results. Probability values in the "deceptive direction" place the result in the algorithm's deceptive reaction zone. PolyScore standardizes data using medians and interquartile ranges rather than means and standard deviations to handle outliers in physiological data [18]Verified Development of a Deep-Learning-Based Computerized Scoring Algorithm
Confirms PolyScore uses linear logistic regression and CPS uses linear discriminant analysis for deception detection
. It is integrated into Lafayette LX software and is used widely in law enforcement and screening contexts.

OSS-3 (Objective Scoring System, Version 3)

The OSS-3 was developed by Raymond Nelson, Donald Krapohl, and Mark Handler as an open-source, objective, and scientifically defensible method for analyzing polygraph data [20]Verified Modern Algorithms in Polygraph Data Analysis
Confirms OSS-3 and PolyScore demonstrated accuracy rates between 85–92% under laboratory conditions
. According to the developers, the OSS-3 algorithm is "based on sound polygraph testing principles derived from existing research" with demonstrable validity across multiple validation samples, including confirmed investigative single-issue polygraphs and cross-validation samples from field investigations [20]Verified Modern Algorithms in Polygraph Data Analysis
Confirms OSS-3 and PolyScore demonstrated accuracy rates between 85–92% under laboratory conditions
.

The OSS-3 demonstrated balanced sensitivity and specificity and outperformed 10 human scorers in a Monte Carlo brute-force comparison study [20]Verified Modern Algorithms in Polygraph Data Analysis
Confirms OSS-3 and PolyScore demonstrated accuracy rates between 85–92% under laboratory conditions
. The computer algorithm provides perfect reliability — the same data will always produce the same result [20]Verified Modern Algorithms in Polygraph Data Analysis
Confirms OSS-3 and PolyScore demonstrated accuracy rates between 85–92% under laboratory conditions
. OSS-3 and PolyScore have demonstrated accuracy rates between 85–92% under laboratory conditions [21]Verified CPSpro Fusion Software
Confirms CPSpro Fusion is Stoelting's polygraph software platform for conducting, scoring, and reporting examinations
. The development and validation of the OSS and ESS represented a paradigm shift toward more objective analysis, replacing subjective hand-scoring and significantly improving inter-scorer reliability [3]Verified Development of the Objective Scoring System (OSS) and Empirical Scoring System (ESS)
Confirms development and validation of OSS and ESS computerized scoring methods that improved inter-scorer reliability and objectivity
.

CPS and CPSpro Fusion (Stoelting / University of Utah)

The Computerized Polygraph System (CPS) was originally developed based on research at the University of Utah Psychology Laboratory [18]Verified Development of a Deep-Learning-Based Computerized Scoring Algorithm
Confirms PolyScore uses linear logistic regression and CPS uses linear discriminant analysis for deception detection
. The CPS scoring algorithm is based on standard multivariate linear discriminant function analysis [18]Verified Development of a Deep-Learning-Based Computerized Scoring Algorithm
Confirms PolyScore uses linear logistic regression and CPS uses linear discriminant analysis for deception detection
. Stoelting's current software platform, CPSpro Fusion, is designed to help examiners conduct, score, and report polygraph examinations [22]Verified CPS Elite Polygraph Systems
Confirms Stoelting has manufactured polygraph instrumentation since 1930 and that Elite systems meet or exceed ASTM standards
. The software includes a question editor, smart subject search, dynamic import capabilities, and numerical scoring functions [22]Verified CPS Elite Polygraph Systems
Confirms Stoelting has manufactured polygraph instrumentation since 1930 and that Elite systems meet or exceed ASTM standards
. Stoelting has been manufacturing polygraph instrumentation since 1930 [23]Verified A Replication and Validation Study on an Empirically Based Manual Scoring System
Confirms ESS replication with independent sample achieving results consistent with original validation studies
, and its CPS Elite system meets or exceeds all ASTM standards for polygraph instrumentation [23]Verified A Replication and Validation Study on an Empirically Based Manual Scoring System
Confirms ESS replication with independent sample achieving results consistent with original validation studies
.

Note that CPSpro Fusion is primarily a comprehensive polygraph software platform that includes manual scoring tools and reporting features, rather than functioning solely as a standalone automated scoring algorithm like PolyScore or OSS-3.

Convergence Between Manual and Automated Scoring

In practice, there is generally strong convergence between manual numerical scoring and automated algorithm outputs — both methods tend to classify the same examinations the same way in the large majority of cases. However, discrepancies can occur, particularly for examinations where the physiological data is marginal or borderline.

APA Standards require examiners to document both manual and automated scores when both are available. When the two methods disagree, the examiner must follow the protocol specified by the validated technique being used. The Blalock, Cushman, and Nelson (2009) replication study confirmed the reliability of the ESS scoring method with an independent sample of confirmed cases, producing results consistent with original validation studies [24]Verified The Polygraph and Lie Detection (Executive Summary)
Confirms NRC finding that polygraph accuracy for screening is almost certainly lower than specific-incident testing
.

Screening vs. Diagnostic Exam Thresholds

Why the Same Data May Produce Different Outcomes

One frequently overlooked aspect of the deceptive reaction zone is that it can shift depending on whether the examination is classified as a screening exam or a diagnostic (evidentiary) exam. This distinction has significant practical implications, particularly in government hiring and security clearance contexts.

Screening examinations cast a wide net — they are used when the base rate of deception in the population being tested is expected to be relatively low. Screening protocols typically use more sensitive (liberal) cutoffs, meaning it is easier for a score to cross into the deceptive reaction zone. This prioritizes identifying potential concerns at the cost of more false positive results among truthful examinees. The NRC (2003) specifically noted that "polygraph accuracy for screening purposes is almost certainly lower than what can be achieved by specific-incident polygraph tests in the field" [25]Verified E2229 Standard Practices for Interpretation of Psychophysiological Detection of Deception (Polygraph) Data
Confirms ASTM E2229 establishes global and numerical procedures for systematic interpretation of PDD data
.

Diagnostic (evidentiary) examinations are used in specific-incident investigations where a particular event has occurred and specific individuals are being tested. Diagnostic protocols typically use more specific (conservative) cutoffs, making it harder for a score to cross into the deceptive reaction zone. This prioritizes protecting truthful examinees from false accusations. Krapohl (2005) published detailed decision rules addressing both evidentiary and paired testing (Marin Protocol) applications [5]Verified Polygraph Decision Rules for Evidentiary and Paired Testing (Marin Protocol) Applications / A Comparison of 3- and 7-Position Scoring Scales with Laboratory Data
Confirms Krapohl's research on scoring scale comparisons and decision rules for evidentiary testing
.

The Sensitivity-Specificity Trade-Off

This relationship between the deceptive reaction zone threshold and the sensitivity/specificity balance is fundamental to understanding polygraph accuracy metrics. Moving the DI cutoff threshold in one direction improves one type of accuracy at the expense of the other.

The APA meta-analysis demonstrated this dynamic clearly: single-issue diagnostic techniques achieved 89% accuracy with an 11% inconclusive rate, while multi-issue techniques (more common in screening) achieved 85% accuracy with a 13% inconclusive rate [1]Verified Meta-Analytic Survey of Criterion Accuracy of Validated Polygraph Techniques
Confirms 89% accuracy for single-issue diagnostic techniques (CI: 83-95%) and 87% overall accuracy (CI: 80-94%) across all validated polygraph techniques
. The Integrated Zone Comparison Technique provided a particularly compelling example — Patton's (2013) field study showed 95% overall accuracy with only a 4% inconclusive rate [12]Verified Integrated Zone Comparison Polygraph Technique Accuracy with Scoring Algorithms
Confirms 100% accuracy for IZCT with all three algorithms when inconclusives excluded; 90% with ASIT Poly Suite and 72% with PolyScore/OSS when inconclusives counted as errors
, while Gordon et al. (2006) found that algorithm performance varied depending on how inconclusives were handled [13]Verified Effectiveness of the Integrated Zone Comparison Technique (IZCT) with Various Scoring Systems in a Mock Crime Experiment by Students
Confirms systematic comparison of three scoring approaches for the IZCT in controlled conditions where ground truth was known
.

Examinees facing screening exams should be aware that the more liberal thresholds can increase the chance of a false DI finding. If you are nervous about your upcoming exam, understanding this distinction can provide important context.

The Examiner's Role in Score Interpretation

Professional Judgment Within Validated Boundaries

While the deceptive reaction zone is defined by mathematical cutoffs, the polygraph examiner plays a critical role in ensuring that the scoring process is conducted correctly and that the resulting classification is valid. Examiners must correctly identify and mark physiological responses, properly pair relevant and comparison questions, and apply the precise scoring criteria specified by the technique.

ASTM International has published multiple standards governing polygraph practice, including E2229 (Standard Practices for Interpretation of Psychophysiological Detection of Deception Data), which establishes "both the global and the numerical procedures for the systematic interpretation and analysis" of PDD data [26]Verified ASTM Forensic Science Standards (Forensic Psychophysiology Section)
Confirms existence of multiple ASTM standards for polygraph including E2439, E2063, E2031, E2035, and E2229
. Additional ASTM standards address instrumentation requirements (E2439), calibration and functionality checks (E2063), quality control procedures (E2031), and examiner requirements (E2065) [27]Verified APA Standards of Practice / Polygraph Validity Research
Confirms APA is the world's leading association for evidence-based credibility assessment with 2,700+ members
. These standards mandate that examiners use the analysis method for which they have been formally trained, and that procedures be correctly matched to the examination format [26]Verified ASTM Forensic Science Standards (Forensic Psychophysiology Section)
Confirms existence of multiple ASTM standards for polygraph including E2439, E2063, E2031, E2035, and E2229
.

The American Polygraph Association, established in 1966, serves as the world's leading professional association, promoting the highest standards of ethical and scientific practices for its 2,700+ members through established standards for techniques, instrumentation, analysis, research, and training [28]Verified E2031 Standard Practice for Quality Control of Psychophysiological Detection of Deception (Polygraph) Examinations
Confirms ASTM E2031 establishes quality control procedures for PDD examinations
.

Quality Control and Blind Scoring

Quality control review is essential for maintaining the integrity of deceptive reaction zone determinations. ASTM Standard E2031 establishes procedures for quality control review of PDD examinations, stating that experienced examiners trained in the PDD format utilized by the original examiner are qualified to conduct such reviews.

Blind scoring — where a second examiner evaluates charts without knowledge of the original examiner's findings or case details — provides an additional layer of reliability. The cross-cultural ESS validation study found that 19 international trainees achieved a mean accuracy of 90.1% when blind-scoring 100 confirmed cases [17]Verified The Polygraph and Lie Detection (Appendix F: Computerized Scoring of Polygraph Data)
Confirms PolyScore developed at Johns Hopkins APL and CPS developed at University of Utah as the two most prominent computerized scoring systems
, demonstrating that well-designed scoring systems produce consistent results across different examiner populations and cultural contexts.

Common Misconceptions About the Deceptive Reaction Zone

Addressing Frequent Misunderstandings

Misconception: The DRZ is the same across all polygraph techniques. Reality: Different validated techniques use different cutoff values, and these thresholds were empirically calibrated for each technique's specific format and question structure.

Misconception: Automated algorithms always agree with manual scoring. Reality: While convergence is typical, borderline cases can produce discrepancies. The OSS-3 demonstrated balanced sensitivity and specificity [20]Verified Modern Algorithms in Polygraph Data Analysis
Confirms OSS-3 and PolyScore demonstrated accuracy rates between 85–92% under laboratory conditions
, but even sophisticated algorithms handle marginal data differently than human scorers. The computer algorithm provides perfect reliability [20]Verified Modern Algorithms in Polygraph Data Analysis
Confirms OSS-3 and PolyScore demonstrated accuracy rates between 85–92% under laboratory conditions
, but human examiners bring contextual understanding that software cannot replicate.

Misconception: A DI result is absolute proof of deception. Reality: A DI finding is a probabilistic determination. The APA meta-analysis found 87% overall accuracy [1]Verified Meta-Analytic Survey of Criterion Accuracy of Validated Polygraph Techniques
Confirms 89% accuracy for single-issue diagnostic techniques (CI: 83-95%) and 87% overall accuracy (CI: 80-94%) across all validated polygraph techniques
, which means that while the vast majority of DI results correctly identify deception, no test achieves perfection. If your score falls within the deceptive reaction zone, you may want to understand what happens next.

Misconception: Nervousness alone can push you into the DRZ. Reality: General anxiety affects the fight-or-flight response but comparison questions are specifically designed to create a baseline of arousal. Only consistent, differential reactivity to relevant questions — not general nervousness — accumulates toward the DI threshold. Certain medications may affect physiological responses, which is why examiners assess fitness for testing beforehand.

Misconception: The spontaneous fluctuation rate will automatically trigger a DRZ finding. Reality: Spontaneous fluctuations in electrodermal activity are normal and accounted for in scoring. Only patterned differences between relevant and comparison question responses contribute to the grand total.

What to Do If Your Score Falls in the Deceptive Reaction Zone

Practical Steps and Options

If your examination result falls within the deceptive reaction zone, several options may be available depending on the context:

Request the examiner explain which validated technique was used and what cutoff criteria define the DRZ for that technique. APA Standards require examiners to document the technique employed in each examination.

Ask whether both manual and automated scoring were conducted. If the two methods produced different outcomes, inquire about which method takes priority under the technique's protocol.

Consider whether a quality control review by a second qualified examiner might be appropriate. ASTM E2031 establishes procedures for such reviews.

In some contexts, retesting may be available. Inconclusive results explicitly allow for retesting, and even DI results can sometimes be reassessed under appropriate circumstances, particularly through high-stakes examination protocols.

Consult an attorney if the results will affect legal proceedings. Understanding the specific scoring methodology used — including whether the examination followed European or American standards — can be relevant in legal challenges.

Remember that a single polygraph examination is one data point. Polygraph results should be considered alongside other evidence and are often used in combination with forensic interviewing and investigative techniques.

Frequently Asked Questions

What is the deceptive reaction zone in polygraph scoring?

The deceptive reaction zone is the specific numerical score range on a polygraph scoring continuum where the accumulated physiological data crosses from an inconclusive outcome into a Deception Indicated (DI) determination. It is defined by validated cutoff thresholds that are specific to each polygraph technique.

Is the deceptive reaction zone the same for all polygraph techniques?

No. Different APA-validated polygraph techniques — including the Federal ZCT, Utah ZCT, Backster ZCT, Integrated Zone Comparison Technique, and others — each define different numerical cutoffs for the deceptive reaction zone. The same physiological data could theoretically produce different outcomes depending on which technique's cutoff criteria are applied.

What is the difference between the 3-position and 7-position scoring scales?

The 3-position scale assigns values of -1, 0, or +1 to each relevant-comparison question pair. The 7-position scale extends from -3 to +3, allowing finer distinctions in response magnitude. Krapohl (1998) directly compared these scales, and research generally shows that the 7-position scale provides better inter-scorer reliability and more precisely defined zone boundaries.

How accurate are polygraph scoring techniques?

The APA's 2011 meta-analytic survey found that single-issue diagnostic techniques achieved 89% accuracy (CI: 83–95%) and all validated techniques combined achieved 87% accuracy (CI: 80–94%). The IZCT demonstrated up to 95% accuracy excluding inconclusives in field studies. Honts' 2021 comprehensive meta-analysis of 138 datasets confirmed that reviews generally produce accuracy estimates over 85%.

What is Backster's Either-Or Rule?

The Either-Or Rule (EOR) is a scoring rule developed by Cleve Backster and used exclusively in the Backster Zone Comparison Technique. According to the rule, if a relevant question does not evoke a physiological reaction, it is scored against the adjacent comparison question with the larger reaction. Research by Meiron, Krapohl, and Ashkenazi (2008) found that the EOR improved accuracy for deceptive cases.

Can automated scoring algorithms produce different results than manual scoring?

Yes. While there is generally strong convergence, discrepancies can occur with borderline data. Automated systems like PolyScore (Johns Hopkins APL), OSS-3 (Nelson, Handler, Krapohl), and CPS (University of Utah) each use their own statistical models and validated cutoffs. APA Standards require documentation of both manual and automated scores when available.

Why do screening and diagnostic exams use different thresholds?

Screening exams use more sensitive (liberal) cutoffs to cast a wider net, prioritizing threat detection at the cost of more false positives. Diagnostic exams use more specific (conservative) cutoffs to protect truthful examinees. The NRC (2003) noted that screening accuracy is almost certainly lower than specific-incident testing accuracy.

What should I do if I receive a Deception Indicated result?

Request an explanation of which validated technique and scoring system were used. Ask whether both manual and automated scoring were conducted. Consider whether a quality control review by a second qualified examiner might be appropriate under ASTM E2031 standards. In some contexts, retesting may be an option, and consulting an attorney is advisable if the results affect legal proceedings.

Can nervousness cause me to enter the deceptive reaction zone?

General nervousness alone is unlikely to produce a DI result. Comparison questions are specifically designed to capture baseline arousal and anxiety. Only consistent, differential reactivity to relevant questions — not general nervousness — accumulates toward the DI threshold. The comparison question methodology accounts for the normal stress of being tested.

What role does ASTM International play in polygraph scoring standards?

ASTM International has published multiple standards governing polygraph practice under Committee E52, including E2229 for interpretation of PDD data, E2439 for instrumentation requirements, E2063 for calibration, E2031 for quality control, and E2035 for terminology. These standards provide the scientific and procedural framework within which scoring thresholds and the deceptive reaction zone are applied.

Sources & References

1

Confirms 89% accuracy for single-issue diagnostic techniques (CI: 83-95%) and 87% overall accuracy (CI: 80-94%) across all validated polygraph techniques

2

Confirms that reviews of the CQT generally produce overall accuracy estimates of over 85% across 138 datasets

3

Confirms development and validation of OSS and ESS computerized scoring methods that improved inter-scorer reliability and objectivity

4

Confirms definitions of 3-position and 7-position scoring scales and the Backster Either-Or Rule

5

Confirms Krapohl's research on scoring scale comparisons and decision rules for evidentiary testing

6

Confirms no statistically significant differences between ESS and seven-position model across 13 dimensions of criterion accuracy

7
The Utah Numerical Scoring System
Brian G. Bell, David C. Raskin, Charles Robert Honts, John C. Kircher (1999) — Polygraph
Verified

Confirms the Utah method as a highly reliable and valid method for scoring specific-incident comparison-question tests using the 7-position scale

8

Confirms the Either-Or Rule is a real scoring rule used in the Backster technique and that both EOR and Federal methods showed higher accuracy than Backster without EOR

9

Confirms comparative analysis of Backster Either-Or Rule versus federal scoring and Matte Quadri-Track system

10

Confirms that the Quadri-Track ZCT became a unique technique with significant changes from the Backster ZCT

11

Confirms IZCT achieved 95% overall accuracy including inconclusives and 98% excluding inconclusives with zero false negatives

12
Integrated Zone Comparison Polygraph Technique Accuracy with Scoring Algorithms
Nathan J. Gordon, Feroze B. Mohamed, Scott H. Faro, Steven M. Platek, Harris Ahmad, J. Michael Williams (2006) — Physiology & Behavior
Verified

Confirms 100% accuracy for IZCT with all three algorithms when inconclusives excluded; 90% with ASIT Poly Suite and 72% with PolyScore/OSS when inconclusives counted as errors

13

Confirms systematic comparison of three scoring approaches for the IZCT in controlled conditions where ground truth was known

14

Confirms substantial main effects of guilt in both OSS2 computer and human scoring for probable-lie and directed-lie CQT variants with 250 participants

15

Confirms field comparison of 3-position and 7-position scoring scales

16

Confirms international validation with 19 trainees achieving mean accuracy of 90.1% (95% CI: 83.8–95.8%) across cultures

17
The Polygraph and Lie Detection (Appendix F: Computerized Scoring of Polygraph Data)
National Research Council (2003) — National Academies Press
Verified

Confirms PolyScore developed at Johns Hopkins APL and CPS developed at University of Utah as the two most prominent computerized scoring systems

18

Confirms PolyScore uses linear logistic regression and CPS uses linear discriminant analysis for deception detection

19

Confirms OSS-3 algorithm description and validation using brute-force methods, demonstrating balanced sensitivity/specificity and outperforming 10 human scorers

20
Modern Algorithms in Polygraph Data AnalysisVerified

Confirms OSS-3 and PolyScore demonstrated accuracy rates between 85–92% under laboratory conditions

21

Confirms CPSpro Fusion is Stoelting's polygraph software platform for conducting, scoring, and reporting examinations

22

Confirms Stoelting has manufactured polygraph instrumentation since 1930 and that Elite systems meet or exceed ASTM standards

23
A Replication and Validation Study on an Empirically Based Manual Scoring System
Benjamin L. Blalock, Barry Cushman, Raymond Nelson (2009) — Polygraph
Verified

Confirms ESS replication with independent sample achieving results consistent with original validation studies

24
The Polygraph and Lie Detection (Executive Summary)
National Research Council (2003) — National Academies Press
Verified

Confirms NRC finding that polygraph accuracy for screening is almost certainly lower than specific-incident testing

25

Confirms ASTM E2229 establishes global and numerical procedures for systematic interpretation of PDD data

26

Confirms existence of multiple ASTM standards for polygraph including E2439, E2063, E2031, E2035, and E2229

27

Confirms APA is the world's leading association for evidence-based credibility assessment with 2,700+ members

28

Confirms ASTM E2031 establishes quality control procedures for PDD examinations

Need to book now? Our online booking system is open 24/7. Speak directly with our team about your test or booking.