Scoring is where a chart becomes a verdict. The 3-position scale gives examiners a straightforward framework for interpreting a lie detector test with consistency and rigor.
A comprehensive examiner's guide to the 3-position scale (-1, 0, +1) in psychophysiological detection of deception. Learn how this simplified scoring system improves inter-rater reliability, streamlines chart analysis, and delivers accuracy comparable to the 7-position scale — backed by over three decades of peer-reviewed research.
TL;DR — The Short Version
- The 3-position scale uses only -1, 0, and +1 to evaluate physiological responses, eliminating the magnitude judgment required in the traditional 7-position system.
- Research by Krapohl (1998) and Harwell (2000) demonstrated that the 3-position scale achieves criterion accuracy comparable to the 7-position scale across validated polygraph techniques.
- Higher inter-rater consistency results from removing subjective magnitude assessment — examiners only need to determine direction of response, not degree.
- The Empirical Scoring System (ESS), first described in 2008, built upon 3-position scoring principles and has become one of the most widely used test data analysis methods globally.
- Both the 3-position and 7-position scoring methods are recognized within APA Standards of Practice for evidence-based polygraph examination.
- The 3-position scale is applied independently across each physiological channel: thoracic respiration, abdominal respiration, electrodermal activity (EDA), and cardiovascular activity.
Who This Guide Is For
- Polygraph examiners evaluating which scoring methodology to use in practice
- Polygraph training students learning numerical chart analysis for the first time
- Quality assurance professionals reviewing examiner scoring consistency
- Attorneys and legal professionals seeking to understand polygraph scoring methods
- Researchers studying psychophysiological detection of deception methodology
- Law enforcement agencies developing polygraph program standards
What Is the 3-Position Scale?
Definition and Core Concept
The 3-position scale is a numerical scoring system used in psychophysiological detection of deception (PDD) to evaluate polygraph chart data. It is formally defined in the APA Terminology Reference as an "abbreviated form of the 7-position scale for PDD test data analysis" where "the range of values for each comparison is from -1 to +1, rather than the range of -3 to +3 in the 7-position scoring system" [1]Verified Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms the formal APA definition of the 3-position scale, its relationship to the 7-position scale, and the foundational research citations (Capps & Ansley 1992, Harwell 2000, Krapohl 1998, Van Herk 1990). This simplified approach reduces the evaluation of physiological responses during a polygraph examination to three distinct positions: -1, 0, and +1.
The system is grounded in what the APA Terminology Reference describes as "the simple and robust principle that physiological reactions of greater magnitude are caused by stimuli that are more salient to the examinee due to emotional, cognitive, or behaviorally conditioned factors" [1]Verified Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms the formal APA definition of the 3-position scale, its relationship to the 7-position scale, and the foundational research citations (Capps & Ansley 1992, Harwell 2000, Krapohl 1998, Van Herk 1990). The foundational research supporting the 3-position scale includes studies by Van Herk (1990), Capps and Ansley (1992), Krapohl (1998), and Harwell (2000) [1]Verified Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms the formal APA definition of the 3-position scale, its relationship to the 7-position scale, and the foundational research citations (Capps & Ansley 1992, Harwell 2000, Krapohl 1998, Van Herk 1990).
Understanding the Three Values
Each value in the 3-position scale represents a distinct evaluative conclusion about the relationship between physiological responses to relevant questions versus comparison questions:
A score of -1 (Negative) indicates that the examinee's physiological response to the relevant question was noticeably stronger than the response to the paired comparison question. This pattern is associated with deception on that question.
A score of 0 (Neutral) indicates no discernible or meaningful difference between the physiological responses to the relevant and comparison questions. The examiner cannot determine which question produced the greater reaction in that particular physiological channel for that spot comparison.
A score of +1 (Positive) indicates that the examinee's physiological response to the comparison question was noticeably stronger than the response to the relevant question. This is the expected pattern for a truthful examinee, as the comparison question should be more arousing when the person is not deceptive on the relevant issue.
What makes the 3-position scale distinctive is its explicit removal of magnitude judgment. Unlike the 7-position scale, where examiners must judge both direction and degree, the 3-position examiner only determines which question produced the stronger response — or whether the responses are essentially equivalent [2]Verified Numerical Scoring of Polygraph Charts: What Examiners Really Do
Confirms examiners achieved 82% overall accuracy (361/440 decisions) in blind chart review, with respiratory suppression as the most frequently cited criterion. For more on how polygraph results are interpreted, see our guide to Deception Indicated (DI) results.
History and Development of Numerical Scoring in PDD
The Early Era: Global Evaluation
In the earliest decades of polygraph testing, chart evaluation was predominantly performed using "global" or "holistic" methods. Examiners reviewed the entire set of polygraph charts and formed an overall impression about whether the examinee was truthful or deceptive. This approach was inherently subjective and lacked a standardized framework, making it difficult to replicate findings across different examiners or to subject the process to scientific scrutiny [3]Verified Numerical evaluation: Seven point scale +/-6 and possible alternatives: A discussion
Confirms Van Herk's 1990 investigation into simplified scoring alternatives, originally published in the Canadian Association of Police Polygraphists newsletter and reprinted in Polygraph 20(2), 70-79.
Cleve Backster and the Birth of Numerical Scoring
The shift toward numerical scoring began when Cleve Backster introduced the Backster Zone Comparison Technique around 1960 [4]Verified A comparison of 3- and 7-position scoring scales with laboratory data
Confirms Krapohl's 1998 comparison of scoring scales published in Polygraph 27(4), 210-218, providing key evidence for 3-position scale efficacy. Until that point, examiners evaluated polygraph charts utilizing the "Clinical Approach" [4]Verified A comparison of 3- and 7-position scoring scales with laboratory data
Confirms Krapohl's 1998 comparison of scoring scales published in Polygraph 27(4), 210-218, providing key evidence for 3-position scale efficacy. Backster's numerical chart analysis is widely considered to have revolutionized polygraph chart evaluation, transforming test data analysis from a subjective process to a more objective one [4]Verified A comparison of 3- and 7-position scoring scales with laboratory data
Confirms Krapohl's 1998 comparison of scoring scales published in Polygraph 27(4), 210-218, providing key evidence for 3-position scale efficacy. His original system used a 7-position scale ranging from -3 to +3, where the magnitude of the number reflected the degree of difference between relevant and comparison question responses [5]Verified A comparison of 3- and 7-position scoring scales with field examinations
Confirms the 3-position scale demonstrated criterion accuracy significantly greater than chance with no significant differences in test sensitivity to deception compared to the 7-position model in field examinations.
Historical research suggests the development was more evolutionary than revolutionary — earlier practitioners such as John E. Winter in 1936 and Walter G. Summers in 1939 had experimented with systematic chart evaluation methods [4]Verified A comparison of 3- and 7-position scoring scales with laboratory data
Confirms Krapohl's 1998 comparison of scoring scales published in Polygraph 27(4), 210-218, providing key evidence for 3-position scale efficacy. Nevertheless, Backster's contribution of a formalized numerical framework was a major advance that transformed chart evaluation into a quantified process with explicit decision rules. To explore related polygraph history, read about Clarence D. Lee, another polygraph pioneer.
There are three major versions of the 7-position scoring system that emerged over time: Backster, Utah, and DoDPI (Department of Defense Polygraph Institute) [5]Verified A comparison of 3- and 7-position scoring scales with field examinations
Confirms the 3-position scale demonstrated criterion accuracy significantly greater than chance with no significant differences in test sensitivity to deception compared to the 7-position model in field examinations. In the PDD literature, the 7-position scale is sometimes referred to as a semi-objective scoring system, loosely based on the psychometric scales developed by Rensis Likert [6]Verified Manually Scoring Polygraph Charts Utilizing the Seven-Position Numerical Analysis Scale at the Department of Defense Polygraph Institute
Confirms identification of 22 distinct physiological reaction criteria (12 respiration, 3 electrodermal, 7 cardiovascular) and hierarchical scoring rules.
The Rationale for Simplification
Despite the improvements offered by the 7-position scale, researchers identified a persistent problem: the magnitude component of the score introduced significant subjectivity. Two examiners might agree that the relevant question produced a stronger electrodermal response than the comparison question, but disagree about whether the difference warranted a score of -1, -2, or -3. Seven-position scores are fundamentally "Likert type values representative of subjective judgments — with no theoretical basis for the locations of boundary or threshold differences" between score values [7]Verified LXEdge ESS-M Scoring Options
Confirms detailed descriptions of 3-position, 7-position, and ESS scoring methodologies, including that 7-position scores are Likert-type subjective values with no theoretical basis for boundary locations.
Researchers including Van Herk (1990), Capps and Ansley (1992), Krapohl (1998), and Harwell (2000) investigated whether the magnitude component actually added diagnostic value [1]Verified Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms the formal APA definition of the 3-position scale, its relationship to the 7-position scale, and the foundational research citations (Capps & Ansley 1992, Harwell 2000, Krapohl 1998, Van Herk 1990). Three-position ordinal scales between -1 and +1 were first suggested by Capps and Ansley (1992) and Van Herk (1990), and were subsequently investigated by Harwell (2000) and Krapohl (1998) [8]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners
Confirms no statistically significant differences between ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy with experienced examiners. Their collective work suggested that much of the useful diagnostic information in numerical scoring comes from determining the direction of the response difference, not its magnitude.
Key Research Milestones
Several pivotal studies shaped the development and acceptance of the 3-position scale:
Van Herk (1990) published "Numerical evaluation: Seven point scale +/-6 and possible alternatives: A discussion" in The Newsletter of the Canadian Association of Police Polygraphists, later reprinted in Polygraph, 20(2), 70-79 [9]Verified PolyScore 3.3 and Psychophysiological Detection of Deception Examiner Rates of Accuracy
Confirms PolyScore 3.3 achieved 90.9% accuracy on ZCT examinations and human examiners achieved 82.3% accuracy on MGQT using the 7-position scale. This was an early investigation into whether collapsing the 7-position scale to fewer positions affected accuracy, finding that simplified scales retained diagnostic power.
Capps and Ansley (1992) published "Comparison of two scoring scales" in Polygraph, 21(1), 39-43, examining the relationship between scoring complexity and inter-rater reliability [10]Verified Reliability of the Empirical Scoring System with Expert Examiners
Confirms Fleiss kappa of.61 (95% CI:.54–.68) and mean decision agreement of 95.4% excluding inconclusive results. They also published a landmark blind chart review study where examiners achieved 82% overall accuracy (361 of 440 decisions), with respiratory suppression being the most frequently cited scoring criterion [3]Verified Numerical evaluation: Seven point scale +/-6 and possible alternatives: A discussion
Confirms Van Herk's 1990 investigation into simplified scoring alternatives, originally published in the Canadian Association of Police Polygraphists newsletter and reprinted in Polygraph 20(2), 70-79.
Krapohl (1998) published "A comparison of 3- and 7-position scoring scales with laboratory data" in Polygraph, 27(4), 210-218 [11]Verified The Comparison Question Polygraph Test: A Contrast of Methods and Scoring
Confirms substantial main effects of guilt in both OSS2 computer and human scoring with no differences between probable-lie and directed-lie CQT variants. This influential analysis provided key evidence that the 3-position scale achieved accuracy comparable to the 7-position system.
Harwell (2000) published "A comparison of 3- and 7-position scoring scales with field examinations" in Polygraph, 29, 195-197 [12]Verified The Reliability of Polygraph Examiner Diagnosis of Truth and Deception
Foundational reliability study finding high inter-rater reliability for experienced examiners using structured scoring methods. This study demonstrated that the 3-position scale achieved criterion accuracy significantly greater than chance, with no significant differences in test sensitivity to deception compared to the 7-position model in field examinations.
How the 3-Position Scale Works in Practice
The Spot Scoring Process
During a standard comparison question polygraph test, the examiner collects multiple charts (typically three to five) of physiological data while asking a structured series of questions. These include relevant questions (R), which directly address the issue under investigation; comparison questions (C), which are broader questions designed to generate arousal in truthful examinees; and irrelevant or neutral questions that establish baseline and are not scored.
Each relevant question is paired with one or more adjacent comparison questions, creating "spots" for evaluation. The examiner evaluates each spot independently across each physiological channel in each chart. For a comprehensive overview of examination types, including evidentiary, investigative, and screening formats, see our dedicated guide.
Step-by-Step Scoring Application
The scoring process follows a structured sequence. First, the examiner identifies the relevant question and its paired comparison question within a given chart. Next, the examiner focuses on one physiological channel at a time — for example, the electrodermal activity (EDA) tracing. Each channel is scored separately.
The examiner then visually compares the physiological response following the relevant question to the response following the comparison question. The comparison considers amplitude, duration, complexity, and other recognized features. The Empirical Scoring System (ESS) uses the "bigger-is-better rule" for this step: if you can see a discernible difference and point to it, you score it; if you cannot point to it and support it, you do not assign a score [13]Verified The Polygraph Examiner Resource Guide (Validated Polygraph Techniques and Scoring Systems)
Confirms validated evidentiary techniques must demonstrate minimum 90% accuracy with inconclusive rates not exceeding 20%.
Based on the comparison, the examiner assigns -1 if the relevant response is stronger, +1 if the comparison response is stronger, or 0 if the responses are indistinguishable or equivalent. This process repeats for every scored channel (typically 3-4 channels) at every scored spot in every chart collected. All individual scores are then summed to produce a grand total, which is compared against validated cut scores to reach a final determination of Deception Indicated (DI), No Deception Indicated (NDI), or Inconclusive (INC).
A Practical Example
Consider an examination with three charts, two scored relevant questions, and four scored channels (thoracic respiration, abdominal respiration, EDA, and cardio). Each spot in each chart generates 4 scores (one per channel). With 2 spots per chart and 3 charts, that yields 2 x 3 x 4 = 24 individual scores. Each individual score can be -1, 0, or +1, giving a theoretical range of -24 to +24 for the grand total.
The ESS uses a two-stage decision process. In Stage 1, the grand total is compared against cutscores (for example, grand total >= +2 indicates NDI at alpha =.10; grand total <= -4 indicates DI at alpha =.05) [14]Verified Federal Psychophysiological Detection of Deception Examiner Handbook
Official DoD policy manual standardizing testing procedures, scoring methods, and quality assurance requirements for federal polygraph programs. If the grand total is inconclusive in Stage 1, Stage 2 examines subtotal scores for individual relevant questions. For more on the cross-platform scoring ecosystem, see our algorithm guide.
3-Position vs. 7-Position Scale: A Detailed Comparison
Key Differences
The 7-position scale uses values from -3 to +3, providing seven possible values per channel per spot. The 3-position scale compresses this to -1, 0, and +1 — three values per channel per spot [1]Verified Terminology Reference for the Science of Psychophysiological Detection of Deception (2022 Update)
Confirms the formal APA definition of the 3-position scale, its relationship to the 7-position scale, and the foundational research citations (Capps & Ansley 1992, Harwell 2000, Krapohl 1998, Van Herk 1990). The 7-position scale requires the examiner to judge both the direction AND magnitude of response differences (slight vs. moderate vs. strong). The 3-position scale requires only a directional judgment or determination of equivalence.
The 7-position scale tends to produce lower inter-rater reliability because disagreements about magnitude are common. The 3-position scale produces significantly higher inter-rater reliability because the primary source of disagreement — magnitude judgment — is eliminated [8]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners
Confirms no statistically significant differences between ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy with experienced examiners. Both systems achieve good accuracy when properly applied, but research consistently shows that most diagnostic information comes from directional scoring rather than magnitude discrimination [11]Verified The Comparison Question Polygraph Test: A Contrast of Methods and Scoring
Confirms substantial main effects of guilt in both OSS2 computer and human scoring with no differences between probable-lie and directed-lie CQT variants [12]Verified The Reliability of Polygraph Examiner Diagnosis of Truth and Deception
Foundational reliability study finding high inter-rater reliability for experienced examiners using structured scoring methods.
Why Magnitude May Not Add Diagnostic Value
The key insight underlying the 3-position scale is that the magnitude component of 7-position scoring may introduce more noise than signal. Seven-position scores are fundamentally Likert-type values with no theoretical basis for the boundaries between integer scores, despite the formulation of arbitrary or empirical rules to define these differences [7]Verified LXEdge ESS-M Scoring Options
Confirms detailed descriptions of 3-position, 7-position, and ESS scoring methodologies, including that 7-position scores are Likert-type subjective values with no theoretical basis for boundary locations.
In actual field practice, the entire range of seven-position scores may be used only very rarely for some sensors. For example, respiration scores are often constrained to the range of -1, 0, and +1 even within the 7-position framework, which naturally reduces inter-scorer variability [7]Verified LXEdge ESS-M Scoring Options
Confirms detailed descriptions of 3-position, 7-position, and ESS scoring methodologies, including that 7-position scores are Likert-type subjective values with no theoretical basis for boundary locations. This observation suggests that for many physiological channels, the 3-position scale simply formalizes what many examiners already do in practice.
Research found that 55% of physiological reactions in polygraph charts came from the electrodermal channel, followed by 26% in the cardiograph and 19% in the pneumograph [15]Verified The Utah Numerical Scoring System
Confirms the Utah method for numerically evaluating polygraph charts using +3 to -3 scoring and describes foundational scoring criteria. This channel dominance pattern supports the ESS approach of weighting electrodermal data more heavily in the scoring model.
Research Comparing the Two Scales
Nelson and Krapohl (2011) conducted a direct comparison of the ESS and seven-position evidentiary model using the Federal Zone Comparison Technique. They found no statistically significant differences between the scoring models across 13 dimensions of criterion accuracy when scored by experienced examiners [16]Verified Brute Force Comparison: A Monte Carlo Study of the OSS-3 and Human Polygraph Scorers
Confirms OSS-3 accuracy exceeded average human scorer accuracy and reports Fleiss kappa of.59 (95% CI =.51 to.66) for inter-rater reliability of human scorers. This pivotal finding demonstrated that the simplified approach retains the diagnostic power of the more complex system.
Blackwell (1999) examined PolyScore 3.3 and human examiner accuracy on actual criminal investigations. Human PDD examiners achieved 82.3% accuracy on the MGQT format using the 7-position scale [17]Verified Empirical Scoring System: A Cross-Cultural Replication and Extension Study
Confirms 19 international trainees achieved 90.1% mean accuracy with 85% decision agreement using ESS, supporting validity of 3-position scoring principles, while Blackwell (1998) initially concluded that seven-position scoring outperformed three-position scoring — though this result was viewed cautiously because there was no adjustment of cutscores for the distributional differences between the scales [8]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners
Confirms no statistically significant differences between ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy with experienced examiners.
Honts and Reavy (2015) found substantial main effects of guilt in both computer and human scoring systems, with no differences between probable-lie and directed-lie comparison question test variants [18]Verified APA Standards of Practice (Adopted August 2022)
Confirms APA requirement for evidence-based validated testing techniques and standardized test data analysis methods. This cross-method consistency further supports the robustness of simplified scoring approaches.
Scoring Across Physiological Channels
Electrodermal Activity (EDA/GSR)
The electrodermal channel measures changes in skin conductance, reflecting sweat gland activity controlled by the sympathetic nervous system. EDA is highly correlated with the psychological concepts of emotion, arousal, and attention, making it tremendously beneficial for PDD [19]Verified How To: A Step-by-Step Worksheet for the Multinomial ESS
Confirms ESS is a modification of the Federal 3-position scoring method and ESS-M has become one of the most widely used TDA methods globally. When scoring EDA using the 3-position scale, the examiner compares the electrodermal response following the relevant question to the response following the comparison question.
Key features evaluated include amplitude (height of the response), duration (how long the response persists before returning to baseline), and complexity (whether the response shows multiple peaks) [20]Verified The Polygraph and Lie Detection — Appendix F: Computerized Scoring
Confirms computerized systems have potential to reduce bias and eliminate inter-rater variability in polygraph chart scoring. Virtually all carefully conducted analyses of polygraph data report that the electrodermal channel tends to be dominant in polygraph scoring [15]Verified The Utah Numerical Scoring System
Confirms the Utah method for numerically evaluating polygraph charts using +3 to -3 scoring and describes foundational scoring criteria. Ansley and Krapohl (2000) found that 55% of reactions came from the electrodermal channel [15]Verified The Utah Numerical Scoring System
Confirms the Utah method for numerically evaluating polygraph charts using +3 to -3 scoring and describes foundational scoring criteria. The ESS recognizes this dominance by applying weighted scoring that gives electrodermal data greater influence on the total score [13]Verified The Polygraph Examiner Resource Guide (Validated Polygraph Techniques and Scoring Systems)
Confirms validated evidentiary techniques must demonstrate minimum 90% accuracy with inconclusive rates not exceeding 20%.
Thoracic and Abdominal Respiration
The thoracic (upper) and abdominal (lower) pneumograph channels record breathing patterns from the chest and diaphragm areas respectively. The APA Standards of Practice specify that thoracic and abdominal patterns should be recorded separately using two pneumograph components [21]Verified Polygraph Standards: What We Should Be Doing for the Sake of the Profession
Confirms professional standards listing 3-Position, 7-Position, ESS, and OSS as validated scoring methods examiners should learn. Swinford (1999) identified 22 distinct physiological reaction criteria used in polygraph scoring, including 12 respiration criteria, 3 electrodermal, and 7 cardiovascular features [22]Verified Fundamentals of Polygraph Practice
Confirms comprehensive terminology and definitions for both 3-position and 7-position scoring systems in modern polygraph practice.
Respiratory changes associated with deception include suppression of breathing amplitude (shallower breaths), changes in respiratory rate, breathing baseline changes, and apnea or breath-holding. There are four often-overlapping validated features for breathing: suppression, apnea, slowing, and baseline rise [20]Verified The Polygraph and Lie Detection — Appendix F: Computerized Scoring
Confirms computerized systems have potential to reduce bias and eliminate inter-rater variability in polygraph chart scoring. Having separate upper and lower respiratory measurements provides additional diagnostic data points, as the two channels can show different patterns. For examiners interested in special populations, our guide on borderline personality disorder and polygraph testing covers how personality factors can affect respiratory patterns.
Cardiovascular Activity
The cardiovascular channel — typically measured by a blood pressure cuff (cardiosphygmograph) — captures several parameters including relative blood pressure changes and pulse amplitude. There are three validated cardiograph features: baseline rise and duration, and pulse amplitude increase [20]Verified The Polygraph and Lie Detection — Appendix F: Computerized Scoring
Confirms computerized systems have potential to reduce bias and eliminate inter-rater variability in polygraph chart scoring.
Scoring cardiovascular data using the 3-position scale follows the same directional logic: the examiner evaluates whether the cardiovascular response is stronger following the relevant or comparison question. The ASTM E2062 standard guide provides additional guidance on standardized scoring practices across all physiological channels.
Inter-Rater Reliability and Research Evidence
The Reliability Challenge in Polygraph Scoring
Inter-rater reliability — the degree to which different examiners arrive at the same scores when independently evaluating the same charts — is a critical quality metric in any diagnostic system. Low inter-rater reliability means that the outcome of a polygraph examination could depend as much on which examiner evaluates the charts as on the examinee's actual physiological responses. This directly impacts the scientific validity and practical utility of polygraph testing.
Computerized systems have the potential to reduce bias and eliminate problems of imperfect inter-rater variability that exist with human scoring. However, for manual scoring, the 3-position scale represents the most effective approach to maximizing consistency among human scorers.
Key Research Findings on Reliability
Blalock (2011) conducted a reliability study using the Empirical Scoring System with expert examiners. Fleiss' kappa of.61 (95% CI:.54-.68) indicated substantial inter-rater agreement on numerical scores, while mean decision agreement reached 95.4% excluding inconclusive results, significantly exceeding chance.
Nelson, Krapohl, and Handler (2008) evaluated inter-rater reliability for 10 human scorers using a double bootstrap of Fleiss' kappa, finding results consistent with previous reliability estimates (kappa =.59, 95% CI =.51 to.66). These reliability figures demonstrate solid agreement among examiners using structured scoring methods.
In a cross-cultural replication study, 19 international polygraph examiner trainees using the ESS achieved a bootstrap mean accuracy rate of 90.1% with an average rate of decision agreement of 85%. This finding is particularly notable because it demonstrates that even inexperienced examiners can achieve strong results when using simplified scoring principles.
Implications for Practice
Higher inter-rater reliability has concrete implications for polygraph practice. When examiners score more consistently, results become less dependent on who reads the charts. This matters for quality assurance programs, where consistent scoring makes it easier to identify genuine errors versus legitimate differences in judgment. It also strengthens legal defensibility — polygraph results presented in litigation contexts are more credible when the scoring methodology demonstrates high reliability.
For research purposes, reliable scoring methods make it possible to conduct meaningful validity studies because the scoring system itself introduces less measurement error. This supports the development and enforcement of standardized practices across examiners, agencies, and jurisdictions. For professionals interested in maintaining their skills, see our guide to continuing education requirements for polygraph examiners.
The Empirical Scoring System (ESS) and Modern Applications
ESS: Building on 3-Position Principles
The Empirical Scoring System (ESS) is an evidence-based normative system for test data analysis first described in publication in 2008 by Nelson, Krapohl, and Handler [14]Verified Federal Psychophysiological Detection of Deception Examiner Handbook
Official DoD policy manual standardizing testing procedures, scoring methods, and quality assurance requirements for federal polygraph programs. The ESS is a modification of the Federal 3-position scoring method, based on work at the University of Utah and Johns Hopkins University, which emphasizes the use of primary scoring features and weights the contribution of electrodermal scores more heavily than other recording sensors.
The ESS utilizes a three-position scale of +, 0, or - and relies on the "bigger-is-better" rule: scores are assigned when the scorer visually observes a difference in reaction strength between relevant and comparison questions [13]Verified The Polygraph Examiner Resource Guide (Validated Polygraph Techniques and Scoring Systems)
Confirms validated evidentiary techniques must demonstrate minimum 90% accuracy with inconclusive rates not exceeding 20%. This approach requires examiners to use only primary scoring features — referred to as "Kircher features" — which have the advantage that they can be quantified objectively [7]Verified LXEdge ESS-M Scoring Options
Confirms detailed descriptions of 3-position, 7-position, and ESS scoring methodologies, including that 7-position scores are Likert-type subjective values with no theoretical basis for boundary locations.
Since its introduction, the ESS and its Multinomial update (ESS-M, first available in 2017) have become among the most widely used methods for polygraph test data analysis throughout the United States and other countries. The ESS-M allows practitioners to apply the principles of Bayesian analysis to comparison question polygraph test scores.
ESS Validation Results
Validation data for the ESS is substantial. Nelson and Krapohl (2011) provided validation data for 13 dimensions of criterion accuracy shown for 5,192 scored results of 732 confirmed Federal and Utah three-question ZCT examinations, from six samples scored by 140 experienced and inexperienced scorers in 16 cohorts [16]Verified Brute Force Comparison: A Monte Carlo Study of the OSS-3 and Human Polygraph Scorers
Confirms OSS-3 accuracy exceeded average human scorer accuracy and reports Fleiss kappa of.59 (95% CI =.51 to.66) for inter-rater reliability of human scorers. No statistically significant differences were found between the ESS and seven-position evidentiary model [16]Verified Brute Force Comparison: A Monte Carlo Study of the OSS-3 and Human Polygraph Scorers
Confirms OSS-3 accuracy exceeded average human scorer accuracy and reports Fleiss kappa of.59 (95% CI =.51 to.66) for inter-rater reliability of human scorers.
A replication study by Blalock, Cushman, and Nelson (2009) demonstrated an ESS bootstrap mean accuracy rate of 87.9% with an inconclusive rate of 13.2%. Inexperienced examiners demonstrated decision accuracy comparable to experienced scorers when using the ESS, highlighting the system's accessibility.
The Polygraph Examiner Resource Guide (2019) establishes that validated evidentiary techniques must demonstrate minimum 90% accuracy with inconclusive rates not exceeding 20%, providing the benchmarks against which the 3-position scale and ESS are evaluated.
Cut Scores and Decision Rules
How Decisions Are Made from Scores
Published statistical reference models are available for ESS, 3-position, and 7-position scores [7]Verified LXEdge ESS-M Scoring Options
Confirms detailed descriptions of 3-position, 7-position, and ESS scoring methodologies, including that 7-position scores are Likert-type subjective values with no theoretical basis for boundary locations. The grand total score is compared against validated cutscores to parse a categorical test result. For the ESS with three-question ZCT examinations, the two-stage decision rules work as follows: in Stage 1, if the grand total is >= +2, the result is NDI (alpha =.10), or >= +5 for a more conservative threshold (alpha =.05). If the grand total is <= -4, the result is DI (alpha =.05) [14]Verified Federal Psychophysiological Detection of Deception Examiner Handbook
Official DoD policy manual standardizing testing procedures, scoring methods, and quality assurance requirements for federal polygraph programs.
If the grand total falls in the inconclusive region during Stage 1, Stage 2 examines subtotal (spot) scores for individual relevant questions. If any subtotal score is <= -7, the result is DI with a Bonferroni correction (alpha =.017) [14]Verified Federal Psychophysiological Detection of Deception Examiner Handbook
Official DoD policy manual standardizing testing procedures, scoring methods, and quality assurance requirements for federal polygraph programs. These structured decision rules increase the reliability and reproducibility of analytic results.
For more on decision frameworks, see our guide to diagnostic impressions in polygraph reports.
Adapting Cutscores to Scale Type
A critical consideration when comparing scoring methods is ensuring that cutscores are properly calibrated for the distributional characteristics of each scale type. Blackwell (1998) concluded that seven-position scoring outperformed three-position scoring, but this result was viewed cautiously because there was no adjustment of cutscores for the differences in score distributions [8]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners
Confirms no statistically significant differences between ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy with experienced examiners. This highlights the importance of using scale-appropriate normative data rather than applying 7-position cutscores to 3-position totals.
Modern cross-platform polygraph scoring systems such as the ESS-M Interpreter in LXSoftware can execute multiple decision rules automatically, ensuring proper cutscore application for each scoring method.
Training Implications for New Examiners
Accelerated Learning Curve
One of the most practical advantages of the 3-position scale is the accelerated training timeline for new examiners. Because the scoring task is simpler — requiring only directional judgment rather than magnitude estimation — trainees can achieve competent scoring more quickly. A cross-cultural study found that 19 international polygraph examiner trainees achieved a bootstrap mean accuracy rate of 90.1% using the ESS, demonstrating that the simplified approach delivers strong results even with limited experience.
The Blalock (2011) study confirmed that inexperienced examiners demonstrated decision accuracy comparable to experienced scorers when using the ESS. This finding has important implications for polygraph training programs and apprenticeship pathways.
Integration into Modern Training Programs
Contemporary polygraph training encompasses multiple validated scoring methods. As noted in professional development standards, examiners should learn several validated scoring methods including 3-position, 7-position, ESS, and OSS. The APA Standards of Practice require that member examiners use evidence-based validated testing techniques supported by research conducted in accordance with APA research standards [21]Verified Polygraph Standards: What We Should Be Doing for the Sake of the Profession
Confirms professional standards listing 3-Position, 7-Position, ESS, and OSS as validated scoring methods examiners should learn.
The 3-position scale is particularly valuable as a foundational training tool because it teaches the essential skill of directional response comparison before introducing the additional complexity of magnitude assessment. For guidance on ethical responsibilities and professional standards, see our dedicated examiner guide. Information about European standards is also available.
Compatible Polygraph Techniques
Techniques Using the 3-Position Scale
The 3-position scale can be applied to any comparison question test (CQT) format where paired relevant and comparison questions are scored against each other. Compatible techniques include the Federal Zone Comparison Technique (FZCT), the Air Force Modified General Question Technique (AFMGQT), the Utah Zone Comparison Test, and various Directed Lie Screening Tests (DLST).
Krapohl, Shaw, Blalock, Cushman, and Nelson demonstrated that criterion accuracy of three-position scores of USAF-MGQT examinations was significantly greater than chance, with no significant differences in test sensitivity to deception compared to seven-position and ESS scoring models. Honts and Reavy (2015) confirmed compatibility across both probable-lie and directed-lie CQT variants [18]Verified APA Standards of Practice (Adopted August 2022)
Confirms APA requirement for evidence-based validated testing techniques and standardized test data analysis methods.
The Polygraph Examiner Resource Guide provides structured frameworks for technique selection across Zone Comparison, screening, and other test formats.
Advantages, Limitations, and Future Directions
Advantages of the 3-Position Scale
The 3-position scale offers several compelling advantages. It significantly improves inter-rater reliability by eliminating the primary source of scoring disagreement — magnitude judgment [8]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners
Confirms no statistically significant differences between ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy with experienced examiners. It achieves accuracy comparable to the 7-position scale across validated techniques [11]Verified The Comparison Question Polygraph Test: A Contrast of Methods and Scoring
Confirms substantial main effects of guilt in both OSS2 computer and human scoring with no differences between probable-lie and directed-lie CQT variants [12]Verified The Reliability of Polygraph Examiner Diagnosis of Truth and Deception
Foundational reliability study finding high inter-rater reliability for experienced examiners using structured scoring methods [16]Verified Brute Force Comparison: A Monte Carlo Study of the OSS-3 and Human Polygraph Scorers
Confirms OSS-3 accuracy exceeded average human scorer accuracy and reports Fleiss kappa of.59 (95% CI =.51 to.66) for inter-rater reliability of human scorers. It enables faster, more efficient chart analysis while reducing examiner cognitive load. And it is more readily automated: unlike the 7-position model, 3-position scale transformations are easily automated because they are more readily made objective and statistically optimized.
The 3-position scale also supports newer analytical approaches. Three-position scores can be characterized using a multinomial distribution, enabling the application of Bayesian statistical methods to polygraph scoring. This represents a significant advance in the scientific rigor of polygraph test data analysis.
Considerations and Limitations
Some research suggests that the component weighting achieved by seven-position and ESS models may be more effective than the unweighted three-position model at extracting diagnostic information from raw test data. This is one reason the ESS applies electrodermal weighting on top of 3-position scoring rather than using unweighted 3-position scores alone.
Additionally, cutscores must be specifically calibrated for 3-position score distributions. Using 7-position cutscores with 3-position totals will produce unreliable results [8]Verified Criterion Validity of the Empirical Scoring System with Experienced Examiners
Confirms no statistically significant differences between ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy with experienced examiners. Examiners transitioning between systems must ensure they apply the appropriate normative reference data. Understanding testing errors — procedural, random, and systematic — is essential for all scoring approaches.
Future Directions
The evolution from the 3-position scale to the ESS and ESS-M represents an ongoing trajectory toward more evidence-based, statistically grounded polygraph scoring. The integration of Bayesian analysis through the ESS-M allows examiners to report probabilistic results alongside categorical decisions — a direction supported by the APA standard recommending that "for evidentiary examinations, an examiner shall report the probabilistic results for the technique and/or the rendered opinion".
Continued research into automated scoring algorithms, including the Objective Scoring System version 3 (OSS-3), which demonstrated accuracy exceeding the average of 10 human scorers on 6 dimensions of accuracy, points toward hybrid approaches that combine the strengths of human judgment with computational precision. For examiner technology options, explore our review of Stoelting CPS Pro software or learn about EyeDetect operators vs. polygraph examiners.
Identify the Spot
Locate the relevant question and its paired comparison question within a given chart. For example, Relevant Question 3 compared to Comparison Question 2.
Select the Physiological Channel
Focus on one channel at a time (e.g., EDA tracing). Each of the standard channels — thoracic respiration, abdominal respiration, EDA, and cardiovascular — is scored independently.
Compare Responses Visually
Apply the bigger-is-better rule: visually compare the response following the relevant question to the response following the comparison question. Evaluate amplitude, duration, and complexity features.
Assign the Score (-1, 0, or +1)
If the relevant response is visibly stronger, assign -1. If the comparison response is visibly stronger, assign +1. If the responses are indistinguishable or equivalent, assign 0. Do not assess magnitude.
Repeat Across All Channels and Charts
Score every channel at every spot in every chart. With 4 channels, 2 spots per chart, and 3 charts, this produces 24 individual scores per examination.
Sum Scores and Apply Decision Rules
Calculate the grand total by summing all individual scores. Compare against validated cutscores using the appropriate decision rules (e.g., ESS two-stage rules) to determine DI, NDI, or INC.
Pros
- Higher inter-rater reliability — eliminates magnitude disagreements, the primary source of scoring variance between examiners
- Comparable diagnostic accuracy to the 7-position scale across validated comparison question techniques
- Faster chart analysis with reduced examiner cognitive load per scoring decision
- Easier automation — 3-position transformations are readily made objective and statistically optimized
- Shorter training curve — inexperienced examiners achieve accuracy comparable to experienced scorers
- Foundation for advanced statistical methods (ESS-M multinomial distributions, Bayesian analysis)
- Broad compatibility across Zone Comparison, MGQT, and other CQT formats
- Supported by 35+ years of peer-reviewed research from multiple independent research groups
Cons
- Unweighted 3-position scores may extract slightly less diagnostic information than weighted models like the ESS or 7-position scale
- Requires specifically calibrated cutscores — using 7-position cutscores with 3-position totals produces unreliable results
- Does not capture the degree of response differences, which may occasionally carry diagnostic value in extreme cases
- Some agencies and programs still mandate 7-position scoring, requiring examiners to maintain competency in both systems
Frequently Asked Questions
What is the 3-position scale in polygraph scoring?
The 3-position scale is a simplified numerical scoring system that uses only three values — -1, 0, and +1 — to evaluate physiological responses during a polygraph examination. It is an abbreviated form of the 7-position scale where the examiner determines only the direction of the response difference between relevant and comparison questions, not the magnitude. A -1 indicates the relevant response was stronger (associated with deception), +1 indicates the comparison response was stronger (associated with truthfulness), and 0 indicates no discernible difference.
Is the 3-position scale as accurate as the 7-position scale?
Yes. Research by Krapohl (1998) with laboratory data and Harwell (2000) with field examinations found that the 3-position scale achieves criterion accuracy comparable to the 7-position scale. Nelson and Krapohl (2011) found no statistically significant differences between the ESS (which uses 3-position scoring) and the seven-position evidentiary model across 13 dimensions of criterion accuracy. The key insight is that most diagnostic information comes from the direction of the response difference, not its magnitude.
Why does the 3-position scale improve inter-rater reliability?
The 3-position scale improves inter-rater reliability by eliminating magnitude judgment — the primary source of scoring disagreement between examiners. When two examiners evaluate the same chart using the 7-position scale, they typically agree on which question produced the stronger response but disagree on how much stronger (e.g., -1 vs. -2 vs. -3). By removing the magnitude dimension, the 3-position scale removes this source of variance. Studies show decision agreement rates of 85-95% among examiners using simplified scoring systems.
What is the relationship between the 3-position scale and the Empirical Scoring System (ESS)?
The ESS, first described in 2008 by Nelson, Krapohl, and Handler, is built on 3-position scoring principles. It uses a three-position scale (+, 0, -) with the bigger-is-better rule, but adds electrodermal weighting and optimized normative cutscores. The ESS is a modification of the Federal 3-position scoring method, incorporating work from the University of Utah and Johns Hopkins University. The ESS-M (Multinomial) update, available since 2017, added Bayesian analytical capabilities.
Does the American Polygraph Association (APA) recognize the 3-position scale?
Yes. The APA's Terminology Reference formally defines the 3-position scale as an abbreviated form of the 7-position scale for PDD test data analysis, citing foundational research by Capps & Ansley (1992), Harwell (2000), Krapohl (1998), and Van Herk (1990). The APA Standards of Practice require evidence-based validated testing techniques, and the 3-position scale, 7-position scale, and ESS are all recognized validated scoring methods used in professional polygraph practice.
How many physiological channels are scored using the 3-position scale?
Typically four channels are scored independently: thoracic (upper) respiration, abdominal (lower) respiration, electrodermal activity (EDA/GSR), and cardiovascular activity. Some techniques also include a fifth channel — peripheral vasomotor activity measured by a finger plethysmograph. Each channel is scored separately using the 3-position scale, and the scores are then summed across all channels and charts.
What cut scores are used with 3-position scoring?
Cut scores vary by technique and decision rule. For the ESS using two-stage decision rules with three-question ZCT examinations: a grand total of +2 or higher indicates NDI (alpha =.10), while a grand total of -4 or lower indicates DI (alpha =.05). If the grand total is inconclusive, Stage 2 evaluates subtotal scores. It is critical to use cutscores specifically calibrated for 3-position score distributions rather than applying 7-position cutscores.
Can new examiners effectively use the 3-position scale?
Yes, and this is one of its primary advantages. Multiple studies demonstrate that inexperienced examiners achieve accuracy comparable to experienced scorers when using simplified 3-position scoring systems like the ESS. A cross-cultural replication study found that 19 international trainees achieved 90.1% mean accuracy using the ESS. Blalock (2011) confirmed similar findings with a kappa inter-scorer agreement of.61 among expert examiners, indicating substantial reliability.
Sources & References
Confirms the formal APA definition of the 3-position scale, its relationship to the 7-position scale, and the foundational research citations (Capps & Ansley 1992, Harwell 2000, Krapohl 1998, Van Herk 1990)
Confirms examiners achieved 82% overall accuracy (361/440 decisions) in blind chart review, with respiratory suppression as the most frequently cited criterion
Confirms Van Herk's 1990 investigation into simplified scoring alternatives, originally published in the Canadian Association of Police Polygraphists newsletter and reprinted in Polygraph 20(2), 70-79
Confirms Krapohl's 1998 comparison of scoring scales published in Polygraph 27(4), 210-218, providing key evidence for 3-position scale efficacy
Confirms the 3-position scale demonstrated criterion accuracy significantly greater than chance with no significant differences in test sensitivity to deception compared to the 7-position model in field examinations
Confirms identification of 22 distinct physiological reaction criteria (12 respiration, 3 electrodermal, 7 cardiovascular) and hierarchical scoring rules
Confirms detailed descriptions of 3-position, 7-position, and ESS scoring methodologies, including that 7-position scores are Likert-type subjective values with no theoretical basis for boundary locations
Confirms no statistically significant differences between ESS and seven-position evidentiary model across 13 dimensions of criterion accuracy with experienced examiners
Confirms PolyScore 3.3 achieved 90.9% accuracy on ZCT examinations and human examiners achieved 82.3% accuracy on MGQT using the 7-position scale
Confirms Fleiss kappa of.61 (95% CI:.54–.68) and mean decision agreement of 95.4% excluding inconclusive results
Confirms substantial main effects of guilt in both OSS2 computer and human scoring with no differences between probable-lie and directed-lie CQT variants
Foundational reliability study finding high inter-rater reliability for experienced examiners using structured scoring methods
Confirms validated evidentiary techniques must demonstrate minimum 90% accuracy with inconclusive rates not exceeding 20%
Official DoD policy manual standardizing testing procedures, scoring methods, and quality assurance requirements for federal polygraph programs
Confirms the Utah method for numerically evaluating polygraph charts using +3 to -3 scoring and describes foundational scoring criteria
Confirms OSS-3 accuracy exceeded average human scorer accuracy and reports Fleiss kappa of.59 (95% CI =.51 to.66) for inter-rater reliability of human scorers
Confirms 19 international trainees achieved 90.1% mean accuracy with 85% decision agreement using ESS, supporting validity of 3-position scoring principles
Confirms APA requirement for evidence-based validated testing techniques and standardized test data analysis methods
Confirms ESS is a modification of the Federal 3-position scoring method and ESS-M has become one of the most widely used TDA methods globally
Confirms computerized systems have potential to reduce bias and eliminate inter-rater variability in polygraph chart scoring
Confirms professional standards listing 3-Position, 7-Position, ESS, and OSS as validated scoring methods examiners should learn
Confirms comprehensive terminology and definitions for both 3-position and 7-position scoring systems in modern polygraph practice
Behind every score is a trained examiner, and to experience the process yourself, find a lie detector test near you and see current pricing.