Professional Examiners Trained to APA Standards
140+ Professional Testing Locations Across the U.S. & Canada
Trusted by 10,000+ Clients, Attorneys & Organizations
LieDetectorTest.com Private & Confidential Polygraph Provider

Spot Analysis in Polygraph: Chart Marking & Scoring

Learn how polygraph examiners use spot analysis to identify, mark, and numerically score physiological reactions on chart data to reach accurate diagnostic conclusions.

Published March 27, 2026 Updated July 24, 2026 37 min read All articles

Spot analysis is how examiners mark and score reactions within specific chart segments; understanding this technique reveals the care behind a lie detector test scored by an experienced professional through LieDetectorTest.com.

Spot analysis is the disciplined process that bridges raw physiological data and a diagnostic opinion. This comprehensive guide explains how polygraph examiners pinpoint reaction sites on chart tracings, assign numerical values using validated scoring scales, and render truthful or deceptive determinations within APA-compliant standards.

92%Blind Review Accuracy
22Scoring Criteria (DoDPI)
15–30sResponse Window
3–5Chart Runs Required

TL;DR — The Short Version

  • Spot analysis is the examiner's systematic process of selecting specific data locations on a polygraph chart where physiological responses to test questions are evaluated and numerically scored.
  • Each 'spot' is a defined time window on one physiological channel (breathing, EDA, cardiovascular) corresponding to a single test question, typically lasting 15 to 30 seconds.
  • Examiners assign numerical values from -3 (strong deceptive response) to +3 (strong truthful indicator) at each identified spot, then sum scores across all charts to reach a diagnostic conclusion.
  • Spots are always evaluated by comparing the response to a relevant question against its paired comparison question — never in isolation.
  • The Swinford (1999) study identified 22 distinct physiological reaction criteria used at DoDPI for manual scoring across respiration, EDA, and cardiovascular channels.
  • Automated algorithms such as OSS-3, PolyScore, and the Empirical Scoring System (ESS-M) supplement manual spot analysis to reduce variability and improve consistency.

Who This Guide Is For

  • Polygraph examinees wanting to understand how their test data will be evaluated
  • Trainee and newly professional polygraph examiners learning chart evaluation skills
  • Defense attorneys who need to challenge or understand polygraph scoring methodology
  • Therapists and treatment providers working with PCSOT clients who receive polygraph results
  • Human resources professionals overseeing pre-employment polygraph programs
  • Students of forensic psychophysiology and lie detection science

What Is Spot Analysis in Polygraph Testing?

Defining the Spot

Spot analysis is the core evaluative technique polygraph examiners use to translate raw physiological data into a diagnostic opinion. When an examiner collects chart data during a polygraph examination, the resulting tracings — representing respiration, electrodermal activity (EDA), and cardiovascular function — are continuous streams of information extending for the full duration of every chart run. Spot analysis is the disciplined process of isolating the precise moments within those tracings that correspond to specific test questions, then scoring those moments to determine whether physiological reactions are more consistent with truth-telling or deception.

During the test, the examiner asks a structured series of questions at regular intervals — typically every 20 to 35 seconds. Each question generates a brief window of physiological data that the examiner later examines. That window is the "spot." It begins roughly at the moment the question is fully asked and extends for a defined period (usually 15 to 30 seconds) during which the autonomic nervous system has time to produce a measurable response.

The NCCA Numerical Evaluation Scoring System defines spot analysis as "a fundamental concept for assigning weighted values, by individual recording channel, based upon comparisons between a relevant question and the applicable comparison question(s)" [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. The word "spot" itself is used across polygraph practitioner circles as shorthand for the specific chart location tied to a specific question. Different training schools may use terms like "reaction site," "response window," or "data evaluation point," but they all describe the same concept: a discrete location on the chart where scoring occurs.

Capps and Ansley (1992) conducted landmark research analysing federal polygraph charts by both spot and chart total methods, demonstrating how these discrete evaluation points drive the overall diagnostic outcome [3]Verified Analysis of federal polygraph charts by spot and chart total
Analysed federal polygraph charts by both spot and chart total methods to determine diagnostic accuracy patterns
.

Why Spot Analysis Matters

Understanding spot analysis is critical because it represents the point at which structured interpretation enters an otherwise mechanical data-collection process. The polygraph instrument itself records data passively. It is the examiner's spot analysis that transforms that data into a conclusion — making it the single most important skill a polygraph professional possesses. For a broader look at the purpose and use of lie detection in the USA, our essential knowledge guide offers further context.

Research confirms that how examiners conduct spot analysis fundamentally determines test accuracy. In a key study, Capps and Ansley (1992) found that examiners achieved 82% overall accuracy in blind chart reviews, with significantly higher accuracy for deceptive cases (89.3%) compared to truthful cases (72.2%), revealing important asymmetries in examiner decision-making patterns [10]Verified Numerical scoring of polygraph charts: What examiners really do
Examiners achieved 82% overall accuracy (361/440 decisions) in blind chart review with respiratory suppression as most cited criterion
. Respiratory suppression was the most frequently cited criterion in their analysis, underscoring the importance of careful spot identification on respiratory channels [10]Verified Numerical scoring of polygraph charts: What examiners really do
Examiners achieved 82% overall accuracy (361/440 decisions) in blind chart review with respiratory suppression as most cited criterion
.

A government compendium of research found that across 11 studies involving independent analyses of 1,609 sets of charts from confirmed field examinations, blind reviewers achieved an average accuracy of 92% [12]Verified Validity and Reliability of Polygraph Testing
Confirmed 92% average accuracy across 11 reliability studies involving 1,609 sets of charts from confirmed field examinations
. This demonstrates that when trained examiners follow consistent spot analysis protocols, the methodology produces highly reliable results. Research also indicates that experienced examiners make an average of 91.4% correct judgments, compared to 77.5% for inexperienced examiners, based on the Horvath and Reid (1971) study [13]Verified Scientific Validity of Polygraph Testing: A Research Review and Evaluation
Experienced examiners achieved 91.4% correct judgments versus 77.5% for inexperienced examiners (Horvath and Reid study)
. This nearly 14-percentage-point gap highlights the extraordinary value of thorough training in spot analysis techniques.

Why Spot Analysis Matters for Polygraph Accuracy

Key Accuracy Factors

The accuracy of any polygraph examination depends not only on the quality of the question format, the calibration of the instrument, and the cooperation of the examinee — it depends fundamentally on how well the examiner conducts spot analysis. Poor spot selection or inconsistent scoring can turn an otherwise well-run exam into an unreliable one. For guidance on how questions themselves are constructed, see Raymond Nelson's guide to polygraph question formulation.

Several factors make spot analysis the linchpin of polygraph accuracy:

Standardization: When examiners follow consistent rules for identifying where to score, inter-rater reliability improves. The Swinford (1999) study at the Department of Defense Polygraph Institute documented 22 distinct physiological reaction criteria — 12 for respiration, 3 for electrodermal activity, and 7 for cardiovascular — and established hierarchical scoring rules for assigning values on the seven-position scale [2]Verified Manually Scoring Polygraph Charts Utilizing the Seven-Position Numerical Analysis Scale at the Department of Defense Polygraph Institute
Documented 22 distinct physiological reaction criteria (12 respiration, 3 EDA, 7 cardiovascular) and established hierarchical scoring rules at DoDPI
. These standardized criteria help ensure that multiple examiners evaluating the same chart arrive at the same conclusion.

Transparency: A properly documented spot analysis creates a clear record of how the examiner reached a conclusion. This is essential for quality assurance reviews, peer review, and any legal challenges to the results. Learn more about documentation best practices in our guide to writing a polygraph examination report.

Error Reduction: Structured spot analysis minimizes the influence of confirmation bias. Research by Krapohl and Dutton (2018) demonstrated that polygraph scores and decisions were significantly shifted in the direction of biasing information — examiners told cases were deceptive scored more deceptively than those told cases were truthful [8]Verified Reducing Bias in Polygraph Scoring: The Effect of Biasing Information on Blind Evaluation of Polygraph Charts
Polygraph scores were significantly shifted by biasing information, demonstrating the need for blind scoring protocols
. By forcing the examiner to evaluate specific data windows using defined criteria, disciplined spot analysis constrains subjective judgment.

Comparison Integrity: In comparison-question test formats, spot analysis requires the examiner to directly compare the physiological response at a relevant question spot against the response at a paired comparison question spot. Without disciplined spot selection, this comparison becomes meaningless.

Training and Validated Techniques

The 2011 APA meta-analysis established new mandatory standards requiring only scientifically validated techniques, effectively ending the era when tradition or personal preference alone could justify testing methods [4]Verified Validated Techniques and Scoring Models for PDD Test Data Analysis – Conclusions from the 2011 APA Report
Confirms the 2011 APA meta-analysis established mandatory standards requiring only scientifically validated techniques
. Examiner adherence to validated spot analysis protocols is now a central requirement for professional practice.

The National Center for Credibility Assessment (NCCA) — formerly the Department of Defense Polygraph Institute (DoDPI) — serves as the U.S. government's polygraph education, oversight, and research centre [12]Verified Validity and Reliability of Polygraph Testing
Confirmed 92% average accuracy across 11 reliability studies involving 1,609 sets of charts from confirmed field examinations
. The NCCA system consists of two numerical evaluation scales: the three-position and seven-position scales [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. Although the three-position scale is more conservative and objective, NCCA advocates use of the seven-position scale because it reduces the number of No Opinion decisions [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. All federal polygraph examiners receive their basic training at NCCA, which currently holds classes in January, April, and August, each lasting approximately three months [14]Verified CBP Polygraph Examiner Career Path — NCCA Training Requirements
Confirms federal polygraph examiners must complete NCCA training, with classes held approximately 3 months each
. This centralised training model promotes consistency and reciprocity across federal agencies.

When comparing scoring approaches, the Senter (2000) study found that when veracity decision rules, chart number, and data channels were equated, DoDPI and Utah scorers showed no significant differences in accuracy [7]Verified Comparison of Utah and DoDPI Scoring Accuracy: Equating Veracity Decision Rule, Chart Rule, and Number of Data Channels Used
When decision rules and channels were equated, DoDPI and Utah scorers showed no significant accuracy differences
. This confirms that standardised spot analysis training — regardless of the specific school — produces comparable results when protocols are properly followed. For more on how examiner training impacts accuracy, see our guides to polygraph training in Kansas, Massachusetts, or Vermont.

Anatomy of a Polygraph Chart: Where Spots Live

Standard Physiological Channels

A modern computerised polygraph chart displays multiple physiological channels simultaneously, each running horizontally across the screen as time passes from left to right. Understanding these channels is essential because the examiner will identify and score separate spots on each channel for every test question.

Upper Pneumograph (Thoracic Respiration): A pneumatic tube worn around the upper chest tracks breathing rate, depth, and pattern. This channel is critical for spot analysis because respiratory changes such as suppression, apnea, or accelerated breathing are primary indicators of autonomic arousal.

Lower Pneumograph (Abdominal Respiration): A second respiratory tube worn around the abdomen. Examiners evaluate both channels because thoracic and abdominal breathing can respond independently to stress stimuli. The NCCA scoring system recognises that although these two sensors often mirror each other, some examinees are predominantly thoracic or abdominal breathers, and because these areas are subject to different innervations (intercostals and phrenic), capturing both is essential [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. Respiratory suppression — a reduction in breathing activity following a question — is one of the most frequently observed deception indicators in polygraph data [10]Verified Numerical scoring of polygraph charts: What examiners really do
Examiners achieved 82% overall accuracy (361/440 decisions) in blind chart review with respiratory suppression as most cited criterion
.

Electrodermal Activity (EDA): Measured through galvanic skin response (GSR) finger plates or electrodes attached to the fingertips. This channel records changes in sweat gland activity. EDA consistently emerges as the most diagnostically powerful channel in research. The Empirical Scoring System (ESS) and its multinomial update (ESS-M) weight EDA scores more heavily than other channels based on empirical evidence [15]Verified Reference Guide for the Empirical Scoring System - Multinomial (ESS-M)
Describes ESS first published in 2008 and ESS-M update available 2017; confirms evidence-based feature extraction and weighted EDA scoring approach
. Weighted EDA scores have been shown to produce an average 49% reduction of inconclusive results across all decision rules [15]Verified Reference Guide for the Empirical Scoring System - Multinomial (ESS-M)
Describes ESS first published in 2008 and ESS-M update available 2017; confirms evidence-based feature extraction and weighted EDA scoring approach
.

Cardio Channel: Typically measured via a blood pressure cuff (sphygmomanometer) on the upper arm or through a cardio activity monitor. This channel records pulse rate, blood pressure changes, and pulse amplitude.

Photoplethysmograph (PPG): Some instruments add a finger pulse oximeter that tracks blood volume changes in peripheral vasculature. Lafayette Instrument Company's LXSoftware is bundled with the OSS-3 scoring algorithm and can record up to ten channels of data [16]Verified OSS-3: Objective Scoring System Version 3
Describes OSS-3 as a probabilistic classifier for diagnostic and screening polygraphs with artifact management features
.

Question Markers and Spot Location

Each test question is time-stamped on the chart. When the examiner asks a relevant question (e.g., R1), a marker appears at that exact moment. The spot for R1 is the data window that begins at that marker and extends forward for the defined evaluation period. The examiner will have separate spots for R1 on every single channel — one on the upper pneumo, one on the lower pneumo, one on the EDA, and one on the cardio. Each channel spot is scored independently.

Multivariate analysis methods can also be applied to polygraph chart data. Minakova and Bozhich (2018) demonstrated that both Principal Component Analysis (PCA) and Partial Least Squares (PLS) methods successfully analysed polygraph data to assess physiological response anomalies, illustrating how modern analytical approaches can supplement traditional spot evaluation [9]Verified Multivariate data analysis of polygraph charts for assessment of biophysical characteristics
PCA and PLS methods successfully analysed polygraph data to assess physiological response anomalies
.

How Examiners Identify and Mark Spots

Defining the Response Window

The response window is the time period after a question during which the examiner expects to see any meaningful physiological change. Most validated techniques define this window as beginning approximately 1 to 3 seconds after the question is fully delivered and extending for 15 to 30 seconds. The exact window depends on the channel:

Respiration: Respiratory responses are typically slower to manifest. Examiners often evaluate one to three full breathing cycles following the question stimulus. A typical evaluation window is 15 to 25 seconds after the question. The NCCA scoring system identifies five respiratory features evaluated using respiratory line length (RLL) methodology: apnea/blocking, decrease in amplitude, progressive decrease in amplitude, decrease in cyclic rate, and inhalation/exhalation ratio changes [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. In 2001, NCCA adopted the RLL concept, which reduces reliance on signature and pattern recognition and provides an objective means of evaluation [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
.

Electrodermal Activity: EDA responses are faster, often appearing within 1 to 5 seconds of the question. However, the full response arc — from onset to peak to recovery — may extend 10 to 20 seconds. The scoring spot captures this entire arc. ESS scoring features focus solely on primary features (referred to as Kircher features) — specifically electrodermal amplitude of increase in phasic activity — because they show stronger links to deception than secondary features [15]Verified Reference Guide for the Empirical Scoring System - Multinomial (ESS-M)
Describes ESS first published in 2008 and ESS-M update available 2017; confirms evidence-based feature extraction and weighted EDA scoring approach
.

Cardiovascular: Blood pressure and pulse rate changes can take 5 to 15 seconds to become apparent and may persist for 20 or more seconds. The Swinford (1999) study documented 7 cardiovascular scoring criteria used at DoDPI [2]Verified Manually Scoring Polygraph Charts Utilizing the Seven-Position Numerical Analysis Scale at the Department of Defense Polygraph Institute
Documented 22 distinct physiological reaction criteria (12 respiration, 3 EDA, 7 cardiovascular) and established hierarchical scoring rules at DoDPI
. For deeper detail on how the cardio channel functions, see our guide to calming techniques and deep breathing during polygraph examinations.

Marking the Spot on the Chart

In computerised polygraph systems — such as the Lafayette LX6 and LX7, Stoelting CPS Elite, or Limestone ParagonX — question markers are placed automatically or semi-automatically during data collection. The examiner typically clicks a marker button at the instant they begin asking each question. This marker anchors the spot. For information about the physical handling of polygraph instruments, see our equipment shipping and handling protocols guide.

Some examiners refine spot boundaries during post-test analysis. If an examinee coughs or moves during a particular question's window, the examiner may annotate that section as an artifact and exclude it from scoring. The OSS-3 algorithm provides the ability to mark artifacted segments that should not be included in the statistical analysis [16]Verified OSS-3: Objective Scoring System Version 3
Describes OSS-3 as a probabilistic classifier for diagnostic and screening polygraphs with artifact management features
. The marking process requires precision — misplacing a marker by even a few seconds can shift the evaluation window and change the score.

Pairing Relevant and Comparison Questions

In comparison-question test formats (such as the Zone Comparison Test, Federal ZCT, or Utah CQT), each relevant question is paired with a specific comparison question. Spot analysis never evaluates a relevant question spot in isolation — it always evaluates it relative to its paired comparison question spot on the same channel, within the same chart run.

For example, if Relevant Question 1 (R1) is paired with Comparison Question 1 (C1), the examiner compares the EDA response at the R1 spot against the EDA response at the C1 spot. If R1 shows a larger EDA response than C1, the examiner assigns a negative score. If C1 shows the larger response, a positive score is assigned. If they appear equal, a score of zero is recorded [2]Verified Manually Scoring Polygraph Charts Utilizing the Seven-Position Numerical Analysis Scale at the Department of Defense Polygraph Institute
Documented 22 distinct physiological reaction criteria (12 respiration, 3 EDA, 7 cardiovascular) and established hierarchical scoring rules at DoDPI
. This pairing principle is foundational to the Control Question Technique and all its validated variants. To understand why the focus on a single issue often yields the strongest results, see our guide to the single-issue polygraph test.

The Numerical Scoring System for Spots

The Seven-Position Scale (-3 to +3)

Once spots have been identified and marked, the examiner assigns a numerical value to each comparison pair on each channel. The most widely used scoring scale is the 7-position scale, which assigns values as follows:

+3 — Strong Truthful Indicator: The comparison question clearly produced a dramatically larger physiological response than the relevant question on this channel.

+2 — Moderate Truthful Indicator: The comparison question produced a noticeably larger response. The difference is clear but not extreme.

+1 — Slight Truthful Indicator: The comparison question produced a slightly larger response. The difference is present but subtle.

0 — No Significant Difference: No meaningful difference exists between the relevant and comparison question responses on this channel.

-1 — Slight Deceptive Indicator: The relevant question produced a slightly larger response than the comparison question.

-2 — Moderate Deceptive Indicator: The relevant question produced a noticeably larger response. The deceptive indicator is clear and consistent.

-3 — Strong Deceptive Indicator: The relevant question produced a dramatically larger response than the comparison question.

The NCCA advocates use of the seven-position scale because it has a tendency to reduce the number of No Opinion decisions by allowing an examiner to assign greater values to each recording channel [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. For EDA specifically, the NCCA uses a ratio method: if one response is visually larger than the comparable response but the ratio is less than 3:1, the examiner defers to the "bigger-is-better" principle and assigns a value of plus or minus one [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
.

How Scores Produce a Diagnostic Opinion

After scoring every comparison pair on every channel across all charts, the examiner sums the scores. The grand total is compared against predetermined cutoff thresholds established by the technique being used. A positive total above the cutoff indicates No Deception Indicated (NDI) — a non-deceptive response pattern. A negative total below the cutoff indicates Deception Indicated (DI). Totals falling between the positive and negative cutoffs result in an Inconclusive (INC) finding.

For example, with the Federal ZCT using the 7-position scale, a grand total score of +6 or greater typically indicates non-deception, while -6 or less indicates deception. Scores between these thresholds are inconclusive. A comprehensive comparison found that the Objective Scoring System and Probability Analysis demonstrated equivalent accuracy with no significant differences in area under the ROC curve [6]Verified Comparison of the Objective Scoring System and Probability Analysis
Both OSS and Probability Analysis demonstrated equivalent accuracy with no significant differences in ROC curve area
.

History of Numerical Scoring

The numerical scoring approach was introduced to the polygraph profession by Cleve Backster in 1959 alongside the Backster Zone Comparison Technique [5]Verified The Numerical Scoring Analysis Predecessors
Confirms numerical scoring was evolutionary, with Backster adapting Arther's check-mark system around 1959–1960
. Backster's 7-position scoring system was an adaptation of the Likert Scale, and was adopted by the U.S. government in the early 1960s [5]Verified The Numerical Scoring Analysis Predecessors
Confirms numerical scoring was evolutionary, with Backster adapting Arther's check-mark system around 1959–1960
. Prior to this, examiners evaluated polygraph charts using the "Clinical Approach" — a subjective, impressionistic method that considered not only physiological responses but also the examinees' verbal and nonverbal behavioural clues observed during the pretest [5]Verified The Numerical Scoring Analysis Predecessors
Confirms numerical scoring was evolutionary, with Backster adapting Arther's check-mark system around 1959–1960
. Backster's numerical analysis transformed the practice by turning test data analysis from subjective to objective evaluation [5]Verified The Numerical Scoring Analysis Predecessors
Confirms numerical scoring was evolutionary, with Backster adapting Arther's check-mark system around 1959–1960
.

However, as Amsel (2023) demonstrates, numerical scoring was more evolutionary than revolutionary. Earlier practitioners had developed proto-scoring systems — for example, John E. Winter practiced a scoring method for breathing curves as early as 1936 [5]Verified The Numerical Scoring Analysis Predecessors
Confirms numerical scoring was evolutionary, with Backster adapting Arther's check-mark system around 1959–1960
. Richard O. Arther developed a check-mark system that Backster likely adapted by replacing check-marks with numerical figures, which led to the birth of the "Numerical Analysis" [5]Verified The Numerical Scoring Analysis Predecessors
Confirms numerical scoring was evolutionary, with Backster adapting Arther's check-mark system around 1959–1960
. These historical insights remind us that the profession's standardisation was built on decades of incremental progress.

7-Position vs. 3-Position Scoring Scales

Comparing the Two Approaches

The NCCA system consists of two numerical evaluation scales: the three-position scale (-1, 0, +1) and the seven-position scale (-3 to +3) [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. The three-position scale is more conservative and objective because there are fewer scoring rules and a narrower range of values, making it easier to learn [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. NCCA introduces trainees to the three-position scale first, then advances them to the seven-position scale once proficiency is achieved [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
.

Three-position scores are unweighted integer scores for all polygraph recording sensors. These scores simply indicate whether a greater change in physiology is observed in response to the relevant question or comparison question within an analysis spot. When differences are based on objective data using primary response features (Kircher features), these are called multinomial three-position scores.

The seven-position scale attempts to capture more information by representing the degree of difference in addition to indicating which question produced the greater response. However, seven-position scores involve pattern recognition methods that are less easily quantified, introducing greater potential for inter-scorer variability.

Research by Robertson (2012) found that the component weighting achieved by the seven-position and ESS models was more effective than the three-position model at extracting diagnostic information from the raw test data. However, no significant differences were observed in unweighted accuracy for the three scoring methods across different numbers of relevant questions.

Automated vs. Manual Spot Analysis

Computer Scoring Algorithms

Modern polygraph analysis leverages standardized computational models designed to objectively evaluate physiological data and provide reproducible results. Several validated automated scoring systems now supplement or verify manual spot analysis:

OSS-3 (Objective Scoring System Version 3): Developed by Raymond Nelson, Mark Handler, and Donald Krapohl, the OSS-3 is a powerful computer algorithm that calculates a probabilistic classifier for both diagnostic and screening polygraphs [16]Verified OSS-3: Objective Scoring System Version 3
Describes OSS-3 as a probabilistic classifier for diagnostic and screening polygraphs with artifact management features
. It outputs a probability score indicating the likelihood of deception, calibrated to empirical accuracy thresholds. The OSS-3 demonstrated balanced sensitivity and specificity and provided significant improvements over previous versions in the form of reduced inconclusive results and increased sensitivity to deception [6]Verified Comparison of the Objective Scoring System and Probability Analysis
Both OSS and Probability Analysis demonstrated equivalent accuracy with no significant differences in ROC curve area
. A key advantage is that the computer algorithm provides perfect reliability — it will always produce the same result from the same data.

PolyScore: Developed at the Johns Hopkins University Applied Physics Laboratory based on criminal case data provided by the Department of Defense Polygraph Institute, PolyScore uses linear discriminant analysis (LDA) and Bayesian probability. A comparison study found that both the Objective Scoring System and Probability Analysis demonstrated equivalent accuracy [6]Verified Comparison of the Objective Scoring System and Probability Analysis
Both OSS and Probability Analysis demonstrated equivalent accuracy with no significant differences in ROC curve area
.

ESS-M (Empirical Scoring System — Multinomial): The ESS was first described in 2008, and the multinomial update (ESS-M) was made first available in 2017 [15]Verified Reference Guide for the Empirical Scoring System - Multinomial (ESS-M)
Describes ESS first published in 2008 and ESS-M update available 2017; confirms evidence-based feature extraction and weighted EDA scoring approach
. The ESS-M is an evidence-based, standardized protocol for analyzing PDD test data that converts changes in physiological activity into numerical scores analysed using a multinomial reference model and a Bayesian classifier [15]Verified Reference Guide for the Empirical Scoring System - Multinomial (ESS-M)
Describes ESS first published in 2008 and ESS-M update available 2017; confirms evidence-based feature extraction and weighted EDA scoring approach
. It has distinguished itself for its criterion accuracy, ease of use, objectivity, and inter-rater reliability. Nelson, Krapohl, and Handler (2008) demonstrated that even inexperienced examiners could produce blind-scoring data with decision accuracy equivalent to that of experienced examiners when using the ESS.

For insights into how AI and technology may further enhance polygraph scoring, see our technology guide.

Human Judgment Still Essential

Despite technological advances, automated algorithms do not replace the human examiner. OSS-3 and similar tools provide a statistical second opinion, but the examiner retains responsibility for identifying artifacts, assessing data quality, and rendering the final professional opinion. The OSS-3 study found that Fleiss' kappa for 10 human scorers was k =.59, consistent with previously reported reliability estimates, compared to the expected perfect reliability of the automated algorithm. However, decision accuracy for trainees using simplified ESS scoring instructions was not statistically different from experienced examiners [16]Verified OSS-3: Objective Scoring System Version 3
Describes OSS-3 as a probabilistic classifier for diagnostic and screening polygraphs with artifact management features
.

Examiners must be able to articulate what decision rules were applied to the analysis and which diagnostic features drove the value assigned to each analysis spot [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. The combination of human expertise and algorithmic precision represents the state of the art in polygraph test data analysis. To understand the ethical obligations governing this process, see our guide on PCSOT polygraph ethics and APA standards.

Common Spot Analysis Errors and How to Avoid Them

Scoring Pitfalls

Even experienced examiners can make errors during spot analysis. Understanding common pitfalls helps examiners and examinees appreciate the importance of quality control:

Misplaced Markers: If a question marker is placed a few seconds early or late, the entire evaluation window shifts, potentially capturing recovery from a previous question or missing the onset of the current response. Computerised systems with automatic timestamping minimise this risk.

Confirmation Bias: The Krapohl and Dutton (2018) study demonstrated that when examiners were given biasing case information before scoring, their scores shifted significantly in the direction of the biasing information [8]Verified Reducing Bias in Polygraph Scoring: The Effect of Biasing Information on Blind Evaluation of Polygraph Charts
Polygraph scores were significantly shifted by biasing information, demonstrating the need for blind scoring protocols
. Blind chart review — where the scorer has no access to case details — is the gold standard for reducing this effect.

Artifact Misidentification: Examiners must distinguish between genuine physiological responses and artifacts caused by movement, coughing, sneezing, or sensor slippage. Failing to mark an artifact can inflate or deflate scores at that spot. The NCCA defines an artifact as "the cause for a change in the examinee's physiological data that is not attributable to an applied stimulus or recovery" [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
.

Inconsistent Application of Scoring Criteria: Using seven-position scoring criteria inconsistently across channels or chart runs introduces noise into the final score total. The 22 physiological reaction criteria documented by Swinford (1999) provide a systematic framework that minimises this problem [2]Verified Manually Scoring Polygraph Charts Utilizing the Seven-Position Numerical Analysis Scale at the Department of Defense Polygraph Institute
Documented 22 distinct physiological reaction criteria (12 respiration, 3 EDA, 7 cardiovascular) and established hierarchical scoring rules at DoDPI
.

Ensuring you can trust your polygraph examiner starts with understanding these quality measures.

Artifacts That Complicate Spot Analysis

Types of Artifacts and Management

Artifacts are non-stimulus physiological changes that can contaminate the data at a spot. Common artifacts include:

Movement Artifacts: Body shifts, arm movements, or leg repositioning create disturbances across multiple channels simultaneously. These are usually obvious on the chart and are excluded from scoring.

Respiratory Artifacts: Deep sighs, yawns, coughs, sneezes, or deliberate breathing manipulation alter the pneumograph tracings. The NCCA scoring guide instructs examiners to distinguish between respiratory patterns attributable to questions and those caused by external events.

Electrodermal Artifacts: Finger movement on the electrode plates, excessive perspiration, or electronic interference can produce EDA spikes unrelated to question stimuli. The OSS-3 algorithm includes the ability to mark these artifacted segments for exclusion [16]Verified OSS-3: Objective Scoring System Version 3
Describes OSS-3 as a probabilistic classifier for diagnostic and screening polygraphs with artifact management features
. It also calculates a supplemental test of proportions to discriminate between countermeasures and random artifacts.

Cardiovascular Artifacts: Premature ventricular contractions (PVCs), sensor slippage, or cuff pressure fluctuations can distort the cardio channel. These require careful examiner judgment to identify and annotate.

Managing artifacts properly is essential because allowing the algorithm to score artifacted responses produces incorrect results. Artifacted segments should always be marked and excluded before allowing an algorithm to calculate a result [16]Verified OSS-3: Objective Scoring System Version 3
Describes OSS-3 as a probabilistic classifier for diagnostic and screening polygraphs with artifact management features
. Understanding how substances like alcohol affect polygraph results can help contextualise some physiological anomalies.

APA Standards Governing Spot Analysis

Professional Requirements

The American Polygraph Association (APA) sets the professional standards that govern polygraph practice, including spot analysis. The 2011 APA meta-analysis was a watershed moment — it established mandatory standards requiring that only scientifically validated techniques be used for polygraph examinations [4]Verified Validated Techniques and Scoring Models for PDD Test Data Analysis – Conclusions from the 2011 APA Report
Confirms the 2011 APA meta-analysis established mandatory standards requiring only scientifically validated techniques
. This means that every technique an examiner uses for spot analysis must have at least two published empirical studies, with original and replicated research, demonstrating adequate accuracy rates with inconclusive results not exceeding 20%.

APA standards also require that all phases of the examination be recorded on audio and video, and that examination files be maintained for a minimum of three years. For evidentiary examinations, examiners must report probabilistic results that support the categorical conclusion of deception or truth-telling. These standards reinforce the transparency and accountability that spot analysis documentation provides.

Training requirements reflect the importance of spot analysis skills. APA-accredited polygraph training programs require a minimum of 400 hours of basic instruction. The NC 14B NCAC 16.0502 polygraph trainee permit guide provides one example of state-level requirements that complement APA standards. Additional state-specific guides are available for Rhode Island and North Dakota.

SPOT Testing: The Searching Peak of Tension Test

What Is SPOT and How Does It Relate to Spot Analysis?

The acronym SPOT also appears in polygraph practice as the Searching Peak of Tension test — a distinct technique that should not be confused with general spot analysis methodology. A SPOT test is a Peak of Tension test in which the examiner does not know the critical item, and it is used to determine information concealed by a guilty examinee. Applications include determining the location of stolen goods or the bodies of murder victims, the amount of cash stolen, or the name of an accomplice.

In a SPOT test, the examiner presents a series of possible answers (e.g., "Was the amount stolen $1,000? $2,000? $3,000?") in a recognisable order. A guilty examinee is expected to show a pattern of responsiveness that increases as the correct alternative approaches in the question sequence and decreases after it passes. The most probable item is placed in the middle of the list during the first of three presentations, and the order is varied with each subsequent presentation.

Jaworski (2015) published detailed observations from the analysis of SPOT charts, demonstrating practical applications of this technique in locating evidence and eliciting confessions in serious criminal cases [1]Verified Observations from the Analysis of Searching Peak of Tension Test (SPOT) Charts
Demonstrates practical applications of SPOT testing in locating evidence and eliciting confessions in serious criminal cases
. His research analysed confirmed chart patterns from field investigations, providing valuable data on the real-world effectiveness of the SPOT format.

While SPOT tests use a fundamentally different analytical approach than comparison-question test spot analysis (they look for a peak-shaped response pattern rather than comparing relevant to comparison question reactions), the underlying principle of identifying specific reaction points on the chart is shared. Both require meticulous examiner attention to the precise physiological data at defined chart locations.

Documenting and Reporting Spot Analysis Results

Creating a Defensible Record

Proper documentation of spot analysis results is essential for professional accountability. Every spot should be documented on a structured test data sheet that records the numerical score assigned to each channel for each question pair on each chart run. The examiner should note the specific diagnostic features that drove each score assignment.

The NCCA instructs that examiners must be able to articulate what decision rules were applied and which diagnostic features drove the value assigned to each analysis spot [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
. Data collection quality directly affects spot analysis defensibility — the NCCA standard advises that examiners should aim to obtain "courtroom quality charts," regardless of whether the examination results ever see the light of a courtroom [11]Verified NCCA Numerical Evaluation Scoring System
Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features
.

For complete guidance on how examination results are formally documented and presented, see our expert guide on how to write a polygraph examination report. The examination report synthesises the numerical spot analysis data into a clear professional narrative. Understanding the potential implications of examination reports helps contextualise the broader landscape explored in our guide to employee rights and lie detector tests.

1

Collect Chart Data

The examiner administers the polygraph test, asking structured questions at regular intervals (every 20–35 seconds) while the instrument records respiration, EDA, and cardiovascular data continuously across 3–5 chart runs.

2

Place Question Markers

During or immediately after data collection, the examiner ensures each question has a time-stamped marker on the chart. Computerised systems like the Lafayette LX6/LX7 allow semi-automatic marker placement.

3

Define Response Windows

For each question marker, the examiner establishes the evaluation window: typically 15–25 seconds for respiration, 10–20 seconds for EDA, and up to 20+ seconds for cardiovascular responses.

4

Identify and Annotate Artifacts

Before scoring, the examiner reviews each spot for artifacts (movement, coughs, sensor issues) and marks them for exclusion. The OSS-3 algorithm provides artifact management tools for this purpose.

5

Pair Relevant and Comparison Questions

Each relevant question spot is paired with its designated comparison question spot on the same channel, within the same chart run. Scoring compares the two responses within each pair.

6

Assign Numerical Scores

Using the 7-position scale (-3 to +3) or the 3-position scale (-1 to +1), the examiner scores each channel at each spot based on which question in the pair produced the greater physiological response.

7

Sum Scores and Apply Decision Rules

All individual spot scores are summed across channels and charts to produce a grand total. This total is compared against validated cutoff thresholds to determine NDI, DI, or INC.

8

Run Automated Verification

The examiner runs the OSS-3 or ESS-M algorithm as a statistical second opinion to supplement the manual analysis. Any discrepancies between manual and automated results are examined.

9

Document and Report

The examiner records all scores, diagnostic features, artifact annotations, and the final diagnostic opinion on the test data sheet and in the formal examination report.

Pros

  • Provides a systematic, evidence-based framework for translating raw physiological data into diagnostic conclusions
  • Standardized scoring criteria (22 DoDPI criteria) promote inter-rater reliability across examiners and agencies
  • Numerical scoring produces a documented, transparent, and reviewable audit trail for every diagnostic opinion
  • Blind review accuracy reaches 92% across validated field studies when proper spot analysis protocols are followed
  • Automated algorithms (OSS-3, ESS-M, PolyScore) supplement manual analysis with perfect reliability and statistical transparency
  • The ESS model enables even inexperienced examiners to produce decision accuracy equivalent to experienced examiners
  • Multiple validated scoring scales (3-position and 7-position) allow flexibility for different testing contexts and precision requirements

Cons

  • Seven-position scoring involves subjective pattern recognition elements that can introduce inter-scorer variability
  • Examiner confirmation bias can influence scores if blind review protocols are not followed, as demonstrated by Krapohl and Dutton (2018)
  • Artifact identification requires considerable examiner judgment and training
  • Manual spot analysis is time-intensive, particularly when using the OSS method without automation
  • Different training schools may apply slightly different scoring criteria, though Senter (2000) showed this does not significantly affect accuracy when protocols are equated

Frequently Asked Questions

What exactly is a 'spot' on a polygraph chart?

A spot is a defined time window on a single physiological channel (such as respiration, EDA, or cardiovascular) that corresponds to one test question. It typically begins when the question is fully asked and extends for 15 to 30 seconds — the period during which the autonomic nervous system produces a measurable response. The NCCA defines spot analysis as assigning weighted values by individual recording channel based upon comparisons between a relevant question and the applicable comparison question.

How accurate is polygraph spot analysis?

Research demonstrates strong accuracy when proper protocols are followed. A government compendium found that blind reviewers achieved an average accuracy of 92% across 11 studies involving 1,609 sets of confirmed field examination charts. Experienced examiners averaged 91.4% correct judgments in the Horvath and Reid study, compared to 77.5% for inexperienced examiners — highlighting the critical importance of thorough training in spot analysis techniques.

What is the difference between the 3-position and 7-position scoring scales?

The 3-position scale uses values of -1, 0, and +1, indicating only the direction of the difference between relevant and comparison question responses. The 7-position scale (-3 to +3) captures both direction and magnitude. The NCCA teaches both but advocates the 7-position scale because it reduces No Opinion outcomes. The 3-position scale is more conservative, objective, and easier to learn, with fewer scoring rules.

What are OSS-3 and ESS-M in polygraph scoring?

OSS-3 (Objective Scoring System Version 3) is a computer algorithm developed by Raymond Nelson, Mark Handler, and Donald Krapohl that calculates a probabilistic classifier for diagnostic and screening polygraphs. ESS-M (Empirical Scoring System — Multinomial) is an evidence-based protocol that converts physiological data into numerical scores analysed using a Bayesian classifier. Both supplement manual scoring to improve consistency and reduce examiner variability.

Can examiner bias affect spot analysis results?

Yes — research by Krapohl and Dutton (2018) demonstrated that when examiners were given biasing case information before scoring charts, their scores and decisions shifted significantly in the biased direction. This is why blind chart review — where the scorer has no access to case details — is considered the gold standard. Structured spot analysis protocols and automated algorithms help further reduce the influence of subjective judgment.

What is the SPOT test and is it the same as spot analysis?

No — the SPOT (Searching Peak of Tension) test is a distinct polygraph technique used to identify concealed information, such as the location of stolen goods or the amount of money taken. It presents possible answers in sequence and looks for a peak response pattern. General spot analysis, by contrast, is the broader methodology of identifying, marking, and numerically scoring reaction sites on comparison-question test charts. Jaworski (2015) published detailed research on SPOT chart analysis in criminal investigations.

How many physiological criteria do examiners use when scoring spots?

The Swinford (1999) study at the Department of Defense Polygraph Institute documented 22 distinct physiological reaction criteria: 12 for respiration, 3 for electrodermal activity, and 7 for cardiovascular responses. These criteria are organised hierarchically, with scoring rules that determine the value assigned based on the number and quality of criteria present at each comparison pair.

What happens if an artifact occurs during a test question?

If a non-stimulus event (such as a cough, body movement, or sensor slippage) occurs during a question's response window, the examiner annotates that section as an artifact and excludes it from scoring. The OSS-3 algorithm provides the ability to mark artifacted segments for exclusion from statistical analysis. Artifacted responses should never be scored, as doing so would produce incorrect results.

Do different polygraph schools produce different spot analysis results?

Research suggests that when protocols are properly followed, different training approaches produce comparable results. The Senter (2000) study found that when veracity decision rules, chart number, and data channels were equated, DoDPI-trained and Utah-trained scorers showed no significant differences in accuracy. Standardised training through institutions like the NCCA promotes consistency across the profession.

Sources & References

1
Observations from the Analysis of Searching Peak of Tension Test (SPOT) Charts
Ryszard Jaworski (2015) — European Polygraph
Verified

Demonstrates practical applications of SPOT testing in locating evidence and eliciting confessions in serious criminal cases

2

Documented 22 distinct physiological reaction criteria (12 respiration, 3 EDA, 7 cardiovascular) and established hierarchical scoring rules at DoDPI

3
Analysis of federal polygraph charts by spot and chart total
Capps, M. H., Ansley, N. (1992) — Polygraph
Verified

Analysed federal polygraph charts by both spot and chart total methods to determine diagnostic accuracy patterns

4

Confirms the 2011 APA meta-analysis established mandatory standards requiring only scientifically validated techniques

5
The Numerical Scoring Analysis Predecessors
Tuvya T. Amsel (2023) — European Polygraph
Verified

Confirms numerical scoring was evolutionary, with Backster adapting Arther's check-mark system around 1959–1960

6
Comparison of the Objective Scoring System and Probability Analysis
Department of Defense Polygraph Institute (2009) — NCJRS (US Government)
Verified

Both OSS and Probability Analysis demonstrated equivalent accuracy with no significant differences in ROC curve area

7

When decision rules and channels were equated, DoDPI and Utah scorers showed no significant accuracy differences

8
Reducing Bias in Polygraph Scoring: The Effect of Biasing Information on Blind Evaluation of Polygraph Charts
Krapohl, D.J., Dutton, D.W. (2018) — Polygraph & Forensic Credibility Assessment
Verified

Polygraph scores were significantly shifted by biasing information, demonstrating the need for blind scoring protocols

9

PCA and PLS methods successfully analysed polygraph data to assess physiological response anomalies

10
Numerical scoring of polygraph charts: What examiners really do
Capps, M. H., Ansley, N. (1992) — Polygraph
Verified

Examiners achieved 82% overall accuracy (361/440 decisions) in blind chart review with respiratory suppression as most cited criterion

11
NCCA Numerical Evaluation Scoring System
National Center for Credibility Assessment (2017) — NCCA Official Document
Verified

Defines NCCA's three- and seven-position scoring scales, spot analysis methodology, RLL features, and 8 primary diagnostic features

12
Validity and Reliability of Polygraph Testing
Office of Justice Programs (2000) — US Department of Justice
Verified

Confirmed 92% average accuracy across 11 reliability studies involving 1,609 sets of charts from confirmed field examinations

13
Scientific Validity of Polygraph Testing: A Research Review and Evaluation
Office of Technology Assessment (1983) — US Congress OTA
Verified

Experienced examiners achieved 91.4% correct judgments versus 77.5% for inexperienced examiners (Horvath and Reid study)

14
CBP Polygraph Examiner Career Path — NCCA Training Requirements
US Customs and Border Protection (2025) — US Government
Verified

Confirms federal polygraph examiners must complete NCCA training, with classes held approximately 3 months each

15
Reference Guide for the Empirical Scoring System - Multinomial (ESS-M)
Raymond Nelson (2017) — Polygraph & Forensic Credibility Assessment
Verified

Describes ESS first published in 2008 and ESS-M update available 2017; confirms evidence-based feature extraction and weighted EDA scoring approach

16
OSS-3: Objective Scoring System Version 3
Raymond Nelson, Mark Handler, Donald Krapohl (2008) — Polygraph
Verified

Describes OSS-3 as a probabilistic classifier for diagnostic and screening polygraphs with artifact management features

Need to book now? Our online booking system is open 24/7. Speak directly with our team about your test or booking.