Professional Examiners Trained to APA Standards
140+ Professional Testing Locations Across the U.S. & Canada
Trusted by 10,000+ Clients, Attorneys & Organizations
LieDetectorTest.com Private & Confidential Polygraph Provider
Research ledger
Effects of differing numerical chart evaluation systems on polygraph examination results

HomePolygraph Research › Effects of differing numerical chart evaluation systems on…

Catalogue entry · Scoring & Test Data Analysis

Effects of differing numerical chart evaluation systems on polygraph examination results

Weaver, R.S. — Polygraph,

1985Published
Key findings

The three systems differed substantially in complexity of rules, scale granularity (3-position vs. 7-position scales), comparison question pairing strategies, and emphasis on specific physiological features (amplitude, duration, baseline changes). These seemingly minor procedural differences had significant implications for classification decisions and examiner outcomes.

Abstract

Weaver conducted a descriptive comparative analysis of three major numerical polygraph chart evaluation systems (Backster, USAMPS, and University of Utah) to document their scoring criteria and procedures. Rather than validating superiority, the study aimed to encourage familiarity with each system and promote further research into standardized, objective scoring methodologies.

Methodology

Weaver conducted a detailed technical comparative analysis systematically documenting and contrasting the scoring criteria, decision thresholds, and operational rules of three established numerical evaluation systems across respiration, electrodermal activity, and cardiovascular response channels.

Detailed summary

Published in 1985, this study addressed the mid-1980s transition from subjective global interpretation to systematic numerical scoring in polygraph chart evaluation. Weaver, a federal government examiner and APA Standards and Ethics Committee Chair, provided technical documentation of three major scoring systems to highlight critical differences in reaction identification criteria, numerical value assignment rules, score summation procedures, and decision thresholds. The analysis revealed that Backster's system employed numerous unvalidated complex rules while the Utah system used simpler empirically-based criteria, and systems differed in scale granularity and comparison question pairing. The study was motivated by the 1983 OTA report questioning polygraph validity and the profession's need for standardized, validated procedures.

Implications for polygraph practice

This comparative documentation of scoring system differences underscores the need for standardization and scientific validation of polygraph procedures, and demonstrates that scoring methodology significantly affects examination outcomes and reliability. The study's emphasis on procedural specificity informed subsequent efforts to establish professional standards and evidence-based scoring criteria within the polygraph profession.

Comprehensive study analysis

An in-depth, original analysis of this research study's methodology, findings, and significance for the polygraph profession.

Background & Context

By the mid-1980s, polygraph chart evaluation had evolved from purely subjective, "global" interpretation methods toward more systematic numerical scoring approaches. This study examined three major numerical evaluation systems: the Backster School of Lie Detection, the U.S. Army Military Police School (USAMPS), and the University of Utah. These systems represented different points along a continuum from clinical judgment to actuarial scoring.

Richard S. Weaver, a federal government examiner and Chairman of the APA Standards and Ethics Committee, sought to document the specific criteria and rules that distinguished these three approaches. At the time, there was considerable debate within the profession about which scoring system was most objective and reliable. The paper explicitly stated it was not intended to critically assess which system was superior, but rather to encourage familiarity with each system and promote further research into scientific objectivity.

This comparative analysis was particularly timely given the 1983 Office of Technology Assessment report questioning polygraph validity, which highlighted the need for standardized, validated scoring procedures. The paper built on Weaver's earlier 1980 work documenting the numerical evaluation methods developed by the United States Army and the University of Utah.

Research Design & Methodology

This was a descriptive comparative analysis that systematically documented and contrasted the scoring criteria and procedures of three established numerical evaluation systems. Rather than an empirical validation study, Weaver conducted a detailed technical comparison of the operational rules and decision criteria embedded in each scoring approach.

The analysis focused on several critical dimensions:

  • Reaction identification criteria for each physiological channel (respiration, electrodermal, cardiovascular)
  • Numerical value assignment rules for scoring the magnitude of responses
  • Score summation procedures for combining scores across questions and charts
  • Decision thresholds for classifying examinees as deceptive, truthful, or inconclusive

The paper emphasized that one of the most profound differences between the three systems was the criterion established for identifying the presence or absence of physiological reactions within each tracing. The analysis summarized the primary criteria each system used in interpreting responses to control and relevant test questions.

Results & Key Findings

The three systems were compared across respiration, electrodermal activity, and cardiovascular activity response areas. While the paper provided detailed technical documentation rather than statistical outcomes, several important distinctions emerged between the scoring systems.

Key differences identified included:

  • Complexity of scoring rules: Backster's system employed numerous complex rules that were not scientifically validated, while the Utah system used simpler, empirically-based criteria
  • Scale granularity: Systems differed in whether they used 3-position or 7-position numerical scales for assigning scores
  • Comparison question pairing: The Utah system at one time scored only the control preceding the relevant question, while other systems used different pairing strategies
  • Scoring criteria specificity: Each system emphasized different physiological features (amplitude, duration, baseline changes) for determining reaction significance

Differences among the systems were noted for the assignment of numerical scores in each physiological response, as well as in the summation of scores. The analysis revealed that seemingly minor procedural differences could lead to different test outcomes when the same chart data was evaluated under different systems.

Later research concluded that differences between the Army and Utah numerical scoring systems were slight, suggesting convergence in federal practice over time.

Discussion & Significance

Weaver's comparative analysis became a foundational reference in polygraph literature, cited extensively in subsequent research on scoring system development and validation. The paper was repeatedly referenced as the authoritative review of semi-objective numerical scoring techniques and helped establish a common vocabulary for discussing chart interpretation methodologies.

The systematic documentation of scoring differences highlighted a critical challenge for the field: the lack of standardization meant that the same physiological data could yield different conclusions depending on which system the examiner applied. This insight reinforced the need for empirical validation studies comparing system accuracy and inter-rater reliability. Research subsequently showed that Backster's numerical scoring system, while an improvement over global approaches, consisted of numerous complex rules that were not validated by scientific research.

The paper's influence extended beyond technical documentation. By encouraging familiarity with each system and calling for further research in improving scientific objectivity and additional research into interpretation criteria and comparison methods, Weaver helped establish a research agenda that shaped polygraph psychophysiology for decades. The comparative framework facilitated later meta-analyses and systematic reviews that could account for scoring system variations when aggregating validity data.

Limitations & Considerations

As a descriptive comparative analysis rather than an empirical validation study, the paper did not provide data on the relative accuracy, reliability, or validity of the three systems. The author explicitly stated it was not the purpose to critically assess which system was most objective, leaving questions about comparative performance unanswered. This limitation meant practitioners had detailed technical documentation but limited empirical guidance for selecting among the systems.

The analysis focused on the formal, published criteria for each system but could not capture variations in how individual examiners actually applied these rules in practice. Subsequent research revealed considerable variability in examiner scoring behavior even when nominally using the same system. Additionally, the paper represented scoring practices as of the mid-1980s; all three systems continued to evolve, with refinements and modifications introduced in later years based on ongoing research and field experience.

Practical Applications

This comparative framework provided essential technical documentation for training programs, accreditation bodies, and quality assurance systems. By clearly articulating the specific criteria distinguishing major scoring approaches, the paper enabled more informed discussions about standardization in federal and state polygraph practice. Many training curricula incorporated Weaver's analysis to help students understand the evolution and diversity of chart interpretation methodologies.

For contemporary examiners and consumers, this work underscores that "numerical scoring" is not monolithic—different systems can produce different results from identical physiological data. Understanding these differences is crucial when evaluating polygraph evidence, comparing studies using different scoring methods, or assessing examiner qualifications. The paper's lasting contribution is demonstrating that transparency about scoring methodology is essential for scientific credibility and forensic defensibility of polygraph results.

Read the original study

The analysis above is original editorial content based on our review of this research. For the complete study including full data, methodology details, and author discussion, access the original publication below.

Related research

Other studies in this category that may be of interest.

Join Our Examiner Network

APA-trained examiners using validated techniques can apply to join the LieDetectorTest.com network.

Apply now →

Keep reading the ledger.

Every peer-reviewed study on polygraph and deception detection we track — catalogued, searchable and citable.

Need to book now? Our online booking system is open 24/7. Speak directly with our team about your test or booking.