🎓 Lesson 7 D4

Scoring Consistency & Calibration Drills

Scoring consistency and calibration drills are structured practice sessions that help blasting engineers make the same accurate risk judgments every time—like using a ruler to measure rock hardness or blast vibration the same way, every shift.

🎯 Learning Objectives

  • Explain the difference between intra-rater and inter-rater reliability in qualitative risk scoring
  • Apply a standardized 5×5 risk matrix to assign consistent likelihood and consequence scores for common blasting hazards (e.g., flyrock, misfire, ground vibration)
  • Analyze discrepancies in team scoring outputs using kappa statistics and identify root causes of bias
  • Design a calibration drill protocol—including scenario prompts, anchor examples, and feedback loops—for a mine site’s pre-blast hazard review meeting

📖 Why This Matters

In mining, two engineers assessing the *same* blast design may rate flyrock risk as 'Medium' or 'High'—not because either is wrong, but because their mental models, experience, or unstated assumptions differ. Without consistent scoring, risk registers become unreliable, critical hazards get deprioritized, and audits reveal systemic gaps—not in engineering, but in communication. Calibration isn’t about conformity; it’s about shared language, traceable logic, and defensible decisions when regulators or courts ask: 'How did you decide this was acceptable?'

📘 Core Principles

Qualitative risk assessment relies on ordinal scales (e.g., 1–5) for likelihood and consequence—but these only work if anchors are explicit, observable, and context-specific. Consistency hinges on three pillars: (1) Criterion-based anchoring (e.g., 'Likelihood 4 = ≥1 event per 100 blasts, documented in site logs'), (2) Cognitive debiasing (mitigating overconfidence from recent near-misses or recency bias), and (3) Structured scoring protocols (e.g., mandatory use of evidence checklist before assigning any score). Calibration drills operationalize these by exposing hidden assumptions through comparative scoring, facilitated discussion, and iterative refinement against benchmarked reference cases.

📐 Inter-Rater Reliability (Cohen’s Kappa)

Cohen’s Kappa (κ) quantifies agreement beyond chance between two raters on categorical judgments—essential for validating calibration outcomes. Used after drill sessions to measure whether scorers converged meaningfully, not just randomly.

Cohen’s Kappa (κ)

κ = (Po − Pe) / (1 − Pe)

Measures inter-rater reliability for categorical data, correcting for chance agreement.

Variables:
SymbolNameUnitDescription
Po Observed proportion of agreement dimensionless Fraction of items where raters assigned identical scores
Pe Expected proportion of agreement by chance dimensionless Probability of agreement if raters scored randomly, based on marginal totals
Typical Ranges:
Untrained team pre-calibration: 0.10 – 0.35
Post-drill calibrated team: 0.61 – 0.80

💡 Worked Example

Problem: Two blasting supervisors independently scored 50 blast-related hazard scenarios using a 5-point likelihood scale. They agreed on 42 scores. Expected agreement by chance = 28 scores.
1. Step 1: Calculate observed agreement (Po) = 42 / 50 = 0.84
2. Step 2: Calculate expected agreement (Pe) = 28 / 50 = 0.56
3. Step 3: Apply κ = (Po − Pe) / (1 − Pe) = (0.84 − 0.56) / (1 − 0.56) = 0.28 / 0.44 = 0.636
Answer: The result is κ = 0.64, which falls within the substantial agreement range (0.61–0.80) per Landis & Koch (1977), indicating effective calibration for this pair.

🏗️ Real-World Application

At Newmont’s Boddington Mine (WA), a 2022 internal audit found <40% agreement among shift supervisors on 'vibration-induced structural damage' consequence ratings during pre-blast reviews. The site implemented biweekly 20-min calibration drills using anonymized past incidents (e.g., cracked portal at Pit 7B, 2021), anchored to MSHA vibration limits (1.0 in/s peak particle velocity for historic structures) and local building code thresholds. After six drills, inter-rater κ improved from 0.31 to 0.72—and subsequent blast permits showed 92% reduction in rework due to inconsistent risk justification.

📋 Case Connection

📋 Urban Tunnel Construction Vibration & Settlement Risk Model

Ground vibration threatening 19th-century masonry structures within 12m of tunnel alignment

📚 References