<- Back to laws

Self-assessment calibration pattern

Dunning-Kruger Effect

Lower performers can show larger positive errors in self-assessment, but the size and mechanism of the pattern depend on measurement, information, task, and statistical design.

Scientific statusContested empirical effect
Predictive formConditional calibration error
DomainSkill and metacognition
EvidenceExperiments + replications
Key limitationStatistical and task dependence
Common misuseNovices are maximally confident
INTERACTIVE MODEL

calibration error = self-assessed performance - measured performance

The original studies compared measured performance with predicted performance or percentile rank. Difference scores, noisy tests, bounded scales, regression to the mean, and shared measurement error can create or amplify apparent group patterns.

The calibration lab separates an evidence-sensitivity account from a noisy-measurement artifact. Its curves are demonstrations, not a universal psychological law or a diagnosis of an individual.

57.5Illustrative self-estimate
(percentile)
0 %100 %
CALIBRATION, NOT A MEME CURVESeparate psychological updating from noisy measurement.
Interactive visual model for Dunning-Kruger Effect.
CALIBRATION ERROR0 ptABSOLUTE CONFIDENCE CLAIMNOT IMPLIED

The diagonal is perfect calibration. The cloud shows why quartile averages and difference scores can be misleading when performance itself is noisy.

CHANGE
Measured performance percentile
WATCH
self-estimate
MEANING
The calibration lab separates an evidence-sensitivity account from a noisy-measurement artifact. Its curves are demonstrations, not a universal psychological law or a diagnosis of an individual.
VISUAL MODEL

The important quantity is error, not a meme-shaped confidence curve.

Measured score and self-estimate are plotted on the same calibrated axes. Noise and weak sensitivity can both enlarge low-score overestimation while producing different evidence patterns.

measured performanceself-estimatecalibration error
01 / MEANING

What it actually says

The defensible core is narrower than the popular story. In several tasks, lower-scoring participants estimated themselves above their measured standing and showed larger positive calibration errors than higher performers. This does not imply that the least skilled are more confident in absolute terms than experts.

Competing explanations include limited metacognitive access, weak updating from task evidence, general above-average beliefs, scale boundaries, test unreliability, and regression to the mean. Serious analysis compares generative models and absolute calibration rather than drawing the familiar unsupported mountain-shaped meme.

Compact formcalibration error = self-assessed performance - measured performance
Best interpretationSkill and metacognition evidence in cognitive biases.
Important cautionStatistical and task dependence.
"A useful law compresses a pattern. It does not erase the conditions that make the pattern true."
02 / ORIGIN

How the idea developed

The modern form emerged through observation, argument, and later refinement. The timeline separates the first insight from the version now used in textbooks and practice.[1]

19991999

Kruger and Dunning report studies in humor, grammar, and logic linking low performance to inflated self-assessment.

2000s2000s

Replications extend the pattern while statistical critiques emphasize regression and measurement artifacts.

2020-212020-21

Large-sample and computational work tests artifact and evidence-sensitivity explanations.

TodayToday

Research focuses on calibration, metacognition, measurement reliability, incentives, and domain specificity.

Historical cautionEponymous laws often change after their first publication. Popular wording may be broader and cleaner than the original evidence.
03 / MECHANISM

How the pattern works

The relation becomes useful only when its mechanism, measurement process, and operating range are visible.

01Imperfect self-evidence

People observe noisy cues about their own performance rather than the latent skill directly.

02Evidence sensitivity

Low performers may update estimates less from item-level evidence or feedback.

03Scale and regression

Grouping on a noisy bounded score mechanically changes expected difference scores.

04Domain knowledge

Knowledge can improve both task execution and recognition of what a correct answer requires.

MODELcalibration error = self-assessed performance - measured performance

The original studies compared measured performance with predicted performance or percentile rank. Difference scores, noisy tests, bounded scales, regression to the mean, and shared measurement error can create or amplify apparent group patterns.

04 / APPLICATIONS

Where it earns its keep

Applications are strongest when the law changes a decision, measurement, model, or experiment rather than merely providing an analogy.

EDUCATION

Measure calibration alongside accuracy

Application

Learners can provide confidence per item and compare it with correctness.

PROFESSIONAL NOTE

Use repeated, reliable measures and feedback rather than labeling students.

PROFESSIONAL TRAINING

Design observable feedback loops

Application

Simulation, peer review, and outcome audits can reveal mismatches between judgment and performance.

PROFESSIONAL NOTE

Confidence is useful information only when task, scale, and consequences are specified.

RESEARCH

Separate mechanisms statistically

Application

Latent-variable, signal-detection, and generative models can test metacognition against artifact accounts.

PROFESSIONAL NOTE

Avoid quartile difference plots as the sole evidence.

05 / LIMITS & MISUSE

Where it stops working

Effects vary by task, scoring method, reliability, incentives, expertise range, culture, and whether participants predict raw scores or percentiles. Some apparent asymmetry follows from bounded measures and regression.

The construct does not license judging a person from disagreement or confidence. Individual diagnosis requires valid domain-specific performance evidence and repeated calibration data.

Misuse

"Beginners are more confident than experts"

Better: The classic result concerns calibration error, not necessarily absolute confidence.
Misuse

"The famous mountain curve came from the original paper"

Better: It did not; that viral curve is a popular invention.
Misuse

"Disagreement proves incompetence"

Better: The claim requires independent performance measurement.
Misuse

"The effect is either entirely real or entirely artifact"

Better: Observed patterns can contain psychological and statistical components simultaneously.
07 / REFERENCES

Sources and further reading

Original publications and serious secondary scholarship are prioritized over summaries.

  1. Kruger and Dunning - Unskilled and Unaware of ItThe original 1999 studies and proposed metacognitive account.https://doi.org/10.1037/0022-3514.77.6.1121
  2. Jansen, Rafferty, and Griffiths - A Rational Model of the Dunning-Kruger EffectLarge-scale replication and computational evidence-sensitivity model.https://doi.org/10.1038/s41562-021-01057-0
  3. Gignac and Zajenkowski - The Dunning-Kruger Effect Is Mostly a Statistical ArtefactReanalysis emphasizing measurement and statistical structure.https://doi.org/10.1016/j.intell.2020.101449
  4. McIntosh et al. - Reevaluating the Dunning-Kruger EffectLarge-data response finding a small residual effect under alternative analysis.https://doi.org/10.1016/j.intell.2022.101717
CONTINUE EXPLORING

Related laws, with the relationship made explicit.

These are editorial connections, not claims that the laws are mathematically equivalent.

CONTINUE READING

Place this law inside the collection.

LAW 039 / 100 PUBLISHED