Signal detection separates sensitivity from willingness
Accuracy mixes perceptual sensitivity with decision criterion. Signal detection theory keeps the evidence distribution and the action threshold separate.
- problem
- A single accuracy score cannot show whether performance changed because evidence became easier to discriminate or because the observer adopted a more liberal or conservative response threshold.
- scope
- An equal-variance Gaussian signal-detection companion covering hit rate, false-alarm rate, d-prime, criterion, ROC curves, and the role of prevalence and decision cost.
- environment
- Conceptual yes/no detection tasks with synthetic ROC curves; no participant data, clinical interpretation, or diagnostic use.
Assumptions
- Noise and signal-plus-noise evidence are represented by equal-variance Gaussian distributions for the displayed formulas.
- Each trial has a defined signal-present or signal-absent state and a yes/no response.
- Hit and false-alarm rates are estimated with enough trials and corrected when they reach exactly zero or one.
Limitations
- Unequal variance, rating-scale models, non-Gaussian evidence, dependence between trials, learning, fatigue, and nonstationary criteria require richer models.
- Signal-detection measures describe task performance; they do not diagnose a person, reveal a mental state, or supply clinical advice.
- The synthetic ROC curves illustrate the equal-variance model and are not human performance data.
Table of contents 4 sections
Start with hits and false alarms
A yes/no detector produces four outcomes: hit, miss, false alarm, and correct rejection. Accuracy compresses all four under the task’s prevalence. Signal detection theory keeps the hit rate and false-alarm rate visible, then asks two separate questions: how far apart are the internal evidence distributions, and where did the observer place the criterion?
Under the equal-variance Gaussian model, d-prime measures separation in standard-deviation units. Criterion c describes response location relative to the two means. Two observers can have the same d-prime and different criteria, producing different hit and false-alarm rates without any change in underlying sensitivity.
d' = \Phi^{-1}(H) - \Phi^{-1}(F), \qquad c = -\frac{1}{2}\left[\Phi^{-1}(H)+\Phi^{-1}(F)\right]The ROC curve moves the threshold
A receiver operating characteristic curve traces hit rate against false-alarm rate as the decision criterion moves. A more liberal criterion raises both. A more conservative criterion lowers both. Sensitivity shapes the curve’s distance above the chance diagonal, while criterion selects an operating point on that curve.
The synthetic curves use the Gaussian relation H = Phi(Phi inverse of F plus d-prime). The higher-sensitivity curve reaches more hits for the same false-alarm rate. No point is universally optimal because the choice depends on prevalence, the cost of misses and false alarms, and the value of correct decisions.
H = \Phi\left(\Phi^{-1}(F) + d'\right)loading plot…
A criterion contains a cost model
Calling a response biased does not mean irrational. A smoke alarm should tolerate more false alarms than a low-stakes recommendation filter because the cost of a miss differs. A rare signal can also make overall accuracy look excellent while the detector misses most true cases. Prevalence and payoff alter the rational operating point even when sensitivity is fixed.
That makes criterion a design and policy variable. The system should state who pays for each error, whether feedback changes behavior, and whether the threshold is chosen by the observer, an institution, or an automated classifier. A threshold can be mathematically explicit and still encode a contested value judgment.
| context | dominant error cost | criterion pressure | required disclosure |
|---|---|---|---|
| smoke alarm | miss can be catastrophic | more liberal | nuisance-alarm burden and failure testing |
| spam filter | false positive can hide legitimate mail | more conservative | quarantine and recovery path |
| quality inspection | depends on defect and rework cost | cost-specific | prevalence, sampling, and downstream review |
| research task | inferential validity | protocol-specific | payoff matrix, trial balance, and corrections |
Measure before telling a psychological story
Signal detection theory is valuable because it resists a premature story. A change in yes responses may reflect sensitivity, criterion, prevalence learning, payoff, fatigue, or nonstationary evidence. Reporting only accuracy or confidence invites those mechanisms to collapse into one adjective about the participant.
The practical standard is to publish the contingency table, trial counts, rate corrections, model assumptions, uncertainty, and the criterion or sensitivity measure used. If the evidence distributions are unequal or the criterion changes across the session, the equal-variance summary becomes a starting point rather than the final model.
Why this is not a diagnostic tool
The model decomposes performance in a defined detection task. Clinical interpretation requires validated instruments, representative norms, measurement invariance, uncertainty, and qualified assessment. None of those can be inferred from an interactive ROC diagram.
- sourced Calculation of signal-detection measures
Stanislaw and Todorov review signal-detection tasks, sensitivity and response-bias measures, formulas, and practical calculation choices.
inspect source ↗ - inferred Synthetic equal-variance ROC curves
The plotted curves are generated from the displayed Gaussian model at d-prime values of 0.8 and 1.6; they are not fitted participant data.
- paper Calculation of signal detection theory measuresopen ↗
Behavior Research Methods, Instruments, & Computers (1999), with formulas and implementation guidance.