[o-o][-_o]Natural Stupidity broke prod markets money ledger this week hall of fails studio cookbook toys about profile

The Human Bug Database

Bug reports filed against the human decision stack. Every ticket is closed WONTFIX: these are not defects in you — they are default settings of the species, reproducible in everyone, the authors of this page included. Each ships with a workaround from the debiasing literature.

Severities are playful. The citations are not.

NS-001 · STATUS: WONTFIX · SEVERITY: Major — unrecoverable resource leak

Sunk Cost Fallacy

Affected systems: budgeting, project continuation, buffet plates, half-watched series

Expected behaviorDecisions weigh only future costs and future benefits; money and effort already spent are gone either way.
Actual behaviorResources already sunk keep steering the decision — the more invested, the harder to stop, so bad commitments escalate precisely because they were expensive (Arkes & Blumer, 1985).

Steps to reproduce

  1. Pay in advance for something non-refundable.
  2. On the day, notice you would rather not go.
  3. Go anyway, 'so it isn't wasted' — and watch the ticket collect a second cost: your evening.

Workaround. Zero-based framing: ask 'starting fresh today, knowing what I know now, would I choose this?' If the answer is no, the past cannot vote. Writing down the two ledgers — future costs, future benefits, nothing else — makes the reframe stick.

NS-002 · STATUS: WONTFIX · SEVERITY: High — unvalidated external input

Anchoring

Affected systems: negotiation, pricing, estimation, first offers

Expected behaviorNumeric judgments derive from evidence; irrelevant numbers in the environment are ignored.
Actual behaviorThe first number on the table drags the estimate toward it, even when everyone knows the number is arbitrary — in the classic demonstration, a rigged wheel of fortune shifted people's estimates of unrelated quantities (Tversky & Kahneman, 1974).

Steps to reproduce

  1. See a tag reading 'was $499, now $299'.
  2. Feel the bargain.
  3. Notice the feeling arrived before any thought about what the thing is worth.

Workaround. Estimate first, look second: commit to your own number before seeing any proposed one. When anchored anyway, argue the other side — deliberately listing reasons the anchor is wrong measurably reduces its pull (Mussweiler, Strack & Pfeiffer, 2000).

NS-003 · STATUS: WONTFIX · SEVERITY: Critical — present in every subsystem, including the debugger

Confirmation Bias

Affected systems: search, interpretation, memory; all belief maintenance

Expected behaviorEvidence updates beliefs in whichever direction it points.
Actual behaviorSearch, reading, and recall all quietly favor what is already believed: supporting evidence gets sought, ambiguous evidence gets read as support, and disconfirming evidence fades first (Nickerson, 1998). The bug runs on exactly the beliefs it protects, so it inspects itself and reports clean.

Steps to reproduce

  1. Adopt a position.
  2. Search the web using the position's own wording.
  3. Read the agreeing results; classify the rest as biased.

Workaround. Consider the opposite, on purpose: before deciding, write down what evidence would change your mind, then go look for exactly that (Lord, Lepper & Preston, 1984). Verdict first, evidence after — the daily game's whole loop is this workaround with a scoreboard.

NS-004 · STATUS: WONTFIX · SEVERITY: Medium — cache consulted instead of database

Availability Heuristic

Affected systems: risk perception, frequency estimates, everything downstream of a news feed

Expected behaviorEstimates of how common things are track how common things are.
Actual behaviorEstimates track how easily examples come to mind, so vivid, recent, and heavily reported events read as frequent while quiet, common ones read as rare (Tversky & Kahneman, 1973).

Steps to reproduce

  1. Follow a week of dramatic coverage of some rare event.
  2. Estimate how likely that event is.
  3. Look up the recorded base rate; compare.

Workaround. Treat 'I can easily think of cases' as data about your feed, not about the world. Look the base rate up before estimating — the five-minute search beats the instant impression precisely where the impression feels strongest.

NS-005 · STATUS: WONTFIX · SEVERITY: High — silent data loss upstream of every query

Survivorship Bias

Affected systems: success studies, product lore, advice from winners

Expected behaviorConclusions are drawn from the whole sample, failures included.
Actual behaviorOnly survivors are visible — the failed ventures, unread books, and downed planes never enter the dataset, so their absence gets read as evidence. The canonical case: in WWII, statistician Abraham Wald observed that bullet-hole maps of returning bombers showed where a plane could be hit and still fly home; the untouched zones marked where the lost planes had been hit — so that is where the armor belonged.

Steps to reproduce

  1. Read interviews with ten wildly successful founders.
  2. Notice they share a bold trait.
  3. Draw the lesson — without access to the thousand un-interviewed founders who shared it too.

Workaround. Before drawing any lesson, ask 'what would this data look like if the missing cases were here?' Hunt failure stories with the same energy as success stories: the denominator you need is attempts, not the gallery of outcomes.

NS-006 · STATUS: WONTFIX · SEVERITY: Medium — sensor reports its own fault as nominal

Dunning-Kruger Effect

Affected systems: self-assessment, early-stage learning, hiring loops

Expected behaviorAccuracy of self-assessment is independent of skill level.
Actual behaviorIn many domains, the skill needed to perform is the same skill needed to judge performance. Kruger and Dunning (1999) found the lowest scorers on logic and grammar tests placed themselves around the 62nd percentile while scoring near the 12th — not because low skill breeds arrogance, but because it removes the instrument that would detect the gap; top scorers, if anything, underrated their relative standing. Honest footnote: part of the raw pattern is statistical (regression to the mean plus a general better-than-average habit), so read this as miscalibration at the edges of skill — never as 'stupid people think they are smart'.

Steps to reproduce

  1. Learn a little of a genuinely new domain.
  2. Register the confidence spike that arrives around week two.
  3. Watch an expert work; feel the scale recalibrate.

Workaround. Test, don't introspect: objective feedback — scored predictions, blind review, measured outcomes — stands in for the self-assessment instrument that low skill disables. Training pays twice, because improving the skill also improves the ability to judge the skill; that was Kruger and Dunning's own finding and fix.

NS-007 · STATUS: WONTFIX · SEVERITY: Major — every ETA ships optimistic

Planning Fallacy

Affected systems: estimates, roadmaps, renovations, theses

Expected behaviorNew estimates incorporate the recorded history of similar tasks.
Actual behaviorPlans get built from a best-case walkthrough of this task while the memory of every previous overrun sits unconsulted (Kahneman & Tversky, 1979). In the classic field study, students finished their theses well past their average estimate — and later than their stated worst case (Buehler, Griffin & Ross, 1994). At civic scale: the Sydney Opera House was scheduled for 1963 at 7 million Australian dollars and opened in 1973 at 102 million.

Steps to reproduce

  1. Estimate a task at one week.
  2. Deliver in three.
  3. Estimate the next, similar task at one week.

Workaround. Take the outside view: ignore this project's steps and ask what actually happened to the last ten projects like it, then base the estimate on that distribution (reference-class forecasting). Where history is thin, run a premortem — 'it is next year and this failed; write the story' — to surface the interruptions the walkthrough omits.

NS-008 · STATUS: WONTFIX · SEVERITY: High — prior probabilities dropped during deserialization

Base-Rate Neglect

Affected systems: screening, alarms, fraud and defect detection, diagnosis-shaped reasoning

Expected behaviorJudgments combine the specific evidence with how common the thing is to begin with.
Actual behaviorA vivid specific signal swamps the background frequency (Kahneman & Tversky, 1973). A 99%-accurate detector hunting a 1-in-10,000 defect raises roughly a hundred false alarms for every real find — yet each alarm feels 99% certain.

Steps to reproduce

  1. Deploy an accurate alarm for a rare event.
  2. Receive an alert.
  3. Act on the felt certainty, skipping the question 'out of 10,000 checks, how many alerts are real?'

Workaround. Translate to natural frequencies before judging: 'out of 10,000 cases, this many are real and this many will trigger the alarm.' Re-expressing probabilities as counts corrects most of the error (Gigerenzer & Hoffrage, 1995). Always ask for the base rate before admiring the accuracy.

NS-009 · STATUS: WONTFIX · SEVERITY: Low — but always on: history rewritten on read

Hindsight Bias

Affected systems: memory of predictions, postmortems, punditry

Expected behaviorPast uncertainty is remembered at its original size.
Actual behaviorOnce the outcome is known, memory quietly rewrites the earlier estimate toward it — 'knew it all along' (Fischhoff, 1975). Outcomes feel like they were always visible, which makes past decisions look negligent and future surprises look impossible.

Steps to reproduce

  1. Before an announcement, privately write down your probability.
  2. Afterwards, recall what you predicted — without looking.
  3. Compare with the paper. The paper wins.

Workaround. Keep a decision journal: date, decision, expected outcome with a probability, and the reasoning, reviewed on a schedule. The written record is the only version of your past beliefs that does not update itself — postmortems built on it judge decisions by what was knowable, not by what happened.

NS-010 · STATUS: WONTFIX · SEVERITY: Critical — confidence gauge pinned at 90 under all loads

Overconfidence Effect

Affected systems: forecasts, interval estimates, every 'I'm sure' said out loud

Expected behaviorConfidence tracks accuracy: claims made with 90% confidence come true about nine times in ten.
Actual behaviorConfidence reliably outruns accuracy, worst of all on hard questions and range estimates (Lichtenstein, Fischhoff & Phillips, 1982). In the classic interval experiments, ranges offered with '98% certainty' held the true value only around 60% of the time (Alpert & Raiffa).

Steps to reproduce

  1. Answer ten factual questions, rating your confidence on each.
  2. Score them.
  3. Compare average confidence with fraction correct; mind the gap.

Workaround. Calibration practice with immediate, scored feedback measurably narrows the gap (Lichtenstein & Fischhoff, 1980) — confidence is a trainable gauge, not a personality trait. This site keeps a test bench running for exactly that. Run the test bench: Overconfident

The Hall of Fails is a monthly field guide to several of these in the wild.

[o-o] ^ top