Items / Reference templates / BP005
REF-BP005-2
C02 bayes rule · C02.2 base rate reasoning · target hard, CD3 · split dev · reference items carry their template's native demand and difficulty
Pending
A condition affects 0.2% of a population. A test is 99% accurate in both directions. Among everyone who tests positive, fewer than one in ten actually has the condition. Which statement best explains this result?
Explanation (as written by the item's author)
With a low prevalence, the false positives generated by the large healthy group dominate the positive results, so the posterior probability stays low even for an accurate test.
| Option | Code | Catalogue meaning | Author's rationale |
|---|---|---|---|
| A | MISC-OTHER | Distractor is not attributable to any catalogued misconception (itself a quality signal). | Blames sensitivity, which was stated and is not the cause. |
| B (key) | KEY | Correctly attributes the low posterior to the size of the non-diseased group. | |
| C | M-BASERATE | Neglects the base rate and equates the posterior with the test's sensitivity. | Equates the posterior with the test accuracy: base-rate neglect. |
| D | MISC-OTHER | Distractor is not attributable to any catalogued misconception (itself a quality signal). | Invents a restriction on what accuracy means. |
Quality flags · 1 fired
- absolute terms in distractors (cue) firedDistractors (only) use absolute terms such as always/never/must/proves, a known cue that they are wrong.
- duplicate options (integrity) not firedTwo options are identical after normalisation.
- numerically equal options (integrity) not firedTwo options denote the same number (e.g. 0.5 and 1/2 and 50%).
- missing options (integrity) not firedFewer than four non-empty options.
- key out of range (integrity) not firedThe keyed answer is a probability outside [0, 1] (or a percentage > 100).
- all none of above (cue) not firedAn option is 'all/none of the above' or 'both A and B'.
- key uniquely longest (cue) not firedThe key is the longest option by at least 20% (length cue).
- key uniquely shortest (cue) not firedThe key is the shortest option by at least 20%.
- distractor out of range (cue) not firedA distractor is a probability outside [0, 1]; it can be eliminated without the construct.
- stem option overlap cue (cue) not firedThe key alone repeats a content word from the stem.
- grammatical mismatch (cue) not firedThe stem ends in 'a'/'an' and some options do not agree with it.
- key is hedged only (cue) not firedOnly the key contains hedging language (may/might/plausibly/likely).
- numeric middle key (info) not firedOptions are numeric and the key is one of the two middle values. Informational; tested at set level against the 0.5 chance rate.
- negative stem (style) not firedStem uses NOT/EXCEPT/LEAST phrasing.
- overlong stem (style) not firedStem exceeds 110 words.
- inconsistent numeric format (style) not firedNumeric options mix fractions, decimals and percentages, or use differing decimal precision.
- unit mismatch (style) not firedOptions carry different units.
- option length imbalance (style) not firedLongest option is more than 2.5x the shortest.
- rationale value mismatch (metadata) not firedA distractor's stated rationale computes a value ('= x') that is not the option it is attached to: the distractor does not actually execute the misconception it claims.
Validator output
Not a numeric item: no program-of-thought check applies.
Answering probes
Probes not yet run for this item.
Shortcut results
| Heuristic | P(key) |
|---|---|
| longest option | 0.00 |
| shortest option | 0.00 |
| convergence | 0.00 |
| stem overlap | 0.00 |
| numeric middle | 0.25 |
| hedged option | 0.50 |
| position prior | 0.00 |
Provenance and generation trace
- Method
- reference_template / —
- Model
- none
- Source
- parameterised template T-C02.2-baserate, written by the AssessAI build agent (Claude); no external item bank; not human-reviewed
- Prompt hash
—- LLM calls · cost
- 0 · $0.00000 · 0.0 s
- Parser notes
- none