A measurement-validity benchmark · introductory probability & statistics
AssessAI asks whether multiple-choice items written by a language model (here Yuu v1) measure the statistics they claim to, or can be answered from their surface form. Every generation method writes items for the same 60 blueprint cells, and every item, reference or generated, goes through one battery of correctness, shortcut and robustness tests.
Status: framework, reference set (180 items) and a complete vertical slice are built; the full generation-and-probe run is pending. No student or human has answered or rated any item.
16constructs
35subskills
60blueprint cells
180items (reference set; generated pending)
Bar length on a log scale.
do generated items measure knowledge or surface form
- 01benchmarkSix generation arms and the reference set on the same 60 blueprint cells. The full run is pending; the reference set and vertical slice are complete.180items
- 02itemsEvery item, including the broken ones: stem, options, key, the evidence for the key, every flag that fired, and the model's own account of its distractors.180browsable
- 03generationZero-shot, blueprint, blueprint with reasoning, critic-and-revise, multi-agent and retrieval-grounded prompts, with the exact prompt text each arm sends.6arms
- 04qualityCorrectness is re-derived, not asked: program-of-thought execution for numeric keys, independent Monte-Carlo checks for reference keys.—key supported, gen.
- 05shortcutsOption-only, context-removed and stem-only probes, seven construct-free heuristics, and a lexical classifier that never sees the stem.—option-only, gen.
- 06robustnessStem paraphrases with a numbers-preserved check, fresh numeric draws with exact keys, and entity swaps with the key held fixed.—paraphrase agree.
- 07difficultyRequested difficulty against a model difficulty proxy. The proxy is an answering-model error rate and is never reported as student difficulty.20cells per level
- 08failuresWrong keys, multiple correct options, ambiguity, unrealistic distractors, construct drift, leakage, rationale errors and format artefacts, per arm.10categories
- 09psychometricsCTT, Rasch and 2PL estimators, verified by recovering planted parameters from simulated responses. No student has answered any item.0.994sim. corr(b)
- 10methods16 constructs, 35 subskills, five operationalised cognitive-demand levels, the pre-registered protocol and every post-freeze change.35subskills
- 11paperThe manuscript, generated from the release file so that every number in it traces to a persisted result.pendingstatus