psychometrics

Classical item difficulty, point-biserial discrimination, Rasch and 2PL IRT are implemented and tested. None has been estimated for an AssessAI item, because no one has answered one.

estimator recovery

2PL · corr(true b, est. b)0.99420 replications · 1000 simulated respondents · 40 items
2PL · corr(true a, est. a)0.949RMSE 0.102
2PL · corr(true θ, EAP θ)0.938
Rasch · corr(true b, est. b)0.99710 replications
SIMULATED: 2PL recovery by number of respondents (40 items, 10 replications each)
Respondentscorr(a)corr(b)RMSE(a)RMSE(b)
1000.6730.9570.2590.304
2500.8150.9750.2030.231
5000.9020.9880.1510.155
10000.9490.9940.1060.117
Line chart: correlation between true and estimated 2PL parameters rises with the number of simulated respondents.
SIMULATED. Recovery of discrimination needs far more respondents than recovery of difficulty; a classroom-sized pilot would estimate b well and a poorly.

planted misconceptions

Simulated students who hold a misconception choose the distractor coded with it, with 80% adherence, on the 60 reference items (instance 0). This checks that distractor coding and response analysis line up end to end.

SIMULATED: selection of the planted distractor
MisconceptionItems with that distractorSelection rate
M-SD-VS-VAR1483%
M-BINOM-AT-MOST684%
M-INV-COND884%
M-JOINT-AS-COND887%
SIMULATED: proportion correct by student profile
ProfileProportion correct
high0.764
medium0.504
low0.238
guesser0.249