psychometrics
Classical item difficulty, point-biserial discrimination, Rasch and 2PL IRT are implemented and tested. None has been estimated for an AssessAI item, because no one has answered one.
estimator recovery
2PL · corr(true b, est. b)0.99420 replications · 1000 simulated respondents · 40 items
2PL · corr(true a, est. a)0.949RMSE 0.102
2PL · corr(true θ, EAP θ)0.938
Rasch · corr(true b, est. b)0.99710 replications
| Respondents | corr(a) | corr(b) | RMSE(a) | RMSE(b) |
|---|---|---|---|---|
| 100 | 0.673 | 0.957 | 0.259 | 0.304 |
| 250 | 0.815 | 0.975 | 0.203 | 0.231 |
| 500 | 0.902 | 0.988 | 0.151 | 0.155 |
| 1000 | 0.949 | 0.994 | 0.106 | 0.117 |
planted misconceptions
Simulated students who hold a misconception choose the distractor coded with it, with 80% adherence, on the 60 reference items (instance 0). This checks that distractor coding and response analysis line up end to end.
| Misconception | Items with that distractor | Selection rate |
|---|---|---|
| M-SD-VS-VAR | 14 | 83% |
| M-BINOM-AT-MOST | 6 | 84% |
| M-INV-COND | 8 | 84% |
| M-JOINT-AS-COND | 8 | 87% |
| Profile | Proportion correct |
|---|---|
| high | 0.764 |
| medium | 0.504 |
| low | 0.238 |
| guesser | 0.249 |