If a student’s exam was scored by the same grader a second time, how similar were the two grades? By a second, but different, human rater? Contrast this with scoring the exam multiple times by AI. The instability of human scoring was one of the driving forces behind the shift to multiple choice exams in which inter-rater reliability was essentially perfect.
What about scoring reliability?
If a student’s exam was scored by the same grader a second time, how similar were the two grades? By a second, but different, human rater? Contrast this with scoring the exam multiple times by AI. The instability of human scoring was one of the driving forces behind the shift to multiple choice exams in which inter-rater reliability was essentially perfect.