An excellent question — how reliable (consistent) are human graders? Would a given paper receive the same grade if it were graded by a given human near the beginning of the set versus near the end of the set? Or would intervening papers affect the grade due to cumulative perception of comparative quality? Or due to cumulative weariness of the human grader? Presumably a specific AI model would grade consistently across those cases because it operates with the same knowledge base. On the other hand, different AI models with different knowledge bases might generate different grades for the same paper — much like different human graders. All in all, this is a very interesting and potentially useful subject!
An excellent question — how reliable (consistent) are human graders? Would a given paper receive the same grade if it were graded by a given human near the beginning of the set versus near the end of the set? Or would intervening papers affect the grade due to cumulative perception of comparative quality? Or due to cumulative weariness of the human grader? Presumably a specific AI model would grade consistently across those cases because it operates with the same knowledge base. On the other hand, different AI models with different knowledge bases might generate different grades for the same paper — much like different human graders. All in all, this is a very interesting and potentially useful subject!