A faculty member discovered unexpected pedagogical value by asking AI to grade student papers while they graded the AI's feedback.
The experiment began during a grinding grading session. Rather than repeat identical comments about APA formatting across multiple assignments, the instructor fed student work to an AI system and let it generate feedback. Then the instructor graded the AI's responses using the same rubric applied to students.
The reversal proved illuminating. Evaluating AI feedback forced the instructor to confront what actually constitutes useful commentary. Many of the AI's responses, while technically accurate, lacked the specificity and context that transform feedback into learning. The AI caught formatting errors but missed the conceptual confusion buried in a paragraph. It praised effort without diagnosing misunderstanding. The exercise revealed how much of good teaching lives in the details and the gaps between what students write and what they meant to say.
This inversion also modeled what students experience when receiving feedback. By grading the AI, the instructor inhabited the student role, noticing which comments felt dismissive versus constructive, which ones prompted revision versus resignation. That perspective shift altered how the instructor now writes feedback for humans.
The broader lesson extends beyond any single classroom. As institutions adopt AI tools for grading assistance or draft feedback, this experiment suggests a useful check: instructors should evaluate the AI output itself. Does the system identify what students actually need to know? Does it treat feedback as a conversation or a verdict? Does it distinguish between surface errors and deeper misunderstandings?
The approach also raises questions about labor. AI can handle certain mechanical tasks, freeing instructors from repetitive work. But the experiment demonstrates that the thinking part of grading, the part that requires understanding a student's specific learning trajectory, remains distinctly human work. Technology works best when it handles what technology does well, not when it replaces the judgment teaching requires.
