Students Accept AI Feedback for Revision but Want Humans to Grade Their Work
Universities are experimenting with ChatGPT and similar tools to provide feedback on student writing. The assumption is that faster, more detailed feedback improves learning. But most studies ask whether AI feedback is useful or whether students like it. A qualitative study from a Saudi public university asks a different question: what happens when students know the AI graded their work? Thirteen male undergraduate computing students completed an in-class handwritten writing task in a technical communication course. Their scanned submissions were evaluated by ChatGPT using a rubric-aligned prompt. Students were then explicitly told that ChatGPT had generated the score and feedback, and asked to reflect on the evaluation in writing. The responses reveal a distinction that has not been clearly named in the prior literature: feedback utility versus evaluative authority.
The Study Design: Transparency as the Intervention
The setup was deliberately simple. Students wrote by hand during class, eliminating concerns about AI-assisted writing. The submissions were scanned and fed to ChatGPT with a structured rubric prompt aligned to the task objectives. Students received the AI-generated score and feedback and were explicitly informed that ChatGPT, not a human instructor, produced both. They then wrote reflections on the evaluation. This transparency is the key intervention. Most prior studies either hide the AI's role or do not emphasize it. Here, students knew exactly what they were evaluating.
The researchers used inductive thematic analysis on the reflection texts, meaning they derived themes from the data rather than testing predefined hypotheses. This is appropriate for a small-sample qualitative study exploring a new phenomenon, but it also means the findings are grounded in a specific context: male computing students at a Saudi university, one course, one task, one AI model. Generalization is not the claim.
Theme 1: Feedback Is Useful for Surface-Level Revision
Students consistently acknowledged that the AI feedback was helpful for identifying mechanical errors, structuring arguments, and improving clarity. They recognized that ChatGPT could point out grammatical issues, suggest reorganization, and flag weak reasoning. This aligns with prior work showing that students find AI feedback practical and actionable for revision tasks. The acceptance was genuine: students did not dismiss the feedback as worthless or reject it on principle.
What students did not accept was the idea that this useful feedback carried evaluative weight. They treated the feedback as a tool, like a grammar checker or a writing guide, not as a judgment. This is the first half of the utility/authority distinction: usefulness does not imply authority.
Theme 2: AI Cannot Understand Context
Students identified a specific limitation: ChatGPT lacks pedagogical and contextual awareness. It does not know the course objectives, the instructor's expectations, the student's progress over the semester, or the norms of the academic community. Several students noted that the AI feedback was generic, applicable to any writing task rather than tailored to the specific assignment. Some pointed out that the AI could not evaluate whether a technical claim was accurate within the course's domain, only whether it was well-articulated.
This is not a complaint about AI capability that will be solved by better models. It is a structural observation: an LLM evaluating a student's work in a specific course context will always lack the contextual knowledge that a human instructor accumulates through teaching the course, reading student work over time, and understanding the pedagogical goals. Students recognized this and rated the feedback accordingly.
Theme 3: Conditional Trust, Not Blanket Acceptance or Rejection
The third theme is the most analytically interesting. Students did not simply trust or distrust the AI. They drew a boundary: they trusted the feedback for specific, bounded purposes (catching errors, suggesting improvements) and withheld trust for grading decisions (assigning a score that affects their academic standing). This is conditional trust, applied selectively based on the task at hand.
The researchers frame this as two analytically separate judgments rather than opposite ends of a single approval scale. A student can think the feedback is helpful and simultaneously believe the AI should not assign the grade. These are not contradictory positions. They are independent evaluations of two different functions: informational utility (does this help me improve?) and evaluative authority (should this determine my grade?).
This has implications for how institutions communicate about AI-assisted assessment. If a university uses AI to generate feedback, students may accept it as a revision tool. If the same university uses AI to assign grades, students are likely to resist, even if the underlying model is the same. The difference is not in the model's capability but in the institutional claim about what the model is authorized to do.
Theme 4: The Human Instructor Remains the Grading Authority
Students consistently positioned the human instructor as the appropriate authority over grading decisions. This was not a statement about AI capability (some students acknowledged that the AI might be more consistent than a human grader) but about institutional legitimacy. The instructor knows the course, the students, and the standards. The instructor is accountable for the grade. The instructor's judgment carries authority in a way that an algorithm's does not.
Some students suggested a division of labor: AI for feedback, humans for grading. Others suggested that the instructor should review and adjust AI-generated scores before releasing them. A few proposed that AI grading might be acceptable for low-stakes assignments but not for major assessments. These are practical compromises, not principled rejections of AI. They reflect students' intuitive understanding that different assessment functions carry different authority requirements.
What This Means for AI in Education
The findings suggest that the current push to use AI for both feedback and grading conflates two functions that students treat as distinct. Feedback is informational: it tells you what to fix. Grading is authoritative: it determines your standing. Students accept the first and resist the second, not because they distrust the technology but because they understand the institutional structure of assessment.
For instructors and administrators, the practical takeaway is: be explicit about what the AI is doing. If it generates feedback, say so. If it assigns grades, say so. Students can handle transparency. What they resist is ambiguity about who is responsible for the judgment. The study also suggests that the most useful deployment of AI in assessment may be as a feedback tool that operates before human grading, giving students actionable input while leaving the authoritative judgment to the instructor.
The study's limitations are clear: 13 participants, one course, one task, one AI model, male students at a Saudi university. The qualitative method provides depth but not breadth. The findings should be treated as a first exploration of a distinction that deserves larger-scale investigation, not as a definitive account of student attitudes toward AI grading. The utility/authority framework, however, is likely to generalize across contexts where students encounter AI-mediated evaluation.