AI in Assessment - Three Case Studies
Video recording of online session
AI generated summary: AI in Assessment — Three Case Studies
- 00:20–05:22 — Human teachers must remain the masters of AI. The inaugural remarks caution that, just as language labs and ICT tools were once seen as solutions to classroom problems, AI should not be treated as a panacea. Teachers should evaluate its usefulness critically and adapt it to Indian classroom contexts.
- 20:51–30:45 — The central problem: assessment workload and AI-era assignments. The speaker argues that conventional written assignments have lost some assessment value because students can readily use generative AI. This creates a need to move from merely assessing written products toward performance-based assessment and digital portfolios, while using AI to manage the resulting workload.
- 31:08–42:40 — Pre-session survey establishes teachers' assessment challenges. Among 103 respondents, MCQs were the most common assessment method (77%), followed by descriptive essays and presentations. Teachers identified lengthy answers, meaningful feedback, identifying strengths/weaknesses, comparing performances, and the time required for assessment as major difficulties. Most participants were comfortable with AI assisting in preliminary evaluation of descriptive answers.
- 43:04–54:55 — The key conceptual shift: AI as assessment assistant, not assessor. Assessment involves reading, interpreting, applying criteria, judging, scoring, and providing feedback—not merely assigning marks. The proposed model is Student Work → AI-Assisted Analysis → Rubric-Based Evidence/Patterns → Teacher Validation → Final Judgment & Feedback. The speaker emphasizes that the rubric is the backbone: the better the assessment design and criteria, the more useful AI assistance becomes.
- 55:21–1:09:40 — Case Study 1: AI-assisted assessment of essay answers. Handwritten literature answers were evaluated using teacher-designed prompts, the BAWE corpus, and CEFR guidelines. AI identified strengths and weaknesses in textual understanding, interpretation, argumentation, evidence, academic language, organization, and critical analysis. Students could then ask AI to show how their own answer might be improved rather than simply receiving an ideal model answer. The teacher remained responsible for validating the AI's evaluation and final score.
- 1:10:11–1:29:25 — Case Study 2: Video essays and performance assessment. Students produced video essays, presentations, literary performances, and short videos. AI-assisted analysis helped examine content, argument, language, delivery, visuals, creativity, coherence, and critical thinking. Tools such as Gemini, Adobe Enhance, and NotebookLM were explored. An important unexpected outcome was that students sometimes challenged AI feedback, encouraging critical engagement with assessment itself. The speaker concludes that AI can observe evidence, but humans must interpret its significance, especially regarding creativity, authenticity, cultural context, and nuance.
- 1:29:25–1:35:47 — Case Study 3: AI-assisted digital portfolio assessment. Portfolios containing essays, videos, presentations, reflections, research, projects, blogs, and creative work were analysed for evidence, reflection, growth, and competency development. AI helped identify patterns in students' learning journeys and possible future competencies/trajectories, while the teacher provided holistic judgment.
- 1:35:47–1:39:44 — The major pedagogical shift: from assessing answers to assessing learning. The teacher's role changes from marker → assessment designer → rubric designer → evidence interpreter → feedback designer → AI validator → ethical decision-maker. The proposed AI Assessment Triangle has three essential components: Assessment Design + Rubric + AI Assistance, resting on Human Judgment. If any component is weak, assessment quality suffers.
- 1:39:44–1:55:51 — Q&A reinforces validity, reliability, bias and human oversight. Questions address free/subscription tools, AI bias, validity and reliability, rubric constraints, language learning, and accessibility. The speaker stresses that AI-generated scores should never be accepted blindly and that anonymizing students as “Student 1, Student 2…” may help reduce potential name-, gender-, or caste-associated bias. AI should be AI-assisted, not AI-determined.
Core takeaway
The session's central message can be reduced to one principle:
The future of assessment is not “Human vs AI” but “Human + AI.”
AI can automate reading at scale, pattern identification, rubric application, evidence extraction, comparison, summarization, and preliminary feedback. The teacher must retain responsibility for context, interpretation, fairness, validity, ethics, final judgment, and academic decisions. This aligns closely with recent Glasp discussions emphasizing that AI can generate signals and support evaluation, but human judgment remains essential for deciding what those signals mean and what action should follow.
The most important five questions before using AI for assessment are:
- What exactly am I assessing?
- What evidence demonstrates it?
- What rubric will I use?
- What can AI reliably assist with?
- What must remain with human judgment?
This framework also resonates with current thinking on AI evaluation: criteria should be observable, AI judgments should be validated against human judgment, and there should always be a mechanism for human disagreement and override.
No comments:
Post a Comment