Archives


 Vol.26 

Human-GenAI Alignment in Italian FL Writing Assessment: Reliability, Consistency, and Pedagogical Implications


Author
Giorgia SFRISO
Synopsis

Following learners’ widespread adoption of GenAI tools, educators are faced with the choice of treating these tools as potential challenges or turning them into instructional assets. While extant research has explored the potential of GenAI in automating tasks, its reliability, consistency, and pedagogical utility in assessing the writing ability of foreign language (FL) students remain critical areas of inquiry, even more so in the case of less commonly studied languages, such as Italian. This exploratory study compares human instructors and three prominent GenAI tools—ChatGPT, Gemini, and Claude—in the context of an Italian FL “Reading and Composition II” course. Five essays representing varying levels of proficiency penned by second-year Italian majors were selected and evaluated by three university instructors and the aforementioned GenAI tools; both the instructors and the GenAI tools all employed a standardized scoring rubric for evaluation. The analysis specifically examines the alignment of GenAI evaluations with predefined criteria, the consistency of its scoring, and the divergence between human and AI-generated assessment. To test for reliability and prompt sensitivity, the GenAI tools were evaluated under three distinct conditions: zero-shot individual assessment, comparative assessment, and sequential assessment. Qualitative thematic analysis was also applied to the generated feedback to identify patterns in tone, error identification, and pedagogical utility. The primary objective was to determine whether GenAI can serve as a reliable assessment assistant for instructors, potentially alleviating instructors’ workload while maintaining feedback quality. The findings indicate that while GenAI is not yet a substitute for human expertise, it demonstrates significant potential in generating specific grading justifications, providing feedback, and identifying error patterns. It may serve as a supplementary aid with the potential to reduce teacher workload and enrich the feedback ecosystem. Practical implications could include using ChatGPT for stable baseline grading, Gemini for providing supportive, student-centered feedback, and Claude for foregrounding positive aspects in weaker student performances when paired with actionable guidance.