前期出版


26

生成式AI與人類教師於義大利文寫作評量之比較研究: 可信度、一致性和教學法相容度 Human-GenAI Alignment in Italian FL Writing Assessment: Reliability, Consistency, and Pedagogical Implications


作者
施喬佳
Author
Giorgia SFRISO
摘要

隨著生成式AI(GenAI)在學生間的廣泛使用,教育工作者面臨著將其從潛在挑戰轉化為教學助手的機遇。儘管現有研究已探索GenAI在自動化任務中的潛力,其在外語系學生的寫作質量評估中的可靠性、一致性與實用性仍是一個亟待探討的關鍵問題。本研究作為一項探索型案例研究,旨在深入比較教師與三種主流GenAI工具(ChatGPT、Gemini、Claude)在外語寫作評量上的表現。研究將選取五篇具備不同義大利語寫作能力的二年級學生作文,由三位大學教師與上述GenAI工具依據共同的評分標準(rubric)進行評量。此分析將聚焦於生成式AI的評量結果是否依據預先設定的評分標准進行評量,並取得一致性的結果,以及教師與生成式AI之評量是否存在顯著的差異。本研究的主要目標是探討GenAI是否能作為教師可靠的寫作評量輔助工具,以期潛在地減輕工作負擔並維持回饋品質,並提供GenAI在外語教學場景中的實務性應用具體指引。結果顯示,雖然GenAI尚不足以取代人類教師,但其在生成具體評分依據、診斷性回饋與錯誤模式識別方面展現出輔助潛力。教師可藉由ChatGPT進行穩定的基準評分、藉由Gemini從以學生為中心的教育概念提供支持型的回饋,藉由Claude在較弱的學生作文中突顯正面特質並搭配具體且可執行的指導建議。

Synopsis

Following learners’ widespread adoption of GenAI tools, educators are faced with the choice of treating these tools as potential challenges or turning them into instructional assets. While extant research has explored the potential of GenAI in automating tasks, its reliability, consistency, and pedagogical utility in assessing the writing ability of foreign language (FL) students remain critical areas of inquiry, even more so in the case of less commonly studied languages, such as Italian. This exploratory study compares human instructors and three prominent GenAI tools—ChatGPT, Gemini, and Claude—in the context of an Italian FL “Reading and Composition II” course. Five essays representing varying levels of proficiency penned by second-year Italian majors were selected and evaluated by three university instructors and the aforementioned GenAI tools; both the instructors and the GenAI tools all employed a standardized scoring rubric for evaluation. The analysis specifically examines the alignment of GenAI evaluations with predefined criteria, the consistency of its scoring, and the divergence between human and AI-generated assessment. To test for reliability and prompt sensitivity, the GenAI tools were evaluated under three distinct conditions: zero-shot individual assessment, comparative assessment, and sequential assessment. Qualitative thematic analysis was also applied to the generated feedback to identify patterns in tone, error identification, and pedagogical utility. The primary objective was to determine whether GenAI can serve as a reliable assessment assistant for instructors, potentially alleviating instructors’ workload while maintaining feedback quality. The findings indicate that while GenAI is not yet a substitute for human expertise, it demonstrates significant potential in generating specific grading justifications, providing feedback, and identifying error patterns. It may serve as a supplementary aid with the potential to reduce teacher workload and enrich the feedback ecosystem. Practical implications could include using ChatGPT for stable baseline grading, Gemini for providing supportive, student-centered feedback, and Claude for foregrounding positive aspects in weaker student performances when paired with actionable guidance.