RESEARCH ARTICLE
Enhancing Alignment Between Large Language Models and Teacher in Open-Ended Assessment through In-Context Learning
Juyeong Lee1* , Euna Lee2, Kangrea Kim3
1Seoul National University
2Osan Wonil Middle School
3Sundong Elementary School
Correspondence to Juyeong Lee, jy9307@snu.ac.kr
Brain, Digital, & Learning. Volume 15, Number 3, 375–401, September 2025. https://doi.org/10.31216/BDL.2025.15.3.4
Received on July 23, 2025, Revised on September 8, 2025, Accepted on September 10, 2025, Published on September 30, 2025.
Copyright © 2025 Institute of Brain based Education, Korea National University of Education This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.
Abstract
This study investigates the effectiveness of in-context learning (ICL) in enhancing the agreement between human teachers and large language models (LLMs) in the context of open-ended assessments. Using a dataset of 485 student responses to six open-ended questions from Korean, Technology, and Social Studies subjects administered in 2024, teacher-generated scores and feedback were collected alongside LLM-generated outputs under varying ICL conditions. Specifically, we provided GPT-4.1 with 0 to 20 examples in prompts to examine whether increasing example count improves agreement between the model and human raters. Quadratic Weighted Kappa (QWK) was used to assess score alignment, and BERTScore measured semantic similarity between teacher and model feedback. Regression and mixed-effects analyses revealed that increasing the number of examples generally improved alignment up to a certain threshold. The strongest improvements occurred with fewer than six examples, beyond which the benefits plateaued or even declined. Additionally, prompt length negatively moderated the effect of example count, suggesting that longer prompts may reduce the model’s capacity to focus on relevant information. These results provide practical guidance for teachers using LLMs in open-ended assessments. Including teacher-generated examples in prompts helps models align more closely with human scoring and feedback. However, the optimal number of examples depends on the type of question and expected answer length: more examples benefit shorter responses, while fewer examples (five or fewer) are more effective for longer or more complex answers.
Keywords
In-context learning, few-shot learning, few-shot prompting, large language models, automated scoring, feedback alignment, ai-assisted assessment, reliability in educational assessment