Brain, Digital, & Learning

 Open access, Peer Reviewed

Indexed in KCI

pISSN 2384-2474
eISSN 2586-7490

RESEARCH ARTICLE

Designing an On-Premise, RAG-Based Rater Training Program for Korean Language Arts Essay Assessment

Korea National University of Education

Correspondence to Jinhwan Son, cdere0903@naver.com

Brain, Digital, & Learning. Volume 16, Number 3, 341–359, September 2026. https://doi.org/10.31216/BDL.2026.16.3.8
Received on July 14, 2026, Revised on September 29, 2026, Accepted on September 29, 2026, Published on September 30, 2026.
Copyright © 2026 Institute of Brain·AI based Education, Korea National University of Education This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

Abstract

The expansion of descriptive and essay-type assessment in South Korea’s 2022 Revised National Curriculum and 2028 College Admissions Reform Plan has intensified the need for reliable scoring by Korean Language Arts teachers. Existing rater training approaches face a persistent dual limitation: face-to-face workshops are constrained by spatiotemporal barriers and their one-off nature, while computer-based systems provide only superficial feedback that fails to address the cognitive processes underlying scoring decisions. This study proposes Retrieval-Augmented Generation (RAG) as a technology capable of bridging this gap and presents the design of a RAG-based rater training program. Drawing on Cognitive Apprenticeship, Frame-of-Reference training, and Deliberate Practice theory, six core design principles were derived through a two-axis analysis integrating theoretical mechanisms with design limitations identified in prior programs. These principles were translated into a seven-stage program following a 2-Round learning path (pretest–intervention–posttest). The program’s RAG system, built on open-source Korean-language LLMs (EXAONE 3.5 7.8B and SOLAR 10.7B) deployed on institutional on-premise servers, differentiates the roles of retrieval and generation across stages: RAG ensures knowledge consistency through precise case retrieval from an expert scoring database, while the LLM performs complex reasoning and personalized feedback synthesis. An adaptive feedback mechanism adjusts content and tone based on diagnosed scoring tendencies (leniency or severity), and a structured What-Where-Why scoring rationale requirement in the second scoring round promotes metacognitive reflection by making the cognitive process of scoring—not merely its numerical outcome—a target of training and feedback. As a design-stage study, this work presents the program’s architecture in implementable detail; empirical effectiveness verification through field deployment with practicing teachers is reserved for subsequent research.
Keywords

rater training, scoring reliability, retrieval-augmented generation, Korean language arts, essay-type assessment, cognitive apprenticeship, large language model

Section