RESEARCH ARTICLE
Development and Validation of a Human-inthe-Loop (HITL) Model for AI-Based Automatic Item Generation in the CSAT Korean Reading Section
Korea National University of Education
Correspondence to Suji Lee, sasa0802@naver.com
Brain, Digital, & Learning. Volume 16, Number 2, 165–187, June 2026. https://doi.org/10.31216/BDL.2026.16.2.4
Received on June 19, 2026, Revised on June 30, 2026, Accepted on June 30, 2026, Published on June 30, 2026.
Copyright © 2026 Institute of Brain based Education, Korea National University of Education This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.
Abstract
This study develops and validates a Human-in-the-Loop (HITL) model for generative AI-based Automatic Item Generation (AIG) of College Scholastic Ability Test (CSAT) Korean reading items. As a high-stakes assessment of higher-order thinking, the CSAT demands rigorous reliability and validity (AERA, APA, & NCME, 2014; Kwon et al., 2017), yet hand-crafted development is resource-intensive and over-relies on developers’ assessment literacy (Jung, 2008; Rudner, 2009). Generative AI can ease these constraints but, used alone, is limited by hallucination and weak distractors, requiring expert supervision (Haladyna et al., 2002; Hommel et al., 2022). The HITL model therefore integrates AI efficiency with the educational and psychometric judgment of human experts (Burke, 2025; Sayin & Gierl, 2024). A 10-step protocol integrated CSAT development principles, the 2015 revised Korean curriculum, AIG theory, and Chain-of-Thought prompting (Wei et al., 2022), pairing an expert-alignment phase with a backward-design phase for creating, reviewing, and refining passages and items. Applying it, one researcher generated three passage-item sets (12 items in total) across three domains in 24 hours. Validation drew on expert ratings from 18 in-service Korean teachers and responses from 401 twelfth-grade students, examining readability, inter-rater reliability, CTT and IRT properties, and congruence between expert judgment and empirical performance. Generated passages showed no significant readability difference from past CSAT passages and scored high for formal completeness (M > 4.8) but lower for difficulty calibration and assessment suitability (p < .05). Items showed acceptable properties (α = .656, mean a = 1.097) but limited discrimination for critical-reading items (mean a = 0.504) and weak distractor attractiveness. Notably, the passage rated highest by experts for educational quality yielded the lowest mean item discrimination (ρ = -1.0). These findings show that AIG improves structural completeness, but human intervention and data-driven validation remain indispensable for psychometric rigor and educational validity, redefining the assessment expert as a psychometric designer and data-driven coordinator.
Keywords
Automatic Item Generation (AIG), Passage Generation, Human-in-the-Loop (HITL), College Scholastic Ability Test (CSAT), Reading Domain, Reading Assessment, Classical Test Theory (CTT), Item Response Theory (IRT)