Brain, Digital, & Learning

 Open access, Peer Reviewed

Indexed in KCI

pISSN 2384-2474
eISSN 2586-7490

RESEARCH ARTICLE

A Stage-wise Analysis of the Effects and Limitations of Small-Scale SFT in Korean CSAT-Style Question Generation

MHResearch

Correspondence to Giljae Kim, advisor-ai@mhresearch.net

Brain, Digital, & Learning. Volume 16, Number 1, 95–116, March 2026. https://doi.org/10.31216/BDL.2026.16.1.7
Received on March 27, 2026, Revised on March 31, 2026, Accepted on March 31, 2026, Published on March 31, 2026.
Copyright © 2026 Institute of Brain based Education, Korea National University of Education This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

Abstract

This study examines the effects and limitations of small-scale supervised fine-tuning (SFT) in Korean CSAT-style reading question generation. We propose a two-stage pipeline in which Stage 1 generates expository passages from keywords and Stage 2 generates multiple-choice questions and distractors from the generated passages. To identify stage-specific effects, we separately evaluate passage quality and item quality using a pairwise comparison framework with LLM-based judges and bootstrap-based statistical estimation. Experiments compare Qwen2.5-7B-Instruct and EXAONE 3.5 7.8B under base, QLoRA-based SFT, expanded-training-set, and JSON-output conditions. Results show that small-scale SFT does not affect all stages equally. In the Qwen family, fine-tuned models repeatedly underperform the base model in item-level evaluation, particularly in distractor plausibility and discriminative quality, while passage-level quality shows no consistent decline. This suggests that small-scale SFT may improve surface-level conformity to exam-style formats without improving substantive item quality. In contrast, the EXAONE family shows a different pattern, with some fine-tuned conditions yielding improvement signals in item quality. These findings indicate that the effectiveness of small-scale SFT is stage-dependent and model-dependent, and that passage generation and item generation should be evaluated separately in high-constraint educational generation tasks.
Keywords

Korean CSAT; automatic question generation; distractor generation; large language models; educational assessment

Section