REVIEW ARTICLE
Utilizing Open Access Publications and Educational Content in LLM-Based Research: Global Trends and Ethical–Policy Considerations
Kuyng Sik Yi1,2, Chi-Hoon Choi1,2*
1Department of Radiology, College of Medicine, Chungbuk National University
2Department of Radiology, Chungbuk National University Hospital
Correspondence to Chi-Hoon Choi, chihoonc@chungbuk.ac.kr
Brain, Digital, & Learning. Volume 15, Number 2, 205–216, June 2025. https://doi.org/10.31216/BDL.2025.15.2.6
Received on May 26, 2025, Revised on June 14, 2025, Accepted on June 17, 2025, Published on June 30, 2025.
Copyright © 2025 Institute of Brain based Education, Korea National University of Education This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.
Abstract
Large language models (LLMs) have become increasingly ingrained in academic research and education. As a result, leveraging open-access publications and publicly available educational materials through LLMs is emerging as a promising strategy to broaden knowledge discovery and enhance teaching and learning. This review surveys global and domestic trends in the application of LLMs within science and education, highlighting notable implementations including Meta’s Galactica, KISTI’s KONI model, and the opensource CleverBee system. These examples showcase the substantial potential of generative AI when applied to openly available knowledge resources, while also revealing important limitations. In this context, we identify six key ethical challenges associated with LLM use—covering issues of copyright and licensing, personal data privacy, hallucinated content, algorithmic bias, plagiarism, and accountability—and examine evolving legal and regulatory responses in the United States, European Union, and South Korea. Additionally, we propose seven practical recommendations for researchers (such as checking data licenses, removing personal identifiers, verifying model outputs, and transparently reporting AI assistance) to foster responsible use of LLMs. Drawing on up-to-date guidelines from organizations like the Committee on Publication Ethics (COPE) and the International Committee of Medical Journal Editors (ICMJE) as well as recent national policy directives, this review offers comprehensive guidance for scholars across disciplines including neuroscience, linguistics, learning sciences, and artificial intelligence. Our aim is to support the responsible and effective integration of LLM technologies into academic practice—advancing innovation in research and education while maintaining high ethical and legal standards.
Keywords
Large language models (LLM), open access data, public educational content, ai ethics, privacy protection