Brain, Digital, & Learning

 Open access, Peer Reviewed

Indexed in KCI

pISSN 2384-2474
eISSN 2586-7490

RESEARCH ARTICLE

Keyword-based Scoring for The Prioritization of Differentially Expressed Genes Using Gene Ontology Annotations and Network Propagation

Laboratory of Animal Physiology and Medicine
Department of Biology Education, Korea National University of Education

Correspondence to Dongsun Park, dvmdpark@knue.ac.kr

Brain, Digital, & Learning. Volume 13, Number 4, 415–451, December 2023. https://doi.org/10.31216/BDL.20230026
Received on November 14, 2023, Revised on December 16, 2023, Accepted on December 17, 2023, Published on December 31, 2023.
Copyright © 2023. Institute of Brain based Education, Korea National University of Education This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

Abstract

Computational gene prioritization provides a basis for the utilization of high-throughput expression data; however, methods for prioritization considering relevance to both biological processes and selected keywords are lacking. In this study, a method was developed to rank differentially expressed genes (DEGs) by utilizing the Gene Ontology (GO) database for keyword searches and propagating the results through a protein–protein interaction network. A scoring system that effectively biases the scores of genes relevant to given keywords for avoidance and preference was established. This scoring method was combined with scoring based on expression characteristic groups (ECGs) with network propagation to obtain a final combined score (cScore). The performance of the new approach was evaluated using a rat middle cerebral artery occlusion (MCAO) dataset, revealing that the method more effectively filtered out DEGs compared with conventional methods based on both significance and fold change values, excluding 76% of genes in average, while retaining genes of interest. Further improvements, including addressing the inability of downward accumulation to balance terms with conflicting hits and the limited number of databases utilized for scoring, can further improve the performance of the method. Overall, the newly developed method can improve the interpretation of DEG data.
Keywords

Gene prioritization, network propagation, gene ontology, protein–protein interaction network, keyword search

Section