한약 처방 논문의 증거 수준 예측을 위한 딥러닝 모델 개발
Development of a Deep Learning Model to Predict the Evidence Levels of Research Papers Addressing Herbal Medicine Formula
Herbal medicine formula (HMF), a crucial treatment method in Korean medicine (KM), is increasingly being addressed in high-quality scientific journals, showing rapid growth both qualitatively and quantitatively. However, much valuable knowledge remains unstructured within a vast number of published papers. Recently, various studies have been conducted to extract knowledge from these unstructured papers by applying deep learning-based natural language processing (NLP) technologies. The levels of evidence, which indicate how reliable a particular study's findings are for making clinical decisions, play a crucial role in practicing evidence-based medicine. However, manually assessing the quality of research and determining its clinical applicability in the rapidly increasing number of papers related to HMF requires significant time and effort. Therefore, in this study, we aim to develop an algorithm to automatically determine the level of evidence of papers related to HMF by applying NLP methods. We constructed a corpus for AI training and testing. First, we selected 740 papers related to HMF and diseases randomly from PubMed using an HMF dictionary and a disease terminology dictionary. Experts in KM annotated the evidence levels to build the corpus. The distribution of evidence levels in the corpus was identified as follows: In-vivo Studies (61.22%), Randomized Control Trials (8.65%), and In-vitro Studies (7.84%). We fine-tuned BERT-based models with the built corpus to create a model that determines the evidence levels of a given paper. By evaluating the performance of four fine-tuned models, we found that SciBERT demonstrated the best performance with 94.59% (micro-F1), 89.11% (macro-F1), and 94.38% (weighted-F1).
목차
Abstract 서론 본론 1. 연구 방법 2. 연구 결과 3. 고찰 결론 감사의 글 참고문헌