년 - 년
발화 의도 예측 및 슬롯 채우기 복합 처리를 위한 한국어 데이터셋 개발 KCI 등재
한국융합학회 한국융합학회논문지 제12권 제1호 2021.01 pp.57-63
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
사람의 발화 내용을 이해하도록 하는 언어 인식 시스템은 주로 영어로 연구되어 왔다. 본 논문에서는 시스템과 사용자의 대화 내용을 수집한 말뭉치를 바탕으로 언어 인식 시스템을 훈련시키고 평가할 때 사용할 수 있는 한국어 데이터셋을 개발하고, 관련 통계를 제시한다. 본 데이터셋은 식당 예약이라는 고정된 주제 안에서 사용자의 발화 의도와 슬롯 채우기를 해야 하는 데이터셋이다. 본 데이터셋은 6857개의 한국어 문장으로 이루어져 있으며, 표기된 단어 슬롯 의 종류는 총 7개이다. 본 데이터셋에서 표기된 발화의 종류는 총 5개이며, 문장의 발화 내용에 따라 최대 2개까지 동시에 기입되어 있다. 영어권에서 연구된 모델을 본 데이터셋에 적용시켜 본 결과, 발화 의도 추측 정확도는 조금 하락 하였고, 슬롯 채우기 F1 점수는 크게 차이나는 모습을 보였다.
Spoken language understanding, which aims to understand utterance as naturally as human would, are mostly focused on English language. In this paper, we construct a Korean language dataset for spoken language understanding, which is based on a conversational corpus between reservation system and its user. The domain of conversation is limited to restaurant reservation. There are 7 types of slot tags and 5 types of intent tags in 6857 sentences. When a model proposed in English-based research is trained with our dataset, intent classification accuracy decreased a little, while slot filling F1 score decreased significantly.
초·중등 인공지능 교육을 위한 PISA 수학 맥락 중심의 데이터셋 개발 KCI 등재
한국정보교육학회 정보교육학회논문지 제27권 제3호 2023.06 pp.255-267
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
AI의 발전은 역사적으로 데이터셋과 깊은 관계가 있다. 최근 AI를 구성하는 데이터셋의 중요성이 강조됨에 따 라 관련 연구가 많이 이루어지고 있지만 AI 교육 측면에서 데이터셋 관련 연구는 상대적으로 부족한 실정이다. 이 에 본 연구는 학생들에게 유의미한 AI 교육용 데이터셋을 제공하기 위해 교수·학습 환경에서 맥락의 중요성을 확 인하고, 컴퓨팅 사고력을 포함하는 PISA 2022 수학의 맥락을 중심으로 AI 교육에서 학생들이 선호하는 데이터셋 의 맥락과 형태 및 교육용 도구를 확인하였다. 이를 바탕으로 AI 교육에 적합한 데이터셋 주제를 탐색하고 다양한 데이터셋을 합성, 수집, 수정 및 개발하였다. 또한 데이터셋의 적합성을 검증하기 위해 전문가 검토를 실시하고 그 결과를 반영하여 25종의 AI 교육용 데이터셋을 도출 및 배포하였다. 본 연구의 결과가 다양한 관점의 AI 교육을 위한 데이터셋 관련 연구에 기반이 되어 학생들의 AI 소양을 기르는데 도움이 될 수 있기를 기대한다.
The development of AI has historically been strongly tied to datasets. Recently, as the importance of datasets in AI has been emphasized, there has been a lot of related research, but there is a relative lack of research on datasets in the context of AI education. In order to provide students with meaningful datasets for AI education, this study identified the importance of context in the teaching-learning environment and identified the context, form, and educational tools of datasets preferred by students in AI education, focusing on the context of PISA 2022 mathematics, which includes computational thinking skills. Based on this, we explored dataset topics suitable for AI education and synthesized, collected, modified, and developed various datasets. We also conducted an expert review to verify the suitability of the datasets, and based on the results, we derived and distributed 25 datasets for AI education. We hope that the results of this study will serve as a basis for research on datasets for AI education from various perspectives and help students develop their AI competency.
본 논문에서는 화재 상황과 화재 현장에 있을 수 있는 객체에 대한 인공지능 학습용 데이터셋을 구축하였다. 학습용 데이터 셋은 화재 현장에서의 불꽃, 연기, 정상 상황을 식별하기 위한 동영상 데이터 5,304개, 이미지 데이터 약 190만개를 구축하였고, 화재 현장에서 식별 가능한 객체 15종을 31,500개씩 총 472,500개의 데이터를 추가에 성공하였다. 구축한 데이터의 품질을 평 가하기 위해 AI 모델로 테스트한 결과, 화재 감지 정확도는 90.98%, 객체 감지 정확도는 95.48%의 결과를 도출하였다.
YOLO 기반 포트홀 탐지 성능 향상을 위한 하이퍼파라미터 최적화 및 데이터셋 확장 비교
한국혁신산업학회 혁신산업기술논문지 제4권 제1호 2026.03 pp.11-17
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 도로 결함(포트홀 및 균열) 탐지를 위한 YOLO 기반 딥러닝 모델 최적화 및 실용성 검증을 목표로 수행되었다. NVIDIA A100 GPU 환경에서 진행된 첫 번째 실험에서는 RDD2022 공개 데이터셋을 활용하여 YOLOv5n/v8n/v10n 모델들의 성능 중 YOLOv8n 모델이 mAP@0.5에서 0.3657로 가장 우수한 성능을 보였다. 두 번째 실험부터는 자체 구축한 실제 도로 이미지를 바탕으로 YOLOv8의 Detection 및 Segmentation 모델을 비교 평가하였으며, YOLOv8l(Detection) 모델이 mAP@0.5에서 0.772로 실무에 적합한 정확도를 달성하였다. 마지막으로 YOLOv8l 모델의 하이퍼파라미터 최적화/데이터 증강 기법/데이터셋 확장을 비교한 결과 데이터셋을 기존 8,737장에서 14,776장으로 확장한 경우 mAP@0.5에서 0.825까지 향상되었다. 이는 학습 데이터가 실제 도로에 얼마나 유사하고 데이터의 양이 모델 성능의 결정적인 영향을 미친다는 것을 입증한 결과로, 향후 도로 결함 탐지 시스템의 실용성 및 즉각 대응 가능성을 확보하는 데 이바지할 것으로 기대된다.
This study aimed to optimize and validate the practicality of a YOLO-based deep learning model for detecting road defects (potholes and cracks). In the first experiment conducted in an NVIDIA A100 GPU environment using the RDD2022 public dataset, the YOLOv8n model demonstrated the best performance among the YOLOv5n/v8n/v10n models, achieving an mAP@0.5 of 0.3657. From the second experiment onwards, the Detection and Segmentation models of YOLOv8 were comparatively evaluated using self-built real-world road images. The YOLOv8l (Detection) model achieved an mAP@0.5 of 0.772, demonstrating accuracy suitable for practical applications. Finally, a comparison of hyperparameter optimization, data augmentation techniques, and dataset expansion for the YOLOv8l model showed that expanding the dataset from 8,737 to 14,776 images improved the mAP@0.5 to 0.825. This demonstrates that the volume of data and the similarity of training data to real-world environments critically influence model performance. These findings are expected to contribute to securing the practicality and immediate responsiveness of future road defect detection systems.
AI 시대 데이터 기반 한시 연구는 어떻게 가능한가 - ‘한국 한시 데이터 아카이브’ 구축에 관하여 - KCI 등재
제주대학교 인문과학연구소 인문학연구 제39집 2025.07 pp.63-82
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
본 연구는 한국 한시 데이터 아카이브 구현을 목표로 설계한 한국 한시 데이터셋의 구축 상황을 소개하고, 이 데이터셋이 추후 데이터 아카이브의 형태를 갖추기 위해 필요하 다고 여겨지는 몇 가지 요소에 대해 논의하는 데 목적을 둔다. 특히 본고에서는 상용 가능한 수준의 기계가독형 한국 한시 데이터셋 구축을 최우선적 인 과제로 보고, 해외 사례인 “chinese-poetry”와 비교하는 방식으로 한국 한시 데이터셋 의 구현 방안을 모색하여 보았다. 나아가 <한국고전종합DB>에 수록된 약 350여 종의 문 집, 17만 수 정도의 한시를 최대한 온전한 형태로 수록하고 이를 공유하였다. 이는 기존의 작업에 비해 시의 형태가 더욱 확실하게 밝혀져 있는 것으로, 이를 통해 수반되는 정보가 앞으로의 문학 연구에서 참고할 만한 것이 되기를 희망한다. 한국 한시 데이터 아카이브의 구현을 위한 또다른 주요 쟁점으로 주제 분류, 시어 전고사전, 인물·장소 등에 대한 문맥정보 및 전거 제어에 대해 논의하였다. 주제 분류의 경우, 조선 전기 분문찬류서 『풍소궤범』의 분류 형태를 기준으로 하여 가능한 개선 방안에 대해 논의하였다. 시어 전고 사전은 특히 『한어대사전』을 예시로 하여, 시어 사전(glossary)이 시어의 태깅에 미치는 실질적인 영향과 그 의의에 대해 논의하였다. 문맥정보 및 전거제어 의 경우, 개체명 인식 등의 대안이 논의되고 있지만 한시의 경우 기계적인 마킹은 쉽지 않을 것으로 여겨지며 결국 전문가의 판단이 필요할 것으로 보인다. 궁극적으로, 디지털 아카이브란 하나의 생태와 같다. 이를 활용하고 관리하는 연구 생태계를 조성하는 것이야말로 한국 한시 데이터 아카이브의 장기적 성공을 위한 핵심 전략 이라 하겠다.
This study aims to introduce the construction status of a Korean Hanshi dataset designed with the goal of implementing a Korean Hanshi Data Archive, and to discuss several elements considered necessary for this dataset to eventually take the form of a data archive. In particular, this paper prioritizes the construction of a commercially viable, machine-readable Korean Hanshi dataset and explores implementation strategies for the Korean Hanshi dataset by examining overseas cases. Furthermore, we have collected and shared approximately 350 literary collections and around 170,000 Hanshi poems from The Korean Classics Comprehensive DB(한국고 전종합DB) in the most complete form possible. We demonstrated through examples that the preservation of complete poetic forms, compared to previous work, has revealed several noteworthy findings. As additional major issues for implementing the Korean Hanshi Data Archive, we discussed subject classification, a dictionary of poetic allusions, and contextual information and authority control for figures and places. For subject classification, we attempted to find solutions in classical classification systems that reflect considerations of Hanshi tradition rather than modern classification approaches. Specifically, we discussed possible improvement measures based on the classification format of Pungsogwebeom(풍소궤 범), an early Joseon period anthology organized by literary genres. For the dictionary of poetic allusions, we discussed the practical impact and significance of glossaries on tagging, using the Han Yu Da Ci Dian(漢語大詞典) as an example. Regarding contextual information and authority control, while alternatives such as named entity recognition are being discussed, complete marking in Hanshi cases is not easily achieved, ultimately requiring expert intervention. Ultimately, a digital archive is like an ecosystem. Creating a research ecosystem that utilizes and manages it is the core strategy for the long-term success of the Korean Hanshi Data Archive.
과적 의심차량 추적에 관한 방법 연구 KCI 등재
한국기계항공기술학회(구 한국기계기술학회) 한국기계항공기술학회지(구 한국기계기술학회지) 제25권 제6호 2023.12 pp.979-984
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
In contemporary society, traffic concerns and road safety are of paramount importance. Overloading and improper loading significantly jeopardize road safety. Leveraging AI Hub's CCTV traffic videos, overloading hazard data, and Yeosu Sunsin Bridge HS-WIM data from Jeollanam-do, we've crafted detection models. These models facilitate the development of a vehicle detection system calibrated for specific scenarios, making the detection of overloading and loading discrepancies more accessible. Such a system promises to streamline detection operations while substituting manual efforts with AI, thereby economizing on both time and resources.
데이터거래소에서의 개인정보 데이터 가격 형성에 대한 연구 : 데이터 공급자 중심으로 KCI 등재
한국산업안보학회(구 한국산업보안연구학회) 한국산업보안연구 제12권 제3호 통권 제25호 2022.12 pp.7-25
※ 기관로그인 시 무료 이용이 가능합니다.
5,400원
최근 축적된 데이터의 활용도가 높아짐에 따라, 필요한 데이터의 확보를 위해 시장을 통해 데 이터 소유자와 사용자 간의 직접 거래가 늘어나고 있다. 다만 수집 데이터 내 정보 주체와 관련 한 항목은 특정 개인을 유추할 수 없도록 비식별 조치가 선행되어야 한다는 규정이 데이터 활용 에 대한 저해 요인으로 지목된다. 데이터 소비자는 데이터 사업 수행 시 양질의 데이터 부족과 함께 불합리한 데이터 가격 정책을 주요 문제점으로 제시한다. 그리고 정부는 데이터 거래 활성 화를 통한 산업화의 기반 조성을 위해 데이터 산업법을 제정하고, 데이터 가치 평가기관 제도를 도입하여 보다 객관적인 가치평가 방법을 제시하기 위해 노력하고 있다. 이에 본 연구는 현재 데 이터 거래소 시장 환경에서 판매자가 제시하는 가격의 통계 분포를 살펴보고, 데이터 소유자의 판매 가격에 영향을 미치는 요인에 대해 알아보고자 한다. 데이터 가격의 분포가 판매자의 자의 적 가치 기준에 의한 결정인지 여부를 판단하여 데이터 가치평가 기관의 도입을 통한 보다 합리 적인 데이터 가격 정책의 가능성을 살펴본다. 이를 위해 데이터의 가격 결정에 영향을 미친다고 알려진 항목과 무형자산의 가격 결정 모델을 살피고, 데이터셋의 분류에 따라 판매자의 제시 가 격 분포를 분석한다. 더불어 국내 데이터거래소에서 개인정보의 익명 처리가 수반된 금융 데이 터의 경우에도 판매자의 제시 가격 분포가 동일한 결과를 보이는지 확인한다.
As the utilization of accumulated data increases, direct transactions between data owners and users are increasing through the market to secure necessary data. However, the regulation requiring items within collected data to undergo de-identification measures to prevent specific individuals from being inferred is considered an obstacle in data usage. This paper aims to analyze the distributions of data owner’s list prices within the actual data market environment. To this end, this paper examines the pricing models of intangible assets and data aspects known to affect data pricing. We also analyze the effect of personal information within a dataset on the seller's sales price presentation.
KIS(KETI Infrastructure Stitching) Dataset : 다수 엣지 기반 카메라/라이다 데이터 스티칭을 통한 인프라 데이터셋
한국ITS학회 한국ITS학회 학술대회 ITS, Connected World 2024.10 pp.267-270
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
A Dataset of Online Handwritten Assamese Characters
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.11 No.3 2015 pp.325-341
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper describes the Tezpur University dataset of online handwritten Assamese characters. The online data acquisition process involves the capturing of data as the text is written on a digitizer with an electronic pen. A sensor picks up the pen-tip movements, as well as pen-up/pen-down switching. The dataset contains 8,235 isolated online handwritten Assamese characters. Preliminary results on the classification of online handwritten Assamese characters using the above dataset are presented in this paper. The use of the support vector machine classifier and the classification accuracy for three different feature vectors are explored in our research.
Effective dataset construction method using Dexofuzzy based on Android malware opcode mining
[NRF 연계] 한국통신학회 ICT Express Vol.8 No.3 2022.09 pp.444-462
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Recently, research on Android malware categorization and detection is increasingly directed toward proposing different learned models based on various features of Android apps and machine learning algorithms. For the implementation of such modeling, properly constructing a dataset is no less important than selecting a suitable algorithm. The present study examines dataset construction using Dexofuzzy and proposes methods to determine the degree of bias and variance in the process and minimize the noise in sample set labeling where there is a possibility that even the same samples can be differently labeled. The method proposed in the present study goes beyond existing dataset construction methods relying on label data provided by AV vendors to include an effective approach to construct new types of datasets built on unified labels combined with opcode morphology. Based on newly constructed datasets, a flexible dataset, which allows overfitting and underfitting to be considered, was obtained via N-Gram and M-Partial Matching. This flexible dataset was then subjected to clustering, and the resultant clustering performance was evaluated.
Waymo Open Dataset 기반 자율차의 주행행태분석을 통한 주행안정성 평가지표 도출 KCI 등재
한국ITS학회 한국ITS학회논문지 제23권 제4호 통권114호 2024.08 pp.94-109
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
무인 자율차의 공도 주행이 허용됨에 따라 연구에 활용가능한 자율차의 실도로 주행 데이터 가 증가하는 추세이다. 따라서 혼합교통류 상황에서 실제 자율차가 교통안전에 미치는 영향을 분석할 수 있게 되었다. 자율차가 교통안전에 미치는 영향을 파악하기 위해서는 자율차의 주행 행태를 효과적으로 반영할 수 있는 평가지표의 활용이 요구된다. 본 연구의 목적은 Waymo Open Dataset을 통해 자율차의 주행행태를 분석하여 단속류 도로 구간별 주행안정성을 평가하 기 위한 주요 지표를 도출하는 것이다. 주성분 분석을 통해 단속류 도로 구간별 데이터에 대한 설명력이 높은 평가지표를 선별하고 주요 평가지표로 정의하였다. 이때, 종방향과 횡방향 주행 안정성을 구분하여 각각에 대한 주요 평가지표를 제시하였다. 이후 동일한 주요 평가지표가 도출된 단속류 도로 구간을 대상으로 주행안정성을 비교하였다. 비신호교차로 대비 곡선 단일 로 구간에서 종방향 주행안정성이 약 35.48% 높게 도출되었다. 횡방향 주행안정성의 경우 비신 호교차로 대비 신호교차로 구간에서 주행안정성이 76.08% 높게 도출되었으며, 직선 단일로가 곡선 단일로에 비해 146.87% 높은 것으로 도출되었다. 본 연구의 결과는 자율차의 실도로 주행 데이터를 활용한 자율차의 교통안전 영향 분석 시 기초 자료로 활용할 수 있을 것으로 기대된다.
As autonomous vehicles are allowed to drive on public roads, there is an increasing amount of on-road data available for research. It has therefore become possible to analyze impacts of autonomous vehicles on traffic safety using real-world data. It is necessary to use indicators that are well-representative of the driving behavior of autonomous vehicles to understand the implications of them on traffic safety. This study aims to derive indicators that effectively reflect the driving stability of autonomous vehicles by analyzing the driving behavior using the Waymo Open Dataset. Principal component analysis was adopted to derive indicators with high explanatory capability for the dataset. Driving stability indicators were separated into longitudinal and lateral ones. The road segments on the dataset were divided into four based on the characteristics of each, which were signalized and unsignalized intersections, tangent road section, and curved road section. The longitudinal driving stability was 35.48% higher in the curved road sections compared to the unsignalized intersections. With regard to the lateral driving stability, the driving stability was 76.08% higher in the signalized intersections than in the unsignalized intersections. The comparison between curved and tangent road segments showed that tangent roads are 146.87% higher regarding lateral driving stability. The results of this study are valuable for the further research to analyze the impact of autonomous vehicles on traffic safety using real-world data.
[NRF 연계] 한국통신학회 ICT Express Vol.12 No.3 2026.06 pp.752-757
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Zero-day attacks threaten IoT security as signature-based detection fails against novel exploits. This paper proposes a hybrid Intrusion Detection System integrating unsupervised anomaly detection, non-parametric Siamese-based cross-dataset dissimilarity filtering, and Proximal Policy Optimization (PPO)-based adaptive defense. Unsupervised models isolate anomalous traffic, Siamese-based correlation extracts structurally rare zero-day candidates, and the PPO agent learns optimal defense policies via environmental feedback. Evaluations on CIC-IoT-2023 and CIC-BCCC-NRC-TabularIoTAttacks-2024 demonstrate 99.28% training accuracy, 99.07% unseen attack accuracy, and 93.94% zero-day detection rate with 0.50 ms latency and 2.21% false-positive rate, providing a scalable, proactive, self-learning defense architecture for autonomous IoT cybersecurity.
Robust cross-dataset deepfake detection with multitask self-supervised learning
[NRF 연계] 한국통신학회 ICT Express Vol.11 No.5 2025.10 pp.858-862
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Deepfake detection is increasingly critical due to the rise of manipulated media. Existing methods often require extensive datasets and struggle with interpretability issues. To address these issues, this study introduces a novel one-class approach for detecting and localizing deepfake artifacts in videos, using authentic images to generate manipulated data for training. By integrating segmentation and leveraging convolutional neural networks with visual transformers, the method predicts both the presence and location of the generated manipulations. Experiments on seven deepfake datasets and emerging diffusion-based manipulations show that our approach consistently outperforms existing methods, demonstrating superior accuracy and localization capabilities.
Manchu Script Letters Dataset Creation and Labeling
[Kisti 연계] 한국정보통신학회 Journal of information and communication convergence engineering Vol.22 No.1 2024 pp.80-87
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The Manchu language holds historical significance, but a complete dataset of Manchu script letters for training optical character recognition machine-learning models is currently unavailable. Therefore, this paper describes the process of creating a robust dataset of extracted Manchu script letters. Rather than performing automatic letter segmentation based on whitespace or the thickness of the central word stem, an image of the Manchu script was manually inspected, and one copy of the desired letter was selected as a region of interest. This selected region of interest was used as a template to match all other occurrences of the same letter within the Manchu script image. Although the dataset in this study contained only 4,000 images of five Manchu script letters, these letters were collected from twenty-eight writing styles. A full dataset of Manchu letters is expected to be obtained through this process. The collected dataset was normalized and trained using a simple convolutional neural network to verify its effectiveness.
TOD: Trash Object Detection Dataset
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.18 No.4 2022 pp.524-534
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this paper, we produce Trash Object Detection (TOD) dataset to solve trash detection problems. A well-organized dataset of sufficient size is essential to train object detection models and apply them to specific tasks. However, existing trash datasets have only a few hundred images, which are not sufficient to train deep neural networks. Most datasets are classification datasets that simply classify categories without location information. In addition, existing datasets differ from the actual guidelines for separating and discharging recyclables because the category definition is primarily the shape of the object. To address these issues, we build and experiment with trash datasets larger than conventional trash datasets and have more than twice the resolution. It was intended for general household goods. And annotated based on guidelines for separating and discharging recyclables from the Ministry of Environment. Our dataset has 10 categories, and around 33K objects were annotated for around 5K images with 1280×720 resolution. The dataset, as well as the pre-trained models, have been released at https://github.com/jms0923/tod.
Knowledge Model for Disaster Dataset Navigation
[Kisti 연계] 한국과학기술정보연구원 Journal of information science theory and practice : JISTaP Vol.9 No.4 2021 pp.35-49
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In a situation where there are multiple diverse datasets, it is essential to have an efficient method to provide users with the datasets they require. To address this suggestion, necessary datasets should be selected on the basis of the relationships between the datasets. In particular, in order to discover the necessary datasets for disaster resolution, we need to consider the disaster resolution stage. In this paper, in order to provide the necessary datasets for each stage of disaster resolution, we constructed a disaster type and disaster management process ontology and designed a method to determine the necessary datasets for each disaster type and disaster management process step. In addition, we introduce a method to determine relationships between datasets necessary for disaster response. We propose a method for discovering datasets based on minimal relationships such as "isA," "sameAs," and "subclassOf." To discover suitable datasets, we designed a knowledge exploration model and collected 651 disaster-related datasets for improving our method. These datasets were categorized by disaster type from the perspective of disaster management. Categorizing actual datasets into disaster types and disaster management types allows a single dataset to be classified as multiple types in both categories. We built a knowledge exploration model on the basis of disaster examples to ensure the configuration of our model.
Construction of a Video Dataset for Face Tracking Benchmarking Using a Ground Truth Generation Tool
[Kisti 연계] 한국콘텐츠학회 International journal of contents Vol.10 No.1 2014 pp.1-11
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In the current generation of smart mobile devices, object tracking is one of the most important research topics for computer vision. Because human face tracking can be widely used for many applications, collecting a dataset of face videos is necessary for evaluating the performance of a tracker and for comparing different approaches. Unfortunately, the well-known benchmark datasets of face videos are not sufficiently diverse. As a result, it is difficult to compare the accuracy between different tracking algorithms in various conditions, namely illumination, background complexity, and subject movement. In this paper, we propose a new dataset that includes 91 face video clips that were recorded in different conditions. We also provide a semi-automatic ground-truth generation tool that can easily be used to evaluate the performance of face tracking systems. This tool helps to maintain the consistency of the definitions for the ground-truth in each frame. The resulting video data set is used to evaluate well-known approaches and test their efficiency.
Construction of a Spatio-Temporal Dataset for Deep Learning-Based Precipitation Nowcasting
[Kisti 연계] 한국과학기술정보연구원 Journal of information science theory and practice : JISTaP Vol.10 No.no.spc 2022 pp.135-142
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Recently, with the development of data processing technology and the increase of computational power, methods to solving social problems using Artificial Intelligence (AI) are in the spotlight, and AI technologies are replacing and supplementing existing traditional methods in various fields. Meanwhile in Korea, heavy rain is one of the representative factors of natural disasters that cause enormous economic damage and casualties every year. Accurate prediction of heavy rainfall over the Korean peninsula is very difficult due to its geographical features, located between the Eurasian continent and the Pacific Ocean at mid-latitude, and the influence of the summer monsoon. In order to deal with such problems, the Korea Meteorological Administration operates various state-of-the-art observation equipment and a newly developed global atmospheric model system. Nevertheless, for precipitation nowcasting, the use of a separate system based on the extrapolation method is required due to the intrinsic characteristics associated with the operation of numerical weather prediction models. The predictability of existing precipitation nowcasting is reliable in the early stage of forecasting but decreases sharply as forecast lead time increases. At this point, AI technologies to deal with spatio-temporal features of data are expected to greatly contribute to overcoming the limitations of existing precipitation nowcasting systems. Thus, in this project the dataset required to develop, train, and verify deep learning-based precipitation nowcasting models has been constructed in a regularized form. The dataset not only provides various variables obtained from multiple sources, but also coincides with each other in spatio-temporal specifications.
Applying Token Tagging to Augment Dataset for Automatic Program Repair
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.18 No.5 2022 pp.628-636
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Automatic program repair (APR) techniques focus on automatically repairing bugs in programs and providing correct patches for developers, which have been investigated for decades. However, most studies have limitations in repairing complex bugs. To overcome these limitations, we developed an approach that augments datasets by utilizing token tagging and applying machine learning techniques for APR. First, to alleviate the data insufficiency problem, we augmented datasets by extracting all the methods (buggy and non-buggy methods) in the program source code and conducting token tagging on non-buggy methods. Second, we fed the preprocessed code into the model as an input for training. Finally, we evaluated the performance of the proposed approach by comparing it with the baselines. The results show that the proposed approach is efficient for augmenting datasets using token tagging and is promising for APR.
Human Detection using Real-virtual Augmented Dataset
[Kisti 연계] 한국정보통신학회 Journal of information and communication convergence engineering Vol.21 No.1 2023 pp.98-102
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper presents a study on how augmenting semi-synthetic image data improves the performance of human detection algorithms. In the field of object detection, securing a high-quality data set plays the most important role in training deep learning algorithms. Recently, the acquisition of real image data has become time consuming and expensive; therefore, research using synthesized data has been conducted. Synthetic data haves the advantage of being able to generate a vast amount of data and accurately label it. However, the utility of synthetic data in human detection has not yet been demonstrated. Therefore, we use You Only Look Once (YOLO), the object detection algorithm most commonly used, to experimentally analyze the effect of synthetic data augmentation on human detection performance. As a result of training YOLO using the Penn-Fudan dataset, it was shown that the YOLO network model trained on a dataset augmented with synthetic data provided high-performance results in terms of the Precision-Recall Curve and F1-Confidence Curve.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.