년 - 년
기계학습을 통한 문제 해결을 위한 단계에서 가장 중요한 부분 중의 하나가 데이터셋을 구축하는 것이다. 기존의 데 이터셋 구축 방법은 트레이닝 과정에서 소량의 이미지에 변화를 주어 데이터를 부풀리는 명령어를 사용하는 것과 ‘google 이미지 검색’에서 이미지 크롤링을 사용하는 것이 대부분이다. 하지만 이러한 방법들은 데이터로서 이미지 의 신뢰성이 보장되어 있지 않는 경우가 자주 발생한다. 기계학습 중 딥러닝 트레이닝을 위해서는 신뢰성이 보장된 데이터셋이 구축되어야 한다. 이에 따라 본 연구에서는 신뢰성을 보장받을 수 있는 데이터셋을 구축하기 위한 도구 를 설계 및 구현한다. 적은 노력으로 큰 데이터셋을 구축할 수 있는 방안을 고려하고, 그 방안을 통해 구축된 데이터 셋의 신뢰성을 테스트한다. 신뢰성을 보증하기 위한 예시로 꽃 이미지 데이터셋을 구축하여 ‘이미지 검색 꽃 사전’에 적용한다. 적용 결과로, 사용자는 딥러닝 트레이닝 과정에서 필요한 데이터셋을 보다 쉽고 경제적으로 구축할 수 있 고, 구축된 데이터셋을 통한 트레이닝에서의 신뢰성과 효율성이 있음을 보인다. 또한 본 연구의 아이디어를 기본으 로 활용한다면, 데이터셋을 구축하고자 하는 대상의 특성이 변하더라도 대량의 이미지 데이터셋을 생산할 수 있을 것이다.
One of the most important steps for solving problems through machine learning is to build a data set. Most previous methods of the data set generation are using commands to inflate data by changing small amounts of images during training, and image crawling in Google Image Search. However, these methods often do not guarantee the reliability of the image as data. For deeplearning training among methods of machine learning, a reliable data set must be established. Accordingly, in this study, a novel tool is designed and implemented to build a data set that can ensure reliability. We propose ways to build large data sets with little effort, and test the reliability of data sets built through them. As an example to ensure reliability, a set of flower image data is established and applied to ‘image search flower dictionary’. As a result, users can more easily and cost-effectively build the data sets needed in the deep-learning training process, and ensure reliability and efficiency in training through the generated data set. In addition, the underlying use of the ideas proposed in this study will enable the production of large sets of image data, even if the characteristics of the target to build the data set will be changed.
Examination Of the Factors That Affect Students Academic Success by Using ANFIS Method
한국AI디지털융합학회(구 한국디지털융합학회) IJICTDC Vol 7 No 2 2022.12 pp.43-51
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
The scientific and technological achievements trigger people’s desire to learn new things and sometimes it turns into a competition among individuals. The competitive environmental so makes an important psychological pressure on individuals and it may manipulate their choices about their educational careers. Artificial neural networks are widely utilized to reach to a short-cut solution and as a method of decision making. The role of the machine learning techniques and data mining algorithms to define the factors affecting students’ success is important. The aim of this research is to make predictions about these factors by using Adaptive-Network Based Fuzzy Inference Systems (ANFIS). Student Performance data set in UCI platform is used for classification part of this study.
Airline Data Set Analysis using Big Data in Cloud Computing
한국경영정보학회 한국경영정보학회 정기 학술대회 지능정보화 시대의 ICT 전략 2017.06 pp.180-184
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
In this paper, 10 years of airline data set in USA is collected. And, the analysis of the airline data set is performed using cloud computing service of Microsoft Azure which runs Hadoop cluster in the cloud. Hadoop Hive QL statements have been used for analyzing the data. Data visualization has been achieved by extracting the output of the Hive code. Using big data technologies like Hadoop and Hive in the cloud, it is easy to analyze the massive data set in SQL like language, export the output results and visualize them excel spreadsheets. Interesting sets of trends and patterns exists in this large data sets, the paper presents the insights that resides between flight diversions and flight distance, flight cancellation and flight distance for the time period.
백화점 고객의 점포만족과 충성도에 관한 연구 - 양자적 자료(Dyadic Data Set)를 이용한 내ㆍ외부 마케팅 통합관점 - KCI 등재후보
한국전략마케팅학회 마케팅논집 제15집 제2권 통권 34호 2007.06 pp.27-64
※ 기관로그인 시 무료 이용이 가능합니다.
8,200원
본 연구는 내외부 마케팅의 통합관점에서 백화점 점포만족과 충성도에 대한 선행요인을 밝히고자 하였다. 실증분석은 판매원과 고객으로부터 수집한 양자적 자료(dyadic data set)를 기반으로 하여 이루어졌다. 내부마케팅 관련변수에는 직무만족, 이에 영향을 미칠 것으로 예상된 내재적 보상과 외재적 보상이 포함되었고, 외부마케팅 관련변수에는 고객 보상프로그램의 가치(경제성, 편의성, 오락성)와 불평관리(접근성, 반응성)가 포함되었다.연구결과, 외부마케팅 관련변수 중에서 보상프로그램의 오락성과 불평에 대한 반응성이, 내부마케팅 요인인 직무만족이 점포만족에 유의한 영향을 미치는 것으로 나타났으며 점포만족은 점포충성도에 긍정적인 영향을 미치는 것으로 나타났다. 그리고 기업이 판매원에게 제공하는 내재적 보상과 외재적 보상 모두 판매원의 직무만족에 유의한 영향을 미치는 것으로 나타났다.
This study investigates antecedents of store satisfaction and store loyalty in department store in the perspectives of internal and external marketing. The analysis is based on a dyadic data set that involves judgements provided by salespeople and customers. Internal marketing related variables consist of salespersons' job satisfaction, intrinsic rewards, extrinsic rewards. External marketing related variables consist of values of customer reward program(utilitarian benefits, convenience, hedonic benefits), complaint handling(accessibility, responsiveness).Results indicate that department store satisfaction is affected significantly by hedonic benefits of rewards program, responsiveness to complaining, salespersons' job satisfaction and store loyalty is influenced by store satisfaction. Futhermore, it is found that intrinsic rewards and extrinsic rewards offered to salespeople influence significantly salespersons' job satisfaction.
Trading Responses to Analyst Reports by Investor Types : A Study of Korea’s Unique Data Set
한국재무학회 한국재무학회 학술대회 2012년 KFA&TFA Joint Conference in Finance 2012.09 pp.149-194
※ 기관로그인 시 무료 이용이 가능합니다.
9,400원
Using trading data from the Korean equity market, we investigate whether there is any difference amongst individual investors, domestic institutional investors, and foreign institutional investors in trading responses to analyst reports. We also examine the determinants of the trading responses of each investor type to analyst reports, where we use firm characteristics as well as analyst characteristics as the determinants of the release date trading volumes of each investor type. Our findings are as follows. First, individual investors are the most responsive investor group to the analyst reports, suggesting that institutional investors (domestic and foreign) perceive analyst reports to be less informative than do the individual investors. In particular, individual investors are more responsive to analyst reports on small firms, firms neglected by analysts, and firms with large inside ownership as well as analyst reports with optimistic forecasts. Second, domestic institutional investors are more responsive to reports on neglected firms and firms with high volatility, which are characterized by large information asymmetry. Furthermore, the increase in trading volume by domestic institutional investors occurs prior to the release date of analyst reports, suggesting that domestic institutional investors are informed traders. Third, foreign institutional investors do not show meaningful responses to research reports by local analysts, suggesting that foreign institutional investors do not find the information provided by local analysts as being informative.
[NRF 연계] 한국기록학회 기록학연구 Vol.74 2022.10 pp.187-222
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구의 목적은 사립대학의 행정정보 데이터세트 운영 현황을 분석하고 개선방안을 제시하는 것이다. 이를 위해 사립대학 178개교의 총 820개 시스템에 대해 시스템의 기능, 개발 유형, 데이터의 생성·정정·삭제 시기 등의 정량적인 분석을 실시하였다. 분석 결과, 통상 1개 이상의 행정정보시스템을 보유하고 학사관리시스템을 공통적으로 사용하고 있으며 대학의 인프라를 통해 시스템을 자체적으로 개발한 사례가 많고 데이터를 수시로 생성하고 정정하며 업무담당자에 의해 데이터가 삭제되고 있으나 데이터의 삭제나 정정에 대한 규정은 명확하지 않다는 문제점이 도출되었다. 이러한 문제점들을 해결하기 위한 개선 방안으로 범정부 EA포털을 현행화하여 사립대학의 행정정보시스템에 대한 보유 현황을 제대로 파악하고, 데이터의 정정이 이루어지지 않는 시스템을 중심으로 기록 관리하며 데이터의 임의 삭제가 이루어지지 않도록 내부 규정의 개정과 교육을 실시할 것 등을 제안하였다.
The aim of this study was to analyze the operation status of administrative information datasets of private universities and present improvement plans. For the system of private universities, the generation, correction, and deletion of functions, development types, and data were analyzed politically. As a result of the analysis, it has one or more administrative information systems and uses the academic management system in common, and the system is often developed on its own through the university’s infrastructure, and data is deleted by the person in charge, but the regulations are not clear. To solve these problems, it was proposed to revise the EA portal to properly investigate the current status of the administrative information system of private universities, manage records centering on systems without data correction, and revise internal regulations to conduct education.
자료집합에서 임의의 자료가 추가 또는 삭제된 경우 변경된 고유벡터 구하는 기법 KCI 등재
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.15 No.1 2019.02 pp.7-14
일반적으로 공분산행렬(covariance matrix)을 이용하여 고유벡터(eigenvector)를 계산한다. 임의의 자료가 추가 되거나, 삭제되면, 전체 자료집합이 변경되고, 공분산행렬의 내용이 바뀐다. 따라서 고유벡터를 다시 계산해야 한 다. 이때 다시 계산하는 과정이 없이, 변경된 고유벡터를 구하는 새로운 방법을 제안한다. 제안한 방법은 기존의 자 료집합에서 얻어진 고유벡터와 첨삭할 자료를 이용하여 변경된 고유벡터를 구하는 방법이다. 제안된 방법은 계산량 을 줄일 수 있다는 장점이 있다.
The eigenvectors are calculated using a covariance matrix. If some data is added or removed, the data set changes, and the covariance matrix is also changed. Therefore, the eigenvectors must be recalculated totally. This paper propose a new method to find the modified eigenvectors without recalculation. To obtain modified eigenvectors, the proposed method use modifying data and already known eigenvectors. The proposed method has the advantage of reducing the computational cost.
병해충 이미지의 데이터 셋과 증강에 대한 최근 동향 조사
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 2022 한국차세대컴퓨팅학회 춘계학술대회 2022.05 pp.405-409
병해충으로 인한 피해가 날이 갈수록 커지고 있다. 특히 고위험 문제 병해충의 피해를 최소화하기 위해서는 병해충의 빠른 진단은 필수이다. 최근 인공지능 기술의 발전으로 많은 병해충 진단 연구에서는 딥러닝 기술을 사용하고 있다. 하지만 병해충은 계절성을 띄고 있고 정확한 병해충 이미지를 수집하기 위해서는 전문가의 도움이 필요하기 때문에 정확한 병해충 이미지 수집에 한계가 있다. 부족한 병해충 이미지를 늘리기 위해 본 논문에서는 현재 제공되고 있는 병해충 이미지 데이터 셋과 최신의 데이터 증강(Data Augmentation) 연구를 조사했다.
그림 기억에 관한 일 연구 : 260개 그림세트의 표준화
[NRF 연계] 한국심리학회 한국심리학회지: 일반 Vol.7 No.2 1988.12 pp.158-186
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 그림자료 처리와 단어자료 처리의 차이점과 유사성을 연구하는 실험에 사용할 수 있는 260개의 그림을 표준화하였다. 표준화한 그림 자료는 Snodgrass 등(1980)이 사용한 단순한 흑백선으로 윤곽만이 그려진 그림을 복사하여 사용하였다. 그림들은 4가지 차원-이름 일치성, 심상 일치성, 친숙성, 시각적 복잡성-에서 표준화하였다.
In this study we present a standardized set of 260 pictures for use in experiments investigating differences and similarities in the processing of pictures and words. The pictures were black-and-white line drawings which were made by Snodgrass et al.. We analyzed and standardized the pictures on four variables of central relevance to memory and cognitive processing : name agreement, image agreement, familiarity and visual complexity.
불균형 데이터 집합에서의 의사결정나무 추론: 종합 병원의 건강 보험료 청구 심사 사례 KCI 등재
한국경영정보학회 경영정보학연구 제9권 제1호 2007.04 pp.45-65
※ 기관로그인 시 무료 이용이 가능합니다.
5,700원
다른 산업과 달리 병원/의료 산업에서는 건강 보험료 심사 평가라는 독특한 검증 과정이 필수적으로 있게 된다. 건강 보험료 심사 평가는 병원의 수익 문제 뿐 아니라 적정한 진료행위를 하는 병원이라는 이미지와도 맞물려 매우 중요한 분야이며, 특히 대형 종합병원일수록 이 부분에 많은 심사관련 인력들을 투입하여, 병원의 수익과 명예를 위해서 업무를 수행하고 있다. 본 논문은 이러한 건강보험료 청구 심사 과정에서, 사전에 수많은 진료 청구 건 중 심사 평가에서 삭감이 될 수 있는 진료 청구 건을 데이터 마이닝을 통해서 발견하여, 사전의 대비를 철저히 하고자 하는 한 국내 대형 종합병원의 사례를 소개하고자 한다. 데이터 마이닝을 적용함에 있어, 주요한 문제점 중 하나는 바로 지도학습 기법을 적용하기에 곤란한 데이터 불균형 문제가 발생하는 것이다. 이런 불균형 문제를 해소하고, 비교 조건 중에 가장 효율적인 삭감 예상 진료 건 탐지 모델을 만들어 내기 위하여, 데이터 불균형 문제의 기본 해법인 Sampling과 오분류 비용의 다양한 혼합적인 적용을 통하여, 적합한 조건을 가지는 의사결정 나무 모델을 도출하였다.
In medical industry, health insurance bill audit is unique and essential process in general hospitals. The health insurance bill audit process is very important because not only for hospital's profit but also hospital's reputation. Particularly, at the large general hospitals many related workers including analysts, nurses, and etc. have engaged in the health insurance bill audit process. This paper introduces a case of health insurance bill audit for finding reducible health insurance bill cases using decision tree induction techniques at a large general hospital in Korea. When supervised learning methods had been tried to be applied, one of major problems was data imbalance problem in the health insurance bill audit data. In other words, there were many normal(passing) cases and relatively small number of reduction cases in a bill audit dataset. To resolve the problem, in this study, well-known methods for imbalanced data sets including over sampling of rare cases, under sampling of major cases, and adjusting the misclassification cost are combined in several ways to find appropriate decision trees that satisfy required conditions in health insurance bill audit situation.
본 연구는 사용자로 하여금 내용 조정 없는 원본 데이터집합을 실시간 접근 가능하게 하면서 프라이버시를 보존하는 방안을 제안한다. 프라이버시 보존을 위한 지금까지의 방법들은 개체노출 방지를 위한 k-익명화와 민감속성노출 방 지를 위한 l-다양화를 적용함에 따른 데이터 사용성 저하가 불가피하였다. 본 연구에서는 개체노출을 허용하는 대신 l-다양화를 통한 민감속성노출 방지만으로 프라이버시를 보존하는 방안을 제안한다. 사용자에게 제공하는 민감속성 값 집합은 민감속성 온톨로지로부터 의미적으로 근접한 내용들을 선택하여 단일 혹은 군집개체에 연결할 내용으로 구성함으로써 민감속성값 집합의 정보손실을 감소시겼다. 실험을 통하여 동일 다양성 기준으로 60% 내외의 민감속 성값 집합 정보손실 감소를 확인하였다. 사용자가 접근의도에 부합하는 내용 획득 정도를 나타내는 정밀도를 기준으 로 k-익명화 수준별로 데이터 사용성을 파악함으로써 개체노출 허용에 따른 데이터 사용성 저하 방지의 유용함을 고찰하였다.
This study proposes an approach to enable users to access the unadjusted original dataset in real time while preserving privacy. While previous methods for privacy preservation have employed k-anonymity for entity disclosure prevention and l-diversity for attribute disclosure prevention, this study suggests an approach that preserves privacy solely through attribute disclosure prevention via l-diversity, excluding entity disclosure prevention. By not implementing k-anonymity, the adjustment of quasi-identifiers in published datasets became unnecessary, thereby preventing a decrease in data usability. By accessing an ontology of sensitive attributes and selecting semantically related values, a set of sensitive attribute values connected to a single entity or an entity cluster was constructed, thus reducing information loss within the provided set of sensitive attribute values for users.
전기 자동차 긴급 충전서비스 데이터셋 표준 개발을 위한 기초 연구
한국ITS학회 한국ITS학회 학술대회 ITS와 함께하는 미래 스마트 시티 2022.06 pp.1043-1047
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
트래픽 유통계획 기반 사이버전 훈련데이터셋 생성방법 설계 및 구현 KCI 등재
한국융합보안학회 융합보안논문지 제20권 제4호 2020.10 pp.71-80
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
사이버전 훈련 시스템에 현실감 있는 트래픽을 제공하기 위해서는 사전에 트래픽 유통계획 작성과 정상/위협 데이터 셋을 이용한 훈련데이터셋 생성이 필요하다. 본 논문은 사이버전 훈련 시스템에 실제 환경과 같은 배경 트래픽을 제공 하기 위한 트래픽 유통계획 저작과 훈련데이터셋을 생성하는 방법의 설계와 구현 결과를 제시한다. 트래픽 유통계획은 트래픽을 유통할 훈련 환경의 네트워크 토폴로지와 실제 및 모의환경에서 수집한 트래픽 속성 정보를 이용하여 저작하 는 방법을 제안한다. 트래픽 유통계획에 따라 훈련데이터셋을 생성하는 방법은 단위트래픽을 이용하는 방법과 프로토콜 의 비율을 이용하는 혼합트래픽 양상 방법을 제안한다. 구현한 도구를 이용하여 트래픽 유통계획을 저작하고, 유통계획 에 따른 훈련데이터셋 생성결과를 확인하였다.
In order to provide realistic traffic to the cyber warfare training system, it is necessary to prepare a traffic distribution plan in advance and to create a training data set using normal/threat data sets. This paper presents the design and implementation results of a method for creating a traffic distribution plan and a training data set to provide background traffic like a real environment to a cyber warfare training system. We propose a method of a traffic distribution plan by using the network topology of the training environment to distribute traffic and the traffic attribute information collected in real and simulated environments. We propose a method of generating a training data set according to a traffic distribution plan using a unit traffic and a mixed traffic method using the ratio of the protocol. Using the implemented tool, a traffic distribution plan was created, and the training data set creation result according to the distribution plan was confirmed.
기계학습 활용을 위한 학습 데이터세트 구축 표준화 방안에 관한 연구 KCI 등재
한국디지털정책학회 디지털융복합연구 제16권 제10호 2018.10 pp.205-212
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
고성능 CPU/GPU의 개발과 심층신경망 등의 인공지능 알고리즘, 그리고 다량의 데이터 확보를 통해 기계학습이 다양한 응용 분야로 확대 적용되고 있다. 특히, 사물인터넷, 사회관계망서비스, 웹페이지, 공공데이터로부터 수집된 다량의 데이터들이 기계학습의 활용에 가속화를 가하고 있다. 기계학습을 위한 학습 데이터세트는 응용 분야와 데이터 종류에 따라 다양한 형식으로 존재하고 있어 효과적으로 데이터를 처리하고 기계학습에 적용하기에 어려움이 따른다. 이에 본 논문은 표준화된 절차에 따라 기계학습을 위한 학습 데이터세트를 구축하기 위한 방안을 연구하였다. 먼저 학습 데이터세트가 갖추 어야할 요구사항을 문제 유형과 데이터 유형별로 분석하였다. 이를 토대로 기계학습 활용을 위한 학습 데이터세트 구축에 관한 참조모델을 제안하였다. 또한 학습 데이터세트 구축 참조모델을 국제 표준으로 개발하기 위해 대상 표준화 기구의 선 정 및 표준화 전략을 제시하였다.
With the development of high performance CPU / GPU, artificial intelligence algorithms such as deep neural networks, and a large amount of data, machine learning has been extended to various applications. In particular, a large amount of data collected from the Internet of Things, social network services, web pages, and public data is accelerating the use of machine learning. Learning data sets for machine learning exist in various formats according to application fields and data types, and thus it is difficult to effectively process data and apply them to machine learning. Therefore, this paper studied a method for building a learning data set for machine learning in accordance with standardized procedures. This paper first analyzes the requirement of learning data set according to problem types and data types. Based on the analysis, this paper presents the reference model to build learning data set for machine learning applications. This paper presents the target standardization organization and a standard development strategy for building learning data set.
연도별 데이터 조합 기계학습을 통한 안드로이드 악성 앱 탐지의 지속 가능성 향상 연구 KCI 등재
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.18 No.5 2022.10 pp.47-57
기계학습을 이용한 안드로이드 악성 앱 탐지 연구가 활발하지만 미래의 악성 앱을 기존의 모델로 예측할 수 있는 지 속 가능성 문제에 대한 연구가 상대적으로 부족하다. 이전 연구는 기존 악성 앱 탐지 연구가 특정기간의 데이터만 사용하였고, 그 결과 지속 가능성이 부족함을 보였다. 본 논문은 그 해결책으로 연도별 데이터를 조합하여 지속 가 능성을 향상하는 실험을 진행하였다. 2014년부터 2020년까지의 연도별 데이터에서 2개 연도 또는 3개 연도 데이 터를 조합하여 기계학습을 실시하고, 모든 연도의 데이터로 모델을 평가하였다. 실험결과 데이터의 분포가 크게 달 라지는 3개 연도 데이터 조합이 가장 좋은 성능을 보였다.
Although research on Android malicious app detection using machine learning is active, research on sustainability issues that can predict future malicious apps with existing models is relatively scarce. Previous studies showed that existing malicious app detection studies only used data for a specific period, and as a result, lack of sustainability. To find a solution, we conducted an experiment to improve sustainability by combining yearly data. We trained machine learning models by combining data from two or three years from year-by-year data from 2014 to 2020, and the model was evaluated with data from all years. As a result of the experiment, the three-year data combination in which the distribution of the data was significantly different showed the best performance.
요양병원 유형에 따른 65세 이상 노인의 입원 의료이용 현황과 특성 : 2018년 건강보험심사평가원 고령환자데이터셋(HIRA-APS)을 사용하여 KCI 등재
보건의료산업학회 보건의료산업학회지 제16권 제1호 2022.03 pp.1-13
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
Objectives: This study aimed to provide basic data for policies related to the classification of long-term care hospitals admitting the elderly, according to receipt or otherwise of specialized rehabilitation treatment, and to analyze the status and characteristics of medical use. Methods: The Health Insurance Review and Assessment Service 2018 Elderly Patient Data was used to analyze 33,821 inpatients who had applied for the daily flat-rate system for long-term care hospitals. Results: The ratio of the paralysis group, cerebrovascular disease group, and musculo-skeletal system group was higher for the rehabilitative long-term care hospital, and the ratios of medical high, medical middle, physical function low, fee-for-service, and medical low groups were also higher than those of general long-term care hospitals. The long-term hospitalization rates were low for patients admitted to rehabilitative long-term care hospitals. Conclusions: It is necessary to establish policies so that patients in need of specialized rehabilitation treatment are not admitted to general long-term care hospitals, which delays their recovery.
IPTV 셋톱박스 로그 데이터 기반 TV 시청률 특성 분석 : IPTV 3사 셋톱박스 전수 데이터 활용을 중심으로 KCI 등재
한국광고학회 광고학연구 제34권 3호 2023.06 pp.67-92
※ 기관로그인 시 무료 이용이 가능합니다.
6,400원
본 연구는 피플미터 기반의 시청률만이 존재하는 현 상황에서 IPTV 셋탑박스 로그데이터를 기반으로 한 새로운 시청률을 실험적으로 제안하기 위해 진행되었다. 구체적으로 본 연 구에서는 IPTV 3사의 셋탑박스 로그데이터를 표본없이 모두 취합하여 시청기록 전수 데이터를 구축하여 시청률을 산출하고 이를 피플미터 기반 방식의 전통적인 시청률 지표인 닐슨 시청률 과 비교하여 검증하였다. 연구결과, IPTV 셋탑박스 로그데이터를 기반으로 한 시청률은 닐슨 시 청률에 비해 평균 약 50~60% 규모로 집계되었으며, 닐슨 시청률과의 상관계수는 .95이상 수 준으로 매우 높게 나타났다. 본 연구에서는 이 같은 결과를 바탕으로 IPTV 셋탑박스 로그데이 터 기반의 시청률지표의 특성을 살펴보고, 산업 내 활용될 수 있는 시청률 지표로서의 가능성을 제안하였다.
The purpose of this study is to propose a new viewership rating based on log data of all IPTV set-top boxes in Korea where only people meter-based viewership ratings exists. Specifically, this study collected all of the set-top box log data from IPTV companies without sampling. This study tried to construct viewership records from the collected data, calculated viewership ratings and also verified the results by comparing them to traditional Nielsen viewership ratings based on people meters. The results showed that viewership ratings, based on IPTV set-top box log data, averaged about 50-60% of Nielsen ratings, with a high correlation coefficient of 0.95 or more. Based on these results, this study examined the characteristics of viewership ratings estimated from the IPTV set-top box log data and proposed their potential as usable viewership ratings in the industry as well.
교통사고 복합 상관성 분석을 위한 데이터 전처리 및 셋 구축 연구
한국ITS학회 한국ITS학회 학술대회 ITS와 함께하는 미래 스마트 시티 2022.06 pp.762-767
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
범주형 데이터의 러프집합 분석을 통한 의사결정 향상기법 KCI 등재
한국디지털정책학회 디지털융복합연구 제13권 제6호 2015.06 pp.157-164
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최적의 의사결정시스템에서 효율적인 정보검색을 위해서는 정보의 감축이 필수적이다. 다양한 종류의 데이 터에 존재하는 유용한 정보를 찾는 데이터 마이닝 기법에 대한 많은 연구가 활발하게 진행되어 왔고 타 산업과의 융복합을 위한 빅데이터 활용이 높아져 가고 있다. 유용한 지식의 발견을 위한 여러 가지 기법들이 특징을 가지고 있지만 단점이 존재하기 마련이다. 따라서 그러한 특징을 수렴하는 하나의 새로운 기법이 필요하다. 본 논문에서는 베이지언 정리를 이용하여 정보의 대수학적인 확률을 측정하고 이 확률에 대하여 정보엔트로피를 계산함으로써 정보 의 불확실성을 계산한다. 제안된 척도를 기반으로 불필요한 속성을 제거하여 최소의 리덕트를 생성하고 이를 기반으 로 결정규칙을 유도하는 알고리즘을 제안한다. 제안된 방법의 효율성을 위하여 콘텍트 렌즈를 결정하는 실험을 통하 여 기존방법과 비교 결과, 본 연구가 의사결정의 유용성면에 있어 일반성이 있음을 보인다.
An efficient retrieval of useful information is a prerequisite of an optimal decision making system. Hence, A research of data mining techniques finding useful patterns from the various forms of data has been progressed with the increase of the application of Big Data for convergence and integration with other industries. Each technique is more likely to have its drawback so that the generalization of retrieving useful information is weak. Another integrated technique is essential for retrieving useful information. In this paper, a uncertainty measure of information is calculated such that algebraic probability is measured by Bayesian theory and then information entropy of the probability is measured. The proposed measure generates the effective reduct set (i.e., reduced set of necessary attributes) and formulating the core of the attribute set. Hence, the optimal decision rules are induced. Through simulation deciding contact lenses, the proposed approach is compared with the equivalence and value-reduct theories. As the result, the proposed is more general than the previous theories in useful decision-making.
인공지능 학습용 데이터 품질에 대한 연구 : 퍼지셋 질적비교분석 KCI 등재
한국경영정보학회 경영정보학연구 제26권 제1호 2024.02 pp.19-56
※ 기관로그인 시 무료 이용이 가능합니다.
8,200원
본 연구는 한국의 인공지능 학습용 데이터 구축 사업과 데이터의 공공 개방에 관한 정책 수행 기관, 데이터 구축 기업, 그리고 이를 활용하는 다양한 기관의 데이터 품질에 대해 이해를 제고하고, 신뢰할 수 있는 인공지능 알고리즘 개발에 있어 가장 중요한 학습용 데이터 품질에 대한 이론적 토대를 만들기 위한 실증적 연구이다. 이를 위해, 데이터의 속성 요인, 데이터 구축환경 요인, 데이터 타입 관련 요인 등 인공지능 학습용 데이터 품질과 관련된 중요 선행요인을 도입하여 이론적 모형을 제안한다. 본 연구는 393명의 인공지능 학습용 데이터 구축 기업과 인공지능 서비스 개발 기업의 실무 담당자를 대상으로 설문조사를 실시하여 데이터를 수집하였다. 데이터 분석은 퍼지셋 질적비교분석 방법과 인공신경망 분석을 통해 이루어졌으며, 분석 결과를 통해 인공지능 학습용 데이터 관련 학술적 및 실무적 시사점을 도출했다.
This study is empirical research to enhance understanding of AI (artificial intelligence) training data project in South Korea. It primarily focuses on the various concerns regarding data quality from policy-executing institutions, data construction companies, and organizations utilizing AI training data to develop the most reliable algorithm for society. For academic contribution, this study suggests a theoretical foundation and research model for understanding AI training data quality and its antecedents, as well as the unique data and ethical aspects of AI. For this purpose, this study proposes a research model with important antecedents related to AI training data quality, such as data attribute factors, data building environmental factors, and data type-related factors. The study collects 393 sample data from actual practitioners and personnel from companies building artificial intelligence training data and companies developing artificial intelligence services. Data analysis was conducted through Fuzzy Set Qualitative Comparative Analysis (fsQCA) and Artificial Neural Network analysis (ANN), presenting academic and practical implications related to the quality of AI training data.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.