년 - 년
헬스케어 정보 수집을 위한 딥 러닝 기반의 서브넷 구축 기법 KCI 등재
한국디지털정책학회 디지털융복합연구 제15권 제3호 2017.03 pp.221-228
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 IoT 기술이 발달함에 따라 헬스케어 서비스를 제공하는 많은 의료기관에서는 IoT 기술을 이용한 의료 서비스가 점점 증가하고 있는 추세이다. 그러나, 사용자 신체에 부착된 IoT 센서의 갯수가 증가하면서 서버에 전달되는 헬스케어 정보는 점점 복잡해져서 서버에서 사용자의 헬스케어 정보를 분석하는 시간이 증가하는 문제점이 발생하고 있다. 본 논문에서는 사용자에 부착된 IoT 장치를 통해 전달되는 다량의 헬스케어 정보를 서버에서 의료 목적에 맞게 헬스케어 정보를 수집 및 처리하기 위한 딥 러닝 기반의 헬스케어 정보 관리 기법을 제안한다. 제안 기법은 서버에 전달된 헬스케어 정보에 속성값을 부여하여 속성값에 따른 서브넷을 구축한 후 서브넷간 연계 정보를 시드로 추출하여 계층적 구조로 그룹화 한다. 제안 기법은 서버에서 그룹화된 헬스케어 정보를 딥 러닝에 적용하여 사용자의 치료 및 처방에 대한 관찰 속도 및 정확도를 향상시킬 수 있어 최적화된 정보를 추출할 수 있는 특징이 있다. 성능평가 결과, 제안기법은 헬스케어 서비스 모델에서 동작되는 의료 서비스의 처리속도가 기존기법에 비해 평균 14.1% 향상되었고, 서버의 오버헤드는 기존 기법보다 평균 6.7% 낮았다. 헬스케어 정보 추출에 대한 정확도는 기존기법보다 평균 10.1% 높은 결과를 얻었다.
With the recent development of IoT technology, medical services using IoT technology are increasing in many medical institutions providing health care services. However, as the number of IoT sensors attached to the user body increases, the healthcare information transmitted to the server becomes complicated, thereby increasing the time required for analyzing the user's healthcare information in the server. In this paper, we propose a deep learning based health care information management method to collect and process healthcare information in a server for a large amount of healthcare information delivered through a user - attached IoT device. The proposed scheme constructs a subnet according to the attribute value by assigning an attribute value to the healthcare information transmitted to the server, and extracts the association information between the subnets as a seed and groups them into a hierarchical structure. The server extracts optimized information that can improve the observation speed and accuracy of user's treatment and prescription by using deep running of grouped healthcare information. As a result of the performance evaluation, the proposed method shows that the processing speed of the medical service operated in the healthcare service model is improved by 14.1% on average and the server overhead is 6.7% lower than the conventional technique. The accuracy of healthcare information extraction was 10.1% higher than the conventional method.
Knowledge Extraction from Academic Journals Using Data Mining Techniques
한국디지털정책학회 디지털융복합연구 제3권 제1호 2005.06 pp.75-88
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
최근 우리는 인접학문 간 그리고 학계와 산업계간의 연구협조가 점차 증가하고 있음을 보아오고 있다. 이러한 현상은 특히 학술저널 간 지식의존성을 촉진하는 계기를 제공하고 있다고 할 수 있다. 본 논문의 목적은 관련저널 간 지식상호 의존성을 규명하고 저널지식의 구조화를 위하여 연관성 (association), 군집화, 링크분석 등 데이터마이닝 기법을 적용하는 방법론을 제시하는 것이다. 제시된 방법을 통하여 기대되는 점들은 1) 논문의 기본 속성인 키워드, 저자, 그리고 인용데이터를 통합하는 규칙 집합을 통하여 논문지식검색기능의 향상, 2) 키워드를 기반으로 관련 저널 간 그리고 저널내부의 군집분석으로 지식동향 파악, 3) Kleinberg (1999)의 권위와 허브 개념을 인용데이터 분석에 활용하여 기존의 양적 평가 기준인 영향력지수 (impact factor)의 문제점을 보완하며, 4) 특정 논문이나 저널의 지식파급과 관련한 영향력을 산출하는 잠재적 지식파급 지수를 제안하는 것이다.
선박 건조 과정에서 2D CAD 데이터 추출을 활용한 케이블 절감 최적화 설계 방안 연구 KCI 등재
한국기계항공기술학회(구 한국기계기술학회) 한국기계항공기술학회지(구 한국기계기술학회지) 제26권 제6호 2024.12 pp.1161-1166
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
In order support the design support system of small and medium-sized shipbuilding companies that carry out designs using 2D CAD, this study developed a system that automatically calculates the cable length by extracting the Y-axis value expressed as text data in 2D CAD. By setting the equipment where the cable starts and ends, the essential route and the installation rate were checked so that the optimal route of the cable could be calculated. As a result, the value calculated based on the optimal route and length of the cable by extracting the data of 2D CAD through this study was the same as the value previously calculated by the actual user, and the installation rate was less than 130% so there was no problem with the on-site installation. In addition, it was confirmed that the cable length calculated through this was reduced by about 7% compared to the existing work.
DTW 거리 기반 kNN을 활용한 시계열 데이터 정보 추출 및 회귀 예측 KCI 등재
한국경영정보학회 경영정보학연구 제26권 제2호 2024.05 pp.83-93
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
본 연구는 도금욕 공정의 완성도 예측을 위한 시계열 데이터의 효과적인 표현을 목표로, Dynamic Time Warping(DTW) 및 k-Nearest Neighbors(kNN) 기반의 전처리 방법론을 제안한다. 제안된 DTW 기반 kNN 전처리 방법을 다양한 회귀 모델에 적용하여 비교한 결과, 기존 결정 나무(Decision tree) 대비 최대 RMSE에서 43%과 MAE에서 24% 개선된 성능 향상을 보였으며, 신경망 구조를 갖는 회귀 모델과 결합했을 때 성능 향상이 두드러졌다. 본 논문에서 제안하는 전처리 방법과 회귀 모델을 결합한 구조는 길이가 긴 시계열 데이터와 제한된 데이터 샘플이 있는 상황에서 적합할 것으로 사료되며, 데이터가 부족한 상황에서도 과적합의 위험을 감소시키며, 합리적인 예측을 가능하게 함을 시사한다. 그러나 DTW 및 kNN 알고리즘은 데이터 샘플이 많아질수록 연산량이 늘어난다는 한계가 존재하며, 향후 연구를 통해 이러한 계산 효율성의 문제를 개선할 수 있는 연구가 필요할 것으로 보인다.
This study proposes a preprocessing methodology based on Dynamic Time Warping (DTW) and k-Nearest Neighbors (kNN) to effectively represent time series data for predicting the completion quality of electroplating baths. The proposed DTW-based kNN preprocessing approach was applied to various regression models and compared. The results demonstrated a performance improvement of up to 43% in maximum RMSE and 24% in MAE compared to traditional decision tree models. Notably, when integrated with neural network-based regression models, the performance improvements were pronounced. The combined structure of the proposed preprocessing method and regression models appears suitable for situations with long time series data and limited data samples, reducing the risk of overfitting and enabling reasonable predictions even with scarce data. However, as the number of data samples increases, the computational load of the DTW and kNN algorithms also increases, indicating a need for future research to improve computational efficiency.
한국코퍼스언어학회 Corpus Linguistics Research Vol. 8 No. 2 2023.12 pp.39-55
※ 기관로그인 시 무료 이용이 가능합니다.
5,100원
Recently, the amount of newly coined words generated in Korean is vast, and the frequency of use in official language media such as the media, broadcasting, and books as well as everyday spoken language is gradually increasing. As the time spent in the Internet space increases, language for communication is created in various forms or its meaning changes to convey new information or values to members of society. In this study, the SNS corpus containing the rapidly changing use of language was analyzed. After selecting new word candidates by constructing a series of pipelines for extracting noun-type new words from the SNS corpus, characteristics and usage were analyzed. At this time, in the natural language processing pipeline that extracts new words, a pipeline including all the processes of rule-based learning using Mecab, unsupervised learning using Soynlp, and user dictionary addition using a correct morpheme analyzer was constructed to extract meaningful tokens. After completing the step of selecting new word candidates, 255 new words were collected. The proportion of sentences including the new word candidate group in the SNS data was 4.799%. Among them, the proportion of sentences in which words belonging to the top 10 appeared was 12.345%. Looking at the ratio of classifying the top 30 new words according to the word formation method, the word formation method that occupied the highest ratio was compound word-synthetic abbreviations (33.3%). The type/token ratio of sentence data including new words was 0.324. The type/token ratio of SNS data was 0.254. Since the type/token ratio of SNS data is lower, it can be said that ototoxicity is higher than that of sentences containing new words. When looking at the collocation relationship and usage of new words such as the initial constant word 'ㄹㅇ', the borrowed word '-특', and the meaning-expanded word '코인', various forms and syntactic uses could be found, and there were many collocations that reflected the social image at the time of data collection. Judging from this phenomenon, the characteristics of corpus, in which initials, borrowings, meeaning-expandings, and special characters are used among newly coined words, become incomplete when simply relying on a dictionary consisting of words or word lists, so a natural language processing dataset containing more diverse social meanings can be constructed by using usage data.
학습 데이터가 없는 모델 탈취 방법에 대한 분석 KCI 등재
한국융합보안학회 융합보안논문지 제23권 제5호 2023.12 pp.57-64
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
딥뉴럴네트워크 모델의 취약점으로 모델 탈취 방법이 있다. 이 방법은 대상 모델에 대하여 여러번의 반복된 쿼 리를 통해서 유사 모델을 생성하여 대상 모델의 예측값과 동일하게 내는 유사 모델을 생성하는 것이다. 본 연구에 서, 학습 데이터가 없이 대상 모델을 탈취하는 방법에 대해서 분석을 하였다. 생성 모델을 이용하여 입력 데이터 를 생성하고 대상 모델과 유사 모델의 예측값이 서로 가까워지도록 손실함수를 정의하여 유사 모델을 생성한다. 이 방법에서 대상 모델의 입력 데이터에 대한 각 클래스의 logit(로직) 값을 이용하여 경사하강법으로 유사 모델 이 그것과 유사하도록 학습하는 과정을 갖는다. 실험 환경으로 pytorch 머신러닝 라이브러리를 이용하였으며, 데 이터셋으로 CIFAR10과 SVHN을 사용하였다. 대상 모델로 ResNet 모델을 이용하였다. 실험 결과로써, 모델 탈취 방법은 CIFAR10에 대해서 86.18%이고 SVHN에 대해서 96.02% 정확도로 대상 모델과 유사한 예측값을 내는 유 사 모델을 생성하는 것을 볼 수가 있었다. 추가적으로 모델 탈취 방법에 대한 고려사항와 한계점에 대한 고찰도 분석하였다.
In this study, we analyzed how to steal the target model without training data. Input data is generated using the generative model, and a similar model is created by defining a loss function so that the predicted values of the target model and the similar model are close to each other. At this time, the target model has a process of learning so that the similar model is similar to it by gradient descent using the logit (logic) value of each class for the input data. The tensorflow machine learning library was used as an experimental environment, and CIFAR10 and SVHN were used as datasets. A similar model was created using the ResNet model as a target model. As a result of the experiment, it was found that the model stealing method generated a similar model with an accuracy of 86.18% for CIFAR10 and 96.02% for SVHN, producing similar predicted values to the target model. In addition, considerations on the model stealing method, military use, and limitations were also analyzed.
공간 빅데이터와 범죄통계자료를 이용한 범죄취약지 추출 KCI 등재
중소기업융합학회 융합정보논문지(구 중소기업융합학회논문지) 제8권 제1호 2018.02 pp.161-171
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
범죄는 특정한 장소나 주변 환경에 따라서 범죄의 유형과 빈도가 매우 밀접한 관계를 갖으며 발생된다. 특히 공간적 으로 범죄는 도심지역, 유흥가, 노상 등에서 많이 발생된다. 이러한 이유로 범죄와 발생장소와의 관계를 분석하는 것은 범죄 를 예측하는데 효과적이며 이를 위해서 다양한 공간분석 기법이 적용되고 있다. 이에 본 논문에서는 범죄 예측에 활용코자 GIS 공간분석 기법을 이용하여 범죄취약지를 추출하였다. 범죄취약지는 범죄통계자료를 이용하여 장소와 용도지역별로 다 르게 발생되는 범죄를 GIS의 핫스팟 분석(Hot Spot Analysis)과 역거리 가중법(IDW)을 이용하여 추출하였다. 또한 셉테드 (CPTED)의 감시요소인 CCTV, 가로등, 지구대, 파출소에 대해서 각각 감시범위와 가중치를 산정하고 범죄취약지도와 중첩 하여 4개 등급(안전, 주의, 경고, 위험)으로 표현된 셉테드 기반의 범죄취약지도를 제작하였다.
This study set out to identify crime vulnerable areas with the GIS spatial analysis technique for the prediction of crimes. Crime vulnerable areas were extracted from the statistics of crimes with the GIS hotspot analysis technique and the inverse distance weighted(IDW) method applied to different crimes according to places and use districts. The scope of surveillance and weight were calculated for each of CPTED surveillance elements including CCTV, streetlamp, patrol division, and police substation. Maps of crime vulnerable areas were overlapped one after another to make a CPTED-based one expressed in four grades(safety, attention, warning, and risk).
전화망 음성 데이터를 이용한 향상된 음성 특징 추출 방법
국제차세대융합기술학회 차세대융합기술학회논문지 제1권 4호 2017.12 pp.181-185
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 음성 인식률 향상을 위한 여러 가지 방법들 중에서 음성특징 파라미터 추출 방법에 관한 한 가지 방법을 제시하였다. 본 논문에서는 청각 특성을 기반으로 한 MFCC(mel frequency cepstrum coefficients)와 성능 향상 을 위한 방법으로 GFCC (gammatone filter frequency cepstrum coefficients)를 제시하고 음성 인식을 수행하여 성능을 비교하였다. MFCC에서 일반적으로 사용하는 임계 대역 필터로 삼각 필터(triangular filter) 대신 청각 구조의 기저막 (basilar membrane)특성을 묘사한 gammatone 대역 통과 필터를 이용하여 특징 파라미터를 추출하였다. 인식 알고리즘 으로 사용하여 LPCC(linear prediction cepstrum coefficient), MFCC와 GFCC의 각 인식 성능을 비교하였다.
In this paper, we study a method extracting the speech feature parameter in many methods for improving the speech recognition rate. MFCC(Mel Frequency Cepstral Coefficient) based on the auditory characteristic and GFCC(Gammatone filter Frequency Cepstrum Coefficient) for improving performance of MFCC are suggested and telephone speech recognition also. Feature parameter is extracted with gammatone bandpass filter depicted basilar membrane characteristic of auditory structure instead of triangular filter generally used for critical bandwidth filter in MFCC. In experiments, we compared LPCC and MFCC. CMS(Cepstral Mean Subtraction) is used for compensating channel distortion in telephone and shows improvement in recognition.
GPS 프로브 차량 속도자료를 이용한 고속도로 사고 위험구간 추출기법 KCI 등재
한국ITS학회 한국ITS학회논문지 제9권 제3호 통권29호 2010.06 pp.73-84
※ 기관로그인 시 무료 이용이 가능합니다.
4,300원
본 연구에서는 고속도로에서 GPS(Global Positioning System)수신기를 장착한 프로브차량을 이용하여 수집한 속도자료를 이용하여 사고 위험구간을 추출하는 방법론을 제시하였다. 위험구간 추출을 사고발생 유ㆍ무를 판단하는 분류문제(Classification)로 정형화하고 베이지안 신경망을 적용하였다. 개별차량의 속도자료를 이용하여 다양한 잠재적 독립변수를 설정하고 이항 로지스틱 회귀분석을 이용하여 통계적으로 유의미한 변수만을 추출하여 베이지안 신경망의 입력자료로 사용하였다. 제안된 방법론의 성능 평가를 위해 사고 발생 경험이 있는 위험구간을 정확히 추출하는 분류정확도를 효과척도로 활용하였다. 본 연구에서 제안한 방법론의 타당성을 60%의 분류정확도를 통해 확인할 수 있었다. 고속도로 신설노선의 교통안전성을 평가하고 사고예방을 위한 대응책 개발 및 적용에 본 연구의 결과가 효과적으로 활용될 것으로 기대된다
This study presents a novel method to identify hazardous segments of freeway using global positioning system(GPS) based probe vehicle data. A variety of candidate contributing factors leading to higher potential of accident occurrence were extracted from the probe vehicle dataset. The research problem was defined as a classification problem, then a well-known classifier, bayesian neural network was adopted to solve the problem. A binary logistic regression technique was also used for selecting salient input variables. Test results showed that the proposed method is promising in extracting hazardous freeway sections. The outcome of this study will be effectively used for evaluating the safety of freeway sections and deriving countermeasures to prevent accidents.
4,200원
지구 온난화로 생활 환경이 급격히 변화하고 있으며, 대형 재난이 증가하고 있다. 이러한 재난 발생 시 복구에 많은 자원을 투입하고 있지만, 재난의 예방 만큼 효과적인 대책은 없을 것이다. 재난전조 정보란 하인리히 법칙에 따라 예고되는 재난에 대한 전조이며, 이에 대한 정보를 자동으로 추출하여 대비할 수 있게 하는 것이 본 논문의 초점이다. 웹에 산재된 정보로부터 전조 정보를 정확히 추출하기 위한 기반이 되는 단어(명사)를 구축하고 이를 기반으로 정확한 데이터를 추출할 수 있는 알고리즘을 연구하였다. 본 연구의 결과물로 도출된 단어는 분석적인 연구결과이기 때문에 장기적으로 실제 데이터를 적용하면서 지속적으로 보완되어야 할 것이다.
Life Environment is rapidly changing and large scale disasters are increasing from the global warming. Although the disaster repair resources are deployed to the disaster fields, the prevention of the disasters is the most effective countermeasures. the disaster sign data is based on the rule of Heinrich. Automatic extraction of the disaster sign data from the web is the focused issues in this paper. We defined the automatic extraction processes and applied information, such as accident nouns, disaster filtering nouns, disaster sign nouns and rules. Using the processes, we implemented the disaster sign data management system. In the future, the applied information must be continuously updated, because the information is only the extracted and analytic result from the some disaster data.
전통 문화 데이터를 이용한 메타 러닝 기반 전역 관계 추출 KCI 등재
한국융합학회 한국융합학회논문지 제9권 제11호 2018.11 pp.23-28
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 존재하는 대부분의 관계 추출 모델은 언급 수준의 관계 추출 모델이다. 이들은 성능은 높지만, 장문의 텍스트에 존재하는 다수의 문장을 처리할 때, 문서 내에 주요 개체 및 여러 문장에 걸쳐서 표현되는 전역적 개체 관계를 파악하지 못한다. 그리고 이러한 높은 수준의 관계를 정의하지 못하는 것은 데이터의 올바른 정형화를 막는 중대한 문제이다. 이 논문 에서는 이러한 문제를 해결하고 전역적 관계를 추출하기 위하여 외부 메모리 신경망 모델을 이용하는 새로운 방식의 전역 관계 추출 모델을 제안한다. 제안하는 모델은 1차적으로는 단편적인 관계 추출을 실행한 뒤, 외부메모리 신경망을 이용하여 단편적인 관계들을 분석 및 종합하여 텍스트 전체로부터 전역적 관계들을 추출한다. 또한 제안된 모델은 외부 메모리를 통 하여 전역적 관계 추출 외에도 주어와 목적어 생략이 잦은 한국어 관계 추출에도 뛰어난 성능을 보인다.
Recent approaches to Relation Extraction methods mostly tend to be limited to mention level relation extractions. These types of methods, while featuring high performances, can only extract relations limited to a single sentence or so. The inability to extract these kinds of data is a terrible amount of information loss. To tackle this problem this paper presents an Augmented External Memory Neural Network model to enable Global Relation Extraction. the proposed model’s Global relation extraction is done by first gathering and analyzing the mention level relation extraction by the Augmented External Memory. Additionally the proposed model shows high level of performances in korean due to the fact it can take the often omitted subjects and objectives into consideration.
DBMS의 웹서비스를 이용한 학습객체 메타데이터 추출 및 통합에 관한 연구 KCI 등재후보
한국정보교육학회 정보교육학회논문지 제7권 제2호 2003.08 pp.199-206
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
LiDAR 데이터 기반 도심부 사고취약상황 내 잠재적 사고 유발 차량 분류 및 추출
한국ITS학회 한국ITS학회 학술대회 SMART MOBILITY : The New Paradigm 2022.11 pp.107-116
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
빅데이터 분석 도구 R을 활용한 효율적인 특허 검색어 추출에 관한 연구
대한안전경영과학회 대한안전경영과학회 학술대회논문집 2013년 대한안전경영과학회 추계학술대회 2013.11 pp.397-401
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
디지털 기술의 발달로 세계가 정보 및 지식이 주도하는 사회로 급변하고, 지식 재산 권의 발전이 급속하게 진행되면서, 각 기업 및 국가들은 그들의 경쟁력을 키우기 위해 지식재산권에 대한 중요성을 강조하고 있다. 이와 같이 지식재산권의 중요성이 강조되 는 현실에서 지식재산권의 확보는 기업의 경쟁력을 좌우하는 요소라 할 수 있다. 따라 서 본 논문에서는 빅데이터 분석 도구인 R을 이용하여 빠른 시간 안에 사용자가 목적 으로 하고 있는 특허검색 결과를 효율적으로 도출할 수 있는 검색어 추출에 관한 연 구를 진행하였다. 이를 위해 다섯 단계의 특허 검색 프로세스를 제안하였고 프로그램 으로 구현하여 검색목적에 맞는 특허의 검색에 필요한 시간을 대폭 단축시키면서 목 표로 하는 특허 검색을 효율적으로 할 수 있었다.
빅데이터를 활용한 중소도시의 생활SOC 결핍지역 추출 연구 - 전라북도 익산시를 중심으로 - KCI 등재
한국농촌건축학회 한국농촌건축학회논문집 제22권 4호 통권 79호 2020.11 pp.43-50
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
The purpose of this study is to extract deficiency areas as basic data of policies and projects in the future Living SOC introduction and planning. In order to extract living SOC deficient areas, accessibility data for living SOC and density data for main users by facility were overlapped, focusing on the living SOC indicators presented in the National Urban Regeneration Basic Policy. According to the analysis of accessibility of the Iksan-si Living SOC, the gap between deficiency in urban and township areas was large in common with the accessibility of the village and local base units. As a result of overlapping life SOC accessibility data and density data analysis of the main users by facility, areas where accessibility is weak but not inhabited by the main users of each facility were extracted. It is meaningful that more accurate deficient areas can be extracted by simultaneously utilizing the density distribution of the main users, rather than simply accessing the facilities.
주성분분석을 이용한 기종점 데이터의 압축 및 주요 패턴 도출에 관한 연구 KCI 등재
한국ITS학회 한국ITS학회논문지 제19권 제4호 통권90호 2020.08 pp.81-99
※ 기관로그인 시 무료 이용이 가능합니다.
5,400원
기종점 데이터는 수요 분석 및 서비스 설계를 위해서 대중교통, 도로운영 등 다양한 분야에 서 저장 및 활용되고 있다. 최근 빅데이터의 활용성이 증대되면서 기종점 데이터의 분석 및 활용에 대한 수요도 함께 증가하고 있다. 기존의 일반적인 교통 정보 데이터가 수집장비 수(n) 에 비례하여 데이터양이 증가(a·n)하는 것과는 다르게, 기종점 데이터는 수집지점 수(n)의 증 가에 따라 수집 데이터의 양이 기하급수적으로 증가(a·n2)하는 경향이 있다. 이로 인하여 기종 점 데이터를 원시 데이터의 형태로 장기간 저장하고 빅데이터 분석에 활용하는 것은 대용량의 저장 공간이 필요하다는 것을 고려할 때 실용적 대안으로 여겨지지 않고 있다. 이와 함께 기종 점 데이터는 0~10 사이의 작은 수요 부분에 패턴화된 형태와 무작위 적인 형태의 데이터가 섞여있어 작은 수요가 그룹화되어 발생하는 주요 패턴을 추출하기에 어려움이 있다. 이러한 기종점 데이터의 저장용량의 한계와 패턴화 분석의 한계를 극복하고자 본 연구에서는 주성분 분석을 활용한 대중교통 기종점 데이터의 압축 및 분석 방법을 제안하였다. 본 연구에서는 서 울시와 세종시의 대중교통 이용 데이터를 활용하여 모빌리티 데이터를 분석하고, 모빌리티 기 종점 데이터에 포함된 무작위 성향이 높은 데이터를 제거하기 위해 주성분분석 기반의 데이터 압축 및 복원에 관한 연구를 수행하였다. 주성분분석으로 분해된 기종점 데이터와 원데이터를 비교하여 주요한 수요 패턴을 찾고 이를 통해 압축률과 복원율을 높일 수 있는 주성분 범위를 제안하였다. 본 연구에서 분석한 결과, 서울시 기준 1~80, 세종시 기준 1~60까지의 주성분을 사용할 경우 주요 이동 데이터의 손실 없이 기종점 데이터에 포함되어있는 노이즈를 제거하고 데이터를 압축 및 복원이 가능하였다.
Origin-destination data have been collected and utilized for demand analysis and service design in various fields such as public transportation and traffic operation. As the utilization of big data becomes important, there are increasing needs to store raw origin-destination data for big data analysis. However, it is not practical to store and analyze the raw data for a long period of time since the size of the data increases by the power of the number of the collection points. To overcome this storage limitation and long-period pattern analysis, this study proposes a methodology for compression and origin-destination data analysis with the compressed data. The proposed methodology is applied to public transit data of Sejong and Seoul. We first measure the reconstruction error and the data size for each truncated matrix. Then, to determine a range of principal components for removing random data, we measure the level of the regularity based on covariance coefficients of the demand data reconstructed with each range of principal components. Based on the distribution of the covariance coefficients, we found the range of principal components that covers the regular demand. The ranges are determined as 1~60 and 1~80 for Sejong and Seoul respectively.
스포츠 경영정보 데이터 추출로서 타액 코디졸 활용에 관한 연구 KCI 등재
한국스포츠학회 한국스포츠학회지 제16권 제2호 2018.06 pp.279-288
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
사람의 무의식적 상황을 측정하여 추출된 정보로 경영에 활용한다는 점에서는 뉴로마케팅과 유사하지만, 실험절 차나 비용 측면에서 상당한 이점이 있다고 판단되는 코티졸을 활용한 스포츠 정보 데이터 추출에 대한 탐구 결론는 다음 과 같다. 기존의 선행 연구들에 대한 이론적 체계와 연구 방법들을 잘 숙지한다면, 연구 설계에는 큰 어려움이 없어 보이 며, 타액채취도 피설험자에 큰 부담을 주지 않고 채취가 가능하나 몇 가지 주의할 점에 대한 세심한 관심과 노력이 필요 할 것으로 보인다. 분석과정은 의학적 전문기술이 필요해 실험자가 직접 진행하기는 힘든 상황으로 보이며, 전문가 또는 전문기관과 협의 및 협조 하에 진행 되어야 할 것으로 보여 진다. 이러한 과정 등을 통해 획득한 데이터는 다음과 같은 스포츠 경영 현장에서 활용이 가능할 것으로 사료된다. 첫째, 스포츠 시설업 경영에 있어 스포츠 센터가 운영하는 프로 그램 등이 현대인의 스트레스 해소에 어떤 긍정적 영향을 미치는가에 대한 정보를 호르몬중심으로 제공 가능하게 한다. 둘째는, 프로 스포츠 경기업에서 경기 현장이든, 미디어를 통한 관람이든, 경기 결과에 따른 팬들의 스트레스 분석 등에 대한 정보 분석을 통해 팬십 측정과 팬 관리에 필요한 유익한 자료로 활용될 수 있을 것이다. 셋째, 아동, 청소년 여가시 간 활동 및 관리에 있어서 모바일 게임 중심 학생과 스포츠 중심 학생간의 연구 등에 대한 데이터들은 아동, 청소년의 올바른 성장에 필요한 최소한의 신체활동과 프로그램을 운영하는 스포츠 경영 활동에 매우 훌륭한 정보가 될 것이다.
The following is the conclusion of an exploration of sports information data extraction using cortisol : If you are familiar with the theoretical framework and research methods of previous research, it appears that there is not much difficulty in the design of research, and saliva sampling by design is not a big burden on the subjects. Data obtained from these processes and others are considered available on the following sports management sites: First, information on how a program operated by a sports center in sports design business affects stress relief by modern people can be determined based on hormones. Second, whether watching a game or watching a game through the media, measuring fans ' stress according to the results of the game can be a good basis for fan management. Third, the information of young adults who will invest a lot of time in managing leisure time in children and teenagers, and those who regularly participate in sports programs, is the information of youth.
Automatic Social Media Data Extraction SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.8 2016.08 pp.39-48
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Opinions, the key influencer of human behavior and activity is ranked as of one of the strong factors that determine the effectiveness of one’s strategy and approach in terms of influential power and trend setting capabilities. This highlights the importance of sentiment analysis done upon the extracted data. Today, statistics have shown significantly that most opinions can be obtained via many social media platforms. Social media has provided a convenient platform for web users to comfortably share their thoughts and to boldly voice up. Having to process such huge amount of data, it is proposed that automated sentiment analysis is done when extracting social media data. Using an effective algorithm which produces meaningful information from raw data, the possibilities of venturing deeper into areas like decision making and influential thinking are simply limitless.
Mobile Security and its Application SCOPUS
보안공학연구지원센터(IJSIA) International Journal of Security and Its Applications Vol.10 No.10 2016.10 pp.89-106
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Ownership of a smartphone has never been easier nowadays, and it is supported by the fact that most of the people around us have a smartphone or an equivalent smart device. What is smartphone and why is it a smart? A smartphone by definition is as the name suggests, it is a phone is smart enough to not only be limited to the features and capabilities of a traditional cellular phone but also perform what a “smart” device can. And in recent years, the device that is deemed as the most intelligent device is the computer as it is the most advance piece of technology that is commercially available to the general public. Why this is so, is because in our opinion it has revolutionized how most if not all of the societies of today work. Hence what makes a smartphone is the mobile operating system that it is built upon, which is similar to a computer. It is becoming more and more of a common sight nowadays and this is because they are being offered at a price where more people are able to afford, hence they are reaching the hands of ceiling of the lower income families, all the way up to the higher income families. Back then, pure play devices were mostly simple in terms of how it function and works, hence if possible, we could suggest that the security aspect was never or rather has never been an issue other than the alteration of data after operation such as the tape of video recorders or images captured but never in the process, in the sense that there were no interruptions during operation, most likely is because it was clear and visible, but nowadays when you combine all of those devices into a complex entity, we tend to leave a hole in the cloth somewhere that we did not or rather can’t see due to the overwhelming amount of other things that we have. In this paper we discuss the current state of the commercially available operating systems of the two biggest names in smart devices, namely iOS and Android; and measure how secure and/or vulnerable (susceptible) are they to malwares and the nature of the mobile ad hoc network. We first analyze the integrity of the core of a smart device, the operating system and then use it to evaluate the effectiveness of their techniques and defenses of preventing and identifying malwares.
A Wrapper for Extracting Information Records from Forums based on Page Segmentation
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.7 No.4 2014.08 pp.11-22
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Foraging information from web forums is still one of the most challenging information retrieval tasks due to various combinations of auto-generated page structural information and user-created contents. Traditional information extraction methods employ either duplicated subtree pattern detection methods, or machine learning methods. Due to the periodical update of forum templates and diversity of page contents, aforementioned approaches do not work very well on forum sites. In this paper, we present a page-segmentation based wrapper specially designed for mining data pattern of web forums, which combines a novel page segmentation algorithm and decision tree classifier together to detect the data pattern in forum. In the segmentation phase, a novel page segmentation algorithm is proposed to identify the records areas in a page, then a classifier is adopted to identify the detailed pattern of each record in the extraction phase. Extensive experiments on various types of web forums are conducted and the results conclude that our wrapper is a more generalized one which requires only few labeled training data.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.