<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>mt_9.log</title>
        <link>https://velog.io/</link>
        <description>아홉 개의 산</description>
        <lastBuildDate>Mon, 27 Sep 2021 04:45:59 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <image>
            <title>mt_9.log</title>
            <url>https://images.velog.io/images/mt_9/profile/7afa7ea7-afa7-4050-a27c-e22c4ece9a30/social.png</url>
            <link>https://velog.io/</link>
        </image>
        <copyright>Copyright (C) 2019. mt_9.log. All rights reserved.</copyright>
        <atom:link href="https://v2.velog.io/rss/mt_9" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[입사 지원 현황]]></title>
            <link>https://velog.io/@mt_9/%EC%9E%85%EC%82%AC-%EC%A7%80%EC%9B%90-%ED%98%84%ED%99%A9</link>
            <guid>https://velog.io/@mt_9/%EC%9E%85%EC%82%AC-%EC%A7%80%EC%9B%90-%ED%98%84%ED%99%A9</guid>
            <pubDate>Mon, 27 Sep 2021 04:45:59 GMT</pubDate>
            <description><![CDATA[<ul>
<li><p>서류 지원</p>
</li>
<li><p>필기 합격</p>
<ul>
<li>한전KPS: 5-&gt;1명, 12/17 발표 (<a href="https://kps.saramin.co.kr/">https://kps.saramin.co.kr/</a>)</li>
</ul>
</li>
<li><p>서류 탈락</p>
<ul>
<li>기술보증기금</li>
<li>한국과학창의재단</li>
<li>한국전력거래소</li>
<li>한국농어촌공사</li>
<li>한국방송광고진흥공사</li>
</ul>
</li>
<li><p>필기 탈락</p>
<ul>
<li>한국특허정보원(포기)</li>
<li>한국산업인력공단</li>
<li>한국산업은행(포기)</li>
<li>IBK기업은행(포기)</li>
<li>한국사회보장정보원(포기)</li>
</ul>
</li>
<li><p>면접 탈락</p>
<ul>
<li>한국소비자원</li>
<li>우체국금융개발원</li>
<li>공무원연금공단</li>
<li>한전KDN</li>
</ul>
</li>
</ul>
]]></description>
        </item>
        <item>
            <title><![CDATA[빅분기 최종 후기]]></title>
            <link>https://velog.io/@mt_9/%EB%B9%85%EB%B6%84%EA%B8%B0-%EC%B5%9C%EC%A2%85-%ED%9B%84%EA%B8%B0</link>
            <guid>https://velog.io/@mt_9/%EB%B9%85%EB%B6%84%EA%B8%B0-%EC%B5%9C%EC%A2%85-%ED%9B%84%EA%B8%B0</guid>
            <pubDate>Fri, 16 Jul 2021 02:54:42 GMT</pubDate>
            <description><![CDATA[<p>오늘 제2회 빅데이터 분석기사이자, 첫 번째 빅데이터 분석기사 결과가 발표되었습니다. 합격하기 위해 노력한 여정을 정리할 겸, 이렇게 후기를 적어봅니다.</p>
<h3 id="필기">필기</h3>
<p>우선 이 시험이 있게 된 사실을 2019년부터 알았습니다. 2020년 초에 제1회 시행 계획이 나왔고, 7-8월에 빅데이터분석기사 관련 참고서로 처음 나온 시대x시 책으로 공부하다가 퀄리티가 너무 안 좋아서 11월에 수x비 책으로 갈아탔습니다. 하루에 1-2시간씩 한 달 간 수x비 책으로 공부했고, 12월 19일에 있을 제1회 빅데이터 분석기사를 대비했습니다.</p>
<p> 그러나 왠걸, 코로나 유행으로 1회가 취소되고, 내년 4월에 제2회 시험으로 연기됐습니다. 하는 수 없이 2021년 4월을 기다렸고, 그 사이 ADsP(데이터분석준전문가) 시험 공부를 했는데 지금 생각해보면 이 시험 준비로 빅데이터분석기사 준비도 되었던 것 같습니다. (결과는 떨어졌지만요.)</p>
<p> 어쨌든 2021년 4월 17일 필기 시험을 봤습니다. 첫 시험이라 참고서에 없었던 문제도 많았지만, 한편으로는 참고서의 도움으로 풀 수 있었던 문제도 있었습니다. (업데준 모평전, 샘탐수모검, 기준분시평 등을 달달 외웠습니다 ㅋㅋ) 만약 3회 시험도 2회 시험처럼 출제된다고 한다면, 빅데이터이나 머신러닝 외에도 통계학 쪽도 공부하면 좋을 것 같습니다.</p>
<p>  <img src="https://images.velog.io/images/mt_9/post/bb103939-0dca-4b77-862a-fa356247b483/%EB%B9%85%EB%B6%84%EA%B8%B0%ED%95%84%EA%B8%B0.png" alt=""></p>
<h3 id="실기">실기</h3>
<p> 필기 시험 결과, 무사히 합격했고, 실기를 준비해야 했습니다. 그러나 실기의 경우에는 참고서도 없어서 어떻게 준비해야 될 지 막막했습니다. 그러나 단답형 예상문제, 작업형 연습문제를 정리해주신 고마운 분이 계셔서 준비를 잘했던 것 같습니다. 단답형 문제의 경우, 빅분기 대비 어플로 공부했고, 전체 문제 파일을 앱 개발자 분께 받아서 정리한 뒤, 달달 외웠습니다. 작업형의 경우에는 캐글의 타이타닉 데이터, 작업형 예시 문제 위주로 연습했습니다. 실제 시험 환경에서는 코드 자동 완성 기능이 없으므로, PyCharm에서 코드 자동 완성 기능을 끄고 연습했습니다.</p>
<p> 그렇게 준비하고 6월 19일 시험을 봤더니, 전체적으로 시험이 쉬워서 맥이 풀렸습니다. 단답형은 거의 다 맞을 정도로 풀 수 있었고, 작업형 제1유형은 코드 작성에 버벅이기도 하고, 함수 이름이 생각이 안 나서 help나 dir를 돌리기도 했지만, 어찌어찌 풀 수 있었습니다. 제2유형은 전처리 과정에서 버벅이긴 했지만 해결하니까 나머지는 쉬웠습니다.</p>
<p> <img src="https://images.velog.io/images/mt_9/post/594ffe5a-6e84-413a-846d-44831cd41f79/%EB%B9%85%EB%B6%84%EA%B8%B0%EC%8B%A4%EA%B8%B0.png" alt=""></p>
<h3 id="소감">소감</h3>
<p>그렇게 빅분기 실기 100점으로 마무리했습니다. 최초의 빅데이터 분석기사 자격증 취득자라는 점에 뿌듯하지만, 한편으로는 출제진에 대해 여러 번 실망했습니다. 제 1회 시험 취소 통보를 2-3일 전에 했다는 점, 필기 시험 출제에 여유가 있었음에도 오류가 있었다는 점, 실기 시험 2유형의 채점 기준이 이상하다는 점 등 이게 기사 시험이 맞는 지 의문이 갈 정도였습니다. 이 점은 출제진이 바뀌지 않는 한, 제3회 시험도 말이 많을 것 같네요.</p>
<p>빅데이터분석기사 시험을 보신 분들, 다들 힘든 상황에도 고생 많으셨습니다. 그리고 앞으로 보실 분들에게 이 글이 도움이 되었으면 좋겠습니다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[빅분기 최종 합격]]></title>
            <link>https://velog.io/@mt_9/%EB%B9%85%EB%B6%84%EA%B8%B0-%EC%B5%9C%EC%A2%85-%ED%95%A9%EA%B2%A9</link>
            <guid>https://velog.io/@mt_9/%EB%B9%85%EB%B6%84%EA%B8%B0-%EC%B5%9C%EC%A2%85-%ED%95%A9%EA%B2%A9</guid>
            <pubDate>Fri, 16 Jul 2021 01:37:38 GMT</pubDate>
            <description><![CDATA[<p><img src="https://images.velog.io/images/mt_9/post/49ffdcfd-dad3-4a2c-b582-15d362f63d02/image.png" alt=""></p>
<p>작업형 제2유형에서 만점이 나올 거라곤 생각 못했다.</p>
<p>출제진의 채점 기준이 심히 궁금하다...</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[빅데이터분석기사 실기 후기]]></title>
            <link>https://velog.io/@mt_9/%EB%B9%85%EB%8D%B0%EC%9D%B4%ED%84%B0%EB%B6%84%EC%84%9D%EA%B8%B0%EC%82%AC-%EC%8B%A4%EA%B8%B0-%ED%9B%84%EA%B8%B0</link>
            <guid>https://velog.io/@mt_9/%EB%B9%85%EB%8D%B0%EC%9D%B4%ED%84%B0%EB%B6%84%EC%84%9D%EA%B8%B0%EC%82%AC-%EC%8B%A4%EA%B8%B0-%ED%9B%84%EA%B8%B0</guid>
            <pubDate>Tue, 22 Jun 2021 02:22:14 GMT</pubDate>
            <description><![CDATA[<p>6월 19일 토요일 빅데이터분석기사를 보고 왔다.</p>
<p>필기에서 많이 데여서, 실기는 여러 방면으로 준비를 했다.</p>
<h3 id="준비">준비</h3>
<p>단답형 : DataManim 님의 빅데이터 대비 앱으로 공부. 전체 문제 파일도 받아서 정리한 뒤, 달달 외웠다.</p>
<p>작업형 : 캐글의 타이타닉 데이터, 작업형 예시 문제 위주로 연습했다. 이때, 실제 시험 환경에서는 코드 자동 완성 기능이 없으므로, PyCharm에서 코드 자동 완성 기능을 끄고 연습했다.</p>
<p>그 외 : 오픈카톡방에 들어가 예상 문제와 성능 개선 방안을 공유해가며 준비했다.</p>
<h3 id="시험">시험</h3>
<p>전체적으로 쉬웠다. 
단답형 : 필기에서 공부했던 것들을 다시 복습해서 그런지, 거의 다 맞을 정도로 풀 수 있었다. </p>
<p>작업형 1유형 : 이상치의 총합, 결측값 제거 전후의 표준편차 차이 등, 데이터를 가공하는 문제 위주로 나왔다. 가끔 코드 작성에 이상이 있어서 버벅이기도 하고, 함수 이름이 생각이 안나 help나 dir를 돌리기도 했지만, 이를 잘 해결할 수 있다면 역시 큰 문제 없이 풀 수 있었다.</p>
<p>작업형 2유형 : 캐글과 같이 실제 데이터셋을 가공하고 분석하여 예측값을 만드는 문제이다. 데이터 분석 전, 전처리 과정에서 2-30분 정도 버벅이긴 했지만, 이를 해결하니 나머지는 쉬웠다. roc-auc 값음 0.669 정도 나왔고, 부분 점수 산출이 어떻게 될 지에 따라 합격 여부가 갈릴 듯 하다.</p>
<h3 id="소감">소감</h3>
<p>전공자의 경우, 필기만 잘 통과한다면 실기는 어렵지 않게 풀 수 있을 것 같다. 다만 실무에서는 코드 자동 완성 기능과 검색을 생활화된 반면, 시험에서는 그게 막혔다는 점에서, 시험 환경이 실무 환경과 동떨어져 있는 점은 아쉽다. 검색을 허용하는 대신, 문제를 더 어렵게 내도 좋지 않을까 싶다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[Data Analysis About Bitcoin Transaction]]></title>
            <link>https://velog.io/@mt_9/Data-Analysis-About-Bitcoin-Transaction</link>
            <guid>https://velog.io/@mt_9/Data-Analysis-About-Bitcoin-Transaction</guid>
            <pubDate>Wed, 09 Jun 2021 01:00:48 GMT</pubDate>
            <description><![CDATA[<p><strong>2021-1 Information Security</strong></p>
<p><strong>Gu San | Yang Yue</strong></p>
<p><a href="https://github.com/MountainNine/bitcoin_fraud_detection">https://github.com/MountainNine/bitcoin_fraud_detection</a></p>
<h3 id="introduction">Introduction</h3>
<p>Blockchain is an intelligent P2P network that uses a decentralized database to identify, disseminate, and record information, also known as the Internet of Value. As such blockchain-based cryptocurrencies become known, their demand is increasing even more. However, as the demand, the number of frauds using cryptocurrency is increasing. In this paper, we will look at previous Bitcoin fraud cases and fraud detection tools. Then, we will look at the distribution of user data based on the transaction data. To the end, we analyze the data.</p>
<h3 id="previous-incidents">Previous Incidents</h3>
<p><strong>1. Value overflow event(August 2010)</strong>
Essentially, when running the code, if the output is too large to overflow in the summation, then the code that checks Bitcoin transactions will be invalid. The hacker used this to create the block and added two transactions of 92 billion BTC. This is an early example of a cryptocurrency &quot;hard fork&quot;, and this is one of the few cases where hackers directly attacked blockchain technology.</p>
<p><strong>2. Bitfinex (August 2016)</strong>
This is the second largest exchange hacking after the MtGox hot wallet was stolen. Ironically, the upgrade designed to improve security has loopholes and was exploited by hackers. Bitfinex uses the software provided by BitGo to establish a multi-signature system to authorize transactions. It is unclear how hackers can easily bypass the need for multiple keys, but the most commonly accepted assumption is that the system on the Bitfinex server is improperly installed. The hacker stole 120,000 BTC, which was worth 72 million U.S. dollars at the time.</p>
<p><strong>3. Coinrail and Bithumb (June 2018)</strong>
In June 2018, the hot wallets of two South Korean exchanges were attacked. Among them, Coinrail lost 5,300 BTC (worth 40 million U.S. dollars), and Bithumb lost 31 million U.S. dollars.</p>
<h3 id="existing-tools">Existing Tools</h3>
<p><strong>1. PARSIQ</strong>
<strong>PARSIQ</strong> is the first blockchain automation platform that allows users to set up &quot;smart triggers&quot; using ParsiQL, its advanced blockchain streaming language. Smart trigger can interact with events on the blockchain. PARSIQ also uses a tool called Mempool, which is a digital asset early warning system. If a major event occurs, such as a large number of digital assets entering the exchange, users will be the first to see it in Mempool.</p>
<p><strong>2. Coral</strong>
<strong>Coral</strong> is using blockchain to provide predictive analysis scores to help protect consumers of cryptocurrencies. Coral continuously builds a database of malicious or suspicious phishing scams and related addresses. Coral is integrated into most wallets and user trading interfaces, which means it can warn ICO investors before providing encrypted files to reckless ICO websites.</p>
<h3 id="dataset">Dataset</h3>
<p><strong>1. Reading Transaction Dataset File</strong>
The transaction dataset file’s size is 26.8 GB, so it is hard to read the file. First, we split the file with file split program. Obviously, split files were broken, but we found that they have the same json format for each line. So we split the file for 100 lines with Windows command, and created 46 split files.</p>
<p><strong>2. Transaction Construction</strong>
<strong>1)    Merkleroot</strong>
<strong>Merkel root</strong> The root of the Merkel tree is the root node of the tree, which is the result of multiple hash calculations of all node pairs in the tree. The block header must include the valid Merkel root calculated from all transaction hashes in the block.</p>
<p><strong>Merkel tree</strong> To generate a complete Merkle tree, you need to recursively hash the pair of hash nodes, and insert the newly generated hash node into the Merkle tree until there is only one hash node, which is the root of the Merkle tree. In Bitcoin, leaf nodes come from transactions in a single block.</p>
<p><strong>2)    Nonce</strong>
Headings The changing number at the end of the sentence is called a <strong>nonce (random number)</strong>.</p>
<p><strong>Nonce</strong> is used to change the output of the encryption function.
There is also a Nonce value in the block header, which records the number of hash recalculations. The Nonce value of the 100,000th block is 274148111, that is, after 274 million calculations, a valid Hash is obtained, and the block can be added to the blockchain.</p>
<p><strong>Random number</strong>
The random number is a 32-bit (4 byte) field in the Bitcoin block. After the value is set, the hash value of the block can be calculated. The hash value starts with multiple 0s. The values of other fields in the block are unchanged because they have certain meanings.</p>
<p><strong>Excess random number</strong>
As the difficulty increases, miners usually fail to find a block after cycling the random value 400 million times. Because the coinbase script can store 2 to 100 bytes of data, miners began to use this storage space as an excess nonce space, allowing them to use a larger range of block header hash values to find valid blocks.</p>
<p><strong>3)    Previousblockhash</strong>
Hash of the previous block.</p>
<p><strong>4)    Hash</strong>
<strong>content</strong>
The Hash of each block is calculated for the &quot;block header&quot; (Head)
<strong>algorithm</strong>
SHA256 is the Hash algorithm of the blockchain
<strong>condition</strong>
Only Hash that meets the conditions will be accepted by the blockchain. This condition is particularly harsh, so that most hashes do not meet the requirements and must be recalculated.
<strong>length</strong>
The hash length of the blockchain is 256 bits
Corollary 1: The Hash of each block is different, and the block can be identified by Hash.
Corollary 2: If the content of a block changes, its Hash will definitely change. </p>
<p><strong>5)    Nextblockhash</strong>
Hash of the next block.</p>
<p><strong>6)    Vin(input) Vout(output)</strong>
In every vin, there is a txid field
For every vout, there are 1 field value, field n, and field scriptPubKey.</p>
<p><strong>Value</strong> is the amount of money transferred to the other party, and n is the number in the output (that is, the field vout in the corresponding vin when the UTXO is spent next time).</p>
<p><strong>scriptPubKey</strong> is the public key of the other party. 
There are two fields in scriptSig, asm and hex, asm is assembly (assembly, or the meaning of assembly), and hex is its hexadecimal expression.</p>
<p><strong>Script Engine</strong>
Simply looking at the scriptPubKey is incomprehensible. We need to splice the scriptPubKey with the next scriptSig of the transaction that will cost it.
scriptPubKey represents the public key information of the payee; the vin in the next transaction that will spend this UTXO must have the previous information of the private key corresponding to the public key, that is, scriptSig.
The two parts are spliced together, with scriptSig in front and scriptPubKey in back (note: scriptPubKey is the current transaction, scriptSig is the next transaction, which will cost this UTXO).
Among them, the contents of the two parts of signature and pub key are similar to the asm part in vin above. Because the string is too long, it is omitted here and replaced by.
The above string of things is a script, a simple DSL language (about what is DSL, Baidu&#39;s own).
This string of scripts is thrown into the Script Engine for execution, and the execution result is TRUE/FALSE, indicating whether the person is eligible to spend the money, or whether the transaction is valid.</p>
<h3 id="processing">Processing</h3>
<p><strong>1. Transaction Count</strong></p>
<p>Every line has nTx, which saves the number of transactions, and medianTime, which the transaction occured. We saved every nTx value and medianTime. At the result, we knew total transaction count of January 2019 is 9,222,178 counts. Following code express how to save transaction count per day.</p>
<pre><code>def save_transaction():
    median_list = []
    nTx = []

    for i in range(1, 47):
        file_name = &quot;D:\\download\coin\coin{}.json&quot;.format(i)
        with open(file_name, &#39;r&#39;) as file_data:
            line = file_data.readline()
            while line:
                data = json.loads(line)
                median_list.append(data[&#39;mediantime&#39;])
                nTx.append(data[&#39;nTx&#39;])
                line = file_data.readline()

            file_data.close()
            print(&quot;coin{} Clear&quot;.format(i))

    df_transaction = pd.DataFrame({&#39;medianTime&#39;: median_list, &#39;nTx&#39;: nTx})
    df_transaction.to_csv(&#39;./transaction.csv&#39;, sep=&#39;,&#39;, na_rep=&#39;NaN&#39;)

def show_transaction():
    data = pd.read_csv(&quot;transaction.csv&quot;)
    data[&#39;medianTime&#39;] = pd.to_datetime(data[&#39;medianTime&#39;], unit=&#39;s&#39;)
    df = data.groupby(by=data[&#39;medianTime&#39;].dt.date).sum()
    y = df.get(&#39;nTx&#39;)
    plt.bar(df.index, y)
    plt.show()</code></pre><p><img src="https://images.velog.io/images/mt_9/post/3c8778e3-aea0-439d-a50e-bef8f9782f2c/Figure_1.png" alt=""></p>
<p><strong>2. Address Count</strong></p>
<p>In tx(transaction) array of each line, there are addresses  that receive Bitcoin, and values of Bitcoin. We extracted those addresses and values, and transaction counts per address in each tx array, and saved to csv file for next processing. Total unique wallet address count of January 2019 is <strong>10,541,781 counts</strong>. Following code express to save unique wallet address.</p>
<pre><code>def get_address():
    sum_nTx = 0
    address = set()

    for i in range(1, 47):
        file_name = &quot;D:\\download\coin\coin{}.json&quot;.format(i)
        with open(file_name, &#39;r&#39;) as file_data:
            line = file_data.readline()
            while line:
                data = json.loads(line)
                for tx in data[&#39;tx&#39;]:
                    for vout in tx[&#39;vout&#39;]:
                        if vout[&#39;scriptPubKey&#39;][&#39;type&#39;] not in (&#39;nulldata&#39;, &#39;nonstandard&#39;):
                            address.update(vout[&#39;scriptPubKey&#39;][&#39;addresses&#39;])
                line = file_data.readline()

            file_data.close()
            print(&quot;coin{} Clear.&quot;.format(i))

    text_file = open(&#39;address.txt&#39;, &#39;w&#39;)
    for a in address:
        text_file.write(a + &quot;\n&quot;)
    text_file.close()</code></pre><p><strong>3. Abusing Address</strong>
There are abusing addresses because of ransomware, fraud, etc. We crawled those abusing users, and compared with previous address data. About 500 addresses are abusing address. Following code represents how we crawled those users, and following graph represents the distribution of  abusing address and not abusing address with transaction count and sum of value per address.</p>
<pre><code>def get_abuse_address():
    import pandas as pd
    from bs4 import BeautifulSoup
    import requests
    import time
    import numpy as np

    abuse_addresses = set()

    for i in range(1, 2248):
        page = requests.get(&quot;https://www.bitcoinabuse.com/reports?page={}&quot;.format(i))
        soup = BeautifulSoup(page.content, &quot;html.parser&quot;)
        for s in soup.select(&quot;div.col-xl-4 a&quot;):
            abuse_addresses.append(s.contents[0])
        print(&quot;{} passed.&quot;.format(i))
        time.sleep(1)

    df = pd.DataFrame(abuse_addresses)
    df_address = pd.read_csv(&#39;users_new.csv&#39;, index_col=False)
    common_address = pd.merge(df, df_address, how=&#39;inner&#39;, left_on=0, right_on=&#39;address&#39;)
    df_address[&#39;is_fraud&#39;] = np.where(df_address[&#39;address&#39;].isin(common_address[&#39;address&#39;]), 1, 0)
    df_address.to_csv(&#39;users_new_2.csv&#39;, index=None)</code></pre><p><img src="https://images.velog.io/images/mt_9/post/9a27fbcf-d8b7-4ea4-bfac-32326c354034/Figure_9.png" alt=""></p>
<h3 id="clustering">Clustering</h3>
<p>We clustered above graph with Gaussian Mixture Model.  Gaussian mixture model is a probabilistic model that assumes all the data points are generated from a mixture of a finite number of Gaussian distributions with unknown parameters. Following code and graph is the clustered graph with Gaussian mixture model(n=9).</p>
<pre><code>def gaussian_mixture():
    from sklearn.mixture import GaussianMixture
    import matplotlib.pyplot as plt
    import pandas as pd
    import numpy as np
    import time

    df = pd.read_csv(&#39;users_new_3.csv&#39;, index_col=False)
    list_gm = pd.DataFrame(map(list, zip(df[&#39;count&#39;], df[&#39;sum&#39;])))

    gm = GaussianMixture(n_components=9, random_state=0).fit_predict(list_gm)

    plt.scatter(list_gm[0], list_gm[1], s=9, c=gm)
    plt.xscale(&#39;log&#39;)
    plt.yscale(&#39;log&#39;)
    plt.show()</code></pre><p><img src="https://images.velog.io/images/mt_9/post/c7b077ff-e95b-4dab-b3ce-9bcdee4f1838/Figure_11.png" alt=""></p>
<p>purple and yellow clusters tend to have a high proportion of abusing address. It seems the higher transaction count or the higher the value, the higher the probability of fraud. </p>
<h3 id="conclusion">Conclusion</h3>
<p>In this paper, we have investigated the incident about Bitcoin fraud and existing fraud detection tools, Bitcoin transaction of January 2019, and abusing address. Then we have analyzed the distribution of address data with Gaussian Mixture Model Clustering. 
We tried to cluster with DBSCAN(Density-based spatial clustering of applications with noise), but it had a memory shortage problem, and the clustering result didn’t appear. We will find the outlier based on density on next study.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[빅분기 필기 합격]]></title>
            <link>https://velog.io/@mt_9/%EB%B9%85%EB%B6%84%EA%B8%B0-%ED%95%84%EA%B8%B0-%ED%95%A9%EA%B2%A9</link>
            <guid>https://velog.io/@mt_9/%EB%B9%85%EB%B6%84%EA%B8%B0-%ED%95%84%EA%B8%B0-%ED%95%A9%EA%B2%A9</guid>
            <pubDate>Fri, 04 Jun 2021 00:14:44 GMT</pubDate>
            <description><![CDATA[<p><img src="https://images.velog.io/images/mt_9/post/bccd4b88-fbc7-44ad-8505-39f00e974444/image.png" alt=""></p>
<p>망했다고 생각했는데, 다행히도 합격했다.</p>
<p>근데 실기는 참고서 같은 것도 없어서 그런지 더 막막하다...</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[제2회 빅데이터 분석기사 필기 후기]]></title>
            <link>https://velog.io/@mt_9/%EC%A0%9C2%ED%9A%8C-%EB%B9%85%EB%8D%B0%EC%9D%B4%ED%84%B0-%EB%B6%84%EC%84%9D%EA%B8%B0%EC%82%AC-%ED%95%84%EA%B8%B0-%ED%9B%84%EA%B8%B0</link>
            <guid>https://velog.io/@mt_9/%EC%A0%9C2%ED%9A%8C-%EB%B9%85%EB%8D%B0%EC%9D%B4%ED%84%B0-%EB%B6%84%EC%84%9D%EA%B8%B0%EC%82%AC-%ED%95%84%EA%B8%B0-%ED%9B%84%EA%B8%B0</guid>
            <pubDate>Sat, 17 Apr 2021 11:42:36 GMT</pubDate>
            <description><![CDATA[<p>제2회 빅데이터분석기사...이지만 제1회는 코로나로 취소되는 바람에 실질적으로 이게 첫 시험이다.</p>
<p>시험에 대해서 복기하고, 내 미련함과 시험에 대해 한탄하는 마음으로 글을 써본다.</p>
<h3 id="시험-소감">시험 소감</h3>
<p>시험을 보고 나온 소감은, 확실하게 불합격이라는 느낌이었다.</p>
<p>문제가 어려웠다기 보다는, 내가 준비한 책에서 나오지 않은 문제들이 6할 정도는 됐던 것 같다.</p>
<p>책을 여러 권 보면서 준비했어야 했는데, 그러지 못한 내 불찰이다.</p>
<p>한편으로는 이게 정말 &#39;빅데이터&#39; 분석기사가 맞는지도 의문이다.</p>
<p>기억나는 문제와 다른 사람들이 올린 문제들을 토대로 정리해보자면,</p>
<ol>
<li>하둡, 빅데이터 관련 소프트웨어 이름을 달달 외웠지만 하나도 안 나왔다.</li>
<li>2과목은 거의 찍다시피 했다. 다른 과목도 그렇지만, 2과목에서 통수를 세게 맞은 기분...</li>
<li>계산 문제가 많이 나왔다.(합성곱 계산, t분포 신뢰도 계산 등)</li>
<li>4과목에서는 혼동 행렬에서 문제가 3-4문제나 나왔던 거로 기억한다.</li>
<li>이건 나중에 안 사실인데, &#39;복수정답일 시 모두 체크하시오.&#39;라는 문장이 유의사항에 있었다고 한다. 1페이지 같이 그냥 넘어갈 부분이 아니라, 문제 부분에 기재했어야 되지 않나?</li>
<li>빅데이터 분석기사인데 3V나 5V 같은 기초개념도 묻지 않은 것도 의외였다.</li>
</ol>
<p>첫 시험이라 혼란이 있을 거라는 건 예상했지만, 막상 당하고 보니 허탈한 기분밖에 안 든다.
지금은 결과를 기다릴 수밖에......</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[서울에서 나에게 있어 급식이 가장 맛있는 학교는?]]></title>
            <link>https://velog.io/@mt_9/%EC%84%9C%EC%9A%B8%EC%97%90%EC%84%9C-%EB%82%98%EC%97%90%EA%B2%8C-%EC%9E%88%EC%96%B4-%EA%B8%89%EC%8B%9D%EC%9D%B4-%EA%B0%80%EC%9E%A5-%EB%A7%9B%EC%9E%88%EB%8A%94-%ED%95%99%EA%B5%90%EB%8A%94</link>
            <guid>https://velog.io/@mt_9/%EC%84%9C%EC%9A%B8%EC%97%90%EC%84%9C-%EB%82%98%EC%97%90%EA%B2%8C-%EC%9E%88%EC%96%B4-%EA%B8%89%EC%8B%9D%EC%9D%B4-%EA%B0%80%EC%9E%A5-%EB%A7%9B%EC%9E%88%EB%8A%94-%ED%95%99%EA%B5%90%EB%8A%94</guid>
            <pubDate>Wed, 03 Mar 2021 07:06:04 GMT</pubDate>
            <description><![CDATA[<h2 id="개요">개요</h2>
<p>학창 시절, 맛있는 식사에 행복해하고, 맛없는 식사에 불행을 느낄 때가 많았다. 그러면서 언제나 맛있는 식사가 나오길 바란 적도 많았다.</p>
<p>그렇다면, 전국 초중고 중에서 언제나 내가 가장 맛있어할 식사가 나올 학교는 어디일까? 이를 R을 이용해 분석했다.</p>
<h2 id="데이터셋">데이터셋</h2>
<p><a href="https://open.neis.go.kr">나이스 교육정보 개방 포털</a> 에서 전국 모든 학교의 기본 정보 및 급식 정보를 조회할 수 있다. 학교 기본 정보는 csv로,
급식 정보는 학교마다 조회해야 되는 번거로움이 있어, 코드 상에서 API 호출을 통해 가져오도록 했다.</p>
<h2 id="코드-설명">코드 설명</h2>
<pre><code>library(httr)
library(jsonlite)
library(dplyr)
Sys.setlocale(&quot;LC_CTYPE&quot;, &quot;.1251&quot;)</code></pre><p>API를 호출해서 JSON 형태의 데이터를 가져와야 하기 때문에 httr과 jsonlite 라이브러리를 로드했다.<br>또한 한글 데이터가 있으므로, 로케일 정보를 .949에서 .1251로 바꿔주었다.</p>
<pre><code>school_info &lt;- read.csv(&quot;./school_info.csv&quot;, encoding = &quot;UTF-8&quot;)
school_info$score &lt;- 0</code></pre><p>기존에 다운로드 받은 학교 기본 정보 csv 파일을 불러와서 dataframe으로 저장해준다. 이때, 인코딩 옵션은 UTF-8로 한다.<br>그 다음, 학교마다 급식에 대한 평균 점수를 저장하기 위해, score 컬럼을 추가한다.</p>
<pre><code>list_pattern &lt;- c(&quot;국수|스파게티|[가-힣]*면$|소바|모밀|우동|치킨|[가-힣]*버거|머핀|카스테라|케이크|케익|케잌|파스타|핫도그|도넛|토스트|피자|와플|쿠키&quot;,
                   &quot;[가-힣]*(수제비|불고기|가스|까스|스테이크|튀김|구이)$|떡국|[가-힣]*(스프|수프|빵|볶음|볶이)|돼지|닭|쥬스|주스|바나나|딸기|귤|수박|참외|멜론|포도|사과|갈비&quot;,
                    &quot;카레$|커리$|(하이|카레|오므|커리|자장|짜장)라이스|비빔밥|볶음밥|만두|필라프|짬뽕|탕수육|훈제&quot;,
                  &quot;덮밥|토마토$|[가-힣]*(찜|잡채|샐러드|전|탕|찌개|죽|떡|말이)$|샌드위치|방울토마토|우유|주물럭|두부|커틀릿|커틀렛&quot;,
                  &quot;[가-힣]*(밥|김치|깍두기|겉절이)$&quot;,&quot;[가-힣]*국$|육개장|나물&quot;, &quot;[가-힣]*조림$&quot;,&quot;[가-힣]*(냉채|생채|무침)$&quot;)
list_score &lt;- c(3,2,2,1,0,-1,-2,-3)</code></pre><p>맛있거나 맛없는 음식재료, 또는 요리를 7점 척도로 나열하고, 정확한 검색을 하기 위해 정규식으로 바꿔주었다.</p>
<pre><code>for (sch_code in school_info[, 3]) {
  df_dish &lt;- NULL
  i &lt;- 1
  while (1) {
    req_url &lt;- paste0(&quot;https://open.neis.go.kr/hub/mealServiceDietInfo?KEY=cee31841befb40f8ab6af2c8a856686d&amp;Type=json&amp;ATPT_OFCDC_SC_CODE=B10&amp;SD_SCHUL_CODE=&quot;,
                      sch_code, &quot;&amp;pSize=1000&amp;MMEAL_SC_CODE=2&amp;pIndex=&quot;, i)
    response &lt;- RETRY(&quot;GET&quot;, url = req_url, timeout(30), times = 5)
    response &lt;- fromJSON(content(response, as=&quot;text&quot;))</code></pre><p>학교마다 다른 URL을 호출해야 하므로, 학교 기본 정보의 3번째 열에 있는 학교 코드를 for문을 통해 호출한다. <br>그리고 각 학교에 저장된 급식의 수가 다르므로, while(1)로 무한반복을 걸었는데,어떻게 탈출하는지는 뒤에 후술하겠다. <br>요청할 URL은 위와 같다. 여기서는 서울에 있는 학교의 중식만 호출하겠다. <br>학교 코드와 페이지 수를 변수로 지정했고, 반복문이 돌아갈 때마다 두 변수들이 바뀔 것이다. <br>마지막으로, 응답 과정에서 timeout error를 방지하기 위해, httr의 RETRY를 통해서 timeout 시간과 반복 횟수를 늘려주었다.</p>
<pre><code>    if (&quot;mealServiceDietInfo&quot; %in% names(response)) {
      df_response &lt;- as.data.frame(response$mealServiceDietInfo$row[2])
      list_dish &lt;- df_response$DDISH_NM
      list_dish_split &lt;- strsplit(list_dish, split = &quot;&lt;br/&gt;&quot;)
      list_dish_only_name &lt;- sapply(list_dish_split, function(x) gsub(&quot;[0-9]+.&quot;, &quot;&quot;, x))
      list_dish_final &lt;- lapply(list_dish_only_name, &#39;length&lt;-&#39;, max(lengths(list_dish_only_name)))
      if (is.null(df_dish)) {
        df_dish &lt;- as.data.frame(list_dish_final)
        df_dish &lt;- data.frame(t(df_dish))
        rownames(df_dish) &lt;- seq_len(nrow(df_dish))

      } else {
        temp_dish &lt;- as.data.frame(list_dish_final)
        temp_dish &lt;- data.frame(t(temp_dish))
        rownames(temp_dish) &lt;- seq_len(nrow(temp_dish))
        df_dish &lt;- bind_rows(df_dish, temp_dish)
      }

      i &lt;- i + 1
    } else {
      break
    }
  }</code></pre><p>다음은 json에서 가져온 데이터를 가공하는 코드이다. 성공한 json request의 경우, mealServiceDietInfo라는 컬럼명이 포함된다. <br>만약 데이터가 하나도 없을 경우에는 이 컬럼이 포함되지 않으므로, break를 걸어서 무한루프에서 벗어나게 해준다.<br>또한 기본적인 급식 데이터의 경우, 다음과 같이 나온다.</p>
<pre><code>영양닭죽(녹두)13.&lt;br/&gt;김말이떡볶이(부식)1.5.6.13.&lt;br/&gt;만두튀김21.5.6.10.12.13.&lt;br/&gt;배추김치9.13.&lt;br/&gt;바나나&lt;br/&gt;급식우유2.</code></pre><p>이 데이터에서 br이라 써진 문자열을 구분자로 하여 문자열을 리스트로 변환시켜준다. 
그 다음, 정규식을 이용해 요리명 뒤의 숫자와 온점을 없애주고, 이를 df_dish에 저장시킨다.
이때, 기존의 df_dish에 데이터가 있을 경우, temp_dish로 저장시켜 bind_rows 함수로 병합시킨다.</p>
<pre><code>  if (!is.null(df_dish)) {
    rm(list_dish, list_dish_only_name, list_dish_split, list_dish_final)


    df_score &lt;- data.frame(list_pattern, list_score)
    df_dish &lt;- as.data.frame(lapply(df_dish,function(x) gsub(&#39;\\([가-힣a-zA-Z0-9\\-\\/ ]*\\)|[\\*\\.\\/\\+\\-\\&amp;\\@][가-힣]*|\\([가-힣]*|\\-[가-힣]*$&#39;, &quot;&quot;, x)))
    df_dish$score &lt;- 0

    for (j in seq_len(nrow(df_dish))) {
      list_temp &lt;- sapply(df_score$list_pattern, grepl, df_dish[j,], ignore.case = TRUE)
      position &lt;- apply(list_temp, 1, function(a) head(c(which(a), 0), n = 1))
      df_dish[j, &quot;score&quot;] &lt;- sum(df_score[position, &quot;list_score&quot;])
    }

    school_info[school_info[3] == sch_code, &quot;score&quot;] &lt;- mean(df_dish$score)
    rm(list_temp, position)
  }

  print(paste(school_info[school_info$X.U.D45C..U.C900..U.D559..U.AD50..U.CF54..U.B4DC. == sch_code, 4], &quot;complete.&quot;))
}</code></pre><p>저장된 급식 식단의 개수가 0이 아니라면 요리명에 섞인 불순물(특수문자, 하이픈 등)을 다시 걸러준 다음, score 컬럼을 추가한다. <br>다시 식단 개수만큼 반복문을 돌려서, df_score에 있는 정규식 패턴에 맞는 요리명을 뽑아와서 list_temp에 저장시킨다.<br>그 다음, 각 요리명의 row를 구해서, 패턴에 맞는 점수 값들을 해당 row의 score 컬럼에 저장시킨다. <br>마지막으로, 반복이 끝나면 각 점수들의 평균을 구하여 school_info의 score에 저장시키고, <br>해당 학교의 점수 산정이 끝났음을 출력한 뒤, 다음 학교로 넘어간다.</p>
<h2 id="결과">결과</h2>
<p>school_info 중, 학교명과 평균 점수만 따로 추출해서 본 결과는 다음과 같았다.</p>
<p><img src="https://images.velog.io/images/mt_9/post/629b9ded-f68c-4f58-a4be-b796b34eb31e/result.PNG" alt=""></p>
<p>점수가 고르게 분포된 것 같지만, 전처리가 제대로 안된건지, 0점도 간간히 보인다.</p>
<p><img src="https://images.velog.io/images/mt_9/post/1c9f3423-1523-4b2f-ba2f-5c00dead61cb/result2.PNG" alt="">
점수를 기준으로 내림차순으로 정렬해보면, 서울 전체 초중고 중에서, 내 기준으로 서울신성초등학교라는 곳의 급식이 가장 맛있다고 나온다. 그렇다면 실제 급식은 어떨까.</p>
<p><img src="https://images.velog.io/images/mt_9/post/1635c974-58c9-4a82-9848-8a9d43c42b2e/result3.png" alt=""></p>
<p>음, 꽤 괜찮아 보이는 식단이다.</p>
<h2 id="느낀-점">느낀 점</h2>
<p>서울의 모든 학교의 급식에 대한 데이터를 가져오고, 처리하는 데만 한 학교에 약 3초, 총 1시간 정도가 소요되었다. <br>이 속도를 개선할 방안을 찾아봐야겠다. 그리고 가장 점수가 낮은 식단을 찾아보니, 생각보다 불순물이 잘 걸러지지 않았다. <br>정규식 필터를 좀 더 다듬어야겠다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[교육용 CreDB를 이용한 금융 데이터 분석 연습]]></title>
            <link>https://velog.io/@mt_9/%EA%B5%90%EC%9C%A1%EC%9A%A9-CreDB%EB%A5%BC-%EC%9D%B4%EC%9A%A9%ED%95%9C-%EA%B8%88%EC%9C%B5-%EB%8D%B0%EC%9D%B4%ED%84%B0-%EB%B6%84%EC%84%9D-%EC%97%B0%EC%8A%B51-%EC%97%B0%EB%A0%B9%EB%8C%80%EB%B3%84-%EB%8C%80%EC%B6%9C%EC%95%A1-%EB%8C%80%ED%91%9C%EA%B0%92-%EA%B5%AC%ED%95%98%EA%B8%B0</link>
            <guid>https://velog.io/@mt_9/%EA%B5%90%EC%9C%A1%EC%9A%A9-CreDB%EB%A5%BC-%EC%9D%B4%EC%9A%A9%ED%95%9C-%EA%B8%88%EC%9C%B5-%EB%8D%B0%EC%9D%B4%ED%84%B0-%EB%B6%84%EC%84%9D-%EC%97%B0%EC%8A%B51-%EC%97%B0%EB%A0%B9%EB%8C%80%EB%B3%84-%EB%8C%80%EC%B6%9C%EC%95%A1-%EB%8C%80%ED%91%9C%EA%B0%92-%EA%B5%AC%ED%95%98%EA%B8%B0</guid>
            <pubDate>Mon, 08 Feb 2021 03:11:08 GMT</pubDate>
            <description><![CDATA[<h2 id="개요와-데이터셋">개요와 데이터셋</h2>
<p>금융 관련 데이터를 이용해 데이터 분석을 연습해보고 싶어서 데이터를 찾아보던 중, 
한국신용정보원에서 제공하는 CreDB 개인신용 교육용DB를 발견했다. <a href="https://www.findatamall.or.kr/fsec/dataProd/generalDataProdDetail.do?cmnx=44&amp;goods_id=dafb9c70-ba89-11ea-9d5e-2d76f6f19fec">(링크)</a></p>
<p>CreDB는 신용정보원에 저장된 신용정보를 표본추출한 뒤, 
이를 비식별 조치하여 교육 목적으로 이용자에게 제공하는 서비스이다. 
내가 받은 데이터는 light 버전으로, 2000명의 대출,연체,카드개설정보가 담겨 있다.</p>
<h2 id="연령별-대출액의-대표값">연령별 대출액의 대표값</h2>
<p>먼저 이 데이터로 연령별 대출액의 대표값을 구하겠다.</p>
<pre><code class="language-library(dplyr)">library(ggplot2)
library(gmodels)

df_id &lt;- read.csv(&quot;./data/ID.csv&quot;)
df_loan &lt;- read.csv(&quot;./data/LN.csv&quot;)
df_delay &lt;- read.csv(&quot;./data/DLQ.csv&quot;)
df_cdopn &lt;- read.csv(&quot;./data/CDOPN.csv&quot;)</code></pre>
<p>먼저 라이브러리와 개인신용 데이터를 import한다.</p>
<pre><code>df_id$age &lt;- as.integer((2018 - df_id$BTH_YR) / 10)
df_loan_new &lt;- df_loan %&gt;% group_by(JOIN_KEY, COM_KEY, SCTR_CD, LN_CD_1, LN_CD_2, LN_YM, LN_AMT) %&gt;%
  summarise(YM = max(YM))
df_loan_sum &lt;- df_loan_new %&gt;% group_by(JOIN_KEY) %&gt;% summarise(SUM_AMT = sum(LN_AMT))
df_id_loan_sum &lt;- merge(df_id, df_loan_sum, by=&quot;JOIN_KEY&quot;, all.y = TRUE)</code></pre><p>순서대로 설명하자면 다음과 같다.</p>
<ol>
<li>기준 년도(2018)와 생년 정보의 차이를 통해 나이를 구한 뒤, 이를 10으로 나누어 연령대를 구한다.</li>
</ol>
<ol start="2">
<li>월마다 대출 정보가 중복되기 때문에, dplyr 라이브러리의 group_by 함수에 YM을 제외한 컬럼을 넣어주고,
summarise 함수로 YM의 최대값으로 함축시켜준다. 그렇게 하면 기존에 58922개였던 데이터가 5382개로 줄어들 것이다.</li>
</ol>
<ol start="3">
<li>중복이 제거된 데이터를 다시 group_by를 통해 JOIN_KEY(회원코드)로 묶어준다. 또한 summarise 함수를 통해,
SUM_AMT라는 컬럼을 추가해주고, 회원별로 LN_AMT(대출액)의 합이 되는 값을 넣어준다.</li>
</ol>
<ol start="4">
<li>연령대 데이터가 들어있는 df_id와 회원별 대출액의 합이 들어있는 df_loan_sum을 JOIN_KEY를 매개로 merge시킨다.
이때, 대출을 하지 않은 사람이 존재할 수도 있으므로, all.y=TRUE를 주어 RIGHT JOIN을 시킨다.</li>
</ol>
<pre><code>df_mean_result &lt;- df_id_loan_sum %&gt;% group_by(age) %&gt;%
  summarise(AMT = round(mean(SUM_AMT)), TYPE = &quot;mean&quot;)
df_median_result &lt;- df_id_loan_sum %&gt;% group_by(age) %&gt;%
  summarise(AMT = round(median(SUM_AMT)), TYPE = &quot;median&quot;)
df_result &lt;- rbind(df_mean_result, df_median_result)</code></pre><p>age(연령대)로 그룹화하여, AMT와 TYPE 컬럼을 생성시킨다. 
여기서 두 개의 데이터프레임을 생성하는 이유는, 시각화할 때 각 대표값의 차이를 명확히 보이게 하기 위함이다.
처음에 한 개의 데이터프레임에 모두 넣어서 시각화를 시도해봤으나, 잘 되지 않았다. 
따라서, 두 개의 데이터프레임에 평균값과 중위값을 각각 넣은 다음, TYPE 컬럼에 평균값이면 mean, 중위값이면 median을 넣어주었다.
마지막으로 rbind 함수로 이 두 데이터프레임을 병합시켰다.</p>
<pre><code>ggplot(df_result, aes(age, AMT, fill=df_result$TYPE)) + geom_bar(stat=&#39;identity&#39;, position = &#39;dodge&#39;) +
  geom_text(aes(y=AMT,label=AMT))</code></pre><p>마지막으로 ggplot을 통해 막대그래프로 시각화시킨 결과, 다음과 같은 그래프가 나왔다.
<img src="https://images.velog.io/images/mt_9/post/0e5433d6-d4a4-49d8-a41c-819081a9d8c4/df_result.png" alt="">
평균값과 중위값의 차이가 꽤나 크다는 점을 볼 수 있다.</p>
<h2 id="연체액과-연체-기간-상관분석">연체액과 연체 기간 상관분석</h2>
<pre><code>rm(&quot;df_id_loan_sum&quot;, &quot;df_loan_new&quot;, &quot;df_loan_sum&quot;, &quot;df_mean_result&quot;, &quot;df_median_result&quot;, &quot;df_result&quot;)</code></pre><p>먼저 위 대표값들을 구하는 과정에서 만든 데이터프레임들을 제거한다.</p>
<pre><code>df_delay_new &lt;- df_delay %&gt;% group_by(JOIN_KEY, COM_KEY, SCTR_CD,DLQ_TYPE, DLQ_CD_1, DLQ_CD_2, DLQ_YM,DLQ_AMT) %&gt;%
  summarise(DLQ_COUNT = n())
df_delay_new$DLQ_AMT &lt;- ifelse(df_delay_new$DLQ_AMT &lt; 500000, df_delay_new$DLQ_AMT, NA)
ggplot(df_delay_new, aes(x=DLQ_COUNT, y=DLQ_AMT)) + geom_point()
cor.test(df_delay_new$DLQ_COUNT, df_delay_new$DLQ_AMT)</code></pre><p>group_by와 summarise 함수를 통해 DLQ_COUNT 컬럼(개별 연체 기간)을 포함한 데이터프레임을 생성한다.
그 다음, DLQ_AMT(개별 연체액)과 DLQ_COUNT 컬럼을 산점도로 표시하면 다음과 같다.
<img src="https://images.velog.io/images/mt_9/post/fbdb266e-2159-445d-baf5-9709143ccfad/scatter01.png" alt="">
이때, 개별 연체액이 1500000인 데이터 때문에 실제와는 다른 결과가 발생할 수 있다.
따라서, 두 번째 줄에서 보이듯이 이상치를 제거해줬다.
이를 산점도로 다시 표시하면 다음과 같고,
<img src="https://images.velog.io/images/mt_9/post/27e69226-1806-49e9-ab5e-b6f04e2faa88/scatter02.png" alt="">
cor.test 함수로 상관분석을 실행하면 다음과 같은 결과가 나온다.</p>
<pre><code>data:  df_delay_new$DLQ_COUNT and df_delay_new$DLQ_AMT
t = 1.162, df = 880, p-value = 0.2455
alternative hypothesis: true correlation is not equal to 0
95 percent confidence interval:
 -0.02694022  0.10488202
sample estimates:
      cor 
0.0391412 </code></pre><p>t검정 값은 1.162가 나오고, p-value(유의확률)의 값은 0.2455로, 95%의 신뢰수준에서 개별 연체액과 개별 연체 횟수 간 상관관계는 낮다고 볼 수 있다.</p>
<p>그렇다면 개인별 총 연체액과 총 연체 기간 간 상관관계는 어떨까. 코드는 다음과 같다.</p>
<pre><code>df_delay_final &lt;- df_delay_new %&gt;% group_by(JOIN_KEY) %&gt;% summarise(DLQ_CNT = sum(DLQ_COUNT), DLQ_AMT = sum(DLQ_AMT))
df_delay_final &lt;- subset(df_delay_final,DLQ_CNT &lt; 100 &amp; DLQ_AMT &lt; 500000)
ggplot(df_delay_final, aes(x=DLQ_CNT, y=DLQ_AMT)) + geom_point()
cor.test(df_delay_final$DLQ_CNT, df_delay_final$DLQ_AMT)</code></pre><p>이번에는 개별 연체액의 총합인 DLQ_CNT 컬럼과 개별 연체 기간의 총합인 DLQ_AMT를 포함한 데이터프레임을 생성했다.
이를 산점도로 표시하면 다음과 같다.
<img src="https://images.velog.io/images/mt_9/post/7c3d45ac-c4b7-408f-916c-95456849c47f/scatter03.png" alt="">
여기서도 이상치가 발생하므로, 두 번째 줄처럼 범위를 지정해서 이를 제거해준 뒤, 산점도로 나타내면 다음과 같고,
<img src="https://images.velog.io/images/mt_9/post/e3bdf51e-e5fd-4739-bd54-e9a8e4712e22/scatter04.png" alt="">
cor.test 함수로 상관분석을 실행하면 다음과 같은 결과가 나온다.</p>
<pre><code>data:  df_delay_final$DLQ_CNT and df_delay_final$DLQ_AMT
t = 4.2706, df = 385, p-value = 2.459e-05
alternative hypothesis: true correlation is not equal to 0
95 percent confidence interval:
 0.1154312 0.3058730
sample estimates:
      cor 
0.2126708 </code></pre><p>t-검정 값은 4.2706이고, p-value 값은 2.459e-05으로, 이는 95% 신뢰수준에서 유의수준인 0.05보다 낮다고 볼 수 있다.
따라서 총 연체액과 총 연체 기간 간 상관관계는 있다고 볼 수가 있다.</p>
<h2 id="느낀-점">느낀 점</h2>
<p>사실 처음에는 개인신용 데이터를 기반으로 신용점수를 계산해보는 코드를 작성하려 했으나, 생각만큼 잘 되지는 않았다.
만약 데이터에 신용점수까지 포함되어 있었다면 회귀분석을 통해 구할 수 있었을까...</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[지도에서 서울, 수도권의 지하철 역은 어떻게 나올까?]]></title>
            <link>https://velog.io/@mt_9/%EC%A7%80%EB%8F%84%EC%97%90%EC%84%9C-%EC%84%9C%EC%9A%B8-%EC%88%98%EB%8F%84%EA%B6%8C%EC%9D%98-%EC%A7%80%ED%95%98%EC%B2%A0-%EC%97%AD%EC%9D%80-%EC%96%B4%EB%96%BB%EA%B2%8C-%EB%82%98%EC%98%AC%EA%B9%8C</link>
            <guid>https://velog.io/@mt_9/%EC%A7%80%EB%8F%84%EC%97%90%EC%84%9C-%EC%84%9C%EC%9A%B8-%EC%88%98%EB%8F%84%EA%B6%8C%EC%9D%98-%EC%A7%80%ED%95%98%EC%B2%A0-%EC%97%AD%EC%9D%80-%EC%96%B4%EB%96%BB%EA%B2%8C-%EB%82%98%EC%98%AC%EA%B9%8C</guid>
            <pubDate>Wed, 20 Jan 2021 03:19:24 GMT</pubDate>
            <description><![CDATA[<h2 id="개요">개요</h2>
<p>우리가 보는 지하철 노선도는 각 노선별 지하철역들을 한 눈에 볼 수 있게 실제 지도와는 다르게 노선도를 표시한다.
그렇다면 실제 지도에서의 지하철 역의 분포는 어떤지, R을 이용해 시각화해보겠다. 
이 작업을 위해 <a href="https://givitallugot.github.io/articles/2020-03/R-visualization-1-seoulmap">이 글</a> 을 참조했다.</p>
<h2 id="데이터셋">데이터셋</h2>
<p>먼저 <a href="http://www.gisdeveloper.co.kr/?p=2332">이 곳</a> 에서 실제 행정구역 공간 데이터를 다운로드할 수 있다. 
이 때, 참조한 글에서는 2017년 시군구 데이터를 다운로드 받으라고 했는데, 최신 데이터의 경우에는 구별로 겹치는 포인트가 있어, 코딩 중 오류가 뜨기 때문이라고 한다.</p>
<p>그 다음, <a href="https://observablehq.com/@taekie/seoul_subway_station_coordinate">이 곳</a> 에서 실제 지하철역의 위도, 경도 좌표를 다운로드할 수 있다.
여기서 주의할 점으로, 이 csv 데이터에서 5호선 양평역의 좌표가 경의중앙선 양평역의 좌표로 되어 있다는 점이다. 따라서 5호선 양평역의 좌표를 (37.525648,126.885778)으로 수정해줬다.</p>
<h2 id="코드-작성">코드 작성</h2>
<pre><code>library(ggmap)
library(ggplot2)
library(raster)
library(rgeos)
library(maptools)
library(rgdal)
Sys.setlocale(&quot;LC_ALL&quot;,&quot;.1251&quot;)</code></pre><p>먼저 위의 라이브러리들을 import한다.
그 다음, 로케일을 위의 마지막 줄 같이 변경하는데, csv를 읽을 때 한글 데이터가 섞여 있어, 제대로 읽어오지 못하거나 한글이 깨지는 현상이 발생했기 때문이다.
이에 관해 여러 가지 해결법을 시도해보았고, 내 경우에는 시스템의 로케일을 바꾸니 해결되었다.</p>
<pre><code>map &lt;- shapefile(&quot;./map/TL_SCCO_SIG.shp&quot;)
map &lt;- spTransform(map, CRSobj = CRS(&#39;+proj=longlat + ellps=WGS84 +datum=WGS84 +no_defs&#39;))
new_map &lt;- fortify(map, region = &#39;SIG_CD&#39;)
new_map$id &lt;- as.numeric(new_map$id)
metro_map &lt;- subset(new_map, new_map$id &lt;= 11740 | (new_map$id &gt;= 41000 &amp; new_map$id &lt; 42000) | (id &gt;= 28000 &amp; id &lt; 29000 &amp; long &gt;= 126))</code></pre><p>다운로드 받은 공간 데이터 중, .shp의 경로를 넣어주고, 참조한 글을 토대로 위와 같이 작성했다. 
지하철이 서울에만 운행하는 건 아니기 때문에, 적어도 경기도와 인천은 포함해야 된다. 
시도코드가 경기도는 41, 인천은 28이라는 점을 감안하여, 마지막 줄을 추가하여, 경기도와 인천을 포함시켰다.</p>
<pre><code>station &lt;- read.csv(&quot;./station_coordinate.csv&quot;,
                    na.strings = c(&quot;&quot;, &quot; &quot;, NA),
                    as.is = TRUE,
                    encoding = &quot;UTF-8&quot;)</code></pre><p>그 다음, 실제 지하철역 좌표 csv를 읽는다. 그냥 읽어오니 인코딩 관련 에러가 발생했고, encoding=UTF-8 파라미터를 추가해주니, 정상적으로 처리되었다.</p>
<pre><code>ggplot(station, aes(x=lng, y=lat)) + geom_polygon(data = metro_map, aes(x=long, y=lat, group = group), fill=&#39;white&#39;, color=&#39;black&#39;) +
  geom_point(aes(color=X.U.FEFF.line))</code></pre><p>마지막으로 ggplot 라이브러리로 실제 지하철역들의 위치를 시각화했다. geom_polygon으로 수도권 지도를 그리고, geom_point로 지하철역 위치를 점으로 찍었다.
이 때, 노선별로 다른 색깔의 점을 찍게 했다.</p>
<p>전체 코드와 시각화된 지도는 다음과 같다.</p>
<pre><code>## 1번 코드
library(ggmap)
library(ggplot2)
library(raster)
library(rgeos)
library(maptools)
library(rgdal)
Sys.setlocale(&quot;LC_ALL&quot;,&quot;.1251&quot;)

##2번 코드
map &lt;- shapefile(&quot;./map/TL_SCCO_SIG.shp&quot;)
map &lt;- spTransform(map, CRSobj = CRS(&#39;+proj=longlat + ellps=WGS84 +datum=WGS84 +no_defs&#39;))
new_map &lt;- fortify(map, region = &#39;SIG_CD&#39;)
new_map$id &lt;- as.numeric(new_map$id)
metro_map &lt;- subset(new_map, new_map$id &lt;= 11740 | (new_map$id &gt;= 41000 &amp; new_map$id &lt; 42000) | (id &gt;= 28000 &amp; id &lt; 29000 &amp; long &gt;= 126))

##3번 코드
station &lt;- read.csv(&quot;./station_coordinate.csv&quot;,
                    na.strings = c(&quot;&quot;, &quot; &quot;, NA),
                    as.is = TRUE,
                    encoding = &quot;UTF-8&quot;)

##4번 코드
ggplot(station, aes(x=lng, y=lat)) + geom_polygon(data = metro_map, aes(x=long, y=lat, group = group), fill=&#39;white&#39;, color=&#39;black&#39;) +
  geom_point(aes(color=X.U.FEFF.line))</code></pre><p><img src="https://images.velog.io/images/mt_9/post/ba8b04d0-5d69-4149-a17e-58f70ace00ec/seoul-metro.png" alt=""></p>
<h2 id="더-좋은-시각화">더 좋은 시각화</h2>
<p>막상 지도를 봐보니, 실제 노선 색깔과 다르며, 어떤 노선인지 구분이 잘 안되었다. 그래서 3번 코드와 4번 코드 사이에, 다음 코드를 추가했다.</p>
<pre><code>metro_palette &lt;- c(&#39;#003DA5&#39;,&#39;#77C4A3&#39;,&#39;#0C8E72&#39;, &#39;#0090D2&#39;,&#39;#A17800&#39;,&#39;#F5A200&#39;,&#39;#81A914&#39;,
                   &#39;#F5A200&#39;,&#39;#D4003B&#39;,&#39;#509F22&#39;,&#39;#B0CE18&#39;,&#39;#FDA600&#39;,&#39;#7CA8D5&#39;,&#39;#ED8B00&#39;,
                   &#39;#0052A4&#39;,&#39;#009D3E&#39;,&#39;#EF7C1C&#39;,&#39;#00A5DE&#39;, &#39;#996CAC&#39;,&#39;#CD7C2F&#39;,&#39;#747F00&#39;,
                   &#39;#EA545D&#39;, &#39;#BB8336&#39;)</code></pre><p>그리고 4번 코드를 다음과 같이 바꿨다.</p>
<pre><code>ggplot(station, aes(x=lng, y=lat)) + geom_polygon(data = metro_map, aes(x=long, y=lat, group = group), fill=&#39;white&#39;, color=&#39;black&#39;) +
  geom_point(aes(color=X.U.FEFF.line)) + scale_color_manual(values = metro_palette)</code></pre><p>그리고 다시 실행해주면 다음과 같은 지도가 나온다.</p>
<p><img src="https://images.velog.io/images/mt_9/post/e7f55977-6c77-4343-87aa-686d0d662fa4/metro_correct_color.png" alt=""></p>
<h2 id="서울만-시각화">서울만 시각화</h2>
<p>서울, 인천, 경기 지도를 그려보니, 서울은 작게 나오고 지하철역의 점들로 거의 보이지 않았다. 이번에는 서울만 그리고, 지하철역도 서울 부근만 나오게 했다.</p>
<p>먼저 서울만 그리게 하기 위해, 2번 코드의 마지막 줄을 다음과 같이 바꿨다.</p>
<pre><code>metro_map &lt;- subset(new_map, new_map$id &lt;= 11740)</code></pre><p>그 다음 서울 부근의 지하철역만 포함시키기 위해, 3번 코드에 다음 코드를 추가했다.</p>
<pre><code>seoul_station &lt;- subset(station, lat&gt;37.41 &amp; lat&lt;37.72 &amp; lng&gt;126.73 &amp; lng&lt;127.27)</code></pre><p>포함되지 않는 노선이 생기므로, 서울 부근 노선에 맞는 팔레트를 추가해줬다.</p>
<pre><code>seoul_metro_palette &lt;- c(&#39;#77C4A3&#39;,&#39;#0C8E72&#39;,&#39;#0090D2&#39;,&#39;#A17800&#39;,&#39;#F5A200&#39;,&#39;#81A914&#39;,&#39;#D4003B&#39;,&#39;#B0CE18&#39;,
                         &#39;#7CA8D5&#39;,&#39;#ED8B00&#39;,&#39;#0052A4&#39;,&#39;#009D3E&#39;,&#39;#EF7C1C&#39;,&#39;#00A5DE&#39;, &#39;#996CAC&#39;,
                         &#39;#CD7C2F&#39;,&#39;#747F00&#39;,&#39;#EA545D&#39;, &#39;#BB8336&#39;)
</code></pre><p>마지막으로, 4번 코드를 다음과 같이 수정했다. 이때, geom_point에 size=3을 추가하여, 점 크기도 키워줬다.</p>
<pre><code>ggplot(seoul_station, aes(x=lng, y=lat)) + geom_polygon(data = metro_map, aes(x=long, y=lat, group = group), fill=&#39;white&#39;, color=&#39;black&#39;) +
  geom_point(aes(color=X.U.FEFF.line), size=3) + scale_color_manual(values = seoul_metro_palette)</code></pre><p>그리고 실행해주면, 다음과 같은 지도가 나온다.</p>
<p><img src="https://images.velog.io/images/mt_9/post/b76c5d78-9398-4f41-9e42-bc508af14cba/well-sized-metro.png" alt=""></p>
<h2 id="느낀-점">느낀 점</h2>
<p>지도 시각화보다 한글 인코딩 처리 과정에서 가장 힘들었다. 그리고 노선별 색깔 매핑에서 노선명 정보에 맞는 색깔을 매핑시키고 싶었지만, 실패하고 순서대로 색깔을 집어넣었다.
여러 가지로, 아직은 많이 부족함을 느꼈다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[로또 당첨번호에 나온 번호는 모두 균일하게 나왔을까?]]></title>
            <link>https://velog.io/@mt_9/%EB%A1%9C%EB%98%90-%EB%8B%B9%EC%B2%A8%EB%B2%88%ED%98%B8%EC%97%90-%EB%82%98%EC%98%A8-%EB%B2%88%ED%98%B8%EB%8A%94-%EB%AA%A8%EB%91%90-%EA%B7%A0%EC%9D%BC%ED%95%98%EA%B2%8C-%EB%82%98%EC%99%94%EC%9D%84%EA%B9%8C</link>
            <guid>https://velog.io/@mt_9/%EB%A1%9C%EB%98%90-%EB%8B%B9%EC%B2%A8%EB%B2%88%ED%98%B8%EC%97%90-%EB%82%98%EC%98%A8-%EB%B2%88%ED%98%B8%EB%8A%94-%EB%AA%A8%EB%91%90-%EA%B7%A0%EC%9D%BC%ED%95%98%EA%B2%8C-%EB%82%98%EC%99%94%EC%9D%84%EA%B9%8C</guid>
            <pubDate>Wed, 13 Jan 2021 08:28:48 GMT</pubDate>
            <description><![CDATA[<h2 id="개요">개요</h2>
<p>로또를 사는 누구나 일확천금을 꿈꾸며 자신이 가지고 있는 번호들이 당첨번호와 일치하길 바란다. 혹자는 지난 회차들의 당첨번호를 보면서 많이 나온 번호들 또는 반대로 적게 나온 번호들로 로또 번호를 찍기도 한다. 이 글에서는 지난 회차들에서 가장 많이 나온 번호와 가장 적게 나온 번호를 나열하고, 이들의 확률이 균일한 편인지 검정해볼 것이다.</p>
<h2 id="데이터-수집">데이터 수집</h2>
<p>2021년 1월 13일 기준으로 현재까지 총 944번의 로또 추첨이 시행되었다. 이 당첨 번호들은 동행복권(dh.lottery.co.kr)에서 제공하는 api를 호출하여 가져와서 csv로 저장했다. 코드는 다음과 같다.</p>
<pre><code>library(httr)
library(jsonlite)
basic_url &lt;- &#39;https://www.dhlottery.co.kr/common.do?method=getLottoNumber&amp;drwNo=&#39;
df_lotto_num &lt;- NULL
for (x in 1:944) {
  req_url &lt;- paste0(basic_url, x)
  df_response &lt;- as.data.frame(fromJSON(req_url))
  df_sub &lt;- subset(df_response, select = c(&quot;drwNo&quot;, &quot;drwtNo1&quot;, &quot;drwtNo2&quot;, &quot;drwtNo3&quot;, &quot;drwtNo4&quot;, &quot;drwtNo5&quot;, &quot;drwtNo6&quot;, &quot;bnusNo&quot;))
  if(x == 1) {
    df_lotto_num &lt;- df_sub
  } else {
      df_lotto_num &lt;- rbind(df_lotto_num, df_sub)
  }
  print(paste(x , &quot;crawled.&quot;))
}
write.csv(df_lotto_num, &quot;./lotto.csv&quot;)</code></pre><h2 id="전처리">전처리</h2>
<p>lotto.csv에서 데이터를 읽어와서, 각 회차별 보너스 번호를 제외한 6개의 번호들만 추출한다. 그 다음, 1~6번째 번호별로 1부터 45까지의 번호가 얼마나 나왔는지 total() 함수로 집계한 뒤, 이를 다시 합쳐서 집계했다. 코드와 집계된 데이터프레임은 다음과 같다.</p>
<pre><code class="language-r">df_lotto_num &lt;- read.csv(&quot;./lotto.csv&quot;)
df_lotto &lt;- subset(df_lotto_num, select = c(&quot;drwtNo1&quot;, &quot;drwtNo2&quot;, &quot;drwtNo3&quot;, &quot;drwtNo4&quot;, &quot;drwtNo5&quot;, &quot;drwtNo6&quot;)) #보너스 번호는 제외
df_count &lt;- apply(df_lotto, 2, table)
df_sum &lt;- NULL
for(item in df_count) {
  if(is.null(df_sum)) {
    df_sum &lt;- as.data.frame(df_count$drwtNo1)
  } else {
    df_item &lt;- as.data.frame(item)
    df_sum &lt;- rbindlist(list(df_sum, df_item))[,lapply(.SD, sum, na.rm=TRUE), by = Var1]
  }
}
df_sum</code></pre>
<pre><code>##     Var1 Freq
##  1:    1  132
##  2:    2  124
##  3:    3  126
##  4:    4  129
##  5:    5  128
##  6:    6  118
##  7:    7  125
##  8:    8  127
##  9:    9   95
## 10:   10  131
## 11:   11  128
## 12:   12  135
## 13:   13  134
## 14:   14  133
## 15:   15  127
## 16:   16  123
## 17:   17  136
## 18:   18  137
## 19:   19  130
## 20:   20  132
## 21:   21  125
## 22:   22  107
## 23:   23  113
## 24:   24  124
## 25:   25  122
## 26:   26  125
## 27:   29  114
## 28:   35  115
## 29:   27  139
## 30:   28  117
## 31:   30  112
## 32:   31  127
## 33:   32  110
## 34:   33  131
## 35:   34  146
## 36:   36  126
## 37:   37  128
## 38:   38  127
## 39:   39  136
## 40:   40  134
## 41:   41  114
## 42:   42  123
## 43:   43  140
## 44:   44  126
## 45:   45  133
##     Var1 Freq</code></pre><h2 id="시각화">시각화</h2>
<p>먼저 로또 번호별 등장 개수를 막대 그래프로 시각화해봤다.</p>
<pre><code class="language-r">library(ggplot2)

ggplot(df_sum, aes(df_sum$Var1, df_sum$Freq))+geom_bar(stat = &#39;identity&#39;)</code></pre>
<p><img src="https://images.velog.io/images/mt_9/post/bf37ddec-5993-444c-9ed8-80d4e2aa00bf/unnamed-chunk-2-1.png" alt="">
이 등장 개수들을 오름차순으로 다시 정렬해보면,</p>
<pre><code class="language-r">ggplot(df_sum, aes(x=reorder(Var1, Freq), y=Freq))+geom_bar(stat = &#39;identity&#39;)</code></pre>
<p><img src="https://images.velog.io/images/mt_9/post/5011245c-e257-4577-9b6c-4d06421dc9fc/unnamed-chunk-3-1.png" alt="">
34번이 가장 많이 등장했고, 9번이 가장 적게 등장했음을 알 수 있다.</p>
<h2 id="검정">검정</h2>
<p>지금까지 나온 번호들이 모두 균일하게 나왔는지에 대해, 카이제곱검정을 통해 검정할 수 있다. chisq.test() 함수로 검정해 본 결과,</p>
<pre><code class="language-r">chisq.test(df_sum$Freq)</code></pre>
<pre><code>## 
##     Chi-squared test for given probabilities
## 
## data:  df_sum$Freq
## X-squared = 32.552, df = 44, p-value = 0.8985</code></pre><p>카이제곱 값(X^2)은 32.552, 자유도(df)는 44, 유의확률(p-value)는 0.8985의 값이 나왔다. 이때, 신뢰수준을 95%로 가정하면, 유의수준(alpha level)은 0.05가 되고, 이는 유의확률보다 매우 작다. 다시 말해, 95%의 신뢰 수준에서, 로또 당첨번호별 등장 확률은 균일하다고 볼 수 있다.</p>
<h2 id="결론">결론</h2>
<p>어느 번호가 더 많고, 어느 번호가 더 적게 나왔지만, 그럼에도 번호별 등장 확률은 균일하다는 검정이 나왔다. 아무래도 로또 1등이 되는 방법은 하늘에 기도하는 방법밖에 없는 모양이다 :(</p>
]]></description>
        </item>
    </channel>
</rss>