<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>j_keun.log</title>
        <link>https://velog.io/</link>
        <description>기록하며 성장하는 개발자</description>
        <lastBuildDate>Mon, 05 Oct 2026 09:39:31 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <image>
            <title>j_keun.log</title>
            <url>https://velog.velcdn.com/images/j_keun/profile/f45b9f87-ff51-4524-8b1a-26a3ab47ea20/image.jpg</url>
            <link>https://velog.io/</link>
        </image>
        <copyright>Copyright (C) 2019. j_keun.log. All rights reserved.</copyright>
        <atom:link href="https://v2.velog.io/rss/j_keun" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[LSTM 논문 과제]]></title>
            <link>https://velog.io/@j_keun/LSTM-%EB%85%BC%EB%AC%B8-%EA%B3%BC%EC%A0%9C</link>
            <guid>https://velog.io/@j_keun/LSTM-%EB%85%BC%EB%AC%B8-%EA%B3%BC%EC%A0%9C</guid>
            <pubDate>Mon, 05 Oct 2026 09:39:31 GMT</pubDate>
            <description><![CDATA[<p>논문 필기본 : <a href="https://drive.google.com/file/d/1hfVOkMnp294IpQXKHi-Hib9hCM_CIwT8/view?usp=sharing">https://drive.google.com/file/d/1hfVOkMnp294IpQXKHi-Hib9hCM_CIwT8/view?usp=sharing</a></p>
<p>논문 정리 : <a href="https://velog.io/@j_keun/LSTM-%EB%85%BC%EB%AC%B8-%EB%A6%AC%EB%B7%B0">https://velog.io/@j_keun/LSTM-%EB%85%BC%EB%AC%B8-%EB%A6%AC%EB%B7%B0</a></p>
<h3 id="1-기존-rnn에는-어떤-문제가-있었는가">1. 기존 RNN에는 어떤 문제가 있었는가?</h3>
<p>기존 RNN(Recurrent Neural Network)은 이전 시점의 은닉 상태를 다음 시점으로 전달하므로, 순서가 있는 데이터를 처리할 수 있다. 예를 들어 문장, 음성, 주가처럼 앞의 정보가 뒤의 결과에 영향을 주는 데이터에 적합하다.</p>
<p>하지만 학습 과정에서 BPTT(Backporpagation Through Time)를 사용하면 기울기 소실과 기울기 폭주 문제가 발생한다.</p>
<ul>
<li>기울기 소실: 과거 시점으로 오차를 전달할수록 값이 계속 작아져 거의 0이 되는 현상</li>
<li>기울기 폭주: 반대로 기울기가 계속 커져 학습이 불안정해지는 현상</li>
</ul>
<p>특히 일반 RNN은 많은 시간 단계를 거쳐 오차를 전달할 때, 활성화 함수의 미분값과 가중치가 계속 곱해진다. 이 곱의 결과가 1보다 작으면 기울기가 빠르게 사라지고, 1보다 크면 지나치게 커진다.</p>
<p>➡ 기존 RNN은 최근 정보는 비교적 잘 반영하지만, 오래전의 중요한 정보를 현재 예측에 활용하는데 어려움이 있다.</p>
<h3 id="2-긴-sequence를-학습할-때-왜-문제가-발생하는가">2. 긴 Sequence를 학습할 때 왜 문제가 발생하는가?</h3>
<p>긴 시퀀스에서 현재 시점의 오차가 과거의 중요한 입력까지 전달되어야 한다. 하지만 현재와 과거 사이에 시간 단계가 많아질수록, 역전파 과정의 기울기 곱셈도 길어진다.</p>
<p>예를 들어 다음 문장을 보자.</p>
<blockquote>
<p>&quot;나는 한국에서 태어나 자랐고, 현재 미국에서 일하고 이싿. 내 모국어는 ___이다.&quot;</p>
</blockquote>
<p>정답을 예측하려면 문장 앞부분의 &quot;한국&quot;이라는 정보를 기억해야 한다. 하지만 그 사이에 많은 단어가 들어가면, 일반 RNN은 앞부분의 정보를 학습 신호로 유지하기 어렵다.</p>
<p>이를 수식적으로 보면, 과거 시점으로 전달되는 기울기는 여러 시점의 미분값과 가중치가 연속적으로 곱해진 결과다.</p>
<pre><code>현재 오차
    ➡ t-1의 가중치 * 미분값
    ➡ t-2의 가중치 * 미분값
    ➡ ...
    ➡ 과거 시점</code></pre><p>따라서 시간 간격이 갈수록 다음 문제가 발생한다.</p>
<ul>
<li>과거의 중요한 정보가 학습 과정에서 사라진다.</li>
<li>현재의 예측 오류가 과거 입력의 가중치까지 도달하지 못한다.</li>
<li>모델이 무엇을 오래 기억해야 하는지 학습하기 어렵다.</li>
<li>잡음이나 관련 없는 입력이 많은 경우 중요한 정보를 유지하기 더 어렵다.</li>
</ul>
<p>➡ 핵심은 &quot;시퀀스가 길다&quot;는 것만이 문제가 아니라, 오래전 정보가 현재 결과에 반드시 필요하지만 중간에 관련 없는 정보가 많이 있는 상황이다.</p>
<h3 id="3-lstm은-기존-rnn의-문제를-어떤-아이디어로-해결하려-했는가">3. LSTM은 기존 RNN의 문제를 어떤 아이디어로 해결하려 했는가?</h3>
<p>LSTM은 메모리 셀 내부에 오차가 일정하게 흐를 수 있는 경로를 만들고, 그 경로를 게이트로 제어하는 방식으로 문제를 해결했다.</p>
<ol>
<li>메모리 셀 내부에 정보를 오래 유지한다.</li>
<li>CEC(Constant Error Carousel)를 통해 기울기가 급격히 사라지거나 커지는 것을 줄인다.</li>
<li>입력 게이트와 출력 게이트를 통해 기억할 정보와 사용할 정보를 선택한다.</li>
</ol>
<p>기존 RNN에서는 새 입력이 계속 들어오면 이전 정보가 쉽게 덮어써지거나, 기억된 정보가 계속 출력되어 다른 계산을 방해할 수 있다. LSTM은 이를 막기 위해 게이트를 둔다.</p>
<pre><code class="language-tex">새 입력
  ↓
입력 게이트: 이 정보를 저장할지 결정
  ↓
메모리 셀: 필요한 정보를 유지
  ↓
출력 게이트: 현재 시점에 꺼내 쓸지 결정
  ↓
다음 계층 또는 최종 예측</code></pre>
<p>➡ LSTM은 필요한 정보는 유지하고, 필요 없는 정보는 저장하거나 출력하지 않도록 학습하는 구조다.</p>
<h3 id="4-lstm의-핵심-구조는-무엇인가">4. LSTM의 핵심 구조는 무엇인가?</h3>
<p><strong>4-1 메모리 셀</strong></p>
<p>메모리 셀은 정보를 장기간 보관하기 위한 내부 저장 공간이다. 일반 RNN의 은닉 상태보다 장기 정보를 유지하기 위한 구조가 강화되어 있다.</p>
<p>논문에서는 메모리 셀의 내부 상태는 이전 상태에 새 후보 정보를 더하는 형태로 표현</p>
<pre><code class="language-tex">현재 셀 상태
= 이전 셀 상태
+ 입력 게이트가 허용한 새 정보</code></pre>
<p>➡ 이전 상태가 매 시점마다 강하게 변형되지 않으며, 긴 시간 동안 정보와 학습 신호를 유지할 수 있다.</p>
<p><strong>4-2 CEC(Constant Error Carousel)</strong></p>
<p>CEC는 메모리 셀 안에서 자신으로 이어지는 연결의 가중치를 <code>1.0</code>으로 유지하는 구조다 .</p>
<p>이 경로에서는 오차가 반복적으로 작아지거나 커지지 않고, 비교적 일정하기 순환할 수 있다.</p>
<p>논문에서는 장기 의존성 학습의 핵심 기반이 CEC라고 말한다.</p>
<pre><code class="language-tex">셀 상태(t-1)
     ↓
   CEC 연결
     ↓
셀 상태(t)</code></pre>
<p>➡ CEC는 먼 과거의 입력까지 학습 신호가 도달할 수 있도록 돕는 통로이다.</p>
<p><strong>4-3 입력 게이트</strong></p>
<p>입력 게이트는 새로 들어온 정보를 메모리 셀이 저장할지 결정한다.</p>
<ul>
<li>값이 1에 가까우면: 현재 입력을 메모리에 반영</li>
<li>값이 0에 가까우면: 현재 입력을 무시하고 기존 메모리를 보호</li>
</ul>
<p>예) 문장에서 인물의 국적 정보를 기억해야 한다면, 국적이 등장하는 시점에는 입력게이트를 열고, 이후의 관련 없는 단어가 들어올 때는 게이트를 닫아 기존 기억이 덮어써지지 않도록 할 수 있다.</p>
<p><strong>4-4 출력 게이트</strong></p>
<p>출력 게이트는 메모리 셀에 저장된 정보를 현재 시점의 계산에 사용할지 결정한다.</p>
<ul>
<li>값이 1에 가까우면: 메모리 내용을 외부로 전달한다.</li>
<li>값이 0에 가까우면: 기억은 유지하되 현재 출력에는 사용하지 않는다.</li>
</ul>
<p>➡ 입력 게이트가 &quot;언제 저장할지&quot;를 결정한다면, 출력 게이트는 &quot;언제 꺼내 사용할지&quot;를 결정한다고 생각하면 된다.</p>
<h3 id="5-논문에서-말하는-장기-의존성long-term-dependency이란-무엇인가">5. 논문에서 말하는 장기 의존성(Long-Term Dependency)이란 무엇인가?</h3>
<p>장기 의존성이란 시퀀스의 앞부분에 등장한 정보가 많은 시간 단계가 지난 뒤의 결과를 결정하는 관계를 말한다.</p>
<p>예)</p>
<pre><code class="language-tex">시점 1: 중요한 신호 A 등장
시점 2~99: 잡음 또는 관련 없는 입력
시점 100: 신호 A를 이용해 결과 예측</code></pre>
<p>모델은 시점 1의 신호 A를 시점 100까지 기억해야 한다. 중간에 입력이 많고 잡음이 섞여 있어도 중요한 정보를 잃으면 안된다.</p>
<p>논문에서는 다음 조건일때 장기 의존성 문제가 더 발생한다고 설명한다.</p>
<ul>
<li>중요한 입력과 결과 사이의 시간 간격이 길다.</li>
<li>중간에 관련 없는 입력이나 잡음이 많다.</li>
<li>중요한 정보가 어느 위치에 나타날지 미리 알기 어렵다.</li>
<li>현재 예측을 위해 여러 과거 정보를 함께 기억해야 한다.</li>
</ul>
<p>➡ LSTM은 CEC와 게이트를  이용해 이러한 장기 의존성을 학습하도록 설계</p>
<h3 id="6-논문을-읽고-새롭게-알게-된-내용">6. 논문을 읽고 새롭게 알게 된 내용</h3>
<p><strong>6-1 LSTM의 핵심은 오차 흐름 제어다.</strong></p>
<p>논문을 읽기 전에는 LSTM이 단순히 정보를 오래 저장하는 모델이라고 생각했다. 하지만 논문에서 다루는 더 핵심적인 문제는, 긴 시퀀스에서 과거까지 학습 신호가 전달되지 않는 기울기 소실 문제였다.</p>
<p>➡ LSTM은 기억 장치이면서 동시에 긴 시간 간격에서도 오차를 전달할 수 있도록 설계된 학습 구조임을 알게되었다.</p>
<p><strong>6-2 LSTM도 모든 문제를 해결하지는 못한다.</strong></p>
<p>논문에서는 LSTM의 한계도 설명해준다. 예를 들어,멀리 떨어진 두 입력의 XOR처럼 두 정보를 모두 기억해야만 오차가 줄어드는 비분해적 문제는 학습하기 어렵다. 하나의 정보만 기억한 상태에서는 성능이 좋아지지 않으므로, 기울기 기반 학습이 어느 방향으로 가야 하는지 찾기 어렵다고 설명한다.</p>
<p>또한 정밀하게 시간 단계를 새어야 하는 문제에는 별도의 카운팅 메커니즘이 필요하다고 설명한다.</p>
<h3 id="7-기타-추가-내용">7. 기타 추가 내용</h3>
<p><strong>7-1 LSTM의 장점</strong></p>
<ul>
<li>긴 시간 간격의 중요한 정보를 학습할 수 있다.</li>
<li>잡음이 있는 시퀀스에서도 필요한 정보를 선택적으로 유지할 수 있다.</li>
<li>시계열 예측, 음성 처리, 자연어 처리 등 순서가 중요한 데이터에 적용할 수 있다.</li>
<li>BPTT와 유사한 수준의 계산 복잡도로 학습할 수 있다.</li>
<li>관련 정보가 시퀀스의 어느 위치에 등장하더라도 일반화 할 수 있다.</li>
</ul>
<p><strong>7-2 LSTM의 한계</strong></p>
<ul>
<li>모든 장기 의존성 문제를 자동으로 해결하지는 못한다.</li>
<li>매우 정확한 시간 단계 계산에는 한계가 있을 수 있다.</li>
<li>게이트와 메모리 셀 때문에 일반 RNN보다 구조와 파라미터가 복잡해질 수 있다.</li>
</ul>
]]></description>
        </item>
        <item>
            <title><![CDATA[LSTM 논문 리뷰]]></title>
            <link>https://velog.io/@j_keun/LSTM-%EB%85%BC%EB%AC%B8-%EB%A6%AC%EB%B7%B0</link>
            <guid>https://velog.io/@j_keun/LSTM-%EB%85%BC%EB%AC%B8-%EB%A6%AC%EB%B7%B0</guid>
            <pubDate>Mon, 05 Oct 2026 09:39:03 GMT</pubDate>
            <description><![CDATA[<blockquote>
<p><strong>원 논문</strong>: Hochreiter, S. &amp; Schmidhuber, J. (1997). <em>Long Short-Term Memory</em>. Neural Computation, 9(8), 1735-1780.  </p>
</blockquote>
<h2 id="목차">목차</h2>
<ol>
<li>Introduction</li>
<li>Previous Work</li>
<li>Constant Error Backprop</li>
<li>The Concept of Long Short-Term Memory</li>
<li>Experiments</li>
<li>Discussion</li>
<li>Conclusion</li>
<li>Acknowledgments</li>
</ol>
<hr>
<h1 id="1-introduction-문제의-출발점">1. Introduction: 문제의 출발점</h1>
<p>RNN(Recurrent Neural Network)은 현재 입력과 이전 상태를 함께 사용하므로, 문장·음성·시계열처럼 순서가 있는 데이터를 처리할 수 있다. 그러나 현재 시점의 오류를 오래전 입력까지 전달해야 할 때 학습이 매우 느려지거나 실패할 수 있다.</p>
<p>논문은 그 원인을 <strong>불충분하고 점차 감소하는 오차 역전파(error backflow)</strong>에서 찾는다. 먼 과거로 갈수록 학습 신호가 약해져, 과거 정보가 현재 결과에 미치는 영향을 모델이 배우기 어렵다는 뜻이다.</p>
<h2 id="논문이-해결하려는-목표">논문이 해결하려는 목표</h2>
<ul>
<li>매우 긴 시간 간격의 입력과 결과를 연결한다.</li>
<li>중간에 잡음이 많아도 중요한 정보를 유지한다.</li>
<li>짧은 시간 간격의 학습 능력을 잃지 않는다.</li>
<li>전체 BPTT보다 실용적인 계산 비용을 유지한다.</li>
</ul>
<h2 id="장기-의존성의-예시">장기 의존성의 예시</h2>
<pre><code class="language-text">시점 1: 중요한 정보 X
시점 2~999: 잡음 또는 관련 없는 입력
시점 1000: X를 활용해야 정답을 낼 수 있는 출력</code></pre>
<p>이 상황에서 모델은 X를 오래 기억해야 할 뿐 아니라, 현재 오류가 시점 1의 파라미터까지 전달되도록 학습해야 한다.</p>
<hr>
<h1 id="2-previous-work-기존-접근법의-한계">2. Previous Work: 기존 접근법의 한계</h1>
<p>2절에서 저자들은 당시의 순환 신경망 학습 방법을 검토한다.</p>
<table>
<thead>
<tr>
<th>접근법</th>
<th>기본 아이디어</th>
<th>긴 시간 지연에서의 한계</th>
</tr>
</thead>
<tbody><tr>
<td>BPTT</td>
<td>시간을 펼친 뒤 역전파</td>
<td>그래디언트 소실·폭주가 누적됨</td>
</tr>
<tr>
<td>RTRL</td>
<td>순환 가중치의 영향을 실시간 계산</td>
<td>계산 비용이 크고 긴 지연에 취약함</td>
</tr>
<tr>
<td>Elman Network</td>
<td>이전 은닉 상태를 문맥으로 사용</td>
<td>중요한 과거 정보의 학습 신호가 약해짐</td>
</tr>
<tr>
<td>고정 시간 상수 기반 구조</td>
<td>일정 기간 상태를 유지</td>
<td>정보 저장·검색을 상황에 맞게 제어하기 어려움</td>
</tr>
</tbody></table>
<p>논문의 핵심 주장은 단순하다. 기존 방법은 짧은 시간 간격에서는 동작할 수 있지만, 최소 시간 간격이 긴 문제에서는 과거 입력까지 충분한 학습 신호를 전달하기 어렵다.</p>
<hr>
<h1 id="3-constant-error-backprop-일정한-오차-흐름의-필요성">3. Constant Error Backprop: 일정한 오차 흐름의 필요성</h1>
<h2 id="31-exponentially-decaying-error">3.1 Exponentially Decaying Error</h2>
<p>원 논문은 시간 <code>t</code>에서 유닛 <code>j</code>로 역전파되는 지역 오차를 다음처럼 표현한다.</p>
<p>$$
\delta_j(t)=f&#39;<em>j(net_j(t))\sum_k \delta_k(t+1)w</em>{kj}
$$</p>
<table>
<thead>
<tr>
<th>기호</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td>$\delta_j(t)$</td>
<td>시간 <code>t</code>에서 유닛 <code>j</code>에 전달되는 오차 신호</td>
</tr>
<tr>
<td>$net_j(t)$</td>
<td>유닛 <code>j</code>의 순입력</td>
</tr>
<tr>
<td>$f&#39;_j$</td>
<td>활성화 함수의 미분값</td>
</tr>
<tr>
<td>$w_{kj}$</td>
<td>유닛 <code>j</code>에서 다음 유닛 <code>k</code>로 향하는 가중치</td>
</tr>
<tr>
<td>$\sum_k$</td>
<td>다음 시간 단계의 연결에서 돌아오는 오차의 합</td>
</tr>
</tbody></table>
<p>시간을 거꾸로 이동할수록 미분값과 가중치가 계속 곱해진다. 이 곱이 평균적으로 1보다 작으면 오차는 지수적으로 감소하고, 1보다 크면 폭주한다.</p>
<pre><code class="language-text">현재 출력 오류
  → 현재 은닉 상태
  → 이전 은닉 상태
  → 더 이전 은닉 상태
  → 먼 과거의 입력</code></pre>
<p>따라서 긴 시퀀스에서는 “과거 정보가 중요하다”는 사실을 알더라도, 그 정보를 저장해야 하는 가중치까지 학습 신호가 닿지 않을 수 있다.</p>
<h2 id="32-constant-error-flow-단순한-접근">3.2 Constant Error Flow: 단순한 접근</h2>
<p>저자들은 자기 연결을 가진 선형 유닛 <code>j</code>를 생각한다. 시간 <code>t</code>에서 오차 흐름이 일정하려면 다음 조건이 필요하다.</p>
<p>$$
f&#39;<em>j(net_j(t))w</em>{jj}=1.0
$$</p>
<p>가장 단순한 방법은 항등 함수와 자기 연결 가중치 <code>1.0</code>을 사용하는 것이다.</p>
<p>$$
f_j(x)=x, \qquad w_{jj}=1.0
$$</p>
<p>이때 상태는 다음처럼 유지된다.</p>
<p>$$
y_j(t+1)=y_j(t)
$$</p>
<p>논문은 이 자기 연결 구조를 <strong>CEC(Constant Error Carousel)</strong>라고 부른다.</p>
<h3 id="단순-cec만으로는-부족한-이유">단순 CEC만으로는 부족한 이유</h3>
<table>
<thead>
<tr>
<th>문제</th>
<th>왜 문제가 되는가</th>
<th>필요한 해결책</th>
</tr>
</thead>
<tbody><tr>
<td>입력 가중치 충돌</td>
<td>중요한 입력은 저장해야 하지만, 이후 잡음은 기존 기억을 덮어쓰지 못하게 해야 함</td>
<td>입력을 선택하는 장치</td>
</tr>
<tr>
<td>출력 가중치 충돌</td>
<td>기억은 필요할 때 사용해야 하지만, 필요 없을 때 다른 유닛을 방해하면 안 됨</td>
<td>출력을 선택하는 장치</td>
</tr>
</tbody></table>
<p>이 두 문제를 해결하기 위해 4절에서 입력 게이트와 출력 게이트가 등장한다.</p>
<hr>
<h1 id="4-the-concept-of-long-short-term-memory">4. The Concept of Long Short-Term Memory</h1>
<h2 id="41-memory-cells-and-gate-units">4.1 Memory Cells and Gate Units</h2>
<p>원 논문의 LSTM 메모리 셀은 다음 요소로 구성된다.</p>
<ul>
<li>중심 선형 유닛과 고정 자기 연결: CEC</li>
<li>입력 게이트 $in_j$</li>
<li>출력 게이트 $out_j$</li>
<li>내부 상태 $s_{c_j}(t)$</li>
</ul>
<h3 id="게이트-활성화">게이트 활성화</h3>
<p>$$
y_{in_j}(t)=f_{in_j}(net_{in_j}(t))
$$</p>
<p>$$
y_{out_j}(t)=f_{out_j}(net_{out_j}(t))
$$</p>
<p>게이트의 순입력은 이전 시점 네트워크 유닛의 출력으로부터 계산된다.</p>
<p>$$
net_{in_j}(t)=\sum_u w_{in_j,u}y_u(t-1)
$$</p>
<p>$$
net_{out_j}(t)=\sum_u w_{out_j,u}y_u(t-1)
$$</p>
<p>여기서 $u$는 입력 유닛, 게이트 유닛, 메모리 셀, 일반 은닉 유닛 등이 될 수 있다.</p>
<h3 id="메모리-셀-상태와-출력">메모리 셀 상태와 출력</h3>
<p>원 논문에서 메모리 셀의 내부 상태는 다음처럼 갱신된다.</p>
<p>$$
s_{c_j}(0)=0
$$</p>
<p>$$
s_{c_j}(t)=s_{c_j}(t-1)+y_{in_j}(t)g(net_{c_j}(t)) \quad (t&gt;0)
$$</p>
<p>메모리 셀의 외부 출력은 다음과 같다.</p>
<p>$$
y_{c_j}(t)=y_{out_j}(t)h(s_{c_j}(t))
$$</p>
<h3 id="수식의-직관적-해석">수식의 직관적 해석</h3>
<pre><code class="language-text">이전 셀 상태 s(t-1)
        +
입력 게이트가 허용한 새 정보
        ↓
현재 셀 상태 s(t)
        ↓
출력 게이트가 허용한 만큼만 외부로 전달
        ↓
셀 출력 y(t)</code></pre>
<table>
<thead>
<tr>
<th>구성 요소</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td>CEC</td>
<td>긴 시간 동안 내부 상태와 오차 흐름을 유지하는 통로</td>
</tr>
<tr>
<td>입력 게이트</td>
<td>새 입력을 기억에 기록할지 결정</td>
</tr>
<tr>
<td>출력 게이트</td>
<td>저장된 기억을 현재 다른 유닛에 공개할지 결정</td>
</tr>
<tr>
<td>$g$</td>
<td>셀에 들어갈 후보 정보를 압축하는 미분 가능한 함수</td>
</tr>
<tr>
<td>$h$</td>
<td>셀 상태를 외부 출력에 적합한 값으로 조절하는 함수</td>
</tr>
</tbody></table>
<h2 id="42-왜-게이트-유닛인가">4.2 왜 게이트 유닛인가?</h2>
<p>입력 게이트는 관련 없는 입력이 CEC 내부 기억을 덮어쓰지 못하게 한다. 출력 게이트는 현재 필요하지 않은 기억이 다른 유닛을 교란하지 못하게 한다.</p>
<pre><code class="language-text">입력 게이트가 닫힘
→ 기존 기억 보호

출력 게이트가 닫힘
→ 기억은 유지하지만 현재 계산에는 노출하지 않음</code></pre>
<p>이로써 LSTM은 다음을 분리한다.</p>
<pre><code class="language-text">무엇을 기록할지     → 입력 게이트
어떻게 오래 유지할지 → CEC
언제 사용할지       → 출력 게이트</code></pre>
<h2 id="43-현대-lstm과의-차이">4.3 현대 LSTM과의 차이</h2>
<p>1997년 원 논문에는 Forget Gate가 없다. 현대 LSTM은 이후 Forget Gate를 추가해 이전 셀 상태를 얼마나 유지할지 직접 조절한다.</p>
<table>
<thead>
<tr>
<th>구분</th>
<th>1997년 원 논문</th>
<th>현대 표준 LSTM</th>
</tr>
</thead>
<tbody><tr>
<td>핵심 구조</td>
<td>CEC, 입력 게이트, 출력 게이트</td>
<td>셀 상태와 입력·망각·출력 게이트</td>
</tr>
<tr>
<td>기존 기억 처리</td>
<td>CEC를 유지하고 입력 게이트로 덮어쓰기를 억제</td>
<td>Forget Gate가 이전 상태 유지 비율을 직접 결정</td>
</tr>
<tr>
<td>Forget Gate</td>
<td>없음</td>
<td>있음</td>
</tr>
</tbody></table>
<p>현대 표준 LSTM의 대표적인 셀 상태 수식은 다음과 같다.</p>
<p>$$
c_t=f_t \odot c_{t-1}+i_t \odot \tilde{c}_t
$$</p>
<p>이 수식은 현대 구현을 설명할 때 사용해야 하며, 1997년 원 논문의 수식과 혼동하지 않아야 한다.</p>
<hr>
<h1 id="5-experiments-lstm은-어떤-과제를-풀었는가">5. Experiments: LSTM은 어떤 과제를 풀었는가?</h1>
<p>논문은 실험 과제를 난도 순으로 배치한다.</p>
<table>
<thead>
<tr>
<th>실험</th>
<th>검증하려는 능력</th>
<th>핵심 입력 형태</th>
</tr>
</thead>
<tbody><tr>
<td>5.1 Embedded Reber Grammar</td>
<td>출력 게이트의 효과, 기존 기법과의 비교</td>
<td>이산 기호 문법</td>
</tr>
<tr>
<td>5.2 Noise-Free and Noisy Sequences</td>
<td>긴 지연과 대량의 방해 기호</td>
<td>이산 기호 + 잡음</td>
</tr>
<tr>
<td>5.3 Noise and Signal on Same Channel</td>
<td>신호와 잡음이 한 채널에 섞인 경우</td>
<td>연속값 단일 채널</td>
</tr>
<tr>
<td>5.4 Adding Problem</td>
<td>긴 시간 동안 연속값 저장·덧셈</td>
<td>실수값 + 마커</td>
</tr>
<tr>
<td>5.5 Multiplication Problem</td>
<td>비적분적 연산</td>
<td>실수값 + 마커</td>
</tr>
<tr>
<td>5.6 Temporal Order</td>
<td>멀리 떨어진 기호의 순서 기억</td>
<td>이산 기호 분류</td>
</tr>
<tr>
<td>5.7 Summary</td>
<td>실험 조건의 종합 비교</td>
<td>표 정리</td>
</tr>
</tbody></table>
<h2 id="51-experiment-1-embedded-reber-grammar">5.1 Experiment 1: Embedded Reber Grammar</h2>
<h3 id="과제-목표">과제 목표</h3>
<p>Embedded Reber Grammar는 정해진 문법 규칙으로 생성되는 기호열을 한 글자씩 입력받고, 다음 기호를 예측하는 벤치마크다.</p>
<pre><code class="language-text">입력: B → T → ... → P → ...
목표: 매 시점마다 다음에 올 수 있는 기호 예측</code></pre>
<p>문자열의 마지막 직전 기호를 올바르게 예측하려면 초반의 두 번째 기호 <code>T</code> 또는 <code>P</code>를 기억해야 한다.</p>
<h3 id="왜-이-과제를-사용했는가">왜 이 과제를 사용했는가?</h3>
<ul>
<li>최소 시간 지연이 약 9단계인 경계 사례라서 기존 방식도 완전히 실패하지는 않는다.</li>
<li>기존 RNN 연구에서 널리 사용된 벤치마크라 비교가 가능하다.</li>
<li>출력 게이트가 왜 필요한지 보여 주기에 적합하다.</li>
</ul>
<h3 id="비교-모델과-대표-결과">비교 모델과 대표 결과</h3>
<p>원 논문 표 1의 대표 결과를 블로그용으로 다시 정리하면 다음과 같다.</p>
<table>
<thead>
<tr>
<th>방법</th>
<th align="right">가중치 수</th>
<th align="right">성공률</th>
<th align="right">성공까지 필요한 시퀀스 수</th>
</tr>
</thead>
<tbody><tr>
<td>RTRL</td>
<td align="right">약 170</td>
<td align="right">일부 성공</td>
<td align="right">173,000</td>
</tr>
<tr>
<td>Elman Network</td>
<td align="right">약 435</td>
<td align="right">0%</td>
<td align="right">200,000 초과</td>
</tr>
<tr>
<td>Recurrent Cascade-Correlation</td>
<td align="right">약 119-198</td>
<td align="right">50%</td>
<td align="right">182,000</td>
</tr>
<tr>
<td>LSTM, 학습률 0.1</td>
<td align="right">264</td>
<td align="right">100%</td>
<td align="right">39,740</td>
</tr>
<tr>
<td>LSTM, 학습률 0.5</td>
<td align="right">276</td>
<td align="right">100%</td>
<td align="right">8,440</td>
</tr>
</tbody></table>
<h3 id="해석">해석</h3>
<p>LSTM은 단순히 문법을 외우는 것이 아니라, 필요한 초기 기호는 저장하되 그 기억이 이후의 쉬운 문법 전이 계산을 방해하지 않게 해야 한다. 이때 출력 게이트가 기억의 공개 시점을 조절한다.</p>
<h2 id="52-experiment-2-noise-free-and-noisy-sequences">5.2 Experiment 2: Noise-Free and Noisy Sequences</h2>
<h3 id="task-2a-잡음-없는-긴-시간-지연">Task 2a: 잡음 없는 긴 시간 지연</h3>
<p>입력은 다음 두 시퀀스 중 하나다.</p>
<p>$$
(x,a_1,a_2,\ldots,a_{p-1},x)
$$</p>
<p>$$
(y,a_1,a_2,\ldots,a_{p-1},y)
$$</p>
<p>모델은 매 시점 다음 기호를 예측한다. 마지막 기호 <code>x</code> 또는 <code>y</code>를 맞히려면 맨 처음의 <code>x</code> 또는 <code>y</code>를 <code>p</code> 단계 동안 기억해야 한다.</p>
<h3 id="task-2b-규칙성이-약한-잡음-시퀀스">Task 2b: 규칙성이 약한 잡음 시퀀스</h3>
<p>중간 기호가 규칙적으로 정렬되지 않고, <code>x</code> 또는 <code>y</code>와 무관한 기호가 무작위로 등장한다. 마지막의 예측 가능한 <code>x</code> 또는 <code>y</code>를 맞히기 위해, 모델은 초반의 핵심 기호를 유지해야 한다.</p>
<h3 id="task-2c-매우-긴-지연과-다수의-방해-기호">Task 2c: 매우 긴 지연과 다수의 방해 기호</h3>
<p>가장 어려운 과제다. 시퀀스는 시작 표식 <code>b</code>, 핵심 기호 <code>x</code> 또는 <code>y</code>, 다수의 무작위 방해 기호, 트리거 <code>e</code>, 최종 답 <code>x</code> 또는 <code>y</code>로 구성된다.</p>
<pre><code class="language-text">b → x 또는 y → 방해 기호 다수 → e(트리거) → x 또는 y</code></pre>
<p>모델은 트리거 <code>e</code>가 등장했을 때 초반의 핵심 기호가 무엇이었는지 출력해야 한다. 방해 기호는 임의 위치에 들어가며, 긴 경우 최소 시간 지연이 1,000단계에 이른다.</p>
<h3 id="원-논문-표-3-task-2c-대표-결과">원 논문 표 3: Task 2c 대표 결과</h3>
<table>
<thead>
<tr>
<th align="right">시간 지연 <code>q+1</code></th>
<th align="right">방해 기호 종류 <code>p</code></th>
<th align="right">방해 기호당 기대 등장 횟수 <code>q/p</code></th>
<th align="right">가중치 수</th>
<th align="right">성공까지 시퀀스 수</th>
</tr>
</thead>
<tbody><tr>
<td align="right">51</td>
<td align="right">50</td>
<td align="right">1</td>
<td align="right">364</td>
<td align="right">30,000</td>
</tr>
<tr>
<td align="right">101</td>
<td align="right">100</td>
<td align="right">1</td>
<td align="right">664</td>
<td align="right">31,000</td>
</tr>
<tr>
<td align="right">201</td>
<td align="right">200</td>
<td align="right">1</td>
<td align="right">1,264</td>
<td align="right">33,000</td>
</tr>
<tr>
<td align="right">501</td>
<td align="right">500</td>
<td align="right">1</td>
<td align="right">3,064</td>
<td align="right">38,000</td>
</tr>
<tr>
<td align="right">1,001</td>
<td align="right">1,000</td>
<td align="right">1</td>
<td align="right">6,064</td>
<td align="right">49,000</td>
</tr>
<tr>
<td align="right">1,001</td>
<td align="right">100</td>
<td align="right">10</td>
<td align="right">664</td>
<td align="right">135,000</td>
</tr>
<tr>
<td align="right">1,001</td>
<td align="right">50</td>
<td align="right">20</td>
<td align="right">364</td>
<td align="right">203,000</td>
</tr>
</tbody></table>
<h3 id="해석-1">해석</h3>
<p>시간 지연이 길어지는 것 자체보다, 같은 종류의 방해 기호가 자주 반복되어 기억을 흔드는 상황에서 학습 시간이 더 증가한다. 그럼에도 논문은 최소 시간 지연 1,000단계 설정을 해결했다고 보고한다.</p>
<h2 id="53-experiment-3-noise-and-signal-on-same-channel">5.3 Experiment 3: Noise and Signal on Same Channel</h2>
<h3 id="task-3a-two-sequence-problem">Task 3a: Two-Sequence Problem</h3>
<p>하나의 실수 입력 채널만 사용한다. 두 클래스 중 하나가 동일 확률로 선택되며, 첫 <code>N</code>개 값이 클래스 정보를 갖는다.</p>
<pre><code class="language-text">클래스 1: 처음 N개 값 = +1.0
클래스 2: 처음 N개 값 = -1.0
이후 값: 평균 0, 분산 0.2의 가우시안 잡음</code></pre>
<p>시퀀스 마지막에서 클래스 1은 <code>1.0</code>, 클래스 2는 <code>0.0</code>을 출력해야 한다.</p>
<h3 id="task-3b-중요한-신호도-잡음에-오염">Task 3b: 중요한 신호도 잡음에 오염</h3>
<p>3a와 달리 클래스 정보를 담은 첫 <code>N</code>개 값에도 잡음이 섞인다. 즉, 모델은 명확한 <code>+1</code>과 <code>-1</code>을 읽는 것이 아니라, 잡음이 섞인 값에서 클래스 신호를 추출해야 한다.</p>
<h3 id="task-3c-잡음이-있는-실수-목표값의-조건부-기대값-학습">Task 3c: 잡음이 있는 실수 목표값의 조건부 기대값 학습</h3>
<p>3c는 단순 추측으로 풀기 어려운 형태로 수정된다.</p>
<pre><code class="language-text">클래스 1의 잡음 없는 목표값: 0.2
클래스 2의 잡음 없는 목표값: 0.8
실제 학습 목표: 목표값에 가우시안 잡음을 더한 값</code></pre>
<p>모델은 잡음이 섞인 목표값을 그대로 외우는 대신, 입력이 주어졌을 때의 <strong>조건부 기대값</strong>을 학습해야 한다.</p>
<h3 id="원-논문-표-6-task-3c-대표-결과">원 논문 표 6: Task 3c 대표 결과</h3>
<table>
<thead>
<tr>
<th align="right">최소 시퀀스 길이 <code>T</code></th>
<th align="right">핵심 신호 개수 <code>N</code></th>
<th align="right">가중치 수</th>
<th align="right">오분류 비율</th>
<th align="right">기대값과의 평균 차이</th>
</tr>
</thead>
<tbody><tr>
<td align="right">100</td>
<td align="right">3</td>
<td align="right">102</td>
<td align="right">0.00558</td>
<td align="right">0.014</td>
</tr>
<tr>
<td align="right">100</td>
<td align="right">1</td>
<td align="right">102</td>
<td align="right">0.00441</td>
<td align="right">0.012</td>
</tr>
</tbody></table>
<h3 id="해석-2">해석</h3>
<p>이 실험은 신호와 잡음이 같은 채널에 있어도 LSTM이 장기 기억을 형성할 수 있는지, 그리고 단순한 분류를 넘어 연속적인 기대값을 예측할 수 있는지 확인한다.</p>
<h2 id="54-experiment-4-adding-problem">5.4 Experiment 4: Adding Problem</h2>
<h3 id="과제-입력">과제 입력</h3>
<p>각 입력 요소는 두 성분의 쌍이다.</p>
<p>$$
(X_t,M_t)
$$</p>
<table>
<thead>
<tr>
<th>성분</th>
<th>설정</th>
</tr>
</thead>
<tbody><tr>
<td>$X_t$</td>
<td>$[-1,1]$에서 무작위로 뽑은 실수</td>
</tr>
<tr>
<td>$M_t$</td>
<td><code>1.0</code>, <code>0.0</code>, <code>-1.0</code> 중 하나인 마커</td>
</tr>
</tbody></table>
<p>시퀀스마다 정확히 두 위치가 <code>M_t=1.0</code>으로 표시된다. 모델은 그 두 위치의 실수 $X_1$, $X_2$를 오랫동안 기억한 뒤, 시퀀스 마지막에서 아래 목표를 출력해야 한다.</p>
<p>$$
target=0.5+\frac{X_1+X_2}{4.0}
$$</p>
<p><code>-1.0</code> 마커는 시퀀스의 시작과 끝을 표시하며, 나머지 대부분의 위치는 <code>0.0</code>이다.</p>
<h3 id="왜-중요한-과제인가">왜 중요한 과제인가?</h3>
<ul>
<li>이산 기호가 아닌 연속 실수값을 저장해야 한다.</li>
<li>두 개의 값을 오랫동안 보존해야 한다.</li>
<li>표시 위치가 고정되지 않는다.</li>
<li>마지막에 덧셈 계산까지 수행해야 한다.</li>
</ul>
<h3 id="원-논문-표-7-adding-problem-결과">원 논문 표 7: Adding Problem 결과</h3>
<table>
<thead>
<tr>
<th align="right">최소 길이 <code>T</code></th>
<th align="right">최소 지연</th>
<th align="right">가중치 수</th>
<th align="right">2,560개 테스트 중 오답</th>
<th align="right">성공까지 시퀀스 수</th>
</tr>
</thead>
<tbody><tr>
<td align="right">100</td>
<td align="right">50</td>
<td align="right">93</td>
<td align="right">1</td>
<td align="right">74,000</td>
</tr>
<tr>
<td align="right">500</td>
<td align="right">250</td>
<td align="right">93</td>
<td align="right">0</td>
<td align="right">209,000</td>
</tr>
<tr>
<td align="right">1,000</td>
<td align="right">500</td>
<td align="right">93</td>
<td align="right">1</td>
<td align="right">853,000</td>
</tr>
</tbody></table>
<h3 id="해석-3">해석</h3>
<p>논문은 LSTM이 긴 시간 동안 연속값을 큰 손실 없이 저장하고, 분산 표현과 수치 연산이 필요한 과제를 해결할 수 있음을 보인다.</p>
<h2 id="55-experiment-5-multiplication-problem">5.5 Experiment 5: Multiplication Problem</h2>
<h3 id="과제-목표-1">과제 목표</h3>
<p>입력 구성과 두 개의 마커 위치는 Adding Problem과 유사하다. 차이는 마지막 목표값이 덧셈이 아니라 곱셈이라는 점이다.</p>
<p>$$
target=X_1 \times X_2
$$</p>
<h3 id="왜-덧셈-다음에-곱셈을-검증했는가">왜 덧셈 다음에 곱셈을 검증했는가?</h3>
<p>CEC는 이전 상태에 새 값을 더하는 구조이므로, “덧셈 과제는 CEC의 누적 특성 때문에 쉽게 풀린 것 아니냐”는 의문이 생길 수 있다.</p>
<p>곱셈은 단순 누적으로 바로 해결할 수 없는 비적분적 연산이다. 이 과제는 LSTM이 단순 누적기가 아니라, 장기 기억을 바탕으로 더 복잡한 연산도 학습할 수 있는지 검증한다.</p>
<h2 id="56-experiment-6-temporal-order">5.6 Experiment 6: Temporal Order</h2>
<h3 id="task-6a-멀리-떨어진-두-기호의-순서">Task 6a: 멀리 떨어진 두 기호의 순서</h3>
<p>시퀀스는 <code>E</code>로 시작하고 <code>B</code>라는 트리거 기호로 끝난다. 그 사이에는 주로 <code>{a,b,c,d}</code>에서 무작위로 선택된 방해 기호가 나타난다. 다만 두 위치에는 <code>X</code> 또는 <code>Y</code>가 등장한다.</p>
<table>
<thead>
<tr>
<th>위치</th>
<th>범위</th>
</tr>
</thead>
<tbody><tr>
<td>전체 시퀀스 길이</td>
<td>100~110</td>
</tr>
<tr>
<td>첫 번째 중요 기호 $t_1$</td>
<td>10~20</td>
</tr>
<tr>
<td>두 번째 중요 기호 $t_2$</td>
<td>50~60</td>
</tr>
</tbody></table>
<p>정답 클래스는 <code>X</code>, <code>Y</code>의 시간 순서로 결정된다.</p>
<table>
<thead>
<tr>
<th>중요 기호 순서</th>
<th>클래스</th>
</tr>
</thead>
<tbody><tr>
<td>X, X</td>
<td>Q</td>
</tr>
<tr>
<td>X, Y</td>
<td>R</td>
</tr>
<tr>
<td>Y, X</td>
<td>S</td>
</tr>
<tr>
<td>Y, Y</td>
<td>U</td>
</tr>
</tbody></table>
<h3 id="task-6b-멀리-떨어진-세-기호의-순서">Task 6b: 멀리 떨어진 세 기호의 순서</h3>
<p>6b에서는 중요 기호가 세 개가 되며, 가능한 순서 조합은 8개다.</p>
<table>
<thead>
<tr>
<th>위치</th>
<th>범위</th>
</tr>
</thead>
<tbody><tr>
<td>첫 번째 중요 기호 $t_1$</td>
<td>10~20</td>
</tr>
<tr>
<td>두 번째 중요 기호 $t_2$</td>
<td>33~43</td>
</tr>
<tr>
<td>세 번째 중요 기호 $t_3$</td>
<td>66~76</td>
</tr>
</tbody></table>
<p>예를 들어 <code>X,Y,X</code>와 <code>Y,X,X</code>는 서로 다른 클래스로 분류해야 한다. 오류 신호는 시퀀스 마지막에만 제공된다.</p>
<h3 id="원-논문-표-9-temporal-order-결과">원 논문 표 9: Temporal Order 결과</h3>
<table>
<thead>
<tr>
<th>과제</th>
<th align="right">가중치 수</th>
<th align="right">2,560개 테스트 중 오답</th>
<th align="right">성공까지 시퀀스 수</th>
</tr>
</thead>
<tbody><tr>
<td>Task 6a: 중요 기호 2개</td>
<td align="right">156</td>
<td align="right">1</td>
<td align="right">31,390</td>
</tr>
<tr>
<td>Task 6b: 중요 기호 3개</td>
<td align="right">308</td>
<td align="right">2</td>
<td align="right">571,100</td>
</tr>
</tbody></table>
<h3 id="해석-4">해석</h3>
<p>이 실험은 LSTM이 단순히 어떤 기호가 등장했는지를 기억하는 수준을 넘어, 멀리 떨어진 여러 입력의 <strong>시간적 순서</strong>까지 추출할 수 있음을 검증한다.</p>
<h2 id="57-summary-of-experimental-conditions">5.7 Summary of Experimental Conditions</h2>
<p>5.7절은 앞선 실험의 조건을 원 논문 표 10, 표 11로 종합한다. 세부 표에는 과제 번호, 최소 길이, 지연, 셀 블록 수, 입력·출력 수, 가중치 수, 게이트 편향, 함수, 학습률 등이 기록되어 있다.</p>
<p>블로그 관점에서 중요한 비교 기준은 다음과 같다.</p>
<table>
<thead>
<tr>
<th>실험군</th>
<th>핵심 난이도</th>
<th>검증한 LSTM 능력</th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td>짧은 지연 문법</td>
<td>출력 게이트의 필요성</td>
</tr>
<tr>
<td>2</td>
<td>매우 긴 지연, 다수의 방해 기호</td>
<td>장기 기억과 잡음 억제</td>
</tr>
<tr>
<td>3</td>
<td>신호·잡음이 같은 채널</td>
<td>신호 추출과 기대값 예측</td>
</tr>
<tr>
<td>4</td>
<td>연속값 2개 저장과 덧셈</td>
<td>연속값 장기 보존</td>
</tr>
<tr>
<td>5</td>
<td>연속값 2개 저장과 곱셈</td>
<td>비적분적 연산</td>
</tr>
<tr>
<td>6</td>
<td>여러 기호의 시간 순서</td>
<td>다중 장기 의존성</td>
</tr>
</tbody></table>
<hr>
<h1 id="6-discussion-한계와-장점">6. Discussion: 한계와 장점</h1>
<h2 id="61-limitations-of-lstm">6.1 Limitations of LSTM</h2>
<h3 id="강하게-지연된-xor-문제">강하게 지연된 XOR 문제</h3>
<p>멀리 떨어진 두 입력의 XOR를 계산하는 문제는 두 정보 중 하나만 기억해서는 오차를 줄일 수 없다. 따라서 쉬운 하위 문제부터 점진적으로 해결하기 어려운 비분해적 문제다.</p>
<h3 id="추가-게이트에-따른-구조-증가">추가 게이트에 따른 구조 증가</h3>
<p>메모리 셀 블록마다 입력 게이트와 출력 게이트가 추가된다. 원 논문은 완전 연결 구조에서 가중치 수가 증가할 수 있음을 인정하지만, 실험에서는 비교 모델과 비슷한 가중치 수를 사용하려 했다.</p>
<h3 id="정밀한-시간-카운팅의-한계">정밀한 시간 카운팅의 한계</h3>
<p>99단계 전과 100단계 전을 반드시 구분해야 하는 문제에는 별도의 카운팅 장치가 필요할 수 있다. 반면 3단계 전과 11단계 전처럼 비교적 큰 시간 차이를 구분하는 것은 가능하다고 설명한다.</p>
<h2 id="62-advantages-of-lstm">6.2 Advantages of LSTM</h2>
<ul>
<li>CEC를 통해 매우 긴 시간 간격의 오차 흐름을 유지할 수 있다.</li>
<li>잡음, 연속값, 분산 표현을 다룰 수 있다.</li>
<li>중요한 입력 위치가 달라져도 일반화할 수 있다.</li>
<li>입력·출력 게이트 편향, 학습률 등 여러 설정 범위에서 비교적 안정적으로 동작한다.</li>
<li>가중치당·시간 단계당 갱신 복잡도가 본질적으로 BPTT와 같은 $O(1)$ 수준이다.</li>
</ul>
<hr>
<h1 id="7-conclusion-논문의-핵심-결론">7. Conclusion: 논문의 핵심 결론</h1>
<p>논문은 메모리 셀 내부의 CEC가 일정한 오차 흐름을 제공하고, 이것이 긴 시간 간격을 연결하는 기반이라고 결론 내린다.</p>
<pre><code class="language-text">CEC
→ 정보와 오차 흐름을 장기간 유지

입력 게이트
→ 관련 없는 입력이 기억을 덮어쓰지 못하게 보호

출력 게이트
→ 현재 필요하지 않은 기억이 다른 유닛을 방해하지 않도록 보호</code></pre>
<p>LSTM의 핵심은 “모든 것을 오래 기억한다”가 아니다. 필요한 정보를 선택적으로 기록하고, 보존하고, 필요한 시점에만 출력하는 구조라는 점이다.</p>
<hr>
<h1 id="최종-정리">최종 정리</h1>
<p>1997년 LSTM 논문은 긴 시퀀스에서 RNN이 실패하는 원인을 그래디언트 흐름 관점에서 분석하고, 이를 CEC와 게이트라는 구조로 해결하려 한 연구다.</p>
<pre><code class="language-text">기존 RNN의 그래디언트 소실·폭주
        ↓
CEC로 일정한 오차 흐름 확보
        ↓
입력 게이트로 기록 제어
        ↓
출력 게이트로 사용 제어
        ↓
긴 지연·잡음·연속값·시간 순서 실험으로 검증</code></pre>
<p>현대 LSTM은 Forget Gate 등을 추가하며 발전했지만, 긴 시간 간격의 정보를 학습하기 위해 <strong>기억 경로와 게이트 제어를 분리한다</strong>는 원 논문의 핵심 아이디어는 그대로 이어지고 있다.</p>
<h2 id="참고-자료">참고 자료</h2>
<ul>
<li>Hochreiter, S., &amp; Schmidhuber, J. (1997). <a href="https://direct.mit.edu/neco/article/9/8/1735/6109/Long-Short-Term-Memory">Long Short-Term Memory</a>. <em>Neural Computation</em>, 9(8), 1735-1780.</li>
<li><a href="https://web.stanford.edu/class/psych209/Readings/HochreiterSchmidthuber97LSTM.pdf">원 논문 PDF</a></li>
</ul>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 카운트 다운]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%B9%B4%EC%9A%B4%ED%8A%B8-%EB%8B%A4%EC%9A%B4</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%B9%B4%EC%9A%B4%ED%8A%B8-%EB%8B%A4%EC%9A%B4</guid>
            <pubDate>Sun, 04 Oct 2026 20:32:02 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>목표 점수 <code>target</code>을 정확히 만들기 위해 다트를 던진다.</p>
<p>최적의 방법은 다음 우선순위로 결정한다.</p>
<ol>
<li>던지는 다트 수를 최소화한다.</li>
<li>다트 수가 같으면 싱글 또는 불을 맞힌 횟수를 최대화한다.</li>
</ol>
<p><code>[최소 다트 수, 최대 싱글 또는 불 횟수]</code>를 반환한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p><code>dp[score]</code>를 정확히 <code>score</code>점을 만드는 최적의 결과라고 정의한다.</p>
<pre><code class="language-text">dp[score] = [최소 다트 수, 그때의 최대 싱글 또는 불 횟수]</code></pre>
<p>마지막 다트로 얻은 점수가 <code>points</code>라면, 그 직전에는 <code>score - points</code>점을 만들었어야 한다.</p>
<pre><code class="language-text">dp[score] = dp[score - points] + 마지막 다트 정보</code></pre>
<p>각 점수에서 가능한 모든 마지막 다트를 비교하면 된다.</p>
<h2 id="다트-점수와-싱글-횟수">다트 점수와 싱글 횟수</h2>
<p>한 번의 다트로 만들 수 있는 점수는 다음과 같다.</p>
<ul>
<li>싱글: 1 ~ 20점, 싱글 또는 불 횟수 1 증가</li>
<li>더블: 2 ~ 40점, 횟수 증가 없음</li>
<li>트리플: 3 ~ 60점, 횟수 증가 없음</li>
<li>불: 50점, 싱글 또는 불 횟수 1 증가</li>
</ul>
<p>같은 점수를 만들 수 있는 방법이 여러 개여도 모두 비교 기준에 따라 처리한다. 예를 들어 6점은 싱글 6, 더블 3, 트리플 2로 만들 수 있고, 다트 수가 같으므로 싱글 6을 선택한다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>한 번에 만들 수 있는 점수와 싱글 또는 불 횟수 증가량을 만든다.</li>
<li><code>dp[0] = [0, 0]</code>으로 초기화한다.</li>
<li>1점부터 <code>target</code>점까지 순서대로 확인한다.</li>
<li>가능한 마지막 다트를 하나씩 적용해 이전 점수의 결과를 갱신한다.</li>
<li>다트 수가 더 적거나, 다트 수가 같으면서 싱글 또는 불 횟수가 더 많을 때만 갱신한다.</li>
</ol>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">def solution(target):
    throws = [(50, 1)]

    for number in range(1, 21):
        throws.append((number, 1))       # 싱글
        throws.append((number * 2, 0))   # 더블
        throws.append((number * 3, 0))   # 트리플

    infinity = target + 1
    dp = [(infinity, -1) for _ in range(target + 1)]
    dp[0] = (0, 0)

    for score in range(1, target + 1):
        for points, single_or_bull in throws:
            if score &lt; points:
                continue

            previous_darts, previous_singles = dp[score - points]

            if previous_darts == infinity:
                continue

            candidate_darts = previous_darts + 1
            candidate_singles = previous_singles + single_or_bull
            current_darts, current_singles = dp[score]

            if (
                candidate_darts &lt; current_darts
                or (
                    candidate_darts == current_darts
                    and candidate_singles &gt; current_singles
                )
            ):
                dp[score] = (candidate_darts, candidate_singles)

    return list(dp[target])</code></pre>
<h2 id="예시">예시</h2>
<h3 id="target--21">target = 21</h3>
<p>7 트리플로 한 번에 21점을 만들 수 있다.</p>
<pre><code class="language-text">다트 수: 1
싱글 또는 불 횟수: 0
결과: [1, 0]</code></pre>
<h3 id="target--58">target = 58</h3>
<p>불 50점과 싱글 8점을 사용하면 2번의 다트로 58점을 만든다.</p>
<pre><code class="language-text">불 50 + 싱글 8 = 58
다트 수: 2
싱글 또는 불 횟수: 2
결과: [2, 2]</code></pre>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p>한 번의 다트로 만들 수 있는 경우는 불을 포함해 61개다.</p>
<p><code>T</code>를 목표 점수라고 하자.</p>
<ul>
<li>시간 복잡도: <code>O(T * 61) = O(T)</code></li>
<li>공간 복잡도: <code>O(T)</code></li>
</ul>
<p><code>target &lt;= 100,000</code>이므로 충분히 빠르게 동작한다.</p>
<h2 id="정리">정리</h2>
<p>각 점수마다 최소 다트 수와 최대 싱글 또는 불 횟수를 함께 저장한다. 다트 수를 우선 비교하고, 동점일 때만 싱글 또는 불 횟수를 비교하면 문제의 우선순위를 그대로 DP에 반영할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 연속 펄스 부분 수열의 합]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%97%B0%EC%86%8D-%ED%8E%84%EC%8A%A4-%EB%B6%80%EB%B6%84-%EC%88%98%EC%97%B4%EC%9D%98-%ED%95%A9</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%97%B0%EC%86%8D-%ED%8E%84%EC%8A%A4-%EB%B6%80%EB%B6%84-%EC%88%98%EC%97%B4%EC%9D%98-%ED%95%A9</guid>
            <pubDate>Sun, 04 Oct 2026 20:22:09 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>수열의 연속 부분 수열 하나를 선택하고, <code>[1, -1, 1, -1, ...]</code> 또는 <code>[-1, 1, -1, 1, ...]</code> 형태의 펄스 수열을 곱한다.</p>
<p>만들어진 연속 펄스 부분 수열의 합 중 최댓값을 구한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p>펄스 수열의 부호 패턴은 두 가지뿐이다.</p>
<pre><code class="language-text">패턴 1:  1, -1,  1, -1, ...
패턴 2: -1,  1, -1,  1, ...</code></pre>
<p>원본 수열의 인덱스 기준으로 두 패턴을 미리 곱한 변환 수열을 생각한다.</p>
<pre><code class="language-text">sequence:  2,  3, -6,  1, ...
변환 1:    2, -3, -6, -1, ...
변환 2:   -2,  3,  6,  1, ...</code></pre>
<p>어떤 연속 구간을 선택하더라도 두 변환 수열 중 하나의 연속 부분합으로 표현된다. 따라서 두 변환 수열 각각에서 <strong>최대 연속 부분합</strong>을 구하고, 둘 중 큰 값을 선택하면 된다.</p>
<p>최대 연속 부분합은 카데인 알고리즘으로 한 번의 순회에 구할 수 있다.</p>
<h2 id="카데인-알고리즘">카데인 알고리즘</h2>
<p>현재 값이 <code>value</code>일 때, 현재 위치에서 끝나는 최대 연속 부분합은 두 선택지 중 큰 값이다.</p>
<pre><code class="language-python">current = max(value, current + value)</code></pre>
<ul>
<li>이전 구간을 버리고 현재 값부터 새로 시작한다.</li>
<li>이전 구간 뒤에 현재 값을 이어 붙인다.</li>
</ul>
<p>두 변환 패턴의 값을 같은 순회에서 함께 계산할 수 있다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>짝수 인덱스에서는 첫 번째 패턴의 부호를 <code>1</code>, 홀수 인덱스에서는 <code>-1</code>로 둔다.</li>
<li>두 번째 패턴은 첫 번째 패턴의 부호를 반대로 한 값이다.</li>
<li>각 변환 값에 카데인 알고리즘을 적용한다.</li>
<li>두 패턴에서 얻은 최대 부분합 중 큰 값을 반환한다.</li>
</ol>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">def solution(sequence):
    current_first = 0
    current_second = 0
    answer = -float(&quot;inf&quot;)

    for index, value in enumerate(sequence):
        # 전체 수열 기준 [1, -1, 1, -1, ...] 패턴
        first_value = value if index % 2 == 0 else -value
        second_value = -first_value

        # 각 패턴에서 현재 인덱스로 끝나는 최대 연속 부분합
        current_first = max(first_value, current_first + first_value)
        current_second = max(second_value, current_second + second_value)

        answer = max(answer, current_first, current_second)

    return answer</code></pre>
<h2 id="예시">예시</h2>
<p><code>sequence = [2, 3, -6, 1, 3, -1, 2, 4]</code>에서 인덱스 1부터 3까지의 구간을 선택해 보자.</p>
<pre><code class="language-text">원본 구간:      [3, -6, 1]
펄스 수열:      [1, -1, 1]
곱한 결과:      [3, 6, 1]
합:             10</code></pre>
<p>이 구간은 두 번째 변환 패턴에서의 연속 부분합으로 계산되며, 전체 최댓값은 <code>10</code>이다.</p>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p><code>N</code>을 수열 길이라고 하자.</p>
<p>수열을 한 번만 순회하고, 두 개의 현재 부분합만 유지한다.</p>
<ul>
<li>시간 복잡도: <code>O(N)</code></li>
<li>공간 복잡도: <code>O(1)</code></li>
</ul>
<h2 id="정리">정리</h2>
<p>펄스 수열의 시작 부호는 두 가지뿐이다. 두 부호 패턴으로 수열을 변환한 뒤 각각의 최대 연속 부분합을 구하면, 모든 연속 펄스 부분 수열의 합 중 최댓값을 선형 시간에 구할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 최고속도]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%B5%9C%EA%B3%A0%EC%86%8D%EB%8F%84</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%B5%9C%EA%B3%A0%EC%86%8D%EB%8F%84</guid>
            <pubDate>Fri, 02 Oct 2026 15:08:05 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>도로는 수평 또는 수직 선분이며, 교차하거나 만나는 지점에서 서로 연결된다. 각 도로의 중앙에는 제한 속도를 가진 카메라가 있다.</p>
<p>1번 도시에서 출발해 각 도시까지 일정한 속도로 이동할 때, 경로에서 통과하는 모든 카메라 제한 속도를 만족하는 최대 속도를 구한다. 카메라를 전혀 지나지 않는 경로가 있으면 제한 없이 이동할 수 있으므로 <code>0</code>을 반환한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<h3 id="1-도로를-중요한-지점에서-분할해-그래프로-만든다">1. 도로를 중요한 지점에서 분할해 그래프로 만든다</h3>
<p>도로 위에서 이동 방향을 바꾸거나 제한 속도가 발생할 수 있는 지점은 다음과 같다.</p>
<ul>
<li>도로의 양 끝점</li>
<li>도시 위치</li>
<li>도로의 교차점 또는 접점</li>
<li>각 도로 중앙의 카메라 위치</li>
</ul>
<p>각 도로에 포함된 중요 지점을 도로 방향으로 정렬하고, 이웃한 지점을 간선으로 연결한다. 같은 좌표의 지점은 하나의 그래프 정점으로 합친다.</p>
<h3 id="2-카메라는-정점-제한으로-처리한다">2. 카메라는 정점 제한으로 처리한다</h3>
<p>카메라가 있는 지점을 지나면 제한 속도를 지켜야 한다. 카메라 정점과 연결되는 모든 간선의 제한 속도를 해당 카메라 제한 속도 이하로 설정하면, 그 정점을 통과하는 모든 경로에 제한을 적용할 수 있다.</p>
<p>한 좌표에 카메라가 여러 개라면 가장 낮은 제한 속도를 사용한다.</p>
<h3 id="3-widest-path를-사용한다">3. Widest Path를 사용한다</h3>
<p>어떤 경로로 이동 가능한 최고 일정 속도는 그 경로에서 만나는 간선 제한 속도의 최솟값이다. 도시마다 이 최솟값을 최대화해야 하므로 최단 경로가 아니라 <strong>Widest Path</strong> 문제다.</p>
<p>다익스트라와 같은 방식으로 우선순위 큐를 사용하되, 거리가 작은 정점 대신 현재까지의 제한 속도가 큰 정점을 먼저 꺼낸다.</p>
<pre><code class="language-text">next_speed = min(current_speed, edge_limit)</code></pre>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">from heapq import heappop, heappush


def solution(city, road):
    INF = 10**18
    road_count = len(road)

    # (x1, y1, x2, y2, limit) 형태로 도로를 저장한다.
    roads = [tuple(info) for info in road]
    road_points = [[] for _ in range(road_count)]
    camera_limit = {}

    def is_horizontal(current_road):
        return current_road[1] == current_road[3]

    def contains(current_road, x, y):
        x1, y1, x2, y2, _ = current_road
        return x1 &lt;= x &lt;= x2 and y1 &lt;= y &lt;= y2

    def add_point(road_index, point):
        road_points[road_index].append(point)

    # 도로의 끝점과 중앙 카메라 위치를 추가한다.
    for index, (x1, y1, x2, y2, limit) in enumerate(roads):
        add_point(index, (x1, y1))
        add_point(index, (x2, y2))

        camera = ((x1 + x2) // 2, (y1 + y2) // 2)
        add_point(index, camera)
        camera_limit[camera] = min(camera_limit.get(camera, INF), limit)

    # 도시가 포함된 모든 도로에 도시 위치를 분할 지점으로 추가한다.
    for x, y in city:
        for index, current_road in enumerate(roads):
            if contains(current_road, x, y):
                add_point(index, (x, y))

    # 모든 도로 쌍의 교차점 또는 접점을 찾는다.
    for first in range(road_count):
        x1, y1, x2, y2, _ = roads[first]
        first_is_horizontal = is_horizontal(roads[first])

        for second in range(first + 1, road_count):
            a1, b1, a2, b2, _ = roads[second]
            second_is_horizontal = is_horizontal(roads[second])
            point = None

            if first_is_horizontal != second_is_horizontal:
                horizontal = roads[first] if first_is_horizontal else roads[second]
                vertical = roads[second] if first_is_horizontal else roads[first]

                hx1, hy, hx2, _, _ = horizontal
                vx, vy1, _, vy2, _ = vertical

                if hx1 &lt;= vx &lt;= hx2 and vy1 &lt;= hy &lt;= vy2:
                    point = (vx, hy)

            elif first_is_horizontal and y1 == b1:
                left = max(x1, a1)
                right = min(x2, a2)
                if left == right:
                    point = (left, y1)

            elif not first_is_horizontal and x1 == a1:
                bottom = max(y1, b1)
                top = min(y2, b2)
                if bottom == top:
                    point = (x1, bottom)

            if point is not None:
                add_point(first, point)
                add_point(second, point)

    # 동일 좌표를 하나의 그래프 정점으로 합친다.
    point_to_index = {}

    def get_index(point):
        if point not in point_to_index:
            point_to_index[point] = len(point_to_index)
        return point_to_index[point]

    # 먼저 모든 중요 지점에 그래프 번호를 부여한다.
    for points in road_points:
        for point in points:
            get_index(point)

    graph = [[] for _ in range(len(point_to_index))]

    # 각 도로의 이웃한 중요 지점을 연결한다.
    for index, points in enumerate(road_points):
        current_road = roads[index]

        if is_horizontal(current_road):
            ordered_points = sorted(set(points), key=lambda point: point[0])
        else:
            ordered_points = sorted(set(points), key=lambda point: point[1])

        for left, right in zip(ordered_points, ordered_points[1:]):
            # 카메라 정점을 지나기 위해 지켜야 하는 가장 낮은 제한 속도
            limit = min(
                camera_limit.get(left, INF),
                camera_limit.get(right, INF),
            )

            left_index = get_index(left)
            right_index = get_index(right)
            graph[left_index].append((right_index, limit))
            graph[right_index].append((left_index, limit))

    start = get_index(tuple(city[0]))
    best_speed = [-1] * len(graph)
    best_speed[start] = INF
    heap = [(-INF, start)]

    # 제한 속도가 큰 경로부터 확정하는 Widest Path 탐색
    while heap:
        negative_speed, current = heappop(heap)
        current_speed = -negative_speed

        if current_speed &lt; best_speed[current]:
            continue

        for next_node, limit in graph[current]:
            next_speed = min(current_speed, limit)

            if next_speed &gt; best_speed[next_node]:
                best_speed[next_node] = next_speed
                heappush(heap, (-next_speed, next_node))

    answer = []
    for x, y in city[1:]:
        speed = best_speed[get_index((x, y))]
        answer.append(0 if speed == INF else speed)

    return answer</code></pre>
<h2 id="정당성-설명">정당성 설명</h2>
<p>그래프의 정점은 도시, 도로 끝점, 교차점, 카메라 위치를 모두 포함한다. 따라서 실제 도로에서 이동 방향을 바꾸거나 카메라 제한이 적용될 수 있는 모든 지점이 그래프에 표현된다. 각 도로에서 이웃한 중요 지점을 연결했으므로 그래프의 경로와 실제 도로 이동 경로는 서로 대응한다.</p>
<p>카메라 정점에 인접한 간선의 제한 속도는 그 지점의 최소 카메라 제한 속도 이하로 설정했다. 그러므로 그래프 경로의 최소 간선 제한 속도는 실제 이동 경로에서 지켜야 하는 가장 낮은 카메라 제한 속도와 같다.</p>
<p>Widest Path 탐색은 현재까지 확보한 제한 속도가 가장 큰 정점을 먼저 처리한다. 어떤 도시를 향하는 경로의 값은 경로 간선 제한 속도의 최솟값이며, 전이식 <code>min(current_speed, edge_limit)</code>은 이 값을 정확히 계산한다. 더 큰 값이 발견될 때만 갱신하므로, 최종 <code>best_speed</code>는 1번 도시에서 각 정점으로 갈 수 있는 최고 일정 속도다.</p>
<p>따라서 카메라를 지나지 않는 경우에는 <code>INF</code>를 <code>0</code>으로 바꾸고, 그 외에는 <code>best_speed</code>를 반환하면 요구한 답을 얻는다.</p>
<h2 id="복잡도-분석">복잡도 분석</h2>
<p>도로 수를 <code>M</code>, 중요 지점과 그래프 간선 수를 각각 <code>V</code>, <code>E</code>라고 하자.</p>
<ul>
<li>도시를 도로에 추가: <code>O(N * M)</code></li>
<li>도로 교차점 검사: <code>O(M^2)</code></li>
<li>도로별 지점 정렬: <code>O(V log V)</code></li>
<li>Widest Path 탐색: <code>O((V + E) log V)</code></li>
</ul>
<p>전체 시간 복잡도는 <code>O(M^2 + (V + E) log V)</code>이며, 공간 복잡도는 <code>O(V + E)</code>다.</p>
<h2 id="주의할-점">주의할 점</h2>
<ul>
<li>도로가 교차하는 좌표뿐 아니라, 도시와 카메라 위치에서도 도로를 분할해야 한다.</li>
<li>같은 카메라 위치에 여러 도로의 카메라가 겹치면 제한 속도는 최솟값을 사용한다.</li>
<li>카메라를 지나지 않는 경로는 제한 속도가 없으므로 내부적으로 큰 값 <code>INF</code>로 표현한 뒤 최종 결과에서 <code>0</code>으로 바꾼다.</li>
<li>최단 거리 다익스트라와 달리, 우선순위 큐에서는 현재 제한 속도가 <strong>큰</strong> 상태를 먼저 꺼낸다.</li>
</ul>
<h2 id="마무리">마무리</h2>
<p>기하 문제처럼 보이지만 도로를 교차점과 특수 지점에서 분할하면 그래프 문제가 된다. 이후에는 경로의 최소 제한 속도를 최대화하는 Widest Path를 적용해 도시별 최고 일정 속도를 구할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[72일차] Denoising AutoEncoder, VAE와 GAN]]></title>
            <link>https://velog.io/@j_keun/72%EC%9D%BC%EC%B0%A8-Denoising-AutoEncoder-VAE%EC%99%80-GAN</link>
            <guid>https://velog.io/@j_keun/72%EC%9D%BC%EC%B0%A8-Denoising-AutoEncoder-VAE%EC%99%80-GAN</guid>
            <pubDate>Fri, 02 Oct 2026 15:03:18 GMT</pubDate>
            <description><![CDATA[<p>기본 AutoEncoder는 입력을 잠재 표현으로 바꾼 뒤 원본과 비슷하게 복원한다. 여기서 입력에 노이즈를 추가하거나 잠재 공간에 제약을 주면 다른 목적의 모델로 확장할 수 있다.</p>
<p>이번에는 <strong>Denoising AutoEncoder</strong>, <strong>Sparse AutoEncoder</strong>, <strong>VAE(Variational AutoEncoder)</strong>를 살펴본다. 이어서 생성자와 판별자가 경쟁하며 새로운 데이터를 만드는 <strong>GAN(Generative Adversarial Network)</strong>을 구현한다.</p>
<hr>
<h2 id="1-latent-representation-다시-확인하기">1. Latent Representation 다시 확인하기</h2>
<p>학습이 끝난 AutoEncoder에서는 Encoder만 따로 사용할 수 있다.</p>
<pre><code class="language-python">x, _ = next(iter(test_loader))
x = x[:8].to(DEVICE)

with torch.no_grad():
    z = model.encoder(x)

print(&quot;입력 Shape:&quot;, x.shape)
print(&quot;Latent Shape:&quot;, z.shape)
print(&quot;원본 값 개수:&quot;, 1 * 28 * 28)
print(&quot;Latent 값 개수:&quot;, z[0].numel())</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">입력 Shape: torch.Size([8, 1, 28, 28])
Latent Shape: torch.Size([8, 32, 7, 7])
원본 값 개수: 784
Latent 값 개수: 1568</code></pre>
<p>공간 해상도는 <code>28 × 28</code>에서 <code>7 × 7</code>로 줄었지만 Channel이 <code>1</code>에서 <code>32</code>로 늘어 전체 값의 개수는 더 많다.</p>
<p>AutoEncoder의 Latent Representation이 반드시 원본보다 적은 수의 값을 가져야 하는 것은 아니다. 합성곱 AutoEncoder는 공간 크기를 줄이는 대신 Channel을 늘려 원본 복원에 유용한 Feature Map으로 변환하기도 한다.</p>
<p>다만 <strong>차원 축소</strong>가 목적이라면 전체 원소 수가 실제로 작은 Bottleneck을 설계해야 한다.</p>
<hr>
<h2 id="2-denoising-autoencoder">2. Denoising AutoEncoder</h2>
<p>Denoising AutoEncoder는 노이즈가 섞인 데이터를 입력으로 받고, 깨끗한 원본을 Target으로 사용한다.</p>
<pre><code class="language-text">노이즈가 섞인 이미지
        ↓
      Encoder
        ↓
Latent Representation
        ↓
      Decoder
        ↓
깨끗하게 복원된 이미지</code></pre>
<p>일반 AutoEncoder와 학습 목표를 비교하면 다음과 같다.</p>
<table>
<thead>
<tr>
<th>모델</th>
<th>입력</th>
<th>Target</th>
</tr>
</thead>
<tbody><tr>
<td>기본 AutoEncoder</td>
<td>깨끗한 이미지</td>
<td>같은 깨끗한 이미지</td>
</tr>
<tr>
<td>Denoising AutoEncoder</td>
<td>노이즈 이미지</td>
<td>깨끗한 원본 이미지</td>
</tr>
</tbody></table>
<h3 id="노이즈-추가하기">노이즈 추가하기</h3>
<pre><code class="language-python">def add_noise(x, noise_factor=0.35):
    noisy = x + noise_factor * torch.randn_like(x)
    return torch.clamp(noisy, 0.0, 1.0)</code></pre>
<ul>
<li><code>torch.randn_like(x)</code>는 <code>x</code>와 같은 Shape의 표준정규분포 노이즈를 만든다.</li>
<li><code>noise_factor</code>는 노이즈의 강도를 조절한다.</li>
<li><code>torch.clamp()</code>는 픽셀값이 <code>[0, 1]</code> 범위를 벗어나지 않도록 제한한다.</li>
</ul>
<pre><code class="language-python">x, _ = next(iter(train_loader))
noisy_x = add_noise(x[:8])</code></pre>
<p><img src="https://velog.velcdn.com/images/j_keun/post/ac1b6dcf-46f9-4696-a5a8-07056cd81748/image.png" alt=""></p>
<p>위쪽은 원본, 아래쪽은 노이즈를 추가한 이미지다.</p>
<h3 id="학습하기">학습하기</h3>
<pre><code class="language-python">denoising_model = ConvAutoencoder().to(DEVICE)
criterion = nn.MSELoss()
optimizer = optim.AdamW(
    denoising_model.parameters(),
    lr=1e-3,
)

def train_denoising_autoencoder(
    model,
    loader,
    criterion,
    optimizer,
    device,
    epochs=5,
):
    for epoch in range(epochs):
        model.train()
        running_loss = 0.0

        for clean_x, _ in loader:
            clean_x = clean_x.to(device)
            noisy_x = add_noise(clean_x)

            _, restored_x = model(noisy_x)
            loss = criterion(restored_x, clean_x)

            optimizer.zero_grad()
            loss.backward()
            optimizer.step()

            running_loss += (
                loss.item() * clean_x.size(0)
            )

        epoch_loss = (
            running_loss / len(loader.dataset)
        )

        print(
            f&quot;Epoch {epoch + 1:02d}/{epochs} | &quot;
            f&quot;Loss: {epoch_loss:.6f}&quot;
        )</code></pre>
<p>중요한 부분은 다음 한 줄이다.</p>
<pre><code class="language-python">loss = criterion(restored_x, clean_x)</code></pre>
<p>모델에는 <code>noisy_x</code>를 입력하지만 Loss는 복원 결과와 깨끗한 <code>clean_x</code> 사이에서 계산한다.</p>
<p>노트북의 실행 결과는 다음과 같다.</p>
<pre><code class="language-text">Epoch 01/5 | Loss: 0.029357
Epoch 02/5 | Loss: 0.008080
Epoch 03/5 | Loss: 0.007619
Epoch 04/5 | Loss: 0.007417
Epoch 05/5 | Loss: 0.007292</code></pre>
<h3 id="복원-결과">복원 결과</h3>
<p><img src="https://velog.velcdn.com/images/j_keun/post/38170e6d-8c4e-4e5a-be3d-841b214639bd/image.png" alt=""></p>
<pre><code class="language-text">위쪽: 노이즈 이미지
가운데: 모델이 복원한 이미지
아래쪽: 깨끗한 원본 이미지</code></pre>
<p>노이즈는 크게 줄었고 숫자 형태도 잘 복원했다. 다만 원본보다 획이 부드럽거나 일부 세부 정보가 손실될 수 있다.</p>
<hr>
<h2 id="3-sparse-autoencoder">3. Sparse AutoEncoder</h2>
<p>일반 AutoEncoder는 Latent Representation의 많은 뉴런을 자유롭게 사용할 수 있다. Sparse AutoEncoder는 <strong>가능하면 적은 수의 뉴런만 강하게 활성화</strong>하도록 제약을 추가한다.</p>
<pre><code class="language-text">일반 Latent
[0.7, 0.5, 0.9, 0.4, 0.6, 0.8]

Sparse Latent
[0.8, 0.0, 1.3, 0.0, 0.9, 0.0]</code></pre>
<p>모든 뉴런을 조금씩 사용하는 대신 일부 뉴런이 특정 특징을 강하게 표현하도록 유도한다.</p>
<pre><code class="language-text">전체 Loss = Reconstruction Loss + Sparsity Penalty</code></pre>
<p>Latent 공간에 어떤 제약을 주느냐에 따라 학습되는 표현의 성질이 달라진다.</p>
<hr>
<h2 id="4-autoencoder를-이용한-이상-탐지">4. AutoEncoder를 이용한 이상 탐지</h2>
<p>정상 데이터만 충분히 학습한 AutoEncoder는 정상 패턴을 잘 복원하도록 최적화된다.</p>
<pre><code class="language-text">정상 데이터
→ 학습한 패턴과 비슷함
→ 잘 복원됨
→ Reconstruction Error가 작음

이상 데이터
→ 학습한 패턴과 다름
→ 잘 복원하지 못함
→ Reconstruction Error가 큼</code></pre>
<p>Reconstruction Error를 기준으로 Threshold를 정하면 이상 여부를 판단할 수 있다.</p>
<pre><code class="language-python">error = ((x - x_hat) ** 2).mean(dim=(1, 2, 3))
is_anomaly = error &gt; threshold</code></pre>
<p>하지만 AutoEncoder가 이상 데이터까지 잘 복원하는 경우도 있다. 따라서 실제 시스템에서는 다음 내용을 함께 고려해야 한다.</p>
<ul>
<li>정상 데이터의 다양성을 충분히 포함했는가?</li>
<li>Validation 데이터로 적절한 Threshold를 정했는가?</li>
<li>이상 유형별 Recall과 False Positive Rate는 어떤가?</li>
<li>다른 이상 탐지 방법과 비교했는가?</li>
</ul>
<p>Reconstruction Error 하나만으로 모든 이상 탐지 문제를 해결할 수 있는 것은 아니다.</p>
<hr>
<h2 id="5-일반-autoencoder의-생성-한계">5. 일반 AutoEncoder의 생성 한계</h2>
<p>일반 AutoEncoder는 입력을 복원하는 작업에는 적합하다.</p>
<pre><code class="language-text">x → Encoder → z → Decoder → x_hat</code></pre>
<p>새 이미지를 생성하려면 입력 이미지 없이 시작점 <code>z</code>를 정해야 한다.</p>
<pre><code class="language-text">? → z → Decoder → 새로운 이미지</code></pre>
<p>하지만 일반 AutoEncoder는 Latent Space가 규칙적인 확률분포가 되도록 학습하지 않는다.</p>
<ul>
<li>학습 데이터가 존재하는 위치에서는 Decoder가 잘 동작할 수 있다.</li>
<li>학습된 위치 사이에는 Decoder가 충분히 학습하지 못한 빈 영역이 생길 수 있다.</li>
<li>임의의 <code>z</code>를 Decoder에 넣는다고 자연스러운 이미지가 보장되지는 않는다.</li>
</ul>
<p>즉, <strong>복원을 잘하는 것과 새로운 데이터를 잘 생성하는 것은 다른 문제</strong>다.</p>
<hr>
<h2 id="6-vae">6. VAE</h2>
<p>VAE(Variational AutoEncoder)는 Latent Space를 확률분포로 표현하는 생성 모델이다.</p>
<p>일반 AutoEncoder는 이미지 하나를 정확한 좌표 하나로 바꾼다.</p>
<pre><code class="language-text">이미지 → z = [2.0, 3.0]</code></pre>
<p>VAE는 정확한 점 하나가 아니라 평균과 분산으로 범위를 표현한다.</p>
<pre><code class="language-text">이미지
  ↓
Encoder
  ↓
평균 μ, 분산 σ²
  ↓
분포에서 z Sampling
  ↓
Decoder
  ↓
복원 이미지</code></pre>
<p>지도에 비유하면 다음과 같다.</p>
<table>
<thead>
<tr>
<th>값</th>
<th>의미</th>
<th>지도 비유</th>
</tr>
</thead>
<tbody><tr>
<td><code>μ</code></td>
<td>분포의 중심</td>
<td>동네의 중심 위치</td>
</tr>
<tr>
<td><code>σ</code></td>
<td>분포의 퍼짐 정도</td>
<td>동네의 넓이</td>
</tr>
<tr>
<td><code>z</code></td>
<td>분포에서 뽑은 한 점</td>
<td>동네 안에서 선택한 한 지점</td>
</tr>
</tbody></table>
<p>각 데이터를 주변 범위까지 포함해 표현하고, 분포들이 지나치게 흩어지지 않도록 제한한다. 이를 통해 임의의 <code>z</code>를 Sampling해도 의미 있는 결과가 나올 가능성을 높인다.</p>
<hr>
<h2 id="7-reparameterization-trick">7. Reparameterization Trick</h2>
<p>Encoder는 <code>μ</code>와 <code>logvar</code>를 출력한다.</p>
<pre><code class="language-text">var = σ²
logvar = log(σ²)</code></pre>
<p>표준편차는 다음처럼 계산한다.</p>
<pre><code class="language-python">std = torch.exp(0.5 * logvar)</code></pre>
<p>그다음 표준정규분포에서 <code>epsilon</code>을 뽑아 <code>z</code>를 만든다.</p>
<pre><code class="language-python">eps = torch.randn_like(std)
z = mu + eps * std</code></pre>
<blockquote>
<p><code>eps</code>는 “아주 작은 값”을 의미하지 않는다. 그리스 문자 epsilon을 변수 이름으로 사용한 것이며, 여기서는 평균 0, 표준편차 1인 정규분포에서 뽑은 무작위 표본이다.</p>
</blockquote>
<h3 id="실행-예제">실행 예제</h3>
<pre><code class="language-python">mu = torch.tensor([[1.0, 2.0]])
logvar = torch.tensor([[0.0, 0.0]])

std = torch.exp(0.5 * logvar)
eps = torch.randn_like(std)
z = mu + eps * std

print(&quot;mu:&quot;, mu)
print(&quot;logvar:&quot;, logvar)
print(&quot;std:&quot;, std)
print(&quot;epsilon:&quot;, eps)
print(&quot;z:&quot;, z)</code></pre>
<p>노트북의 실행 결과:</p>
<pre><code class="language-text">mu:      tensor([[1., 2.]])
logvar:  tensor([[0., 0.]])
std:     tensor([[1., 1.]])
epsilon: tensor([[-1.1235, 1.1595]])
z:       tensor([[-0.1235, 3.1595]])</code></pre>
<h3 id="왜-이-방식이-필요한가">왜 이 방식이 필요한가?</h3>
<p>단순히 분포에서 직접 Sampling하면 무작위 연산을 통과해 <code>μ</code>와 <code>σ</code>를 만든 Encoder까지 Gradient를 전달하기 어렵다.</p>
<pre><code class="language-text">z = μ + ε × σ</code></pre>
<p>처럼 표현하면 무작위성은 <code>ε</code>에 분리되고, <code>z</code>는 <code>μ</code>와 <code>σ</code>에 대한 미분 가능한 계산으로 만들어진다. 이를 <strong>Reparameterization Trick</strong>이라고 한다.</p>
<hr>
<h2 id="8-vae-loss">8. VAE Loss</h2>
<p>VAE는 두 Loss를 함께 사용한다.</p>
<pre><code class="language-text">VAE Loss = Reconstruction Loss + KL Loss</code></pre>
<p>수식으로 표현하면 다음과 같다.</p>
<pre><code class="language-text">L_VAE = L_reconstruction + L_KL</code></pre>
<h3 id="reconstruction-loss">Reconstruction Loss</h3>
<p>원본과 복원 이미지의 차이를 줄인다.</p>
<pre><code class="language-python">recon_loss = F.binary_cross_entropy(
    x_hat,
    x,
    reduction=&quot;sum&quot;,
)</code></pre>
<h3 id="kl-divergence">KL Divergence</h3>
<p>Encoder가 만든 분포가 기준 분포인 표준정규분포 <code>N(0, I)</code>에서 지나치게 멀어지지 않도록 한다.</p>
<pre><code class="language-python">kl_loss = -0.5 * torch.sum(
    1 + logvar - mu.pow(2) - logvar.exp()
)</code></pre>
<p>KL Loss는 Latent Space가 무작위 Sampling에 적합한 규칙적인 형태를 갖도록 유도한다.</p>
<pre><code class="language-python">def vae_loss_function(x_hat, x, mu, logvar):
    recon_loss = F.binary_cross_entropy(
        x_hat,
        x,
        reduction=&quot;sum&quot;,
    )

    kl_loss = -0.5 * torch.sum(
        1 + logvar - mu.pow(2) - logvar.exp()
    )

    batch_size = x.size(0)

    total_loss = (
        recon_loss + kl_loss
    ) / batch_size

    recon_loss = recon_loss / batch_size
    kl_loss = kl_loss / batch_size

    return total_loss, recon_loss, kl_loss</code></pre>
<p><code>reduction=&quot;sum&quot;</code>으로 784개 픽셀의 BCE를 모두 더한 뒤 Batch 크기로 나누므로 Loss 값이 기본 AutoEncoder의 평균 MSE보다 크게 나타난다. 서로 다른 계산 방식의 Loss 숫자를 직접 비교하면 안 된다.</p>
<hr>
<h2 id="9-합성곱-vae-구현">9. 합성곱 VAE 구현</h2>
<p>이번 VAE는 2차원 Latent Space를 사용한다.</p>
<pre><code class="language-python">LATENT_DIM = 2</code></pre>
<h3 id="전체-shape-흐름">전체 Shape 흐름</h3>
<pre><code class="language-text">x: [B, 1, 28, 28]
        ↓ Encoder
h: [B, 64, 7, 7]
        ↓ Flatten
h: [B, 3136]
        ├─ fc_mu     → μ:      [B, 2]
        └─ fc_logvar → logvar: [B, 2]
                           ↓ Sampling
                         z: [B, 2]
                           ↓ fc_decode
                       [B, 3136]
                           ↓ Reshape
                       [B, 64, 7, 7]
                           ↓ Decoder
x_hat: [B, 1, 28, 28]</code></pre>
<h3 id="모델-코드">모델 코드</h3>
<pre><code class="language-python">class ConvVAE(nn.Module):
    def __init__(self, latent_dim=2):
        super().__init__()

        self.encoder = nn.Sequential(
            nn.Conv2d(
                1, 32,
                kernel_size=3,
                stride=2,
                padding=1,
            ),
            nn.ReLU(),
            nn.Conv2d(
                32, 64,
                kernel_size=3,
                stride=2,
                padding=1,
            ),
            nn.ReLU(),
        )

        self.flatten_dim = 64 * 7 * 7

        self.fc_mu = nn.Linear(
            self.flatten_dim,
            latent_dim,
        )
        self.fc_logvar = nn.Linear(
            self.flatten_dim,
            latent_dim,
        )

        self.fc_decode = nn.Linear(
            latent_dim,
            self.flatten_dim,
        )

        self.decoder = nn.Sequential(
            nn.ConvTranspose2d(
                64, 32,
                kernel_size=4,
                stride=2,
                padding=1,
            ),
            nn.ReLU(),
            nn.ConvTranspose2d(
                32, 1,
                kernel_size=4,
                stride=2,
                padding=1,
            ),
            nn.Sigmoid(),
        )

    def encode(self, x):
        h = self.encoder(x)
        h = h.view(x.size(0), -1)

        mu = self.fc_mu(h)
        logvar = self.fc_logvar(h)

        return mu, logvar

    def reparameterize(self, mu, logvar):
        std = torch.exp(0.5 * logvar)
        eps = torch.randn_like(std)
        return mu + eps * std

    def decode(self, z):
        h = self.fc_decode(z)
        h = h.view(-1, 64, 7, 7)
        return self.decoder(h)

    def forward(self, x):
        mu, logvar = self.encode(x)
        z = self.reparameterize(mu, logvar)
        x_hat = self.decode(z)

        return x_hat, mu, logvar</code></pre>
<hr>
<h2 id="10-vae-학습-결과">10. VAE 학습 결과</h2>
<pre><code class="language-python">optimizer = optim.AdamW(
    vae.parameters(),
    lr=1e-3,
)</code></pre>
<p>학습 중에는 Total, Reconstruction, KL Loss를 각각 기록한다.</p>
<p>노트북의 주요 결과는 다음과 같다.</p>
<table>
<thead>
<tr>
<th align="right">Epoch</th>
<th align="right">Total</th>
<th align="right">Reconstruction</th>
<th align="right">KL</th>
</tr>
</thead>
<tbody><tr>
<td align="right">1</td>
<td align="right">194.5616</td>
<td align="right">188.3023</td>
<td align="right">6.2593</td>
</tr>
<tr>
<td align="right">2</td>
<td align="right">170.0631</td>
<td align="right">165.0527</td>
<td align="right">5.0104</td>
</tr>
<tr>
<td align="right">5</td>
<td align="right">160.4446</td>
<td align="right">155.2527</td>
<td align="right">5.1919</td>
</tr>
<tr>
<td align="right">10</td>
<td align="right">156.8558</td>
<td align="right">151.3831</td>
<td align="right">5.4727</td>
</tr>
</tbody></table>
<p><img src="https://velog.velcdn.com/images/j_keun/post/a796fd7e-005a-48e8-9954-3fa6da2b4442/image.png" alt=""></p>
<ul>
<li>Total Loss와 Reconstruction Loss가 감소하며 복원 성능이 좋아졌다.</li>
<li>KL Loss는 단조롭게 감소하지 않고 Reconstruction 목표와 균형을 이루며 변한다.</li>
<li>KL Loss가 지나치게 빠르게 <code>0</code>에 가까워지면 Decoder가 Latent 정보를 거의 사용하지 않는 <strong>Posterior Collapse</strong>가 발생할 수 있다.</li>
</ul>
<p>따라서 KL Loss는 무조건 작을수록 좋다고 해석하면 안 된다.</p>
<h3 id="vae-복원-결과">VAE 복원 결과</h3>
<p><img src="https://velog.velcdn.com/images/j_keun/post/def31210-d2c7-4b37-98c0-22b43a9c8b47/image.png" alt=""></p>
<p>위쪽은 원본, 아래쪽은 복원 이미지다. 숫자의 전체적인 형태는 복원하지만 기본 AutoEncoder보다 부드럽거나 흐리게 보이는 결과가 있다.</p>
<p>VAE는 Latent에서 확률적으로 <code>z</code>를 Sampling하며, 단순한 픽셀 단위 Loss와 Decoder 구조도 흐릿한 복원에 영향을 줄 수 있다.</p>
<hr>
<h2 id="11-현대-생성-모델에서-vae의-역할">11. 현대 생성 모델에서 VAE의 역할</h2>
<p>고품질 이미지 생성은 단순 VAE만 사용하는 방식보다 Diffusion, Transformer, Flow 계열 모델과 결합하는 방향으로 발전했다.</p>
<p>Latent 기반 생성 모델은 고해상도 픽셀 공간을 직접 처리하는 대신 이미지를 작은 Latent Representation으로 압축한 뒤 그 공간에서 생성 작업을 수행할 수 있다.</p>
<pre><code class="language-text">Image
  ↓
VAE Encoder
  ↓
Latent Representation
  ↓
Diffusion 또는 Transformer
  ↓
Generated Latent
  ↓
VAE Decoder
  ↓
Generated Image</code></pre>
<p>VAE는 확률적인 Latent Space와 Sampling을 이해하는 기초가 된다.</p>
<hr>
<h2 id="12-gan">12. GAN</h2>
<p>GAN(Generative Adversarial Network)은 생성자와 판별자라는 두 신경망이 서로 경쟁하며 학습하는 생성 모델이다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/38231097-c8d8-4e87-8c5d-4406a35dfb8c/image.png" alt=""></p>
<table>
<thead>
<tr>
<th>모델</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td>Generator(G)</td>
<td>무작위 Noise로 진짜처럼 보이는 가짜 데이터를 만든다.</td>
</tr>
<tr>
<td>Discriminator(D)</td>
<td>입력 데이터가 실제 데이터인지 생성된 가짜인지 판별한다.</td>
</tr>
</tbody></table>
<pre><code class="language-text">Noise z
   ↓
Generator
   ↓
Fake Image ──┐
             ├─ Discriminator → Real 또는 Fake 판단
Real Image ──┘</code></pre>
<p>GAN은 특정 학습 이미지를 그대로 복사하는 것이 아니라 학습 데이터의 분포와 특징을 익혀 새로운 Sample을 생성하는 것을 목표로 한다.</p>
<hr>
<h2 id="13-gan의-경쟁-학습">13. GAN의 경쟁 학습</h2>
<h3 id="판별자-학습">판별자 학습</h3>
<p>판별자는 실제 이미지에는 <code>1</code>, 생성자가 만든 이미지에는 <code>0</code>을 출력하도록 학습한다.</p>
<pre><code class="language-text">실제 이미지 x
→ D(x)
→ 1에 가깝게

가짜 이미지 G(z)
→ D(G(z))
→ 0에 가깝게</code></pre>
<h3 id="생성자-학습">생성자 학습</h3>
<p>생성자는 판별자가 가짜 이미지를 진짜라고 판단하도록 학습한다.</p>
<pre><code class="language-text">z → G(z) → D(G(z)) → 1에 가깝게</code></pre>
<p>여기서 Target <code>1</code>은 생성 이미지가 실제 정답 이미지라는 뜻이 아니다. 생성자가 판별자를 속이기 위해 원하는 목표값이다.</p>
<p>두 모델을 번갈아 학습하면서 생성자는 실제 데이터 분포와 비슷한 Sample을 만들고, 판별자는 진짜와 가짜를 더 잘 구별하도록 발전한다.</p>
<hr>
<h2 id="14-gan-데이터-준비">14. GAN 데이터 준비</h2>
<pre><code class="language-python">BATCH_SIZE = 128

transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.5,), (0.5,)),
])

train_dataset = datasets.MNIST(
    root=&quot;./data&quot;,
    train=True,
    download=True,
    transform=transform,
)

train_loader = DataLoader(
    train_dataset,
    batch_size=BATCH_SIZE,
    shuffle=True,
    num_workers=4,
)</code></pre>
<p><code>ToTensor()</code>가 만든 <code>[0, 1]</code> 범위의 픽셀을 <code>Normalize((0.5,), (0.5,))</code>로 변환한다.</p>
<pre><code class="language-text">정규화 결과 = (x - 0.5) / 0.5

0 → -1
1 →  1</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">images: torch.Size([128, 1, 28, 28])
min/max: -1.0 1.0</code></pre>
<p>Generator 마지막 층의 <code>Tanh</code>도 <code>[-1, 1]</code> 범위를 출력하므로 실제 이미지와 생성 이미지의 범위를 맞춘다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/1d5bc607-4ee1-4c0e-852a-d547ec4bdf68/image.png" alt=""></p>
<hr>
<h2 id="15-generator">15. Generator</h2>
<p>Generator는 이미지가 아니라 <code>z</code>라는 무작위 벡터를 입력받는다.</p>
<pre><code class="language-text">z [B, 100]
   ↓ Linear
[B, 256]
   ↓ Linear
[B, 512]
   ↓ Linear
[B, 784]
   ↓ Reshape
[B, 1, 28, 28]</code></pre>
<pre><code class="language-python">LATENT_DIM = 100

class Generator(nn.Module):
    def __init__(self, latent_dim=100):
        super().__init__()

        self.net = nn.Sequential(
            nn.Linear(latent_dim, 256),
            nn.LeakyReLU(0.2, inplace=True),

            nn.Linear(256, 512),
            nn.BatchNorm1d(512),
            nn.LeakyReLU(0.2, inplace=True),

            nn.Linear(512, 28 * 28),
            nn.Tanh(),
        )

    def forward(self, z):
        x = self.net(z)
        return x.view(-1, 1, 28, 28)

G = Generator(LATENT_DIM).to(DEVICE)</code></pre>
<h3 id="leakyrelu"><code>LeakyReLU</code></h3>
<p>일반 ReLU는 음수 입력을 모두 <code>0</code>으로 만든다. <code>LeakyReLU(0.2)</code>는 음수 영역에도 작은 기울기를 남겨 Gradient가 완전히 끊어지는 현상을 줄인다.</p>
<pre><code class="language-text">x &gt;= 0 → x
x &lt; 0  → 0.2x</code></pre>
<h3 id="생성-결과-shape">생성 결과 Shape</h3>
<pre><code class="language-python">z = torch.randn(
    8,
    LATENT_DIM,
    device=DEVICE,
)

fake = G(z)

print(&quot;Noise:&quot;, z.shape)
print(&quot;Fake:&quot;, fake.shape)
print(&quot;Range:&quot;, fake.min().item(), fake.max().item())</code></pre>
<p>노트북 실행 결과:</p>
<pre><code class="language-text">Noise: torch.Size([8, 100])
Fake: torch.Size([8, 1, 28, 28])
Range: -0.9718 0.9251</code></pre>
<hr>
<h2 id="16-discriminator">16. Discriminator</h2>
<p>Discriminator는 이미지 한 장을 받아 진짜 여부를 나타내는 Logit 하나를 출력한다.</p>
<pre><code class="language-text">[B, 1, 28, 28]
   ↓ Flatten
[B, 784]
   ↓ Linear
[B, 512]
   ↓ Linear
[B, 256]
   ↓ Linear
[B, 1] Logit</code></pre>
<pre><code class="language-python">class Discriminator(nn.Module):
    def __init__(self):
        super().__init__()

        self.net = nn.Sequential(
            nn.Flatten(),

            nn.Linear(28 * 28, 512),
            nn.LeakyReLU(0.2, inplace=True),
            nn.Dropout(0.3),

            nn.Linear(512, 256),
            nn.LeakyReLU(0.2, inplace=True),
            nn.Dropout(0.3),

            nn.Linear(256, 1),
        )

    def forward(self, x):
        return self.net(x)

D = Discriminator().to(DEVICE)</code></pre>
<p>마지막에 <code>Sigmoid</code>를 넣지 않는다. <code>BCEWithLogitsLoss</code>가 내부에서 Sigmoid와 BCE를 수치적으로 안정적인 방식으로 함께 계산하기 때문이다.</p>
<pre><code class="language-python">criterion = nn.BCEWithLogitsLoss()</code></pre>
<p>확률이 필요할 때만 출력 Logit에 <code>torch.sigmoid()</code>를 적용한다.</p>
<pre><code class="language-python">with torch.no_grad():
    logits = D(fake)
    probabilities = torch.sigmoid(logits)</code></pre>
<hr>
<h2 id="17-optimizer와-fixed-noise">17. Optimizer와 Fixed Noise</h2>
<p>Generator와 Discriminator는 서로 다른 Optimizer를 사용한다.</p>
<pre><code class="language-python">optimizer_D = torch.optim.Adam(
    D.parameters(),
    lr=2e-4,
    betas=(0.5, 0.999),
)

optimizer_G = torch.optim.Adam(
    G.parameters(),
    lr=2e-4,
    betas=(0.5, 0.999),
)</code></pre>
<p>학습 변화를 비교할 고정 Noise도 만든다.</p>
<pre><code class="language-python">fixed_noise = torch.randn(
    64,
    LATENT_DIM,
    device=DEVICE,
)</code></pre>
<p>Epoch마다 새로운 Noise를 사용하면 이미지가 달라진 이유가 학습 때문인지 입력 Noise 때문인지 구분하기 어렵다. 같은 <code>fixed_noise</code>를 반복 사용하면 <strong>같은 잠재 입력이 학습에 따라 어떻게 바뀌는지</strong> 비교할 수 있다.</p>
<hr>
<h2 id="18-discriminator-학습-과정">18. Discriminator 학습 과정</h2>
<pre><code class="language-python">optimizer_D.zero_grad()

# 실제 이미지의 Target은 1
real_logits = D(real)
real_targets = torch.ones_like(real_logits)
d_real_loss = criterion(
    real_logits,
    real_targets,
)

# 생성 이미지의 Target은 0
z = torch.randn(
    batch_size,
    LATENT_DIM,
    device=DEVICE,
)

fake = G(z)
fake_logits = D(fake.detach())
fake_targets = torch.zeros_like(fake_logits)
d_fake_loss = criterion(
    fake_logits,
    fake_targets,
)

d_loss = d_real_loss + d_fake_loss
d_loss.backward()
optimizer_D.step()</code></pre>
<h3 id="fakedetach가-필요한-이유"><code>fake.detach()</code>가 필요한 이유</h3>
<p><code>fake</code>는 Generator의 연산 결과이므로 원래는 Generator의 계산 그래프와 연결되어 있다.</p>
<p>Discriminator를 학습하는 단계에서는 Generator를 업데이트하면 안 된다.</p>
<pre><code class="language-python">fake.detach()</code></pre>
<p>는 Tensor 값은 그대로 사용하면서 Generator 방향의 Gradient 연결을 끊는다.</p>
<pre><code class="language-text">G(z) ──X── Gradient 차단
  ↓
D(fake) → D만 학습</code></pre>
<hr>
<h2 id="19-generator-학습-과정">19. Generator 학습 과정</h2>
<p>Generator는 판별자가 가짜 이미지를 <code>1</code>, 즉 진짜로 판단하도록 학습한다.</p>
<pre><code class="language-python">optimizer_G.zero_grad()

z = torch.randn(
    batch_size,
    LATENT_DIM,
    device=DEVICE,
)

fake = G(z)
fake_logits = D(fake)

g_targets = torch.ones_like(fake_logits)
g_loss = criterion(fake_logits, g_targets)

g_loss.backward()
optimizer_G.step()</code></pre>
<p>이 단계에서는 <code>fake.detach()</code>를 사용하지 않는다. <code>D(fake)</code>에서 계산한 Gradient가 Discriminator 연산을 거쳐 Generator까지 전달되어야 하기 때문이다.</p>
<pre><code class="language-text">G Parameter
   ↑ Gradient
G(z) → D(G(z)) → BCE Target 1</code></pre>
<p>Generator와 Discriminator 학습을 한 Batch 안에서 차례로 반복한다.</p>
<hr>
<h2 id="20-gan-생성-결과">20. GAN 생성 결과</h2>
<h3 id="epoch-1">Epoch 1</h3>
<p><img src="https://velog.velcdn.com/images/j_keun/post/cefaa5bd-050b-4813-b464-10e7f417b063/image.png" alt=""></p>
<p>숫자와 비슷한 형태가 나타나기 시작했지만 노이즈가 많고 경계가 불분명하다.</p>
<h3 id="epoch-10">Epoch 10</h3>
<p><img src="https://velog.velcdn.com/images/j_keun/post/a8d17966-bd14-4327-ae32-f7dcd146bb33/image.png" alt=""></p>
<p>Epoch 1보다 숫자 형태가 선명해지고 배경 노이즈도 줄었다. 다만 일부 이미지는 어떤 숫자인지 불명확하며 획이 깨진 Sample도 남아 있다.</p>
<p>이 예제는 <code>Linear</code> 층으로 만든 간단한 GAN이므로 학습 시간이 짧고 생성 품질에도 한계가 있다.</p>
<hr>
<h2 id="21-gan-loss-해석">21. GAN Loss 해석</h2>
<p>노트북의 Epoch별 마지막 Batch Loss는 다음과 같다.</p>
<table>
<thead>
<tr>
<th align="right">Epoch</th>
<th align="right">D Loss</th>
<th align="right">G Loss</th>
</tr>
</thead>
<tbody><tr>
<td align="right">1</td>
<td align="right">1.1470</td>
<td align="right">0.8597</td>
</tr>
<tr>
<td align="right">2</td>
<td align="right">1.0611</td>
<td align="right">1.0561</td>
</tr>
<tr>
<td align="right">4</td>
<td align="right">1.0332</td>
<td align="right">1.4288</td>
</tr>
<tr>
<td align="right">7</td>
<td align="right">1.0145</td>
<td align="right">1.5309</td>
</tr>
<tr>
<td align="right">10</td>
<td align="right">0.9822</td>
<td align="right">1.1602</td>
</tr>
</tbody></table>
<p><img src="https://velog.velcdn.com/images/j_keun/post/580ebf43-3ee8-487c-a4fa-86a24f18d1b8/image.png" alt=""></p>
<p>일반적인 분류 모델은 Loss가 계속 낮아지면 학습이 잘된다고 해석하는 경우가 많다. GAN은 두 모델이 서로 경쟁하므로 G Loss와 D Loss가 단조롭게 감소하지 않아도 된다.</p>
<p>다음 내용을 함께 확인해야 한다.</p>
<ol>
<li>고정 Noise에서 생성 이미지가 점차 좋아지는가?</li>
<li>여러 숫자가 생성되며 다양성이 유지되는가?</li>
<li>거의 같은 이미지만 반복해서 생성하지 않는가?</li>
<li>Discriminator가 너무 강해 Generator의 학습이 멈추지 않는가?</li>
<li>생성 이미지에 심한 Artifact가 남아 있지 않은가?</li>
</ol>
<blockquote>
<p>위 표의 값은 각 Epoch 전체 평균이 아니라 학습 코드가 출력한 마지막 Batch의 Loss다. 전체 경향을 비교하려면 Epoch 평균 Loss를 별도로 계산하는 편이 더 안정적이다.</p>
</blockquote>
<hr>
<h2 id="22-mode-collapse">22. Mode Collapse</h2>
<p>Mode Collapse는 Generator가 다양한 Sample을 만들지 못하고 비슷한 결과만 반복해서 생성하는 현상이다.</p>
<p>예를 들어 MNIST에는 0부터 9까지 다양한 숫자가 있지만 Generator가 판별자를 잘 속이는 특정 모양의 <code>3</code>만 반복해서 만들 수 있다.</p>
<pre><code class="language-text">서로 다른 Noise z
  ↓
Generator
  ↓
거의 동일한 이미지들</code></pre>
<p>생성 이미지가 선명해 보여도 다양성이 부족하면 좋은 생성 모델이라고 보기 어렵다.</p>
<hr>
<h2 id="23-생성-품질-평가">23. 생성 품질 평가</h2>
<p>GAN의 Loss만으로 이미지 품질을 완전히 판단하기 어렵다. 실제 생성 결과를 시각화하고 필요하면 생성 품질 지표를 사용한다.</p>
<p>대표적인 지표로 FID(Fréchet Inception Distance)가 있다.</p>
<ul>
<li>실제 이미지와 생성 이미지의 Feature 분포를 비교한다.</li>
<li>일반적으로 값이 낮을수록 두 분포가 비슷하다고 해석한다.</li>
<li>충분한 Sample 수와 일관된 평가 설정이 필요하다.</li>
</ul>
<p>MNIST처럼 단순한 데이터에서는 숫자 분류 모델을 이용해 생성 이미지의 인식률과 클래스 다양성을 함께 확인할 수도 있다.</p>
<hr>
<h2 id="24-autoencoder-vae-gan-비교">24. AutoEncoder, VAE, GAN 비교</h2>
<table>
<thead>
<tr>
<th>항목</th>
<th>AutoEncoder</th>
<th>VAE</th>
<th>GAN</th>
</tr>
</thead>
<tbody><tr>
<td>주요 목적</td>
<td>입력 복원과 특징 학습</td>
<td>규칙적인 Latent Space 학습과 생성</td>
<td>실제와 비슷한 Sample 생성</td>
</tr>
<tr>
<td>입력</td>
<td>실제 데이터</td>
<td>실제 데이터</td>
<td>무작위 Noise</td>
</tr>
<tr>
<td>Target</td>
<td>입력 데이터 자체</td>
<td>입력 데이터 자체</td>
<td>판별자를 속이는 목표</td>
</tr>
<tr>
<td>Latent 표현</td>
<td>결정적인 값</td>
<td><code>μ</code>, <code>logvar</code>로 표현한 확률분포</td>
<td>입력 Noise <code>z</code></td>
</tr>
<tr>
<td>핵심 Loss</td>
<td>Reconstruction Loss</td>
<td>Reconstruction + KL Loss</td>
<td>Generator와 Discriminator의 적대적 Loss</td>
</tr>
<tr>
<td>생성 방식</td>
<td>임의 Sampling이 어려울 수 있음</td>
<td>분포에서 <code>z</code>를 Sampling</td>
<td>Noise를 Generator에 입력</td>
</tr>
<tr>
<td>대표적 특징</td>
<td>안정적으로 복원한다.</td>
<td>연속적인 Latent Space를 학습한다.</td>
<td>비교적 선명한 Sample을 만들 수 있다.</td>
</tr>
<tr>
<td>주의점</td>
<td>Latent 공간이 불규칙할 수 있다.</td>
<td>복원 결과가 흐릴 수 있다.</td>
<td>학습이 불안정하고 Mode Collapse가 생길 수 있다.</td>
</tr>
</tbody></table>
<p>VAE와 GAN은 모두 생성 모델이지만 학습 방식은 다르다.</p>
<pre><code class="language-text">VAE
분포를 명시적으로 정리하며 Reconstruction과 KL Loss를 최소화

GAN
생성자와 판별자의 경쟁을 통해 실제 데이터 분포를 간접적으로 학습</code></pre>
<hr>
<h2 id="25-핵심-정리">25. 핵심 정리</h2>
<ol>
<li>Denoising AutoEncoder는 노이즈 이미지를 입력하고 깨끗한 원본을 Target으로 사용한다.</li>
<li>Sparse AutoEncoder는 적은 수의 Latent 뉴런만 강하게 사용하도록 제약한다.</li>
<li>정상 데이터로 학습한 AutoEncoder의 Reconstruction Error를 이상 탐지에 활용할 수 있다.</li>
<li>일반 AutoEncoder는 복원을 잘해도 임의의 Latent Vector에서 자연스러운 데이터를 생성한다고 보장할 수 없다.</li>
<li>VAE는 하나의 좌표 대신 <code>μ</code>와 <code>σ</code>로 Latent 분포를 표현한다.</li>
<li>Reparameterization Trick은 무작위성을 <code>epsilon</code>에 분리해 Encoder까지 Gradient가 흐르게 한다.</li>
<li>VAE Loss는 Reconstruction Loss와 KL Loss의 합이다.</li>
<li>KL Loss가 지나치게 <code>0</code>에 가까워지면 Latent를 사용하지 않는 Posterior Collapse가 생길 수 있다.</li>
<li>GAN의 Generator는 Noise로 가짜 이미지를 만들고, Discriminator는 진짜와 가짜를 구별한다.</li>
<li><code>BCEWithLogitsLoss</code>를 사용하면 Discriminator 마지막 층에 Sigmoid를 따로 넣지 않는다.</li>
<li>Discriminator 학습에서 <code>fake.detach()</code>는 Generator 방향의 Gradient를 차단한다.</li>
<li>GAN Loss는 경쟁 관계 때문에 단조롭게 감소하지 않으며 생성 이미지와 다양성을 함께 평가해야 한다.</li>
</ol>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 고고학 최고의 발견]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EA%B3%A0%EA%B3%A0%ED%95%99-%EC%B5%9C%EA%B3%A0%EC%9D%98-%EB%B0%9C%EA%B2%AC</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EA%B3%A0%EA%B3%A0%ED%95%99-%EC%B5%9C%EA%B3%A0%EC%9D%98-%EB%B0%9C%EA%B2%AC</guid>
            <pubDate>Thu, 01 Oct 2026 17:45:36 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p><code>n x n</code> 시계 격자에서 한 시계를 조작하면 자기 자신과 상하좌우 인접 시계가 시계 방향으로 90도 회전한다. 모든 시곗바늘을 12시 방향인 <code>0</code>으로 만들기 위한 최소 조작 횟수를 구한다.</p>
<p>시계 방향은 <code>0, 1, 2, 3</code>으로 표현하며 모든 변화는 <code>mod 4</code> 연산으로 처리할 수 있다.</p>
<h2 id="핵심-관찰">핵심 관찰</h2>
<h3 id="1-같은-시계는-최대-3번만-조작하면-된다">1. 같은 시계는 최대 3번만 조작하면 된다</h3>
<p>한 시계를 4번 조작하면 영향을 받는 모든 시계가 한 바퀴 돌아 원래 상태로 돌아온다.</p>
<p>따라서 최적해에서 각 시계의 조작 횟수는 <code>0</code>, <code>1</code>, <code>2</code>, <code>3</code>번 중 하나만 고려하면 충분하다.</p>
<h3 id="2-첫-행을-제외한-행의-조작은-강제로-결정된다">2. 첫 행을 제외한 행의 조작은 강제로 결정된다</h3>
<p><code>r</code>행의 시계를 조작할 수 있는 시점에는 <code>r - 1</code>행을 더 이상 바꿀 수 있는 방법이 거의 없다.</p>
<p><code>r - 1</code>행의 <code>c</code>열 시계에 영향을 주는 아래 행의 조작은 오직 <code>(r, c)</code>의 조작뿐이다. 그러므로 <code>(r - 1, c)</code>를 <code>0</code>으로 만들기 위해 <code>(r, c)</code>를 몇 번 조작해야 하는지는 다음처럼 유일하게 결정된다.</p>
<pre><code class="language-python">count = (-board[r - 1][c]) % 4</code></pre>
<p>첫 행의 조작 횟수만 정하면, 둘째 행부터 마지막 행까지의 조작 횟수는 모두 자동으로 정해진다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>첫 행의 각 칸을 <code>0~3</code>번 조작하는 모든 경우를 탐색한다.</li>
<li>첫 행 조작을 격자에 적용한다.</li>
<li><code>r = 1</code>부터 마지막 행까지 순회한다.</li>
<li>바로 위 행의 각 시계를 0으로 만들기 위해 현재 행의 같은 열 시계를 필요한 횟수만큼 조작한다.</li>
<li>마지막 행이 모두 0인 경우, 조작 횟수 최솟값을 갱신한다.</li>
</ol>
<p>첫 행의 경우의 수는 최대 <code>4^8 = 65,536</code>개이므로 충분히 탐색할 수 있다.</p>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">from itertools import product


def solution(clockHands):
    n = len(clockHands)
    answer = float(&quot;inf&quot;)

    dr = [0, -1, 1, 0, 0]
    dc = [0, 0, 0, -1, 1]

    def rotate(board, row, col, count):
        &quot;&quot;&quot;(row, col)을 count번 조작한 결과를 보드에 반영한다.&quot;&quot;&quot;
        for direction in range(5):
            nr = row + dr[direction]
            nc = col + dc[direction]

            if 0 &lt;= nr &lt; n and 0 &lt;= nc &lt; n:
                board[nr][nc] = (board[nr][nc] + count) % 4

    # 첫 행의 모든 조작 조합을 시도한다.
    for first_row in product(range(4), repeat=n):
        board = [row[:] for row in clockHands]
        operation_count = sum(first_row)

        # 첫 행 조작을 적용한다.
        for col, count in enumerate(first_row):
            if count:
                rotate(board, 0, col, count)

        # 현재 행을 조작해 바로 위 행을 0으로 확정한다.
        for row in range(1, n):
            for col in range(n):
                count = (-board[row - 1][col]) % 4

                if count:
                    rotate(board, row, col, count)
                    operation_count += count

        # 마지막 행도 모두 0이면 유효한 해다.
        if all(clock == 0 for clock in board[-1]):
            answer = min(answer, operation_count)

    return answer</code></pre>
<h2 id="정당성-설명">정당성 설명</h2>
<p>첫 행의 각 칸은 0~3번만 조작하면 되므로, 알고리즘은 가능한 첫 행 조작 조합을 모두 탐색한다.</p>
<p>임의의 첫 행 조작 조합을 고정하자. <code>r - 1</code>행 <code>c</code>열에 영향을 주는 아직 선택되지 않은 조작은 <code>(r, c)</code>뿐이다. 따라서 이 시계를 <code>0</code>으로 만들려면 <code>(r, c)</code>를 <code>(-board[r - 1][c]) mod 4</code>번 조작해야 하며, 다른 횟수로는 <code>r - 1</code>행을 0으로 확정할 수 없다.</p>
<p>그러므로 고정된 첫 행 조합에 대해 알고리즘이 계산한 이후 행의 조작은 유일하며, 그 조합에서 가능한 해가 존재한다면 반드시 찾아낸다. 모든 첫 행 조합을 탐색하므로 모든 가능한 해를 검토하고, 그중 최소 조작 횟수를 반환한다.</p>
<h2 id="복잡도-분석">복잡도 분석</h2>
<p>첫 행의 조작 조합은 <code>4^n</code>개다. 각 조합마다 <code>n^2</code>개의 칸을 처리하고, 한 번의 조작 반영은 최대 5개 칸에만 영향을 준다.</p>
<ul>
<li>시간 복잡도: <code>O(4^n * n^2)</code></li>
<li>공간 복잡도: <code>O(n^2)</code></li>
</ul>
<p><code>n &lt;= 8</code>이므로 최대 <code>65,536</code>개의 첫 행 조합만 확인하면 된다.</p>
<h2 id="주의할-점">주의할 점</h2>
<ul>
<li>조작 횟수는 <code>board[row - 1][col]</code>이 아니라 그 음수의 <code>mod 4</code> 값이다. 현재 값이 <code>3</code>이면 한 번 조작해 <code>0</code>으로 만들어야 한다.</li>
<li>현재 행을 처리할 때는 반드시 <strong>바로 위 행</strong>을 0으로 확정해야 한다.</li>
<li>마지막 행은 아래에서 보정할 행이 없으므로, 모든 값이 0인지 직접 검사해야 한다.</li>
<li>보드마다 독립적으로 시뮬레이션해야 하므로, 첫 행 조합을 바꿀 때 원본 배열을 깊은 복사해야 한다.</li>
</ul>
<h2 id="마무리">마무리</h2>
<p>행 단위로 보면 첫 행만 선택이고 나머지는 강제 전이라는 구조가 보인다. 첫 행의 작은 경우의 수를 완전탐색하고, 이후 행을 위에서 아래로 확정해 나가면 최소 조작 횟수를 구할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[71일차] Transformer 구조와 AutoEncoder]]></title>
            <link>https://velog.io/@j_keun/71%EC%9D%BC%EC%B0%A8-Transformer-%EA%B5%AC%EC%A1%B0%EC%99%80-AutoEncoder</link>
            <guid>https://velog.io/@j_keun/71%EC%9D%BC%EC%B0%A8-Transformer-%EA%B5%AC%EC%A1%B0%EC%99%80-AutoEncoder</guid>
            <pubDate>Thu, 01 Oct 2026 17:36:45 GMT</pubDate>
            <description><![CDATA[<p>Transformer는 RNN처럼 Token을 순서대로 처리하지 않고, <strong>Attention</strong>을 이용해 시퀀스 전체의 관계를 한 번에 계산한다. AutoEncoder는 입력 데이터를 압축한 뒤 다시 복원하면서 데이터의 핵심 특징을 학습한다.</p>
<p>이번에는 Transformer의 Multi-Head Attention, 위치 정보, Encoder와 Decoder 구조를 살펴본다. 이어서 MNIST 이미지에 합성곱 AutoEncoder를 적용해 원본 이미지를 복원하는 과정을 정리한다.</p>
<hr>
<h2 id="1-padding-mask">1. Padding Mask</h2>
<p>배치 학습에서는 길이가 서로 다른 문장을 하나의 Tensor로 묶기 위해 짧은 문장 뒤에 <code>&lt;PAD&gt;</code> Token을 추가한다.</p>
<pre><code class="language-text">커피 한잔 어때
안녕 &lt;PAD&gt; &lt;PAD&gt;</code></pre>
<p><code>&lt;PAD&gt;</code>는 길이를 맞추기 위한 자리일 뿐 실제 의미가 없다. Attention이 이 위치를 참고하면 의미 없는 정보가 문맥에 섞이므로 <strong>Padding Mask</strong>로 제외해야 한다.</p>
<pre><code class="language-python">import torch

sample_scores = torch.tensor([[2.0, 1.0, 0.5, 3.0]])
padding_mask = torch.tensor([1, 1, 1, 0])

masked_scores = sample_scores.masked_fill(
    padding_mask == 0,
    float(&quot;-inf&quot;),
)

weights_before = torch.softmax(sample_scores, dim=-1)
weights_after = torch.softmax(masked_scores, dim=-1)

print(&quot;Mask 전:&quot;, weights_before)
print(&quot;Mask 후:&quot;, weights_after)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">Mask 전: tensor([[0.2321, 0.0854, 0.0518, 0.6308]])
Mask 후: tensor([[0.6285, 0.2312, 0.1402, 0.0000]])</code></pre>
<p>Mask가 적용된 Score는 <code>-inf</code>가 된다. <code>softmax(-inf)</code>의 결과는 <code>0</code>이므로 해당 위치의 Value가 최종 출력에 반영되지 않는다.</p>
<pre><code class="language-text">Score:  [2.0, 1.0, 0.5, -inf]
                         ↓ Softmax
Weight: [0.6285, 0.2312, 0.1402, 0.0000]</code></pre>
<hr>
<h2 id="2-transformer">2. Transformer</h2>
<p>Transformer는 2017년 논문 <a href="https://arxiv.org/abs/1706.03762">Attention Is All You Need</a>에서 제안된 신경망 구조다.</p>
<p>RNN 계열 모델은 이전 시점의 계산이 끝나야 다음 시점을 처리할 수 있다. Transformer는 Self-Attention으로 모든 Token의 관계를 직접 계산하므로 병렬 연산에 유리하며, 멀리 떨어진 Token 사이의 관계도 직접 연결할 수 있다.</p>
<pre><code class="language-text">RNN
x1 → h1 → x2 → h2 → x3 → h3

Transformer
x1 ─┬─ x1, x2, x3과 관계 계산
x2 ─┼─ x1, x2, x3과 관계 계산
x3 ─┴─ x1, x2, x3과 관계 계산</code></pre>
<p>대표적인 Transformer 기반 모델에는 BERT, GPT, T5 등이 있다.</p>
<hr>
<h2 id="3-multi-head-attention">3. Multi-Head Attention</h2>
<p>Self-Attention을 Q, K, V 한 세트로만 계산하면 하나의 표현 공간에서 관계를 학습한다. <strong>Multi-Head Attention</strong>은 여러 Head가 서로 다른 Q, K, V 투영을 학습해 다양한 관계를 동시에 포착하도록 만든다.</p>
<p>학습 결과에 따라 어떤 Head는 가까운 Token 관계에, 다른 Head는 문법적 관계나 의미적 관계에 반응할 수 있다. 각 Head의 역할을 사람이 미리 지정하는 것은 아니다.</p>
<pre><code class="language-text">d_model = 512
num_heads = 8
head_dim = 512 / 8 = 64</code></pre>
<p><code>d_model</code>은 <code>num_heads</code>로 나누어떨어져야 한다.</p>
<h3 id="pytorch로-확인하기">PyTorch로 확인하기</h3>
<pre><code class="language-python">import torch
import torch.nn as nn

embed_dim = 8
num_heads = 2
head_dim = embed_dim // num_heads

print(&quot;전체 Embedding 차원:&quot;, embed_dim)
print(&quot;Head 개수:&quot;, num_heads)
print(&quot;Head 하나의 차원:&quot;, head_dim)

# batch=1, token=3, embed_dim=8
x = torch.randn(1, 3, embed_dim)

mha = nn.MultiheadAttention(
    embed_dim=embed_dim,
    num_heads=num_heads,
    batch_first=True,
)

attn_output, attn_weights = mha(
    query=x,
    key=x,
    value=x,
    need_weights=True,
)

print(&quot;입력 Shape:&quot;, x.shape)
print(&quot;출력 Shape:&quot;, attn_output.shape)
print(&quot;Attention Weight Shape:&quot;, attn_weights.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">전체 Embedding 차원: 8
Head 개수: 2
Head 하나의 차원: 4
입력 Shape: torch.Size([1, 3, 8])
출력 Shape: torch.Size([1, 3, 8])
Attention Weight Shape: torch.Size([1, 3, 3])</code></pre>
<p>입력과 출력의 마지막 차원은 모두 <code>8</code>이다. 내부에서는 두 Head가 각각 4차원 공간에서 Attention을 계산한 뒤, 결과를 합쳐 다시 8차원 표현을 만든다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/a09a0894-ae11-49de-bcd5-51369ada1655/image.png" alt=""><img src="https://velog.velcdn.com/images/j_keun/post/6a049111-d97f-4935-9c7b-b4e095b03929/image.png" alt="">
<img src="https://velog.velcdn.com/images/j_keun/post/addc9fc0-42d6-4eff-896c-29359e317265/image.png" alt=""><img src="https://velog.velcdn.com/images/j_keun/post/199a335f-5a7b-4f8d-9a88-1f24fa2c8215/image.png" alt=""></p>
<h3 id="head별-attention-weight-확인하기">Head별 Attention Weight 확인하기</h3>
<p><code>nn.MultiheadAttention</code>은 기본적으로 여러 Head의 Weight를 평균해 반환한다.</p>
<pre><code class="language-text">[batch, target_length, source_length]</code></pre>
<p><code>average_attn_weights=False</code>를 지정하면 Head별 Weight를 따로 확인할 수 있다.</p>
<pre><code class="language-python">attn_output, head_weights = mha(
    query=x,
    key=x,
    value=x,
    need_weights=True,
    average_attn_weights=False,
)

print(head_weights.shape)
print(&quot;Head 0:\n&quot;, head_weights[0, 0])
print(&quot;Head 1:\n&quot;, head_weights[0, 1])</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">torch.Size([1, 2, 3, 3])</code></pre>
<p>차원의 의미는 다음과 같다.</p>
<pre><code class="language-text">[batch, num_heads, target_length, source_length]
[1,     2,         3,             3]</code></pre>
<p>Self-Attention에서는 Query와 Key가 같은 시퀀스에서 만들어지므로 보통 <code>target_length</code>와 <code>source_length</code>가 같다.</p>
<h3 id="attention-weight-시각화">Attention Weight 시각화</h3>
<p><img src="https://velog.velcdn.com/images/j_keun/post/70c4097a-4426-4031-ab40-3436781105bc/image.png" alt=""></p>
<ul>
<li>세로축은 정보를 찾는 Query Token이다.</li>
<li>가로축은 참고 대상인 Key Token이다.</li>
<li>색이 밝을수록 해당 Key를 더 많이 참고한다.</li>
<li>Weight는 학습 과정에서 계속 달라지므로 한 번 출력한 값만으로 Head의 역할을 단정해서는 안 된다.</li>
</ul>
<hr>
<h2 id="4-transformer-encoder-전체-흐름">4. Transformer Encoder 전체 흐름</h2>
<p>Transformer Encoder의 한 Block은 다음 순서로 동작한다.</p>
<pre><code class="language-text">문장
  ↓
Tokenization
  ↓
Token ID
  ↓
Token Embedding + Positional Encoding
  ↓
Multi-Head Self-Attention
  ↓
Residual Connection + LayerNorm
  ↓
Feed Forward Network
  ↓
Residual Connection + LayerNorm
  ↓
Encoder 출력</code></pre>
<p>각 단계는 입력과 출력의 Shape을 유지하면서 Token 표현을 점차 문맥에 맞게 바꾼다.</p>
<hr>
<h2 id="5-token-embedding">5. Token Embedding</h2>
<p>먼저 Token을 정수 ID로 바꾼다.</p>
<pre><code class="language-python">vocab = {
    &quot;&lt;PAD&gt;&quot;: 0,
    &quot;커피&quot;: 1,
    &quot;한잔&quot;: 2,
    &quot;어때&quot;: 3,
}

tokens = [&quot;커피&quot;, &quot;한잔&quot;, &quot;어때&quot;]
token_ids = torch.tensor([[1, 2, 3]])

print(token_ids)
print(token_ids.shape)</code></pre>
<pre><code class="language-text">tensor([[1, 2, 3]])
torch.Size([1, 3])</code></pre>
<p><code>nn.Embedding</code>은 각 Token ID를 <code>d_model</code>차원의 실수 벡터로 바꾼다.</p>
<pre><code class="language-python">vocab_size = len(vocab)
d_model = 8

embedding = nn.Embedding(
    vocab_size,
    d_model,
    padding_idx=0,
)

X = embedding(token_ids)
print(X.shape)</code></pre>
<pre><code class="language-text">torch.Size([1, 3, 8])</code></pre>
<p>Shape은 다음처럼 변한다.</p>
<pre><code class="language-text">[batch, sequence_length]
[1, 3]
      ↓ Embedding
[batch, sequence_length, d_model]
[1, 3, 8]</code></pre>
<hr>
<h2 id="6-positional-encoding">6. Positional Encoding</h2>
<p>Self-Attention은 모든 Token의 관계를 한 번에 계산하므로 RNN처럼 계산 순서에서 위치 정보를 얻지 못한다. 따라서 입력에 Token의 위치 정보를 별도로 넣어야 한다.</p>
<p>원래 Transformer 논문은 고정된 Sin과 Cos 함수를 이용한 Positional Encoding을 사용한다.</p>
<pre><code class="language-text">최종 입력 = Token Embedding + Positional Encoding</code></pre>
<p>예를 들어 다음처럼 같은 차원의 두 벡터를 더한다.</p>
<pre><code class="language-text">Token Embedding:      [0.20, -0.40, 0.70, ...]
Position 0 Encoding:  [0.00,  1.00, 0.00, ...]
최종 입력:            [0.20,  0.60, 0.70, ...]</code></pre>
<p>모델이 각 숫자를 Token 정보와 위치 정보로 따로 읽는 것은 아니다. 같은 <code>d_model</code> 공간에서 두 벡터를 더해 <strong>의미와 위치가 함께 반영된 표현</strong>을 만든다.</p>
<h3 id="sincos-positional-encoding-구현">Sin/Cos Positional Encoding 구현</h3>
<pre><code class="language-python">import math

class PositionalEncoding(nn.Module):
    def __init__(self, d_model, max_len=100):
        super().__init__()

        position = torch.arange(
            max_len,
            dtype=torch.float32,
        ).unsqueeze(1)

        div_term = torch.exp(
            torch.arange(
                0,
                d_model,
                2,
                dtype=torch.float32,
            )
            * (-math.log(10000.0) / d_model)
        )

        pe = torch.zeros(max_len, d_model)
        pe[:, 0::2] = torch.sin(position * div_term)
        pe[:, 1::2] = torch.cos(position * div_term)

        self.register_buffer(&quot;pe&quot;, pe.unsqueeze(0))

    def forward(self, x):
        seq_len = x.size(1)
        return x + self.pe[:, :seq_len]</code></pre>
<p><code>pe</code>는 학습으로 바뀌는 Parameter가 아니라 공식으로 계산한 고정값이다. <code>register_buffer()</code>로 등록하면 다음 특성을 얻는다.</p>
<ul>
<li><code>model.to(device)</code>를 호출할 때 모델과 함께 GPU나 MPS로 이동한다.</li>
<li><code>state_dict()</code>에 포함되어 저장된다.</li>
<li>Optimizer의 학습 대상 Parameter에는 포함되지 않는다.</li>
</ul>
<pre><code class="language-python">pos_encoding = PositionalEncoding(d_model)
X_pos = pos_encoding(X)

print(&quot;Embedding:&quot;, X.shape)
print(&quot;Position 추가:&quot;, X_pos.shape)</code></pre>
<pre><code class="language-text">Embedding: torch.Size([1, 3, 8])
Position 추가: torch.Size([1, 3, 8])</code></pre>
<p>위치 정보를 더해도 Shape은 변하지 않는다.</p>
<h3 id="rope">RoPE</h3>
<p>RoPE(Rotary Positional Embedding)는 위치 벡터를 입력 Embedding에 단순히 더하지 않는다. Attention에 사용하는 Q와 K를 Token 위치에 따라 회전시켜 상대적인 위치 관계가 점수에 반영되도록 만든다.</p>
<pre><code class="language-text">Sin/Cos Positional Encoding
Embedding에 위치 벡터를 더한다.

RoPE
Q와 K의 Attention 계산에 위치 관계를 반영한다.</code></pre>
<hr>
<h2 id="7-multi-head-self-attention-적용">7. Multi-Head Self-Attention 적용</h2>
<p>위치 정보가 더해진 <code>X_pos</code>를 Query, Key, Value에 모두 사용하면 Self-Attention이 된다.</p>
<pre><code class="language-python">num_heads = 2

mha = nn.MultiheadAttention(
    embed_dim=d_model,
    num_heads=num_heads,
    batch_first=True,
)

attn_out, attn_weights = mha(
    query=X_pos,
    key=X_pos,
    value=X_pos,
    need_weights=True,
    average_attn_weights=False,
)

print(&quot;입력:&quot;, X_pos.shape)
print(&quot;출력:&quot;, attn_out.shape)
print(&quot;Head별 Weight:&quot;, attn_weights.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">입력: torch.Size([1, 3, 8])
출력: torch.Size([1, 3, 8])
Head별 Weight: torch.Size([1, 2, 3, 3])</code></pre>
<p>Attention은 Token 사이의 정보를 섞지만 <code>[batch, sequence_length, d_model]</code> Shape은 유지한다.</p>
<hr>
<h2 id="8-residual-connection">8. Residual Connection</h2>
<p>Residual Connection은 Attention 결과만 다음 단계로 보내지 않고 원래 입력을 다시 더하는 구조다.</p>
<pre><code class="language-text">y = x + F(x)</code></pre>
<ul>
<li><code>x</code>는 Attention에 들어가기 전의 원래 정보다.</li>
<li><code>F(x)</code>는 Attention이 계산한 변화 정보다.</li>
<li><code>x + F(x)</code>는 원래 정보를 유지하면서 새로운 문맥 정보를 추가한 결과다.</li>
</ul>
<pre><code class="language-python">residual = X_pos + attn_out
print(residual.shape)</code></pre>
<pre><code class="language-text">torch.Size([1, 3, 8])</code></pre>
<p>원소별 덧셈을 수행하므로 <code>x</code>와 <code>F(x)</code>의 Shape이 같아야 한다. Residual Connection은 깊은 신경망에서 정보와 Gradient가 전달될 경로를 제공해 학습을 돕는다.</p>
<hr>
<h2 id="9-layer-normalization">9. Layer Normalization</h2>
<p><code>nn.LayerNorm(d_model)</code>은 각 Token 벡터의 Feature 차원을 기준으로 값을 정규화한다.</p>
<p>입력 Shape이 <code>[batch, sequence_length, d_model]</code>이면 각 Token의 <code>(d_model,)</code> 벡터를 독립적으로 정규화한다.</p>
<pre><code class="language-python">norm1 = nn.LayerNorm(d_model)
x1 = norm1(residual)

print(&quot;평균:&quot;, x1[0, 0].mean())
print(&quot;분산:&quot;, x1[0, 0].var(unbiased=False))</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">평균: tensor(0.)
분산: tensor(1.000)</code></pre>
<p>LayerNorm은 각 샘플과 Token 안에서 정규화하므로 BatchNorm보다 배치 크기의 영향을 덜 받는다. 가변 길이 시퀀스를 다루는 Transformer에 잘 맞는다.</p>
<hr>
<h2 id="10-feed-forward-network">10. Feed Forward Network</h2>
<p>Attention과 FFN은 서로 다른 역할을 한다.</p>
<table>
<thead>
<tr>
<th>구성 요소</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td>Attention</td>
<td>다른 Token에서 어떤 정보를 가져올지 계산한다.</td>
</tr>
<tr>
<td>FFN</td>
<td>모은 정보를 각 Token 위치에서 독립적으로 비선형 변환한다.</td>
</tr>
</tbody></table>
<pre><code class="language-python">ffn = nn.Sequential(
    nn.Linear(d_model, 32),
    nn.ReLU(),
    nn.Linear(32, d_model),
)

ffn_out = ffn(x1)

print(&quot;입력:&quot;, x1.shape)
print(&quot;출력:&quot;, ffn_out.shape)</code></pre>
<pre><code class="language-text">입력: torch.Size([1, 3, 8])
출력: torch.Size([1, 3, 8])</code></pre>
<p>동일한 FFN이 모든 Token 위치에 적용되지만, 각 Token은 다른 Token과 섞이지 않고 자신의 Feature만 변환한다.</p>
<p>Encoder Block에는 두 개의 Sub-layer가 있으며 각각 Residual Connection과 LayerNorm이 붙는다.</p>
<pre><code class="language-text">Self-Attention
  ↓
Residual + LayerNorm
  ↓
FFN
  ↓
Residual + LayerNorm</code></pre>
<pre><code class="language-python">norm2 = nn.LayerNorm(d_model)
encoder_output = norm2(x1 + ffn_out)

print(encoder_output.shape)</code></pre>
<pre><code class="language-text">torch.Size([1, 3, 8])</code></pre>
<hr>
<h2 id="11-encoder에서-padding-mask-사용하기">11. Encoder에서 Padding Mask 사용하기</h2>
<p>PyTorch의 <code>nn.MultiheadAttention</code>은 <code>key_padding_mask</code>에서 다음 규칙을 사용한다.</p>
<pre><code class="language-text">False: 참고할 수 있는 실제 Token
True:  참고 대상에서 제외할 PAD Token</code></pre>
<pre><code class="language-python">padded_ids = torch.tensor([
    [1, 2, 3],
    [1, 3, 0],
])

padding_mask = padded_ids.eq(0)

print(padded_ids)
print(padding_mask)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">tensor([[1, 2, 3],
        [1, 3, 0]])

tensor([[False, False, False],
        [False, False,  True]])</code></pre>
<p>Mask를 Multi-Head Attention에 전달한다.</p>
<pre><code class="language-python">padded_x = pos_encoding(embedding(padded_ids))

masked_out, masked_weights = mha(
    padded_x,
    padded_x,
    padded_x,
    key_padding_mask=padding_mask,
    need_weights=True,
)

print(masked_weights[1])</code></pre>
<p>두 번째 문장의 Attention Weight는 다음과 같다.</p>
<pre><code class="language-text">tensor([[0.392, 0.608, 0.000],
        [0.372, 0.628, 0.000],
        [0.414, 0.586, 0.000]])</code></pre>
<p>마지막 열은 <code>&lt;PAD&gt;</code>가 있는 Key 위치다. 모든 Query에서 가중치가 <code>0</code>이므로 참고 대상에서 제외된 것을 확인할 수 있다.</p>
<blockquote>
<p><code>key_padding_mask</code>는 PAD를 Key와 Value의 참고 대상에서 제외한다. PAD 위치의 Query 출력까지 자동으로 삭제하는 것은 아니므로, 이후 Loss를 계산할 때도 PAD 위치를 제외해야 한다.</p>
</blockquote>
<hr>
<h2 id="12-causal-mask">12. Causal Mask</h2>
<p>Padding Mask와 Causal Mask는 목적이 다르다.</p>
<table>
<thead>
<tr>
<th>Mask</th>
<th>가리는 대상</th>
<th>사용 목적</th>
</tr>
</thead>
<tbody><tr>
<td>Padding Mask</td>
<td>의미 없는 <code>&lt;PAD&gt;</code> 위치</td>
<td>길이가 다른 문장을 배치로 처리한다.</td>
</tr>
<tr>
<td>Causal Mask</td>
<td>현재보다 뒤에 있는 미래 Token</td>
<td>다음 Token의 정답 누출을 막는다.</td>
</tr>
</tbody></table>
<p>문장 <code>나는 커피를 마신다</code>를 학습할 때 각 위치가 볼 수 있는 범위는 다음과 같다.</p>
<pre><code class="language-text">Query  | 나는 | 커피를 | 마신다
-------|------|--------|-------
나는   |  O   |   X    |   X
커피를 |  O   |   O    |   X
마신다 |  O   |   O    |   O</code></pre>
<p>추론할 때는 Token을 하나씩 생성하지만, 학습할 때는 정답 문장 전체를 알고 있다. Mask 없이 병렬 계산하면 앞쪽 위치가 미래의 정답 Token을 미리 볼 수 있으므로 Causal Mask가 필요하다.</p>
<pre><code class="language-python">seq_len = 4

causal_mask = torch.triu(
    torch.ones(seq_len, seq_len, dtype=torch.bool),
    diagonal=1,
)

print(causal_mask)</code></pre>
<pre><code class="language-text">tensor([[False,  True,  True,  True],
        [False, False,  True,  True],
        [False, False, False,  True],
        [False, False, False, False]])</code></pre>
<p><code>True</code>인 오른쪽 위 영역이 미래 Token에 해당한다.</p>
<p>실제 Attention Weight는 다음처럼 미래 위치가 <code>0</code>이 된다.</p>
<pre><code class="language-text">tensor([[1.000, 0.000, 0.000, 0.000],
        [0.640, 0.360, 0.000, 0.000],
        [0.315, 0.311, 0.374, 0.000],
        [0.166, 0.159, 0.303, 0.372]])</code></pre>
<hr>
<h2 id="13-transformer-encoder-block-구현">13. Transformer Encoder Block 구현</h2>
<p>지금까지 확인한 구성 요소를 하나의 Encoder Block으로 묶는다.</p>
<pre><code class="language-python">class SimpleTransformerEncoderBlock(nn.Module):
    def __init__(
        self,
        d_model=8,
        num_heads=2,
        d_ff=32,
        dropout=0.1,
    ):
        super().__init__()

        self.self_attention = nn.MultiheadAttention(
            d_model,
            num_heads,
            dropout=dropout,
            batch_first=True,
        )

        self.norm1 = nn.LayerNorm(d_model)
        self.norm2 = nn.LayerNorm(d_model)

        self.ffn = nn.Sequential(
            nn.Linear(d_model, d_ff),
            nn.ReLU(),
            nn.Dropout(dropout),
            nn.Linear(d_ff, d_model),
        )

        self.dropout1 = nn.Dropout(dropout)
        self.dropout2 = nn.Dropout(dropout)

    def forward(self, x, padding_mask=None):
        attn_out, weights = self.self_attention(
            x,
            x,
            x,
            key_padding_mask=padding_mask,
            need_weights=True,
        )

        x = self.norm1(
            x + self.dropout1(attn_out)
        )

        ffn_out = self.ffn(x)
        x = self.norm2(
            x + self.dropout2(ffn_out)
        )

        return x, weights</code></pre>
<pre><code class="language-python">block = SimpleTransformerEncoderBlock()
out, weights = block(X_pos)

print(&quot;입력:&quot;, X_pos.shape)
print(&quot;출력:&quot;, out.shape)
print(&quot;Weight:&quot;, weights.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">입력: torch.Size([1, 3, 8])
출력: torch.Size([1, 3, 8])
Weight: torch.Size([1, 3, 3])</code></pre>
<p>이 구현은 Attention 뒤에 정규화하는 <strong>Post-Norm</strong> 구조다.</p>
<pre><code class="language-text">x → Attention → Add → LayerNorm → FFN → Add → LayerNorm</code></pre>
<hr>
<h2 id="14-transformer-decoder">14. Transformer Decoder</h2>
<p>Encoder는 입력 Token 전체를 Self-Attention으로 처리해 문맥이 반영된 표현을 만든다. Decoder는 이 정보를 이용해 출력 문장을 생성한다.</p>
<p>Decoder의 주요 구성은 다음과 같다.</p>
<ol>
<li><strong>Masked Self-Attention</strong>: 미래 출력 Token을 미리 보지 못하게 한다.</li>
<li><strong>Cross-Attention</strong>: Encoder 출력 중 현재 생성에 필요한 부분을 참고한다.</li>
<li><strong>FFN</strong>: 각 Decoder Token 표현을 비선형 변환한다.</li>
</ol>
<pre><code class="language-text">입력 문장
   ↓
Encoder
   ↓
Encoder Memory ──────────────┐
                            │ K, V
출력 입력                    │
   ↓                        │
Masked Self-Attention        │
   ↓                        │
Cross-Attention ◀────────────┘
   ↓
FFN
   ↓
Decoder 출력</code></pre>
<h3 id="cross-attention">Cross-Attention</h3>
<p>Cross-Attention에서는 Q와 K, V의 출처가 다르다.</p>
<pre><code class="language-text">Q = Decoder의 현재 표현
K = Encoder 출력
V = Encoder 출력</code></pre>
<pre><code class="language-python">cross_attention = nn.MultiheadAttention(
    embed_dim=8,
    num_heads=2,
    batch_first=True,
)

encoder_memory = torch.randn(1, 5, 8)
decoder_state = torch.randn(1, 3, 8)

cross_out, cross_weights = cross_attention(
    query=decoder_state,
    key=encoder_memory,
    value=encoder_memory,
)

print(&quot;Q(Decoder):&quot;, decoder_state.shape)
print(&quot;K/V(Encoder):&quot;, encoder_memory.shape)
print(&quot;출력:&quot;, cross_out.shape)
print(&quot;Weight:&quot;, cross_weights.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">Q(Decoder): torch.Size([1, 3, 8])
K/V(Encoder): torch.Size([1, 5, 8])
출력: torch.Size([1, 3, 8])
Weight: torch.Size([1, 3, 5])</code></pre>
<p><code>[1, 3, 5]</code>는 Decoder의 Token 3개가 Encoder Token 5개를 각각 얼마나 참고했는지 나타낸다.</p>
<p>GPT 같은 Decoder-only 모델은 별도의 Encoder와 Cross-Attention 없이 Causal Self-Attention을 중심으로 구성한다.</p>
<hr>
<h2 id="15-autoencoder">15. AutoEncoder</h2>
<p>AutoEncoder는 입력을 작은 잠재 표현으로 바꾸고, 그 표현으로 원본을 다시 복원하는 신경망이다.</p>
<pre><code class="language-text">입력 x
  ↓
Encoder
  ↓
Latent Representation z
  ↓
Decoder
  ↓
복원 결과 x_hat</code></pre>
<p><img src="https://velog.velcdn.com/images/j_keun/post/16c2d43b-239a-4f0f-9fb3-db25b7b1e09f/image.png" alt=""></p>
<p>입력과 복원 결과의 차이가 작아지도록 학습한다.</p>
<pre><code class="language-text">입력:  숫자 7 이미지
정답:  같은 숫자 7 이미지
출력:  AutoEncoder가 복원한 숫자 7 이미지</code></pre>
<h3 id="비지도-학습과-자기지도-학습">비지도 학습과 자기지도 학습</h3>
<p>이미지 분류 모델은 <code>강아지</code>, <code>고양이</code>와 같은 사람이 만든 클래스 라벨을 필요로 한다. AutoEncoder는 별도의 클래스 라벨 없이 입력 데이터 자체를 정답으로 사용한다.</p>
<pre><code class="language-python">for x, _ in loader:
    _, x_hat = model(x)
    loss = criterion(x_hat, x)</code></pre>
<p>DataLoader가 제공한 <code>label</code>을 <code>_</code>로 받아 사용하지 않고, 입력 <code>x</code>가 그대로 학습 Target이 된다. 전통적으로 비지도 학습으로 분류하며, 입력에서 정답을 자동으로 만드는 관점에서는 자기지도 학습의 한 형태로도 볼 수 있다.</p>
<hr>
<h2 id="16-encoder-latent-representation-decoder">16. Encoder, Latent Representation, Decoder</h2>
<h3 id="encoder">Encoder</h3>
<p>Encoder는 입력에서 복원에 필요한 특징을 추출한다.</p>
<pre><code class="language-text">입력 이미지
[1, 28, 28]
      ↓ Encoder
잠재 표현
[32, 7, 7]</code></pre>
<h3 id="latent-representation">Latent Representation</h3>
<p>Latent Representation은 Encoder가 만든 내부 표현이다. 원본을 그대로 복사하는 대신 Decoder가 원본을 다시 만들 때 필요한 특징을 담는다.</p>
<p>이번 합성곱 모델은 공간 크기를 <code>28 × 28</code>에서 <code>7 × 7</code>로 줄이는 대신 Channel을 <code>1</code>에서 <code>32</code>로 늘린다.</p>
<pre><code class="language-text">원본 원소 수:  1 × 28 × 28 = 784
Latent 원소 수: 32 × 7 × 7 = 1,568</code></pre>
<p>따라서 이 예제는 <strong>가로와 세로 크기는 압축하지만 전체 원소 수가 더 작은 엄격한 차원 축소 구조는 아니다.</strong> 더 강한 Bottleneck이 필요하다면 Channel 수를 줄이거나 <code>Linear</code> 층으로 더 작은 Latent Vector를 만들 수 있다.</p>
<h3 id="decoder">Decoder</h3>
<p>Decoder는 Latent Representation만 보고 원본과 비슷한 데이터를 복원한다.</p>
<pre><code class="language-text">[32, 7, 7]
      ↓ Decoder
[1, 28, 28]</code></pre>
<hr>
<h2 id="17-reconstruction-loss">17. Reconstruction Loss</h2>
<p>AutoEncoder는 원본과 복원본의 차이를 <strong>Reconstruction Loss</strong>로 계산한다.</p>
<p>예를 들어 한 픽셀의 원본값이 <code>0.8</code>, 복원값이 <code>0.6</code>이라면 제곱 오차는 다음과 같다.</p>
<pre><code class="language-text">(0.8 - 0.6)² = 0.04</code></pre>
<p>전체 픽셀의 평균 제곱 오차를 줄이도록 Encoder와 Decoder의 Parameter를 함께 업데이트한다.</p>
<pre><code class="language-python">criterion = nn.MSELoss()</code></pre>
<hr>
<h2 id="18-mnist-데이터-준비">18. MNIST 데이터 준비</h2>
<pre><code class="language-python">import torch
import torch.nn as nn
import torch.optim as optim
import matplotlib.pyplot as plt

from torch.utils.data import DataLoader
from torchvision import datasets, transforms
from torchvision.utils import make_grid

torch.manual_seed(2026)</code></pre>
<p><code>ToTensor()</code>는 MNIST 이미지를 Tensor로 바꾸고 픽셀 범위를 <code>[0, 1]</code>로 변환한다.</p>
<pre><code class="language-python">transform = transforms.ToTensor()

train_dataset = datasets.MNIST(
    root=&quot;./data&quot;,
    train=True,
    download=True,
    transform=transform,
)

test_dataset = datasets.MNIST(
    root=&quot;./data&quot;,
    train=False,
    download=True,
    transform=transform,
)

train_loader = DataLoader(
    train_dataset,
    batch_size=128,
    shuffle=True,
    num_workers=4,
)

test_loader = DataLoader(
    test_dataset,
    batch_size=128,
    shuffle=False,
    num_workers=4,
)</code></pre>
<pre><code class="language-python">images, labels = next(iter(train_loader))

print(&quot;images:&quot;, images.shape)
print(&quot;labels:&quot;, labels.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">images: torch.Size([128, 1, 28, 28])
labels: torch.Size([128])</code></pre>
<p><code>images</code>의 차원은 <code>[batch, channel, height, width]</code> 순서다.</p>
<hr>
<h2 id="19-합성곱-autoencoder-구현">19. 합성곱 AutoEncoder 구현</h2>
<pre><code class="language-python">class ConvAutoencoder(nn.Module):
    def __init__(self):
        super().__init__()

        self.encoder = nn.Sequential(
            # [B, 1, 28, 28] → [B, 16, 14, 14]
            nn.Conv2d(
                1,
                16,
                kernel_size=3,
                stride=2,
                padding=1,
            ),
            nn.ReLU(),

            # [B, 16, 14, 14] → [B, 32, 7, 7]
            nn.Conv2d(
                16,
                32,
                kernel_size=3,
                stride=2,
                padding=1,
            ),
            nn.ReLU(),
        )

        self.decoder = nn.Sequential(
            # [B, 32, 7, 7] → [B, 16, 14, 14]
            nn.ConvTranspose2d(
                32,
                16,
                kernel_size=3,
                stride=2,
                padding=1,
                output_padding=1,
            ),
            nn.ReLU(),

            # [B, 16, 14, 14] → [B, 1, 28, 28]
            nn.ConvTranspose2d(
                16,
                1,
                kernel_size=3,
                stride=2,
                padding=1,
                output_padding=1,
            ),
            nn.Sigmoid(),
        )

    def forward(self, x):
        z = self.encoder(x)
        x_hat = self.decoder(z)
        return z, x_hat</code></pre>
<h3 id="shape-변화">Shape 변화</h3>
<table>
<thead>
<tr>
<th>단계</th>
<th>연산</th>
<th>출력 Shape</th>
</tr>
</thead>
<tbody><tr>
<td>입력</td>
<td>MNIST 이미지</td>
<td><code>[B, 1, 28, 28]</code></td>
</tr>
<tr>
<td>Encoder 1</td>
<td><code>Conv2d(1, 16)</code></td>
<td><code>[B, 16, 14, 14]</code></td>
</tr>
<tr>
<td>Encoder 2</td>
<td><code>Conv2d(16, 32)</code></td>
<td><code>[B, 32, 7, 7]</code></td>
</tr>
<tr>
<td>Decoder 1</td>
<td><code>ConvTranspose2d(32, 16)</code></td>
<td><code>[B, 16, 14, 14]</code></td>
</tr>
<tr>
<td>Decoder 2</td>
<td><code>ConvTranspose2d(16, 1)</code></td>
<td><code>[B, 1, 28, 28]</code></td>
</tr>
</tbody></table>
<p><code>stride=2</code>인 <code>Conv2d</code>는 공간 크기를 절반으로 줄인다. <code>ConvTranspose2d</code>는 반대로 공간 크기를 키운다.</p>
<p>마지막 <code>Sigmoid</code>는 복원 픽셀값을 <code>0</code>과 <code>1</code> 사이로 만든다. 입력도 <code>ToTensor()</code>로 <code>[0, 1]</code> 범위이므로 출력 범위가 잘 맞는다.</p>
<hr>
<h2 id="20-autoencoder-학습">20. AutoEncoder 학습</h2>
<pre><code class="language-python">model = ConvAutoencoder().to(DEVICE)

criterion = nn.MSELoss()
optimizer = optim.AdamW(
    model.parameters(),
    lr=1e-3,
)</code></pre>
<p>학습 함수는 입력 <code>x</code>와 복원 결과 <code>x_hat</code>의 차이를 계산한다.</p>
<pre><code class="language-python">def train_autoencoder(
    model,
    loader,
    criterion,
    optimizer,
    device,
    epochs=5,
):
    history = []

    for epoch in range(epochs):
        model.train()
        running_loss = 0.0

        for x, _ in loader:
            x = x.to(device)

            _, x_hat = model(x)
            loss = criterion(x_hat, x)

            optimizer.zero_grad()
            loss.backward()
            optimizer.step()

            running_loss += loss.item() * x.size(0)

        epoch_loss = running_loss / len(loader.dataset)
        history.append(epoch_loss)

        print(
            f&quot;Epoch {epoch + 1:02d}/{epochs} | &quot;
            f&quot;Loss: {epoch_loss:.6f}&quot;
        )

    return history</code></pre>
<p>노트북의 실행 결과는 다음과 같다.</p>
<pre><code class="language-text">Epoch 01/5 | Loss: 0.031372
Epoch 02/5 | Loss: 0.001786
Epoch 03/5 | Loss: 0.001211
Epoch 04/5 | Loss: 0.000970
Epoch 05/5 | Loss: 0.000810</code></pre>
<p><img src="https://velog.velcdn.com/images/j_keun/post/7ce2cb55-8814-4a0c-a2d2-862ad8d000fe/image.png" alt=""></p>
<p>Reconstruction Loss가 지속해서 감소하므로 원본을 더 비슷하게 복원하도록 학습되고 있음을 알 수 있다.</p>
<hr>
<h2 id="21-복원-결과-확인">21. 복원 결과 확인</h2>
<pre><code class="language-python">@torch.no_grad()
def show_reconstructions(model, loader, device, n=8):
    model.eval()

    x, _ = next(iter(loader))
    x = x[:n].to(device)

    _, x_hat = model(x)

    comparison = torch.cat(
        [x.cpu(), x_hat.cpu()],
        dim=0,
    )

    grid = make_grid(
        comparison,
        nrow=n,
        padding=2,
    )

    plt.figure(figsize=(14, 4))
    plt.imshow(
        grid.permute(1, 2, 0).squeeze(),
        cmap=&quot;gray&quot;,
    )
    plt.axis(&quot;off&quot;)
    plt.title(&quot;Top: Original / Bottom: Reconstruction&quot;)
    plt.show()</code></pre>
<p><img src="https://velog.velcdn.com/images/j_keun/post/ce48e111-d18e-4e6b-b5d3-5cc71e689dd8/image.png" alt=""></p>
<ul>
<li>위쪽 행은 원본 이미지다.</li>
<li>아래쪽 행은 AutoEncoder가 복원한 이미지다.</li>
<li>전체적인 숫자 모양은 잘 복원하지만 가장자리 일부가 부드럽거나 흐리게 표현될 수 있다.</li>
</ul>
<p>학습 Loss가 낮아도 새로운 데이터의 복원 품질이 항상 좋다는 뜻은 아니다. Test 데이터의 Reconstruction Loss와 복원 이미지를 함께 확인해야 한다.</p>
<hr>
<h2 id="22-autoencoder-활용-분야">22. AutoEncoder 활용 분야</h2>
<table>
<thead>
<tr>
<th>활용 분야</th>
<th>사용 방법</th>
</tr>
</thead>
<tbody><tr>
<td>특징 추출</td>
<td>Encoder가 만든 Latent Representation을 다른 모델의 입력으로 사용한다.</td>
</tr>
<tr>
<td>차원 축소</td>
<td>원본보다 작은 Bottleneck을 만들어 핵심 정보를 압축한다.</td>
</tr>
<tr>
<td>노이즈 제거</td>
<td>노이즈가 섞인 이미지를 입력하고 깨끗한 이미지를 Target으로 사용한다.</td>
</tr>
<tr>
<td>이상 탐지</td>
<td>정상 데이터로 학습한 뒤 Reconstruction Error가 큰 데이터를 이상으로 판단한다.</td>
</tr>
<tr>
<td>생성 모델 기초</td>
<td>VAE처럼 Latent Space를 확률적으로 모델링하는 구조로 확장한다.</td>
</tr>
</tbody></table>
<p>일반 AutoEncoder는 입력을 복원하도록 학습하지만, 임의의 Latent Vector에서 항상 자연스러운 데이터를 생성할 수 있는 것은 아니다. 생성 목적이라면 Latent Space의 분포를 학습하는 VAE 같은 구조가 더 적합할 수 있다.</p>
<hr>
<h2 id="23-핵심-정리">23. 핵심 정리</h2>
<ol>
<li>Padding Mask는 <code>&lt;PAD&gt;</code>의 Attention Weight를 <code>0</code>으로 만들어 의미 없는 위치를 참고하지 않게 한다.</li>
<li>Multi-Head Attention은 여러 표현 공간에서 Token 관계를 동시에 학습한다.</li>
<li><code>average_attn_weights=False</code>를 사용하면 Head별 Attention Weight를 확인할 수 있다.</li>
<li>Self-Attention은 순서를 모르므로 Token Embedding에 Positional Encoding을 더한다.</li>
<li>Residual Connection은 원래 정보에 Attention이나 FFN의 변화 정보를 더한다.</li>
<li>LayerNorm은 각 Token의 Feature 차원을 기준으로 정규화한다.</li>
<li>FFN은 각 Token 위치에 독립적으로 같은 비선형 변환을 적용한다.</li>
<li>Causal Mask는 미래 Token을 가려 다음 Token 예측의 정답 누출을 방지한다.</li>
<li>Decoder의 Cross-Attention은 Decoder 표현을 Q로, Encoder 출력을 K와 V로 사용한다.</li>
<li>AutoEncoder는 입력 자체를 Target으로 사용해 Encoder와 Decoder를 함께 학습한다.</li>
<li>MNIST 합성곱 AutoEncoder는 <code>28 × 28 → 7 × 7 → 28 × 28</code>로 공간 크기를 줄였다가 복원한다.</li>
<li>AutoEncoder는 Reconstruction Loss뿐 아니라 실제 복원 결과도 함께 평가해야 한다.</li>
</ol>
]]></description>
        </item>
        <item>
            <title><![CDATA[[70일차] LSTM, GRU와 트랜스포머의 Attention]]></title>
            <link>https://velog.io/@j_keun/70%EC%9D%BC%EC%B0%A8-LSTM-GRU%EC%99%80-%ED%8A%B8%EB%9E%9C%EC%8A%A4%ED%8F%AC%EB%A8%B8%EC%9D%98-Attention</link>
            <guid>https://velog.io/@j_keun/70%EC%9D%BC%EC%B0%A8-LSTM-GRU%EC%99%80-%ED%8A%B8%EB%9E%9C%EC%8A%A4%ED%8F%AC%EB%A8%B8%EC%9D%98-Attention</guid>
            <pubDate>Wed, 30 Sep 2026 15:32:18 GMT</pubDate>
            <description><![CDATA[<p> 기본 RNN은 이전 시점의 정보를 <code>hidden state</code>에 담아 다음 시점으로 전달한다. 하지만 시퀀스가 길어지면 오래전 정보가 제대로 학습되지 않는 <strong>장기 의존성 문제</strong>가 발생할 수 있다.</p>
<p>이번에는 기본 RNN의 한계를 보완한 <strong>LSTM</strong>과 <strong>GRU</strong>를 살펴보고, 입력 시퀀스를 출력 시퀀스로 바꾸는 <strong>Seq2Seq</strong>, 필요한 정보에 직접 집중하는 <strong>Attention</strong>까지 이어서 정리한다.</p>
<hr>
<h2 id="1-lstm">1. LSTM</h2>
<p>LSTM(Long Short-Term Memory)은 기본 RNN의 장기 의존성 문제를 완화하기 위해 만든 순환 신경망이다.</p>
<p>기본 RNN은 주로 하나의 <code>hidden state</code>로 과거 정보를 전달하지만, LSTM은 다음 두 상태를 함께 사용한다.</p>
<table>
<thead>
<tr>
<th>상태</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td><code>hidden state(h_t)</code></td>
<td>현재 시점의 출력과 다음 시점의 계산에 직접 사용되는 단기 표현이다.</td>
</tr>
<tr>
<td><code>cell state(C_t)</code></td>
<td>중요한 정보를 여러 시점 동안 비교적 안정적으로 전달하는 기억 통로다.</td>
</tr>
</tbody></table>
<p>LSTM은 세 개의 게이트로 정보의 흐름을 조절한다.</p>
<table>
<thead>
<tr>
<th>게이트</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td>Forget Gate</td>
<td>과거 기억을 얼마나 남길지 결정한다.</td>
</tr>
<tr>
<td>Input Gate</td>
<td>현재 입력에서 새로운 정보를 얼마나 저장할지 결정한다.</td>
</tr>
<tr>
<td>Output Gate</td>
<td>갱신한 기억 중 현재 출력으로 얼마나 보낼지 결정한다.</td>
</tr>
</tbody></table>
<p><img src="https://velog.velcdn.com/images/j_keun/post/73011227-a282-4430-9c3a-1b526e36d0af/image.png" alt=""></p>
<h3 id="lstm에-들어오는-값">LSTM에 들어오는 값</h3>
<p>시점 <code>t</code>의 LSTM은 다음 정보를 받는다.</p>
<pre><code class="language-text">x_t       : 현재 시점의 입력
h_(t-1)   : 이전 시점의 hidden state
C_(t-1)   : 이전 시점의 cell state</code></pre>
<p>게이트는 사람이 직접 정한 규칙이 아니다. 현재 입력과 이전 상태를 사용해 계산한 뒤, 학습 과정에서 가중치와 편향을 수정하면서 필요한 정보의 비율을 스스로 결정한다.</p>
<hr>
<h2 id="2-lstm의-게이트-동작">2. LSTM의 게이트 동작</h2>
<h3 id="forget-gate">Forget Gate</h3>
<p>Forget Gate는 이전 <code>cell state</code>에서 각 정보를 얼마나 유지할지 결정한다.</p>
<pre><code class="language-text">f_t = sigmoid(W_f · [h_(t-1), x_t] + b_f)</code></pre>
<p><code>sigmoid</code>의 결과는 <code>0</code>과 <code>1</code> 사이의 값이다.</p>
<ul>
<li><code>0</code>에 가까우면 해당 정보를 거의 전달하지 않는다.</li>
<li><code>1</code>에 가까우면 해당 정보를 대부분 유지한다.</li>
<li>중간값이면 그 비율만큼 정보를 전달한다.</li>
</ul>
<p>따라서 Forget Gate는 단순히 기억과 삭제 중 하나만 고르는 스위치가 아니라, 정보를 <strong>얼마나 남길지 조절하는 벡터</strong>다.</p>
<h3 id="input-gate">Input Gate</h3>
<p>Input Gate는 현재 입력에서 어떤 정보를 새 기억으로 저장할지 결정한다.</p>
<pre><code class="language-text">i_t       = sigmoid(W_i · [h_(t-1), x_t] + b_i)
C_tilde_t = tanh(W_c · [h_(t-1), x_t] + b_c)</code></pre>
<ul>
<li><code>i_t</code>는 새 정보를 얼마나 반영할지 나타낸다.</li>
<li><code>C_tilde_t</code>는 현재 입력으로 만든 새로운 기억의 후보값이다.</li>
</ul>
<p>이전 기억과 새 기억을 합쳐 현재 <code>cell state</code>를 만든다.</p>
<pre><code class="language-text">C_t = f_t ⊙ C_(t-1) + i_t ⊙ C_tilde_t</code></pre>
<p>여기서 <code>⊙</code>는 같은 위치의 원소끼리 곱하는 연산을 의미한다.</p>
<h3 id="output-gate">Output Gate</h3>
<p>Output Gate는 갱신된 <code>cell state</code> 중 현재 <code>hidden state</code>로 전달할 정보를 선택한다.</p>
<pre><code class="language-text">o_t = sigmoid(W_o · [h_(t-1), x_t] + b_o)
h_t = o_t ⊙ tanh(C_t)</code></pre>
<p>전체 흐름은 다음과 같다.</p>
<pre><code class="language-text">이전 기억 C_(t-1)
        │
        ├── Forget Gate ── 과거 기억 중 남길 부분
        │
현재 입력 x_t
        ├── Input Gate  ── 새로 저장할 부분
        │
        ▼
새 Cell State C_t
        │
        └── Output Gate ── 현재 Hidden State h_t</code></pre>
<h3 id="장기-기억과-단기-기억의-기준">장기 기억과 단기 기억의 기준</h3>
<p>LSTM 내부에 <code>6시간이 지나면 장기 기억</code>과 같은 고정 규칙은 없다. 정보가 얼마나 오래되었는지보다 <strong>현재 예측에 얼마나 도움이 되는지</strong>가 중요하다.</p>
<p>예를 들어 전력 사용량이 매일 저녁 비슷한 시간에 증가한다면 오래전 시점에서 만들어진 정보도 유용할 수 있다. 반대로 1시간 전 정보라도 잡음에 가깝다면 게이트가 영향력을 줄일 수 있다.</p>
<hr>
<h2 id="3-가정-전력-사용량-예측">3. 가정 전력 사용량 예측</h2>
<p>LSTM을 이용해 한 가정의 다음 1시간 전력 사용량을 예측한다.</p>
<ul>
<li>원본 데이터는 약 4년 동안 1분 간격으로 측정한 전력 사용량이다.</li>
<li><code>Global_active_power</code> 한 개의 Feature를 사용한다.</li>
<li>1분 단위 데이터를 1시간 단위 평균으로 변환한다.</li>
<li>직전 24시간을 입력으로 사용해 다음 1시간을 예측한다.</li>
</ul>
<pre><code class="language-text">입력 X: t-24, t-23, ..., t-1 시점의 전력 사용량
정답 y: t 시점의 전력 사용량</code></pre>
<h3 id="데이터-불러오기">데이터 불러오기</h3>
<pre><code class="language-python">import pandas as pd

DATA_PATH = &quot;./data/household_power_consumption.txt&quot;

df = pd.read_csv(
    DATA_PATH,
    sep=&quot;;&quot;,
    na_values=&quot;?&quot;,
    low_memory=False,
)

print(df.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">(2075259, 9)</code></pre>
<p>주요 컬럼은 다음과 같다.</p>
<table>
<thead>
<tr>
<th>컬럼</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td><code>Date</code>, <code>Time</code></td>
<td>측정 날짜와 시간이다.</td>
</tr>
<tr>
<td><code>Global_active_power</code></td>
<td>전체 유효전력(kW)이다.</td>
</tr>
<tr>
<td><code>Global_reactive_power</code></td>
<td>무효전력이다.</td>
</tr>
<tr>
<td><code>Voltage</code></td>
<td>평균 전압(V)이다.</td>
</tr>
<tr>
<td><code>Global_intensity</code></td>
<td>평균 전류 세기(A)다.</td>
</tr>
<tr>
<td><code>Sub_metering_1</code></td>
<td>주방 일부 기기의 사용량이다.</td>
</tr>
<tr>
<td><code>Sub_metering_2</code></td>
<td>세탁실 관련 기기의 사용량이다.</td>
</tr>
<tr>
<td><code>Sub_metering_3</code></td>
<td>전기 온수기와 에어컨 관련 사용량이다.</td>
</tr>
</tbody></table>
<h3 id="날짜와-시간-전처리">날짜와 시간 전처리</h3>
<p><code>Date</code>와 <code>Time</code>을 하나의 <code>datetime</code>으로 합친 뒤 시간 순서로 정렬한다.</p>
<pre><code class="language-python">df[&quot;datetime&quot;] = pd.to_datetime(
    df[&quot;Date&quot;] + &quot; &quot; + df[&quot;Time&quot;],
    format=&quot;%d/%m/%Y %H:%M:%S&quot;,
    errors=&quot;coerce&quot;,
)

df = df.dropna(subset=[&quot;datetime&quot;])
df = df.sort_values(&quot;datetime&quot;)
df = df.set_index(&quot;datetime&quot;)</code></pre>
<p>예측에 사용할 <code>Global_active_power</code>만 남기고 숫자로 변환한다.</p>
<pre><code class="language-python">power = df[[&quot;Global_active_power&quot;]].copy()
power[&quot;Global_active_power&quot;] = pd.to_numeric(
    power[&quot;Global_active_power&quot;],
    errors=&quot;coerce&quot;,
)
power = power.dropna()

print(power.shape)</code></pre>
<pre><code class="language-text">(2049280, 1)</code></pre>
<h3 id="1시간-단위로-변환하기">1시간 단위로 변환하기</h3>
<pre><code class="language-python">hourly = power.resample(&quot;1h&quot;).mean()

print(
    &quot;시간 단위 결측 개수:&quot;,
    hourly[&quot;Global_active_power&quot;].isna().sum(),
)

hourly = hourly.ffill().dropna()
print(&quot;hourly shape:&quot;, hourly.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">시간 단위 결측 개수: 421
hourly shape: (34589, 1)</code></pre>
<p><code>resample(&quot;1h&quot;).mean()</code>은 같은 시간에 속한 1분 단위 관측값의 평균을 계산한다. 시간 단위로 비어 있는 구간은 <code>ffill()</code>을 사용해 직전 관측값으로 채운다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/cdcd5a36-03b1-4a08-91e2-d59c5a2b5e6e/image.png" alt=""></p>
<hr>
<h2 id="4-전력-데이터를-시퀀스로-만들기">4. 전력 데이터를 시퀀스로 만들기</h2>
<p>데이터는 시간 순서를 유지한 채 Train 70%, Validation 15%, Test 15%로 나눈다. <code>StandardScaler</code>는 미래 정보가 학습 데이터에 섞이지 않도록 Train 구간에만 <code>fit()</code>한다.</p>
<pre><code class="language-python">import numpy as np
from sklearn.preprocessing import StandardScaler

values = hourly[&quot;Global_active_power&quot;].values.reshape(-1, 1)

n = len(values)
train_end = int(n * 0.70)
val_end = int(n * 0.85)

scaler = StandardScaler()
scaler.fit(values[:train_end])
scaled_values = scaler.transform(values).astype(np.float32)</code></pre>
<h3 id="최근-24시간과-다음-1시간-묶기">최근 24시간과 다음 1시간 묶기</h3>
<pre><code class="language-python">SEQ_LENGTH = 24

def make_sequence(data, seq_length):
    X = []
    y = []
    target_indices = []

    for i in range(seq_length, len(data)):
        X.append(data[i - seq_length:i])
        y.append(data[i])
        target_indices.append(i)

    return (
        np.array(X, dtype=np.float32),
        np.array(y, dtype=np.float32),
        np.array(target_indices),
    )

X, y, indices = make_sequence(scaled_values, SEQ_LENGTH)

print(X.shape)
print(y.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">(34565, 24, 1)
(34565, 1)</code></pre>
<p>각 차원의 의미는 다음과 같다.</p>
<pre><code class="language-text">X.shape = (샘플 수, 시퀀스 길이, Feature 수)
        = (34565, 24, 1)</code></pre>
<p>첫 번째 샘플은 <code>2006-12-16 17:00</code>부터 <code>2006-12-17 16:00</code>까지의 24시간을 입력으로 사용하고, <code>2006-12-17 17:00</code>의 값을 정답으로 사용한다.</p>
<p>시퀀스의 정답 시점을 기준으로 데이터를 나눈 결과는 다음과 같다.</p>
<table>
<thead>
<tr>
<th>구간</th>
<th><code>X</code> Shape</th>
<th><code>y</code> Shape</th>
</tr>
</thead>
<tbody><tr>
<td>Train</td>
<td><code>(24188, 24, 1)</code></td>
<td><code>(24188, 1)</code></td>
</tr>
<tr>
<td>Validation</td>
<td><code>(5188, 24, 1)</code></td>
<td><code>(5188, 1)</code></td>
</tr>
<tr>
<td>Test</td>
<td><code>(5189, 24, 1)</code></td>
<td><code>(5189, 1)</code></td>
</tr>
</tbody></table>
<hr>
<h2 id="5-pytorch-lstm-모델">5. PyTorch LSTM 모델</h2>
<pre><code class="language-python">import torch
import torch.nn as nn

class PowerLSTM(nn.Module):
    def __init__(
        self,
        input_size=1,
        hidden_size=64,
        num_layers=2,
        dropout=0.2,
    ):
        super().__init__()

        self.lstm = nn.LSTM(
            input_size=input_size,
            hidden_size=hidden_size,
            num_layers=num_layers,
            batch_first=True,
            dropout=dropout if num_layers &gt; 1 else 0.0,
        )
        self.fc = nn.Linear(hidden_size, 1)

    def forward(self, x):
        out, (h_n, c_n) = self.lstm(x)
        last_hidden = out[:, -1, :]
        prediction = self.fc(last_hidden)
        return prediction</code></pre>
<p>모델의 처리 과정은 다음과 같다.</p>
<pre><code class="language-text">[batch, 24, 1]
        │
        ▼
2층 LSTM, hidden_size=64
        │
        ▼
마지막 시점의 hidden state: [batch, 64]
        │
        ▼
Linear(64, 1)
        │
        ▼
다음 1시간의 전력 사용량: [batch, 1]</code></pre>
<h3 id="lstm의-출력-shape">LSTM의 출력 Shape</h3>
<p><code>batch_first=True</code>일 때 입력은 <code>[batch, sequence, feature]</code> 순서다.</p>
<pre><code class="language-python">out, (h_n, c_n) = model.lstm(sample_x)

print(&quot;입력:&quot;, sample_x.shape)
print(&quot;out:&quot;, out.shape)
print(&quot;h_n:&quot;, h_n.shape)
print(&quot;c_n:&quot;, c_n.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">입력: torch.Size([4, 24, 1])
out:  torch.Size([4, 24, 64])
h_n:  torch.Size([2, 4, 64])
c_n:  torch.Size([2, 4, 64])</code></pre>
<table>
<thead>
<tr>
<th>출력</th>
<th>Shape</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td><code>out</code></td>
<td><code>[batch, seq_length, hidden_size]</code></td>
<td>마지막 LSTM 층이 각 시점에서 만든 hidden state다.</td>
</tr>
<tr>
<td><code>h_n</code></td>
<td><code>[num_layers, batch, hidden_size]</code></td>
<td>각 층의 마지막 hidden state다.</td>
</tr>
<tr>
<td><code>c_n</code></td>
<td><code>[num_layers, batch, hidden_size]</code></td>
<td>각 층의 마지막 cell state다.</td>
</tr>
</tbody></table>
<p>단방향 LSTM의 마지막 층에서는 <code>out[:, -1, :]</code>와 <code>h_n[-1]</code>이 같은 마지막 시점의 hidden state를 나타낸다.</p>
<hr>
<h2 id="6-학습과-평가-결과">6. 학습과 평가 결과</h2>
<p>학습에는 평균 제곱 오차와 AdamW Optimizer를 사용한다.</p>
<pre><code class="language-python">criterion = nn.MSELoss()
optimizer = torch.optim.AdamW(
    model.parameters(),
    lr=1e-3,
    weight_decay=1e-4,
)

MAX_EPOCHS = 50
PATIENCE = 8
CLIP_NORM = 1.0</code></pre>
<ul>
<li>Validation Loss가 좋아진 모델의 상태를 저장한다.</li>
<li>8번 연속 좋아지지 않으면 Early Stopping을 적용한다.</li>
<li><code>clip_grad_norm_()</code>으로 기울기 폭주를 완화한다.</li>
</ul>
<p>노트북에서는 29 Epoch에 Early Stopping이 발생했고, 가장 좋은 가중치를 다시 불러와 Test 데이터를 평가했다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/d72be1a4-6678-4075-b3d6-2013fc79e2d6/image.png" alt=""></p>
<h3 id="원래-단위로-복원하기">원래 단위로 복원하기</h3>
<p>모델의 출력은 표준화된 값이므로 <code>inverse_transform()</code>으로 원래 전력 단위로 복원한다.</p>
<pre><code class="language-python">pred_scaled, true_scaled = predict(model, test_loader, DEVICE)

pred = scaler.inverse_transform(pred_scaled).ravel()
true = scaler.inverse_transform(true_scaled).ravel()</code></pre>
<h3 id="persistence-baseline과-비교하기">Persistence Baseline과 비교하기</h3>
<p>Persistence Baseline은 복잡한 모델 없이 <strong>마지막 입력값이 다음 시점에도 그대로 이어진다</strong>고 예측한다.</p>
<pre><code class="language-python">last_input_scaled = X_test[:, -1, :]
baseline_pred = scaler.inverse_transform(last_input_scaled).ravel()</code></pre>
<p>노트북의 실행 결과는 다음과 같다.</p>
<table>
<thead>
<tr>
<th>모델</th>
<th align="right">RMSE</th>
<th align="right">MAE</th>
</tr>
</thead>
<tbody><tr>
<td>Persistence Baseline</td>
<td align="right">0.5759</td>
<td align="right">0.3729</td>
</tr>
<tr>
<td>LSTM</td>
<td align="right"><strong>0.4915</strong></td>
<td align="right"><strong>0.3423</strong></td>
</tr>
</tbody></table>
<p>RMSE와 MAE는 낮을수록 좋다. 이 실행에서는 LSTM이 단순 Baseline보다 두 지표 모두 낮았다. 특히 RMSE는 약 14.7%, MAE는 약 8.2% 감소했다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/b895af04-a3d7-42ab-98a8-24223be8b836/image.png" alt=""></p>
<p>그래프를 볼 때는 다음 내용을 함께 확인해야 한다.</p>
<ol>
<li>실제값의 상승과 하락 방향을 따라가는가?</li>
<li>급격한 Peak를 지나치게 평평하게 예측하는가?</li>
<li>예측이 실제값보다 한 시점 늦게 따라가는가?</li>
</ol>
<p>위 결과에서는 전반적인 주기와 방향은 따라가지만, 일부 급격한 Peak를 실제보다 낮게 예측한다. 따라서 시계열 모델은 하나의 평가지표뿐 아니라 Baseline과 시각화 결과를 함께 확인해야 한다.</p>
<hr>
<h2 id="7-gru">7. GRU</h2>
<p>GRU(Gated Recurrent Unit)는 LSTM과 마찬가지로 장기 의존성 문제를 완화하기 위한 순환 신경망이다. LSTM보다 구조가 단순하며 별도의 <code>cell state</code>를 사용하지 않는다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/03b29547-c611-4d3b-8b06-06ead61193c3/image.png" alt=""></p>
<p>GRU는 두 게이트를 사용한다.</p>
<table>
<thead>
<tr>
<th>게이트</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td>Update Gate</td>
<td>이전 정보를 얼마나 유지하고 새 정보로 얼마나 갱신할지 결정한다.</td>
</tr>
<tr>
<td>Reset Gate</td>
<td>새 후보 상태를 만들 때 과거 정보를 얼마나 참고할지 결정한다.</td>
</tr>
</tbody></table>
<h3 id="lstm과-gru-비교">LSTM과 GRU 비교</h3>
<table>
<thead>
<tr>
<th>항목</th>
<th>LSTM</th>
<th>GRU</th>
</tr>
</thead>
<tbody><tr>
<td>상태</td>
<td><code>hidden state</code>, <code>cell state</code></td>
<td><code>hidden state</code></td>
</tr>
<tr>
<td>게이트</td>
<td>Input, Forget, Output</td>
<td>Update, Reset</td>
</tr>
<tr>
<td>구조</td>
<td>비교적 복잡하다.</td>
<td>비교적 단순하다.</td>
</tr>
<tr>
<td>파라미터와 연산량</td>
<td>상대적으로 많다.</td>
<td>상대적으로 적다.</td>
</tr>
<tr>
<td>특징</td>
<td>긴 의존 관계를 세밀하게 제어한다.</td>
<td>학습이 빠르고 적은 데이터에서도 효율적일 수 있다.</td>
</tr>
</tbody></table>
<p>어느 모델이 항상 더 좋다고 정할 수는 없다. 데이터 길이와 크기, 필요한 문맥 범위에 따라 Validation 성능과 학습 시간을 비교해 선택해야 한다.</p>
<h3 id="pytorch에서-gru-사용하기">PyTorch에서 GRU 사용하기</h3>
<p>LSTM과 생성 방식은 거의 같다.</p>
<pre><code class="language-python">gru = nn.GRU(
    input_size=1,
    hidden_size=64,
    num_layers=2,
    batch_first=True,
    dropout=0.2,
)

out, h_n = gru(x)</code></pre>
<p>LSTM은 다음처럼 두 상태를 반환한다.</p>
<pre><code class="language-python">out, (h_n, c_n) = lstm(x)</code></pre>
<p>GRU는 별도의 <code>cell state</code>가 없으므로 <code>out</code>과 <code>h_n</code>만 반환한다.</p>
<hr>
<h2 id="8-seq2seq">8. Seq2Seq</h2>
<p>Seq2Seq(Sequence-to-Sequence)는 입력 시퀀스를 받아 다른 출력 시퀀스를 생성하는 구조다. 기계 번역, 문서 요약, 질의응답 등에 사용할 수 있다.</p>
<pre><code class="language-text">I love coffee
      │
      ▼
   Encoder
      │
문장 정보 압축
      │
      ▼
   Decoder
      │
      ▼
&lt;BOS&gt; 나는 → 커피를 → 좋아해 &lt;EOS&gt;</code></pre>
<ul>
<li><code>&lt;BOS&gt;</code>는 문장 생성의 시작을 알리는 특별 Token이다.</li>
<li><code>&lt;EOS&gt;</code>는 문장의 끝을 알리는 특별 Token이다.</li>
<li>Decoder는 이전에 생성한 Token을 이용해 다음 Token을 순서대로 예측한다.</li>
<li>이처럼 이전 출력을 다음 입력으로 사용하는 방식을 <strong>자기회귀(Autoregressive)</strong>라고 한다.</li>
</ul>
<p><img src="https://velog.velcdn.com/images/j_keun/post/a764763c-cf82-4d03-934d-5e1063d1d47f/image.png" alt=""></p>
<h3 id="encoder와-decoder">Encoder와 Decoder</h3>
<table>
<thead>
<tr>
<th>구성 요소</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td>Encoder</td>
<td>입력 Token을 순서대로 읽으며 문장 정보를 내부 상태에 누적한다.</td>
</tr>
<tr>
<td>Decoder</td>
<td>Encoder가 만든 정보를 바탕으로 출력 Token을 한 개씩 생성한다.</td>
</tr>
</tbody></table>
<p>고전적인 Seq2Seq의 Encoder와 Decoder에는 RNN, LSTM, GRU 등을 사용할 수 있다.</p>
<h3 id="seq2seq의-문제점">Seq2Seq의 문제점</h3>
<p>초기 Seq2Seq는 Encoder의 마지막 <code>hidden state</code> 하나를 Decoder에 전달하는 방식을 주로 사용했다.</p>
<p>문장이 길어질수록 앞부분부터 뒷부분까지의 모든 정보를 하나의 고정된 크기 벡터에 압축해야 한다. 이 과정에서 중요한 정보가 손실되는 현상을 <strong>정보 병목(Bottleneck)</strong>이라고 한다.</p>
<p>또한 RNN 계열은 이전 시점의 계산이 끝나야 다음 시점을 계산할 수 있다.</p>
<pre><code class="language-text">t=1 계산 → t=2 계산 → t=3 계산 → ...</code></pre>
<p>따라서 긴 시퀀스에서는 학습이 느리고 병렬 처리에도 불리하다.</p>
<hr>
<h2 id="9-attention-메커니즘">9. Attention 메커니즘</h2>
<p>Attention은 현재 처리하는 Token이 다른 Token을 얼마나 참고해야 하는지 점수로 계산하는 방법이다.</p>
<pre><code class="language-text">&quot;나는 오늘 학교에서 수학 시험을 봤다&quot;</code></pre>
<p>어떤 Token이 중요한지는 항상 같지 않다. 현재 어떤 Token을 처리하는지에 따라 참고해야 할 대상이 달라진다. Attention은 이러한 관계를 학습된 점수로 표현한다.</p>
<h3 id="token을-embedding-vector로-바꾸기">Token을 Embedding Vector로 바꾸기</h3>
<p>신경망은 문자열을 그대로 계산할 수 없으므로 Token을 ID로 바꾸고, <code>nn.Embedding</code>을 이용해 학습 가능한 실수 벡터로 변환한다.</p>
<pre><code class="language-python">import torch
import torch.nn as nn

torch.manual_seed(2026)

tokens = [&quot;커피&quot;, &quot;한잔&quot;, &quot;어때&quot;]
token_ids = torch.tensor([0, 1, 2])

embedding = nn.Embedding(
    num_embeddings=3,
    embedding_dim=4,
)

X = embedding(token_ids)

print(&quot;Token ID Shape:&quot;, token_ids.shape)
print(&quot;Embedding Shape:&quot;, X.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">Token ID Shape: torch.Size([3])
Embedding Shape: torch.Size([3, 4])</code></pre>
<p>Token ID <code>0</code>, <code>1</code>, <code>2</code>는 단어의 의미나 크기를 나타내는 숫자가 아니다. Embedding Table에서 각 Token의 벡터를 찾기 위한 행 번호다.</p>
<hr>
<h2 id="10-self-attention과-q-k-v">10. Self-Attention과 Q, K, V</h2>
<p>Self-Attention은 같은 시퀀스 안의 모든 Token이 서로의 관계를 계산하는 Attention이다.</p>
<pre><code class="language-text">커피 → 커피, 한잔, 어때를 참고
한잔 → 커피, 한잔, 어때를 참고
어때 → 커피, 한잔, 어때를 참고</code></pre>
<p>Token이 3개라면 관계 점수표의 Shape은 <code>3 × 3</code>이 된다.</p>
<h3 id="query-key-value">Query, Key, Value</h3>
<p>Embedding <code>X</code>를 서로 다른 세 개의 선형층에 통과시켜 Query, Key, Value를 만든다.</p>
<table>
<thead>
<tr>
<th>벡터</th>
<th>직관적인 의미</th>
<th>계산</th>
</tr>
</thead>
<tbody><tr>
<td>Query(Q)</td>
<td>현재 Token이 어떤 정보를 찾는가?</td>
<td><code>XW_Q</code></td>
</tr>
<tr>
<td>Key(K)</td>
<td>각 Token이 어떤 특징을 가졌는가?</td>
<td><code>XW_K</code></td>
</tr>
<tr>
<td>Value(V)</td>
<td>선택되었을 때 실제로 전달할 정보는 무엇인가?</td>
<td><code>XW_V</code></td>
</tr>
</tbody></table>
<pre><code class="language-python">embed_dim = 4

W_Q = nn.Linear(embed_dim, embed_dim, bias=False)
W_K = nn.Linear(embed_dim, embed_dim, bias=False)
W_V = nn.Linear(embed_dim, embed_dim, bias=False)

Q = W_Q(X)
K = W_K(X)
V = W_V(X)

print(&quot;X:&quot;, X.shape)
print(&quot;Q:&quot;, Q.shape)
print(&quot;K:&quot;, K.shape)
print(&quot;V:&quot;, V.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">X: torch.Size([3, 4])
Q: torch.Size([3, 4])
K: torch.Size([3, 4])
V: torch.Size([3, 4])</code></pre>
<h3 id="q와-k로-관계-점수-계산하기">Q와 K로 관계 점수 계산하기</h3>
<p>현재 Token의 Query와 모든 Token의 Key를 내적해 관련도 점수를 계산한다.</p>
<pre><code class="language-python">scores = Q @ K.T
print(scores.shape)</code></pre>
<p>Shape의 변화는 다음과 같다.</p>
<pre><code class="language-text">Q:   (3, 4)
K.T: (4, 3)
----------------
결과: (3, 3)</code></pre>
<ul>
<li>행은 정보를 찾는 Query Token이다.</li>
<li>열은 참고 대상인 Key Token이다.</li>
<li><code>scores[0, 2]</code>는 <code>커피</code>의 Query와 <code>어때</code>의 Key를 비교한 점수다.</li>
</ul>
<h3 id="scaled-dot-product-attention">Scaled Dot-Product Attention</h3>
<p>Transformer에서는 점수가 지나치게 커지는 것을 줄이기 위해 Key 벡터의 차원 <code>d_k</code>의 제곱근으로 나눈다. 그다음 <code>softmax</code>로 각 행의 합이 <code>1</code>인 Attention Weight를 만든다.</p>
<pre><code class="language-python">import math

d_k = Q.size(-1)

scores = (Q @ K.T) / math.sqrt(d_k)
attention_weights = torch.softmax(scores, dim=-1)
attention_output = attention_weights @ V

print(&quot;Score Shape:&quot;, scores.shape)
print(&quot;Weight Shape:&quot;, attention_weights.shape)
print(&quot;Output Shape:&quot;, attention_output.shape)</code></pre>
<p>수식으로 표현하면 다음과 같다.</p>
<pre><code class="language-text">Attention(Q, K, V)
    = softmax(QK^T / √d_k)V</code></pre>
<p>처리 과정은 다음과 같이 정리할 수 있다.</p>
<pre><code class="language-text">Token IDs
    │
    ▼
Embedding X
    │
    ├── Linear ── Q
    ├── Linear ── K
    └── Linear ── V

Q × K.T
    │
    ▼
관련도 점수
    │
√d_k로 나누기
    │
Softmax
    │
    ▼
Attention Weight
    │
Weight × V
    │
    ▼
문맥을 반영한 출력</code></pre>
<hr>
<h2 id="11-attention과-transformer의-연결">11. Attention과 Transformer의 연결</h2>
<p>Attention을 사용하면 각 Token이 다른 Token을 직접 참고할 수 있다. RNN처럼 정보를 시점별로 차례로 전달할 필요가 없으므로 학습할 때 여러 Token의 관계 계산을 병렬화하기 쉽다.</p>
<p>다만 Self-Attention만으로는 Token의 순서를 자동으로 알 수 없다. 실제 Transformer는 Token Embedding에 위치 정보를 더하고, Multi-Head Attention과 Feed Forward Network 등을 함께 사용한다.</p>
<pre><code class="language-text">RNN 계열
앞 시점의 상태를 다음 시점으로 순차 전달

Transformer
각 Token이 다른 Token과의 관계를 직접 계산</code></pre>
<p>이번 예제는 Transformer 전체 구조를 한 번에 구현하기보다, 그 핵심 연산인 <strong>Embedding → Q·K·V → 관계 점수</strong>가 어떻게 만들어지는지를 확인하는 단계다.</p>
<hr>
<h2 id="12-모델별-핵심-비교">12. 모델별 핵심 비교</h2>
<table>
<thead>
<tr>
<th>모델</th>
<th>정보 전달 방식</th>
<th>장점</th>
<th>주의할 점</th>
</tr>
</thead>
<tbody><tr>
<td>RNN</td>
<td>하나의 hidden state를 순차적으로 전달한다.</td>
<td>구조가 단순하다.</td>
<td>긴 시퀀스에서 장기 의존성 학습이 어렵다.</td>
</tr>
<tr>
<td>LSTM</td>
<td>cell state와 세 게이트로 기억을 조절한다.</td>
<td>오래된 정보를 비교적 안정적으로 유지한다.</td>
<td>구조가 복잡하고 연산량이 많다.</td>
</tr>
<tr>
<td>GRU</td>
<td>hidden state와 두 게이트로 기억을 조절한다.</td>
<td>LSTM보다 단순하고 빠를 수 있다.</td>
<td>데이터에 따라 LSTM보다 성능이 낮을 수 있다.</td>
</tr>
<tr>
<td>Seq2Seq</td>
<td>Encoder가 입력을 읽고 Decoder가 출력을 생성한다.</td>
<td>서로 다른 길이의 시퀀스를 변환할 수 있다.</td>
<td>하나의 Context Vector에 의존하면 병목이 생긴다.</td>
</tr>
<tr>
<td>Self-Attention</td>
<td>모든 Token 사이의 관계를 직접 계산한다.</td>
<td>긴 거리의 관계를 직접 반영하고 병렬화하기 쉽다.</td>
<td>시퀀스가 길수록 관계 점수 행렬의 크기가 커진다.</td>
</tr>
</tbody></table>
<hr>
<h2 id="13-핵심-정리">13. 핵심 정리</h2>
<ol>
<li>LSTM은 <code>hidden state</code>와 <code>cell state</code>를 사용한다.</li>
<li>Forget, Input, Output Gate가 과거 기억의 유지와 새 정보의 저장을 조절한다.</li>
<li>전력 예측 예제는 최근 24시간을 이용해 다음 1시간을 예측한다.</li>
<li>LSTM의 출력은 <code>out</code>, <code>h_n</code>, <code>c_n</code>이며 각 Shape과 역할이 다르다.</li>
<li>모델 성능은 RMSE와 MAE뿐 아니라 단순 Baseline, 예측 그래프와 함께 평가해야 한다.</li>
<li>GRU는 별도의 <code>cell state</code> 없이 Update Gate와 Reset Gate를 사용한다.</li>
<li>Seq2Seq는 Encoder와 Decoder로 구성되며, 긴 문장을 하나의 벡터로 압축할 때 정보 병목이 생길 수 있다.</li>
<li>Self-Attention은 Query와 Key로 관계 점수를 만들고, Value를 가중합해 문맥이 반영된 출력을 만든다.</li>
<li>Transformer의 핵심 Attention 연산은 <code>softmax(QK^T / √d_k)V</code>로 표현할 수 있다.</li>
</ol>
]]></description>
        </item>
        <item>
            <title><![CDATA[[69일차]  텍스트 벡터화와 RNN 시계열 예측]]></title>
            <link>https://velog.io/@j_keun/69%EC%9D%BC%EC%B0%A8-%ED%85%8D%EC%8A%A4%ED%8A%B8-%EB%B2%A1%ED%84%B0%ED%99%94%EC%99%80-RNN-%EC%8B%9C%EA%B3%84%EC%97%B4-%EC%98%88%EC%B8%A1</link>
            <guid>https://velog.io/@j_keun/69%EC%9D%BC%EC%B0%A8-%ED%85%8D%EC%8A%A4%ED%8A%B8-%EB%B2%A1%ED%84%B0%ED%99%94%EC%99%80-RNN-%EC%8B%9C%EA%B3%84%EC%97%B4-%EC%98%88%EC%B8%A1</guid>
            <pubDate>Tue, 29 Sep 2026 16:29:16 GMT</pubDate>
            <description><![CDATA[<p><strong>텍스트 벡터화와 워드 임베딩</strong>을 살펴보고, 순서가 중요한 데이터를 처리하는 <strong>RNN(Recurrent Neural Network)</strong>으로 다음 거래일의 KOSPI 종가를 예측한다.</p>
<hr>
<h2 id="1-텍스트-벡터화">1. 텍스트 벡터화</h2>
<p>벡터화(Vectorization)는 텍스트를 머신러닝과 딥러닝 모델이 계산할 수 있는 숫자 벡터로 바꾸는 과정이다.</p>
<p>텍스트를 숫자로 표현하는 대표적인 방법은 다음과 같다.</p>
<table>
<thead>
<tr>
<th>방법</th>
<th>핵심 기준</th>
<th>특징</th>
</tr>
</thead>
<tbody><tr>
<td>원-핫 인코딩</td>
<td>특정 Token의 위치</td>
<td>단순하지만 희소하고 의미 관계를 표현하지 못한다.</td>
</tr>
<tr>
<td>Bag of Words</td>
<td>문서 안의 단어 빈도</td>
<td>구현이 쉽지만 단어 순서를 잃는다.</td>
</tr>
<tr>
<td>TF-IDF</td>
<td>문서별 단어 중요도</td>
<td>여러 문서에 흔한 단어의 가중치를 낮춘다.</td>
</tr>
<tr>
<td>Word Embedding</td>
<td>학습된 밀집 벡터</td>
<td>단어 사이의 의미적 관계를 벡터 공간에 표현한다.</td>
</tr>
</tbody></table>
<h3 id="원-핫-인코딩">원-핫 인코딩</h3>
<p>원-핫 인코딩(One-Hot Encoding)은 Vocabulary 크기만큼의 벡터를 만들고, 표현하려는 Token 위치만 <code>1</code>, 나머지는 모두 <code>0</code>으로 채운다.</p>
<pre><code class="language-python">import torch
import torch.nn.functional as F

token_id = torch.tensor(2)
vocab_size = 5

one_hot = F.one_hot(token_id, num_classes=vocab_size)

print(one_hot)
print(one_hot.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">tensor([0, 0, 1, 0, 0])
torch.Size([5])</code></pre>
<p>Vocabulary 크기가 10,000이면 Token 하나를 나타내기 위해 길이가 10,000인 벡터가 필요하다. 그중 대부분이 <code>0</code>이므로 이러한 벡터를 <strong>희소 벡터(Sparse Vector)</strong>라고 한다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/97abcc97-d965-4b49-ab1a-09111ccc7f05/image.png" alt=""></p>
<h3 id="bag-of-words">Bag of Words</h3>
<p>Bag of Words(BOW)는 문서에 각 단어가 몇 번 등장했는지를 벡터로 표현한다.</p>
<pre><code class="language-text">Vocabulary: [&quot;I&quot;, &quot;love&quot;, &quot;Python&quot;]
문장: &quot;I love Python Python&quot;
BOW: [1, 1, 2]</code></pre>
<p>단어 빈도를 간단하게 표현할 수 있지만 다음 두 문장을 구별하기 어렵다.</p>
<pre><code class="language-text">나는 너를 좋아한다
너는 나를 좋아한다</code></pre>
<p>두 문장에 등장하는 단어가 같다면 BOW 벡터도 비슷해진다. 즉, 단어의 <strong>순서와 문맥을 반영하지 못한다.</strong></p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/5fa4698c-893a-4e8a-8470-0708bfef7f89/image.png" alt=""></p>
<h3 id="tf-idf">TF-IDF</h3>
<p>TF-IDF(Term Frequency-Inverse Document Frequency)는 단어가 한 문서에서 자주 나타날수록 가중치를 높이고, 전체 문서에서 흔하게 나타날수록 가중치를 낮춘다.</p>
<pre><code class="language-text">TF-IDF = TF × IDF</code></pre>
<ul>
<li>TF: 특정 문서에서 해당 단어가 등장한 빈도다.</li>
<li>IDF: 전체 문서에서 드물게 등장할수록 커지는 값이다.</li>
</ul>
<p>모든 문서에 자주 나타나는 단어보다 특정 문서의 내용을 잘 나타내는 단어에 높은 가중치를 줄 수 있다. 다만 BOW와 마찬가지로 단어 순서와 문맥을 직접 학습하지는 않는다.</p>
<hr>
<h2 id="2-word-embedding">2. Word Embedding</h2>
<p>Word Embedding은 Token을 저차원의 <strong>밀집 벡터(Dense Vector)</strong>로 표현하는 방법이다.</p>
<pre><code class="language-text">Token ID
   ↓
Embedding Layer
   ↓
[0.21, -0.72, 0.13, 0.84, ...]</code></pre>
<p>원-핫 벡터는 각 단어가 서로 독립적이지만, Embedding은 학습을 통해 비슷한 문맥에서 사용되는 단어가 가까운 벡터를 갖도록 만들 수 있다.</p>
<h3 id="nnembedding"><code>nn.Embedding</code></h3>
<p>PyTorch의 <code>nn.Embedding</code>은 Token ID에 대응하는 벡터를 Embedding Table에서 가져오는 층이다.</p>
<pre><code class="language-python">import torch.nn as nn

embedding = nn.Embedding(
    num_embeddings=10000,
    embedding_dim=128,
    padding_idx=0,
)

print(embedding)
print(embedding.weight.shape)</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">Embedding(10000, 128, padding_idx=0)
torch.Size([10000, 128])</code></pre>
<table>
<thead>
<tr>
<th>인자</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td><code>num_embeddings</code></td>
<td>Vocabulary에 포함된 Token 개수다.</td>
</tr>
<tr>
<td><code>embedding_dim</code></td>
<td>Token 하나를 나타낼 벡터 차원이다.</td>
</tr>
<tr>
<td><code>padding_idx</code></td>
<td>Padding Token의 ID다. 해당 행은 학습에서 갱신하지 않는다.</td>
</tr>
</tbody></table>
<p>입력과 출력 Shape는 다음처럼 변한다.</p>
<pre><code class="language-python">input_ids = torch.tensor([
    [1, 2, 3],
    [4, 2, 5],
], dtype=torch.long)

embedding = nn.Embedding(
    num_embeddings=10,
    embedding_dim=4,
    padding_idx=0,
)

output = embedding(input_ids)

print(&quot;입력 Shape:&quot;, input_ids.shape)
print(&quot;출력 Shape:&quot;, output.shape)
print(&quot;학습 여부:&quot;, embedding.weight.requires_grad)
print(&quot;Parameter 수:&quot;, embedding.weight.numel())</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">입력 Shape: torch.Size([2, 3])
출력 Shape: torch.Size([2, 3, 4])
학습 여부: True
Parameter 수: 40</code></pre>
<p>입력 Shape <code>[batch_size, sequence_length]</code>의 뒤에 <code>embedding_dim</code>이 추가된다.</p>
<pre><code class="language-text">[2, 3] → [2, 3, 4]</code></pre>
<p>Embedding 벡터의 의미를 사람이 직접 지정하는 것은 아니다. 처음에는 초기값으로 시작하고, 모델의 다른 가중치와 함께 학습된다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/34f78fd4-f721-46bb-9ee5-a4eb13658b4b/image.png" alt=""></p>
<hr>
<h2 id="3-word2vec">3. Word2Vec</h2>
<p>Word2Vec은 주변 문맥을 이용해 단어의 벡터를 학습한다. 비슷한 문맥에서 자주 사용되는 단어가 벡터 공간에서도 가까워지도록 만든다.</p>
<h3 id="cbow와-skip-gram">CBOW와 Skip-gram</h3>
<table>
<thead>
<tr>
<th>방식</th>
<th>입력</th>
<th>예측 대상</th>
<th>특징</th>
</tr>
</thead>
<tbody><tr>
<td>CBOW</td>
<td>주변 단어</td>
<td>중심 단어</td>
<td>비교적 빠르고 자주 등장하는 단어 학습에 유리하다.</td>
</tr>
<tr>
<td>Skip-gram</td>
<td>중심 단어</td>
<td>주변 단어</td>
<td>계산량은 많지만 드문 단어 학습에 유리하다.</td>
</tr>
</tbody></table>
<p>문장이 다음과 같다고 가정한다.</p>
<pre><code class="language-text">나는 오늘 공원에서 귀여운 강아지를 산책시키며 행복한 시간을 보냈다.</code></pre>
<p>중심 단어가 <code>강아지를</code>이고 Window 크기가 2라면 주변 문맥은 다음과 같다.</p>
<pre><code class="language-text">[&quot;공원에서&quot;, &quot;귀여운&quot;, &quot;산책시키며&quot;, &quot;행복한&quot;]</code></pre>
<ul>
<li>CBOW는 주변 네 단어로 <code>강아지를</code>을 예측한다.</li>
<li>Skip-gram은 <code>강아지를</code>로 주변 네 단어를 예측한다.</li>
</ul>
<p><img src="https://velog.velcdn.com/images/j_keun/post/045a09c5-edce-4fbf-90c6-eba8ca26352e/image.png" alt=""></p>
<h3 id="sgns">SGNS</h3>
<p>SGNS(Skip-Gram with Negative Sampling)는 Skip-gram의 계산량을 줄이는 방법이다.</p>
<ul>
<li>실제 중심 단어와 주변 단어의 조합은 관련이 있는 Positive Sample로 학습한다.</li>
<li>무작위로 선택한 단어 조합은 관련이 없는 Negative Sample로 학습한다.</li>
<li>Vocabulary 전체에 대한 확률을 매번 계산하지 않고 일부 단어만 비교한다.</li>
</ul>
<h3 id="nsmc-데이터-전처리">NSMC 데이터 전처리</h3>
<p>한국어 Word2Vec 학습을 위해 네이버 영화 리뷰 데이터(NSMC)를 사용한다.</p>
<p>노트북의 macOS 환경에서는 MeCab을 다음과 같이 연결했다.</p>
<pre><code class="language-python">import os
from konlpy.tag import Mecab

os.environ[&quot;MECABRC&quot;] = &quot;/opt/homebrew/etc/mecabrc&quot;

mecab = Mecab(
    dicpath=&quot;/opt/homebrew/lib/mecab/dic/mecab-ko-dic&quot;,
)</code></pre>
<p><code>MECABRC</code>와 사전 경로는 설치 방법과 운영체제에 따라 달라진다.</p>
<pre><code class="language-python">import urllib.request
import pandas as pd

urllib.request.urlretrieve(
    &quot;https://raw.githubusercontent.com/e9t/nsmc/master/ratings.txt&quot;,
    filename=&quot;ratings.txt&quot;,
)

train_data = pd.read_table(&quot;ratings.txt&quot;)
train_data = train_data[:20000]</code></pre>
<p>한글과 공백을 제외한 문자를 제거한다.</p>
<pre><code class="language-python">train_data[&quot;document&quot;] = train_data[&quot;document&quot;].str.replace(
    &quot;[^ㄱ-ㅎㅏ-ㅣ가-힣 ]&quot;,
    &quot;&quot;,
    regex=True,
)</code></pre>
<p>MeCab으로 형태소를 분리하고 불용어를 제거한다.</p>
<pre><code class="language-python">stopwords = [
    &quot;도&quot;, &quot;는&quot;, &quot;다&quot;, &quot;의&quot;, &quot;가&quot;, &quot;이&quot;, &quot;은&quot;, &quot;한&quot;, &quot;에&quot;, &quot;하&quot;,
    &quot;고&quot;, &quot;을&quot;, &quot;를&quot;, &quot;인&quot;, &quot;듯&quot;, &quot;과&quot;, &quot;와&quot;, &quot;네&quot;, &quot;들&quot;, &quot;지&quot;,
    &quot;임&quot;, &quot;게&quot;,
]

tokenized_data = []

for sentence in train_data[&quot;document&quot;]:
    tokens = mecab.morphs(sentence)
    tokens = [word for word in tokens if word not in stopwords]
    tokenized_data.append(tokens)</code></pre>
<p>MeCab은 붙어 있는 한국어 문장을 형태소 단위로 나눈다.</p>
<pre><code class="language-python">mecab.morphs(&quot;아버지가방에들어가신다&quot;)</code></pre>
<pre><code class="language-text">[&#39;아버지&#39;, &#39;가&#39;, &#39;방&#39;, &#39;에&#39;, &#39;들어가&#39;, &#39;신다&#39;]</code></pre>
<h3 id="gensim으로-word2vec-학습하기">Gensim으로 Word2Vec 학습하기</h3>
<pre><code class="language-python">from gensim.models import Word2Vec

model = Word2Vec(
    sentences=tokenized_data,
    vector_size=100,
    window=5,
    min_count=5,
    workers=2,
    sg=1,
    negative=5,
)</code></pre>
<table>
<thead>
<tr>
<th>인자</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td><code>vector_size</code></td>
<td>단어 벡터 차원이다.</td>
</tr>
<tr>
<td><code>window</code></td>
<td>중심 단어 앞뒤로 확인할 단어 범위다.</td>
</tr>
<tr>
<td><code>min_count</code></td>
<td>이 횟수 이상 등장한 단어만 학습한다.</td>
</tr>
<tr>
<td><code>workers</code></td>
<td>학습에 사용할 CPU Worker 수다.</td>
</tr>
<tr>
<td><code>sg</code></td>
<td><code>0</code>은 CBOW, <code>1</code>은 Skip-gram이다.</td>
</tr>
<tr>
<td><code>negative</code></td>
<td>Negative Sampling에 사용할 단어 수다.</td>
</tr>
</tbody></table>
<p>학습 결과는 3,990개 단어를 각각 100차원 벡터로 표현했다.</p>
<pre><code class="language-python">print(model.wv.vectors.shape)</code></pre>
<pre><code class="language-text">(3990, 100)</code></pre>
<p><code>most_similar()</code>는 코사인 유사도를 기준으로 가까운 단어를 찾는다.</p>
<pre><code class="language-python">model.wv.most_similar(&quot;블록버스터&quot;, topn=5)</code></pre>
<p>데이터를 20,000개 리뷰로 제한했기 때문에 결과 품질은 전체 데이터와 학습 설정에 따라 달라질 수 있다.</p>
<hr>
<h2 id="4-fasttext와-한글-자모-분해">4. FastText와 한글 자모 분해</h2>
<p>Word2Vec은 Vocabulary에 없는 단어의 벡터를 바로 만들기 어렵다. FastText는 단어를 문자 n-gram 단위로 나누어 학습하므로, 처음 보는 단어나 오타도 이미 학습한 부분 문자열을 조합해 표현할 수 있다.</p>
<pre><code class="language-text">playing
→ &lt;pl, pla, lay, ayi, yin, ing, ng&gt;</code></pre>
<p>한국어는 조사와 어미가 붙어 단어 형태가 다양하다. 자모 단위로 분해하면 비슷한 철자를 가진 단어와 오타가 더 많은 부분을 공유한다.</p>
<h3 id="hgtk로-한글-분해와-조합하기"><code>hgtk</code>로 한글 분해와 조합하기</h3>
<pre><code class="language-python">import hgtk

print(hgtk.letter.decompose(&quot;김&quot;))
print(hgtk.letter.compose(&quot;ㄱ&quot;, &quot;ㅣ&quot;, &quot;ㅁ&quot;))</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">(&#39;ㄱ&#39;, &#39;ㅣ&#39;, &#39;ㅁ&#39;)
김</code></pre>
<p>받침이 없는 글자는 빈 문자열 대신 <code>-</code>를 넣어 항상 세 글자 단위로 표현한다.</p>
<pre><code class="language-python">def word_to_jamo(token):
    def to_special_token(jamo):
        return jamo if jamo else &quot;-&quot;

    decomposed_token = &quot;&quot;

    for char in token:
        try:
            cho, jung, jong = hgtk.letter.decompose(char)
            decomposed_token += (
                to_special_token(cho)
                + to_special_token(jung)
                + to_special_token(jong)
            )
        except Exception as error:
            if type(error).__name__ == &quot;NotHangulException&quot;:
                decomposed_token += char

    return decomposed_token</code></pre>
<pre><code class="language-python">print(word_to_jamo(&quot;남동생&quot;))
print(word_to_jamo(&quot;여동생&quot;))</code></pre>
<pre><code class="language-text">ㄴㅏㅁㄷㅗㅇㅅㅐㅇ
ㅇㅕ-ㄷㅗㅇㅅㅐㅇ</code></pre>
<p>MeCab 형태소 분석과 자모 분해를 연결한다.</p>
<pre><code class="language-python">def tokenize_by_jamo(sentence):
    return [word_to_jamo(token) for token in mecab.morphs(sentence)]</code></pre>
<p>FastText 실습에는 네이버 쇼핑 리뷰 20만 건을 사용했다.</p>
<pre><code class="language-python">urllib.request.urlretrieve(
    &quot;https://raw.githubusercontent.com/bab2min/corpus/master/sentiment/naver_shopping.txt&quot;,
    filename=&quot;ratings_total.txt&quot;,
)

total_data = pd.read_table(
    &quot;ratings_total.txt&quot;,
    names=[&quot;ratings&quot;, &quot;reviews&quot;],
)</code></pre>
<p>각 리뷰를 형태소로 나눈 뒤 자모 단위로 변환한다.</p>
<pre><code class="language-python">from tqdm import tqdm

tokenized_data = []

for review in tqdm(total_data[&quot;reviews&quot;].to_list()):
    tokenized_sample = tokenize_by_jamo(review)
    tokenized_data.append(tokenized_sample)</code></pre>
<p>분해한 Token을 텍스트 파일로 저장하고 FastText를 학습한다.</p>
<pre><code class="language-python">import fasttext

with open(&quot;tokenized_data.txt&quot;, &quot;w&quot;, encoding=&quot;utf8&quot;) as output_file:
    for line in tokenized_data:
        output_file.write(&quot; &quot;.join(line) + &quot;\n&quot;)

model = fasttext.train_unsupervised(
    &quot;tokenized_data.txt&quot;,
    model=&quot;cbow&quot;,
)

model.save_model(&quot;fasttext.bin&quot;)</code></pre>
<p>노트북에서는 <code>남동쉥</code>, <code>남동셍ㅋ</code>, <code>제품^^</code>처럼 변형되거나 오타가 있는 입력에서도 원래 단어와 관련된 결과가 나타나는지 확인했다.</p>
<pre><code class="language-text">남동쉥 → 남동생, 남친, 남매, 남짓, 남녀 ...
제품^^ → 제품, 제풍, 반제품, 완제품, 상품 ...</code></pre>
<p>이는 FastText가 단어 전체만 외우지 않고 문자 조각을 함께 학습하기 때문이다.</p>
<hr>
<h2 id="5-시퀀스-데이터">5. 시퀀스 데이터</h2>
<p>시퀀스 데이터(Sequence Data)는 값의 <strong>순서와 앞뒤 관계가 중요한 데이터</strong>다.</p>
<table>
<thead>
<tr>
<th>데이터</th>
<th>순서가 중요한 이유</th>
</tr>
</thead>
<tbody><tr>
<td>자연어 문장</td>
<td>단어 순서가 달라지면 문장의 의미가 달라진다.</td>
</tr>
<tr>
<td>주가·날씨</td>
<td>이전 시점의 변화가 다음 시점과 연결된다.</td>
</tr>
<tr>
<td>음성</td>
<td>시간에 따른 파형의 흐름이 의미를 만든다.</td>
</tr>
<tr>
<td>영상</td>
<td>연속된 Frame의 변화가 움직임을 나타낸다.</td>
</tr>
</tbody></table>
<p>일반적인 MLP는 각 입력을 독립적으로 처리한다. 반면 RNN은 이전 시점의 정보를 Hidden State에 담아 다음 시점 계산에 전달한다.</p>
<hr>
<h2 id="6-rnn의-동작-원리">6. RNN의 동작 원리</h2>
<p>RNN(Recurrent Neural Network)은 현재 입력과 이전 Hidden State를 함께 사용해 새로운 Hidden State를 만든다.</p>
<pre><code class="language-text">h_t = tanh(W_xh x_t + W_hh h_(t-1) + b)</code></pre>
<table>
<thead>
<tr>
<th>기호</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td><code>x_t</code></td>
<td>현재 시점의 입력이다.</td>
</tr>
<tr>
<td><code>h_(t-1)</code></td>
<td>이전 시점까지의 정보를 담은 Hidden State다.</td>
</tr>
<tr>
<td><code>h_t</code></td>
<td>현재 입력까지 반영한 새로운 Hidden State다.</td>
</tr>
<tr>
<td><code>W_xh</code>, <code>W_hh</code></td>
<td>학습되는 가중치다.</td>
</tr>
<tr>
<td><code>tanh</code></td>
<td>값을 <code>-1~1</code> 범위로 변환하는 활성화 함수다.</td>
</tr>
</tbody></table>
<p>Hidden State는 가중치 자체가 아니라, 현재 시점까지 본 정보 중 다음 계산에 필요한 내용을 압축한 <strong>작업 메모리</strong>다.</p>
<p>RNN은 시점마다 새로운 Cell을 학습하는 것이 아니다. 같은 RNN Cell과 같은 가중치를 전체 시간축에서 반복해서 사용한다.</p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/4188e76c-9c14-4494-9c32-5e0c6f80fc5d/image.png" alt=""></p>
<p><img src="https://velog.velcdn.com/images/j_keun/post/4e993366-1317-41a2-91e7-03163dd4b620/image.png" alt=""></p>
<h3 id="입력-shape">입력 Shape</h3>
<p><code>batch_first=True</code>인 RNN의 입력 Shape는 다음과 같다.</p>
<pre><code class="language-text">[batch_size, sequence_length, input_size]</code></pre>
<p>예를 들어 <code>[32, 20, 10]</code>은 다음을 의미한다.</p>
<ul>
<li>한 번에 32개 시퀀스를 처리한다.</li>
<li>시퀀스 하나는 연속된 20개 시점으로 구성된다.</li>
<li>각 시점은 10개 Feature를 가진다.</li>
</ul>
<h3 id="장기-의존성-문제">장기 의존성 문제</h3>
<p>기본 RNN은 가까운 과거의 패턴을 처리하는 데 유용하지만, 오래전 정보를 마지막까지 유지하기 어렵다. 같은 가중치를 시간축에서 반복 적용하며 역전파하기 때문에 긴 시퀀스에서는 기울기 소실 또는 기울기 폭주가 발생하기 쉽다.</p>
<p>이 문제를 완화하기 위해 LSTM과 GRU 같은 구조가 사용된다.</p>
<hr>
<h2 id="7-rnn으로-다음-kospi-종가-예측하기">7. RNN으로 다음 KOSPI 종가 예측하기</h2>
<p>최근 20거래일의 정보를 입력으로 사용해 다음 거래일의 종가를 예측한다.</p>
<pre><code class="language-text">1일차 입력 + 초기 Hidden State → h1
2일차 입력 + h1                → h2
...
20일차 입력 + h19              → h20
h20 → Linear Layer → 21일차 종가 예측</code></pre>
<h3 id="데이터-확인과-정제">데이터 확인과 정제</h3>
<p>원본 데이터는 4,513행이며 다음 열을 가진다.</p>
<table>
<thead>
<tr>
<th>열</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td><code>Open</code></td>
<td>시가</td>
</tr>
<tr>
<td><code>High</code></td>
<td>고가</td>
</tr>
<tr>
<td><code>Low</code></td>
<td>저가</td>
</tr>
<tr>
<td><code>Close</code></td>
<td>종가</td>
</tr>
<tr>
<td><code>Adj Close</code></td>
<td>배당과 주식 분할 등을 반영한 수정 종가</td>
</tr>
<tr>
<td><code>Volume</code></td>
<td>거래량</td>
</tr>
</tbody></table>
<p>날짜를 <code>datetime</code>으로 바꾸고 필수 열의 결측치를 제거한 뒤 날짜순으로 정렬한다.</p>
<pre><code class="language-python">df[&quot;Date&quot;] = pd.to_datetime(
    df[&quot;Date&quot;],
    dayfirst=True,
    errors=&quot;coerce&quot;,
)

numeric_cols = [&quot;Open&quot;, &quot;High&quot;, &quot;Low&quot;, &quot;Close&quot;, &quot;Adj Close&quot;, &quot;Volume&quot;]

for column in numeric_cols:
    df[column] = pd.to_numeric(df[column], errors=&quot;coerce&quot;)

required_cols = [&quot;Date&quot;, &quot;Open&quot;, &quot;High&quot;, &quot;Low&quot;, &quot;Close&quot;]
df_copy = df.dropna(subset=required_cols).copy()
df_copy = df_copy.sort_values(&quot;Date&quot;).reset_index(drop=True)</code></pre>
<p>정제 후 2000년 1월 4일부터 2018년 1월 22일까지 총 4,452행이 남았다.</p>
<h3 id="feature-engineering">Feature Engineering</h3>
<p>당일 가격과 최근 가격 흐름을 함께 표현하도록 파생변수를 만든다.</p>
<pre><code class="language-python">data = df_copy.copy()

data[&quot;Return_1d&quot;] = data[&quot;Close&quot;].pct_change()
data[&quot;Range_pct&quot;] = (data[&quot;High&quot;] - data[&quot;Low&quot;]) / data[&quot;Close&quot;]
data[&quot;OC_pct&quot;] = (data[&quot;Close&quot;] - data[&quot;Open&quot;]) / data[&quot;Open&quot;]

data[&quot;MA5&quot;] = data[&quot;Close&quot;].rolling(5).mean()
data[&quot;MA20&quot;] = data[&quot;Close&quot;].rolling(20).mean()

data[&quot;MA5_ratio&quot;] = data[&quot;Close&quot;] / data[&quot;MA5&quot;]
data[&quot;MA20_ratio&quot;] = data[&quot;Close&quot;] / data[&quot;MA20&quot;]
data[&quot;Volatility_10&quot;] = data[&quot;Return_1d&quot;].rolling(10).std()

data[&quot;Target_Close&quot;] = data[&quot;Close&quot;].shift(-1)
data[&quot;Target_Date&quot;] = data[&quot;Date&quot;].shift(-1)</code></pre>
<p>사용하는 Feature는 총 10개다.</p>
<pre><code class="language-python">FEATURES = [
    &quot;Open&quot;,
    &quot;High&quot;,
    &quot;Low&quot;,
    &quot;Close&quot;,
    &quot;Return_1d&quot;,
    &quot;Range_pct&quot;,
    &quot;OC_pct&quot;,
    &quot;MA5_ratio&quot;,
    &quot;MA20_ratio&quot;,
    &quot;Volatility_10&quot;,
]</code></pre>
<p><code>shift(-1)</code>을 사용하면 현재 행의 Target에 다음 거래일 종가가 들어간다. 이동평균과 다음 날 Target을 만드는 과정에서 생긴 결측치를 제거한 결과 학습에 사용할 데이터는 4,432행이다.</p>
<hr>
<h2 id="8-시계열-데이터-분리와-scaling">8. 시계열 데이터 분리와 Scaling</h2>
<p>시계열 데이터를 무작위로 나누면 미래 데이터가 Train에 포함되고 과거 데이터가 Test에 포함될 수 있다. 따라서 날짜순으로 다음과 같이 분리한다.</p>
<pre><code class="language-python">n = len(data)

train_end = int(n * 0.70)
val_end = int(n * 0.85)</code></pre>
<table>
<thead>
<tr>
<th>분할</th>
<th align="right">행 개수</th>
<th>Target 기간</th>
</tr>
</thead>
<tbody><tr>
<td>Train</td>
<td align="right">3,102</td>
<td>2000-02-01 ~ 2012-08-20</td>
</tr>
<tr>
<td>Validation</td>
<td align="right">665</td>
<td>2012-08-21 ~ 2015-05-04</td>
</tr>
<tr>
<td>Test</td>
<td align="right">665</td>
<td>2015-05-06 ~ 2018-01-22</td>
</tr>
</tbody></table>
<p>Feature마다 값의 크기가 다르므로 <code>StandardScaler</code>로 평균 0, 표준편차 1에 가깝게 변환한다.</p>
<pre><code class="language-python">from sklearn.preprocessing import StandardScaler

scaler_x = StandardScaler()
scaler_y = StandardScaler()

scaler_x.fit(data.loc[:train_end - 1, FEATURES])
scaler_y.fit(data.loc[:train_end - 1, [&quot;Target_Close&quot;]])

X_all = scaler_x.transform(data[FEATURES]).astype(np.float32)
y_all = scaler_y.transform(data[[&quot;Target_Close&quot;]]).astype(np.float32)</code></pre>
<p>Scaler는 반드시 Train 데이터에만 <code>fit()</code>한다. Validation과 Test를 함께 사용하면 미래 데이터의 평균과 표준편차를 학습 단계에서 미리 알게 되는 데이터 누수가 발생한다.</p>
<pre><code class="language-text">X_all Shape: (4432, 10)
y_all Shape: (4432, 1)</code></pre>
<hr>
<h2 id="9-연속된-20일을-하나의-시퀀스로-만들기">9. 연속된 20일을 하나의 시퀀스로 만들기</h2>
<p>RNN은 한 행씩 독립적으로 보는 대신 연속된 여러 행을 하나의 입력으로 받는다.</p>
<pre><code class="language-python">SEQUENCE_LENGTH = 20

def make_sequence(X, y, sequence_length):
    X_seq = []
    y_seq = []
    target_indices = []

    for index in range(sequence_length - 1, len(X)):
        start = index - sequence_length + 1
        X_seq.append(X[start:index + 1])
        y_seq.append(y[index])
        target_indices.append(index)

    return (
        np.asarray(X_seq, dtype=np.float32),
        np.asarray(y_seq, dtype=np.float32),
        np.asarray(target_indices, dtype=np.int64),
    )</code></pre>
<p>Window는 하루씩 이동한다.</p>
<pre><code class="language-text">첫 번째 입력: 0~19번 행 → 19번 행에 연결된 다음 날 종가
두 번째 입력: 1~20번 행 → 20번 행에 연결된 다음 날 종가
세 번째 입력: 2~21번 행 → 21번 행에 연결된 다음 날 종가</code></pre>
<p>전체 데이터에서 먼저 Window를 만들고 Target 시점의 인덱스를 기준으로 Train, Validation, Test를 나눈다. 이렇게 하면 Validation과 Test의 첫 번째 예측에서도 그 이전 날짜를 입력 이력으로 사용할 수 있다.</p>
<pre><code class="language-text">전체 Sequence Shape: (4413, 20, 10)
전체 Target Shape: (4413, 1)</code></pre>
<p>분리 결과는 다음과 같다.</p>
<pre><code class="language-text">Train: (3083, 20, 10), (3083, 1)
Validation: (665, 20, 10), (665, 1)
Test: (665, 20, 10), (665, 1)</code></pre>
<hr>
<h2 id="10-dataset과-dataloader">10. Dataset과 DataLoader</h2>
<pre><code class="language-python">from torch.utils.data import Dataset, DataLoader

class SequenceDataset(Dataset):
    def __init__(self, X, y):
        self.X = torch.tensor(X, dtype=torch.float32)
        self.y = torch.tensor(y, dtype=torch.float32)

    def __len__(self):
        return len(self.X)

    def __getitem__(self, index):
        return self.X[index], self.y[index]</code></pre>
<pre><code class="language-python">BATCH_SIZE = 64

train_loader = DataLoader(
    SequenceDataset(X_train, y_train),
    batch_size=BATCH_SIZE,
    shuffle=True,
)

val_loader = DataLoader(
    SequenceDataset(X_val, y_val),
    batch_size=BATCH_SIZE,
    shuffle=False,
)

test_loader = DataLoader(
    SequenceDataset(X_test, y_test),
    batch_size=BATCH_SIZE,
    shuffle=False,
)</code></pre>
<p>Train의 <code>shuffle=True</code>는 20일로 구성된 Window들의 배치 순서를 섞는다. 각 Window 내부의 날짜 순서는 바뀌지 않는다.</p>
<pre><code class="language-text">입력 Batch Shape: torch.Size([64, 20, 10])
정답 Batch Shape: torch.Size([64, 1])</code></pre>
<hr>
<h2 id="11-rnn-회귀-모델-만들기">11. RNN 회귀 모델 만들기</h2>
<pre><code class="language-python">class RNNRegressor(nn.Module):
    def __init__(self, input_size, hidden_size=64, num_layers=2, dropout=0.10):
        super().__init__()

        self.rnn = nn.RNN(
            input_size=input_size,
            hidden_size=hidden_size,
            num_layers=num_layers,
            nonlinearity=&quot;tanh&quot;,
            batch_first=True,
            dropout=dropout if num_layers &gt; 1 else 0.0,
        )

        self.fc = nn.Linear(hidden_size, 1)

    def forward(self, x):
        out, h_n = self.rnn(x)
        last_hidden = out[:, -1, :]
        prediction = self.fc(last_hidden)
        return prediction</code></pre>
<pre><code class="language-python">model = RNNRegressor(
    input_size=len(FEATURES),
    hidden_size=64,
    num_layers=2,
    dropout=0.10,
).to(DEVICE)</code></pre>
<p>모델 구조는 다음과 같다.</p>
<pre><code class="language-text">RNNRegressor(
  (rnn): RNN(10, 64, num_layers=2, batch_first=True, dropout=0.1)
  (fc): Linear(in_features=64, out_features=1, bias=True)
)</code></pre>
<h3 id="out과-h_n"><code>out</code>과 <code>h_n</code></h3>
<p>4개 샘플을 모델의 RNN 층에 전달한 결과는 다음과 같다.</p>
<pre><code class="language-text">입력:          [4, 20, 10]
RNN out:       [4, 20, 64]
RNN h_n:       [2, 4, 64]
마지막 시점:   [4, 64]</code></pre>
<ul>
<li><code>out</code>은 마지막 RNN Layer가 모든 시점에서 만든 Hidden State다.</li>
<li><code>h_n</code>은 각 RNN Layer의 마지막 Hidden State다.</li>
<li>다음 날 종가 예측에는 <code>out[:, -1, :]</code>로 마지막 시점의 정보를 사용한다.</li>
<li><code>Linear(64, 1)</code>은 64개 Hidden 값을 종가 예측값 하나로 바꾼다.</li>
</ul>
<hr>
<h2 id="12-학습-설정">12. 학습 설정</h2>
<pre><code class="language-python">criterion = nn.HuberLoss(delta=1.0)

optimizer = torch.optim.AdamW(
    model.parameters(),
    lr=1e-3,
    weight_decay=1e-4,
)

scheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(
    optimizer=optimizer,
    mode=&quot;min&quot;,
    factor=0.5,
    patience=5,
)

MAX_EPOCHS = 200
PATIENCE = 15
CLIP_NORM = 1.0</code></pre>
<table>
<thead>
<tr>
<th>설정</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td><code>HuberLoss</code></td>
<td>작은 오차에는 제곱 오차를, 큰 오차에는 선형 오차를 적용해 이상치의 영향을 완화한다.</td>
</tr>
<tr>
<td><code>AdamW</code></td>
<td>Parameter를 업데이트하고 Weight Decay를 적용한다.</td>
</tr>
<tr>
<td><code>ReduceLROnPlateau</code></td>
<td>Validation Loss가 개선되지 않으면 Learning Rate를 줄인다.</td>
</tr>
<tr>
<td>Early Stopping</td>
<td>Validation Loss가 일정 기간 개선되지 않으면 학습을 종료한다.</td>
</tr>
<tr>
<td>Gradient Clipping</td>
<td>기울기 크기를 제한해 RNN의 기울기 폭주를 완화한다.</td>
</tr>
</tbody></table>
<pre><code class="language-python">loss.backward()
torch.nn.utils.clip_grad_norm_(
    model.parameters(),
    max_norm=CLIP_NORM,
)
optimizer.step()</code></pre>
<p>Validation Loss가 가장 작을 때 모델 상태를 저장한다.</p>
<pre><code class="language-python">if val_loss &lt; best_val_loss:
    best_val_loss = val_loss
    best_state = copy.deepcopy(model.state_dict())
    wait = 0
else:
    wait += 1</code></pre>
<p>학습이 끝난 뒤에는 최적 상태를 다시 불러온 다음 Test 예측을 수행한다.</p>
<pre><code class="language-python">model.load_state_dict(best_state)</code></pre>
<hr>
<h2 id="13-학습-결과와-예측값-복원">13. 학습 결과와 예측값 복원</h2>
<p>노트북에서는 36번째 Epoch에서 Early Stopping이 실행됐다.</p>
<pre><code class="language-text">Epoch 001 | Train 0.046967 | Valid 0.001290 | LR 1.00e-03
Epoch 010 | Train 0.001861 | Valid 0.000449 | LR 1.00e-03
Epoch 020 | Train 0.001287 | Valid 0.000699 | LR 2.50e-04
Epoch 030 | Train 0.001110 | Valid 0.000472 | LR 1.25e-04
Early Stopping: epoch 36</code></pre>
<p>모델은 표준화된 종가를 예측하므로 <code>inverse_transform()</code>으로 실제 종가 단위로 되돌린다.</p>
<pre><code class="language-python">pred_scaled, true_scaled = predict(model, test_loader, DEVICE)

pred_close = scaler_y.inverse_transform(pred_scaled).ravel()
true_close = scaler_y.inverse_transform(true_scaled).ravel()</code></pre>
<p>실행 결과:</p>
<pre><code class="language-text">예측 개수: 665
예측값 앞 5개: [2130.90, 2109.48, 2097.60, 2090.25, 2100.17]
실제값 앞 5개: [2104.58, 2091.00, 2085.52, 2097.38, 2096.77]</code></pre>
<p>앞의 몇 개 값이 비슷해 보인다는 것만으로 모델 성능을 판단할 수는 없다. 전체 Test 데이터에 대해 MAE와 RMSE를 계산하고, 실제값과 예측값의 시계열 그래프도 함께 확인해야 한다.</p>
<pre><code class="language-python">from sklearn.metrics import mean_absolute_error, mean_squared_error

mae = mean_absolute_error(true_close, pred_close)
rmse = mean_squared_error(true_close, pred_close) ** 0.5

print(&quot;MAE:&quot;, mae)
print(&quot;RMSE:&quot;, rmse)</code></pre>
<hr>
<h2 id="14-전체-흐름-정리">14. 전체 흐름 정리</h2>
<pre><code class="language-text">텍스트 데이터
  ↓
형태소 분석과 불용어 제거
  ↓
원-핫·BOW·TF-IDF 또는 Word Embedding
  ↓
Word2Vec / FastText로 단어 관계 학습

시계열 데이터
  ↓
Feature Engineering
  ↓
시간순 Train / Validation / Test 분리
  ↓
Train 기준 Scaling
  ↓
연속된 20일 Window 생성
  ↓
RNN → 마지막 Hidden State → Linear
  ↓
다음 거래일 종가 예측</code></pre>
<ul>
<li>원-핫 인코딩과 BOW는 단순하지만 고차원 희소 벡터가 되고 문맥을 충분히 반영하지 못한다.</li>
<li><code>nn.Embedding</code>은 Token ID를 학습 가능한 밀집 벡터로 바꾼다.</li>
<li>Word2Vec은 주변 문맥으로 단어의 의미 관계를 학습한다.</li>
<li>FastText는 문자 n-gram을 사용해 OOV와 오타에 비교적 강하다.</li>
<li>RNN은 이전 Hidden State를 다음 시점 계산에 전달해 순서 정보를 처리한다.</li>
<li>시계열은 무작위 분할하지 않고 시간순으로 분리해야 한다.</li>
<li>Scaling 기준은 Train 데이터에서만 계산해야 데이터 누수를 막을 수 있다.</li>
<li>기본 RNN은 장기 의존성과 기울기 문제를 가지며, LSTM과 GRU가 이를 보완한다.</li>
</ul>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 디펜스 게임]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EB%94%94%ED%8E%9C%EC%8A%A4-%EA%B2%8C%EC%9E%84</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EB%94%94%ED%8E%9C%EC%8A%A4-%EA%B2%8C%EC%9E%84</guid>
            <pubDate>Mon, 28 Sep 2026 14:33:17 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>병사 <code>n</code>명으로 순서대로 등장하는 적을 막는다.</p>
<ul>
<li>한 라운드를 일반적으로 막으면 적 수만큼 병사가 줄어든다.</li>
<li>무적권은 최대 <code>k</code>번 사용할 수 있고, 사용한 라운드에서는 병사가 줄지 않는다.</li>
</ul>
<p>최대로 막을 수 있는 라운드 수를 구한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p>어떤 시점까지 막은 라운드들 중에서 무적권은 적 수가 가장 많은 <code>k</code>개 라운드에 쓰는 것이 항상 최적이다.</p>
<p>적 수가 작은 라운드에 무적권을 쓰고 큰 라운드에 병사를 쓰는 경우가 있다면, 두 라운드의 무적권 사용 여부를 바꾸면 병사 소모는 줄어들거나 같아진다.</p>
<p>따라서 각 라운드에서 다음 상태를 유지한다.</p>
<ul>
<li>최소 힙: 지금까지 등장한 공격 중 무적권을 사용할 <code>k</code>개의 큰 공격</li>
<li>병사: 무적권을 쓰지 않는 나머지 공격을 막는 데 사용</li>
</ul>
<p>힙 크기가 <code>k</code>를 넘으면, 힙에서 가장 작은 공격을 꺼내 병사로 막는다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>적 수를 최소 힙에 넣는다.</li>
<li>힙의 크기가 <code>k</code>보다 크면, 가장 작은 적 수를 꺼내 병사에서 뺀다.</li>
<li>병사가 음수가 되면 현재 라운드를 막을 수 없으므로, 현재 인덱스를 반환한다.</li>
<li>모든 라운드를 처리하면 전체 라운드 수를 반환한다.</li>
</ol>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">import heapq


def solution(n, k, enemy):
    invincible_rounds = []

    for round_index, enemy_count in enumerate(enemy):
        # 현재까지는 이 라운드에도 무적권을 쓴다고 가정한다.
        heapq.heappush(invincible_rounds, enemy_count)

        # 무적권 수를 넘으면 가장 작은 공격은 병사로 막는다.
        if len(invincible_rounds) &gt; k:
            n -= heapq.heappop(invincible_rounds)

        if n &lt; 0:
            return round_index

    return len(enemy)</code></pre>
<h2 id="예시">예시</h2>
<p><code>n = 7</code>, <code>k = 3</code>, <code>enemy = [4, 2, 4, 5, 3, 3, 1]</code>인 경우를 보자.</p>
<p>5라운드까지 처리한 뒤 무적권을 사용할 공격은 가장 큰 <code>5, 4, 4</code>다.</p>
<pre><code class="language-text">무적권 사용: 5, 4, 4
병사로 처리: 2 + 3 = 5
남은 병사: 7 - 5 = 2</code></pre>
<p>6라운드의 적 수 3까지 포함하면, 무적권을 제외하고 병사로 막아야 하는 적 수의 합이 10이 된다. 병사가 부족하므로 5라운드까지만 막을 수 있다.</p>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p><code>N</code>을 전체 라운드 수라고 하자.</p>
<p>힙에는 최대 <code>k</code>개의 원소만 유지한다.</p>
<ul>
<li>힙 삽입과 삭제: <code>O(log k)</code></li>
<li>전체 라운드 처리: <code>O(N log k)</code></li>
<li>공간 복잡도: <code>O(k)</code></li>
</ul>
<h2 id="정리">정리</h2>
<p>매 순간 무적권을 가장 큰 공격들에 배정한다고 생각하면 된다. 최소 힙에서 가장 작은 값을 병사로 처리하면서, 지금까지의 최적 무적권 배정을 계속 유지할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 자동완성]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%9E%90%EB%8F%99%EC%99%84%EC%84%B1</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%9E%90%EB%8F%99%EC%99%84%EC%84%B1</guid>
            <pubDate>Sun, 27 Sep 2026 18:45:27 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>중복 없는 학습 단어들이 주어졌을 때, 각 단어를 다른 단어와 구분해 자동완성하려면 몇 글자를 입력해야 하는지 구한다.</p>
<p>모든 단어에 필요한 입력 글자 수의 합을 반환한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p>단어를 사전순으로 정렬하면, 어떤 단어와 가장 긴 접두사를 공유할 수 있는 단어는 정렬된 목록에서 바로 앞 또는 바로 뒤에 있다.</p>
<p>따라서 현재 단어가 필요한 입력 글자 수는 다음처럼 구할 수 있다.</p>
<pre><code class="language-text">max(앞 단어와의 공통 접두사 길이, 뒤 단어와의 공통 접두사 길이) + 1</code></pre>
<p>단, 현재 단어 자체가 다른 단어의 접두사라면 더 입력할 문자가 없으므로 단어 전체를 입력해야 한다.</p>
<pre><code class="language-python">required = min(len(word), longest_common_prefix + 1)</code></pre>
<h2 id="왜-인접한-단어만-비교할까">왜 인접한 단어만 비교할까?</h2>
<p>사전순 정렬에서 같은 접두사를 가진 단어들은 항상 연속해서 모인다.</p>
<p>예를 들어 <code>word</code>와 <code>wor...</code>로 시작하는 단어들은 모두 한 구간에 모인다. 현재 단어와 가장 긴 접두사를 공유하는 단어는 그 구간에서 바로 앞이나 바로 뒤에 있으므로, 두 이웃만 비교하면 충분하다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>단어 목록을 사전순으로 정렬한다.</li>
<li>각 단어와 앞 단어의 최장 공통 접두사 길이를 구한다.</li>
<li>각 단어와 뒤 단어의 최장 공통 접두사 길이를 구한다.</li>
<li>둘 중 큰 값에 1을 더한다.</li>
<li>현재 단어 길이를 넘지 않도록 제한한 값을 답에 더한다.</li>
</ol>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">def common_prefix_length(first, second):
    limit = min(len(first), len(second))
    index = 0

    while index &lt; limit and first[index] == second[index]:
        index += 1

    return index


def solution(words):
    words.sort()
    total = 0

    for index, word in enumerate(words):
        longest_common_prefix = 0

        if index &gt; 0:
            longest_common_prefix = max(
                longest_common_prefix,
                common_prefix_length(word, words[index - 1])
            )

        if index + 1 &lt; len(words):
            longest_common_prefix = max(
                longest_common_prefix,
                common_prefix_length(word, words[index + 1])
            )

        # word가 다른 단어의 접두사인 경우에는 단어 전체를 입력해야 한다.
        total += min(len(word), longest_common_prefix + 1)

    return total</code></pre>
<h2 id="예시">예시</h2>
<p><code>words = [&quot;go&quot;, &quot;gone&quot;, &quot;guild&quot;]</code>는 이미 사전순으로 정렬되어 있다.</p>
<table>
<thead>
<tr>
<th>단어</th>
<th>이웃 단어와의 가장 긴 공통 접두사</th>
<th>필요한 입력</th>
</tr>
</thead>
<tbody><tr>
<td>go</td>
<td>go (gone과 공유)</td>
<td>2</td>
</tr>
<tr>
<td>gone</td>
<td>go</td>
<td>3</td>
</tr>
<tr>
<td>guild</td>
<td>g</td>
<td>2</td>
</tr>
</tbody></table>
<p><code>go</code>는 <code>gone</code>의 접두사다. 공통 접두사 길이는 2이지만 <code>go</code> 뒤에 입력할 글자가 없으므로, 단어 전체인 2글자를 입력해야 한다.</p>
<p>총 입력 글자 수는 <code>2 + 3 + 2 = 7</code>이다.</p>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p><code>N</code>을 단어 개수, <code>L</code>을 모든 단어 길이의 합, <code>M</code>을 가장 긴 단어 길이라고 하자.</p>
<ul>
<li><p>단어 정렬: 최대 <code>O(N log N * M)</code></p>
</li>
<li><p>인접 단어 공통 접두사 계산: 각 단어는 앞뒤 단어와만 비교하므로 <code>O(L)</code></p>
</li>
<li><p>시간 복잡도: <code>O(N log N * M + L)</code></p>
</li>
<li><p>공간 복잡도: 정렬 구현을 포함해 <code>O(N)</code></p>
</li>
</ul>
<h2 id="정리">정리</h2>
<p>자동완성에 필요한 글자 수는 해당 단어와 가장 비슷한 다른 단어를 구분할 수 있는 위치로 결정된다. 사전순 정렬 후 앞뒤 단어의 공통 접두사만 비교하면, 트라이 없이도 필요한 총 입력 횟수를 구할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 숫자 게임]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%88%AB%EC%9E%90-%EA%B2%8C%EC%9E%84</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%88%AB%EC%9E%90-%EA%B2%8C%EC%9E%84</guid>
            <pubDate>Sun, 27 Sep 2026 18:31:14 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>A팀의 출전 순서는 이미 정해져 있고, B팀은 출전 순서를 자유롭게 정할 수 있다.</p>
<p>B팀 선수가 A팀 선수보다 큰 숫자를 낼 때만 1점을 얻는다. B팀이 얻을 수 있는 최대 승점을 구한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p>실제 출전 순서보다 어떤 숫자가 어떤 숫자를 이기는지가 중요하다. 따라서 A와 B를 모두 오름차순 정렬한다.</p>
<p>B의 작은 숫자부터 확인하면서, 아직 이기지 않은 A 숫자 중 가장 작은 숫자를 이길 수 있는지 확인한다.</p>
<ul>
<li><code>B의 숫자 &gt; A의 가장 작은 미매칭 숫자</code>: 이 A 선수를 이기고 승점을 얻는다.</li>
<li>그렇지 않다: 이 B 선수는 어떤 남은 A 선수도 이길 수 없으므로 승점 없이 넘긴다.</li>
</ul>
<p>이길 수 있는 B 선수를 더 큰 A 숫자에 쓰면 더 작은 A 숫자를 이길 기회를 잃을 수 있다. 따라서 가능한 가장 작은 A 숫자와 매칭하는 선택이 항상 최적이다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>A와 B를 오름차순 정렬한다.</li>
<li>A에서 아직 이기지 않은 가장 작은 숫자의 인덱스를 <code>a_index</code>로 둔다.</li>
<li>정렬된 B를 앞에서부터 확인한다.</li>
<li>현재 B 숫자가 <code>A[a_index]</code>보다 크면 승점을 1 늘리고 <code>a_index</code>를 다음 A 선수로 옮긴다.</li>
<li>모든 B 선수를 확인한 뒤 승점을 반환한다.</li>
</ol>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">def solution(A, B):
    A.sort()
    B.sort()

    score = 0
    a_index = 0

    for b_number in B:
        # 현재 B 선수가 가장 작은 미매칭 A 선수를 이길 수 있다.
        if a_index &lt; len(A) and b_number &gt; A[a_index]:
            score += 1
            a_index += 1

    return score</code></pre>
<h2 id="예시">예시</h2>
<pre><code class="language-text">A = [5, 1, 3, 7] -&gt; [1, 3, 5, 7]
B = [2, 2, 6, 8] -&gt; [2, 2, 6, 8]</code></pre>
<table>
<thead>
<tr>
<th>B 숫자</th>
<th>이길 A 숫자</th>
<th>승점</th>
</tr>
</thead>
<tbody><tr>
<td>2</td>
<td>1</td>
<td>1</td>
</tr>
<tr>
<td>2</td>
<td>이길 수 없음</td>
<td>1</td>
</tr>
<tr>
<td>6</td>
<td>3</td>
<td>2</td>
</tr>
<tr>
<td>8</td>
<td>5</td>
<td>3</td>
</tr>
</tbody></table>
<p>따라서 B팀은 최대 <code>3</code>점을 얻는다.</p>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p><code>N</code>을 팀 인원 수라고 하자.</p>
<ul>
<li><p>A, B 정렬: <code>O(N log N)</code></p>
</li>
<li><p>B 배열 순회: <code>O(N)</code></p>
</li>
<li><p>시간 복잡도: <code>O(N log N)</code></p>
</li>
<li><p>공간 복잡도: 정렬 구현에 따라 <code>O(N)</code></p>
</li>
</ul>
<h2 id="정리">정리</h2>
<p>이길 수 있는 B 숫자는 남은 A 숫자 중 가장 작은 수를 이기는 데 사용한다. 이 단순한 그리디 기준을 정렬된 두 배열에 적용하면 B팀의 최대 승점을 구할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 단어 퍼즐]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EB%8B%A8%EC%96%B4-%ED%8D%BC%EC%A6%90</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EB%8B%A8%EC%96%B4-%ED%8D%BC%EC%A6%90</guid>
            <pubDate>Sun, 27 Sep 2026 18:27:16 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>주어진 단어 조각을 원하는 만큼 사용해 문자열 <code>t</code>를 완성한다.</p>
<p>문자열을 완성하는 데 필요한 단어 조각 수의 최솟값을 구하고, 만들 수 없다면 <code>-1</code>을 반환한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p><code>dp[i]</code>를 <code>t</code>의 앞에서부터 <code>i</code>글자까지 완성하는 데 필요한 최소 조각 수라고 정의한다.</p>
<pre><code class="language-text">dp[0] = 0</code></pre>
<p>어떤 위치 <code>i</code>까지 만들 수 있고, 그 위치부터 시작하는 조각 <code>piece</code>가 <code>t</code>와 일치한다면 다음 위치를 갱신한다.</p>
<pre><code class="language-python">dp[i + len(piece)] = min(
    dp[i + len(piece)],
    dp[i] + 1
)</code></pre>
<p>각 조각을 무한히 사용할 수 있으므로, 같은 조각을 여러 위치에서 반복해서 사용해도 된다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>길이가 <code>len(t) + 1</code>인 DP 배열을 만들고, 도달할 수 없는 값은 큰 값으로 초기화한다.</li>
<li><code>dp[0] = 0</code>으로 시작한다.</li>
<li>문자열의 각 위치에서, 해당 위치부터 일치하는 모든 단어 조각을 확인한다.</li>
<li>일치하는 조각의 끝 위치를 최소 조각 수로 갱신한다.</li>
<li><code>dp[len(t)]</code>가 여전히 큰 값이면 <code>-1</code>을 반환한다.</li>
</ol>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">def solution(strs, t):
    length = len(t)
    infinity = length + 1

    # dp[i]: t의 앞 i글자를 만드는 데 필요한 최소 조각 수
    dp = [infinity] * (length + 1)
    dp[0] = 0

    for start in range(length):
        if dp[start] == infinity:
            continue

        for piece in strs:
            end = start + len(piece)

            if end &lt;= length and t.startswith(piece, start):
                dp[end] = min(dp[end], dp[start] + 1)

    return -1 if dp[length] == infinity else dp[length]</code></pre>
<h2 id="예시">예시</h2>
<p><code>strs = [&quot;ba&quot;, &quot;na&quot;, &quot;n&quot;, &quot;a&quot;]</code>, <code>t = &quot;banana&quot;</code>인 경우를 보자.</p>
<pre><code class="language-text">dp[0] = 0
&quot;ba&quot; 사용  -&gt; dp[2] = 1
&quot;na&quot; 사용  -&gt; dp[4] = 2
&quot;na&quot; 사용  -&gt; dp[6] = 3</code></pre>
<p>따라서 <code>&quot;ba&quot; + &quot;na&quot; + &quot;na&quot;</code>로 문자열을 만들 수 있고, 필요한 조각 수는 <code>3</code>이다.</p>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p><code>N</code>을 <code>t</code>의 길이, <code>S</code>를 단어 조각 개수, <code>L</code>을 조각의 최대 길이라고 하자.</p>
<p>각 위치에서 모든 조각을 확인하고, 문자열 비교는 최대 <code>L</code>글자를 확인한다.</p>
<ul>
<li>시간 복잡도: <code>O(N * S * L)</code></li>
<li>공간 복잡도: <code>O(N)</code></li>
</ul>
<p>이 문제에서는 <code>N &lt;= 20,000</code>, <code>S &lt;= 100</code>, <code>L &lt;= 5</code>이므로 충분히 빠르게 동작한다.</p>
<h2 id="정리">정리</h2>
<p>문자열을 앞에서부터 완성하는 최소 비용 문제로 바꾸면 된다. <code>dp[i]</code>가 도달 가능한 위치인지 확인하고, 그 위치에 이어 붙일 수 있는 단어 조각으로 다음 상태를 갱신한다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[Malware Detection 정리]]></title>
            <link>https://velog.io/@j_keun/Malware-Detection-%EC%A0%95%EB%A6%AC</link>
            <guid>https://velog.io/@j_keun/Malware-Detection-%EC%A0%95%EB%A6%AC</guid>
            <pubDate>Sun, 27 Sep 2026 11:28:46 GMT</pubDate>
            <description><![CDATA[<h2 id="1-project-overview-프로젝트-개요">1. Project Overview (프로젝트 개요)</h2>
<h3 id="11-프로젝트-명과-개발-배경">1.1 프로젝트 명과 개발 배경</h3>
<table>
<thead>
<tr>
<th>항목</th>
<th>내용</th>
</tr>
</thead>
<tbody><tr>
<td>프로젝트 명</td>
<td><strong>Malware Detection</strong></td>
</tr>
<tr>
<td>서비스 형태</td>
<td>React 기반 웹 UI와 FastAPI 기반 분석 API로 구성된 Windows PE 정적 분석 서비스</td>
</tr>
<tr>
<td>분석 대상</td>
<td>단일 PE 파일 또는 폴더 안의 다중 파일</td>
</tr>
<tr>
<td>분석 원칙</td>
<td>업로드 파일을 실행·치료·수정·삭제하지 않고 바이트와 PE 헤더를 읽기 전용으로 처리</td>
</tr>
<tr>
<td>핵심 접근</td>
<td>1차 XGBoost 선별 → 필요 시 2차 MalConv2 계열 패밀리 분류</td>
</tr>
<tr>
<td>최종 목적</td>
<td>정상으로 보이는 파일과 위험 의심 파일을 구분하고, 검토 우선순위와 유형 정보를 제공</td>
</tr>
</tbody></table>
<p>Windows 실행 파일은 <code>.exe</code>처럼 보이는 확장자만으로 신뢰할 수 없다. 또한 다운로드 폴더나 공유 폴더에는 문서·이미지·실행 파일이 섞여 있어, 사용자는 어떤 파일부터 확인해야 하는지 빠르게 판단하기 어렵다.</p>
<p>이 프로젝트는 이러한 문제를 대상으로 다음 질문에 답하는 보조 분석 서비스를 목표로 한다.</p>
<ol>
<li>입력 파일이 실제 Windows PE 구조를 가진 파일인가?</li>
<li>유효한 PE라면 1차 정적 특징 기준에서 위험 신호가 낮은가, 의심되는가, 높은가?</li>
<li>위험 방향의 결과라면 2차 원시 바이트 모델은 어떤 악성 유형을 가장 높게 보는가?</li>
<li>두 모델이 불일치하거나 유형 점수가 낮을 때, 이를 단정하지 않고 어떻게 표현할 것인가?</li>
</ol>
<p>따라서 본 서비스는 Windows Defender·EDR·샌드박스를 대체하는 보안 제품이 아니다. 파일 실행 전 <strong>선별과 결과 해석을 돕는 교육용·보조 정적 분석 서비스</strong>로 범위를 명확히 둔다.</p>
<h3 id="12-문제-정의">1.2 문제 정의</h3>
<p>초기 문제는 “악성코드를 탐지할 수 있는 분류 모델을 만든다”였지만, 최종적으로는 아래처럼 구체화했다.</p>
<blockquote>
<p><strong>정상 파일과 악성 의심 PE 파일을 읽기 전용 정적 분석으로 우선 선별하고, 위험 파일에 한해 유형 분류와 불확실성 정보를 함께 제공한다.</strong></p>
</blockquote>
<p>이 정의에는 두 가지 제약을 포함한다.</p>
<ul>
<li>정적 분석만으로 실제 실행 행위, 감염 여부, 정보 유출, 네트워크 통신을 확정하지 않는다.</li>
<li>모델의 점수는 보정(calibration) 전 분류 점수이므로 실제 악성 확률이나 자동 차단 근거로 표현하지 않는다.</li>
</ul>
<hr>
<h2 id="2-기획-변경-및-문제-해결-과정">2. 기획 변경 및 문제 해결 과정</h2>
<h3 id="21-초기-접근법-파일-이미지화-기반-분류">2.1 초기 접근법: 파일 이미지화 기반 분류</h3>
<p>초기 기획은 실행 파일의 바이트를 Grayscale 이미지 등 2차원 이미지로 변환한 뒤, 컴퓨터 비전 분류 모델로 정상·악성 또는 악성 유형을 구분하는 방식이었다.</p>
<p>이 방식은 바이트를 <code>0~255</code>의 픽셀 값으로 대응시키고, 일정한 폭으로 줄바꿈해 이미지로 재배열한다. CNN이나 Vision Transformer는 이미지에서 질감·반복 패턴·국소 구조를 학습하므로, 원시 바이트를 즉시 이미지 모델에 넣을 수 있다는 장점이 있다.</p>
<p>초기에는 서로 다른 귀납 편향(inductive bias)을 비교하기 위해 아래 모델을 후보로 검토했다.</p>
<table>
<thead>
<tr>
<th>후보 모델</th>
<th>핵심 구조</th>
</tr>
</thead>
<tbody><tr>
<td>EfficientNetV2</td>
<td>Fused-MBConv·MBConv, 복합 스케일링</td>
</tr>
<tr>
<td>ResNet</td>
<td>Skip Connection을 가진 잔차 블록</td>
</tr>
<tr>
<td>Swin Transformer</td>
<td>Window Attention·Shifted Window</td>
</tr>
<tr>
<td>VGG16</td>
<td>작은 3×3 합성곱을 깊게 쌓는 단순 CNN</td>
</tr>
<tr>
<td>ConvNeXt</td>
<td>현대화한 CNN 블록·큰 커널·LayerNorm</td>
</tr>
</tbody></table>
<h3 id="22-초기-방식의-한계">2.2 초기 방식의 한계</h3>
<p>이미지화 접근은 모델을 빠르게 실험할 수 있지만, 파일 보안 분석 서비스의 핵심 질문에는 다음 한계가 있었다.</p>
<table>
<thead>
<tr>
<th>관찰된 한계</th>
<th>기술적 이유</th>
<th>서비스 관점의 영향</th>
</tr>
</thead>
<tbody><tr>
<td>PE 구조 의미의 약화</td>
<td>바이트를 줄 단위로 재배열하면 헤더·섹션·Import Table·Overlay 같은 파일 구조의 경계가 이미지 좌표에서 명시적으로 보존되지 않음</td>
<td>“왜 이 파일이 위험한가?”를 PE 구조와 연결해 설명하기 어려움</td>
</tr>
<tr>
<td>이미지 폭 선택에 따른 표현 변화</td>
<td>같은 바이트라도 이미지 너비·리사이즈·정규화 방식에 따라 서로 다른 시각 패턴이 생성됨</td>
<td>모델 결과가 파일 의미보다 변환 규칙에 영향을 받을 수 있음</td>
</tr>
<tr>
<td>Grad-CAM 해석의 한계</td>
<td>중요도 지도는 이미지상 영향 위치를 보여 줄 뿐, 해당 픽셀이 실제 어떤 PE 코드·데이터·행위를 뜻하는지는 바로 알 수 없음</td>
<td>악성 행위를 증명하는 근거처럼 제시할 수 없음</td>
</tr>
<tr>
<td>단일 분류 결과의 정보 부족</td>
<td>정상/악성 한 번의 결과만으로는 검토 우선순위, 유형, 불확실성을 표현하기 어려움</td>
<td>사용자가 후속 확인을 결정하기 어려움</td>
</tr>
<tr>
<td>기존 이미지 분류와의 차별성 부족</td>
<td>악성 바이트 이미지 분류는 기존 연구·예제가 많은 접근</td>
<td>프로젝트만의 파일 분석 흐름과 기술 의사결정이 약하게 보일 수 있음</td>
</tr>
</tbody></table>
<p>특히 Grad-CAM 결과는 “모델이 이미지의 어디에 반응했는가”에는 도움이 되지만, 그 위치가 실제 악성 행위나 특정 악성 코드 패턴이라는 사실을 보장하지 않는다. 이 문제는 단순히 더 높은 정확도의 이미지 모델을 선택해서 해결되지 않는다.</p>
<h3 id="23-피드백-수용과-최종-기획-전환">2.3 피드백 수용과 최종 기획 전환</h3>
<p>초기 피드백의 핵심은 다음과 같았다.</p>
<blockquote>
<p>단순히 파일을 이미지로 변환해 분류하는 방식만으로는 기존 접근과 차별성이 부족하며, 파일 구조와 원시 데이터의 의미를 더 직접적으로 반영할 필요가 있다.</p>
</blockquote>
<p>이에 따라 프로젝트는 “이미지 분류 성능 비교”에서 “PE 파일을 안전하게 선별하고, 위험 결과를 단계적으로 해석하는 서비스”로 범위를 재정의했다.</p>
<table>
<thead>
<tr>
<th>초기 방향</th>
<th>최종 방향</th>
<th>전환의 논리</th>
</tr>
</thead>
<tbody><tr>
<td>바이트 이미지를 하나의 이미지 모델에 입력</td>
<td>PE 구조 검증 후 1차·2차 계층형 추론</td>
<td>유효하지 않은 파일을 먼저 분리하고, 비용이 큰 다중 클래스 모델을 필요한 파일에만 실행</td>
</tr>
<tr>
<td>이미지 패턴 중심 분류</td>
<td>1차는 PE 정적 특징, 2차는 원시 바이트</td>
<td>구조적 신호와 원시 바이트 문맥을 역할에 맞게 분리</td>
</tr>
<tr>
<td>Normal/Malware 단일 결론</td>
<td><code>Normal</code>, <code>악성 유형 의심</code>, <code>판정 충돌</code>, <code>분류 불확실</code></td>
<td>모델 불일치와 낮은 신뢰도를 숨기지 않음</td>
</tr>
<tr>
<td>Grad-CAM 기반 이미지 중요도</td>
<td>요청 시 단계적 가림(occlusion) 재추론</td>
<td>모델 점수 변화가 큰 파일 오프셋 후보를 제시하되 실제 행위 증거로 과장하지 않음</td>
</tr>
<tr>
<td>모델 실험 중심</td>
<td>React·FastAPI 웹 서비스와 분석 흐름 중심</td>
<td>사용자가 파일 선택부터 결과 확인·방어적 안내까지 경험할 수 있게 구성</td>
</tr>
</tbody></table>
<h3 id="24-malconv2-학습-문제와-데이터-관점의-개선">2.4 MalConv2 학습 문제와 데이터 관점의 개선</h3>
<p>2차 다중 클래스 모델에서는 원시 바이트 입력의 용량과 클래스 분포가 중요한 과제였다. 원시 PE 파일은 길이가 균일하지 않고, 악성 유형별 표본 수 차이가 클 수 있어 다음 문제가 발생한다.</p>
<ul>
<li>긴 원시 바이트 입력은 GPU 메모리·I/O·학습 시간을 크게 사용한다.</li>
<li>다중 클래스 데이터가 불균형하면 빈도가 높은 유형 중심으로 학습되어 macro F1이 낮아질 수 있다.</li>
<li>악성 유형별 파일 길이·패킹 방식·수집 시점 차이가 모델이 학습하는 분포에 영향을 준다.</li>
</ul>
<p>이에 팀은 <strong>BODMAS를 2차 분류의 기준 데이터셋</strong>으로 유지했다. 현재 웹 서비스에 연결된 2차 모델 번들도 <code>config.json</code>의 <code>DATASET_NAME</code>과 <code>class_mapping.txt</code> 기준으로 BODMAS의 <code>benign</code> 1개와 악성 유형 16개, 총 17개 클래스 체계를 사용한다. 다만 BODMAS 내부에서 유형별 표본 수가 고르지 않아, 일부 클래스가 충분한 학습 신호를 갖지 못하는 문제가 있었다.</p>
<p>이를 보완하기 위해 <strong><a href="https://github.com/CS-and-AI/RawMal-TF">RawMal-TF</a></strong>를 함께 사용했다. RawMal-TF는 악성코드 타입과 패밀리 라벨이 부여된 원시 Windows PE 악성 바이너리와 사전 추출 특징 벡터를 제공하는 <strong>공개 데이터셋</strong>이다. 팀은 RawMal-TF에서 BODMAS의 부족한 악성 유형에 대응하는 샘플을 선별하고, BODMAS의 17개 클래스 매핑에 맞춰 라벨을 정렬한 뒤 학습 데이터에 보완했다. 즉, BODMAS는 최종 분류 체계와 번들의 기준이며, RawMal-TF는 클래스 불균형을 완화하기 위한 보완 데이터 소스다.</p>
<p>현재 저장된 2차 학습 실행 정보는 다음을 확인할 수 있다.</p>
<table>
<thead>
<tr>
<th>항목</th>
<th>기록된 값</th>
<th>해석</th>
</tr>
</thead>
<tbody><tr>
<td>클래스 수</td>
<td>17개</td>
<td><code>benign</code> 1개와 악성 유형 16개</td>
</tr>
<tr>
<td>데이터 크기</td>
<td>약 29.8GB</td>
<td>원시 바이트 기반 학습 데이터 규모</td>
</tr>
<tr>
<td>학습 환경</td>
<td>NVIDIA A100 80GB, bf16</td>
<td>대용량 원시 바이트 학습을 위한 GPU 환경</td>
</tr>
<tr>
<td>학습 설정</td>
<td>30 epoch, batch size 128</td>
<td>현재 저장된 학습 실행 기준</td>
</tr>
<tr>
<td>최적 epoch</td>
<td>29</td>
<td>검증 macro F1이 가장 높았던 시점</td>
</tr>
<tr>
<td>최고 검증 macro F1</td>
<td>0.8674</td>
<td>클래스 불균형을 고려하는 거시 평균 F1 기준</td>
</tr>
</tbody></table>
<hr>
<h2 id="3-시스템-아키텍처-및-파이프라인">3. 시스템 아키텍처 및 파이프라인</h2>
<h3 id="31-전체-데이터-흐름">3.1 전체 데이터 흐름</h3>
<pre><code class="language-mermaid">flowchart TD
    A[사용자: 단일 파일 또는 폴더 업로드] --&gt; B[FastAPI: 임시 사본 저장]
    B --&gt; C{PE 구조 검증\nMZ / e_lfanew / PE\0\0 / COFF}
    C -- 비 PE --&gt; U[미지원 결과]
    C -- 손상 또는 읽기 실패 --&gt; E[오류 결과]
    C -- 유효 PE --&gt; D[1차 XGBoost\n341개 PE 정적 특징]
    D -- 점수 &lt; 0.10 --&gt; N[Normal]
    D -- 0.10 이상 --&gt; F[2차 MalConv2 계열\n원시 바이트 전체]
    F --&gt; G{최고 클래스 점수}
    G -- 0.70 미만 --&gt; Q[Unknown\n분류 불확실]
    G -- benign --&gt; R[판정 충돌\n1차 위험 / 2차 정상]
    G -- 악성 유형 --&gt; S[악성 유형 의심]
    S --&gt; T[선택: 악성 유형 안내 / Gemini 조언]
    S --&gt; V[선택: 단계적 전체 가림 재추론\n64KB → 4KB → 256B]
    V --&gt; W[상위 3개 영향 후보 구간]</code></pre>
<p>분석 작업은 FastAPI의 메모리 기반 작업 관리자가 순차적으로 처리한다. 폴더 분석에서 한 파일이 비-PE·손상·추론 오류여도 다음 파일의 분석은 계속한다. 프런트엔드는 REST API로 작업을 만들고, 초기 분석 진행률은 SSE(Server-Sent Events)로 수신한다.</p>
<h3 id="32-입력-검증과-읽기-전용-정책">3.2 입력 검증과 읽기 전용 정책</h3>
<p>분석 전에는 확장자를 신뢰하지 않고 다음을 순서대로 확인한다.</p>
<ol>
<li>DOS 헤더의 <code>MZ</code> 서명</li>
<li>DOS 헤더 <code>0x3C</code> 위치의 <code>e_lfanew</code>가 가리키는 PE 헤더 오프셋 범위</li>
<li>NT 헤더의 <code>PE\0\0</code> 시그니처</li>
<li>COFF 헤더의 Machine 필드</li>
</ol>
<p>검증 결과는 다음처럼 분리한다.</p>
<table>
<thead>
<tr>
<th>결과</th>
<th>의미</th>
<th>모델 실행 여부</th>
</tr>
</thead>
<tbody><tr>
<td>유효 PE</td>
<td>최소 PE 헤더 구조가 확인됨</td>
<td>1차 모델 실행</td>
</tr>
<tr>
<td>미지원</td>
<td>MZ 서명이 없거나 PE 형식이 아님</td>
<td>실행하지 않음</td>
</tr>
<tr>
<td>오류</td>
<td>MZ는 있으나 헤더가 손상됐거나, 읽기·추론에 실패함</td>
<td>해당 파일만 중단</td>
</tr>
</tbody></table>
<h3 id="33-1차-모델-xgboost-기반-정상악성-선별">3.3 1차 모델: XGBoost 기반 정상/악성 선별</h3>
<h4 id="채택-이유">채택 이유</h4>
<p>1차는 모든 유효 PE 파일을 빠르게 통과하는 관문이므로, 구조화된 정적 특징에 강하고 추론 비용이 낮은 XGBoost를 채택했다.</p>
<ul>
<li><strong>표 형식 특징에 적합</strong>: 파일 크기·엔트로피·헤더·섹션·Import/Export·데이터 디렉터리·Overlay·바이트 히스토그램처럼 서로 성격이 다른 수치 특징을 함께 다룰 수 있다.</li>
<li><strong>비선형 상호작용 학습</strong>: 단일 특징 하나가 아니라 “섹션 수, 엔트로피, 실행 권한, Import 구조” 등의 조합에서 나타나는 위험 신호를 트리 앙상블로 학습한다.</li>
<li><strong>실무형 선별 비용에 적합</strong>: 2차 원시 바이트 딥러닝보다 가벼워, 모든 PE에 먼저 적용하고 위험 방향 결과만 다음 단계로 넘길 수 있다.</li>
<li><strong>특징 계약 검증</strong>: 실제 서비스는 학습 당시의 341개 특징 이름·순서·특징 추출기 해시를 번들에서 확인한다. 웹 서버에서 특징을 임의로 다시 구현해 잘못된 열 순서로 예측하는 위험을 줄인다.</li>
</ul>
<h4 id="341개-정적-특징-구성">341개 정적 특징 구성</h4>
<table>
<thead>
<tr>
<th>특징 그룹</th>
<th align="right">개수</th>
<th>예시</th>
</tr>
</thead>
<tbody><tr>
<td>일반 파일 통계</td>
<td align="right">5</td>
<td>파일 크기, 전체 엔트로피, 0 바이트 비율</td>
</tr>
<tr>
<td>헤더·Entry Point</td>
<td align="right">20</td>
<td>Machine, 섹션 수, Entry Point 위치, 정렬 크기</td>
</tr>
<tr>
<td>섹션 통계</td>
<td align="right">19</td>
<td>섹션 크기·엔트로피·권한·RWX 개수</td>
</tr>
<tr>
<td>Import·Export</td>
<td align="right">5</td>
<td>Import DLL 수, 심볼 수, Export 수</td>
</tr>
<tr>
<td>데이터 디렉터리</td>
<td align="right">32</td>
<td>Import·Resource·TLS·IAT 등 존재 여부와 크기</td>
</tr>
<tr>
<td>Overlay</td>
<td align="right">4</td>
<td>섹션 뒤 후행 데이터 존재·크기·엔트로피</td>
</tr>
<tr>
<td>정규화 바이트 히스토그램</td>
<td align="right">256</td>
<td>값 <code>0~255</code> 각각의 빈도</td>
</tr>
<tr>
<td><strong>합계</strong></td>
<td align="right"><strong>341</strong></td>
<td>학습·추론에서 동일한 순서 유지</td>
</tr>
</tbody></table>
<p><img src="https://velog.velcdn.com/images/j_keun/post/1256691b-b40b-427d-837f-617172202b8d/image.png" alt=""></p>
<blockquote>
<p>XGBoost는 파일의 원시 바이트 전체를 바로 신경망에 넣지 않는다. PE 헤더·섹션·Import/Export·Overlay·바이트 히스토그램에서 만든 341개 수치 특징을 입력으로 받고, 여러 결정트리가 앞선 트리의 오차를 순차적으로 보완한 뒤 위험 점수를 계산한다. 이 점수를 임계값과 비교해 <code>Normal</code>·<code>Suspicious</code>·<code>Malware</code>로 나누며, 뒤의 두 결과만 MalConv2 2차 분류로 전달한다.</p>
</blockquote>
<h4 id="1차-라우팅-정책">1차 라우팅 정책</h4>
<table>
<thead>
<tr>
<th>점수 구간</th>
<th>1차 결과</th>
<th>2차 전달</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td><code>&lt; 0.10</code></td>
<td><code>Normal</code></td>
<td>전달하지 않음</td>
<td>현재 정적 특징 기준 위험 점수가 낮음</td>
</tr>
<tr>
<td><code>0.10 이상 ~ 0.90 미만</code></td>
<td><code>Suspicious</code></td>
<td>전달</td>
<td>의심 구간이므로 유형 분류로 추가 확인</td>
</tr>
<tr>
<td><code>0.90 이상</code></td>
<td><code>Malware</code></td>
<td>전달</td>
<td>1차 모델의 위험 점수가 높은 구간</td>
</tr>
</tbody></table>
<p>임계값은 validation 분할에서 결정됐으며, 1차 점수는 확률 보정 전 값이다. 따라서 UI에서는 점수 자체보다 구간 결과와 후속 단계 여부를 중심으로 해석한다.</p>
<h3 id="34-2차-모델-malconv2-계열-악성-유형-분류">3.4 2차 모델: MalConv2 계열 악성 유형 분류</h3>
<h4 id="채택-이유-1">채택 이유</h4>
<p>2차 모델은 1차에서 <code>Suspicious</code> 또는 <code>Malware</code>로 분류된 파일만 대상으로 한다. 이 단계는 사람이 설계한 PE 정적 특징만으로 충분하지 않을 수 있는 원시 바이트 문맥을 학습해, 세부 유형을 구분하는 역할을 맡는다.</p>
<ul>
<li><strong>원시 바이트 직접 입력</strong>: 이미지를 만들거나 수작업 특징만으로 치환하지 않고 파일 바이트 <code>0~255</code>를 직접 입력한다.</li>
<li><strong>긴 파일 처리</strong>: <code>LowMemConv</code> 계열의 청크 스캔 방식을 사용한다. 현재 <code>65,536 bytes</code>는 파일을 그 길이로 자르는 최대 길이가 아니라 내부 스캔 청크 크기다.</li>
<li><strong>다중 클래스 분류</strong>: <code>benign</code>과 16개 악성 유형, 총 17개 클래스의 softmax 점수를 계산한다.</li>
<li><strong>모델 구조 검증</strong>: 가중치·클래스 매핑·설정의 일치를 확인한 뒤 추론한다. 현재 체크포인트에 <code>context_net</code> 가중치가 있으면 Global Context 계열 MalConv 구조를 선택해 strict loading을 수행한다.</li>
</ul>
<p><img src="https://velog.velcdn.com/images/j_keun/post/e6cb17bf-d52b-4056-82f6-287397010709/image.png" alt=""></p>
<blockquote>
<p>MalConv2-GCG는 파일을 이미지로 바꾸지 않고 원시 바이트를 토큰으로 입력한다. Context Path는 파일 전체의 문맥을 요약하고, Feature Path는 1D Convolution으로 지역 바이트 패턴을 추출한다. 전역 문맥 게이트는 유형 분류에 덜 중요한 채널을 낮추고 중요한 패턴을 상대적으로 강조한다. 이후 시간 축 최대 풀링과 완전연결 계층을 거쳐 17개 클래스 점수를 계산한다.</p>
</blockquote>
<h4 id="원시-바이트-전처리와-내부-구조">원시 바이트 전처리와 내부 구조</h4>
<pre><code class="language-text">파일 원시 바이트 0~255
        ↓
토큰 1~256으로 이동 (0은 Padding 전용으로 예약)
        ↓
Byte Embedding
        ↓
병렬 1D Convolution / Gating 또는 Context 경로
        ↓
Temporal Max Pooling
        ↓
Fully Connected Layer
        ↓
17개 클래스 Softmax</code></pre>
<p>현재 서비스는 원본 파일 전체 바이트를 보존해 모델 어댑터에 전달한다. 바이트 값에 <code>1</code>을 더해 <code>1~256</code> 토큰으로 바꾸고, <code>0</code>은 모델의 padding에만 사용한다. 이 처리는 실제 <code>0x00</code> 바이트와 패딩이 혼동되는 문제를 방지한다.</p>
<h4 id="2차-결과와-최종-상태-정책">2차 결과와 최종 상태 정책</h4>
<table>
<thead>
<tr>
<th>1차 결과</th>
<th>2차 최고 클래스·점수</th>
<th>최종 표시</th>
<th>해석</th>
</tr>
</thead>
<tbody><tr>
<td><code>Normal</code></td>
<td>2차 미실행</td>
<td><code>Normal</code></td>
<td>1차 기준 위험도가 낮음</td>
</tr>
<tr>
<td><code>Suspicious</code>/<code>Malware</code></td>
<td>악성 유형, 점수 <code>0.70 이상</code></td>
<td><code>악성 유형 의심</code></td>
<td>두 단계가 위험 방향으로 이어짐</td>
</tr>
<tr>
<td><code>Suspicious</code>/<code>Malware</code></td>
<td><code>benign</code>, 점수 <code>0.70 이상</code></td>
<td><code>판정 충돌</code></td>
<td>1차 위험 결과와 2차 정상 최고 클래스가 불일치</td>
</tr>
<tr>
<td><code>Suspicious</code>/<code>Malware</code></td>
<td>최고 점수 <code>&lt; 0.70</code></td>
<td><code>분류 불확실</code></td>
<td>유형을 특정할 근거가 부족하여 <code>Unknown</code> 처리</td>
</tr>
</tbody></table>
<p>2차 <code>0.70</code> 임계값은 현재 검증 데이터 기반 확률 보정 전의 Unknown 처리 기준이다. 이 값은 설정 파일로 분리되어 있어 향후 calibration 또는 재평가 결과에 따라 API를 바꾸지 않고 조정할 수 있다.</p>
<h3 id="35-결과-해석-보조-기능">3.5 결과 해석 보조 기능</h3>
<p>악성 유형 의심 결과에는 아래 보조 기능을 제공한다.</p>
<table>
<thead>
<tr>
<th>기능</th>
<th>입력</th>
<th>동작</th>
<th>제한</th>
</tr>
</thead>
<tbody><tr>
<td>악성 유형 안내</td>
<td>2차 유형·최종 상태</td>
<td>서버가 유형별 일반 방어 안내를 반환</td>
<td>실제 감염·행위를 확정하지 않음</td>
</tr>
<tr>
<td>Gemini AI 전문가 조언</td>
<td>1·2차 결과와 사용자가 고른 제한된 상황 메타데이터</td>
<td>사용자가 요청할 때만 방어적 조언 생성</td>
<td>원시 바이트·파일명·경로·문자열을 외부 LLM에 보내지 않음</td>
</tr>
<tr>
<td>단계적 영향 구간 분석</td>
<td>악성 유형 의심 파일과 2차 기준 점수</td>
<td>64KB → 4KB → 256B 단위로 가리고 재추론해 상위 3개 후보 반환</td>
<td>영향 위치는 모델 점수 변화 후보일 뿐, 실제 악성 행위 증거가 아님</td>
</tr>
</tbody></table>
<h3 id="36-fastapi-백엔드-구현-구조">3.6 FastAPI 백엔드 구현 구조</h3>
<p>백엔드는 React 화면에서 업로드한 단일 파일 또는 폴더 파일 목록을 받아, PE 검증과 두 단계 모델 추론을 순서대로 수행한다. 분석 시간이 파일마다 다를 수 있으므로 업로드 요청을 받자마자 분석 작업 ID를 반환하고, 이후 진행 상태와 결과를 작업 단위로 조회하는 방식으로 구성했다.</p>
<pre><code class="language-text">React 브라우저
  │ 파일·폴더 업로드 / 작업 상태·진행률 요청
  ▼
FastAPI Router
  │ /api/v1/analyses/file, /folder
  ▼
AnalysisManager
  ├─ PE Validator
  ├─ 1차 XGBoost
  ├─ 필요 시 2차 MalConv2-GCT
  └─ 최종 상태 조합
  ▼
작업 결과 반환 → React 결과 화면</code></pre>
<h4 id="앱-시작점과-api-연결">앱 시작점과 API 연결</h4>
<p><code>main.py</code>는 FastAPI 애플리케이션을 만들고, React 개발 서버의 요청을 위한 CORS와 분석 라우터를 등록한다. <code>lifespan</code>에는 프로세스 전체가 공유할 분석 관리자를 연결해 요청마다 모델 관리 객체가 새로 만들어지지 않도록 했다.</p>
<pre><code class="language-python">app = FastAPI(
    title=resolved_settings.title,
    version=resolved_settings.version,
    lifespan=create_analysis_lifespan(manager_factory),
)

app.add_middleware(CORSMiddleware, ...)
app.include_router(analyze_router, prefix=&quot;/api/v1&quot;)</code></pre>
<h4 id="분석-작업-생성-api">분석 작업 생성 API</h4>
<p>단일 파일 분석 API는 파일을 받은 뒤 전체 추론이 끝날 때까지 HTTP 연결을 붙잡지 않는다. 대신 작업을 생성해 <code>202 Accepted</code>와 작업 정보를 반환한다. 폴더 분석도 같은 방식으로 여러 파일과 상대 경로를 작업 대기열에 넣는다.</p>
<table>
<thead>
<tr>
<th>엔드포인트</th>
<th>역할</th>
<th>응답</th>
</tr>
</thead>
<tbody><tr>
<td><code>POST /api/v1/analyses/file</code></td>
<td>단일 파일 분석 작업 생성</td>
<td>작업 ID와 초기 진행 상태, <code>202 Accepted</code></td>
</tr>
<tr>
<td><code>POST /api/v1/analyses/folder</code></td>
<td>폴더의 다중 파일 분석 작업 생성</td>
<td>작업 ID와 초기 진행 상태, <code>202 Accepted</code></td>
</tr>
<tr>
<td><code>GET /api/v1/analyses/{job_id}</code></td>
<td>현재 진행 상태와 결과 조회</td>
<td>파일별 결과·집계·오류 상태</td>
</tr>
<tr>
<td><code>GET /api/v1/analyses/{job_id}/events</code></td>
<td>진행 상태 변경 이벤트 수신</td>
<td>SSE(Server-Sent Events) 스트림</td>
</tr>
</tbody></table>
<pre><code class="language-python">@router.post(&quot;/file&quot;, response_model=AnalysisJob,
             status_code=status.HTTP_202_ACCEPTED)
async def create_single_file_analysis(
    manager: AnalysisManagerDependency,
    file: UploadFile = File(...),
) -&gt; AnalysisJob:
    return await manager.create_job(
        input_kind=&quot;file&quot;,
        uploads=[file],
        relative_paths=[file.filename or &quot;uploaded-file&quot;],
    )</code></pre>
<h4 id="읽기-전용-pe-검증">읽기 전용 PE 검증</h4>
<p>모델 호출 이전에 <code>MZ</code> 서명, DOS 헤더의 PE 오프셋(<code>e_lfanew</code>), <code>PE\0\0</code> 시그니처를 확인한다. 따라서 확장자만 <code>.exe</code>로 변경한 문서 파일은 모델까지 전달되지 않는다.</p>
<pre><code class="language-python">if file_bytes[:2] != b&quot;MZ&quot;:
    return PeValidationResult(is_valid=False)

pe_offset = int.from_bytes(file_bytes[0x3C:0x40], &quot;little&quot;)

if file_bytes[pe_offset : pe_offset + 4] != b&quot;PE\x00\x00&quot;:
    return PeValidationResult(is_valid=False)</code></pre>
<h4 id="1차·2차-모델-연결과-최종-상태-결정">1차·2차 모델 연결과 최종 상태 결정</h4>
<p>유효한 PE 파일은 먼저 1차 XGBoost로 분석한다. 1차 결과가 <code>Normal</code>이면 종료하고, <code>Suspicious</code> 또는 <code>Malware</code>이면 원시 바이트 시퀀스를 준비해 2차 MalConv2-GCT로 전달한다. 이후 두 모델 결과를 조합해 <code>Normal</code>, <code>악성 유형 의심</code>, <code>판정 충돌</code>, <code>분류 불확실</code> 중 하나의 최종 상태로 반환한다.</p>
<pre><code class="language-python">validation = validate_pe(file_bytes)
stage1 = self.stage1.analyze_file(file_path)

if stage1.needs_stage2:
    sequence = prepare_malconv2_byte_sequence(file_bytes)
    stage2 = self._get_stage2()
    family_class, family_confidence, is_unknown = \
        stage2.predict_stage2(sequence)

final_status = self._final_status(
    stage1.result, family_class, is_unknown
)</code></pre>
<p>이 흐름은 <a href="../backend/app/main.py">main.py</a>, <a href="../backend/app/routers/analyze.py">analyze.py</a>, <a href="../backend/app/services/pe_validator.py">pe_validator.py</a>, <a href="../backend/app/services/analysis_service.py">analysis_service.py</a>를 기준으로 정리했다.</p>
<hr>
<h2 id="4-모델-분석-및-시각화-가이드">4. 모델 분석 및 시각화 가이드</h2>
<h3 id="41-모델별-핵심-구조·장단점-비교">4.1 모델별 핵심 구조·장단점 비교</h3>
<table>
<thead>
<tr>
<th>모델</th>
<th>핵심 구조</th>
<th>강점</th>
<th>한계</th>
<th>프로젝트에서의 위치</th>
</tr>
</thead>
<tbody><tr>
<td>EfficientNetV2</td>
<td>Fused-MBConv·MBConv과 효율적 스케일링</td>
<td>정확도·속도 균형, 이미지 분류에 효율적</td>
<td>이미지 변환 규칙에 따라 입력 의미가 변할 수 있음</td>
<td>초기 이미지 기반 후보</td>
</tr>
<tr>
<td>ResNet</td>
<td>잔차 연결(Residual Connection)</td>
<td>깊은 네트워크 학습 안정성, 풍부한 기준선</td>
<td>바이트 이미지의 픽셀 중요도를 PE 의미로 바로 해석하기 어려움</td>
<td>초기 이미지 기반 후보</td>
</tr>
<tr>
<td>Swin Transformer</td>
<td>Shifted Window Self-Attention</td>
<td>지역 창과 창 간 문맥을 계층적으로 학습</td>
<td>데이터·GPU 비용이 크며 이미지화 한계는 남음</td>
<td>초기 이미지 기반 후보</td>
</tr>
<tr>
<td>VGG16</td>
<td>반복되는 3×3 Convolution</td>
<td>구조가 직관적이고 비교 기준으로 좋음</td>
<td>파라미터·연산량이 크고 현대 모델 대비 효율이 낮음</td>
<td>초기 이미지 기반 후보</td>
</tr>
<tr>
<td>ConvNeXt</td>
<td>현대화한 CNN 블록, 큰 커널·LayerNorm</td>
<td>CNN 기반에서 강한 표현력</td>
<td>여전히 2D 이미지 표현에 의존</td>
<td>초기 이미지 기반 후보</td>
</tr>
<tr>
<td>XGBoost</td>
<td>Gradient Boosted Decision Trees</td>
<td>341개 구조화 특징의 비선형 조합 학습, 빠른 선별, 특징 계약 검증</td>
<td>사람이 정의한 특징 범위 밖의 원시 바이트 문맥은 제한적</td>
<td><strong>최종 1차 기본 모델</strong></td>
</tr>
<tr>
<td>MalConv2 계열</td>
<td>Byte Embedding·1D Conv·Gating/Context·Temporal Max Pooling</td>
<td>긴 원시 바이트를 직접 보고 악성 유형 다중 분류</td>
<td>학습 비용이 크고 클래스 불균형·분포 변화에 민감</td>
<td><strong>최종 2차 패밀리 모델</strong></td>
</tr>
</tbody></table>
<hr>
<h2 id="5-결론-및-회고">5. 결론 및 회고</h2>
<h3 id="51-프로젝트를-통해-얻은-기술적-인사이트">5.1 프로젝트를 통해 얻은 기술적 인사이트</h3>
<ol>
<li><p><strong>입력 표현은 모델 선택만큼 중요하다.</strong></p>
<ul>
<li>이미지 모델의 성능을 비교하는 것보다, 바이트를 왜 이미지로 바꾸는지와 그 과정에서 어떤 의미가 사라지는지를 먼저 검토해야 했다.</li>
</ul>
</li>
<li><p><strong>악성코드 분석은 단일 정확도보다 역할 분리가 중요하다.</strong></p>
<ul>
<li>1차는 빠르게 위험 파일을 선별하고, 2차는 비용이 큰 원시 바이트 다중 클래스 분류를 담당한다. 두 모델의 목적이 다르므로 동일한 지표만으로 선택하지 않았다.</li>
</ul>
</li>
<li><p><strong>불확실성을 결과로 설계해야 한다.</strong></p>
<ul>
<li>2차 최고 클래스가 <code>benign</code>인 경우를 단순 정상으로 덮지 않고 <code>판정 충돌</code>로, 최고 점수가 낮은 경우를 <code>Unknown</code>·<code>분류 불확실</code>로 표현했다. 이는 모델 결과를 과신하지 않기 위한 제품 정책이다.</li>
</ul>
</li>
<li><p><strong>설명 가능성은 곧 행위 증명이 아니다.</strong></p>
<ul>
<li>Grad-CAM이나 단계적 가림 분석은 모델이 반응한 위치를 보조적으로 보여 줄 수 있다. 그러나 해당 위치가 실제 감염 사실이나 악성 행위를 증명한다고 해석해서는 안 된다.</li>
</ul>
</li>
<li><p><strong>운영 가능한 ML 서비스에는 모델 외 계층이 필요하다.</strong></p>
<ul>
<li>PE 검증, 업로드 임시 파일 관리, 오류 분리, 모델 번들 무결성 검사, 진행률, API 계약, UI 상태 표현, 외부 LLM 전송 최소화가 모델 추론만큼 중요했다.</li>
</ul>
</li>
</ol>
<h3 id="52-향후-발전-방향">5.2 향후 발전 방향</h3>
<table>
<thead>
<tr>
<th>개선 방향</th>
<th>기대 효과</th>
</tr>
</thead>
<tbody><tr>
<td>동적 분석 샌드박스 연계</td>
<td>프로세스·파일·레지스트리·네트워크 행위를 관찰해 정적 분석 한계 보완</td>
</tr>
<tr>
<td>점수 보정과 임계값 재평가</td>
<td>점수를 현실 확률처럼 오해하지 않도록 calibration·운영 비용 기반 임계값 정책 도입</td>
</tr>
<tr>
<td>시간 분리·외부 데이터 평가</td>
<td>특정 데이터셋에 맞춘 과적합과 데이터 분포 변화를 더 엄격히 검증</td>
</tr>
<tr>
<td>클래스 불균형 대응 고도화</td>
<td>재표본화, class weight, focal loss, 유형별 F1·혼동행렬 분석으로 소수 클래스 성능 개선</td>
</tr>
<tr>
<td>모델 버전·실험 추적 체계화</td>
<td>데이터셋 버전, 특징 계약, 가중치 해시, 임계값, 평가 리포트를 한 실행 단위로 보존</td>
</tr>
<tr>
<td>영향 구간 분석 최적화</td>
<td>재추론 횟수와 세분화 기준을 모델·파일 크기별로 조절해 응답 시간 개선</td>
</tr>
<tr>
<td>결과 이력의 영속 저장</td>
<td>사용자 동의·보존 정책을 전제로 분석 이력과 모델 버전을 감사 가능하게 관리</td>
</tr>
</tbody></table>
<h3 id="53-최종-회고">5.3 최종 회고</h3>
<p>이 프로젝트의 핵심 성과는 “이미지 분류 모델 하나를 학습했다”는 데 있지 않다. 피드백을 통해 초기 접근의 한계를 인정하고, <strong>PE 구조 검증 → 1차 선별 → 2차 원시 바이트 분류 → 불확실성 표현 → 방어적 안내</strong>라는 전체 분석 경험으로 문제를 재정의했다는 데 있다.</p>
<p>향후에는 동적 분석과 외부 보안 도구 검증을 결합해 실제 악성 행위 판단에 가까워질 수 있다. 그러나 현재 단계에서도 모델 결과를 과장하지 않고, 어떤 파일을 먼저 검토해야 하는지와 무엇을 추가 확인해야 하는지를 제공하는 정적 분석 보조 서비스로서 명확한 역할을 가진다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[68일차] 자연어 처리 기초와 토큰화]]></title>
            <link>https://velog.io/@j_keun/68%EC%9D%BC%EC%B0%A8-%EC%9E%90%EC%97%B0%EC%96%B4-%EC%B2%98%EB%A6%AC-%EA%B8%B0%EC%B4%88%EC%99%80-%ED%86%A0%ED%81%B0%ED%99%94-6va8r19g</link>
            <guid>https://velog.io/@j_keun/68%EC%9D%BC%EC%B0%A8-%EC%9E%90%EC%97%B0%EC%96%B4-%EC%B2%98%EB%A6%AC-%EA%B8%B0%EC%B4%88%EC%99%80-%ED%86%A0%ED%81%B0%ED%99%94-6va8r19g</guid>
            <pubDate>Sun, 27 Sep 2026 04:29:35 GMT</pubDate>
            <description><![CDATA[<p>자연어 처리(Natural Language Processing, NLP)는 컴퓨터가 사람이 사용하는 언어를 분석하고 이해하며 생성할 수 있게 만드는 인공지능 분야다. 컴퓨터는 문장을 그대로 이해하지 못하므로, 텍스트를 토큰(Token)으로 나누고 숫자 ID로 바꾸는 전처리 과정을 거친다.</p>
<p>이번에는 자연어 처리의 기본 단위인 Corpus, Document, Token, Vocabulary를 살펴보고, 직접 만든 사전과 Hugging Face의 <code>klue/bert-base</code> 토크나이저를 비교한다.</p>
<hr>
<h2 id="1-자연어와-자연어-처리">1. 자연어와 자연어 처리</h2>
<p>자연어는 사람이 일상에서 말하고 쓰는 언어다. 문법과 어휘뿐 아니라 문맥, 표현 방식, 뉘앙스처럼 규칙만으로 처리하기 어려운 요소를 포함한다.</p>
<p>자연어 처리는 텍스트 또는 음성 데이터를 모델이 다룰 수 있는 형태로 바꾸고, 그 안에서 의미 있는 결과를 찾거나 새로운 문장을 만드는 기술이다.</p>
<table>
<thead>
<tr>
<th>활용 분야</th>
<th>설명</th>
</tr>
</thead>
<tbody><tr>
<td>텍스트 분류</td>
<td>문서를 뉴스, 스포츠, 경제처럼 범주로 나눈다.</td>
</tr>
<tr>
<td>감정 분석</td>
<td>리뷰나 댓글의 긍정·부정 성향을 분석한다.</td>
</tr>
<tr>
<td>문서 요약</td>
<td>긴 문서에서 핵심 내용을 짧게 정리한다.</td>
</tr>
<tr>
<td>기계 번역</td>
<td>한 언어의 문장을 다른 언어로 변환한다.</td>
</tr>
<tr>
<td>질의응답</td>
<td>질문에 맞는 답변을 찾거나 생성한다.</td>
</tr>
<tr>
<td>정보 검색</td>
<td>사용자의 검색어와 관련된 문서를 찾는다.</td>
</tr>
<tr>
<td>챗봇·문장 생성</td>
<td>사용자 입력에 맞춰 자연스러운 문장을 생성한다.</td>
</tr>
<tr>
<td>RAG·LLM Agent</td>
<td>외부 문서를 검색하거나 도구를 사용해 답변을 보완한다.</td>
</tr>
</tbody></table>
<p>최근에는 텍스트뿐 아니라 이미지, 음성, 영상과 결합하는 멀티모달 AI로도 확장되고 있다.</p>
<hr>
<h2 id="2-corpus와-document">2. Corpus와 Document</h2>
<h3 id="corpus">Corpus</h3>
<p>Corpus(말뭉치)는 자연어 처리 모델을 분석하거나 학습하기 위해 모은 전체 텍스트 데이터 집합이다.</p>
<p>예를 들어 영화 리뷰 10만 개를 모았다면, 리뷰 10만 개 전체가 하나의 Corpus가 된다.</p>
<h3 id="document">Document</h3>
<p>Document는 Corpus를 구성하는 개별 텍스트 하나다. 반드시 긴 글일 필요는 없다.</p>
<pre><code class="language-text">Corpus: 영화 리뷰 10만 개
 ├── Document 1: &quot;재미있고 배우들의 연기가 좋다.&quot;
 ├── Document 2: &quot;전개가 느려서 아쉬웠다.&quot;
 └── Document 3: &quot;다시 보고 싶은 영화다.&quot;</code></pre>
<p>뉴스 기사 하나, 리뷰 하나, 댓글 하나, 짧은 문장 하나도 모두 Document가 될 수 있다.</p>
<hr>
<h2 id="3-tokenization">3. Tokenization</h2>
<p>Tokenization(토큰화)은 텍스트를 모델이 처리할 수 있는 작은 단위인 Token으로 나누는 과정이다.</p>
<pre><code class="language-text">&quot;자연어 처리를 공부합니다.&quot;
            ↓ Tokenization
[&quot;자연&quot;, &quot;##어&quot;, &quot;처리&quot;, &quot;##를&quot;, &quot;공부&quot;, &quot;##합니다&quot;, &quot;.&quot;]
            ↓ Token ID 변환
[3941, 2051, 4211, 2138, 4244, 11800, 18]
            ↓ Tensor 변환
tensor([3941, 2051, 4211, 2138, 4244, 11800, 18])</code></pre>
<p>토큰은 항상 단어 하나와 같지 않다. 단어 전체일 수도 있고, 단어의 일부, 구두점, 공백과 관련된 단위일 수도 있다.</p>
<h3 id="단순-공백-기준-토큰화">단순 공백 기준 토큰화</h3>
<p>가장 간단한 방법은 공백을 기준으로 문장을 나누는 것이다.</p>
<pre><code class="language-python">text = &quot;나는 자연어를 공부한다&quot;
tokens = text.split()

print(tokens)</code></pre>
<pre><code class="language-text">[&#39;나는&#39;, &#39;자연어를&#39;, &#39;공부한다&#39;]</code></pre>
<p>이 방식은 이해하기 쉽지만, 한국어의 조사나 어미를 분리하지 못하고 구두점 처리도 세밀하지 않다. 예를 들어 <code>자연어를</code>은 하나의 Token으로 남고, 처음 보는 단어는 사전에 없을 가능성이 커진다.</p>
<hr>
<h2 id="4-vocabulary와-token-id">4. Vocabulary와 Token ID</h2>
<p>Vocabulary는 토크나이저가 사용할 수 있는 Token의 집합이다. 모델은 문자열 자체가 아니라 Vocabulary에 등록된 정수 ID를 입력으로 받는다.</p>
<pre><code class="language-python">vocab = {
    &quot;&lt;PAD&gt;&quot;: 0,
    &quot;&lt;UNK&gt;&quot;: 1,
    &quot;나는&quot;: 2,
    &quot;자연어를&quot;: 3,
    &quot;공부한다&quot;: 4,
}

text = &quot;나는 자연어를 공부한다&quot;
tokens = text.split()

token_ids = [vocab.get(token, vocab[&quot;&lt;UNK&gt;&quot;]) for token in tokens]
input_ids = torch.tensor(token_ids, dtype=torch.long)

print(&quot;원문:&quot;, text)
print(&quot;Token:&quot;, tokens)
print(&quot;Token ID:&quot;, token_ids)
print(&quot;PyTorch Tensor:&quot;, input_ids)
print(&quot;dtype:&quot;, input_ids.dtype)</code></pre>
<p>실행 결과는 다음과 같다.</p>
<pre><code class="language-text">원문: 나는 자연어를 공부한다
Token: [&#39;나는&#39;, &#39;자연어를&#39;, &#39;공부한다&#39;]
Token ID: [2, 3, 4]
PyTorch Tensor: tensor([2, 3, 4])
dtype: torch.int64</code></pre>
<p><code>2</code>, <code>3</code>, <code>4</code>라는 숫자 자체에는 언어적 의미가 없다. 이 숫자는 Vocabulary 안에서 특정 Token을 찾기 위한 식별자다. 이후 모델은 이 ID를 Embedding 층에 전달해 학습 가능한 벡터로 바꾼다.</p>
<h3 id="special-token">Special Token</h3>
<table>
<thead>
<tr>
<th>Token</th>
<th>의미</th>
<th>용도</th>
</tr>
</thead>
<tbody><tr>
<td><code>&lt;PAD&gt;</code></td>
<td>Padding Token</td>
<td>문장 길이를 맞출 때 빈자리를 채운다.</td>
</tr>
<tr>
<td><code>&lt;UNK&gt;</code></td>
<td>Unknown Token</td>
<td>Vocabulary에 없는 Token을 나타낸다.</td>
</tr>
<tr>
<td><code>[CLS]</code></td>
<td>Classification Token</td>
<td>BERT 계열에서 문장 전체를 대표하는 특별한 시작 Token이다.</td>
</tr>
<tr>
<td><code>[SEP]</code></td>
<td>Separator Token</td>
<td>문장 또는 문장 쌍의 경계를 구분한다.</td>
</tr>
</tbody></table>
<p>모델은 보통 한 번에 여러 문장을 Batch로 처리한다. 문장 길이가 다르면 Tensor 모양을 맞추기 어려우므로 짧은 문장 뒤에 <code>&lt;PAD&gt;</code>를 추가한다.</p>
<pre><code class="language-text">문장 A: [2, 3, 4]
문장 B: [2, 4]

Padding 후
문장 A: [2, 3, 4]
문장 B: [2, 4, 0]  # 0은 &lt;PAD&gt;</code></pre>
<hr>
<h2 id="5-oov-문제와-unk">5. OOV 문제와 <code>&lt;UNK&gt;</code></h2>
<p>OOV(Out-Of-Vocabulary)는 입력된 Token이 Vocabulary에 존재하지 않는 상황이다.</p>
<pre><code class="language-python">vocab = {
    &quot;&lt;PAD&gt;&quot;: 0,
    &quot;&lt;UNK&gt;&quot;: 1,
    &quot;나는&quot;: 2,
    &quot;파이토치를&quot;: 3,
    &quot;공부한다&quot;: 4,
}

tokens = [&quot;나는&quot;, &quot;트랜스포머를&quot;, &quot;공부한다&quot;]
token_ids = [vocab.get(token, vocab[&quot;&lt;UNK&gt;&quot;]) for token in tokens]

print(token_ids)</code></pre>
<pre><code class="language-text">[2, 1, 4]</code></pre>
<p><code>트랜스포머를</code>은 Vocabulary에 없으므로 <code>&lt;UNK&gt;</code>의 ID인 <code>1</code>로 바뀐다. 단어 단위 사전만 사용하면 새로운 단어, 오타, 활용형을 자주 <code>&lt;UNK&gt;</code>로 처리하게 되어 정보가 사라질 수 있다.</p>
<blockquote>
<p>노트북 원본은 <code>tokens</code>를 중괄호 <code>{}</code>로 만든 <code>set</code>으로 작성했다. <code>set</code>은 순서를 보장하지 않으므로, 문장의 Token 순서를 유지해야 할 때는 위 예제처럼 대괄호 <code>[]</code>를 사용하는 <code>list</code>가 알맞다.</p>
</blockquote>
<hr>
<h2 id="6-subword-tokenization">6. Subword Tokenization</h2>
<p>현대 NLP 모델은 OOV 문제를 줄이기 위해 Subword Tokenization을 많이 사용한다. 단어 전체를 하나로만 처리하지 않고, 자주 등장하는 부분 단위로 나눈다.</p>
<pre><code class="language-text">&quot;자연어 처리를 공부합니다.&quot;
        ↓
[&#39;자연&#39;, &#39;##어&#39;, &#39;처리&#39;, &#39;##를&#39;, &#39;공부&#39;, &#39;##합니다&#39;, &#39;.&#39;]</code></pre>
<p><code>##</code>는 BERT 계열 WordPiece 토크나이저에서 앞 Token에 이어 붙는 부분이라는 표시다. 예를 들어 <code>자연</code>과 <code>##어</code>를 합치면 <code>자연어</code>가 되고, <code>공부</code>와 <code>##합니다</code>를 합치면 <code>공부합니다</code>가 된다.</p>
<p>처음 보는 단어라도 이미 알고 있는 부분 Token의 조합으로 표현할 가능성이 높다. 이 때문에 단어 전체만 사용하는 방식보다 <code>&lt;UNK&gt;</code> 발생을 줄일 수 있다.</p>
<hr>
<h2 id="7-hugging-face-autotokenizer-사용하기">7. Hugging Face <code>AutoTokenizer</code> 사용하기</h2>
<p>Hugging Face는 자연어 처리뿐 아니라 이미지, 음성, 멀티모달 모델과 관련 도구를 제공하는 생태계다. <code>AutoTokenizer</code>를 사용하면 모델 저장소에 맞는 Tokenizer 설정과 Vocabulary를 불러올 수 있다.</p>
<pre><code class="language-python">from transformers import AutoTokenizer

# Hugging Face Hub의 klue/bert-base 토크나이저 불러오기
tokenizer = AutoTokenizer.from_pretrained(&quot;klue/bert-base&quot;)

text = &quot;자연어 처리를 공부합니다.&quot;
tokens = tokenizer.tokenize(text)
token_ids = tokenizer.convert_tokens_to_ids(tokens)

print(&quot;Token:&quot;, tokens)
print(&quot;Token ID:&quot;, token_ids)</code></pre>
<p>실행 결과는 다음과 같다.</p>
<pre><code class="language-text">Token: [&#39;자연&#39;, &#39;##어&#39;, &#39;처리&#39;, &#39;##를&#39;, &#39;공부&#39;, &#39;##합니다&#39;, &#39;.&#39;]
Token ID: [3941, 2051, 4211, 2138, 4244, 11800, 18]</code></pre>
<p><code>klue/bert-base</code>의 Vocabulary 크기와 주요 Special Token ID는 다음과 같이 확인했다.</p>
<pre><code class="language-python">print(&quot;Vocabulary 크기:&quot;, tokenizer.vocab_size)
print(&quot;[PAD] ID:&quot;, tokenizer.pad_token_id)
print(&quot;[UNK] ID:&quot;, tokenizer.unk_token_id)
print(&quot;[CLS] ID:&quot;, tokenizer.cls_token_id)
print(&quot;[SEP] ID:&quot;, tokenizer.sep_token_id)</code></pre>
<pre><code class="language-text">Vocabulary 크기: 32000
[PAD] ID: 0
[UNK] ID: 1
[CLS] ID: 2
[SEP] ID: 3</code></pre>
<blockquote>
<p>노트북에서는 <code>[PAD] ID</code>를 출력할 때 <code>tokenizer.pad_token_type_id</code>를 사용했다. 이 값은 이 모델에서 우연히 <code>0</code>으로 출력되지만, PAD Token의 Vocabulary ID를 확인하려면 의미에 맞는 <code>tokenizer.pad_token_id</code>를 사용해야 한다.</p>
</blockquote>
<h3 id="tokenize와-모델-입력-만들기의-차이"><code>tokenize()</code>와 모델 입력 만들기의 차이</h3>
<p><code>tokenizer.tokenize()</code>는 사람이 확인하기 좋은 Token 문자열 목록을 반환한다. 실제 모델 입력을 만들 때는 보통 토큰화와 ID 변환, Special Token 추가, Padding, Attention Mask 생성을 한 번에 처리한다.</p>
<pre><code class="language-python">encoded = tokenizer(
    &quot;자연어 처리를 공부합니다.&quot;,
    return_tensors=&quot;pt&quot;,
)

print(encoded[&quot;input_ids&quot;])
print(encoded[&quot;attention_mask&quot;])</code></pre>
<p><code>input_ids</code>는 모델에 전달할 Token ID이고, <code>attention_mask</code>는 실제 Token과 Padding 위치를 구분하는 값이다. 이처럼 모델용 입력을 만들 때는 <code>tokenizer(...)</code> 호출이 더 편리하다.</p>
<hr>
<h2 id="8-macos와-windows-환경-차이">8. macOS와 Windows 환경 차이</h2>
<p>토큰화와 <code>transformers</code> API는 운영체제와 관계없이 거의 동일하다. 하지만 가상환경 활성화 명령어, 경로 표기, 사용할 수 있는 GPU 가속 방식은 macOS와 Windows에서 다르다.</p>
<h3 id="81-한눈에-비교하기">8.1 한눈에 비교하기</h3>
<table>
<thead>
<tr>
<th>항목</th>
<th>macOS</th>
<th>Windows</th>
</tr>
</thead>
<tbody><tr>
<td>기본 터미널 예시</td>
<td>Terminal, iTerm2, VS Code Terminal</td>
<td>PowerShell, 명령 프롬프트, VS Code Terminal</td>
</tr>
<tr>
<td>가상환경 활성화</td>
<td><code>source .venv/bin/activate</code></td>
<td>PowerShell: <code>.\.venv\Scripts\Activate.ps1</code><br>명령 프롬프트: <code>.venv\Scripts\activate.bat</code></td>
</tr>
<tr>
<td>경로 구분자</td>
<td><code>/</code></td>
<td><code>\</code>를 주로 사용하지만 <code>/</code>도 대부분의 Python 코드에서 사용할 수 있다.</td>
</tr>
<tr>
<td>대표 GPU 가속</td>
<td>Apple Silicon의 <code>mps</code></td>
<td>NVIDIA GPU의 <code>cuda</code></td>
</tr>
<tr>
<td>GPU 확인</td>
<td><code>torch.backends.mps.is_available()</code></td>
<td><code>torch.cuda.is_available()</code></td>
</tr>
<tr>
<td>Hugging Face 기본 캐시</td>
<td><code>~/.cache/huggingface/hub</code></td>
<td><code>C:\Users\&lt;사용자명&gt;\.cache\huggingface\hub</code></td>
</tr>
</tbody></table>
<p>경로를 문자열로 이어 붙이기보다 <code>pathlib.Path</code>를 사용하면 운영체제마다 다른 구분자를 신경 쓰지 않아도 된다.</p>
<pre><code class="language-python">from pathlib import Path

model_dir = Path(&quot;models&quot;) / &quot;klue-bert&quot;
print(model_dir)</code></pre>
<h3 id="82-공통-준비">8.2 공통 준비</h3>
<p><code>transformers</code>와 PyTorch는 프로젝트마다 독립된 가상환경에 설치하는 편이 좋다. 프로젝트마다 라이브러리 버전이 달라도 서로 영향을 주지 않는다.</p>
<pre><code class="language-text">프로젝트 폴더
├── .venv/          # 현재 프로젝트 전용 Python 환경
├── notebook.ipynb
└── requirements.txt</code></pre>
<p>Hugging Face의 <code>from_pretrained()</code>는 처음 실행할 때 모델 또는 토크나이저 파일을 Hub에서 내려받아 로컬 캐시에 저장한다. 이후에는 같은 버전이 캐시에 있으면 로컬 파일을 사용한다.</p>
<h3 id="83-macos-apple-silicon과-mps">8.3 macOS: Apple Silicon과 MPS</h3>
<p>Apple Silicon Mac에서는 PyTorch의 MPS(Metal Performance Shaders) 백엔드를 사용할 수 있다. 노트북의 실행 결과도 <code>mps</code>였다.</p>
<pre><code class="language-zsh">python3 -m venv .venv
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install torch torchvision transformers</code></pre>
<p>설치 후에는 MPS 사용 가능 여부를 확인한다.</p>
<pre><code class="language-python">import torch

print(&quot;MPS 사용 가능:&quot;, torch.backends.mps.is_available())
print(&quot;MPS 빌드 여부:&quot;, torch.backends.mps.is_built())</code></pre>
<ul>
<li><code>is_built()</code>가 <code>False</code>이면 현재 PyTorch 빌드에 MPS 지원이 포함되지 않은 상태다.</li>
<li><code>is_built()</code>는 <code>True</code>인데 <code>is_available()</code>이 <code>False</code>이면 현재 장치나 macOS 환경에서 MPS를 사용할 수 없는 상태다.</li>
<li>MPS를 사용할 수 없으면 <code>cpu</code>로 실행하면 된다. 기능이 틀린 것이 아니라 가속 장치만 달라지는 것이다.</li>
</ul>
<p>macOS에서 <code>cuda</code>는 NVIDIA CUDA를 뜻하므로 일반적인 Apple Silicon Mac에서는 선택되지 않는다. 따라서 <code>torch.cuda.is_available()</code>이 <code>False</code>여도 오류가 아니다.</p>
<h3 id="84-windows-nvidia-gpu와-cuda">8.4 Windows: NVIDIA GPU와 CUDA</h3>
<p>Windows에서는 NVIDIA GPU가 있고 CUDA용 PyTorch를 설치했다면 보통 <code>cuda</code>를 사용한다. GPU가 없거나 CPU 전용 PyTorch를 설치했다면 <code>cpu</code>로 실행된다.</p>
<pre><code class="language-powershell">py -m venv .venv
.\.venv\Scripts\Activate.ps1

py -m pip install --upgrade pip
py -m pip install transformers</code></pre>
<p>PyTorch는 Windows 환경에서 GPU와 CUDA 버전에 맞는 설치 명령을 선택해야 한다. PyTorch의 설치 선택 화면에서 <code>Windows</code>, <code>Pip</code>, 사용 중인 CUDA 또는 <code>CPU</code>를 고른 뒤 제시된 명령을 사용한다.</p>
<pre><code class="language-powershell"># CPU 환경 예시
py -m pip install torch torchvision transformers

# NVIDIA GPU 환경은 PyTorch 공식 선택 화면에서
# 현재 CUDA 조합에 맞는 명령을 확인해 설치한다.</code></pre>
<p>설치가 끝나면 Python에서 GPU를 확인한다.</p>
<pre><code class="language-python">import torch

print(&quot;CUDA 사용 가능:&quot;, torch.cuda.is_available())
if torch.cuda.is_available():
    print(&quot;GPU 이름:&quot;, torch.cuda.get_device_name(0))</code></pre>
<p><code>torch.cuda.is_available()</code>이 <code>False</code>라면 NVIDIA GPU가 없거나, GPU 드라이버·CUDA 조합과 PyTorch 패키지가 맞지 않거나, CPU 전용 빌드를 설치했을 가능성이 있다. 이 경우 PyTorch 공식 설치 선택 화면에서 환경에 맞는 명령을 다시 확인한다.</p>
<h3 id="85-운영체제와-관계없이-동작하는-장치-선택-코드">8.5 운영체제와 관계없이 동작하는 장치 선택 코드</h3>
<p>노트북의 흐름을 조금 보완하면 macOS, Windows, CPU 환경에서 모두 안전하게 사용할 수 있다.</p>
<pre><code class="language-python">import torch

if torch.cuda.is_available():
    DEVICE = torch.device(&quot;cuda&quot;)       # 주로 Windows 또는 Linux의 NVIDIA GPU
elif torch.backends.mps.is_available():
    DEVICE = torch.device(&quot;mps&quot;)        # Apple Silicon Mac
elif hasattr(torch, &quot;xpu&quot;) and torch.xpu.is_available():
    DEVICE = torch.device(&quot;xpu&quot;)        # 지원되는 Intel GPU 환경
else:
    DEVICE = torch.device(&quot;cpu&quot;)

print(&quot;사용 장치:&quot;, DEVICE)</code></pre>
<p><code>hasattr(torch, &quot;xpu&quot;)</code> 검사는 PyTorch 버전에 따라 <code>torch.xpu</code> 속성이 없는 환경에서도 오류 없이 다음 조건으로 넘어가게 한다.</p>
<hr>
<h2 id="9-pytorch-tensor와-실행-장치">9. PyTorch Tensor와 실행 장치</h2>
<p>Token ID는 정수 인덱스이므로 <code>torch.long</code>, 즉 일반적으로 <code>torch.int64</code> 자료형 Tensor로 만든다.</p>
<pre><code class="language-python">input_ids = torch.tensor(token_ids, dtype=torch.long)</code></pre>
<p>노트북에서는 실행 장치를 다음 순서로 선택했고, 현재 환경에서는 <code>mps</code>가 선택됐다.</p>
<pre><code class="language-python">if torch.cuda.is_available():
    DEVICE = torch.device(&quot;cuda&quot;)
elif torch.backends.mps.is_available():
    DEVICE = torch.device(&quot;mps&quot;)
elif hasattr(torch, &quot;xpu&quot;) and torch.xpu.is_available():
    DEVICE = torch.device(&quot;xpu&quot;)
else:
    DEVICE = torch.device(&quot;cpu&quot;)

print(DEVICE)</code></pre>
<pre><code class="language-text">mps</code></pre>
<p><code>mps</code>는 Apple Silicon Mac의 GPU 가속을 위한 PyTorch 장치다. 학습이나 추론 시 모델과 입력 Tensor를 같은 장치에 올려야 한다.</p>
<pre><code class="language-python">model = model.to(DEVICE)
encoded = {key: value.to(DEVICE) for key, value in encoded.items()}</code></pre>
<p>모델은 <code>mps</code>에 있는데 입력은 CPU Tensor인 상태처럼 장치가 다르면 오류가 발생한다. <code>model</code>과 <code>input_ids</code>, <code>attention_mask</code> 등 입력 Tensor를 모두 같은 <code>DEVICE</code>로 옮긴다.</p>
<hr>
<h2 id="10-자주-만나는-환경-오류">10. 자주 만나는 환경 오류</h2>
<table>
<thead>
<tr>
<th>상황</th>
<th>원인</th>
<th>해결 방법</th>
</tr>
</thead>
<tbody><tr>
<td><code>ModuleNotFoundError: No module named &#39;transformers&#39;</code></td>
<td>현재 실행 중인 Python 환경에 <code>transformers</code>가 설치되지 않았다.</td>
<td>가상환경을 활성화한 뒤 <code>python -m pip install transformers</code>를 실행한다.</td>
</tr>
<tr>
<td>Jupyter에서는 import 오류, 터미널에서는 정상</td>
<td>Notebook Kernel이 다른 가상환경을 사용한다.</td>
<td>Notebook의 Kernel을 <code>.venv</code> 환경으로 바꾸고, <code>import sys; print(sys.executable)</code>로 경로를 확인한다.</td>
</tr>
<tr>
<td><code>MPS is not available</code></td>
<td>Apple Silicon, macOS, PyTorch 빌드 중 하나가 MPS 조건을 만족하지 않는다.</td>
<td><code>torch.backends.mps.is_built()</code>와 <code>is_available()</code>을 확인하고, 사용 불가하면 CPU로 실행한다.</td>
</tr>
<tr>
<td><code>torch.cuda.is_available()</code>이 <code>False</code></td>
<td>Windows에서 CPU 전용 PyTorch를 설치했거나 NVIDIA GPU·드라이버·CUDA 조합이 맞지 않는다.</td>
<td>PyTorch 공식 설치 선택 화면에서 Windows와 CUDA 조합에 맞는 명령을 다시 확인한다.</td>
</tr>
<tr>
<td><code>IProgress not found</code> 경고</td>
<td>Jupyter의 진행 표시 위젯이 설치되지 않았다.</td>
<td>현재 가상환경에 <code>python -m pip install ipywidgets</code>를 설치하고 Kernel을 다시 시작한다.</td>
</tr>
<tr>
<td>Hugging Face 비인증 요청 경고</td>
<td>공개 모델은 불러올 수 있지만 인증되지 않은 요청이라 제한 관련 안내가 표시된다.</td>
<td>학습용 공개 모델 실습에서는 경고만으로 중단되지는 않는다. 반복 사용 시 Hugging Face 계정 인증과 캐시 사용을 검토한다.</td>
</tr>
</tbody></table>
<p>가상환경에서 <code>pip</code>가 다른 Python을 가리키는 문제를 줄이려면 <code>pip install ...</code>보다 <code>python -m pip install ...</code> 형식을 사용하는 편이 안전하다. <code>python</code>이 바로 현재 활성화된 가상환경의 Python을 가리키기 때문이다.</p>
<hr>
<h2 id="11-전체-흐름-정리">11. 전체 흐름 정리</h2>
<pre><code class="language-text">Corpus
  ↓
Document 선택
  ↓
Tokenization
  ↓
Vocabulary에서 Token ID 조회
  ↓
Padding·Attention Mask 생성
  ↓
PyTorch Tensor로 변환
  ↓
NLP 모델 입력</code></pre>
<ul>
<li>Corpus는 학습·분석에 사용하는 전체 텍스트 집합이고, Document는 그 안의 개별 텍스트다.</li>
<li>Token은 모델이 다루는 입력 단위이며, 단어 하나와 항상 같지 않다.</li>
<li>Vocabulary는 Token과 정수 ID의 대응표다.</li>
<li><code>&lt;UNK&gt;</code>는 사전에 없는 Token을 처리하지만, Subword Tokenization은 OOV 문제를 줄이는 데 도움이 된다.</li>
<li><code>AutoTokenizer</code>는 모델에 맞는 Vocabulary와 Tokenization 규칙을 불러온다.</li>
<li>세밀한 NLP 모델 입력에는 Token ID뿐 아니라 Special Token, Padding, Attention Mask도 함께 필요하다.</li>
<li>macOS에서는 MPS, Windows의 NVIDIA GPU 환경에서는 CUDA를 우선 확인하고, GPU가 없거나 사용할 수 없으면 CPU로 실행한다.</li>
</ul>
<hr>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 카드 짝 맞추기]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%B9%B4%EB%93%9C-%EC%A7%9D-%EB%A7%9E%EC%B6%94%EA%B8%B0</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%B9%B4%EB%93%9C-%EC%A7%9D-%EB%A7%9E%EC%B6%94%EA%B8%B0</guid>
            <pubDate>Sun, 20 Sep 2026 22:01:47 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>4 x 4 보드에서 같은 그림 카드 두 장을 선택해 제거한다. 방향키 이동, Ctrl + 방향키 이동, Enter 입력은 각각 1회 조작으로 센다.</p>
<p>현재 커서 위치에서 모든 카드 쌍을 제거하는 최소 조작 횟수를 구한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p>이 문제에는 두 종류의 탐색이 필요하다.</p>
<ol>
<li><strong>BFS</strong>: 현재 남은 카드 상태에서 한 커서 위치에서 다른 위치까지의 최소 이동 횟수를 구한다.</li>
<li><strong>DFS + 메모이제이션</strong>: 어떤 카드 쌍부터 제거할지, 그 카드 쌍에서 어느 카드를 먼저 선택할지 탐색한다.</li>
</ol>
<p>카드 종류는 최대 6개이므로 제거 순서는 최대 <code>6!</code>개다. 각 쌍은 두 카드 중 어느 쪽을 먼저 선택할지 2가지 경우가 있으므로, 모든 경우를 탐색해도 충분하다.</p>
<h2 id="남은-카드-상태를-비트마스크로-표현하기">남은 카드 상태를 비트마스크로 표현하기</h2>
<p>카드 숫자는 1부터 6까지다. 어떤 카드 종류가 제거되었는지를 비트로 표현한다.</p>
<pre><code class="language-text">mask의 k번 비트가 1: 숫자 k 카드 쌍은 이미 제거됨
mask의 k번 비트가 0: 숫자 k 카드 쌍이 아직 남아 있음</code></pre>
<p>보드를 직접 수정하지 않아도 <code>mask</code>만으로 현재 칸에 카드가 남아 있는지 판단할 수 있다.</p>
<pre><code class="language-python">def has_card(row, col, removed_mask):
    value = board[row][col]
    return value != 0 and (removed_mask &amp; (1 &lt;&lt; value)) == 0</code></pre>
<h2 id="ctrl-이동-처리">Ctrl 이동 처리</h2>
<p>Ctrl + 방향키는 해당 방향으로 이동하다가 다음 중 하나를 만나면 멈춘다.</p>
<ul>
<li>가장 가까운 남은 카드</li>
<li>보드의 끝</li>
</ul>
<p>따라서 한 칸씩 전진하면서 카드 또는 경계를 만날 때까지 확인한다.</p>
<h2 id="dfs-상태">DFS 상태</h2>
<p><code>dfs(removed_mask, row, col)</code>은 현재 커서 위치와 제거된 카드 상태에서, 남은 카드를 모두 제거하는 최소 조작 횟수다.</p>
<p>아직 남은 카드 종류 하나를 골라 두 장을 제거한다.</p>
<ul>
<li>첫 번째 카드 → 두 번째 카드 순서</li>
<li>두 번째 카드 → 첫 번째 카드 순서</li>
</ul>
<p>두 경우를 모두 계산한다. 카드 한 쌍을 선택하려면 Enter가 두 번 필요하므로 이동 횟수에 2를 더한다.</p>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">from collections import deque
from functools import lru_cache


def solution(board, r, c):
    positions = {}

    for row in range(4):
        for col in range(4):
            value = board[row][col]

            if value != 0:
                positions.setdefault(value, []).append((row, col))

    card_numbers = list(positions)
    full_mask = sum(1 &lt;&lt; number for number in card_numbers)
    directions = [(-1, 0), (1, 0), (0, -1), (0, 1)]

    def has_card(row, col, removed_mask):
        value = board[row][col]
        return value != 0 and (removed_mask &amp; (1 &lt;&lt; value)) == 0

    def ctrl_move(row, col, dr, dc, removed_mask):
        while True:
            next_row = row + dr
            next_col = col + dc

            # 해당 방향의 끝 칸에서 멈춘다.
            if not (0 &lt;= next_row &lt; 4 and 0 &lt;= next_col &lt; 4):
                return row, col

            row, col = next_row, next_col

            if has_card(row, col, removed_mask):
                return row, col

    def move_distance(start_row, start_col, target_row, target_col, removed_mask):
        visited = [[False] * 4 for _ in range(4)]
        visited[start_row][start_col] = True
        queue = deque([(start_row, start_col, 0)])

        while queue:
            row, col, distance = queue.popleft()

            if (row, col) == (target_row, target_col):
                return distance

            for dr, dc in directions:
                # 일반 방향키 이동
                next_row = row + dr
                next_col = col + dc

                if (
                    0 &lt;= next_row &lt; 4
                    and 0 &lt;= next_col &lt; 4
                    and not visited[next_row][next_col]
                ):
                    visited[next_row][next_col] = True
                    queue.append((next_row, next_col, distance + 1))

                # Ctrl + 방향키 이동
                next_row, next_col = ctrl_move(
                    row,
                    col,
                    dr,
                    dc,
                    removed_mask
                )

                if not visited[next_row][next_col]:
                    visited[next_row][next_col] = True
                    queue.append((next_row, next_col, distance + 1))

    @lru_cache(None)
    def dfs(removed_mask, row, col):
        if removed_mask == full_mask:
            return 0

        minimum = float(&quot;inf&quot;)

        for number in card_numbers:
            bit = 1 &lt;&lt; number

            if removed_mask &amp; bit:
                continue

            first, second = positions[number]
            next_mask = removed_mask | bit

            # first 카드를 먼저 선택하는 경우
            first_to_second = (
                move_distance(row, col, *first, removed_mask)
                + move_distance(*first, *second, removed_mask)
                + 2
                + dfs(next_mask, *second)
            )

            # second 카드를 먼저 선택하는 경우
            second_to_first = (
                move_distance(row, col, *second, removed_mask)
                + move_distance(*second, *first, removed_mask)
                + 2
                + dfs(next_mask, *first)
            )

            minimum = min(minimum, first_to_second, second_to_first)

        return minimum

    return dfs(0, r, c)</code></pre>
<h2 id="코드-설명">코드 설명</h2>
<p>카드 쌍을 제거하기 전까지 두 카드는 보드에 남아 있어야 한다. 따라서 첫 카드에서 두 번째 카드로 이동할 때도 <code>removed_mask</code>는 아직 바꾸지 않는다.</p>
<pre><code class="language-python">move_distance(*first, *second, removed_mask)</code></pre>
<p>두 카드를 모두 선택하고 Enter를 두 번 누른 뒤에만 <code>next_mask</code>로 바꾼다.</p>
<pre><code class="language-python">+ 2
+ dfs(next_mask, *second)</code></pre>
<p>이 순서가 중요하다. 먼저 카드를 제거해 버리면 Ctrl 이동이 실제 게임 규칙과 달라진다.</p>
<h2 id="예시">예시</h2>
<p>첫 번째 예시에서 카드 쌍 제거 순서와 각 쌍의 선택 순서에 따라 Ctrl 이동 결과가 달라진다.</p>
<p>예를 들어 어떤 카드 쌍을 먼저 제거하면 이후 빈 칸이 늘어나므로, Ctrl 이동이 더 멀리 이동할 수 있다. DFS는 가능한 제거 순서를 모두 비교하고, BFS는 각각의 현재 보드 상태에서 실제 최소 이동 횟수를 계산한다.</p>
<p>따라서 두 카드가 남아 있는 순서까지 고려한 최솟값을 구할 수 있다.</p>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p>카드 종류를 <code>K</code>라고 하자. <code>K &lt;= 6</code>이다.</p>
<p>DFS 상태는 제거된 카드 조합과 커서 위치로 이루어진다.</p>
<ul>
<li>제거 상태 수: 최대 <code>2^K</code></li>
<li>커서 위치 수: 16</li>
<li>한 상태에서 확인할 카드 순서: 최대 <code>2K</code></li>
<li>BFS 한 번: 보드 칸이 16개이므로 <code>O(1)</code></li>
</ul>
<p>제한된 4 x 4 보드와 최대 6개 카드 쌍에서는 충분히 빠르게 동작한다.</p>
<ul>
<li>시간 복잡도: <code>O(2^K * 16 * K)</code></li>
<li>공간 복잡도: <code>O(2^K * 16)</code></li>
</ul>
<h2 id="정리">정리</h2>
<p>카드 제거 순서가 Ctrl 이동 경로를 바꾸므로, 제거 순서와 카드 선택 순서를 모두 탐색해야 한다. 제거 상태는 비트마스크로 저장하고, 각 상태의 커서 이동은 BFS로 계산하면 최소 조작 횟수를 구할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 트리 트리오 중간값]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%ED%8A%B8%EB%A6%AC-%ED%8A%B8%EB%A6%AC%EC%98%A4-%EC%A4%91%EA%B0%84%EA%B0%92</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%ED%8A%B8%EB%A6%AC-%ED%8A%B8%EB%A6%AC%EC%98%A4-%EC%A4%91%EA%B0%84%EA%B0%92</guid>
            <pubDate>Sat, 19 Sep 2026 15:55:07 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>트리에서 서로 다른 세 정점 <code>a</code>, <code>b</code>, <code>c</code>를 골랐을 때, 세 쌍의 거리의 중간값을 <code>f(a, b, c)</code>라고 한다.</p>
<p>모든 세 정점 조합 중 <code>f</code>의 최댓값을 구한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p>트리에서 가장 먼 두 정점 사이의 거리를 <strong>지름</strong>이라고 하자. 어떤 세 정점을 골라도 두 정점 사이의 거리는 지름을 넘을 수 없으므로, <code>f</code>의 최댓값도 지름을 넘을 수 없다.</p>
<p>지름의 양 끝을 <code>u</code>, <code>v</code>, 지름 길이를 <code>D</code>라고 하자.</p>
<h3 id="답이-d가-되는-경우">답이 D가 되는 경우</h3>
<p><code>u</code>에서 거리가 <code>D</code>인 정점이 둘 이상 있다면, 그중 두 정점과 <code>u</code>를 고를 수 있다.</p>
<pre><code class="language-text">u와 x의 거리 = D
u와 y의 거리 = D
x와 y의 거리 &lt;= D</code></pre>
<p>세 거리의 중간값은 <code>D</code>이므로 답은 <code>D</code>다.</p>
<p><code>u</code>에서는 가장 먼 정점이 하나뿐이어도, <code>v</code>에서 거리가 <code>D</code>인 정점이 둘 이상이라면 같은 이유로 답은 <code>D</code>다.</p>
<h3 id="답이-d---1이-되는-경우">답이 D - 1이 되는 경우</h3>
<p>지름 양 끝 모두에서 가장 먼 정점이 하나뿐이면, 지름 길이 <code>D</code>를 두 번 포함하는 세 정점을 만들 수 없다. 이때 최댓값은 <code>D - 1</code>이 된다.</p>
<p>따라서 다음만 확인하면 된다.</p>
<ol>
<li>지름의 한 끝 <code>u</code>를 찾는다.</li>
<li><code>u</code>에서 가장 먼 정점들의 개수를 센다.</li>
<li>그 개수가 2 이상이면 <code>D</code>를 반환한다.</li>
<li>그렇지 않으면 반대 끝 <code>v</code>에서도 같은 검사를 한다.</li>
<li>둘 다 하나뿐이면 <code>D - 1</code>을 반환한다.</li>
</ol>
<h2 id="지름의-한-끝-찾기">지름의 한 끝 찾기</h2>
<p>트리의 임의의 정점에서 BFS를 수행해 가장 먼 정점을 찾으면, 그 정점은 지름의 한 끝이 된다.</p>
<p>그 정점에서 다시 BFS를 수행하면 지름 길이와 반대쪽 끝을 찾을 수 있다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>간선 정보로 인접 리스트를 만든다.</li>
<li>1번 정점에서 BFS를 수행해 지름의 한 끝 <code>u</code>를 찾는다.</li>
<li><code>u</code>에서 BFS를 수행해 지름 길이 <code>D</code>와 최장 거리 정점 개수를 구한다.</li>
<li>최장 거리 정점이 2개 이상이면 <code>D</code>를 반환한다.</li>
<li>반대쪽 지름 끝 <code>v</code>에서 BFS를 수행한다.</li>
<li><code>v</code>에서도 최장 거리 정점이 2개 이상이면 <code>D</code>, 아니면 <code>D - 1</code>을 반환한다.</li>
</ol>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">from collections import deque


def solution(n, edges):
    graph = [[] for _ in range(n)]

    for left, right in edges:
        left -= 1
        right -= 1
        graph[left].append(right)
        graph[right].append(left)

    def bfs(start):
        distance = [-1] * n
        distance[start] = 0
        queue = deque([start])

        farthest_node = start
        max_distance = 0
        farthest_count = 1

        while queue:
            node = queue.popleft()

            for next_node in graph[node]:
                if distance[next_node] != -1:
                    continue

                distance[next_node] = distance[node] + 1
                queue.append(next_node)

                if distance[next_node] &gt; max_distance:
                    max_distance = distance[next_node]
                    farthest_node = next_node
                    farthest_count = 1
                elif distance[next_node] == max_distance:
                    farthest_count += 1

        return farthest_node, max_distance, farthest_count

    # 임의의 정점에서 가장 먼 정점은 지름의 한 끝이다.
    diameter_end, _, _ = bfs(0)

    # 지름의 한 끝에서 지름 길이와 최장 거리 정점 개수를 구한다.
    other_end, diameter, farthest_count = bfs(diameter_end)

    if farthest_count &gt;= 2:
        return diameter

    # 반대쪽 끝에서도 최장 거리 정점이 여러 개인지 확인한다.
    _, _, farthest_count = bfs(other_end)

    if farthest_count &gt;= 2:
        return diameter

    return diameter - 1</code></pre>
<h2 id="예시">예시</h2>
<h3 id="예시-1-일직선-트리">예시 1: 일직선 트리</h3>
<pre><code class="language-text">1 - 2 - 3 - 4</code></pre>
<p>지름은 <code>1</code>과 <code>4</code> 사이의 거리 3이다. 지름 끝인 1에서 가장 먼 정점은 4 하나뿐이고, 4에서도 1 하나뿐이다.</p>
<p>따라서 답은 <code>3 - 1 = 2</code>다.</p>
<h3 id="예시-2-별-모양-트리">예시 2: 별 모양 트리</h3>
<pre><code class="language-text">1   2
 \ /
  5
 / \
3   4</code></pre>
<p>어떤 잎 정점에서 시작해도 다른 잎 정점 여러 개가 거리 2로 가장 멀다. 지름 길이 2가 두 번 포함되는 세 정점을 만들 수 있으므로 답은 <code>2</code>다.</p>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p>정점 수를 <code>N</code>이라고 하자.</p>
<p>BFS는 간선과 정점을 각각 한 번씩 방문하므로 <code>O(N)</code>이다. BFS를 세 번 수행한다.</p>
<ul>
<li>시간 복잡도: <code>O(N)</code></li>
<li>공간 복잡도: <code>O(N)</code></li>
</ul>
<p>재귀 DFS 대신 BFS를 사용하므로, 정점 수가 25만 개인 긴 일자 트리에서도 재귀 깊이 제한에 걸리지 않는다.</p>
<h2 id="정리">정리</h2>
<p>최댓값은 지름 길이 <code>D</code> 또는 <code>D - 1</code> 중 하나다. 지름 끝에서 최장 거리 정점이 여러 개인지 확인하면, 세 정점이 지름 길이를 두 번 만들 수 있는지 판단할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스] 선입 선출 스케줄링 ]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%84%A0%EC%9E%85-%EC%84%A0%EC%B6%9C-%EC%8A%A4%EC%BC%80%EC%A4%84%EB%A7%81</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4-%EC%84%A0%EC%9E%85-%EC%84%A0%EC%B6%9C-%EC%8A%A4%EC%BC%80%EC%A4%84%EB%A7%81</guid>
            <pubDate>Sat, 19 Sep 2026 15:47:05 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>처리 시간이 서로 다른 여러 CPU 코어에 작업을 순서대로 배정한다.</p>
<ul>
<li>시작 시각에는 모든 코어가 비어 있으므로, 앞 번호 코어부터 작업을 하나씩 받는다.</li>
<li>어떤 코어의 작업이 끝나면 즉시 다음 작업을 받는다.</li>
<li>같은 시각에 여러 코어가 비면 번호가 작은 코어부터 작업을 받는다.</li>
</ul>
<p><code>n</code>번째, 즉 마지막 작업을 처리하는 코어 번호를 반환한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p>시간 <code>time</code>이 주어졌을 때, 그 시각까지 배정된 작업 수를 계산할 수 있다.</p>
<pre><code class="language-text">처음 시각에 배정되는 작업 수: 코어 개수
이후 각 코어가 처리한 작업 수: time // 코어 처리 시간</code></pre>
<p>따라서 전체 작업 수는 다음과 같다.</p>
<pre><code class="language-python">len(cores) + sum(time // core for core in cores)</code></pre>
<p>시간이 증가할수록 이 값은 줄어들지 않는다. 그러므로 <code>n</code>번째 작업이 배정되는 최소 시각을 이분 탐색으로 찾을 수 있다.</p>
<h2 id="마지막-코어-찾기">마지막 코어 찾기</h2>
<p>최소 시각을 <code>time</code>이라고 하자.</p>
<p><code>time - 1</code>까지 배정된 작업 수를 구하면, <code>time</code>에 동시에 비는 코어들 중 몇 번째 코어가 <code>n</code>번째 작업을 받는지 알 수 있다.</p>
<pre><code class="language-text">remaining = n - (time - 1초까지 배정된 작업 수)</code></pre>
<p>앞 번호 코어부터 확인하면서 <code>time % core == 0</code>인 코어를 만날 때마다 <code>remaining</code>을 하나씩 줄인다. 값이 0이 되는 코어가 답이다.</p>
<h2 id="풀이-과정">풀이 과정</h2>
<ol>
<li>작업 수가 코어 수 이하라면 시작 시각에 배정되므로 답은 <code>n</code>번 코어다.</li>
<li>이분 탐색으로 <code>n</code>번째 작업이 배정되는 최소 시각을 찾는다.</li>
<li>그 직전 시각까지 배정된 작업 수를 구한다.</li>
<li>최소 시각에 비는 코어를 번호순으로 확인해 마지막 작업의 코어를 찾는다.</li>
</ol>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">def solution(n, cores):
    core_count = len(cores)

    # 시작 시각에 앞 번호 코어부터 하나씩 작업을 받는다.
    if n &lt;= core_count:
        return n

    def assigned_work_count(time):
        return core_count + sum(time // core for core in cores)

    left = 0
    right = max(cores) * (n - core_count)

    # n번째 작업이 배정되는 최소 시각을 찾는다.
    while left &lt; right:
        mid = (left + right) // 2

        if assigned_work_count(mid) &gt;= n:
            right = mid
        else:
            left = mid + 1

    time = left
    assigned_before = assigned_work_count(time - 1)
    remaining = n - assigned_before

    # 같은 시각에 비는 코어는 번호순으로 다음 작업을 받는다.
    for index, core in enumerate(cores, start=1):
        if time % core == 0:
            remaining -= 1

            if remaining == 0:
                return index</code></pre>
<h2 id="예시">예시</h2>
<p><code>n = 6</code>, <code>cores = [1, 2, 3]</code>인 경우를 보자.</p>
<table>
<thead>
<tr>
<th>시각</th>
<th>비는 코어</th>
<th>배정되는 작업</th>
</tr>
</thead>
<tbody><tr>
<td>0</td>
<td>1, 2, 3</td>
<td>1, 2, 3</td>
</tr>
<tr>
<td>1</td>
<td>1</td>
<td>4</td>
</tr>
<tr>
<td>2</td>
<td>1, 2</td>
<td>5, 6</td>
</tr>
</tbody></table>
<p>6번째 작업은 시각 2에 비는 코어 중 두 번째인 2번 코어에 배정되므로 답은 <code>2</code>다.</p>
<h2 id="시간-복잡도">시간 복잡도</h2>
<p><code>C</code>를 코어 수, <code>T</code>를 탐색 시간의 최댓값이라고 하자.</p>
<ul>
<li><p>작업 수를 계산하는 데 <code>O(C)</code></p>
</li>
<li><p>이분 탐색은 <code>O(log T)</code></p>
</li>
<li><p>마지막 코어를 찾는 데 <code>O(C)</code></p>
</li>
<li><p>시간 복잡도: <code>O(C log T)</code></p>
</li>
<li><p>공간 복잡도: <code>O(1)</code></p>
</li>
</ul>
<h2 id="정리">정리</h2>
<p>작업 하나씩을 직접 시뮬레이션할 필요가 없다. <code>n</code>번째 작업이 배정되는 시각을 먼저 찾고, 그 시각에 비는 코어를 번호순으로 세면 마지막 작업을 처리하는 코어를 구할 수 있다.</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[[프로그래머스]빛의 경로 사이클]]></title>
            <link>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4%EB%B9%9B%EC%9D%98-%EA%B2%BD%EB%A1%9C-%EC%82%AC%EC%9D%B4%ED%81%B4</link>
            <guid>https://velog.io/@j_keun/%ED%94%84%EB%A1%9C%EA%B7%B8%EB%9E%98%EB%A8%B8%EC%8A%A4%EB%B9%9B%EC%9D%98-%EA%B2%BD%EB%A1%9C-%EC%82%AC%EC%9D%B4%ED%81%B4</guid>
            <pubDate>Mon, 14 Sep 2026 13:22:32 GMT</pubDate>
            <description><![CDATA[<h2 id="문제-요약">문제 요약</h2>
<p>각 칸에 <code>S</code>, <code>L</code>, <code>R</code>이 적힌 격자에서 빛은 현재 칸의 지시에 따라 직진, 좌회전, 우회전한 뒤 다음 칸으로 이동한다. 격자 밖으로 나가면 반대편으로 이어지는 토러스 구조다.</p>
<p>모든 빛의 경로 사이클 길이를 구해 오름차순으로 반환해야 한다.</p>
<h2 id="핵심-아이디어">핵심 아이디어</h2>
<p>빛의 경로는 위치만으로 결정되지 않는다. 같은 칸에 있어도 어느 방향으로 들어왔는지에 따라 다음 이동이 달라진다.</p>
<p>따라서 아래 세 값을 하나의 상태로 관리한다.</p>
<pre><code class="language-text">(행, 열, 방향)</code></pre>
<p>행이 <code>N</code>, 열이 <code>M</code>이면 가능한 상태는 총 <code>N * M * 4</code>개다. 각 상태에서 다음 상태는 정확히 하나로 정해진다.</p>
<p>또한 이동 규칙은 역방향으로도 상태를 하나만 복원할 수 있으므로, 모든 상태는 어떤 경로 사이클에 포함된다. 방문하지 않은 상태에서 시작해 같은 상태로 돌아올 때까지 이동하면 한 사이클의 길이를 셀 수 있다.</p>
<h2 id="이동-규칙">이동 규칙</h2>
<p>방향을 <code>0: 상</code>, <code>1: 우</code>, <code>2: 하</code>, <code>3: 좌</code>로 정의한다.</p>
<pre><code class="language-python">dr = [-1, 0, 1, 0]
dc = [0, 1, 0, -1]</code></pre>
<p>현재 칸의 문자에 따라 방향을 바꾼다.</p>
<ul>
<li><code>S</code>: 방향 유지</li>
<li><code>L</code>: <code>(direction - 1) % 4</code></li>
<li><code>R</code>: <code>(direction + 1) % 4</code></li>
</ul>
<p>다음 위치는 나머지 연산으로 계산하면 격자 경계를 자연스럽게 넘어갈 수 있다.</p>
<pre><code class="language-python">next_row = (row + dr[direction]) % row_count
next_col = (col + dc[direction]) % col_count</code></pre>
<h2 id="python-코드">Python 코드</h2>
<pre><code class="language-python">def solution(grid):
    row_count = len(grid)
    col_count = len(grid[0])

    # visited[row][col][direction]
    visited = [
        [[False] * 4 for _ in range(col_count)]
        for _ in range(row_count)
    ]

    dr = [-1, 0, 1, 0]  # 상, 우, 하, 좌
    dc = [0, 1, 0, -1]
    cycles = []

    for start_row in range(row_count):
        for start_col in range(col_count):
            for start_direction in range(4):
                if visited[start_row][start_col][start_direction]:
                    continue

                row = start_row
                col = start_col
                direction = start_direction
                length = 0

                # 시작 상태로 되돌아올 때까지 한 경로를 추적한다.
                while not visited[row][col][direction]:
                    visited[row][col][direction] = True
                    length += 1

                    if grid[row][col] == &quot;L&quot;:
                        direction = (direction - 1) % 4
                    elif grid[row][col] == &quot;R&quot;:
                        direction = (direction + 1) % 4

                    row = (row + dr[direction]) % row_count
                    col = (col + dc[direction]) % col_count

                cycles.append(length)

    return sorted(cycles)</code></pre>
<h2 id="정당성-설명">정당성 설명</h2>
<p>각 상태 <code>(행, 열, 방향)</code>에서 현재 칸의 문자와 격자 크기는 고정되어 있으므로, 다음 상태는 정확히 하나로 결정된다.</p>
<p>방문하지 않은 상태에서 이동을 시작하면 유한한 상태 집합 안에서 계속 이동하므로 언젠가 이미 방문한 상태에 도달한다. 이 문제의 이동은 방향 전환과 토러스 이동으로 이루어져 역방향 상태도 유일하게 정할 수 있다. 따라서 경로 중간의 다른 사이클으로 합류하는 경우가 없고, 처음 방문 상태로 되돌아온 구간 전체가 하나의 경로 사이클이다.</p>
<p>알고리즘은 사이클을 추적하며 포함된 모든 상태를 방문 처리한다. 이후 해당 상태에서 다시 시작하지 않으므로 각 상태는 정확히 한 번만 어떤 사이클 길이에 포함된다. 따라서 모든 경로 사이클의 길이를 중복이나 누락 없이 구한다.</p>
<h2 id="복잡도-분석">복잡도 분석</h2>
<p>상태 수는 <code>4 * N * M</code>개다.</p>
<ul>
<li>경로 탐색: <code>O(N * M)</code></li>
<li>결과 정렬까지 포함한 시간 복잡도: <code>O(N * M * log(N * M))</code></li>
<li>공간 복잡도: <code>O(N * M)</code></li>
</ul>
<p>각 상태는 한 번만 방문한다. 다만 사이클 길이를 오름차순으로 반환해야 하므로, 최악의 경우 결과 배열 정렬 비용이 추가된다.</p>
<h2 id="주의할-점">주의할 점</h2>
<ul>
<li><code>visited[row][col]</code>처럼 위치만 방문 처리하면 서로 다른 방향의 상태가 섞여 잘못된 결과가 나온다.</li>
<li>방향 전환은 <strong>현재 칸에서</strong> 수행한 뒤 이동해야 한다.</li>
<li>Python에서 음수 인덱스 대신 <code>(direction - 1) % 4</code>를 사용하면 좌회전을 명확하게 표현할 수 있다.</li>
<li>행과 열의 경계 이동에는 <code>% row_count</code>, <code>% col_count</code>를 각각 사용해야 한다.</li>
</ul>
<h2 id="마무리">마무리</h2>
<p>이 문제는 격자를 그래프로 직접 만들 필요 없이, 위치와 방향을 합친 상태 그래프로 생각하면 간단해진다. 전체 상태를 한 번씩만 순회하는 시뮬레이션으로 모든 빛의 경로 사이클을 찾을 수 있다.</p>
]]></description>
        </item>
    </channel>
</rss>