<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>def_uk.log</title>
        <link>https://velog.io/</link>
        <description>크아앙</description>
        <lastBuildDate>Mon, 28 Sep 2026 07:48:26 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <image>
            <title>def_uk.log</title>
            <url>https://velog.velcdn.com/images/def_uk/profile/c4e6737e-a860-4bb6-ace8-fd4457336d7f/image.png</url>
            <link>https://velog.io/</link>
        </image>
        <copyright>Copyright (C) 2019. def_uk.log. All rights reserved.</copyright>
        <atom:link href="https://v2.velog.io/rss/def_uk" rel="self" type="application/rss+xml"/>
        <item>
            <title><![CDATA[ResNet (Residual Network)]]></title>
            <link>https://velog.io/@def_uk/ResNet-Residual-Network</link>
            <guid>https://velog.io/@def_uk/ResNet-Residual-Network</guid>
            <pubDate>Mon, 28 Sep 2026 07:48:26 GMT</pubDate>
            <description><![CDATA[<blockquote>
<p>He et al., <em>Deep Residual Learning for Image Recognition</em>, 2015 (ILSVRC 2015 우승)</p>
</blockquote>
<hr>
<h2 id="1-배경-왜-나왔나-problem">1. 배경: 왜 나왔나? (Problem)</h2>
<p>VGG 이후로 &quot;층을 깊게 쌓을수록 성능이 좋아진다&quot;는 게 거의 정설이었다. 그런데 실제로 계속 쌓아보니 이상한 일이 생겼다.</p>
<ul>
<li>56층 네트워크가 20층 네트워크보다 <strong>test error뿐 아니라 train error도 더 높게</strong> 나왔다.</li>
<li>train error까지 높다는 건 <strong>overfitting이 아니라는 뜻</strong>이다. 모델이 아예 학습을 제대로 못 하고 있는 것.</li>
<li>Batch Normalization 등으로 vanishing gradient는 어느 정도 잡은 상태였는데도 이런 현상이 나타났다.</li>
</ul>
<p>이걸 <strong>Degradation Problem</strong>이라고 부른다.</p>
<p>논리적으로 보면 말이 안 된다. 20층 모델에 <strong>아무것도 안 하는 층(identity)</strong> 36개만 붙여도 56층 모델은 최소한 20층만큼은 나와야 한다. 즉, 문제는 표현력이 아니라 <strong>&quot;깊은 네트워크가 identity 같은 단순한 함수조차 학습하기 어렵다&quot;는 최적화 문제</strong>였다.</p>
<ul>
<li>identity: 항등함수</li>
</ul>
<hr>
<h2 id="2-다른-구조와-뭐가-다른가">2. 다른 구조와 뭐가 다른가?</h2>
<table>
<thead>
<tr>
<th>구분</th>
<th>Plain Network (VGG 등)</th>
<th>ResNet</th>
</tr>
</thead>
<tbody><tr>
<td>학습 대상</td>
<td>원하는 출력 <code>H(x)</code>를 직접 학습</td>
<td>차이값 <code>F(x) = H(x) - x</code>만 학습</td>
</tr>
<tr>
<td>층의 출력</td>
<td><code>y = F(x)</code></td>
<td><code>y = F(x) + x</code></td>
</tr>
<tr>
<td>identity가 최적일 때</td>
<td>여러 비선형 층으로 identity를 근사해야 함 (어려움)</td>
<td><code>F(x) → 0</code>으로 만들면 끝 (쉬움)</td>
</tr>
<tr>
<td>gradient 흐름</td>
<td>층을 지날 때마다 곱해지며 약해짐</td>
<td>shortcut을 통해 직접 전달되는 경로가 있음</td>
</tr>
</tbody></table>
<p>비슷한 시기의 <strong>Highway Network</strong>도 shortcut을 썼지만, 여기엔 학습 가능한 gate가 있어서 파라미터가 늘고 gate가 닫히면 shortcut이 막힐 수 있다. ResNet의 identity shortcut은 <strong>파라미터 0개, 항상 열려 있다</strong>는 점이 다르다.</p>
<p>Highway Network와 ResNet 모두 입력을 뒤쪽으로 전달하는 <strong>shortcut(skip connection)</strong>을 사용한다. 차이는 shortcut을 조절하는 방식이다. Highway Network는 $y=T(x)H(x)+(1-T(x))x$처럼 학습 가능한 gate를 곱해 원래 입력 $x$를 얼마나 전달할지 결정한다. Gate 값에 따라 shortcut이 거의 막힐 수도 있다. 반면 ResNet의 identity shortcut은 $y=F(x)+x$처럼 입력을 그대로 더한다. Shortcut 자체에는 학습할 파라미터가 없고, gate에 의해 닫히지도 않는다.</p>
<hr>
<h2 id="3-그래서-어떻게-풀었나-solve">3. 그래서 어떻게 풀었나? (Solve)</h2>
<h3 id="핵심-residual-block--skip-connection">핵심: Residual Block + Skip Connection</h3>
<pre><code>x ──────────────────┐
│                   │ (identity shortcut)
▼                   │
Conv → BN → ReLU    │
│                   │
Conv → BN           │
│                   │
▼                   │
(+) ◄───────────────┘
│
ReLU
▼
y = F(x) + x</code></pre><h3 id="왜-잘-되나">왜 잘 되나?</h3>
<ol>
<li><p><strong>학습 목표가 쉬워진다</strong>
출력 전체를 새로 만드는 대신 입력에서 &quot;얼마나 바꿀지&quot;만 배운다. 필요 없는 층은 가중치를 0 근처로 두면 그냥 통과시킬 수 있다.</p>
</li>
<li><p><strong>gradient가 잘 흐른다</strong>
$y = F(x) + x$를 미분하면
$$\frac{\partial y}{\partial x} = \frac{\partial F}{\partial x} + 1$$
$+1$ 덕분에 gradient가 아무리 깊어도 사라지지 않고 앞쪽 층까지 전달된다.</p>
</li>
</ol>
<h3 id="세부-설계">세부 설계</h3>
<ul>
<li><strong>차원이 안 맞을 때</strong>: stride나 채널 수가 바뀌어 <code>x</code>와 <code>F(x)</code>의 shape이 다르면, shortcut에 <code>1x1 conv</code>(projection)를 넣어 맞춘다.</li>
<li><strong>Basic Block</strong> (ResNet-18/34): <code>3x3 → 3x3</code></li>
<li><strong>Bottleneck Block</strong> (ResNet-50/101/152): <code>1x1(채널 축소) → 3x3 → 1x1(채널 복원)</code>
깊게 쌓으면서도 연산량을 줄이기 위한 구조.</li>
</ul>
<p>결과적으로 <strong>152층</strong>까지 안정적으로 학습시켰고, ImageNet top-5 error 3.57%로 당시 1위를 차지했다.</p>
<hr>
<h2 id="4-코드로는-어떻게-쓰나">4. 코드로는 어떻게 쓰나?</h2>
<h3 id="1-residual-block-직접-구현-pytorch">(1) Residual Block 직접 구현 (PyTorch)</h3>
<pre><code class="language-python">import torch.nn as nn

class BasicBlock(nn.Module):
    def __init__(self, in_ch, out_ch, stride=1):
        super().__init__()
        self.conv1 = nn.Conv2d(in_ch, out_ch, 3, stride, 1, bias=False)
        self.bn1   = nn.BatchNorm2d(out_ch)
        self.conv2 = nn.Conv2d(out_ch, out_ch, 3, 1, 1, bias=False)
        self.bn2   = nn.BatchNorm2d(out_ch)
        self.relu  = nn.ReLU(inplace=True)

        # shape이 달라지면 1x1 conv로 shortcut을 맞춰줌
        self.shortcut = nn.Identity()
        if stride != 1 or in_ch != out_ch:
            self.shortcut = nn.Sequential(
                nn.Conv2d(in_ch, out_ch, 1, stride, bias=False),
                nn.BatchNorm2d(out_ch),
            )

    def forward(self, x):
        out = self.relu(self.bn1(self.conv1(x)))
        out = self.bn2(self.conv2(out))
        out = out + self.shortcut(x)   # 핵심: F(x) + x
        return self.relu(out)</code></pre>
<h3 id="2-사전학습-모델-가져와서-fine-tuning">(2) 사전학습 모델 가져와서 Fine-tuning</h3>
<p>직접 구현하기보다 ImageNet으로 학습된 가중치를 가져와 <strong>분류기(fc)만 바꿔서</strong> 쓰는 경우가 대부분이다.</p>
<pre><code class="language-python">import torch.nn as nn
from torchvision.models import resnet50, ResNet50_Weights

model = resnet50(weights=ResNet50_Weights.DEFAULT)

# (선택) 백본은 고정하고 분류기만 학습
for p in model.parameters():
    p.requires_grad = False

num_classes = 5
model.fc = nn.Linear(model.fc.in_features, num_classes)  # 새 분류기는 학습됨</code></pre>
<p>데이터가 충분하면 백본 일부(<code>layer4</code> 등)를 풀어서 함께 학습시키면 성능이 더 오른다.</p>
<h3 id="3-개념은-어디에나-있다">(3) 개념은 어디에나 있다</h3>
<p>Skip connection은 이미지 분류를 넘어서 이후 거의 모든 딥러닝 구조의 기본 부품이 됐다.</p>
<ul>
<li><strong>Transformer</strong>: 각 sublayer마다 <code>x + Attention(x)</code>, <code>x + FFN(x)</code></li>
<li><strong>U-Net</strong>: encoder-decoder 사이 skip connection</li>
<li><strong>DenseNet</strong>: 더하기 대신 이어붙이기(concat)로 확장</li>
</ul>
<blockquote>
<p>한 줄 요약: <strong>&quot;깊게 쌓으면 학습이 안 된다&quot; → &quot;차이만 배우게 하고 입력은 그대로 더해주자&quot;</strong></p>
</blockquote>
]]></description>
        </item>
        <item>
            <title><![CDATA[GoogLeNet (Inception v1)]]></title>
            <link>https://velog.io/@def_uk/GoogLeNet-Inception-v1</link>
            <guid>https://velog.io/@def_uk/GoogLeNet-Inception-v1</guid>
            <pubDate>Mon, 28 Sep 2026 07:02:44 GMT</pubDate>
            <description><![CDATA[<blockquote>
<p>2014, Google (Szegedy et al.) · ILSVRC 2014 분류 부문 우승 (Top-5 error 6.67%)</p>
</blockquote>
<hr>
<h2 id="1-왜-나왔나--problem">1. 왜 나왔나? — Problem</h2>
<p>2012년 AlexNet 이후 CNN 연구의 흐름은 <strong>&quot;네트워크를 더 깊고 넓게 만들면 성능이 오른다&quot;</strong> 였다. 하지만 이 방식에는 분명한 한계가 있었다.</p>
<ul>
<li><strong>연산량과 파라미터의 폭발</strong>: 층과 채널을 늘리면 연산량이 제곱 단위로 늘어난다. 모델이 커질수록 학습·추론 비용이 감당하기 어려워진다.</li>
<li><strong>과적합</strong>: 파라미터가 많을수록 데이터가 충분하지 않으면 과적합되기 쉽다. 특히 마지막 Fully Connected 층에 파라미터가 몰려 있었다.</li>
<li><strong>필터 크기 선택의 딜레마</strong>: 이미지 속 객체의 크기는 제각각이다. 작은 필터(3×3)는 세밀한 특징을, 큰 필터(5×5)는 넓은 문맥을 잘 잡는데, 층마다 어떤 크기를 쓸지 사람이 정해야 했다.</li>
<li><strong>기울기 소실</strong>: 깊어질수록 앞쪽 층까지 gradient가 잘 전달되지 않는다.</li>
</ul>
<p>즉, <strong>&quot;깊게 쌓고 싶은데 비용은 감당할 수 있게&quot;</strong> 라는 문제를 풀어야 했다.</p>
<hr>
<h2 id="2-다른-모델과-뭐가-다른가">2. 다른 모델과 뭐가 다른가?</h2>
<table>
<thead>
<tr>
<th></th>
<th>AlexNet (2012)</th>
<th>VGG-16 (2014)</th>
<th><strong>GoogLeNet (2014)</strong></th>
</tr>
</thead>
<tbody><tr>
<td>깊이</td>
<td>8층</td>
<td>16층</td>
<td><strong>22층</strong></td>
</tr>
<tr>
<td>파라미터 수</td>
<td>약 6,000만</td>
<td>약 1억 3,800만</td>
<td><strong>약 500만</strong></td>
</tr>
<tr>
<td>설계 철학</td>
<td>CNN의 가능성 증명</td>
<td>3×3 conv를 단순하게 깊이 쌓기</td>
<td><strong>여러 크기 필터를 병렬로, 효율적으로</strong></td>
</tr>
<tr>
<td>네트워크 형태</td>
<td>한 줄로 쌓음</td>
<td>한 줄로 쌓음</td>
<td><strong>층 안에서 가지가 갈라졌다 합쳐짐</strong></td>
</tr>
<tr>
<td>마지막 분류부</td>
<td>큰 FC 층 3개</td>
<td>큰 FC 층 3개</td>
<td><strong>Global Average Pooling + FC 1개</strong></td>
</tr>
</tbody></table>
<p>VGG가 &quot;단순함&quot;으로 깊이를 얻었다면, GoogLeNet은 <strong>&quot;구조 설계&quot;로 깊이와 효율을 동시에</strong> 얻었다. VGG보다 깊으면서도 파라미터는 약 1/27이다.</p>
<hr>
<h2 id="3-그래서-어떻게-풀었나--solve">3. 그래서 어떻게 풀었나? — Solve</h2>
<h3 id="①-inception-모듈-→-필터-크기-딜레마-해결">① Inception 모듈 → 필터 크기 딜레마 해결</h3>
<p>한 층에서 필터 크기를 하나만 고르지 않고, <strong>1×1 / 3×3 / 5×5 conv와 3×3 max pooling을 병렬로 수행한 뒤 채널 방향으로 이어 붙인다.</strong> 어떤 스케일의 정보가 중요한지는 네트워크가 학습으로 알아서 조합한다. 이 모듈을 9개 쌓아 네트워크 몸통을 만든다.</p>
<pre><code>            ┌─ 1×1 conv ─────────────────┐
            ├─ 1×1 conv → 3×3 conv ──────┤
입력 ───────┤                            ├─ Concat → 출력
            ├─ 1×1 conv → 5×5 conv ──────┤
            └─ 3×3 maxpool → 1×1 conv ───┘</code></pre><h3 id="②-1×1-convolution-병목-→-연산량-폭발-해결">② 1×1 Convolution 병목 → 연산량 폭발 해결</h3>
<p>병렬로 여러 필터를 돌리면 연산량이 커지므로, 3×3·5×5 연산 <strong>앞에</strong> 1×1 conv를 두어 채널 수를 먼저 줄인다.</p>
<table>
<thead>
<tr>
<th>입력 28×28×192 → 5×5 conv로 32채널 출력</th>
<th>곱셈 연산량</th>
</tr>
</thead>
<tbody><tr>
<td>5×5 conv 바로 적용</td>
<td>약 1억 2,000만 회</td>
</tr>
<tr>
<td>1×1 conv로 16채널 축소 → 5×5 conv</td>
<td>약 1,240만 회</td>
</tr>
</tbody></table>
<p>약 <strong>10배</strong> 절감된다. 1×1 conv 뒤에도 ReLU가 붙어 비선형성이 추가되는 효과도 있다.</p>
<h3 id="③-global-average-pooling-→-파라미터·과적합-해결">③ Global Average Pooling → 파라미터·과적합 해결</h3>
<p>마지막 feature map을 채널별로 평균 내어 1×1×1024 벡터로 만든다. 거대한 FC 층을 없애서 파라미터를 크게 줄이고 과적합도 완화했다.</p>
<h3 id="④-보조-분류기auxiliary-classifier-→-기울기-소실-완화">④ 보조 분류기(Auxiliary Classifier) → 기울기 소실 완화</h3>
<p>깊은 신경망은 마지막 층에서 계산한 오차가 앞쪽 층까지 전달되는 과정에서 학습 신호가 약해질 수 있다. 이를 완화하기 위해 중간 층에 작은 분류기를 붙여 추가 loss를 계산한다.</p>
<pre><code>전체 loss = 메인 loss + 0.3 × (보조 loss1 + 보조 loss2)</code></pre><p>추론할 때는 떼어낸다. (이후 Batch Normalization이 등장하면서 필요성은 줄었다.)</p>
<hr>
<h2 id="4-코드로-어떻게-활용하나">4. 코드로 어떻게 활용하나?</h2>
<h3 id="①-inception-모듈-직접-구현-pytorch">① Inception 모듈 직접 구현 (PyTorch)</h3>
<pre><code class="language-python">import torch
import torch.nn as nn

class Inception(nn.Module):
    def __init__(self, in_ch, c1, c3_red, c3, c5_red, c5, pool_proj):
        super().__init__()
        # 1×1
        self.b1 = nn.Sequential(nn.Conv2d(in_ch, c1, 1), nn.ReLU(inplace=True))
        # 1×1 병목 → 3×3
        self.b2 = nn.Sequential(
            nn.Conv2d(in_ch, c3_red, 1), nn.ReLU(inplace=True),
            nn.Conv2d(c3_red, c3, 3, padding=1), nn.ReLU(inplace=True))
        # 1×1 병목 → 5×5
        self.b3 = nn.Sequential(
            nn.Conv2d(in_ch, c5_red, 1), nn.ReLU(inplace=True),
            nn.Conv2d(c5_red, c5, 5, padding=2), nn.ReLU(inplace=True))
        # 3×3 maxpool → 1×1
        self.b4 = nn.Sequential(
            nn.MaxPool2d(3, stride=1, padding=1),
            nn.Conv2d(in_ch, pool_proj, 1), nn.ReLU(inplace=True))

    def forward(self, x):
        # 네 갈래의 출력을 채널 방향(dim=1)으로 이어 붙임
        return torch.cat([self.b1(x), self.b2(x), self.b3(x), self.b4(x)], dim=1)

# 논문의 inception(3a) 설정
m = Inception(192, 64, 96, 128, 16, 32, 32)
x = torch.randn(1, 192, 28, 28)
print(m(x).shape)  # torch.Size([1, 256, 28, 28])  ← 64+128+32+32</code></pre>
<p>padding으로 공간 크기(28×28)를 맞춰두었기 때문에 채널 방향으로 concat이 가능하다.</p>
<h3 id="②-사전학습-모델로-전이학습-실무에서-가장-흔한-방식">② 사전학습 모델로 전이학습 (실무에서 가장 흔한 방식)</h3>
<pre><code class="language-python">import torch.nn as nn
from torchvision import models
from torchvision.models import GoogLeNet_Weights

weights = GoogLeNet_Weights.DEFAULT
model = models.googlenet(weights=weights)   # ImageNet 사전학습, 보조 분류기는 제거된 상태
preprocess = weights.transforms()           # 학습 때와 같은 전처리

# 백본 고정 (데이터가 적을 때)
for p in model.parameters():
    p.requires_grad = False

# 마지막 분류기만 내 클래스 수에 맞게 교체
model.fc = nn.Linear(model.fc.in_features, 10)  # in_features = 1024</code></pre>
<h3 id="③-처음부터-학습할-때-보조-분류기-loss-활용">③ 처음부터 학습할 때: 보조 분류기 loss 활용</h3>
<pre><code class="language-python">model = models.googlenet(weights=None, aux_logits=True,
                         num_classes=10, init_weights=True)
criterion = nn.CrossEntropyLoss()

model.train()
out = model(images)  # GoogLeNetOutputs(logits, aux_logits2, aux_logits1)
loss = criterion(out.logits, labels) \
     + 0.3 * (criterion(out.aux_logits1, labels) + criterion(out.aux_logits2, labels))

model.eval()
pred = model(images)  # 추론 시에는 메인 출력만 반환</code></pre>
<hr>
<h2 id="한-줄-정리">한 줄 정리</h2>
<blockquote>
<p>GoogLeNet은 <strong>여러 크기의 필터를 병렬로 쓰는 Inception 모듈</strong>과 <strong>1×1 conv 병목</strong>으로, 깊이는 늘리면서 연산량과 파라미터는 줄인 모델이다. 이 병목 설계는 이후 ResNet의 bottleneck block, MobileNet 같은 경량 모델로 이어진다.</p>
</blockquote>
]]></description>
        </item>
        <item>
            <title><![CDATA[VGGNet]]></title>
            <link>https://velog.io/@def_uk/VGGNet-%EC%A0%95%EB%A6%AC</link>
            <guid>https://velog.io/@def_uk/VGGNet-%EC%A0%95%EB%A6%AC</guid>
            <pubDate>Mon, 28 Sep 2026 06:56:48 GMT</pubDate>
            <description><![CDATA[<h2 id="1-왜-나왔나-problem">1. 왜 나왔나? (Problem)</h2>
<p>2012년 AlexNet이 ImageNet에서 큰 성공을 거두면서, 연구자들은 <strong>&quot;CNN을 어떻게 설계해야 성능이 더 좋아질까?&quot;</strong>를 고민하기 시작했다.</p>
<p>당시에는 필터 크기, stride, 층 수 등을 이것저것 바꿔보는 식이어서, <strong>어떤 요소가 성능을 결정하는지</strong> 명확하지 않았다. 특히 이런 질문이 남아 있었다.</p>
<ul>
<li>네트워크를 <strong>깊게</strong> 만들면 성능이 좋아지는가?</li>
<li>그런데 층을 늘리면 파라미터와 연산량이 폭발하는데, 이걸 어떻게 감당하나?</li>
</ul>
<p>2014년 옥스퍼드 대학 Visual Geometry Group(Simonyan &amp; Zisserman)은 다른 조건은 최대한 단순하게 고정하고 <strong>깊이 하나만 바꿔가며</strong> 이 질문에 답하려 했다.</p>
<h2 id="2-다른-모델과-뭐가-다른가">2. 다른 모델과 뭐가 다른가?</h2>
<table>
<thead>
<tr>
<th>구분</th>
<th>AlexNet (2012)</th>
<th>VGGNet (2014)</th>
<th>GoogLeNet (2014)</th>
</tr>
</thead>
<tbody><tr>
<td>필터 크기</td>
<td>11×11, 5×5, 3×3 혼용</td>
<td><strong>3×3으로 통일</strong></td>
<td>1×1, 3×3, 5×5 병렬 (Inception)</td>
</tr>
<tr>
<td>깊이</td>
<td>8층</td>
<td><strong>16~19층</strong></td>
<td>22층</td>
</tr>
<tr>
<td>구조</td>
<td>층마다 설정이 제각각</td>
<td><strong>단순하고 규칙적</strong></td>
<td>복잡한 모듈 구조</td>
</tr>
<tr>
<td>파라미터</td>
<td>약 6천만</td>
<td>약 1억 3,800만</td>
<td>약 700만</td>
</tr>
</tbody></table>
<ul>
<li><strong>AlexNet 대비</strong>: 큰 필터를 버리고 작은 필터를 깊게 쌓았다.</li>
<li><strong>GoogLeNet 대비</strong>: 성능은 약간 낮았지만(ILSVRC 2014 분류 2위), 구조가 훨씬 단순해서 이해·구현·변형이 쉬웠다. 그래서 오히려 더 널리 쓰였다.</li>
</ul>
<h2 id="3-어떻게-풀었나-solve">3. 어떻게 풀었나? (Solve)</h2>
<h3 id="핵심-3×3-합성곱을-여러-번-쌓는다">핵심: 3×3 합성곱을 여러 번 쌓는다</h3>
<p>3×3 합성곱을 연속으로 쌓으면 큰 필터와 같은 <strong>수용 영역(receptive field)</strong>을 얻는다.</p>
<table>
<thead>
<tr>
<th>구성</th>
<th>수용 영역</th>
<th>파라미터 수 (채널 C 기준)</th>
</tr>
</thead>
<tbody><tr>
<td>7×7 합성곱 1층</td>
<td>7×7</td>
<td>49C²</td>
</tr>
<tr>
<td>3×3 합성곱 3층</td>
<td>7×7</td>
<td><strong>27C²</strong></td>
</tr>
</tbody></table>
<p>이렇게 하면 깊이를 늘리면서도 다음 이점을 얻는다.</p>
<ol>
<li><strong>파라미터 감소</strong>: 같은 수용 영역을 더 적은 파라미터로 얻는다.</li>
<li><strong>비선형성 증가</strong>: 층마다 ReLU가 들어가므로 표현력이 풍부해진다.</li>
</ol>
<h3 id="전체-구조">전체 구조</h3>
<p>입력 224×224 → 합성곱 블록 5개 → FC 3개 → Softmax</p>
<ul>
<li>각 블록 = 3×3 합성곱(stride 1, padding 1) 여러 개 + 2×2 Max Pooling(stride 2)</li>
<li>풀링마다 공간 크기는 절반, 채널은 두 배 (64 → 128 → 256 → 512 → 512)</li>
<li>마지막은 FC 4096 → FC 4096 → FC 1000</li>
<li><strong>VGG16</strong>: 합성곱 13층 + FC 3층 / <strong>VGG19</strong>: 합성곱 16층 + FC 3층</li>
</ul>
<h3 id="결과와-남은-문제">결과와 남은 문제</h3>
<p>깊이가 깊어질수록 성능이 좋아진다는 것을 실험으로 보여주었다. 다만 한계도 분명했다.</p>
<ul>
<li>파라미터 대부분이 FC층에 몰려 있어 <strong>모델이 무겁다</strong> (첫 FC층만 약 1억 개).</li>
<li>19층 이상 쌓으면 <strong>기울기 소실</strong>로 학습이 어려워진다 → 이후 <strong>ResNet</strong>이 잔차 연결로 해결.</li>
</ul>
<h2 id="4-코드로-어떻게-활용하나-pytorch">4. 코드로 어떻게 활용하나? (PyTorch)</h2>
<h3 id="4-1-구조-직접-구현해보기">4-1. 구조 직접 구현해보기</h3>
<p>블록 구조가 규칙적이라 설정 리스트만으로 네트워크를 만들 수 있다.</p>
<pre><code class="language-python">import torch.nn as nn

# 숫자 = 3×3 합성곱의 출력 채널, &#39;M&#39; = Max Pooling
VGG16_CFG = [64, 64, &#39;M&#39;, 128, 128, &#39;M&#39;, 256, 256, 256, &#39;M&#39;,
             512, 512, 512, &#39;M&#39;, 512, 512, 512, &#39;M&#39;]

def make_features(cfg):
    layers, in_ch = [], 3
    for v in cfg:
        if v == &#39;M&#39;:
            layers.append(nn.MaxPool2d(kernel_size=2, stride=2))
        else:
            layers += [nn.Conv2d(in_ch, v, kernel_size=3, padding=1),
                       nn.ReLU(inplace=True)]
            in_ch = v
    return nn.Sequential(*layers)

class VGG16(nn.Module):
    def __init__(self, num_classes=1000):
        super().__init__()
        self.features = make_features(VGG16_CFG)
        self.classifier = nn.Sequential(
            nn.Flatten(),
            nn.Linear(512 * 7 * 7, 4096), nn.ReLU(True), nn.Dropout(0.5),
            nn.Linear(4096, 4096), nn.ReLU(True), nn.Dropout(0.5),
            nn.Linear(4096, num_classes),
        )

    def forward(self, x):
        return self.classifier(self.features(x))</code></pre>
<h3 id="4-2-전이학습-가장-흔한-활용">4-2. 전이학습 (가장 흔한 활용)</h3>
<p>실무에서는 직접 학습하기보다 ImageNet으로 사전학습된 모델을 가져와 내 데이터에 맞게 마지막 층만 바꾼다.</p>
<pre><code class="language-python">import torch.nn as nn
from torchvision import models

model = models.vgg16(weights=models.VGG16_Weights.IMAGENET1K_V1)

# 특징 추출부는 고정
for p in model.features.parameters():
    p.requires_grad = False

# 마지막 FC층을 내 클래스 수에 맞게 교체 (예: 5개 클래스)
model.classifier[6] = nn.Linear(4096, 5)</code></pre>
<h3 id="4-3-특징-추출기로-활용-perceptual-loss-등">4-3. 특징 추출기로 활용 (Perceptual Loss 등)</h3>
<p>VGG의 중간층 출력은 이미지의 질감·형태를 잘 담고 있어서, 스타일 트랜스퍼나 초해상도 모델에서 <strong>&quot;두 이미지가 사람 눈에 얼마나 비슷한가&quot;</strong>를 측정하는 데 쓰인다.</p>
<pre><code class="language-python">import torch.nn as nn
from torchvision import models

vgg = models.vgg16(weights=models.VGG16_Weights.IMAGENET1K_V1).features[:16].eval()
for p in vgg.parameters():
    p.requires_grad = False

def perceptual_loss(pred, target):
    # 픽셀 값 대신 VGG 특징 맵끼리 비교
    return nn.functional.mse_loss(vgg(pred), vgg(target))</code></pre>
<h2 id="정리">정리</h2>
<ul>
<li><strong>Problem</strong>: CNN에서 깊이가 성능에 어떤 영향을 주는지 불분명했다.</li>
<li><strong>차이점</strong>: 큰 필터 대신 3×3 필터만 쓰고, 구조를 단순·규칙적으로 만들었다.</li>
<li><strong>Solve</strong>: 작은 필터를 깊게 쌓아 파라미터는 줄이고 비선형성은 늘렸다.</li>
<li><strong>활용</strong>: 지금은 분류 모델 자체보다 전이학습 백본, 특징 추출기로 주로 쓰인다.</li>
</ul>
]]></description>
        </item>
        <item>
            <title><![CDATA[AlexNet]]></title>
            <link>https://velog.io/@def_uk/AlexNet</link>
            <guid>https://velog.io/@def_uk/AlexNet</guid>
            <pubDate>Mon, 28 Sep 2026 06:30:39 GMT</pubDate>
            <description><![CDATA[<blockquote>
<p>Krizhevsky, Sutskever, Hinton. <em>ImageNet Classification with Deep Convolutional Neural Networks</em>. NeurIPS 2012.</p>
</blockquote>
<hr>
<h2 id="1-왜-나왔나-problem">1. 왜 나왔나? (Problem)</h2>
<p>2012년 이전의 이미지 분류는 <strong>사람이 특징을 설계</strong>하는 방식이었다. SIFT, HOG 같은 특징 추출기로 이미지를 벡터로 바꾸고, 그 위에 SVM 같은 분류기를 얹었다. 이 방식은 두 가지 벽에 부딪혀 있었다.</p>
<ul>
<li><strong>성능 정체</strong>: 특징을 사람이 만들다 보니 복잡한 이미지에서 한계가 뚜렷했다. ImageNet 대회 top-5 오류율은 25% 근처에서 맴돌았다.</li>
<li><strong>신경망은 &quot;이론상 좋지만 실제로는 안 되는&quot; 기술</strong>: CNN은 LeNet(1998)부터 있었지만 손글씨 숫자 같은 작은 문제에서만 통했다. 네트워크를 깊고 크게 만들면<ul>
<li>학습이 너무 느리고 (sigmoid/tanh의 기울기 포화, CPU 연산 한계)</li>
<li>데이터 대비 파라미터가 많아 <strong>과적합</strong>이 심했다.</li>
</ul>
</li>
</ul>
<p>즉 풀어야 할 문제는 이것이었다.
<strong>&quot;120만 장, 1000개 클래스짜리 대규모 데이터에서, 거대한 CNN을 현실적인 시간 안에, 과적합 없이 학습시킬 수 있는가?&quot;</strong></p>
<hr>
<h2 id="2-다른-방법과-뭐가-다른가">2. 다른 방법과 뭐가 다른가?</h2>
<table>
<thead>
<tr>
<th>비교 대상</th>
<th>그쪽 방식</th>
<th>AlexNet</th>
</tr>
</thead>
<tbody><tr>
<td>전통 방식 (SIFT + SVM)</td>
<td>특징을 사람이 설계</td>
<td>특징을 <strong>데이터로부터 학습</strong> (end-to-end)</td>
</tr>
<tr>
<td>LeNet-5 (1998)</td>
<td>작은 입력(32×32), 층 얕음, 파라미터 약 6만, CPU, tanh</td>
<td>224×224 컬러, Conv 5 + FC 3, 파라미터 <strong>약 6,000만</strong>, GPU, ReLU</td>
</tr>
<tr>
<td><div align="center"></td>
<td></td>
<td></td>
</tr>
<tr>
<td><img src="https://velog.velcdn.com/images/def_uk/post/610184ae-5be2-4958-ade3-4b1e99ff5357/image.png" alt="LeNet-5와 AlexNet"></td>
<td></td>
<td></td>
</tr>
<tr>
<td><p>좌측은 LeNet-5, 우측은 AlexNet입니다.<br></td>
<td></td>
<td></td>
</tr>
<tr>
<td>이미지 출처: <a href="https://daeun-computer-uneasy.tistory.com/33">daeun-computer-uneasy.tistory.com</a></td>
<td></td>
<td></td>
</tr>
<tr>
<td></p></td>
<td></td>
<td></td>
</tr>
<tr>
<td></div></td>
<td></td>
<td></td>
</tr>
</tbody></table>
<p>핵심 차이는 &quot;새로운 층을 발명했다&quot;가 아니라, <strong>규모를 키웠을 때 생기는 문제들을 실전적으로 해결해서 깊은 CNN이 실제로 동작함을 증명했다</strong>는 점이다.</p>
<hr>
<h2 id="3-그래서-어떻게-풀었나-solve">3. 그래서 어떻게 풀었나? (Solve)</h2>
<h3 id="구조-conv-5층--fc-3층">구조: Conv 5층 + FC 3층</h3>
<table>
<thead>
<tr>
<th>층</th>
<th>구성</th>
<th>비고</th>
</tr>
</thead>
<tbody><tr>
<td>Conv1</td>
<td>96개, 11×11, stride 4</td>
<td>ReLU → LRN → MaxPool</td>
</tr>
<tr>
<td>Conv2</td>
<td>256개, 5×5</td>
<td>ReLU → LRN → MaxPool</td>
</tr>
<tr>
<td>Conv3</td>
<td>384개, 3×3</td>
<td>ReLU</td>
</tr>
<tr>
<td>Conv4</td>
<td>384개, 3×3</td>
<td>ReLU</td>
</tr>
<tr>
<td>Conv5</td>
<td>256개, 3×3</td>
<td>ReLU → MaxPool</td>
</tr>
<tr>
<td>FC6</td>
<td>4096</td>
<td>ReLU + Dropout</td>
</tr>
<tr>
<td>FC7</td>
<td>4096</td>
<td>ReLU + Dropout</td>
</tr>
<tr>
<td>FC8</td>
<td>1000</td>
<td>Softmax</td>
</tr>
</tbody></table>
<h3 id="문제별-해결책">문제별 해결책</h3>
<p><strong>학습이 느리다 → ReLU + GPU</strong></p>
<ul>
<li><code>max(0, x)</code>는 양수 구간에서 기울기가 1로 유지돼 포화가 없다. 논문 실험에서 tanh 대비 같은 오류율 도달이 약 6배 빨랐다.</li>
<li>GTX 580(3GB) 한 장에 모델이 안 들어가서 <strong>네트워크를 절반씩 두 GPU에 나눠</strong> 학습했다. 특정 층에서만 GPU끼리 통신한다. (구조도가 위아래 두 갈래인 이유)</li>
</ul>
<p><strong>과적합이 심하다 → Dropout + Data Augmentation</strong></p>
<ul>
<li><strong>Dropout (p=0.5)</strong>: FC6, FC7에서 뉴런을 무작위로 끈다. 파라미터가 몰린 FC층의 과적합을 직접 겨냥했다.</li>
<li><strong>Data Augmentation</strong>: 256×256에서 무작위 224×224 크롭 + 좌우 반전, PCA 기반 색상 변형으로 데이터를 사실상 크게 늘렸다.</li>
</ul>
<p><strong>일반화 성능을 조금 더 → Overlapping Pooling, LRN</strong></p>
<ul>
<li>3×3 풀링을 stride 2로 겹쳐서 적용.</li>
<li>LRN으로 인접 채널 간 정규화. (단, 이후 효과가 미미하다고 밝혀져 BatchNorm에 자리를 내줌)</li>
</ul>
<p><strong>LRN(Local Response Normalization)이란?</strong>
LRN은 AlexNet에서 사용한 정규화 기법이다. 같은 공간 위치에 있는 인접 채널들의 활성값을 이용해 각 값을 조정한다. 주변 채널의 활성값이 클수록 더 큰 값으로 나누기 때문에, 해당 채널의 출력이 작아진다.
여기서 값을 <strong>‘누른다’</strong>는 것은 0으로 만든다는 뜻이 아니라 크기를 줄인다는 뜻이다. 큰 활성값도 LRN을 거치면 작아질 수 있으며, 채널 사이의 반응을 상대적으로 비교하는 효과가 있다.
AlexNet에서는 일부 합성곱 층의 ReLU 뒤, Max Pooling 앞에 LRN을 적용했다. 오늘날에는 LRN보다 Batch Normalization 같은 정규화 방법이 주로 사용된다.</p>
<p><strong>Overlapping Pooling이란?</strong>
풀링 창을 이동할 때 인접한 창이 서로 겹치도록 하는 방식이다. 풀링 창의 크기보다 이동 간격(stride)이 작으면 겹침이 생긴다.
예를 들어 AlexNet은 3×3 Max Pooling에 stride 2를 사용했다. 첫 번째 창이 가로 방향의 1<del>3번째 값을 보면, 다음 창은 3</del>5번째 값을 본다. 3번째 값이 두 창에 모두 포함되는 것이다.
반대로 2×2 창을 stride 2로 이동하면 창끼리 겹치지 않는다. AlexNet 논문에서는 overlapping pooling을 사용했을 때 오류율이 소폭 낮아졌다고 보고했다.</p>
<p><strong>학습 설정</strong>: SGD + momentum 0.9, weight decay 0.0005, batch 128, lr 0.01에서 시작해 정체 시 1/10, 약 90 epoch, 5~6일.</p>
<h3 id="결과">결과</h3>
<p>ILSVRC 2012 top-5 오류율 <strong>15.3%</strong>. 2위는 26.2%. 이 격차로 컴퓨터 비전의 주류가 하루아침에 딥러닝으로 넘어갔다.</p>
<hr>
<h2 id="4-코드로는-어떻게-쓰나">4. 코드로는 어떻게 쓰나?</h2>
<p>실무에서 AlexNet을 처음부터 학습시킬 일은 거의 없다. 주로 <strong>구조를 직접 구현하며 CNN을 이해하는 용도</strong>나, <strong>사전학습 모델로 전이학습을 연습하는 용도</strong>로 쓴다.</p>
<blockquote>
<p>참고: <code>torchvision</code>의 AlexNet은 원 논문이 아니라 단일 GPU 버전이다. Conv1 채널이 64개이고, LRN이 없다.</p>
</blockquote>
<h3 id="1-직접-구현-pytorch">(1) 직접 구현 (PyTorch)</h3>
<pre><code class="language-python">import torch
import torch.nn as nn

class AlexNet(nn.Module):
    def __init__(self, num_classes=1000):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(3, 96, kernel_size=11, stride=4, padding=2),
            nn.ReLU(inplace=True),
            nn.LocalResponseNorm(size=5),
            nn.MaxPool2d(kernel_size=3, stride=2),     # overlapping pooling

            nn.Conv2d(96, 256, kernel_size=5, padding=2),
            nn.ReLU(inplace=True),
            nn.LocalResponseNorm(size=5),
            nn.MaxPool2d(kernel_size=3, stride=2),

            nn.Conv2d(256, 384, kernel_size=3, padding=1),
            nn.ReLU(inplace=True),
            nn.Conv2d(384, 384, kernel_size=3, padding=1),
            nn.ReLU(inplace=True),
            nn.Conv2d(384, 256, kernel_size=3, padding=1),
            nn.ReLU(inplace=True),
            nn.MaxPool2d(kernel_size=3, stride=2),
        )
        self.classifier = nn.Sequential(
            nn.Dropout(0.5),
            nn.Linear(256 * 6 * 6, 4096),
            nn.ReLU(inplace=True),
            nn.Dropout(0.5),
            nn.Linear(4096, 4096),
            nn.ReLU(inplace=True),
            nn.Linear(4096, num_classes),
        )

    def forward(self, x):
        x = self.features(x)
        x = torch.flatten(x, 1)
        return self.classifier(x)

model = AlexNet()
print(model(torch.randn(1, 3, 224, 224)).shape)  # torch.Size([1, 1000])</code></pre>
<h3 id="2-사전학습-모델로-추론">(2) 사전학습 모델로 추론</h3>
<pre><code class="language-python">from torchvision.models import alexnet, AlexNet_Weights
from PIL import Image

weights = AlexNet_Weights.IMAGENET1K_V1
model = alexnet(weights=weights).eval()
preprocess = weights.transforms()   # 리사이즈·크롭·정규화를 한 번에

img = preprocess(Image.open(&quot;dog.jpg&quot;)).unsqueeze(0)
with torch.no_grad():
    pred = model(img).softmax(dim=1)

idx = pred.argmax().item()
print(weights.meta[&quot;categories&quot;][idx], pred[0, idx].item())</code></pre>
<h3 id="3-전이학습-내-데이터예-5개-클래스에-맞추기">(3) 전이학습: 내 데이터(예: 5개 클래스)에 맞추기</h3>
<pre><code class="language-python">model = alexnet(weights=AlexNet_Weights.IMAGENET1K_V1)

# 특징 추출부는 고정
for p in model.features.parameters():
    p.requires_grad = False

# 마지막 FC층만 내 클래스 수로 교체
model.classifier[6] = nn.Linear(4096, 5)

optimizer = torch.optim.SGD(
    filter(lambda p: p.requires_grad, model.parameters()),
    lr=1e-3, momentum=0.9, weight_decay=5e-4,   # 논문 설정과 같은 계열
)</code></pre>
]]></description>
        </item>
        <item>
            <title><![CDATA[데이터 과학을 위한 통계(1)]]></title>
            <link>https://velog.io/@def_uk/%EB%8D%B0%EC%9D%B4%ED%84%B0-%EA%B3%BC%ED%95%99%EC%9D%84-%EC%9C%84%ED%95%9C-%ED%86%B5%EA%B3%84-1</link>
            <guid>https://velog.io/@def_uk/%EB%8D%B0%EC%9D%B4%ED%84%B0-%EA%B3%BC%ED%95%99%EC%9D%84-%EC%9C%84%ED%95%9C-%ED%86%B5%EA%B3%84-1</guid>
            <pubDate>Fri, 25 Sep 2026 13:19:58 GMT</pubDate>
            <description><![CDATA[<h2 id="정형화된-데이터의-요소">정형화된 데이터의 요소</h2>
<p>센서 측정, 이미지, 비디오, 텍스트, 이벤트 등을 통해 졍형화되어 있지 않은 데이터를 받는다. <strong>데이터 과학</strong>은 이러한 데이터들의 <strong><em>폭발적인 양의 원시(가공되지 않은) 데이터</em></strong>를 활용 가능하도록 정형화된 형태로 변환해야 한다.
정형 데이터 중 가장 일반적인 형태는 열이 있는 <strong>테이블 형태</strong>이다.</p>
<h3 id="정형데이터의-종류">정형데이터의 종류</h3>
<h4 id="수치형-데이터">수치형 데이터</h4>
<ul>
<li><strong>연속형 데이터</strong>: 풍속이나 지속 시간</li>
<li><strong>이산 데이터</strong>: 사건의 발생 빈도<h4 id="범주형-데이터">범주형 데이터</h4>
도시명(대전, 부산, 서울)과 같이 범위가 정해진 값들을 같은 경우</li>
<li><strong>이진 데이터</strong>: 범주형 중에서도 참/거짓, 0과 1, 예/아니요 같은 두 값 중 하나를 갖는 특수한 경우</li>
<li><strong>순서형 데이터</strong>: 평점(1, 2, 3, 4, 5)이 대표적</li>
</ul>
<h3 id="왜-정형화된-데이터로-바꿀까">왜 정형화된 데이터로 바꿀까?</h3>
<p>데이터를 분석하고 예측을 모델링할 때, 시각화, 해석, 통계 모델 결정 등에 데이터 종류가 중요한 역할을 하기 때문이다.</p>
<hr>
<h2 id="테이블-데이터">테이블 데이터</h2>
<p>가장 대표적으로 사용되는 객체의 형태로 엑셀이나 데이터베의 테이블과 같은 테이블 데이터이다. 기본적으로 레코드를 나타내는 행과 피쳐(변수)를 나타내는 열로 이루어진 이차원 행렬을 의미한다.</p>
<blockquote>
<p>지표 변수가 결과 변수가 되기도 한다.</p>
</blockquote>
<h3 id="테이블-형식이-아닌-데이터-구조">테이블 형식이 아닌 데이터 구조</h3>
<ul>
<li><strong>시계열 데이터</strong>: 동일한 변수 안에 연속적인 측정값을 가진다.</li>
<li><strong>공간 데이터</strong>: 객체를 표현할 때는, 어떤 객체(주택)와 그것의 공간 좌표가 데이터의 중심이 된다. 반면 필드 정보는 공간을 나타내는 작은 단위들과 적당한 측정 기준값(픽셀의 밝기)에 중점을 둔다</li>
<li><strong>그래프(혹은 네트워크) 데이터</strong>: 물리적 관계, 사회적 관계, 그리고 다소 추상적인 관계들을 표현하기 위해 사용</li>
</ul>
<hr>
<h2 id="위치-추정">위치 추정</h2>
<p>데이터를 살펴보는 가장 기초적인 단계는 각 피처(변수)의 &#39;대푯값&#39;을 구하는 것이다. 대부분의 값이 어디에 위치하는지(중심경향성)을 나타내는 추정값</p>
<ul>
<li><strong>평균(mean)</strong>: 모든 값의 총합을 개수로 나눈 값</li>
<li><strong>가중평균(weighted mean)</strong>: 가중치를 곱한 값의 총합을 가중치의 총합으로 나눈 값</li>
<li><strong>중간값(median)</strong>: 데이터에서 가장 가운데 위치한 값, 로버스트한 위치 추정 방버</li>
<li><strong>백분위수(percentile)</strong>: 전체 데이터의 $P%$를 아래에 두는 값(크기순으로 나열한 데이터를 100등분하여 특정 위치의 값을 나타내는 통계적 수치)</li>
<li><strong>가중 중간값(weighted median)**</strong>: 데이터를 정렬한 후, 각 가중치 값을 위에서부터 더할 때, 총합의 중간이 위치하는 데이터 값</li>
<li><strong>절사평균(trimmed mean)</strong>: 정해진 개수의 극단값을 제외한 나머지 값들의 평균</li>
<li><strong>로버스트(robust)</strong>: 극단값들에 민감하지 않다는 것을 의미</li>
<li><strong>특잇값(outlier)</strong>: 대부분의 값과 매우 다른 데이터 값(유의어: 극단값)</li>
</ul>
<blockquote>
<p><strong>지수 가중 평균</strong>: 데이터의 이동 평균을 구할 때, 오래된 데이터가 미치는 영향을 지수적으로 감쇠(exponential decay) 하도록 만들어 주는 방법</p>
</blockquote>
<hr>
<h2 id="변이-추정">변이 추정</h2>
<p>데이터 값이 얼마나 밀집해 있는지 혹은 퍼져 있는지를 나타내는 산포도를 나타낸다. </p>
<ul>
<li><strong>편차(deviation)</strong>: 관측값과 위치 추정값 사이의 차이(유의어: 오차, 잔차)</li>
<li><strong>분산(variance)</strong>: 평균과의 편차를 제곱한 값들의 합을 n-1로 나눈 값, n은 데이터 개수(유의어: 평균제곱오차)</li>
<li><strong>표준편차(standard deviation)</strong>: 분산의 제곱근</li>
<li><strong>평균절대편차(mean absolute deviation)</strong>: 평균과의 편차의 절댓갑의 평균(유의어: L1 노름, 맨해튼 노름, 평균절대오차)</li>
<li><strong>중간값의 중위절대편차(MAD)</strong>: 중간값과의 편차의 절댓갑의 중간값</li>
<li><strong>순서 통계량(order statistics)</strong>: 최소에서 최대까지 정렬된 데이터 값에 따른 계량형</li>
</ul>
<h4 id="출처">출처</h4>
<p>데이터 과학을 위한 통계(p.19-38)</p>
]]></description>
        </item>
        <item>
            <title><![CDATA[ 제약이 능력을 만든다: PyTorch로 이해하는 AE, DAE, VAE]]></title>
            <link>https://velog.io/@def_uk/%EC%A0%9C%EC%95%BD%EC%9D%B4-%EB%8A%A5%EB%A0%A5%EC%9D%84-%EB%A7%8C%EB%93%A0%EB%8B%A4-PyTorch%EB%A1%9C-%EC%9D%B4%ED%95%B4%ED%95%98%EB%8A%94-AE-DAE-VAE</link>
            <guid>https://velog.io/@def_uk/%EC%A0%9C%EC%95%BD%EC%9D%B4-%EB%8A%A5%EB%A0%A5%EC%9D%84-%EB%A7%8C%EB%93%A0%EB%8B%A4-PyTorch%EB%A1%9C-%EC%9D%B4%ED%95%B4%ED%95%98%EB%8A%94-AE-DAE-VAE</guid>
            <pubDate>Tue, 15 Sep 2026 05:41:12 GMT</pubDate>
            <description><![CDATA[<h1 id="제약이-능력을-만든다-pytorch로-이해하는-ae-dae-vae">제약이 능력을 만든다: PyTorch로 이해하는 AE, DAE, VAE</h1>
<p>오토인코더를 처음 접하면 목표부터 이상하게 느껴진다. 입력한 이미지를 그대로 출력하도록 학습한다니, <code>return x</code> 한 줄이면 끝나는 것 아닐까?</p>
<p>하지만 입력이 좁은 통로를 지나야 한다면 이야기가 달라진다. 784개 픽셀을 더 적은 수의 값으로 표현하고 다시 복원하려면, 어떤 특징을 남길지 학습해야 한다.</p>
<p>이번 글에서는 MNIST 손글씨 이미지로 세 가지 모델을 살펴본다.</p>
<table>
<thead>
<tr>
<th>모델</th>
<th>학습 방식</th>
<th>살펴볼 질문</th>
</tr>
</thead>
<tbody><tr>
<td>AE · Autoencoder</td>
<td>입력을 압축한 뒤 복원</td>
<td>작은 잠재 벡터에 무엇이 남을까?</td>
</tr>
<tr>
<td>DAE · Denoising Autoencoder</td>
<td>손상된 입력에서 원본 복원</td>
<td>노이즈를 넣으면 무엇을 배우게 될까?</td>
</tr>
<tr>
<td>VAE · Variational Autoencoder</td>
<td>이미지별 잠재 분포로 복원하고 기준 분포에 가깝게 유도</td>
<td>새 이미지를 만들 좌표는 어디서 뽑을까?</td>
</tr>
</tbody></table>
<blockquote>
<p>아래 수치와 그림은 실습 노트북에 저장된 실행 결과다. 모델 초기화, 난수, 실행 환경에 따라 재실행 결과는 달라질 수 있다. 본문에는 핵심 코드를 싣고, 긴 비교·시각화 코드는 마지막 부록에 모았다.</p>
</blockquote>
<h2 id="1-실습-준비-28×28-이미지를-784차원-벡터로">1. 실습 준비: 28×28 이미지를 784차원 벡터로</h2>
<p>MNIST 이미지는 28×28 크기의 흑백 이미지다. <code>ToTensor()</code>로 픽셀을 0~1 범위의 텐서로 바꾸고, 완전연결층에 넣기 전에 784차원으로 펼친다.</p>
<p>필요한 패키지는 다음과 같이 설치한다.</p>
<pre><code class="language-bash">pip install torch torchvision matplotlib</code></pre>
<p>다음은 원본 실습의 환경 설정이다. Apple Silicon의 MPS를 사용할 수 있으면 활용하고, 그렇지 않으면 CPU에서 실행한다.</p>
<pre><code class="language-python">import torch
import torch.nn as nn
import torch.nn.functional as F
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
import matplotlib.pyplot as plt

# 모델 초기화와 노이즈 생성 등에 사용할 공통 시드입니다.
SEED = 42
torch.manual_seed(SEED)

device = &quot;mps&quot; if torch.backends.mps.is_available() else &quot;cpu&quot;

train_set = datasets.MNIST(&quot;./data&quot;, train=True, download=True,
                           transform=transforms.ToTensor())
train_loader = DataLoader(train_set, batch_size=256, shuffle=True)</code></pre>
<p><code>DataLoader</code>는 학습 데이터를 256장씩 묶어 제공한다. 한 번에 처리하는 양과 업데이트 횟수를 조절하기 위한 선택이다.</p>
<p><code>torch.flatten(images, start_dim=1)</code>은 배치 차원을 유지하면서 이미지 차원만 합친다.</p>
<pre><code class="language-text">[배치 크기, 1, 28, 28] → [배치 크기, 784]</code></pre>
<h2 id="2-ae-복원을-배우게-하는-병목">2. AE: 복원을 배우게 하는 병목</h2>
<p>AE는 인코더와 디코더로 구성된다.</p>
<ul>
<li><strong>인코더:</strong> 입력 <code>x</code>를 잠재 벡터 <code>z</code>로 변환한다.</li>
<li><strong>디코더:</strong> 잠재 벡터 <code>z</code>로 복원 이미지 <code>x̂</code>를 만든다.</li>
</ul>
<p>이번 모델의 구조는 다음과 같다.</p>
<pre><code class="language-text">입력 784 → 은닉층 128 → 잠재 벡터 32 → 은닉층 128 → 출력 784</code></pre>
<p>잠재 벡터는 이미지를 복원하는 데 사용할 요약 표현이다. 병목은 정보를 그대로 전달하기 어렵게 만들지만, 사람이 원하는 의미적 특징을 자동으로 보장하지는 않는다.</p>
<pre><code class="language-python">class AutoEncoder(nn.Module):
    #잠재표현 층을 몇 차원으로 줄건지 latent_dim으로 받음.
    def __init__(self, latent_dim):
        super().__init__()
        #encoder
        self.encoder = nn.Sequential(
            nn.Linear(784, 128), 
            nn.ReLU(),
            nn.Linear(128, latent_dim)
        )
        #decoder
        self.decoder = nn.Sequential(
            nn.Linear(latent_dim, 128), 
            nn.ReLU(),
            nn.Linear(128, 784), 
            nn.Sigmoid()
        )
    #forward = 순전파
    def forward(self, x):
        z = self.encoder(x)
        return self.decoder(z)</code></pre>
<p>마지막 <code>Sigmoid</code>는 출력 픽셀을 0~1 범위로 제한한다. <code>forward()</code>에서는 입력을 인코딩한 뒤 바로 디코딩한다.</p>
<h3 id="숫자-레이블-대신-입력-자체를-목표로-사용한다">숫자 레이블 대신 입력 자체를 목표로 사용한다</h3>
<p>학습할 때 숫자 클래스 레이블은 사용하지 않는다. 대신 복원 결과와 입력 이미지 사이의 MSE, 즉 픽셀별 평균 제곱 오차를 줄인다.</p>
<p>여기서 “레이블을 쓰지 않는다”는 “학습 목표가 없다”는 뜻이 아니다. <strong>입력 이미지 자체가 복원의 목표</strong>다. MSE는 이번 실습에서 선택한 복원 손실이며, 픽셀이 연속값이라는 이유로 반드시 MSE만 써야 하는 것은 아니다.</p>
<pre><code class="language-python">model = AutoEncoder(latent_dim=32).to(device)
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

for epoch in range(5):
    total = 0
    for images, _ in train_loader:
        #start_dim=1: 1번 차원부터 합치기를 시작합니다. (보통 0번 차원은 배치 크기(Batch Size)이므로 유지합니다.)
        x = torch.flatten(images, start_dim=1).to(device)

        output = model(x)
        loss = criterion(output, x)

        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        total += loss.item()
    print(f&quot;epoch {epoch+1}: {total/len(train_loader):.4f}&quot;)</code></pre>
<p>저장된 학습 로그는 다음과 같다. 이 값은 각 epoch의 배치별 평균 손실을 다시 평균한 학습 지표다.</p>
<pre><code class="language-text">epoch 1: 0.0645
epoch 2: 0.0304
epoch 3: 0.0210
epoch 4: 0.0172
epoch 5: 0.0151</code></pre>
<p>손실은 감소했다. 실제 복원 이미지도 확인해보자.</p>
<p><img src="https://velog.velcdn.com/images/def_uk/post/393594cf-fa38-4d4e-9232-5a5e8370ed3a/image.png" alt="위: 원본 이미지 · 아래: 잠재 차원 32인 AE의 복원 결과. 학습 데이터에서 뽑은 예시다."></p>
<p><em>위: 원본 이미지 · 아래: 잠재 차원 32인 AE의 복원 결과. 학습 데이터에서 뽑은 예시다.</em></p>
<h3 id="잠재-차원을-784로-늘리면-완벽하게-복사할까">잠재 차원을 784로 늘리면 완벽하게 복사할까?</h3>
<p>현재 모델에서는 잠재 차원만 늘려도 다음 구조가 된다.</p>
<pre><code class="language-text">784 → 128 → 784 → 128 → 784</code></pre>
<p>여전히 128차원 층이 남아 있다. 따라서 잠재 차원이 입력과 같아졌다는 이유만으로 병목이 없어졌다고 말할 수 없다.</p>
<p>중간 층까지 넓히고 구조와 활성화 함수가 허용하면 항등 함수를 표현할 수 있지만, 학습이 실제로 그 해에 도달하는지는 별개의 문제다. 또 처음 보는 입력까지 그대로 반환하는 항등 함수와 훈련 데이터만 외우는 과적합도 구분해야 한다.</p>
<p>핵심은 <strong>복원을 잘하는 것과 다른 작업에도 유용한 표현을 배우는 것은 서로 다른 평가 대상</strong>이라는 점이다.</p>
<h2 id="3-잠재-차원을-줄이면-무엇을-잃을까">3. 잠재 차원을 줄이면 무엇을 잃을까?</h2>
<p>잠재 차원을 <code>2, 4, 8, 16, 32, 64</code>로 바꾸어 비교했다. 모든 모델을 5 epoch 학습하고, 같은 순서의 학습 배치를 사용했다. 평가는 테스트 이미지 10,000장 전체의 픽셀당 평균 MSE로 계산했다.</p>
<pre><code class="language-python"># 모든 잠재 차원을 동일한 테스트 데이터로 평가합니다.
test_set = datasets.MNIST(&quot;./data&quot;, train=False, download=True,
                          transform=transforms.ToTensor())
test_loader = DataLoader(test_set, batch_size=1000, shuffle=False)
latent_dims = [2, 4, 8, 16, 32, 64]


def evaluate_mse(model, loader):
    model.eval()
    squared_error = 0.0
    pixel_count = 0
    with torch.no_grad():
        for images, _ in loader:
            x = torch.flatten(images, start_dim=1).to(device)
            recon = model(x)
            squared_error += F.mse_loss(
                recon, x, reduction=&quot;sum&quot;
            ).item()
            pixel_count += x.numel()
    return squared_error / pixel_count


results = {}

for d in latent_dims:
    model = AutoEncoder(latent_dim=d).to(device)
    criterion = nn.MSELoss()
    optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
    # 모델 크기별 초기화 난수 소비량과 무관하게 같은 배치 순서를 사용합니다.
    comparison_loader = DataLoader(
        train_set, batch_size=256, shuffle=True,
        generator=torch.Generator().manual_seed(SEED)
    )

    for epoch in range(5):
        model.train()
        for images, _ in comparison_loader:
            x = torch.flatten(images, start_dim=1).to(device)
            loss = criterion(model(x), x)
            optimizer.zero_grad()
            loss.backward()
            optimizer.step()

    test_mse = evaluate_mse(model, test_loader)
    results[d] = {&quot;model&quot;: model, &quot;test_mse&quot;: test_mse}
    print(f&quot;latent {d:&gt;2}: test MSE {test_mse:.4f}&quot;)</code></pre>
<table>
<thead>
<tr>
<th align="right">잠재 차원</th>
<th align="right">테스트 MSE</th>
</tr>
</thead>
<tbody><tr>
<td align="right">2</td>
<td align="right">0.0488</td>
</tr>
<tr>
<td align="right">4</td>
<td align="right">0.0357</td>
</tr>
<tr>
<td align="right">8</td>
<td align="right">0.0249</td>
</tr>
<tr>
<td align="right">16</td>
<td align="right">0.0188</td>
</tr>
<tr>
<td align="right">32</td>
<td align="right">0.0135</td>
</tr>
<tr>
<td align="right">64</td>
<td align="right">0.0125</td>
</tr>
</tbody></table>
<p>이번 실행에서는 차원이 커질수록 테스트 MSE가 낮아졌다. 32에서 64로 늘렸을 때도 개선은 있었지만, 그 차이는 0.0010이었다.</p>
<p><img src="https://velog.velcdn.com/images/def_uk/post/994ad51b-d38e-4677-838f-38b7ccfed32d/image.png" alt="같은 테스트 이미지 8장을 잠재 차원별로 복원한 결과."></p>
<p><em>같은 테스트 이미지 8장을 잠재 차원별로 복원한 결과.</em></p>
<p>처음에는 숫자를 알아보려면 64차원 정도는 필요하지 않을까 예상했다. 하지만 그림에서는 2차원에서도 1이나 0처럼 읽을 수 있는 예시가 있고, 8~16차원에서는 더 많은 숫자의 형태가 드러난다. 차원을 더 늘리면 획과 필체도 원본에 가까워지는 모습이 보인다.</p>
<p>다만 두 질문은 나누어 봐야 한다.</p>
<ol>
<li>어떤 숫자인지 사람이 알아볼 수 있는가?</li>
<li>원본의 획 두께, 기울기, 필체까지 보존하는가?</li>
</ol>
<p>MSE는 픽셀 오차를 측정하므로 사람의 판독 가능성과 반드시 일치하지 않는다. 또한 차원을 늘리면 파라미터 수와 학습 양상도 달라진다. 한 번의 초기화와 5 epoch 결과만으로 특정 차원을 MNIST의 “실질적인 정보량”이라고 해석할 수는 없다.</p>
<h2 id="4-2차원-잠재-공간을-직접-들여다보기">4. 2차원 잠재 공간을 직접 들여다보기</h2>
<p>잠재 차원이 2라면 이미지 하나를 평면의 점 하나로 표시할 수 있다. 테스트 이미지를 인코더에 넣어 <code>z1</code>, <code>z2</code>를 얻고, 해석을 위해 숫자 레이블로 색을 칠했다.</p>
<pre><code class="language-python">test_set = datasets.MNIST(&quot;./data&quot;, train=False, download=True,
                          transform=transforms.ToTensor())
test_loader = DataLoader(test_set, batch_size=1000)

model2 = results[2][&quot;model&quot;]
model2.eval()

zs, ys = [], []
with torch.no_grad():
    for images, labels in test_loader:
        x = torch.flatten(images, start_dim=1).to(device)
        zs.append(model2.encoder(x).cpu())
        ys.append(labels)

z = torch.cat(zs); y = torch.cat(ys)

plt.figure(figsize=(8, 7))
sc = plt.scatter(z[:, 0], z[:, 1], c=y, cmap=&quot;tab10&quot;, s=3, alpha=0.6)
plt.colorbar(sc, ticks=range(10))
plt.xlabel(&quot;z1&quot;); plt.ylabel(&quot;z2&quot;)
plt.show()</code></pre>
<p><img src="https://velog.velcdn.com/images/def_uk/post/257aafb6-9832-4980-8453-71d2b7febc4c/image.png" alt="2차원 AE의 잠재 벡터. 색은 학습에 사용하지 않은 숫자 레이블이다."></p>
<p><em>2차원 AE의 잠재 벡터. 색은 학습에 사용하지 않은 숫자 레이블이다.</em></p>
<p>같은 숫자의 이미지들이 일부 가까이 모이지만 서로 다른 숫자가 겹치는 영역도 있다. 복원을 위해 배운 형태적 특징이 숫자 구분과 연결되는 부분이 있다는 정도로 해석할 수 있다. 이 결과만으로 모델이 숫자의 의미를 이해했다고 단정하지는 않는다.</p>
<h3 id="관측된-점이-없는-곳을-디코딩하면">관측된 점이 없는 곳을 디코딩하면?</h3>
<p>잠재 공간에서 한 숫자가 우세한 위치, 레이블이 섞인 위치, 관측점에서 떨어진 위치를 골라 디코더에 넣어보았다.</p>
<p><img src="https://velog.velcdn.com/images/def_uk/post/5a7a55c8-1be6-4265-9369-1035af103f1f/image.png" alt="왼쪽의 X 좌표를 디코딩한 결과가 오른쪽에 표시되어 있다."></p>
<p><em>왼쪽의 X 좌표를 디코딩한 결과가 오른쪽에 표시되어 있다.</em></p>
<p>그림의 표시값은 다음과 같이 읽는다.</p>
<ul>
<li><code>dominant share</code>: 가까운 30개 점에서 가장 많은 숫자의 비율이다. 디코더의 분류 확률이 아니다.</li>
<li><code>mixed labels</code>: 주변 점들의 숫자 레이블이 섞인 후보 위치다.</li>
<li><code>nearby gap</code>: 관측된 테스트 잠재 벡터에서 떨어진 후보 위치다. 학습 데이터까지 전혀 없었던 곳이라는 뜻은 아니다.</li>
</ul>
<p>원본 실습에서는 빈 공간 후보에서도 9와 1처럼 보이는 출력이 관찰되었다. 즉, “AE의 빈 공간에서는 숫자가 나오지 않는다”는 설명은 적절하지 않다.</p>
<p>남는 질문은 따로 있다. <strong>새 이미지를 생성하려면 좌표를 어디서, 어떤 확률로 뽑아야 할까?</strong> 일반 AE의 복원 목표만으로는 그 분포가 정해지지 않는다. 이 질문은 뒤에서 VAE로 이어진다.</p>
<h2 id="5-dae-손상된-입력에서-원본을-추정하기">5. DAE: 손상된 입력에서 원본을 추정하기</h2>
<p>이번에는 입력에 노이즈를 더해보자. 일반 AE와 DAE의 차이는 모델 구조보다 입력과 목표의 관계에 있다.</p>
<table>
<thead>
<tr>
<th>모델</th>
<th>학습 입력</th>
<th>복원 목표</th>
</tr>
</thead>
<tbody><tr>
<td>AE</td>
<td>깨끗한 원본 <code>x</code></td>
<td>깨끗한 원본 <code>x</code></td>
</tr>
<tr>
<td>DAE</td>
<td>노이즈를 더한 입력 <code>x̃</code></td>
<td>깨끗한 원본 <code>x</code></td>
</tr>
</tbody></table>
<p>핵심 코드는 간단하다.</p>
<pre><code class="language-python">def add_noise(x, factor=0.3):
    return torch.clamp(x + factor * torch.randn_like(x), 0., 1.)

loss = criterion(dae(add_noise(x)), x)</code></pre>
<p>입력에만 노이즈가 있고, 손실 함수의 목표는 여전히 깨끗한 원본이다. 모델은 손상된 입력을 그대로 복사하는 대신 원본을 추정해야 한다.</p>
<h3 id="같은-노이즈-입력으로-ae와-dae-비교하기">같은 노이즈 입력으로 AE와 DAE 비교하기</h3>
<p>두 모델은 잠재 차원 16, 같은 초기 가중치, 같은 배치 순서, 같은 학습 횟수인 50 epoch를 사용했다. 평가 시에는 두 모델 모두 동일한 노이즈 테스트 이미지를 입력받았다.</p>
<table>
<thead>
<tr>
<th>평가 대상</th>
<th align="right">원본 대비 테스트 MSE</th>
</tr>
</thead>
<tbody><tr>
<td>노이즈 입력 자체</td>
<td align="right">0.0467</td>
</tr>
<tr>
<td>일반 AE의 출력</td>
<td align="right">0.0514</td>
</tr>
<tr>
<td>DAE의 출력</td>
<td align="right">0.0146</td>
</tr>
</tbody></table>
<p>세 값 모두 테스트 이미지 10,000장 전체의 픽셀당 평균 제곱 오차다.</p>
<p><img src="https://velog.velcdn.com/images/def_uk/post/6d5d85da-82b2-47a5-ab0d-8dd110bea417/image.png" alt="위에서부터 원본, 노이즈 입력, 일반 AE 출력, DAE 출력."></p>
<p><em>위에서부터 원본, 노이즈 입력, 일반 AE 출력, DAE 출력.</em></p>
<p>이번 결과에서 일반 AE는 노이즈 입력 자체보다 MSE가 높았다. 깨끗한 입력으로 복원을 학습했다고 해서 손상된 입력에서도 잘 복원하는 것은 아니라는 사례다.</p>
<p>반면 DAE는 노이즈를 줄이고 숫자 형태를 더 잘 복원했다. 다만 획과 필체까지 완벽하게 보존하지는 않았다. <strong>노이즈를 얼마나 없앴는지와 원본의 특징을 얼마나 남겼는지를 함께 확인해야 한다.</strong></p>
<p>이 비교는 현재 노이즈 종류와 세기에 대한 결과다. 모든 손상에 같은 성능을 보인다는 뜻은 아니다.</p>
<h2 id="6-vae-새-좌표를-뽑을-분포까지-정하기">6. VAE: 새 좌표를 뽑을 분포까지 정하기</h2>
<p>AE에서는 입력을 잠재 벡터 하나로 바꾸었다. 이번 VAE에서는 입력마다 잠재 벡터의 <strong>평균과 로그 분산</strong>을 계산한다.</p>
<pre><code class="language-text">입력 이미지 → 인코더 → mu, logvar → 잠재 벡터 z 샘플링 → 디코더 → 복원</code></pre>
<p>이미지별 잠재 분포를 이용해 원본을 복원하면서, 그 분포가 기준 분포인 표준정규분포에서 너무 멀어지지 않도록 학습한다.</p>
<p>$$
q_\phi(z\mid x)=\mathcal{N}\left(\mu_\phi(x),\operatorname{diag}(\sigma_\phi^2(x))\right),\qquad p(z)=\mathcal{N}(0,I)
$$</p>
<p>분포는 숫자 클래스마다 하나씩 만들어지는 것이 아니라 <strong>입력 이미지마다</strong> 만들어진다. 생성할 때는 입력 이미지 없이 기준 분포에서 <code>z</code>를 뽑아 디코더에 넣는다.</p>
<h3 id="reparameterization-trick">Reparameterization trick</h3>
<p>인코더가 출력한 <code>logvar</code>는 로그 분산이다. 이를 표준편차로 바꾼 뒤 표준정규분포 난수와 결합한다.</p>
<p>$$
\sigma=\exp\left(\frac{1}{2}\operatorname{logvar}\right),\qquad
\epsilon\sim\mathcal{N}(0,I),\qquad
z=\mu+\sigma\odot\epsilon
$$</p>
<p>난수는 <code>eps</code>에서 뽑고, <code>mu</code>와 <code>std</code>에는 미분 가능한 연산을 적용한다. 이 경로를 통해 복원 손실의 기울기가 인코더까지 전달된다.</p>
<pre><code class="language-python">class VAE(nn.Module):
    def __init__(self, latent_dim=2):
        super().__init__()
        self.fc = nn.Sequential(nn.Linear(784, 128), nn.ReLU())
        # 이미지마다 잠재 좌표의 중심 mu를 계산
        # latent_dim=2라면 [mu1, mu2]처럼 숫자 2개를 출력
        self.fc_mu = nn.Linear(128, latent_dim)
        # 로그 연산을 하는 층이 아니라, 출력값을 로그 분산으로 사용하는 층
        # 아래 표준편차 계산과 KL 손실에서 이 의미로 사용함
        self.fc_logvar = nn.Linear(128, latent_dim)
        self.decoder = nn.Sequential(
            nn.Linear(latent_dim, 128), nn.ReLU(),
            nn.Linear(128, 784), nn.Sigmoid()
        )

    def encode(self, x):
        h = self.fc(x)
        return self.fc_mu(h), self.fc_logvar(h)

    def reparameterize(self, mu, logvar):
        # 로그 분산을 표준편차로 변환: std = exp(logvar / 2)
        std = torch.exp(0.5 * logvar)
        # 평균 0, 표준편차 1인 정규분포 난수
        # 각 좌표에서 중심으로부터 표준편차의 몇 배만큼 이동할지 정함
        eps = torch.randn_like(std)
        # 중심 + 이동량으로 잠재 좌표 계산
        return mu + eps * std

    def forward(self, x):
        mu, logvar = self.encode(x)
        z = self.reparameterize(mu, logvar)
        return self.decoder(z), mu, logvar</code></pre>
<h3 id="복원-손실과-kl-항">복원 손실과 KL 항</h3>
<p>이번 실습의 손실은 복원 오차와 KL 항을 더한 것이다.</p>
<p>$$
\mathcal{L}=\mathcal{L}<em>{\mathrm{reconstruction}}+D</em>{\mathrm{KL}}\left(q_\phi(z\mid x)\Vert p(z)\right)
$$</p>
<p>복원 항은 원본을 잘 되살리도록 하고, KL 항은 이미지별 분포의 평균과 분산이 기준 분포에 가까워지도록 작용한다.</p>
<pre><code class="language-python"># reconstruction: 원본과 복원 이미지의 픽셀별 제곱 오차 합
# KL: 이미지별 잠재 분포가 기준 분포에서 벗어난 정도
# 로그에서는 두 항과 합계를 각각 이미지 수로 나눕니다.
# reconstruction은 픽셀 평균 MSE가 아니라 이미지당 픽셀 오차 합입니다.
def vae_loss(recon, x, mu, logvar, return_components=False):
    recon_loss = F.mse_loss(recon, x, reduction=&quot;sum&quot;)
    kl = -0.5 * torch.sum(1 + logvar - mu.pow(2) - logvar.exp())
    total = recon_loss + kl
    if return_components:
        return total, recon_loss, kl
    return total</code></pre>
<p>여기서 복원 항은 <code>reduction=&quot;sum&quot;</code>을 사용한다. 로그를 출력할 때는 이미지 수로 나누므로, 앞에서 사용한 <strong>픽셀당 평균 MSE와 단위가 다르다.</strong> 복원 항과 KL 항의 합산 방식은 두 목표의 상대적인 크기에도 영향을 준다.</p>
<p>KL 항을 빼면 어떻게 될까? 모델은 복원 오차만 줄이면 되므로 표준편차를 작게 만들어 난수의 영향을 줄이는 방향으로 학습될 수 있다. 또한 이미지별 분포를 기준 분포에 맞출 이유가 사라져, 생성 시 표준정규분포에서 뽑은 좌표가 좋은 출력을 만든다고 기대하기 어려워진다.</p>
<p>그렇다고 KL 항이 모든 빈 공간을 채우거나 이미지를 반드시 선명하게 만드는 것은 아니다. 복원 목표와 분포 제약 사이에 절충이 생긴다.</p>
<h3 id="2차원-vae-학습하기">2차원 VAE 학습하기</h3>
<pre><code class="language-python">vae = VAE(latent_dim=2).to(device)
optimizer = torch.optim.Adam(vae.parameters(), lr=1e-3)

for epoch in range(20):
    vae.train()
    recon_total, kl_total, sample_count = 0.0, 0.0, 0
    for images, _ in train_loader:
        x = images.view(images.size(0), -1).to(device)
        recon, mu, logvar = vae(x)
        loss, recon_term, kl_term = vae_loss(
            recon, x, mu, logvar, return_components=True
        )
        optimizer.zero_grad(); loss.backward(); optimizer.step()
        recon_total += recon_term.item()
        kl_total += kl_term.item()
        sample_count += x.size(0)
    if (epoch+1) % 5 == 0:
        print(f&quot;epoch {epoch+1} | per image: &quot;
              f&quot;reconstruction {recon_total/sample_count:.4f} | &quot;
              f&quot;KL {kl_total/sample_count:.4f} | &quot;
              f&quot;total {(recon_total+kl_total)/sample_count:.4f}&quot;)</code></pre>
<table>
<thead>
<tr>
<th align="right">Epoch</th>
<th align="right">복원 항 / 이미지</th>
<th align="right">KL / 이미지</th>
<th align="right">합계 / 이미지</th>
</tr>
</thead>
<tbody><tr>
<td align="right">5</td>
<td align="right">39.5887</td>
<td align="right">3.2838</td>
<td align="right">42.8726</td>
</tr>
<tr>
<td align="right">10</td>
<td align="right">37.2238</td>
<td align="right">3.6590</td>
<td align="right">40.8828</td>
</tr>
<tr>
<td align="right">15</td>
<td align="right">36.0197</td>
<td align="right">3.8960</td>
<td align="right">39.9157</td>
</tr>
<tr>
<td align="right">20</td>
<td align="right">35.2477</td>
<td align="right">4.0587</td>
<td align="right">39.3064</td>
</tr>
</tbody></table>
<p>학습 중 복원 항은 감소하고 KL 항은 증가했다. 합계는 감소했으므로 두 항을 함께 최적화하는 과정에서 나타난 변화로 읽을 수 있다. 각 항이 매번 동시에 감소해야 하는 것은 아니다.</p>
<h2 id="7-입력-이미지-없이-생성하기">7. 입력 이미지 없이 생성하기</h2>
<p>학습이 끝나면 인코더를 거치지 않고 표준정규분포에서 잠재 벡터를 뽑는다.</p>
<pre><code class="language-python"># 입력 이미지와 인코더 없이, 기준 분포 N(0, I)에서 생성합니다.
vae.eval()
with torch.no_grad():
    z_new = torch.randn(16, vae.fc_mu.out_features, device=device)
    generated = vae.decoder(z_new).cpu().view(16, 28, 28)

fig, axes = plt.subplots(4, 4, figsize=(6, 6))
for i, ax in enumerate(axes.flat):
    ax.imshow(generated[i], cmap=&quot;gray&quot;, vmin=0, vmax=1)
    ax.axis(&quot;off&quot;)
fig.suptitle(&quot;VAE: samples from N(0, I)&quot;)
plt.tight_layout()
plt.show()</code></pre>
<p><img src="https://velog.velcdn.com/images/def_uk/post/2313ee87-0661-4670-b72b-85ac48fb8ace/image.png" alt="2차원 VAE의 기준 분포에서 잠재 벡터 16개를 뽑아 생성한 이미지."></p>
<p><em>2차원 VAE의 기준 분포에서 잠재 벡터 16개를 뽑아 생성한 이미지.</em></p>
<p>이 결과는 주어진 원본의 복원이 아니다. 새로운 잠재 좌표를 디코더에 넣은 생성 결과다. 숫자 형태가 나타나지만 모든 결과가 선명하거나 명확한 것은 아니다.</p>
<h3 id="잠재-차원을-16으로-늘리면">잠재 차원을 16으로 늘리면?</h3>
<p>같은 데이터, 배치 크기, 학습률, 손실 함수로 16차원 VAE를 별도로 20 epoch 학습했다.</p>
<p><img src="https://velog.velcdn.com/images/def_uk/post/fc852b3a-03d5-440c-96cb-26a41d1837c6/image.png" alt="왼쪽: 잠재 차원 2 · 오른쪽: 잠재 차원 16. 각 모델의 잠재 공간은 서로 다르다."></p>
<p><em>왼쪽: 잠재 차원 2 · 오른쪽: 잠재 차원 16. 각 모델의 잠재 공간은 서로 다르다.</em></p>
<p>이 예시에서는 오른쪽에 경계가 더 뚜렷한 숫자가 여러 개 보인다. 하지만 같은 칸에 같은 숫자가 나와야 하는 비교는 아니다. 두 모델의 잠재 좌표는 서로 대응하는 의미를 갖지 않는다.</p>
<p>20 epoch의 학습 로그도 비교하면 다음과 같다.</p>
<table>
<thead>
<tr>
<th align="right">잠재 차원</th>
<th align="right">복원 항 / 이미지</th>
<th align="right">KL / 이미지</th>
<th align="right">합계 / 이미지</th>
</tr>
</thead>
<tbody><tr>
<td align="right">2</td>
<td align="right">35.2477</td>
<td align="right">4.0587</td>
<td align="right">39.3064</td>
</tr>
<tr>
<td align="right">16</td>
<td align="right">20.1170</td>
<td align="right">11.2927</td>
<td align="right">31.4097</td>
</tr>
</tbody></table>
<p>16차원 모델에서 복원 항은 작고 KL 항은 컸다. 차원이 늘면 파라미터 수와 KL을 합산하는 차원 수도 달라지고, 이번 두 모델의 초기화와 학습 순서도 같지 않다. 따라서 차원을 늘리면 반드시 생성 품질이 좋아진다는 결론 대신, 현재 설정에서의 탐색적 비교로 받아들이는 것이 적절하다.</p>
<h3 id="잠재-좌표를-조금씩-이동해보기">잠재 좌표를 조금씩 이동해보기</h3>
<p><img src="https://velog.velcdn.com/images/def_uk/post/3f3a1235-cd60-4dc0-bb8d-98dc39930acc/image.png" alt="왼쪽: AE의 관측 잠재 범위 · 오른쪽: VAE의 각 축 -2.5~2.5 범위. 균일 간격의 좌표 격자다."></p>
<p><em>왼쪽: AE의 관측 잠재 범위 · 오른쪽: VAE의 각 축 -2.5~2.5 범위. 균일 간격의 좌표 격자다.</em></p>
<p>가로로 이동하면 <code>z1</code>, 세로로 이동하면 <code>z2</code>가 변한다. 이웃한 좌표에서 획과 숫자 형태가 어떻게 바뀌는지 볼 수 있다.</p>
<p>이 그림은 <strong>균일한 간격으로 좌표를 훑은 결과</strong>이며, 정규분포에서 무작위로 뽑은 결과가 아니다. AE와 VAE의 탐색 범위, 좌표 의미, 학습 횟수도 다르므로 이 그림 하나로 생성 성능의 우열을 판단하기는 어렵다.</p>
<h2 id="8-정리-어떤-제약을-주느냐가-학습을-바꾼다">8. 정리: 어떤 제약을 주느냐가 학습을 바꾼다</h2>
<p>처음 질문은 “입력을 그대로 출력할 거라면 왜 학습할까?”였다. 실험을 따라가며 중요한 것은 어떤 경로로 복원하게 만드는지라는 점을 확인했다.</p>
<table>
<thead>
<tr>
<th>모델</th>
<th>이번 실습에서 준 제약</th>
<th>확인한 점</th>
</tr>
</thead>
<tbody><tr>
<td>AE</td>
<td>좁은 잠재 공간을 거쳐 복원</td>
<td>차원에 따라 복원 오차와 보존되는 형태가 달라졌다.</td>
</tr>
<tr>
<td>DAE</td>
<td>손상된 입력으로 깨끗한 원본 추정</td>
<td>해당 노이즈 조건에서 일반 AE보다 낮은 복원 오차를 보였다.</td>
</tr>
<tr>
<td>VAE</td>
<td>이미지별 잠재 분포를 기준 분포에 가깝게 유도</td>
<td>기준 분포에서 좌표를 뽑아 입력 없이 이미지를 생성했다.</td>
</tr>
</tbody></table>
<p>이 과정에서 구분해야 할 것도 남았다.</p>
<ul>
<li>낮은 복원 손실만으로 잠재 표현의 유용성을 판단할 수 없다.</li>
<li>노이즈 제거 성능과 원본의 세부 특징 보존은 함께 확인해야 한다.</li>
<li>몇 개의 그럴듯한 생성 이미지와 안정적인 생성 성능은 다르다.</li>
</ul>
<p>이번 비교는 저장된 단일 실행 결과를 바탕으로 한다. 테스트 결과를 반복해서 보며 차원이나 학습 횟수를 선택하려면 별도의 검증 세트를 두어야 한다.</p>
<p>오토인코더의 흥미로운 점은 “복원”이라는 비슷한 목표에서도 병목, 손상된 입력, 확률분포라는 조건을 어떻게 주느냐에 따라 배우는 표현과 활용 방식이 달라진다는 것이다.</p>
<hr>
<h2 id="부록-비교-실험과-시각화-코드">부록: 비교 실험과 시각화 코드</h2>
<p>아래 코드는 본문에서 정의한 변수와 모델을 이어서 사용한다. DAE 학습과 16차원 VAE 학습은 환경에 따라 시간이 걸릴 수 있다.</p>
<details>
<summary>AE 복원 그림</summary>

<pre><code class="language-python">model.eval()
with torch.no_grad():
    images, _ = next(iter(train_loader))
    x = torch.flatten(images, start_dim=1).to(device)
    recon = model(x).cpu().view(-1, 28, 28)

fig, axes = plt.subplots(2, 8, figsize=(12, 3))
for i in range(8):
    # 원본
    axes[0, i].imshow(images[i].squeeze(), cmap=&quot;gray&quot;); axes[0, i].axis(&quot;off&quot;)
    # 복원  
    axes[1, i].imshow(recon[i], cmap=&quot;gray&quot;); axes[1, i].axis(&quot;off&quot;)
plt.show()</code></pre>
</details>

<details>
<summary>잠재 차원별 복원 그림</summary>

<pre><code class="language-python"># 섞지 않은 테스트 데이터의 첫 8장을 모든 모델에 똑같이 사용합니다.
images, _ = next(iter(test_loader))
images = images[:8]
x = torch.flatten(images, start_dim=1).to(device)

fig, axes = plt.subplots(len(latent_dims) + 1, 8, figsize=(12, 11))
for i in range(8):
    axes[0, i].imshow(images[i].squeeze(), cmap=&quot;gray&quot;, vmin=0, vmax=1)
    axes[0, i].axis(&quot;off&quot;)
axes[0, 0].set_title(&quot;original&quot;, loc=&quot;left&quot;)

for row, d in enumerate(latent_dims, start=1):
    model_d = results[d][&quot;model&quot;]
    model_d.eval()
    with torch.no_grad():
        recon = model_d(x).cpu().view(-1, 28, 28)
    for i in range(8):
        axes[row, i].imshow(recon[i], cmap=&quot;gray&quot;, vmin=0, vmax=1)
        axes[row, i].axis(&quot;off&quot;)
    axes[row, 0].set_title(f&quot;latent {d}&quot;, loc=&quot;left&quot;)
plt.tight_layout()
plt.show()</code></pre>
</details>

<details>
<summary>AE 잠재 좌표 선택과 디코딩</summary>

<pre><code class="language-python">model2 = results[2][&quot;model&quot;]
model2.eval()


def select_latent_points(z, y, neighbors=30):
    # CPU에서 현재 산점도의 이웃을 조사합니다. 레이블은 위치 해석에만 씁니다.
    coords = z.detach().cpu().float()
    labels = y.detach().cpu().long()
    k = min(neighbors, len(coords))
    candidates = coords[torch.linspace(0, len(coords) - 1,
                                       min(600, len(coords))).long()]
    distances, indices = torch.cdist(candidates, coords, compute_mode=&quot;donot_use_mm_for_euclid_dist&quot;).topk(k, largest=False)
    proportions = F.one_hot(labels[indices], num_classes=10).float().mean(1)
    purity, dominant = proportions.max(1)
    entropy = -(proportions * proportions.clamp_min(1e-12).log()).sum(1)
    radius = distances[:, -1]
    dense = radius &lt;= torch.quantile(radius, 0.6)
    typical_radius = radius.median().clamp_min(1e-6)
    separation = (coords.max(0).values - coords.min(0).values).norm() * 0.12

    def pick_spaced(order, pool, count=2):
        chosen = []
        for index in order.tolist():
            if not chosen or torch.cdist(pool[index:index+1], pool[chosen]).min() &gt;= separation:
                chosen.append(index)
            if len(chosen) == count:
                break
        # 분포가 좁아도 같은 후보를 중복 선택하지 않습니다.
        for index in order.tolist():
            if len(chosen) == count:
                break
            if index not in chosen:
                chosen.append(index)
        return chosen

    # 밀집한 후보에서 서로 다른 숫자가 우세한 위치를 고릅니다.
    dense_indices = torch.where(dense)[0]
    pure_order = dense_indices[torch.argsort(purity[dense_indices], descending=True)]
    first = pure_order[0].item()
    other_classes = pure_order[dominant[pure_order] != dominant[first]]
    second = other_classes[0].item() if len(other_classes) else pure_order[1].item()
    pure_indices = [first, second]

    # 밀집한 후보 중 주변 레이블이 많이 섞인 위치를 고릅니다.
    mixed_order = dense_indices[torch.argsort(entropy[dense_indices], descending=True)]
    mixed_order = mixed_order[(mixed_order != first) &amp; (mixed_order != second)]
    high_entropy = mixed_order[entropy[mixed_order] &gt;= 0.5 * entropy[mixed_order].max()]
    if len(high_entropy) &gt;= 2:
        mixed_order = high_entropy
    mixed_indices = pick_spaced(mixed_order, candidates)

    # 관측 범위보다 조금 넓은 격자에서 실제 점과 적당히 떨어진 위치를 찾습니다.
    lower, upper = coords.min(0).values, coords.max(0).values
    padding = (upper - lower).clamp_min(1e-6) * 0.05
    gx = torch.linspace(lower[0] - padding[0], upper[0] + padding[0], 35)
    gy = torch.linspace(lower[1] - padding[1], upper[1] + padding[1], 35)
    xx, yy = torch.meshgrid(gx, gy, indexing=&quot;ij&quot;)
    grid = torch.stack([xx.flatten(), yy.flatten()], dim=1)
    nearest = torch.cat([
        torch.cdist(chunk, coords, compute_mode=&quot;donot_use_mm_for_euclid_dist&quot;).min(1).values
        for chunk in grid.split(128)
    ])
    # 가장 가까운 실제 점까지의 거리가 대표적인 이웃 반경의 1.5배에 가까운 후보를 고릅니다.
    gap_order = torch.argsort((nearest - 1.5 * typical_radius).abs())
    gap_indices = pick_spaced(gap_order, grid)

    points = torch.cat([candidates[pure_indices], candidates[mixed_indices], grid[gap_indices]])
    names = [f&quot;dominant {dominant[j].item()}&quot; for j in pure_indices]
    names += [&quot;mixed labels&quot; if entropy[j] &gt; 0.1 else &quot;low mixing&quot; for j in mixed_indices]
    names += [&quot;nearby gap&quot;, &quot;nearby gap&quot;]
    notes = [f&quot;dominant share: {purity[j]:.0%}&quot; for j in pure_indices + mixed_indices]
    notes += [f&quot;nearest sample: {nearest[j]:.2f}&quot; for j in gap_indices]
    return points, names, notes


points, point_names, point_notes = select_latent_points(z, y)
with torch.no_grad():
    images = model2.decoder(points.to(device)).cpu().view(-1, 28, 28).numpy()

fig = plt.figure(figsize=(14, 8))
gs = fig.add_gridspec(3, 4, width_ratios=[1, 1, 0.85, 0.85])
ax = fig.add_subplot(gs[:, :2])
sc = ax.scatter(z[:, 0], z[:, 1], c=y, cmap=&quot;tab10&quot;, s=3, alpha=0.6,
                vmin=-0.5, vmax=9.5)
fig.colorbar(sc, ax=ax, ticks=range(10), fraction=0.04)
for i, p in enumerate(points):
    ax.scatter(*p.tolist(), s=160, marker=&quot;X&quot;, c=&quot;white&quot;, edgecolors=&quot;black&quot;, zorder=5)
    ax.annotate(str(i), p.tolist(), xytext=(6, 6), textcoords=&quot;offset points&quot;,
                bbox=dict(boxstyle=&quot;round,pad=0.2&quot;, fc=&quot;white&quot;, alpha=0.8))
ax.set(xlabel=&quot;z1&quot;, ylabel=&quot;z2&quot;, title=&quot;Points selected from the current latent space&quot;)
for i in range(6):
    row, col = divmod(i, 2)
    im_ax = fig.add_subplot(gs[row, col + 2])
    im_ax.imshow(images[i], cmap=&quot;gray&quot;, vmin=0, vmax=1)
    px, py = points[i].tolist()
    im_ax.set_title(f&quot;{i}: {point_names[i]} ({px:.1f}, {py:.1f})\n{point_notes[i]}&quot;, fontsize=9)
    im_ax.axis(&quot;off&quot;)
plt.tight_layout()
plt.show()</code></pre>
</details>

<details>
<summary>AE·DAE 동일 조건 학습과 평가</summary>

<pre><code class="language-python"># 입력에만 노이즈를 더하고, 복원 목표는 원본으로 둡니다.
def add_noise(x, factor=0.3):
    return torch.clamp(x + factor * torch.randn_like(x), 0., 1.)


def evaluate_denoising(ae, dae, loader, factor, seed):
    ae.eval()
    dae.eval()
    # 평가할 때마다 같은 난수열로 각 이미지에 동일한 노이즈를 더합니다.
    noise_generator = torch.Generator().manual_seed(seed)
    totals = {&quot;noisy input&quot;: 0.0, &quot;AE&quot;: 0.0, &quot;DAE&quot;: 0.0}
    pixels = 0
    with torch.no_grad():
        for images, _ in loader:
            clean_cpu = torch.flatten(images, start_dim=1)
            noise = torch.randn(clean_cpu.shape, generator=noise_generator)
            noisy = (clean_cpu + factor * noise).clamp(0, 1).to(device)
            clean = clean_cpu.to(device)
            outputs = {&quot;noisy input&quot;: noisy, &quot;AE&quot;: ae(noisy), &quot;DAE&quot;: dae(noisy)}
            for name, output in outputs.items():
                totals[name] += F.mse_loss(output, clean, reduction=&quot;sum&quot;).item()
            pixels += clean.numel()
    return {name: total / pixels for name, total in totals.items()}


latent_dim = 16
epochs = 50
noise_factor = 0.3
torch.manual_seed(SEED)
ae_baseline = AutoEncoder(latent_dim=latent_dim).to(device)
dae = AutoEncoder(latent_dim=latent_dim).to(device)
# 동일한 초기 가중치에서 시작해 입력에 노이즈를 더하는 효과를 비교합니다.
dae.load_state_dict(ae_baseline.state_dict())
ae_optimizer = torch.optim.Adam(ae_baseline.parameters(), lr=1e-3)
dae_optimizer = torch.optim.Adam(dae.parameters(), lr=1e-3)
criterion = nn.MSELoss()
paired_loader = DataLoader(
    train_set, batch_size=256, shuffle=True,
    generator=torch.Generator().manual_seed(SEED)
)
test_loader_eval = DataLoader(test_set, batch_size=1000, shuffle=False)
train_hist, test_hist = [], []
ae_train_hist, ae_test_hist = [], []

for epoch in range(epochs):
    ae_baseline.train()
    dae.train()
    ae_total, dae_total, train_pixels = 0.0, 0.0, 0
    for images, _ in paired_loader:
        x = torch.flatten(images, start_dim=1).to(device)
        # 일반 AE: 원본 → 원본
        ae_loss = criterion(ae_baseline(x), x)
        ae_optimizer.zero_grad()
        ae_loss.backward()
        ae_optimizer.step()
        # DAE: 노이즈 입력 → 원본. 같은 미니배치로 같은 횟수만큼 업데이트합니다.
        dae_loss = criterion(dae(add_noise(x, noise_factor)), x)
        dae_optimizer.zero_grad()
        dae_loss.backward()
        dae_optimizer.step()
        ae_total += ae_loss.item() * x.numel()
        dae_total += dae_loss.item() * x.numel()
        train_pixels += x.numel()

    metrics = evaluate_denoising(ae_baseline, dae, test_loader_eval, noise_factor, SEED + 1)
    ae_train_hist.append(ae_total / train_pixels)
    train_hist.append(dae_total / train_pixels)
    ae_test_hist.append(metrics[&quot;AE&quot;])
    test_hist.append(metrics[&quot;DAE&quot;])
    if (epoch + 1) % 10 == 0 or epoch + 1 == epochs:
        print(f&quot;epoch {epoch+1}: fixed noisy test MSE | &quot;
              f&quot;input {metrics[&#39;noisy input&#39;]:.4f} | &quot;
              f&quot;AE {metrics[&#39;AE&#39;]:.4f} | DAE {metrics[&#39;DAE&#39;]:.4f}&quot;)</code></pre>
<pre><code class="language-python"># 평가와 동일한 테스트 이미지와 노이즈에서 첫 8장을 표시합니다.
images, _ = next(iter(test_loader_eval))
clean_cpu = torch.flatten(images, start_dim=1)
noise_generator = torch.Generator().manual_seed(SEED + 1)
noise = torch.randn(clean_cpu.shape, generator=noise_generator)
noisy_cpu = (clean_cpu + noise_factor * noise).clamp(0, 1)
x = clean_cpu[:8].to(device)
noisy = noisy_cpu[:8].to(device)

ae_baseline.eval()
dae.eval()
with torch.no_grad():
    ae_out = ae_baseline(noisy).cpu().view(-1, 28, 28)
    out = dae(noisy).cpu().view(-1, 28, 28)

rows = [images[:8].squeeze(1), noisy.cpu().view(-1, 28, 28), ae_out, out]
names = [&quot;original&quot;, &quot;noisy input&quot;, &quot;AE reconstruction&quot;, &quot;DAE reconstruction&quot;]
fig, axes = plt.subplots(4, 8, figsize=(12, 7))
for row, (batch, name) in enumerate(zip(rows, names)):
    for i in range(8):
        axes[row, i].imshow(batch[i], cmap=&quot;gray&quot;, vmin=0, vmax=1)
        axes[row, i].axis(&quot;off&quot;)
    axes[row, 0].set_title(name, loc=&quot;left&quot;, fontsize=9)
plt.tight_layout()
plt.show()

# 표시한 8장만이 아니라 테스트 세트 전체의 평균 MSE입니다.
comparison_mse = evaluate_denoising(
    ae_baseline, dae, test_loader_eval, noise_factor, SEED + 1
)
print(f&quot;Test samples: {len(test_set)} | latent dim: {latent_dim} | epochs: {epochs}&quot;)
for name, mse in comparison_mse.items():
    print(f&quot;{name:&gt;11}: MSE {mse:.4f}&quot;)</code></pre>
</details>

<details>
<summary>16차원 VAE 학습과 생성 비교</summary>

<pre><code class="language-python"># 기존 vae(2차원)는 유지하고 16차원 모델을 별도로 학습합니다.
torch.manual_seed(SEED)
vae16 = VAE(latent_dim=16).to(device)
optimizer16 = torch.optim.Adam(vae16.parameters(), lr=1e-3)
comparison_epochs = 20  # 위 2차원 VAE와 같은 학습 횟수
vae16_loader = DataLoader(
    train_set, batch_size=train_loader.batch_size, shuffle=True,
    generator=torch.Generator().manual_seed(SEED)
)

for epoch in range(comparison_epochs):
    vae16.train()
    recon_total16, kl_total16, sample_count16 = 0.0, 0.0, 0
    for images, _ in vae16_loader:
        x = torch.flatten(images, start_dim=1).to(device)
        recon, mu16, logvar16 = vae16(x)
        loss16, recon_term16, kl_term16 = vae_loss(
            recon, x, mu16, logvar16, return_components=True
        )
        optimizer16.zero_grad()
        loss16.backward()
        optimizer16.step()
        recon_total16 += recon_term16.item()
        kl_total16 += kl_term16.item()
        sample_count16 += x.size(0)
    if (epoch + 1) % 5 == 0 or epoch + 1 == comparison_epochs:
        print(f&quot;latent 16 | epoch {epoch+1} | per image: &quot;
              f&quot;reconstruction {recon_total16/sample_count16:.4f} | &quot;
              f&quot;KL {kl_total16/sample_count16:.4f} | &quot;
              f&quot;total {(recon_total16+kl_total16)/sample_count16:.4f}&quot;)

# 고정된 난수로 생성 예시를 비교합니다. 각 모델의 z는 서로 다른 공간입니다.
sample_generator = torch.Generator().manual_seed(SEED + 2)
z16_compare = torch.randn(16, 16, generator=sample_generator)
z2_compare = z16_compare[:, :2]
vae.eval()
vae16.eval()
with torch.no_grad():
    samples2 = vae.decoder(z2_compare.to(device)).cpu().view(16, 28, 28)
    samples16 = vae16.decoder(z16_compare.to(device)).cpu().view(16, 28, 28)

fig, axes = plt.subplots(4, 8, figsize=(12, 6))
for i in range(16):
    row, col = divmod(i, 4)
    for offset, samples in [(0, samples2), (4, samples16)]:
        ax = axes[row, col + offset]
        ax.imshow(samples[i], cmap=&quot;gray&quot;, vmin=0, vmax=1)
        ax.axis(&quot;off&quot;)
fig.suptitle(&quot;VAE generation: latent 2 (left) / latent 16 (right)&quot;)
plt.tight_layout()
plt.show()</code></pre>
</details>

<details>
<summary>VAE의 평균 좌표와 숫자별 중심 디코딩</summary>

<pre><code class="language-python">vae.eval()
zs, ys = [], []
with torch.no_grad():
    for imgs_b, labels in test_loader:

        xb = imgs_b.view(imgs_b.size(0), -1).to(device)
        mu, _ = vae.encode(xb)
        zs.append(mu.cpu()); ys.append(labels)
z_vae = torch.cat(zs); y_vae = torch.cat(ys)

print(&quot;std:&quot;, z_vae.std(dim=0).tolist())

# 1) 좌표를 먼저 정한다 — 숫자별 군집 중심
targets = [0, 1, 3, 7, 8, 9]
points, labels_txt = [], []
for d in targets:
    c = z_vae[y_vae == d].mean(dim=0)
    points.append([c[0].item(), c[1].item()])
    labels_txt.append(f&quot;center of {d}&quot;)

# 2) 그 좌표로 이미지를 만든다
with torch.no_grad():
    pts = torch.tensor(points, dtype=torch.float32).to(device)
    gen = vae.decoder(pts).cpu().view(-1, 28, 28).numpy()

# 3) 그린다
fig = plt.figure(figsize=(13, 7))
ax = fig.add_axes([0.05, 0.08, 0.55, 0.85])

sc = ax.scatter(z_vae[:, 0], z_vae[:, 1], c=y_vae, cmap=&quot;tab10&quot;, s=3, alpha=0.6)
fig.colorbar(sc, ax=ax, ticks=range(10), fraction=0.04)

for i, p in enumerate(points):
    ax.scatter(p[0], p[1], s=200, marker=&quot;X&quot;,
               c=&quot;white&quot;, edgecolors=&quot;black&quot;, linewidths=1.5, zorder=5)
    ax.annotate(labels_txt[i], p, xytext=(6, 6), textcoords=&quot;offset points&quot;,
                fontsize=9, bbox=dict(boxstyle=&quot;round,pad=0.3&quot;, fc=&quot;white&quot;, alpha=0.8))

ax.set_xlabel(&quot;z1&quot;); ax.set_ylabel(&quot;z2&quot;); ax.set_title(&quot;VAE latent space&quot;)

for i in range(6):
    row, col = divmod(i, 2)
    im_ax = fig.add_axes([0.68 + col * 0.15, 0.68 - row * 0.28, 0.13, 0.22])
    im_ax.imshow(gen[i], cmap=&quot;gray&quot;)
    im_ax.set_title(f&quot;{labels_txt[i]} ({points[i][0]:.1f},{points[i][1]:.1f})&quot;,
                    fontsize=8)
    im_ax.axis(&quot;off&quot;)

plt.show()</code></pre>
<p><img src="https://velog.velcdn.com/images/def_uk/post/f52f50d1-573f-416b-8602-7f91c4e7501f/image.png" alt=""></p>
</details>

<details>
<summary>AE·VAE 잠재 격자 시각화</summary>

<pre><code class="language-python">import numpy as np


def latent_grid(model_decoder, x_range, y_range, n=15, device=device):
    grid_x = np.linspace(*x_range, n)
    grid_y = np.linspace(*y_range, n)
    canvas = np.zeros((28*n, 28*n))
    with torch.no_grad():
        for i, yi in enumerate(grid_y[::-1]):
            pts = torch.tensor([[xj, yi] for xj in grid_x],
                               dtype=torch.float32, device=device)
            imgs = model_decoder(pts).cpu().view(-1, 28, 28).numpy()
            for j in range(n):
                canvas[i*28:(i+1)*28, j*28:(j+1)*28] = imgs[j]
    return canvas


ae2 = results[2][&quot;model&quot;]
ae2.eval()
vae.eval()
# 다른 셀의 z를 재사용하지 않고 현재 AE로 테스트 잠재 좌표를 계산합니다.
ae_grid_codes = []
with torch.no_grad():
    for images, _ in test_loader:
        x = torch.flatten(images, start_dim=1).to(device)
        ae_grid_codes.append(ae2.encoder(x).cpu())
ae_grid_codes = torch.cat(ae_grid_codes)
ae_lower = ae_grid_codes.min(dim=0).values
ae_upper = ae_grid_codes.max(dim=0).values
ae_x_range = (ae_lower[0].item(), ae_upper[0].item())
ae_y_range = (ae_lower[1].item(), ae_upper[1].item())

# VAE는 관측된 mu 범위가 아니라 기준 분포 N(0,I)의 중심 주변을 탐색합니다.
vae_x_range = (-2.5, 2.5)
vae_y_range = (-2.5, 2.5)
grid_n = 15
ae_canvas = latent_grid(ae2.decoder, ae_x_range, ae_y_range, n=grid_n)
vae_canvas = latent_grid(vae.decoder, vae_x_range, vae_y_range, n=grid_n)

fig, axes = plt.subplots(1, 2, figsize=(16, 8))
panels = [
    (ae_canvas, &quot;AE: observed test range&quot;, ae_x_range, ae_y_range),
    (vae_canvas, &quot;VAE: prior N(0, I), +/-2.5 per axis&quot;, vae_x_range, vae_y_range),
]
tick_indices = np.array([0, grid_n // 2, grid_n - 1])
# 각 타일의 중심에 실제 디코딩한 좌표를 표시합니다.
tick_positions = tick_indices * 28 + 13.5
for ax, (canvas, title, x_range, y_range) in zip(axes, panels):
    ax.imshow(canvas, cmap=&quot;gray&quot;, vmin=0, vmax=1, interpolation=&quot;nearest&quot;)
    ax.set_xticks(tick_positions, [f&quot;{v:.2f}&quot; for v in np.linspace(*x_range, grid_n)[tick_indices]])
    ax.set_yticks(tick_positions, [f&quot;{v:.2f}&quot; for v in np.linspace(*y_range, grid_n)[::-1][tick_indices]])
    ax.set_xlabel(&quot;z1&quot;)
    ax.set_ylabel(&quot;z2&quot;)
    ax.set_title(title, fontsize=12)
fig.suptitle(&quot;Uniform coordinate grid — not random sampling&quot;, fontsize=14)
plt.tight_layout()
plt.show()</code></pre>
</details>
]]></description>
        </item>
        <item>
            <title><![CDATA[Optimizer는 무엇을 보완해 왔는가?]]></title>
            <link>https://velog.io/@def_uk/Gradient-Descent</link>
            <guid>https://velog.io/@def_uk/Gradient-Descent</guid>
            <pubDate>Mon, 07 Sep 2026 04:23:49 GMT</pubDate>
            <description><![CDATA[<blockquote>
<h2 id="학습-목표">학습 목표</h2>
<p>이 글을 읽고 나면 다음을 할 수 있다.</p>
<ol>
<li><code>loss.backward()</code>로 계산된 Gradient가 parameter update에 어떻게 쓰이는지 설명할 수 있다.</li>
<li><code>batch_size</code>에 따라 Batch GD, SGD, Mini-batch GD를 코드로 구분할 수 있다.</li>
<li>Momentum, AdaGrad, RMSProp, Adadelta, Adam의 핵심 update를 직접 구현할 수 있다.</li>
<li>같은 training loop에서 PyTorch Optimizer를 바꿔 가며 비교할 수 있다.</li>
</ol>
</blockquote>
<hr>
<h2 id="목차">목차</h2>
<ol>
<li>공통 실습 코드</li>
<li>Gradient Descent — Gradient로 parameter 직접 업데이트하기</li>
<li>Batch GD / SGD / Mini-batch — <code>batch_size</code>로 확인하기</li>
<li>기본 Gradient Descent의 문제 — 코드로 재현하기</li>
<li>Momentum — 이전 방향을 저장하기</li>
<li>AdaGrad — parameter별 누적 보폭 만들기</li>
<li>RMSProp / Adadelta — 오래된 Gradient를 잊기</li>
<li>Adam — 방향과 크기를 함께 저장하기</li>
<li><code>torch.optim</code>으로 같은 실험 실행하기</li>
<li>Trade-off와 전체 흐름 정리</li>
</ol>
<br>

<div align="center">
  <strong>
    Optimizer의 차이는 결국<br>
    어떤 상태값을 저장하고 parameter를 어떻게 업데이트하는가에 있다.
  </strong>
</div>

<br>

<h2 id="1-공통-실습-코드">1. 공통 실습 코드</h2>
<p>먼저 모든 Optimizer에서 함께 사용할 간단한 선형 회귀 데이터를 만든다.</p>
<p>정답 관계는 대략 $y=3x+2$이며, 모델은 이 관계의 weight와 bias를 학습한다.</p>
<pre><code class="language-python">import torch
from torch import nn
from torch.utils.data import DataLoader, TensorDataset

torch.manual_seed(42)

# y = 3x + 2에 작은 noise를 추가한 데이터
X = torch.linspace(-2, 2, 256).unsqueeze(1)
y = 3 * X + 2 + 0.2 * torch.randn_like(X)

dataset = TensorDataset(X, y)


def make_model():
    torch.manual_seed(42)
    return nn.Linear(1, 1)


criterion = nn.MSELoss()</code></pre>
<p>모델의 예측식은 다음과 같다.</p>
<p>$$
\hat{y}=wx+b
$$</p>
<p>학습해야 할 parameter는 <code>model.weight</code>와 <code>model.bias</code>다.</p>
<pre><code class="language-python">model = make_model()

for name, parameter in model.named_parameters():
    print(name, parameter.shape)

# weight torch.Size([1, 1])
# bias   torch.Size([1])</code></pre>
<p>이제 같은 모델과 데이터에서 update 코드만 바꾸며 Optimizer의 차이를 확인해 보자.</p>
<hr>
<h2 id="2-gradient-descent--gradient로-parameter-직접-업데이트하기">2. Gradient Descent — Gradient로 parameter 직접 업데이트하기</h2>
<p>Gradient Descent의 update 식은 다음과 같다.</p>
<p>$$
\theta_{t+1}
=
\theta_t-\eta\nabla_\theta J(\theta_t)
$$</p>
<p>이 식은 PyTorch 코드로 거의 그대로 옮길 수 있다.</p>
<pre><code class="language-python">model = make_model()

pred = model(X)
loss = criterion(pred, y)

# 각 parameter의 gradient를 계산해 .grad에 저장한다.
loss.backward()

learning_rate = 0.1

with torch.no_grad():
    for parameter in model.parameters():
        parameter -= learning_rate * parameter.grad

# 다음 step의 gradient가 누적되지 않도록 비운다.
model.zero_grad()</code></pre>
<p>핵심은 세 줄이다.</p>
<pre><code class="language-python">loss.backward()                              # gradient 계산
parameter -= learning_rate * parameter.grad # gradient 반대 방향으로 이동
model.zero_grad()                            # 이전 gradient 제거</code></pre>
<h3 id="parametergrad에는-무엇이-들어-있을까"><code>parameter.grad</code>에는 무엇이 들어 있을까?</h3>
<p><code>loss.backward()</code>를 호출하면 각 parameter를 Loss에 대해 편미분한 값이 <code>.grad</code>에 저장된다.</p>
<pre><code class="language-python">model = make_model()
loss = criterion(model(X), y)
loss.backward()

print(model.weight.grad)  # ∂J / ∂w
print(model.bias.grad)    # ∂J / ∂b</code></pre>
<p>두 값을 하나의 벡터로 모으면 Gradient가 된다.</p>
<p>$$
\nabla J(w,b)
=
\begin{bmatrix}
\frac{\partial J}{\partial w}\
\frac{\partial J}{\partial b}
\end{bmatrix}
$$</p>
<blockquote>
<p><strong>Gradient는 현재 위치에서 목적함수 $J$가 가장 빠르게 증가하는 방향이다.</strong></p>
</blockquote>
<p>Loss를 줄이려면 반대 방향으로 이동해야 하므로 update 코드에 <code>-</code>가 들어간다.</p>
<pre><code class="language-python">parameter -= learning_rate * parameter.grad</code></pre>
<h3 id="왜-매-step마다-gradient를-다시-계산할까">왜 매 step마다 Gradient를 다시 계산할까?</h3>
<p>parameter가 바뀌면 모델의 예측값과 Loss가 달라지고, 새로운 위치에서의 Gradient도 달라진다.</p>
<p>그래서 학습 코드는 다음 과정을 반복한다.</p>
<pre><code class="language-python">for step in range(100):
    pred = model(X)               # 현재 parameter로 예측
    loss = criterion(pred, y)     # 현재 위치의 Loss

    loss.backward()               # 현재 위치의 Gradient

    with torch.no_grad():
        for parameter in model.parameters():
            parameter -= 0.1 * parameter.grad

    model.zero_grad()             # 다음 step을 위해 초기화</code></pre>
<p><img src="https://velog.velcdn.com/images/def_uk/post/f60f4604-a800-4f72-bf68-122cbeb3c5b8/image.png" alt="step마다 gradient를 계산하는 이유"></p>
<hr>
<h2 id="3-batch-gd--sgd--mini-batch--batch_size로-확인하기">3. Batch GD / SGD / Mini-batch — <code>batch_size</code>로 확인하기</h2>
<p>Batch GD, SGD, Mini-batch의 update 원리는 같다. 코드에서 가장 눈에 띄는 차이는 <code>DataLoader</code>의 <code>batch_size</code>다.</p>
<pre><code class="language-python"># Batch Gradient Descent: 전체 데이터로 1번 update
batch_loader = DataLoader(
    dataset,
    batch_size=len(dataset),
    shuffle=True,
)

# Stochastic Gradient Descent: 데이터 1개마다 update
sgd_loader = DataLoader(
    dataset,
    batch_size=1,
    shuffle=True,
)

# Mini-batch Gradient Descent: 32개마다 update
mini_batch_loader = DataLoader(
    dataset,
    batch_size=32,
    shuffle=True,
)</code></pre>
<p>한 epoch의 update 횟수도 바로 확인할 수 있다.</p>
<pre><code class="language-python">print(len(batch_loader))       # 1
print(len(sgd_loader))         # 256
print(len(mini_batch_loader))  # 8</code></pre>
<p>training loop은 세 경우 모두 같다.</p>
<pre><code class="language-python">def train_one_epoch(model, loader, optimizer):
    model.train()

    for x_batch, y_batch in loader:
        optimizer.zero_grad()

        pred = model(x_batch)
        loss = criterion(pred, y_batch)

        loss.backward()
        optimizer.step()</code></pre>
<p>차이는 한 번의 <code>optimizer.step()</code>이 어떤 데이터를 보고 계산된 Gradient를 사용하는가다.</p>
<table>
<thead>
<tr>
<th>방식</th>
<th align="right"><code>batch_size</code></th>
<th align="right">한 번의 update에 쓰는 데이터</th>
<th>특징</th>
</tr>
</thead>
<tbody><tr>
<td>Batch GD</td>
<td align="right">전체 데이터 수</td>
<td align="right">전체</td>
<td>안정적이지만 update 비용이 큼</td>
</tr>
<tr>
<td>SGD</td>
<td align="right"><code>1</code></td>
<td align="right">1개</td>
<td>자주 update하지만 매우 noisy함</td>
</tr>
<tr>
<td>Mini-batch</td>
<td align="right"><code>32</code>, <code>64</code>, <code>128</code> 등</td>
<td align="right">일부</td>
<td>안정성과 계산 효율성의 절충</td>
</tr>
</tbody></table>
<h3 id="mini-batch-gradient는-어떻게-만들어질까">Mini-batch Gradient는 어떻게 만들어질까?</h3>
<p>batch에 데이터가 $B$개 있다면 각 sample의 Gradient를 평균낸 값이 사용된다.</p>
<p>$$
g_B=\frac{1}{B}\sum_{i=1}^{B}g_i
$$</p>
<p>PyTorch에서는 loss를 평균내는 <code>nn.MSELoss()</code>의 기본 설정과 <code>backward()</code>가 이 계산을 처리한다.</p>
<pre><code class="language-python">x_batch, y_batch = next(iter(mini_batch_loader))

pred = model(x_batch)
loss = criterion(pred, y_batch)  # batch의 평균 MSE
loss.backward()                  # 평균 Loss에 대한 Gradient</code></pre>
<p>Mini-batch는 여러 sample을 평균내 순수한 SGD보다 안정적이며, 행렬 연산을 통해 GPU를 효율적으로 사용할 수 있다.</p>
<p>반면 batch가 너무 크면 메모리 사용량이 증가하고 한 epoch의 update 횟수가 감소한다. 너무 작으면 Gradient noise가 커진다.</p>
<blockquote>
<p><strong>Batch Size 역시 계산 효율성과 Gradient 안정성 사이의 trade-off다.</strong></p>
</blockquote>
<hr>
<h2 id="4-기본-gradient-descent의-문제--코드로-재현하기">4. 기본 Gradient Descent의 문제 — 코드로 재현하기</h2>
<h3 id="41-learning-rate가-너무-작거나-큰-경우">4.1 Learning Rate가 너무 작거나 큰 경우</h3>
<p>간단한 함수에서 Learning Rate만 바꿔 보자.</p>
<p>$$
J(w)=(w-3)^2
$$</p>
<p>이 함수의 minimum은 $w=3$이다.</p>
<pre><code class="language-python">def run_gradient_descent(lr, steps=10):
    w = torch.tensor(0.0, requires_grad=True)
    history = []

    for step in range(steps):
        loss = (w - 3) ** 2
        loss.backward()

        with torch.no_grad():
            w -= lr * w.grad

        history.append({
            &quot;step&quot;: step,
            &quot;w&quot;: w.item(),
            &quot;loss&quot;: loss.item(),
        })

        w.grad.zero_()

    return history


slow = run_gradient_descent(lr=0.05)      # 조금씩 이동
oscillating = run_gradient_descent(lr=0.8) # minimum 양쪽을 오가며 수렴
diverging = run_gradient_descent(lr=1.1)   # 진폭이 커지며 발산</code></pre>
<p>Gradient가 방향을 알려줘도 실제 이동량은 <code>lr</code>이 결정한다.</p>
<pre><code class="language-python">update = learning_rate * gradient</code></pre>
<ul>
<li><code>lr</code>이 너무 작으면 학습이 느리다.</li>
<li><code>lr</code>이 너무 크면 minimum을 지나쳐 진동하거나 발산할 수 있다.</li>
</ul>
<h3 id="42-gradient가-0이면-항상-minimum일까">4.2 Gradient가 0이면 항상 minimum일까?</h3>
<p>Saddle Point는 한 방향에서는 minimum, 다른 방향에서는 maximum처럼 보이는 지점이다.</p>
<p>$$
f(x,y)=x^2-y^2
$$</p>
<p>원점에서 Gradient를 코드로 계산하면 두 성분이 모두 0이다.</p>
<pre><code class="language-python">point = torch.tensor([0.0, 0.0], requires_grad=True)
x, y = point[0], point[1]

value = x**2 - y**2
value.backward()

print(point.grad)  # tensor([0., 0.])</code></pre>
<p>하지만 $x$ 방향으로는 값이 증가하고 $y$ 방향으로는 감소하므로 원점은 minimum이 아니다.</p>
<h3 id="43-모든-parameter에-같은-learning-rate를-적용한다">4.3 모든 parameter에 같은 Learning Rate를 적용한다</h3>
<p>기본 SGD의 코드를 보면 모든 parameter에 동일한 <code>lr</code>을 곱한다.</p>
<pre><code class="language-python">for parameter in model.parameters():
    parameter -= lr * parameter.grad</code></pre>
<p>그러나 각 parameter의 Gradient 크기와 update 빈도는 다를 수 있다. 이후 AdaGrad, RMSProp, Adam은 parameter마다 다른 실제 보폭을 만들기 위해 별도의 상태값을 저장한다.</p>
<hr>
<h2 id="5-momentum--이전-방향을-저장하기">5. Momentum — 이전 방향을 저장하기</h2>
<p>Mini-batch SGD는 batch마다 Gradient가 달라져 좁고 긴 Loss Surface에서 지그재그로 움직일 수 있다.</p>
<p>기본 SGD는 현재 Gradient만 사용한다.</p>
<pre><code class="language-python">parameter -= lr * parameter.grad</code></pre>
<p>Momentum은 이전까지의 이동 방향을 <code>velocity</code>에 저장한다.</p>
<p>$$
v_t=\beta v_{t-1}+g_t
$$</p>
<p>$$
\theta_{t+1}=\theta_t-\eta v_t
$$</p>
<p>이를 직접 구현하면 다음과 같다.</p>
<pre><code class="language-python">model = make_model()
parameters = list(model.parameters())

# parameter마다 같은 모양의 velocity를 하나씩 만든다.
velocity = {
    parameter: torch.zeros_like(parameter)
    for parameter in parameters
}


@torch.no_grad()
def momentum_step(parameters, velocity, lr=0.01, beta=0.9):
    for parameter in parameters:
        velocity[parameter].mul_(beta).add_(parameter.grad)
        parameter.add_(velocity[parameter], alpha=-lr)</code></pre>
<p>training loop에서는 직접 만든 <code>momentum_step()</code>을 <code>optimizer.step()</code> 자리에 호출한다.</p>
<pre><code class="language-python">for x_batch, y_batch in mini_batch_loader:
    model.zero_grad()

    loss = criterion(model(x_batch), y_batch)
    loss.backward()

    momentum_step(parameters, velocity, lr=0.01, beta=0.9)</code></pre>
<p>코드의 핵심은 이 부분이다.</p>
<pre><code class="language-python">velocity[parameter].mul_(beta).add_(parameter.grad)</code></pre>
<ul>
<li>Gradient가 같은 방향으로 반복되면 <code>velocity</code>에 누적된다.</li>
<li>Gradient가 반대 방향으로 바뀌면 이전 값과 일부 상쇄된다.</li>
</ul>
<blockquote>
<p><strong>Momentum은 일관된 방향의 이동은 강화하고, 반복되는 진동은 줄인다.</strong></p>
</blockquote>
<p>다만 <code>beta</code>가 너무 크면 관성이 강해져 minimum을 지나치는 overshooting이 발생할 수 있다.</p>
<p>PyTorch에서는 같은 동작을 다음처럼 사용한다.</p>
<pre><code class="language-python">optimizer = torch.optim.SGD(
    model.parameters(),
    lr=0.01,
    momentum=0.9,
)</code></pre>
<hr>
<h2 id="6-adagrad--parameter별-누적-보폭-만들기">6. AdaGrad — parameter별 누적 보폭 만들기</h2>
<p>AdaGrad의 출발점은 간단하다.</p>
<blockquote>
<p><strong>자주 크게 움직인 parameter는 천천히, 드물게 움직인 parameter는 상대적으로 크게 움직이자.</strong></p>
</blockquote>
<p>이를 위해 parameter마다 과거 Gradient 제곱의 누적값을 저장한다.</p>
<p>$$
G_t=G_{t-1}+g_t^2
$$</p>
<p>$$
\theta_{t+1}
=
\theta_t-\frac{\eta}{\sqrt{G_t}+\epsilon}g_t
$$</p>
<pre><code class="language-python">model = make_model()
parameters = list(model.parameters())

squared_grad_sum = {
    parameter: torch.zeros_like(parameter)
    for parameter in parameters
}


@torch.no_grad()
def adagrad_step(parameters, squared_grad_sum, lr=0.01, eps=1e-10):
    for parameter in parameters:
        grad = parameter.grad

        # G_t = G_{t-1} + g_t²
        squared_grad_sum[parameter].addcmul_(grad, grad)

        # parameter마다 서로 다른 실제 보폭
        denominator = squared_grad_sum[parameter].sqrt().add_(eps)
        parameter.addcdiv_(grad, denominator, value=-lr)</code></pre>
<p>핵심은 <code>denominator</code>가 parameter별로 다르다는 점이다.</p>
<pre><code class="language-python">denominator = squared_grad_sum[parameter].sqrt() + eps
effective_lr = lr / denominator</code></pre>
<p>Gradient가 계속 컸던 parameter는 <code>squared_grad_sum</code>이 커지고 실제 Learning Rate는 작아진다.</p>
<p>문제는 이 값이 계속 더해지기만 한다는 것이다.</p>
<pre><code class="language-python">squared_grad_sum[parameter] += grad * grad</code></pre>
<p>Gradient 제곱은 음수가 아니므로 누적값은 줄어들지 않는다. 학습이 길어지면 실제 Learning Rate가 거의 0에 가까워질 수 있다.</p>
<blockquote>
<p><strong>AdaGrad는 과거를 너무 오래 기억하기 때문에 학습이 지나치게 빨리 느려질 수 있다.</strong></p>
</blockquote>
<p>PyTorch에서는 다음과 같이 사용한다.</p>
<pre><code class="language-python">optimizer = torch.optim.Adagrad(
    model.parameters(),
    lr=0.01,
)</code></pre>
<hr>
<h2 id="7-rmsprop--adadelta--오래된-gradient를-잊기">7. RMSProp / Adadelta — 오래된 Gradient를 잊기</h2>
<p>AdaGrad가 과거 Gradient를 모두 더하는 것이 문제라면, 오래된 정보의 영향력을 줄이면 된다.</p>
<h3 id="rmsprop">RMSProp</h3>
<p>RMSProp은 단순 누적합 대신 Gradient 제곱의 지수 이동 평균을 저장한다.</p>
<p>$$
v_t=\beta v_{t-1}+(1-\beta)g_t^2
$$</p>
<pre><code class="language-python">model = make_model()
parameters = list(model.parameters())

square_avg = {
    parameter: torch.zeros_like(parameter)
    for parameter in parameters
}


@torch.no_grad()
def rmsprop_step(parameters, square_avg, lr=0.001, beta=0.9, eps=1e-8):
    for parameter in parameters:
        grad = parameter.grad

        # v_t = beta * v_{t-1} + (1 - beta) * g_t²
        square_avg[parameter].mul_(beta).addcmul_(
            grad,
            grad,
            value=1 - beta,
        )

        denominator = square_avg[parameter].sqrt().add_(eps)
        parameter.addcdiv_(grad, denominator, value=-lr)</code></pre>
<p>AdaGrad와 비교하면 한 줄의 차이가 핵심이다.</p>
<pre><code class="language-python"># AdaGrad: 과거 값을 그대로 유지하며 계속 더한다.
accumulator += grad**2

# RMSProp: 과거 값은 beta만큼만 남긴다.
square_avg = beta * square_avg + (1 - beta) * grad**2</code></pre>
<p>예를 들어 <code>beta=0.9</code>라면 오래된 Gradient의 영향은 시간이 지나며 다음처럼 작아진다.</p>
<p>$$
1\rightarrow0.9\rightarrow0.9^2\rightarrow0.9^3\rightarrow\cdots
$$</p>
<blockquote>
<p><strong>RMSProp은 최근 Gradient의 크기를 더 중요하게 보면서 parameter별 보폭을 조절한다.</strong></p>
</blockquote>
<p>PyTorch 코드는 다음과 같다.</p>
<pre><code class="language-python">optimizer = torch.optim.RMSprop(
    model.parameters(),
    lr=0.001,
    alpha=0.9,
)</code></pre>
<h3 id="adadelta">Adadelta</h3>
<p>Adadelta도 Gradient 제곱의 지수 이동 평균을 사용한다. 여기에 최근 parameter update의 제곱 평균도 함께 저장한다.</p>
<pre><code class="language-python">model = make_model()
parameters = list(model.parameters())

square_avg = {
    parameter: torch.zeros_like(parameter)
    for parameter in parameters
}
acc_delta = {
    parameter: torch.zeros_like(parameter)
    for parameter in parameters
}


@torch.no_grad()
def adadelta_step(parameters, square_avg, acc_delta, rho=0.9, eps=1e-6):
    for parameter in parameters:
        grad = parameter.grad

        # 최근 Gradient²의 이동 평균
        square_avg[parameter].mul_(rho).addcmul_(
            grad,
            grad,
            value=1 - rho,
        )

        # 최근 update²와 Gradient²를 이용해 update 크기 결정
        std = square_avg[parameter].add(eps).sqrt()
        delta = acc_delta[parameter].add(eps).sqrt().div(std).mul(grad)

        parameter.sub_(delta)

        # 방금 적용한 update²의 이동 평균
        acc_delta[parameter].mul_(rho).addcmul_(
            delta,
            delta,
            value=1 - rho,
        )</code></pre>
<p>RMSProp이 <code>square_avg</code> 하나를 중심으로 보폭을 조절한다면, Adadelta는 <code>square_avg</code>와 <code>acc_delta</code>를 함께 사용한다.</p>
<pre><code class="language-python">optimizer = torch.optim.Adadelta(
    model.parameters(),
    rho=0.9,
)</code></pre>
<table>
<thead>
<tr>
<th>AdaGrad</th>
<th>RMSProp / Adadelta</th>
</tr>
</thead>
<tbody><tr>
<td>Gradient²를 계속 누적</td>
<td>오래된 Gradient의 영향 감소</td>
</tr>
<tr>
<td>실제 Learning Rate가 계속 감소</td>
<td>최근 정보를 중심으로 보폭 조절</td>
</tr>
<tr>
<td>장기 학습에서 지나치게 느려질 수 있음</td>
<td>AdaGrad의 감소 문제 완화</td>
</tr>
</tbody></table>
<hr>
<h2 id="8-adam--방향과-크기를-함께-저장하기">8. Adam — 방향과 크기를 함께 저장하기</h2>
<p>지금까지 코드에서 저장한 상태값을 다시 보자.</p>
<pre><code class="language-python"># Momentum
velocity       # Gradient 방향을 누적

# RMSProp
square_avg     # Gradient²의 이동 평균</code></pre>
<p>Adam은 두 아이디어를 함께 사용한다.</p>
<pre><code class="language-python">first_moment   # Gradient의 지수 이동 평균
second_moment  # Gradient²의 지수 이동 평균</code></pre>
<h3 id="first-moment">First Moment</h3>
<p>$$
m_t=\beta_1m_{t-1}+(1-\beta_1)g_t
$$</p>
<p>Momentum과 비슷하게 Gradient의 방향을 안정화한다.</p>
<h3 id="second-moment">Second Moment</h3>
<p>$$
v_t=\beta_2v_{t-1}+(1-\beta_2)g_t^2
$$</p>
<p>RMSProp과 비슷하게 parameter별 보폭을 조절한다.</p>
<h3 id="bias-correction까지-직접-구현하기">Bias Correction까지 직접 구현하기</h3>
<p>두 moment는 0으로 초기화되므로 학습 초반에는 0에 가깝게 편향된다.</p>
<p>$$
\hat{m}_t=\frac{m_t}{1-\beta_1^t},\qquad
\hat{v}_t=\frac{v_t}{1-\beta_2^t}
$$</p>
<p>보정한 값을 사용하는 전체 update 코드는 다음과 같다.</p>
<pre><code class="language-python">model = make_model()
parameters = list(model.parameters())

first_moment = {
    parameter: torch.zeros_like(parameter)
    for parameter in parameters
}
second_moment = {
    parameter: torch.zeros_like(parameter)
    for parameter in parameters
}


@torch.no_grad()
def adam_step(
    parameters,
    first_moment,
    second_moment,
    step,
    lr=0.001,
    beta1=0.9,
    beta2=0.999,
    eps=1e-8,
):
    for parameter in parameters:
        grad = parameter.grad

        # 1차 모멘트: Gradient의 이동 평균
        first_moment[parameter].mul_(beta1).add_(
            grad,
            alpha=1 - beta1,
        )

        # 2차 모멘트: Gradient²의 이동 평균
        second_moment[parameter].mul_(beta2).addcmul_(
            grad,
            grad,
            value=1 - beta2,
        )

        # 0으로 초기화되어 생기는 초반 편향 보정
        m_hat = first_moment[parameter] / (1 - beta1**step)
        v_hat = second_moment[parameter] / (1 - beta2**step)

        parameter.addcdiv_(
            m_hat,
            v_hat.sqrt().add_(eps),
            value=-lr,
        )</code></pre>
<p>training loop에서는 <code>step</code>을 1부터 증가시킨다.</p>
<pre><code class="language-python">step = 0

for x_batch, y_batch in mini_batch_loader:
    model.zero_grad()

    loss = criterion(model(x_batch), y_batch)
    loss.backward()

    step += 1
    adam_step(
        parameters,
        first_moment,
        second_moment,
        step,
    )</code></pre>
<p>코드와 역할을 연결하면 다음과 같다.</p>
<table>
<thead>
<tr>
<th>코드의 상태값</th>
<th>수식</th>
<th>기억하는 것</th>
<th>역할</th>
</tr>
</thead>
<tbody><tr>
<td><code>first_moment</code></td>
<td>$m_t$</td>
<td>Gradient</td>
<td>방향 안정화</td>
</tr>
<tr>
<td><code>second_moment</code></td>
<td>$v_t$</td>
<td>Gradient²</td>
<td>parameter별 보폭 조절</td>
</tr>
<tr>
<td><code>m_hat</code>, <code>v_hat</code></td>
<td>$\hat{m}_t$, $\hat{v}_t$</td>
<td>보정된 moment</td>
<td>초기 0 편향 보정</td>
</tr>
</tbody></table>
<p>PyTorch에서는 다음 한 줄로 같은 핵심 로직을 사용한다.</p>
<pre><code class="language-python">optimizer = torch.optim.Adam(
    model.parameters(),
    lr=0.001,
    betas=(0.9, 0.999),
)</code></pre>
<blockquote>
<p><strong>Adam은 Gradient의 방향과 크기를 모두 추적하고, parameter마다 adaptive한 update를 수행한다.</strong></p>
</blockquote>
<hr>
<h2 id="9-torchoptim으로-같은-실험-실행하기">9. <code>torch.optim</code>으로 같은 실험 실행하기</h2>
<p>직접 구현한 update 식은 서로 다르지만 PyTorch의 training loop는 거의 같다.</p>
<pre><code class="language-python">def train(model, loader, optimizer, epochs=100):
    losses = []

    for epoch in range(epochs):
        model.train()
        epoch_loss = 0.0

        for x_batch, y_batch in loader:
            optimizer.zero_grad()

            pred = model(x_batch)
            loss = criterion(pred, y_batch)

            loss.backward()
            optimizer.step()

            epoch_loss += loss.item() * len(x_batch)

        losses.append(epoch_loss / len(loader.dataset))

    return losses</code></pre>
<p>Optimizer 생성 부분만 함수로 분리해 바꿔 가며 실행할 수 있다.</p>
<pre><code class="language-python">def make_optimizer(name, model):
    if name == &quot;sgd&quot;:
        return torch.optim.SGD(model.parameters(), lr=0.01)

    if name == &quot;momentum&quot;:
        return torch.optim.SGD(
            model.parameters(),
            lr=0.01,
            momentum=0.9,
        )

    if name == &quot;adagrad&quot;:
        return torch.optim.Adagrad(model.parameters(), lr=0.01)

    if name == &quot;rmsprop&quot;:
        return torch.optim.RMSprop(
            model.parameters(),
            lr=0.001,
            alpha=0.9,
        )

    if name == &quot;adadelta&quot;:
        return torch.optim.Adadelta(
            model.parameters(),
            rho=0.9,
        )

    if name == &quot;adam&quot;:
        return torch.optim.Adam(
            model.parameters(),
            lr=0.001,
            betas=(0.9, 0.999),
        )

    raise ValueError(f&quot;Unknown optimizer: {name}&quot;)</code></pre>
<p>같은 초기 parameter와 같은 Mini-batch 조건에서 비교한다.</p>
<pre><code class="language-python">results = {}

for name in [
    &quot;sgd&quot;,
    &quot;momentum&quot;,
    &quot;adagrad&quot;,
    &quot;rmsprop&quot;,
    &quot;adadelta&quot;,
    &quot;adam&quot;,
]:
    model = make_model()
    loader = DataLoader(
        dataset,
        batch_size=32,
        shuffle=True,
        generator=torch.Generator().manual_seed(42),
    )
    optimizer = make_optimizer(name, model)

    losses = train(model, loader, optimizer, epochs=100)

    results[name] = {
        &quot;final_loss&quot;: losses[-1],
        &quot;weight&quot;: model.weight.item(),
        &quot;bias&quot;: model.bias.item(),
    }

for name, result in results.items():
    print(name, result)</code></pre>
<p>이 실험에서 중요한 것은 순위를 한 번 정하는 것이 아니다. Optimizer마다 Learning Rate의 적절한 범위가 다르므로 값 하나만으로 공정한 우열을 결정할 수 없다.</p>
<p>대신 다음을 확인하는 코드로 사용하면 좋다.</p>
<ul>
<li>초기 Loss가 얼마나 빠르게 감소하는가?</li>
<li>Loss가 크게 진동하는가?</li>
<li>같은 epoch 이후 $w\approx3$, $b\approx2$에 얼마나 가까운가?</li>
<li>Learning Rate를 바꾸면 결과가 얼마나 민감하게 변하는가?</li>
</ul>
<h3 id="optimizer가-실제로-저장하는-상태-확인하기">Optimizer가 실제로 저장하는 상태 확인하기</h3>
<p><code>optimizer.state</code>를 보면 각 알고리즘이 무엇을 기억하는지 확인할 수 있다. 최소 한 번 <code>optimizer.step()</code>을 실행한 뒤 확인해야 한다.</p>
<pre><code class="language-python">parameter = next(model.parameters())
state = optimizer.state[parameter]

print(state.keys())</code></pre>
<p>대표적인 state는 다음과 같다.</p>
<table>
<thead>
<tr>
<th>Optimizer</th>
<th>주요 state</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td>SGD</td>
<td>없음</td>
<td>현재 Gradient만 사용</td>
</tr>
<tr>
<td>SGD + Momentum</td>
<td><code>momentum_buffer</code></td>
<td>이전 방향 누적</td>
</tr>
<tr>
<td>AdaGrad</td>
<td><code>sum</code></td>
<td>Gradient² 누적합</td>
</tr>
<tr>
<td>RMSProp</td>
<td><code>square_avg</code></td>
<td>Gradient² 이동 평균</td>
</tr>
<tr>
<td>Adadelta</td>
<td><code>square_avg</code>, <code>acc_delta</code></td>
<td>Gradient²와 update² 이동 평균</td>
</tr>
<tr>
<td>Adam</td>
<td><code>exp_avg</code>, <code>exp_avg_sq</code>, <code>step</code></td>
<td>1차·2차 모멘트와 step</td>
</tr>
</tbody></table>
<p>결국 같은 <code>optimizer.step()</code> 안에서 Optimizer마다 서로 다른 state를 읽고 갱신한다.</p>
<hr>
<h2 id="10-trade-off와-전체-흐름-정리">10. Trade-off와 전체 흐름 정리</h2>
<p>Adam의 코드가 가장 많은 정보를 사용한다고 해서 모든 문제에서 항상 가장 좋은 것은 아니다.</p>
<table>
<thead>
<tr>
<th>Optimizer</th>
<th align="right">추가 상태값</th>
<th>얻는 것</th>
<th>주의할 점</th>
</tr>
</thead>
<tbody><tr>
<td>SGD</td>
<td align="right">0개</td>
<td>단순한 update, 적은 메모리</td>
<td>Learning Rate와 진동에 민감</td>
</tr>
<tr>
<td>Momentum SGD</td>
<td align="right">parameter당 1개</td>
<td>방향 누적, 진동 감소</td>
<td>강한 관성은 overshooting 가능</td>
</tr>
<tr>
<td>AdaGrad</td>
<td align="right">parameter당 1개</td>
<td>parameter별 Learning Rate</td>
<td>보폭이 계속 작아질 수 있음</td>
</tr>
<tr>
<td>RMSProp</td>
<td align="right">parameter당 1개</td>
<td>최근 Gradient 중심의 보폭</td>
<td>decay rate 설정 필요</td>
</tr>
<tr>
<td>Adadelta</td>
<td align="right">parameter당 2개</td>
<td>Gradient와 update 크기 반영</td>
<td>더 많은 state 필요</td>
</tr>
<tr>
<td>Adam</td>
<td align="right">parameter당 2개 이상</td>
<td>방향과 보폭을 함께 조절</td>
<td>메모리 증가, 항상 최고는 아님</td>
</tr>
</tbody></table>
<p>Optimizer의 발전을 코드의 state 변화로 연결하면 다음과 같다.</p>
<pre><code class="language-text">Gradient Descent / SGD
└─ parameter.grad만 사용
   │
   │ Mini-batch Gradient가 noisy하고 진동한다.
   ▼
Momentum
└─ velocity 추가
   │
   │ 모든 parameter에 같은 Learning Rate를 사용한다.
   ▼
AdaGrad
└─ squared_grad_sum 추가
   │
   │ 과거 Gradient²가 끝없이 누적된다.
   ▼
RMSProp / Adadelta
└─ square_avg, acc_delta로 오래된 정보의 영향 감소
   │
   │ 방향 정보도 함께 사용하면 어떨까?
   ▼
Adam
├─ first_moment: Gradient 방향의 이동 평균
└─ second_moment: Gradient²의 이동 평균</code></pre>
<p>코드에서 다시 보면 차이는 더욱 선명하다.</p>
<pre><code class="language-python"># SGD
parameter -= lr * grad

# Momentum
velocity = beta * velocity + grad
parameter -= lr * velocity

# AdaGrad
grad_sum += grad**2
parameter -= lr * grad / (sqrt(grad_sum) + eps)

# RMSProp
square_avg = beta * square_avg + (1 - beta) * grad**2
parameter -= lr * grad / (sqrt(square_avg) + eps)

# Adam
first_moment = beta1 * first_moment + (1 - beta1) * grad
second_moment = beta2 * second_moment + (1 - beta2) * grad**2
parameter -= lr * m_hat / (sqrt(v_hat) + eps)</code></pre>
<blockquote>
<p><strong>Optimizer의 발전은 단순히 더 빠르게 내려가기 위한 경쟁이 아니다.</strong><br><strong>학습 과정에서 만나는 문제에 따라 어떤 정보를 추가로 저장하고 update에 반영할지 발전시켜 온 과정이다.</strong></p>
</blockquote>
<hr>
<h2 id="reference">Reference</h2>
<ul>
<li>Sebastian Ruder, <em>An overview of gradient descent optimization algorithms</em>, <a href="https://arxiv.org/abs/1609.04747">arXiv:1609.04747</a>.</li>
<li><a href="https://pytorch.org/docs/stable/optim.html">PyTorch Optimizer documentation</a></li>
</ul>
]]></description>
        </item>
        <item>
            <title><![CDATA[Random Forest와 Extra Trees]]></title>
            <link>https://velog.io/@def_uk/Random-Forest%EC%99%80-Extra-Trees</link>
            <guid>https://velog.io/@def_uk/Random-Forest%EC%99%80-Extra-Trees</guid>
            <pubDate>Tue, 01 Sep 2026 14:14:16 GMT</pubDate>
            <description><![CDATA[<blockquote>
<h3 id="학습지표">학습지표</h3>
<p><strong>Decision Tree</strong>의 높은 variance를 줄이기 위해 <strong>Bagging</strong>은 데이터에 randomness를 주고, <strong>Random Forest</strong>는 feature 후보에도 randomness를 추가하며, <strong>Extra Trees</strong>는 분할 threshold까지 randomness를 추가한다.</p>
</blockquote>
<p>1편에서는 Decision Tree의 높은 variance를 Bagging이 Bootstrap과 Aggregation으로 완화하는 과정을 살펴봤다.</p>
<p>하지만 Bootstrap만으로 모든 Tree가 충분히 달라지는 것은 아니다. 이번 글에서는 Bagging 뒤에도 남는 <strong>Tree 간 오류 상관</strong>에서 출발해 Random Forest와 Extra Trees의 차이를 정리한다.</p>
<hr>
<h2 id="1-bagging-뒤에도-남는-문제-tree들이-너무-비슷하다">1. Bagging 뒤에도 남는 문제: Tree들이 너무 비슷하다</h2>
<p>Bootstrap 표본이 다르더라도 모든 Tree는 보통 전체 Feature에서 가장 좋은 분할을 찾는다. 하나의 Feature가 매우 강력하다면 대부분의 Tree가 루트나 상위 노드에서 같은 Feature를 선택할 수 있다.</p>
<pre><code class="language-text">Bootstrap A → Fare로 첫 분할
Bootstrap B → Fare로 첫 분할
Bootstrap C → Fare로 첫 분할
Bootstrap D → Fare로 첫 분할</code></pre>
<p>학습 표본은 조금씩 달라도 Tree 구조와 오류가 비슷해질 수 있다. 모델 수만 늘리는 것으로 평균 효과가 계속 커지지 않는 이유다.</p>
<p>같은 variance <code>σ²</code>을 가진 모델 <code>B</code>개의 모든 쌍이 같은 상관계수 <code>ρ</code>를 가진다고 단순화하면, 평균 예측의 variance는 다음처럼 쓸 수 있다.</p>
<p>$$
\operatorname{Var}(\bar f)
=\sigma^2\left(\rho+\frac{1-\rho}{B}\right)
$$</p>
<ul>
<li><code>ρ=0</code>이면 평균 variance가 <code>σ²/B</code>까지 감소한다.</li>
<li><code>ρ</code>가 크면 모델 수를 늘려도 감소 폭이 제한된다.</li>
<li><code>B→∞</code>여도 식은 <code>ρσ²</code>에 가까워진다.</li>
</ul>
<p>이 식은 단순한 가정 아래의 직관이지만 중요한 메시지를 준다.</p>
<blockquote>
<p>앙상블에서는 모델 수뿐 아니라 구성원 오류가 얼마나 다르게 움직이는지도 중요하다.</p>
</blockquote>
<p>Random Forest는 이 상관을 낮추기 위해 Feature 선택에도 무작위성을 넣는다.</p>
<hr>
<h2 id="2-random-forest-노드마다-다른-feature-후보를-본다">2. Random Forest: 노드마다 다른 Feature 후보를 본다</h2>
<p>scikit-learn의 전형적인 Random Forest는 다음 세 요소를 결합한다.</p>
<pre><code class="language-text">Bootstrap sample
  +
노드별 Feature 후보 무작위화
  +
Tree 예측의 평균</code></pre>
<p>핵심은 <code>max_features</code>다. Tree를 만들 때 Feature 묶음을 한 번만 뽑아 끝까지 고정하는 방식이 아니다. <strong>각 노드에서 분할을 찾을 때마다</strong> Feature 후보 집합을 새로 선택한다.</p>
<pre><code class="language-text">루트 노드:      [Age, Fare, Sex] 중 최적 분할 선택
왼쪽 자식 노드: [Pclass, Fare, SibSp] 중 최적 분할 선택
오른쪽 자식:    [Age, Parch, Pclass] 중 최적 분할 선택</code></pre>
<p>강력한 Feature가 후보에서 빠진 노드에서는 다른 Feature가 분할 기회를 얻는다. 그 결과 Tree 구조와 오류 패턴이 더 다양해질 수 있다.</p>
<h3 id="21-random-forest의-목적은-단순히-더-랜덤하게가-아니다">2.1 Random Forest의 목적은 단순히 “더 랜덤하게”가 아니다</h3>
<p>Feature 후보 수에는 trade-off가 있다.</p>
<ul>
<li>후보가 많으면 개별 Tree는 좋은 분할을 찾기 쉽지만 Tree들이 비슷해질 수 있다.</li>
<li>후보가 적으면 Tree 다양성은 커질 수 있지만 중요한 Feature를 자주 놓칠 수 있다.</li>
</ul>
<p>Breiman의 Random Forest 논문은 이를 개별 Tree의 <strong>strength</strong>와 Tree 사이의 <strong>correlation</strong> 사이의 균형으로 설명한다.</p>
<pre><code class="language-text">좋은 Random Forest를 만드는 방향
= 충분한 개별 Tree의 예측력
+ 지나치게 높지 않은 Tree 간 오류 상관</code></pre>
<p>따라서 <code>max_features</code>는 “작을수록 좋다”는 파라미터가 아니다. 데이터에 맞게 검증해야 하는 조절 장치다.</p>
<h3 id="22-bagged-trees와-random-forest의-차이">2.2 Bagged Trees와 Random Forest의 차이</h3>
<pre><code class="language-text">Bagged Trees
Bootstrap sample
  → 보통 모든 Feature에서 최적 분할 탐색
  → Tree 생성
  → 예측 집계

Random Forest
Bootstrap sample
  → 매 노드에서 Feature 일부를 무작위 선택
  → 선택된 후보 안에서 최적 분할 탐색
  → Tree 생성
  → 예측 집계</code></pre>
<p>단, scikit-learn의 범용 <code>BaggingClassifier</code>도 <code>max_features</code>와 <code>bootstrap_features</code>를 통해 Feature sampling을 지원한다. 차이는 Random Forest가 <strong>Tree의 각 노드에서</strong> Feature 후보를 다시 선택한다는 점이다.</p>
<h3 id="23-독립적인-tree라는-표현을-조심해야-한다">2.3 “독립적인 Tree”라는 표현을 조심해야 한다</h3>
<p>Random Forest의 Tree들은 앞 Tree의 결과를 다음 Tree가 이어받지 않으므로 별도·병렬 학습이 가능하다. 그러나 예측 오류가 통계적으로 독립이라는 뜻은 아니다.</p>
<p>오히려 Random Forest가 해결하려는 문제가 바로 <strong>Tree 오류의 상관</strong>이다. 따라서 다음 표현이 더 정확하다.</p>
<blockquote>
<p>Tree들은 순차 의존 없이 학습할 수 있으며, Feature Randomness는 Tree들이 덜 비슷하게 움직이도록 설계된 장치다.</p>
</blockquote>
<hr>
<h2 id="3-extra-trees-분할-임계값까지-무작위화한다">3. Extra Trees: 분할 임계값까지 무작위화한다</h2>
<p><strong>Extra Trees(Extremely Randomized Trees)</strong>는 Random Forest보다 분할 과정에 더 강한 무작위성을 넣는다.</p>
<pre><code class="language-text">Random Forest
→ Feature 일부를 무작위 선택
→ 각 후보 Feature에서 좋은 Threshold 탐색
→ 후보 안에서 최적 분할 선택

Extra Trees
→ Feature 일부를 무작위 선택
→ 각 후보 Feature의 Threshold를 무작위 생성
→ 무작위로 생성된 분할 중 가장 좋은 것 선택</code></pre>
<p>예를 들어 <code>Age</code>가 후보 Feature라면 Random Forest는 가능한 분할점들을 비교해 불순도를 가장 많이 줄이는 값을 찾는다. Extra Trees는 <code>Age &lt; 31.7</code>과 같은 임계값을 무작위로 생성하고, 다른 후보 Feature에서 생성한 무작위 분할과 비교한다.</p>
<p>중요한 점은 Extra Trees가 아무 분할이나 그대로 채택하는 것은 아니라는 것이다.</p>
<blockquote>
<p>후보 Feature마다 무작위 Threshold를 만들지만, 그 무작위 후보들 중에서는 가장 좋은 분할을 선택한다.</p>
</blockquote>
<h3 id="31-더-강한-무작위성의-trade-off">3.1 더 강한 무작위성의 trade-off</h3>
<p>Extra Trees의 무작위화는 Tree 오류의 상관과 ensemble variance를 더 낮출 수 있다. 모든 가능한 Threshold를 탐색하지 않아 학습이 빨라질 수도 있다.</p>
<p>반면 유용한 최적 분할을 놓치면 bias가 커질 수 있다.</p>
<pre><code class="language-text">무작위성 증가
  → Tree 다양성 증가 가능
  → 오류 상관과 variance 감소 가능
  → 개별 Tree의 최적 분할을 놓쳐 bias 증가 가능</code></pre>
<p>이 방향은 <strong>일반적인 경향이지 성능 보장이나 고정 순위가 아니다.</strong></p>
<h3 id="32-scikit-learn의-중요한-기본값-차이">3.2 scikit-learn의 중요한 기본값 차이</h3>
<p>scikit-learn에서 두 모델의 기본값은 다음과 다르다.</p>
<ul>
<li><code>RandomForestClassifier</code>: <code>bootstrap=True</code></li>
<li><code>ExtraTreesClassifier</code>: <code>bootstrap=False</code></li>
</ul>
<p>따라서 Extra Trees는 기본적으로 각 Tree가 전체 학습 표본을 사용한다. Tree 차이는 주로 노드별 Feature 후보와 Threshold 무작위성에서 나온다.</p>
<p>이 때문에 Extra Trees까지 엄밀한 의미의 “Bootstrap Bagging”으로 부르면 정확하지 않다. Bagging, Random Forest, Extra Trees를 함께 묶어 말하고 싶다면 다음 표현이 안전하다.</p>
<blockquote>
<p>순차 의존 없이 여러 randomized Tree를 만들고 평균하는 ensemble</p>
</blockquote>
<p>또는 scikit-learn 문서의 표현처럼 <strong>perturb-and-combine 방식의 randomized tree ensemble</strong>이라고 부를 수 있다.</p>
<hr>
<h2 id="4-random-forest와-extra-trees-비교">4. Random Forest와 Extra Trees 비교</h2>
<table>
<thead>
<tr>
<th>구분</th>
<th>Random Forest</th>
<th>Extra Trees</th>
</tr>
</thead>
<tbody><tr>
<td>학습 관계</td>
<td>Tree 간 순차 의존 없음</td>
<td>Tree 간 순차 의존 없음</td>
</tr>
<tr>
<td>표본</td>
<td>기본값 Bootstrap</td>
<td>기본값 전체 학습 표본</td>
</tr>
<tr>
<td>Feature 후보</td>
<td>노드마다 무작위 일부</td>
<td>노드마다 무작위 일부</td>
</tr>
<tr>
<td>Threshold</td>
<td>후보 Feature 안에서 좋은 값 탐색</td>
<td>Feature별 무작위 값 중 최선 선택</td>
</tr>
<tr>
<td>무작위성</td>
<td>강함</td>
<td>더 강함</td>
</tr>
<tr>
<td>계산</td>
<td>정확한 분할 탐색 비용</td>
<td>더 빠를 수 있음</td>
</tr>
<tr>
<td>bias–variance</td>
<td>데이터·설정에 따라 결정</td>
<td>variance가 더 낮고 bias가 더 높아질 수 있음</td>
</tr>
<tr>
<td>OOB</td>
<td>기본 설정에서 사용 가능</td>
<td><code>bootstrap=True</code>로 바꿔야 사용 가능</td>
</tr>
</tbody></table>
<p>표의 bias–variance 행은 방향성이다. 다음처럼 확정적인 화살표 순위를 붙이면 과도한 단순화가 된다.</p>
<pre><code class="language-text">Bagging variance ↓
Random Forest variance ↓↓
Extra Trees variance ↓↓↓</code></pre>
<p>실제 결과는 표본 수, 신호 강도, Feature 수, 잡음, Tree 깊이, <code>max_features</code> 등에 따라 달라진다.</p>
<hr>
<h2 id="5-scikit-learn으로-비교하기">5. scikit-learn으로 비교하기</h2>
<h3 id="51-공통-데이터">5.1 공통 데이터</h3>
<pre><code class="language-python">from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    stratify=y,
    random_state=42,
)</code></pre>
<h3 id="52-randomforestclassifier">5.2 RandomForestClassifier</h3>
<pre><code class="language-python">from sklearn.ensemble import RandomForestClassifier

random_forest = RandomForestClassifier(
    n_estimators=300,
    max_features=&quot;sqrt&quot;,
    bootstrap=True,
    oob_score=True,
    n_jobs=-1,
    random_state=42,
)

random_forest.fit(X_train, y_train)

print(&quot;RF test accuracy:&quot;, random_forest.score(X_test, y_test))
print(&quot;RF OOB accuracy:&quot;, random_forest.oob_score_)</code></pre>
<p><code>max_features=&quot;sqrt&quot;</code>는 각 노드의 분할을 찾을 때 대략 전체 Feature 수의 제곱근만큼을 후보로 고려한다는 뜻이다. 후보 집합은 노드마다 다시 무작위로 선택된다.</p>
<p>scikit-learn은 유효한 분할을 찾기 위해 필요한 경우 실질적으로 <code>max_features</code>보다 더 많은 Feature를 검사할 수 있다고 문서에 명시한다. 따라서 “어떤 상황에서도 정확히 그 개수만 본다”라고 표현하지 않는 편이 안전하다.</p>
<h3 id="53-extratreesclassifier">5.3 ExtraTreesClassifier</h3>
<pre><code class="language-python">from sklearn.ensemble import ExtraTreesClassifier

extra_trees = ExtraTreesClassifier(
    n_estimators=300,
    max_features=&quot;sqrt&quot;,
    bootstrap=False,
    n_jobs=-1,
    random_state=42,
)

extra_trees.fit(X_train, y_train)

print(&quot;Extra Trees test accuracy:&quot;, extra_trees.score(X_test, y_test))</code></pre>
<p><code>bootstrap=False</code>는 scikit-learn의 기본 동작을 명시한 것이다. 모든 Tree가 전체 학습 표본을 사용하더라도 Feature와 Threshold 무작위성 때문에 서로 다른 Tree가 만들어질 수 있다.</p>
<p>OOB 점수가 필요하다면 Bootstrap을 명시적으로 켠다.</p>
<pre><code class="language-python">extra_trees_oob = ExtraTreesClassifier(
    n_estimators=300,
    max_features=&quot;sqrt&quot;,
    bootstrap=True,
    oob_score=True,
    n_jobs=-1,
    random_state=42,
)</code></pre>
<h3 id="54-같은-평가-조건에서-비교하기">5.4 같은 평가 조건에서 비교하기</h3>
<p>한 번의 train/test split보다 같은 Cross Validation 조건에서 평균과 fold별 변동성을 함께 보는 편이 낫다.</p>
<pre><code class="language-python">from sklearn.model_selection import StratifiedKFold, cross_val_score

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

for name, model in {
    &quot;Random Forest&quot;: random_forest,
    &quot;Extra Trees&quot;: extra_trees,
}.items():
    scores = cross_val_score(model, X, y, cv=cv, scoring=&quot;accuracy&quot;)
    print(f&quot;{name:14s}: {scores.mean():.3f} ± {scores.std():.3f}&quot;)</code></pre>
<p>여기서 표준편차는 해당 fold 구성에서 측정한 <strong>성능 점수의 변동성</strong>이다. bias–variance 분해에서 말하는 모델 variance와 같은 값은 아니다.</p>
<hr>
<h2 id="6-선택할-때-주의할-점">6. 선택할 때 주의할 점</h2>
<ul>
<li>먼저 조정할 값은 <code>n_estimators</code>, <code>max_features</code>, <code>max_depth</code>, <code>min_samples_leaf</code>다.</li>
<li><code>max_features</code>가 작을수록 variance가 단조롭게 줄거나 성능이 좋아진다는 법칙은 없다.</li>
<li>“Clean이면 RF, Noisy면 Extra Trees”는 비교 가설일 뿐 모델 선택 규칙이 아니다.</li>
<li>Feature Randomness는 Tree 오류 상관을 낮출 수 있지만 개별 Tree의 예측력도 낮출 수 있다.</li>
<li>Tree 수를 늘리면 예측은 안정되는 경향이 있지만 개선은 포화되고 계산 비용은 증가한다.</li>
<li>Extra Trees는 <code>bootstrap=False</code>가 기본이므로 전형적인 Bootstrap Bagging과 같지 않다.</li>
<li>scikit-learn의 Random Forest 분류는 Tree별 class probability를 평균해 최종 class를 정한다.</li>
</ul>
<blockquote>
<p>이론은 비교할 가설을 제공하고, 실제 선택은 동일한 평가 조건의 Cross Validation 결과가 결정한다.</p>
</blockquote>
<hr>
<h2 id="7-이번-글의-핵심">7. 이번 글의 핵심</h2>
<h3 id="random-forest">Random Forest</h3>
<blockquote>
<p>노드별 Feature 후보 무작위화로 Tree들이 덜 비슷하게 움직이도록 만들고 평균의 효과를 키우려 한다.</p>
</blockquote>
<h3 id="extra-trees">Extra Trees</h3>
<blockquote>
<p>Feature뿐 아니라 분할 Threshold까지 무작위화해 더 강한 bias–variance trade-off를 만든다.</p>
</blockquote>
<h3 id="모델-선택">모델 선택</h3>
<blockquote>
<p>Random Forest와 Extra Trees의 우열은 데이터에 따라 달라지므로 같은 Cross Validation 조건에서 성능과 비용을 비교한다.</p>
</blockquote>
<p>Bagging, Random Forest, Extra Trees는 Tree들을 순차 의존 없이 만들고 예측을 평균하는 방식이다. 다음 글에서는 이들과 다른 질문을 던지는 <strong>Boosting</strong>을 살펴본다.</p>
<blockquote>
<p>현재 모델이 아직 틀리는 방향을 다음 Tree가 어떻게 보완할까?</p>
</blockquote>
<hr>
<h2 id="참고-자료">참고 자료</h2>
<ul>
<li><a href="https://www.stat.berkeley.edu/~breiman/randomforest2001.pdf">Leo Breiman, “Random Forests” (2001)</a></li>
<li><a href="https://people.montefiore.uliege.be/ernst/uploads/news/id63/extremely-randomized-trees.pdf">Geurts, Ernst &amp; Wehenkel, “Extremely randomized trees” (2006)</a></li>
<li><a href="https://scikit-learn.org/stable/modules/ensemble.html">scikit-learn: Ensembles user guide</a></li>
<li><a href="https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassifier.html">scikit-learn: RandomForestClassifier</a></li>
<li><a href="https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.ExtraTreesClassifier.html">scikit-learn: ExtraTreesClassifier</a></li>
</ul>
]]></description>
        </item>
        <item>
            <title><![CDATA[Decision Tree, Bagging과 OOB]]></title>
            <link>https://velog.io/@def_uk/Decision-Tree-Bagging%EA%B3%BC-OOB</link>
            <guid>https://velog.io/@def_uk/Decision-Tree-Bagging%EA%B3%BC-OOB</guid>
            <pubDate>Tue, 01 Sep 2026 13:59:30 GMT</pubDate>
            <description><![CDATA[<blockquote>
<h3 id="학습지표">학습지표</h3>
<p><strong>Decision Tree</strong>의 높은 variance를 줄이기 위해 <strong>Bagging</strong>은 데이터에 randomness를 주고, <strong>Random Forest</strong>는 feature 후보에도 randomness를 추가하며, <strong>Extra Trees</strong>는 분할 threshold까지 randomness를 추가한다.</p>
</blockquote>
<p>Decision Tree는 비선형 관계와 변수 사이의 상호작용을 자연스럽게 표현한다. 전처리가 비교적 적고, 사람이 규칙을 따라갈 수 있다는 장점도 있다.</p>
<p>그러나 한 가지 약점이 크다. <strong>학습 표본이 조금만 달라져도 Tree 구조와 예측이 크게 달라질 수 있다.</strong></p>
<p>이번 글에서는 이 문제에서 출발해 다음 질문에 답한다.</p>
<ol>
<li>Decision Tree는 왜 흔들리는가?</li>
<li>Bagging은 왜 Tree를 여러 개 만드는가?</li>
<li>Bootstrap과 Aggregation은 각각 무슨 역할을 하는가?</li>
<li>OOB는 어디서 생기고 무엇을 평가하는가?</li>
</ol>
<hr>
<h2 id="1-decision-tree는-왜-불안정할까">1. Decision Tree는 왜 불안정할까?</h2>
<h3 id="11-상위-분할-하나가-전체-구조를-바꾼다">1.1 상위 분할 하나가 전체 구조를 바꾼다</h3>
<p>Decision Tree는 각 노드에서 데이터를 가장 잘 나누는 질문을 탐욕적(Greedy Algorithm)으로 선택한다.</p>
<pre><code class="language-text">petal width &lt; 0.8인가?</code></pre>
<p>최적 후보들의 점수가 비슷하다면 학습 데이터 몇 개만 추가되거나 빠져도 선택되는 분할이 달라질 수 있다. 상위 노드의 분할이 바뀌면 자식 노드가 받는 데이터도 달라지므로 이후 가지 구조가 연쇄적으로 변한다.</p>
<pre><code class="language-text">원본 데이터 D      → Tree A → Decision Boundary A
조금 바뀐 데이터 D&#39; → Tree B → Decision Boundary B</code></pre>
<p>이 성질을 통계적인 언어로는 <strong>variance가 높다</strong>고 표현한다.</p>
<p>여기서 variance는 데이터 열의 분산이 아니다.</p>
<blockquote>
<p>같은 모집단에서 학습 표본을 다시 뽑아 모델을 재학습한다면 예측이 얼마나 달라지는가?</p>
</blockquote>
<p>깊은 Tree는 복잡한 패턴을 잡아 bias가 낮을 수 있지만, 표본의 우연한 변화와 잡음까지 따라가기 쉬워 variance가 커질 수 있다.</p>
<pre><code class="language-text">복잡하고 깊은 Tree: bias가 낮을 수 있지만 variance가 커질 수 있음
단순하고 얕은 Tree: variance는 작아질 수 있지만 bias가 커질 수 있음</code></pre>
<p>이는 절대적인 법칙이라기보다 모델 복잡도를 조절할 때 나타나는 대표적인 trade-off다.</p>
<h3 id="12-tree-하나를-단순하게-만들면-되지-않을까">1.2 Tree 하나를 단순하게 만들면 되지 않을까?</h3>
<p><code>max_depth</code>, <code>min_samples_leaf</code> 같은 제약으로 Tree를 단순하게 만들면 흔들림을 줄일 수 있다. 그러나 지나치게 제한하면 중요한 비선형 관계까지 놓쳐 underfitting이 생길 수 있다.</p>
<p>Bagging은 다른 질문을 던진다.</p>
<blockquote>
<p>개별 Tree의 표현력을 크게 포기하지 않으면서 여러 Tree의 흔들림을 결합 과정에서 상쇄할 수는 없을까?</p>
</blockquote>
<p>이 질문에 대한 답이 <strong>Bootstrap Aggregating</strong>, 줄여서 <strong>Bagging</strong>이다.</p>
<hr>
<h2 id="2-bagging-서로-다르게-흔들리게-한-뒤-평균낸다">2. Bagging: 서로 다르게 흔들리게 한 뒤 평균낸다</h2>
<p>Bagging은 같은 종류의 기본 모델을 서로 다른 Bootstrap 표본에 학습하고 예측을 결합한다.</p>
<pre><code class="language-text">원본 학습 데이터
  ├─ Bootstrap 표본 1 → Tree 1 ─┐
  ├─ Bootstrap 표본 2 → Tree 2 ─┼→ 평균 또는 투표
  └─ Bootstrap 표본 3 → Tree 3 ─┘</code></pre>
<p>Tree들은 앞 Tree의 결과를 다음 Tree가 이어받지 않는다. 따라서 <strong>순차 의존 없이 별도로 학습할 수 있고 병렬화하기 쉽다.</strong></p>
<p>이를 “Tree들이 통계적으로 독립적이다”라고 표현하면 정확하지 않다. 같은 원본 데이터와 Feature를 공유하므로 예측 오류는 서로 상관될 수 있다. 이 상관 문제는 다음 글의 Random Forest가 다룬다.</p>
<h3 id="21-같은-tree를-복사하는-것만으로는-부족하다">2.1 같은 Tree를 복사하는 것만으로는 부족하다</h3>
<p>같은 데이터, 같은 알고리즘, 같은 설정으로 Tree를 100개 만들면 동일하거나 매우 비슷한 Tree가 만들어질 수 있다.</p>
<pre><code class="language-text">같은 학습 데이터
  ├─ Tree 1 ─┐
  ├─ Tree 2 ─┤ 비슷한 예측
  └─ Tree 3 ─┘</code></pre>
<p>같은 실수를 하는 모델 100개를 모아도 평균의 효과는 제한된다. Bagging은 먼저 학습 표본을 바꿔 각 모델에 서로 다른 학습 경험을 제공한다.</p>
<hr>
<h2 id="3-bootstrap-sampling-서로-다른-학습-표본-만들기">3. Bootstrap Sampling: 서로 다른 학습 표본 만들기</h2>
<p>Bootstrap은 원본 학습 데이터에서 <strong>복원 추출(with replacement)</strong>하는 방법이다.</p>
<p>원본 데이터가 다음과 같다고 하자.</p>
<pre><code class="language-text">[A, B, C, D, E]</code></pre>
<p>크기가 5인 Bootstrap 표본은 다음처럼 만들어질 수 있다.</p>
<pre><code class="language-text">표본 1: [A, A, C, D, E]
표본 2: [B, B, C, C, E]
표본 3: [A, B, D, D, D]</code></pre>
<p>한 번 뽑은 관측치를 다시 뽑을 수 있으므로 중복이 생긴다. 반대로 어떤 관측치는 한 번도 선택되지 않을 수 있다.</p>
<h3 id="max_samples08의-정확한-의미"><code>max_samples=0.8</code>의 정확한 의미</h3>
<p>학습 데이터가 1,000개이고 다음과 같이 설정했다고 하자.</p>
<pre><code class="language-python">max_samples=0.8
bootstrap=True</code></pre>
<p>각 모델은 원본의 80%에 해당하는 <strong>800번을 복원 추출</strong>한다. 서로 다른 관측치 800개가 반드시 포함된다는 뜻은 아니다. 중복 때문에 실제 고유 관측치 수는 800개보다 적다.</p>
<p>Bootstrap의 목적은 개별 Tree 하나를 반드시 더 정확하게 만드는 것이 아니다.</p>
<blockquote>
<p>Tree들을 서로 다르게 흔들리게 만들어, 결합할 때 오차가 일부 상쇄될 조건을 만드는 것이다.</p>
</blockquote>
<hr>
<h2 id="4-aggregation-흔들림을-안정성으로-바꾸기">4. Aggregation: 흔들림을 안정성으로 바꾸기</h2>
<p>Bootstrap 표본 <code>D_1, D_2, ..., D_B</code>로 모델 <code>f_1, f_2, ..., f_B</code>를 학습했다고 하자.</p>
<p>회귀에서는 예측값을 평균한다.</p>
<p>$$
\hat f_{\text{bag}}(x)=\frac{1}{B}\sum_{b=1}^{B}f_b(x)
$$
$$
(B) = Tree 개수
$$
$$
(f_b(x)) = b번째 Tree가 입력 (x)에 대해 낸 예측값
$$
$$
(\sum_{b=1}^{B} f_b(x)) = 모든 Tree 예측을 다 더함
$$
(\frac{1}{B}) = Tree 개수로 나눔
$$
(\hat f_{\text{bag}}(x)) = Bagging의 최종 예측
$$</p>
<p>예를 들어 실제값이 50이고 다섯 Tree의 예측이 다음과 같다면:</p>
<pre><code class="language-text">Tree 1 → 43
Tree 2 → 58
Tree 3 → 48
Tree 4 → 55
Tree 5 → 46
평균   → 50</code></pre>
<p>개별 Tree의 오차 방향이 완전히 같지 않다면 평균 과정에서 일부 오차가 상쇄된다. 그래서 Bagging을 대표적인 <strong>variance reduction 기법</strong>이라고 부른다.</p>
<p>분류의 결합 방식은 설명 층위를 구분할 필요가 있다.</p>
<ul>
<li>Breiman의 원래 Bagging 설명: 가장 많은 표를 받은 클래스를 선택하는 plurality vote</li>
<li>scikit-learn <code>BaggingClassifier</code>: 기본 모델이 <code>predict_proba()</code>를 제공하면 클래스 확률을 평균하고, 제공하지 않으면 voting 사용</li>
</ul>
<p>즉 “분류는 무조건 hard voting”이라고 외우기보다 사용하는 구현의 결합 방식을 확인해야 한다.</p>
<hr>
<h2 id="5-oob-bootstrap에서-자연스럽게-생기는-내부-추정치">5. OOB: Bootstrap에서 자연스럽게 생기는 내부 추정치</h2>
<p>원본 데이터가 <code>N</code>개이고 크기가 <code>N</code>인 Bootstrap 표본을 만든다고 하자. 특정 관측치 하나가 한 번도 선택되지 않을 확률은 다음과 같다.</p>
<p>$$
\left(1-\frac{1}{N}\right)^N \xrightarrow[N\to\infty]{} \frac{1}{e}\approx0.368
$$</p>
<p>따라서 표준적인 Bootstrap 표본 하나에는 평균적으로 다음과 같은 구성이 나타난다.</p>
<pre><code class="language-text">한 번 이상 등장한 고유 관측치: 약 63.2%
한 번도 등장하지 않은 관측치: 약 36.8%</code></pre>
<p>한 Tree의 Bootstrap 표본에 포함되지 않은 관측치를 그 Tree의 <strong>Out-of-Bag(OOB) 표본</strong>이라고 한다.</p>
<h3 id="51-oob-평가는-어떻게-만들어질까">5.1 OOB 평가는 어떻게 만들어질까?</h3>
<p>각 학습 관측치에 대해 그 관측치를 보지 않고 학습한 Tree들의 예측만 모은다.</p>
<pre><code class="language-text">관측치 x_i
  ↓
x_i를 학습에 사용하지 않은 Tree만 선택
  ↓
그 Tree들의 예측을 집계
  ↓
y_i와 비교 → OOB 성능 추정치</code></pre>
<p>OOB는 별도의 모델 재학습 없이 얻을 수 있는 편리한 신호다. 그러나 역할을 과장하면 안 된다.</p>
<ul>
<li>36.8%는 <code>N</code>개에서 <code>N</code>번 복원 추출할 때의 근사값이다.</li>
<li><code>max_samples</code>가 달라지면 OOB 비율도 달라진다.</li>
<li>OOB는 원래 학습 풀 내부에서 만든 추정치다.</li>
<li>최종 독립 테스트 세트인 final holdout과 같지 않다.</li>
<li>모든 상황에서 Cross Validation을 대체한다는 규칙도 아니다.</li>
</ul>
<p>OOB는 “공짜 테스트 세트”보다 <strong>Bootstrap 과정에서 얻는 내부 보조 성능 추정치</strong>라고 부르는 편이 정확하다.</p>
<hr>
<h2 id="6-scikit-learn으로-bagging과-oob-확인하기">6. scikit-learn으로 Bagging과 OOB 확인하기</h2>
<h3 id="61-데이터-준비">6.1 데이터 준비</h3>
<pre><code class="language-python">from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    stratify=y,
    random_state=42,
)</code></pre>
<h3 id="62-단일-decision-tree">6.2 단일 Decision Tree</h3>
<pre><code class="language-python">from sklearn.tree import DecisionTreeClassifier

tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)

print(&quot;Decision Tree test accuracy:&quot;, tree.score(X_test, y_test))</code></pre>
<p>한 번의 정확도만으로 모델의 안정성을 판단할 수는 없다. 여러 데이터 분할에서 예측이나 성능이 얼마나 달라지는지도 확인해야 한다.</p>
<h3 id="63-baggingclassifier">6.3 BaggingClassifier</h3>
<pre><code class="language-python">from sklearn.ensemble import BaggingClassifier
from sklearn.tree import DecisionTreeClassifier

bagging = BaggingClassifier(
    estimator=DecisionTreeClassifier(random_state=42),
    n_estimators=100,
    max_samples=1.0,
    bootstrap=True,
    oob_score=True,
    n_jobs=-1,
    random_state=42,
)

bagging.fit(X_train, y_train)

print(&quot;Bagging test accuracy:&quot;, bagging.score(X_test, y_test))
print(&quot;Bagging OOB accuracy:&quot;, bagging.oob_score_)</code></pre>
<table>
<thead>
<tr>
<th>파라미터</th>
<th>의미</th>
</tr>
</thead>
<tbody><tr>
<td><code>estimator</code></td>
<td>반복해서 학습할 기본 모델</td>
</tr>
<tr>
<td><code>n_estimators</code></td>
<td>기본 모델 수</td>
</tr>
<tr>
<td><code>max_samples</code></td>
<td>각 모델을 위해 뽑을 표본 수 또는 비율</td>
</tr>
<tr>
<td><code>bootstrap=True</code></td>
<td>표본을 복원 추출</td>
</tr>
<tr>
<td><code>oob_score=True</code></td>
<td>OOB 표본으로 내부 점수 계산</td>
</tr>
<tr>
<td><code>n_jobs=-1</code></td>
<td>사용 가능한 CPU 코어 활용</td>
</tr>
<tr>
<td><code>random_state</code></td>
<td>무작위 과정 재현</td>
</tr>
</tbody></table>
<p>scikit-learn 1.2 이전 코드에서는 <code>estimator</code> 대신 <code>base_estimator</code>가 보일 수 있다. 현재 API에서는 <code>estimator</code>를 사용한다.</p>
<p>Iris처럼 작고 쉬운 데이터에서는 두 모델의 한 번의 test accuracy가 같을 수도 있다. 이 결과 하나로 “Bagging이 항상 더 좋다”거나 “효과가 없다”고 결론 내리면 안 된다.</p>
<hr>
<h2 id="7-이번-글의-핵심">7. 이번 글의 핵심</h2>
<h3 id="decision-tree">Decision Tree</h3>
<blockquote>
<p>학습 표본이 조금 달라질 때 예측이 크게 바뀔 수 있는 고분산 모델이다.</p>
</blockquote>
<h3 id="bagging">Bagging</h3>
<blockquote>
<p>불안정한 모델을 버리는 대신 서로 다른 Bootstrap 표본에서 다르게 흔들리게 만들고, 예측을 집계해 흔들림을 줄인다.</p>
</blockquote>
<h3 id="oob">OOB</h3>
<blockquote>
<p>각 모델이 학습 중 보지 않은 표본의 예측만 모아 만든 학습 풀 내부의 보조 성능 추정치다.</p>
</blockquote>
<p>Bagging은 Tree뿐 아니라 여러 기본 추정기에 적용할 수 있다. 다만 표본 변화에 민감한 고분산 모델에서 특히 유용해 Decision Tree와 자주 결합한다. Tree 수를 늘리면 예측은 안정되는 경향이 있지만 개선은 포화되고 계산 비용은 증가한다.</p>
<p>그러나 Bootstrap을 사용해도 강력한 Feature 때문에 Tree들이 계속 비슷한 분할을 선택할 수 있다. Tree들의 오류가 비슷하면 평균 효과도 제한된다.</p>
<p>다음 글에서는 이 문제를 해결하기 위해 <strong>노드별 Feature 후보를 무작위화하는 Random Forest</strong>와 <strong>분할 임계값까지 무작위화하는 Extra Trees</strong>를 비교한다.</p>
<hr>
<h2 id="참고-자료">참고 자료</h2>
<ul>
<li><a href="https://doi.org/10.1007/BF00058655">Leo Breiman, “Bagging predictors” (1996)</a></li>
<li><a href="https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.BaggingClassifier.html">scikit-learn: BaggingClassifier</a></li>
<li><a href="https://scikit-learn.org/stable/auto_examples/ensemble/plot_ensemble_oob.html">scikit-learn: OOB Errors for Random Forests</a></li>
<li><a href="https://scikit-learn.org/stable/modules/cross_validation.html">scikit-learn: Cross-validation guide</a></li>
</ul>
]]></description>
        </item>
    </channel>
</rss>