LLM Systems 06: Transformer

0. Recap

  • Learning parameters of an NN needs gradient calculation
  • Computation Graph
    • to perform computation: topological traversal along the DAG
  • Auto Differentiation
    • building backward computation graph for gradient calculation
  • Put together: Deep Learning Framework
    • Define program i.e., symbolic computation graph w/ placeholders/variable/operation nodes
    • Executes (optimized) computation graph on a set of available devices

1. Encoder-decoder Architecture

以翻译任务为例,encoder-decoder 范式的序列转换模型核心是条件概率建模与自回归生成:


  • Conditional text generation: directly learning a function mapping from source sequence to target sequence:

\[ p_\theta(y|x) = \prod_t p(y_t | x, y_{1:t-1}; \theta) \]

  • Previous encoder/decoder: LSTM or GRU
    • RNN 循环生成
    • 存在的问题:RNN 的循环结构限制了并行计算。
    • Motivation:使用支持并行计算(并行 encoding,不依赖 RNN 逐个 token 循环计算隐藏表示)的模型结构
      • Full context and parallel: use Attention in both encoder and decoder
      • no recurrent ==> concurrent encoding

2. Transformer Model


整体流程:

  • 输入 tokens 进入(encoder) Embedding Layer,加上 positional embedding,得到 tokens 向量表示;
  • 进入 encoder Layer(MHA + FFN)进行 forward 计算,得到输入 tokens 的隐藏表示;
  • decoder 侧,开始输入 bos 特殊 token,得到 embeddings
  • 将 encoder 输出以及 decoder 侧 embeddings 输入 decoder MHA + FFN 层进行 forward 计算,经过 Softmax 预测下一个 token 概率;
  • 将 decoder 输出 token 拼接,进行自回归生成,直到生成 eos 特殊 token 或者达到最大输出长度。

Embedding

  • Token Embedding:
    • Shared (tied) input and output embedding from lookup table
  • Positional Embedding:
    • to distinguish words in different position, Map position labels to embedding, dimension is same as Tok Emb, for t-th pos, i-th dim:

\[ \begin{aligned} PE_{t, 2i} &= \sin\left(\frac{t}{1000^{2i/d}}\right) \\ PE_{t, 2i+1} &= \cos\left(\frac{t}{1000^{2i/d}}\right) \end{aligned} \]

位置编码可视化:

Multihead Attention, Decoder Self-Attention

  • Instead of one vector for each token
  • break into multiple heads
  • each head perform attention

\[ \text{head}_i=\text{Attention}(QW_i^Q,KW_i^K,VW_i^V) \] \[ \text{MultiHead}(Q,K,V)=\text{Concat}(\text{head}_1,\dots,\text{head}_h)W^O \]

Decoder Masked Attention

  • Maskout right side before softmax (-inf)

FFN

\[ \text{FFN}\left(x\right) = \max\left( 0, xW_1 + b_1 \right)W_2 + b_2 \]

Put Together

LayerNorm

  • Residual Connection
  • LayerNorm: Make it zero mean and unit variance within layer
  • Post-norm:下面图左侧,Transformer paper 原始方法。
    • 缺点:在深层网络中(如 24 层、48 层甚至更深),由于每一层都要经过 LayerNorm,梯度在反向传播时容易受到归一化操作的影响,导致梯度消失或训练不稳定。
  • Pre-norm:下面图右侧,现代大语言模型(如 GPT-2、GPT-3、LLaMA 等)普遍采用的改进架构。


此外,还有 Sandwich-norm,在 FFN/Attention 前后都加入 norm。三种 norm 位置方式对比:

架构 公式 残差连接位置 特点
Post-norm (原论文) \(\text{LN}(x + \text{Sublayer}(x))\) 跨过子层,但穿过 LayerNorm 表达能力强,但深层梯度易消失,依赖 Warmup
Pre-norm (现代主流) \(x + \text{Sublayer}(\text{LN}(x))\) 完全直通,不穿过任何 LayerNorm 训练极稳,易堆叠极深,但方差易累积
Sandwich-norm \(x + \text{LN}(\text{Sublayer}(\text{LN}(x)))\) 完全直通,不穿过任何 LayerNorm 兼顾稳定性与表达能力,但计算开销稍大

3. Training Techniques and Performance of Transformer

Training Objective & Loss

以翻译任务为例,训练目标是最大化目标序列的条件概率:

\[ P(Y|X) = \prod_t P(y_t | y_{<t}, x) \]

训练损失采用交叉熵损失(Cross-Entropy Loss):

\[ l = -\sum_n \sum_t \log f_\theta(x_n, y_{n,1}, \dots, y_{n,t-1}) \]

Teacher-forcing during training:训练时采用 Teacher-forcing,即 “pretend to know groundtruth for prefix”。无论模型前一步预测结果如何,当前步都直接使用真实的 prefix 作为 decoder 输入,以稳定训练并加速收敛。

注:训练时输入 decoder 的 target 序列会进行 shifted right,并配合 Masked Attention 避免看到未来信息。

Training Techniques (Regularization)

Dropout - 应用于每个子层的输出(在残差连接之前)以及 embedding 和 positional embedding 的和。 - 概率 \(p=0.1 \sim 0.3\)。

Label Smoothing - 假设 \(y \in \mathbb{R}^n\) 是 one-hot 编码:\(y_i = 1\) if belongs to class \(i\),else \(0\)。 - 直接用 softmax 逼近 0/1 值很困难(容易导致模型过于自信和过拟合)。 - 引入平滑后的目标分布,通常取 \(\epsilon = 0.1\):

\[ y_i = \begin{cases} 1 - \epsilon & \text{if belongs to class } i \\ \frac{\epsilon}{n - 1} & \text{otherwise} \end{cases} \]

Training Setup (Hardware & Batching)

Batch - 按近似句子长度分组(group by approximate sentence length)。 - 仍然需要 shuffling 避免同质化批次。

Hardware - 一台机器 8 GPUs (in 2017 paper)。 - Base model: 100k steps (12 hours)。 - Large model: 300k steps (3.5 days)。

Optimizer (Adam with Warmup)

使用 Adam 优化器,并在训练初期采用 warmup 策略(increase learning rate during warmup, then decrease):

\[ \eta = \frac{1}{\sqrt{d}} \min\left( \frac{1}{\sqrt{t}}, \frac{t}{\sqrt{t_0^3}} \right) \]

其中 \(d\) 是 model dimension,\(t\) 是当前 step,\(t_0\) 是 warmup steps。

Adam Optimizer Formulas

Adam 优化器的具体计算过程如下:

\[ \begin{aligned} m_{t+1} &= \beta_1 m_t + (1 - \beta_1) \nabla \ell(x_t) \\ v_{t+1} &= \beta_2 v_t + (1 - \beta_2) (\nabla \ell(x_t))^2 \\ \hat{m}_{t+1} &= \frac{m_{t+1}}{1 - \beta_1^{t+1}} \\ \hat{v}_{t+1} &= \frac{v_{t+1}}{1 - \beta_2^{t+1}} \\ x_{t+1} &= x_t - \frac{\eta}{\sqrt{\hat{v}_{t+1}} + \epsilon} \hat{m}_{t+1} \end{aligned} \]

注:原论文中 Adam 超参数为 \(\beta_1 = 0.9, \beta_2 = 0.98, \epsilon = 10^{-9}\),warmup steps \(t_0 = 4000\)。

4. Code walkthrough

5. Summary

  • Sequence-to-sequence encoder-decoder framework for conditional generation, including Machine Translation
  • Key components in Transformer (why each?)
    • Positional Embedding (to distinguish tokens at different pos)
    • Multihead attention:进行 token 间 hidden states 融合操作(上下文信息交互)
    • Residual connection
    • layer norm
    • FFN:进行 token 内 hidden state 变换(单个 token 特征加工与模式表达)

LLM Systems 06: Transformer
https://arcsin2.cloud/posts/2026/10/584140393/
作者
arcsin2
发布于
2026年10月5日
许可协议