LLM Systems 06: Transformer
0. Recap
- Learning parameters of an NN needs gradient calculation
- Computation Graph
- to perform computation: topological traversal along the DAG
- Auto Differentiation
- building backward computation graph for gradient calculation
- Put together: Deep Learning Framework
- Define program i.e., symbolic computation graph w/ placeholders/variable/operation nodes
- Executes (optimized) computation graph on a set of available devices
1. Encoder-decoder Architecture
以翻译任务为例,encoder-decoder 范式的序列转换模型核心是条件概率建模与自回归生成:
- Conditional text generation: directly learning a function mapping from source sequence to target sequence:
\[ p_\theta(y|x) = \prod_t p(y_t | x, y_{1:t-1}; \theta) \]
- Previous encoder/decoder: LSTM or GRU

- RNN 循环生成
- 存在的问题:RNN 的循环结构限制了并行计算。
- Motivation:使用支持并行计算(并行 encoding,不依赖 RNN 逐个 token
循环计算隐藏表示)的模型结构
- Full context and parallel: use Attention in both encoder and decoder
- no recurrent ==> concurrent encoding
2. Transformer Model
整体流程:
- 输入 tokens 进入(encoder) Embedding Layer,加上 positional embedding,得到 tokens 向量表示;
- 进入 encoder Layer(MHA + FFN)进行 forward 计算,得到输入 tokens 的隐藏表示;
- decoder 侧,开始输入
bos特殊 token,得到 embeddings - 将 encoder 输出以及 decoder 侧 embeddings 输入 decoder MHA + FFN 层进行 forward 计算,经过 Softmax 预测下一个 token 概率;
- 将 decoder 输出 token 拼接,进行自回归生成,直到生成
eos特殊 token 或者达到最大输出长度。
Embedding
- Token Embedding:
- Shared (tied) input and output embedding from lookup table
- Positional Embedding:
- to distinguish words in different position, Map position labels to embedding, dimension is same as Tok Emb, for t-th pos, i-th dim:
\[ \begin{aligned} PE_{t, 2i} &= \sin\left(\frac{t}{1000^{2i/d}}\right) \\ PE_{t, 2i+1} &= \cos\left(\frac{t}{1000^{2i/d}}\right) \end{aligned} \]
位置编码可视化:
Multihead Attention, Decoder Self-Attention
- Instead of one vector for each token
- break into multiple heads
- each head perform attention
\[ \text{head}_i=\text{Attention}(QW_i^Q,KW_i^K,VW_i^V) \] \[ \text{MultiHead}(Q,K,V)=\text{Concat}(\text{head}_1,\dots,\text{head}_h)W^O \]
Decoder Masked Attention
- Maskout right side before softmax (-inf)
FFN
\[ \text{FFN}\left(x\right) = \max\left( 0, xW_1 + b_1 \right)W_2 + b_2 \]
Put Together
LayerNorm
- Residual Connection
- LayerNorm: Make it zero mean and unit variance within layer
- Post-norm:下面图左侧,Transformer paper 原始方法。
- 缺点:在深层网络中(如 24 层、48 层甚至更深),由于每一层都要经过 LayerNorm,梯度在反向传播时容易受到归一化操作的影响,导致梯度消失或训练不稳定。
- Pre-norm:下面图右侧,现代大语言模型(如 GPT-2、GPT-3、LLaMA 等)普遍采用的改进架构。
此外,还有 Sandwich-norm,在 FFN/Attention 前后都加入 norm。三种 norm 位置方式对比:
| 架构 | 公式 | 残差连接位置 | 特点 |
|---|---|---|---|
| Post-norm (原论文) | \(\text{LN}(x + \text{Sublayer}(x))\) | 跨过子层,但穿过 LayerNorm | 表达能力强,但深层梯度易消失,依赖 Warmup |
| Pre-norm (现代主流) | \(x + \text{Sublayer}(\text{LN}(x))\) | 完全直通,不穿过任何 LayerNorm | 训练极稳,易堆叠极深,但方差易累积 |
| Sandwich-norm | \(x + \text{LN}(\text{Sublayer}(\text{LN}(x)))\) | 完全直通,不穿过任何 LayerNorm | 兼顾稳定性与表达能力,但计算开销稍大 |
3. Training Techniques and Performance of Transformer
Training Objective & Loss
以翻译任务为例,训练目标是最大化目标序列的条件概率:
\[ P(Y|X) = \prod_t P(y_t | y_{<t}, x) \]
训练损失采用交叉熵损失(Cross-Entropy Loss):
\[ l = -\sum_n \sum_t \log f_\theta(x_n, y_{n,1}, \dots, y_{n,t-1}) \]
Teacher-forcing during training:训练时采用 Teacher-forcing,即 “pretend to know groundtruth for prefix”。无论模型前一步预测结果如何,当前步都直接使用真实的 prefix 作为 decoder 输入,以稳定训练并加速收敛。
注:训练时输入 decoder 的 target 序列会进行 shifted right,并配合 Masked Attention 避免看到未来信息。
Training Techniques (Regularization)
Dropout - 应用于每个子层的输出(在残差连接之前)以及 embedding 和 positional embedding 的和。 - 概率 \(p=0.1 \sim 0.3\)。
Label Smoothing - 假设 \(y \in \mathbb{R}^n\) 是 one-hot 编码:\(y_i = 1\) if belongs to class \(i\),else \(0\)。 - 直接用 softmax 逼近 0/1 值很困难(容易导致模型过于自信和过拟合)。 - 引入平滑后的目标分布,通常取 \(\epsilon = 0.1\):
\[ y_i = \begin{cases} 1 - \epsilon & \text{if belongs to class } i \\ \frac{\epsilon}{n - 1} & \text{otherwise} \end{cases} \]
Training Setup (Hardware & Batching)
Batch - 按近似句子长度分组(group by approximate sentence length)。 - 仍然需要 shuffling 避免同质化批次。
Hardware - 一台机器 8 GPUs (in 2017 paper)。 - Base model: 100k steps (12 hours)。 - Large model: 300k steps (3.5 days)。
Optimizer (Adam with Warmup)
使用 Adam 优化器,并在训练初期采用 warmup 策略(increase learning rate during warmup, then decrease):
\[ \eta = \frac{1}{\sqrt{d}} \min\left( \frac{1}{\sqrt{t}}, \frac{t}{\sqrt{t_0^3}} \right) \]
其中 \(d\) 是 model dimension,\(t\) 是当前 step,\(t_0\) 是 warmup steps。
Adam Optimizer Formulas
Adam 优化器的具体计算过程如下:
\[ \begin{aligned} m_{t+1} &= \beta_1 m_t + (1 - \beta_1) \nabla \ell(x_t) \\ v_{t+1} &= \beta_2 v_t + (1 - \beta_2) (\nabla \ell(x_t))^2 \\ \hat{m}_{t+1} &= \frac{m_{t+1}}{1 - \beta_1^{t+1}} \\ \hat{v}_{t+1} &= \frac{v_{t+1}}{1 - \beta_2^{t+1}} \\ x_{t+1} &= x_t - \frac{\eta}{\sqrt{\hat{v}_{t+1}} + \epsilon} \hat{m}_{t+1} \end{aligned} \]
注:原论文中 Adam 超参数为 \(\beta_1 = 0.9, \beta_2 = 0.98, \epsilon = 10^{-9}\),warmup steps \(t_0 = 4000\)。
4. Code walkthrough
5. Summary
- Sequence-to-sequence encoder-decoder framework for conditional generation, including Machine Translation
- Key components in Transformer (why each?)
- Positional Embedding (to distinguish tokens at different pos)
- Multihead attention:进行 token 间 hidden states 融合操作(上下文信息交互)
- Residual connection
- layer norm
- FFN:进行 token 内 hidden state 变换(单个 token 特征加工与模式表达)