The Annotated Transformer
本文仅摘录原始 blog 部分代码并添加部分注释,不涉及原文全部内容,推荐阅读原文。
Model Architecture
1 | |
1 | |
Encoder and Decoder Stacks
Encoder
1 | |
1 | |
1 | |
Decoder
1 | |
1 | |
Decoder 自注意力层 mask 生成:
1 | |
Attention
1 | |
MHA:
1 | |
Position-wise Feed-Forward Networks
1 | |
Embeddings and Softmax
1 | |
Positional Encoding
1 | |
Full Model
1 | |
Inference
1 | |
输出: 1
2
3
4
5
6
7
8
9
10Example Untrained Model Prediction: tensor([[0, 0, 0, 0, 0, 0, 0, 0, 0, 0]])
Example Untrained Model Prediction: tensor([[0, 3, 4, 4, 4, 4, 4, 4, 4, 4]])
Example Untrained Model Prediction: tensor([[ 0, 10, 10, 10, 3, 2, 5, 7, 9, 6]])
Example Untrained Model Prediction: tensor([[ 0, 4, 3, 6, 10, 10, 2, 6, 2, 2]])
Example Untrained Model Prediction: tensor([[ 0, 9, 0, 1, 5, 10, 1, 5, 10, 6]])
Example Untrained Model Prediction: tensor([[ 0, 1, 5, 1, 10, 1, 10, 10, 10, 10]])
Example Untrained Model Prediction: tensor([[ 0, 1, 10, 9, 9, 9, 9, 9, 1, 5]])
Example Untrained Model Prediction: tensor([[ 0, 3, 1, 5, 10, 10, 10, 10, 10, 10]])
Example Untrained Model Prediction: tensor([[ 0, 3, 5, 10, 5, 10, 4, 2, 4, 2]])
Example Untrained Model Prediction: tensor([[0, 5, 6, 2, 5, 6, 2, 6, 2, 2]])
Model Training
Batches and Masking
1 | |
Training Loop
1 | |
Optimizer
1 | |
Regularization
1 | |
Loss Computation
1 | |
原文后续内容在此忽略,本文重点关注模型结构。
Summary:自底向上视角回顾
论文是自顶向下描述,但实际写代码通常自底向上:先造基础模块,再逐层组装。
实现顺序与关键点
- 基础工具
clones:deepcopyN 个独立层,放进nn.ModuleList,避免参数共享。LayerNorm:对最后一维归一化,a_2/b_2可学习,eps防除零。
- 输入表示
Embeddings:整数索引 →d_model,乘sqrt(d_model)平衡尺度;可与输出层权重共享。PositionalEncoding:固定 sin/cos,register_buffer不更新梯度;对数空间算div_term,与 embedding 相加后 dropout。
- 核心计算
attention:QK^T / sqrt(d_k)→ mask 填-1e9→ softmax → 加权V。subsequent_mask:下三角 True,上三角 False,防止看未来。MultiHeadedAttention:4 个线性层;view + transpose拆头为(batch, h, seq, d_k);并行 attention;再transpose + contiguous + view合并,最后过W^O。PositionwiseFeedForward:Linear → ReLU → Dropout → Linear,逐位置独立。
- 层组装
SublayerConnection:博客采用 Pre-norm:x + Dropout(Sublayer(LayerNorm(x)));原论文是 Post-norm。EncoderLayer:Self-Attn + FFN,各包一个SublayerConnection。DecoderLayer:Masked Self-Attn + Cross-Attn(Q 来自 decoder,K/V 来自 encoder memory)+ FFN,三个SublayerConnection。
- 堆叠与顶层
Encoder/Decoder:clones(layer, N)串行,最后加 LayerNorm。Generator:Linear(d_model, vocab)+log_softmax。EncoderDecoder:只调度encode/decode。make_model:组装所有组件;各层用deepcopy保证独立;Xavier 初始化只对dim > 1参数。
依赖关系图
1 | |