跳转至

Attention Is All You Need

Ashish Vaswani et al.
NIPS 2017


Problem Statement

  • Traditional sequence models rely on RNNs/CNNs with sequential computation
  • Recurrent networks have inherent parallelization limitations
  • Convolutional models require increasing layers for long-range dependencies
  • Attention mechanisms previously used alongside recurrence/convolution

Key Motivation: Self-Attention Advantages

  • Parallelization: O(1) sequential operations vs O(n) for RNNs
  • Long-range dependencies: Constant path length between positions
  • Computational efficiency: Better than RNNs for typical sequence lengths
  • Interpretability: Attention distributions reveal linguistic structure

Transformer Architecture Overview


Encoder-decoder structure with stacked self-attention and feed-forward layers


Encoder Structure

  • Stack of 6 identical layers with two sub-layers: multi-head self-attention and position-wise feed-forward network
  • Residual connections + layer normalization around each sub-layer
  • Input: Embeddings + positional encodings (dmodel=512)

Decoder Structure

  • Stack of 6 identical layers with three sub-layers: masked self-attention, encoder-decoder attention, and feed-forward network
  • Masked self-attention prevents access to future positions
  • Same residual connections and layer normalization as encoder

Attention Mechanism

  • Maps queries, keys, values to output via weighted sum of values
  • Two common approaches: additive (feed-forward) and dot-product (faster)
  • Transformer uses Scaled Dot-Product Attention: Attention(Q,K,V) = softmax((QKT)/√dk)V, with scaling to prevent gradient issues

Multi-Head Attention

  • Projects Q, K, V h times with different linear projections (h=8 heads)
  • Performs attention in parallel on projected subspaces
  • Concatenates results and applies final linear projection
  • dk=dv=dmodel/h=64 in base model, enabling diverse dependency modeling

Attention Applications

  • Encoder self-attention: Each position attends to all encoder positions
  • Decoder self-attention: Each position attends to previous decoder positions
  • Encoder-decoder attention: Decoder positions attend to all encoder positions

Positional Encoding

  • Injects sequence order information (no recurrence/convolution)
  • Uses sine/cosine functions: PE(pos,2i)=sin(pos/10000^(2i/dmodel)), PE(pos,2i+1)=cos(...)
  • Allows model to learn relative position relationships

Training Setup

  • Datasets: WMT 2014 EN-DE (4.5M pairs), EN-FR (36M pairs)
  • Hardware: 8 NVIDIA P100 GPUs (12h for base, 3.5d for big model)
  • Optimizer: Adam with β1=0.9, β2=0.98, and warmup learning rate
  • Regularization: Residual dropout (Pdrop=0.1) and label smoothing (ϵls=0.1)

Machine Translation Results

Model EN-DE BLEU EN-FR BLEU Training Cost (FLOPs)
ConvS2S Ensemble 26.36 41.29 7.7×10¹⁹
GNMT + RL Ensemble 26.30 41.16 1.8×10²⁰
Transformer (big) 28.4 41.8 2.3×10¹⁹
  • 2+ BLEU improvement over previous state-of-the-art on EN-DE
  • New single-model state-of-the-art on EN-FR with 1/4 training cost

Model Variations (Ablation Study)

Configuration Dev PPL Dev BLEU
Base model 4.92 25.8
Single attention head 5.29 24.9
No dropout 4.67 25.3
Big model (dmodel=1024) 4.33 26.4
  • Multi-head attention critical for performance
  • Dropout prevents overfitting
  • Larger model dimensions improve translation quality

Long-Distance Dependency Attention


Encoder self-attention (layer 5) showing 'making' attending to distant 'difficult'


Anaphora Resolution Attention


Attention heads 5 and 6 resolving 'its' reference to 'The Law'


English Constituency Parsing

Parser Training WSJ 23 F1
RNN Grammar [8] WSJ only 91.7
Transformer (4 layers) WSJ only 91.3
Previous semi-supervised Semi-supervised 92.1
Transformer (4 layers) Semi-supervised 92.7
  • Strong performance without task-specific tuning
  • Outperforms RNN sequence-to-sequence models in small-data regime

Computational Complexity

Layer Type Complexity Sequential Ops Max Path Length
Self-Attention O(n²·d) O(1) O(1)
Recurrent O(n·d²) O(n) O(n)
Convolutional O(k·n·d²) O(1) O(logk(n))
  • Self-attention enables better parallelization than RNNs
  • Constant path length improves long-range dependency learning

Limitations and Future Work

  • Limitations: Quadratic complexity with sequence length; less effective for very long sequences
  • Future directions: Extend to other modalities (images, audio, video); develop local/restricted attention
  • Key achievement: Transformer achieves new SOTA in translation with less training time
  • Broader impact: Self-attention replaces recurrence/convolution successfully and generalizes to other tasks