Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention
Research finds relative positional encodings enable transformer models to generalize better to longer sequences than absolute encodings.