bigscience-workshop / Megatron-DeepSpeed

Ongoing research training transformer language models at scale, including: BERT & GPT-2
Other
1.3k stars 211 forks source link

Add xPos embeddings #370

Open janEbert opened 1 year ago

janEbert commented 1 year ago

See https://arxiv.org/abs/2212.10554.