Awesome AI Papers
Section: Historical Papers
Entry
Appears in 5 awesome lists
bidirectional transformer pretraining; foundation for most encoder-based NLP work since 2018. Read online with section navigation and the ACL source attached.
Section: Historical Papers
Section: Pretraining and Adaptation · bidirectional transformer pretraining; foundation for most encoder-based NLP work since 2018. Read online with section navigation and the ACL source attached.
Section: Recent Language Models · , Jacob Devlin, et al., NAACL 2019, 2018.
Section: Papers · by Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova.
Section: The technical principle of ChatGPT · This paper ushered in the era of pre-training in NLP. BERT came out of nowhere.
Announcement of ChatGPT, a conversational model trained to answer follow-up questions, admit mistakes, challenge incorrect premises, and reject inappropriate requests. OpenAI blog, November 30, 2022.
(Alibaba, 2024-2025) - strong multilingual coverage, especially Chinese; often top open model on multilingual benchmarks.
and FLAN-T5 - text-to-text framing for NLP tasks; strong instruction-tuned encoder-decoder baselines.
Nature, 2021. [All Versions]. This paper provides the first computational method that can regularly predict protein structures with atomic accuracy even in cases in which no similar structure is known. This approach is a canonical application of observation- and explanation- based method for…
(Meta, 2024-2025) - widely adopted open-weight family; default base for fine-tuning across NLP tasks.
robustly optimized BERT pretraining; common encoder baseline.
(AI2, 2025) - fully open: weights, training data, code; reproducibility benchmark.
encoder vs decoder vs encoder-decoder for NLP transfer.