Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment
In plain terms
This paper reviews Transformer-based language models, which are the powerful AI systems behind tools like ChatGPT. It first categorizes different types of these models, explaining how they are built and what makes them unique. The authors then discuss recent breakthroughs since 2023, such as methods for training models to follow instructions better (instruction tuning) and using human feedback to improve their responses (reinforcement learning from human feedback). A key part of the paper surveys how these AI models are used in various fields like healthcare, finance, and importantly, education, linking specific model capabilities to their real-world uses. Finally, the paper critically assesses these models, comparing their architectures, energy costs, and how we measure their 'state-of-the-art' performance, concluding with open research questions.