Short answer
A large language model is an AI model trained on very large amounts of text to predict the next piece of text, called a token. That simple objective, at huge scale, lets it write, summarise, translate, reason through problems and follow instructions. Most modern LLMs use the transformer architecture. ChatGPT, Gemini and Claude are products built on LLMs.
By Chandan Maheshwari, AI mentor & consultant, founder of School of AI · Last updated
Key takeaways
- LLMs predict the next token based on the context they are given.
- The context window limits how much they can consider at once.
- Transformers are the architecture behind most LLMs.
- Good context produces better answers than clever wording.
Key terms
The vocabulary you will meet most often.
- Token — a chunk of text, roughly part of a word
- Context window — how much text the model can consider at once
- Transformer — the neural network architecture behind most LLMs
- Embedding — a numeric representation of meaning, used for search
- Hallucination — a fluent but false output
Learning this with School of AI
GenAI Foundation Program
Live, mentor-led and online, open to learners anywhere in India.