School of AI

Learn

What are large language models (LLMs)?

What are large language models (LLMs)?

Short answer

A large language model is an AI model trained on very large amounts of text to predict the next piece of text, called a token. That simple objective, at huge scale, lets it write, summarise, translate, reason through problems and follow instructions. Most modern LLMs use the transformer architecture. ChatGPT, Gemini and Claude are products built on LLMs.

By Chandan Maheshwari, AI mentor & consultant, founder of School of AI · Last updated

Key takeaways

  • LLMs predict the next token based on the context they are given.
  • The context window limits how much they can consider at once.
  • Transformers are the architecture behind most LLMs.
  • Good context produces better answers than clever wording.

Key terms

The vocabulary you will meet most often.

  • Token — a chunk of text, roughly part of a word
  • Context window — how much text the model can consider at once
  • Transformer — the neural network architecture behind most LLMs
  • Embedding — a numeric representation of meaning, used for search
  • Hallucination — a fluent but false output

Learning this with School of AI

GenAI Foundation Program

Live, mentor-led and online, open to learners anywhere in India.

Further reading

Related questions

No. They learn patterns from training data up to a cutoff date and may be wrong or out of date unless given current information.

Your AI journey can start with one conversation.

Tell us what you want to learn, build or solve. We'll help you find the right starting point.

We usually reply within a few working hours, Monday to Saturday.

Talk on WhatsApp