Current Section

Overview

0%

← Back to AI Foundations
AI Foundations · Chapter 5

What is an LLM?

Understand how Large Language Models process language, predict tokens, generate responses, work with context, and power modern AI applications.

Beginner11–13 min readLarge Language Models

What you will learn

✓What the words Large, Language, and Model mean
✓How an LLM turns a prompt into a response
✓What tokens are and why they matter
✓The difference between training and inference
✓How context windows and external memory work
✓Why LLMs hallucinate and require safeguards

30-second explanation

A Large Language Model is a Deep Learning system trained to recognise and generate patterns in language.

When you enter a prompt, the model analyses the available context and predicts one token at a time until it produces a complete response.

Prompt → Tokens → Transformer processing → Next-token prediction → Response

Understand

What does “Large Language Model” mean?

The name describes the scale, the type of information, and the mathematical system involved.

Scale

Large

Modern LLMs are trained using enormous datasets, large numbers of parameters, and significant computing power.

  • ✓Huge text and code datasets
  • ✓Millions or billions of learned parameters
  • ✓Powerful GPU infrastructure
Information

Language

The model learns patterns found in human language, programming languages, documents, conversations, and structured text.

  • ✓Natural-language conversations
  • ✓Software code
  • ✓Articles, reports, and documentation
Mathematics

Model

An LLM is a mathematical system that has learned relationships between tokens and uses those relationships to generate predictions.

  • ✓Learns probabilities
  • ✓Predicts likely next tokens
  • ✓Produces outputs from learned patterns
In simple terms: An LLM is a large mathematical prediction system trained on language and code.

Visualize

What happens after you press Enter?

The model does not instantly retrieve a prewritten answer. It processes the prompt and generates the response step by step.

01

Prompt

You provide a question, instruction, document, code sample, or business task.

02

Tokenise

The input is divided into smaller pieces called tokens.

03

Process context

Transformer layers analyse relationships between the tokens.

04

Predict

The model calculates probabilities for possible next tokens.

05

Generate

Tokens are produced repeatedly until the response is complete.

Response architecture

Your prompt

Instructions and context

Transformer layers

Analyse token relationships

Probability calculation

Score possible next tokens

Generated response

Produce tokens repeatedly

Tokens

LLMs do not read exactly like humans

Before processing text, the system divides it into smaller units called tokens. A token may be a whole word, part of a word, a punctuation mark, or a code fragment.

Original input

Artificial Intelligence

Possible token split

Artificial Intelligence

A phrase may be split into complete words or word pieces depending on the tokenizer.

Original input

unbelievable

Possible token split

unbelievable

Long or uncommon words may be broken into smaller reusable pieces.

Original input

console.log(value);

Possible token split

console.log(value);

Programming syntax can also be divided into words, symbols, and fragments.

Cost

Many AI services calculate usage and pricing partly from input and output token counts.

Context

The context window is measured in tokens, which limits how much information the model can process.

Response length

Generating more tokens usually takes more time and consumes more model capacity.

Simple example

How next-token prediction creates a sentence

Suppose the prompt contains:

The capital of Belgium is

The model scores possible next tokens based on patterns learned during training and the current context.

Brussels

High probability

Antwerp

Lower probability

Paris

Very low probability

because

Very low probability

The model selects a token and continues.

After generating “Brussels,” it predicts the next token using the updated sentence. This process repeats rapidly until the answer is complete.

Compare

Training an LLM vs using an LLM

Training and inference are two very different stages.

Training

The model learns its parameters

Text + Code + Documents
↓
Large-scale GPU training
↓
Learned model parameters
  • • Can take weeks or months
  • • Requires large datasets
  • • Uses significant computing power
  • • Adjusts billions of internal values

Inference

The trained model generates an answer

User prompt
↓
Trained model
↓
Generated response
  • • Usually takes seconds
  • • Uses existing model parameters
  • • Processes the current context
  • • Produces one or more outputs

Generate

One language engine, many kinds of output

An LLM is not only a chatbot. The same underlying model can generate many forms of language and structured content.

Example source information

A software release was delayed by two days because final security testing found a configuration issue. The fix is complete, and the new release date is Friday.

Professional email

Draft an email confirming a project timeline and requesting approval.

Python code

Generate a function that validates and transforms incoming JSON data.

Document summary

Turn a long report into decisions, risks, and next actions.

Structured JSON

Extract customer name, issue category, urgency, and recommended action.

Context

Context window and memory

The context window is the amount of information an LLM can consider during one request.

Inside the context window

1System instructions
2Conversation messages
3Uploaded document content
4Retrieved knowledge
5Current user prompt
Context-window limit

Model context

Information included in the current request can directly influence the response.

External memory

Applications may store preferences or earlier information in a database and add it to future prompts when required.

Forgotten context

Information outside the available context is not automatically visible to the model.

Evolution

Rule-based chatbot vs LLM assistant

Rule-based chatbot

  • • Uses fixed scripts and decision trees
  • • Handles predefined questions
  • • Produces approved standard responses
  • • Struggles with unexpected wording
  • • Requires manual rule updates

LLM assistant

  • • Generates responses dynamically
  • • Handles many forms of natural language
  • • Can summarise, explain, and transform information
  • • Supports broader and more flexible tasks
  • • Requires stronger evaluation and safeguards

Avoid

Why do LLMs hallucinate?

An LLM is designed to generate likely language, not automatically verify every claim against a trusted source.

01

Question

The user asks for information.

02

Missing knowledge

The answer is absent, unclear, private, or outdated.

03

Language prediction

The model still predicts a plausible continuation.

04

Fluent response

The answer sounds natural and confident.

05

Incorrect claim

The generated content may not be factual.

Limitations

Where LLMs can fail

Hallucinations

An LLM can generate information that sounds convincing but is inaccurate or completely invented.

Knowledge limitations

The model may not know private information, recent developments, or specialised company knowledge.

Prompt sensitivity

Small changes in wording or context can sometimes produce noticeably different answers.

Limited reasoning reliability

The model may solve a problem correctly once and fail on a similar problem later.

Privacy and security

Sensitive information must be protected when prompts, documents, or code are sent to an AI system.

No automatic accountability

The model cannot take responsibility for business, legal, medical, or financial decisions.

Use

Where LLMs are used in real work

The strongest applications combine language models with trusted information, software tools, workflows, and human judgement.

Knowledge assistant

Answer employee questions using trusted company documents and internal knowledge.

Customer support copilot

Summarise cases, retrieve policies, draft responses, and recommend next actions.

Software engineering

Generate code, explain existing systems, create tests, review changes, and assist debugging.

Document processing

Extract information, classify documents, summarise content, and generate structured outputs.

Research support

Compare sources, organise findings, create summaries, and identify open questions.

Business automation

Connect language understanding to workflows, APIs, tools, databases, and operational systems.

Enterprise architecture

A production LLM system is more than a chatbot

Companies normally place an LLM inside a controlled architecture rather than allowing it to operate alone.

User

Provides a request

Application

Adds rules and context

Knowledge

Retrieves trusted data

LLM

Generates or reasons

Tools

Call APIs or systems

Review

Validate the outcome

Trusted knowledge

Connect the model to approved documents, databases, and knowledge systems.

Tools and APIs

Allow the model to retrieve data or perform approved actions through controlled integrations.

Permissions

Restrict which users and models can access specific information or capabilities.

Evaluation

Measure accuracy, relevance, safety, latency, cost, and business outcomes.

Human review

Keep people responsible for high-impact outputs and important decisions.

Monitoring

Track failures, model behaviour, usage patterns, security risks, and quality over time.

Consider an LLM when

  • ✓ The task involves language, code, or documents.
  • ✓ Multiple acceptable outputs can exist.
  • ✓ Relevant context can be provided.
  • ✓ The output can be evaluated or reviewed.
  • ✓ Flexibility creates measurable value.

Use caution when

  • • Every result must be mathematically exact.
  • • The decision carries serious legal or safety impact.
  • • No reliable source or context is available.
  • • Sensitive data cannot be properly protected.
  • • No human or automated validation is possible.

Key takeaway

An LLM is a powerful language prediction engine, not a guaranteed source of truth.

Large Language Models generate impressive responses by processing context and predicting tokens. Their real value appears when they are connected to trusted knowledge, clear instructions, controlled tools, evaluation, security, and human oversight.

Remember the basic pattern:

Convert the prompt into tokens → process relationships through transformer layers → predict the next token → repeat until the response is complete.

Continue Learning

Related Lessons & Next Steps

Explore more practical AI guides from AIMates.

Stay in the loop

Get practical AI tutorials, frameworks, and real-work insights.