Current Section

Overview

0%

← Back to AI Foundations
AI Foundations · Chapter 8

Tokens Explained Simply

Understand how AI models divide language into tokens and why those tokens affect context limits, cost, speed, RAG design, and application performance.

Beginner10–12 min readLLM Engineering

What you will learn

✓What an AI token is in simple language
✓How sentences, technical words, and code become tokens
✓How input and output tokens are processed
✓What information fills an LLM context window
✓How tokens affect API pricing and latency
✓How token limits influence RAG and agent design

30-second explanation

Tokens are the small units of text that an AI model processes internally.

A token may be a complete word, part of a word, punctuation, a number, or a code fragment. The model receives tokens as input and generates new tokens as output.

Human text → Tokens → LLM processing → Output tokens → Human-readable response

Understand

What exactly is a token?

Humans think in words and sentences. Language models process smaller units selected by a tokenizer.

Human view

We naturally read a sentence as complete words carrying meaning.

Hello, how are you today?

Possible model view

A tokenizer may divide the same sentence into several separate units.

Hello, how are you today?
Important: Token boundaries are not universal. Different models and tokenizers may split the same text differently.

Visualize

How different content becomes tokens

Tokenisation depends on the vocabulary learned by the tokenizer, which means ordinary text, technical terms, and code may be split differently.

Original input

Hello, how are you?

Possible tokens

Hello, how are you?

Common words and punctuation may become separate tokens. The exact split depends on the model's tokenizer.

Original input

unbelievable

Possible tokens

unbelievable

Longer or less common words may be divided into smaller reusable pieces.

Original input

Retrieval-Augmented Generation

Possible tokens

Retrieval-Augmented Generation

Technical phrases may be divided into full words, punctuation, and partial words.

Original input

console.log(userId);

Possible tokens

console.log(userId);

Code is also tokenised into identifiers, symbols, keywords, and fragments.

Process

How an LLM processes tokens

Tokens provide the bridge between human-readable language and the numerical calculations performed by a neural network.

01

Receive text

The application receives a prompt, document, message, or code sample.

02

Tokenise

A tokenizer divides the content into smaller units called tokens.

03

Convert

Each token is represented internally using a numerical identifier.

04

Process

The model analyses relationships between the tokens in the context.

05

Generate

The model predicts and produces output tokens one at a time.

Internal architecture

Prompt

Human-readable text

Tokenizer

Split text into units

Token IDs

Numerical representations

Neural network

Process and predict

Usage

Input tokens vs output tokens

A model request normally consumes tokens on both sides of the interaction.

Input tokens

Information sent to the model

  • • System instructions
  • • User prompt
  • • Conversation history
  • • Retrieved documents
  • • Tool results and structured data

Output tokens

Information generated by the model

  • • Text responses
  • • Code
  • • JSON or structured output
  • • Summaries and explanations
  • • Tool-call instructions

Simple usage example

500

Input tokens

300

Output tokens

800

Total processed tokens

Providers may price input and output tokens differently, so the actual cost depends on the selected model and service.

Context window

What fills an LLM context window?

The context window represents the total amount of tokenised information the model can consider during a request.

01

System instructions

Rules describing the model's role, behaviour, restrictions, and output expectations.

02

Conversation history

Earlier user and assistant messages included to maintain continuity.

03

Retrieved knowledge

Relevant document chunks, database results, or search information added by the application.

04

Current prompt

The user's latest question, instruction, or task.

05

Generated response

The output tokens being created also consume available context capacity.

Context-window boundary

When the combined content exceeds the available limit, the application must shorten, remove, summarise, or retrieve information more selectively.

Inside the current context

The model can use the information

Instructions, messages, and documents included in the current request can directly influence the generated response.

Outside the current context

The model cannot automatically see it

Older messages or external information must be stored and added again by the application when they are needed.

Cost and performance

Why tokens affect cost and speed

Token usage is not only a technical limit. It directly influences the economics and responsiveness of an AI application.

Smaller request

Relevant instructions
+ Selected context
+ Concise output
  • ✓ Lower token consumption
  • ✓ Often lower cost
  • ✓ Often faster processing
  • ✓ Less irrelevant context

Oversized request

Repeated instructions
+ Entire documents
+ Long unnecessary output
  • • Higher token consumption
  • • Greater cost
  • • Longer processing time
  • • More potential distraction

RAG architecture

Why RAG systems do not send every document

A large document collection may contain millions of tokens. A RAG system searches the collection and sends only the most relevant content to the model.

Documents

Large knowledge base

Chunking

Split into sections

Embeddings

Represent meaning

Retrieval

Find relevant chunks

Context

Add selected text

LLM

Generate answer

Without retrieval

The application may send too much irrelevant content.

With retrieval

Only the most useful document chunks enter the context.

Result

Lower token usage and more focused answers.

Use

Why tokens matter in real AI systems

API pricing

Many model providers calculate usage separately for input and output tokens.

Context management

Token limits determine how much conversation, documentation, and retrieved knowledge can be processed.

RAG design

Document chunks must fit within the available context together with prompts and generated answers.

Application latency

Larger prompts and longer outputs generally require more processing time.

Agent workflows

Every planning step, tool result, observation, and response may add additional tokens.

Scalability

Small inefficiencies become expensive when an application handles thousands or millions of requests.

Optimise

Practical token optimisation

Good token optimisation removes waste without removing information the model genuinely needs.

Remove repeated instructions

Avoid sending the same background information multiple times within one request.

Retrieve only relevant chunks

Use search, metadata filters, and reranking instead of sending an entire knowledge base.

Summarise old history

Compress older conversation messages while retaining important decisions and facts.

Set output limits

Ask for a response length appropriate to the task rather than generating unnecessary detail.

Use structured context

Clear headings, fields, and formats help reduce repetition and improve model interpretation.

Choose the right model

Not every request requires the largest or most expensive available model.

Do not optimise blindly: the cheapest prompt is not useful if it removes the context required for a reliable answer.

Avoid

Common token mistakes

Sending an entire document

Large documents consume context, increase cost, and may distract the model with irrelevant information.

Ignoring output tokens

A short prompt can still become expensive when the application generates very long answers.

Repeating conversation history

Sending unnecessary old messages consumes context without improving the current response.

Using oversized RAG chunks

Very large chunks can contain mixed topics and leave less room for other useful context.

Assuming words equal tokens

Token counts vary by language, punctuation, formatting, code, and the tokenizer used.

Optimising only for cost

Reducing too much context can damage answer quality, so cost and usefulness must be balanced.

AI for Real Work

Good AI engineers optimise the complete request

Token management is not simply about shortening prompts. It means selecting the right instructions, retrieving relevant knowledge, controlling output length, and balancing quality, cost, latency, and scalability.

Less waste

Remove duplicate and irrelevant information.

Lower cost

Reduce unnecessary input and output processing.

Better speed

Keep requests focused and responses appropriate.

Greater scale

Make every request more efficient before traffic grows.

Key takeaway

Words are for humans. Tokens are the units processed by language models.

Tokens influence how much information an LLM can consider, how long its answers can be, how much an API request costs, how quickly it responds, and how RAG and agent systems must be designed.

Remember the basic pattern:

Text becomes tokens → tokens fill the context window → the model processes them → output tokens become the final response.

Continue Learning

Related Lessons & Next Steps

Explore more practical AI guides from AIMates.

Stay in the loop

Get practical AI tutorials, frameworks, and real-work insights.