Tokens Explained Simply
Understand how AI models divide language into tokens and why those tokens affect context limits, cost, speed, RAG design, and application performance.
What you will learn
30-second explanation
Tokens are the small units of text that an AI model processes internally.
A token may be a complete word, part of a word, punctuation, a number, or a code fragment. The model receives tokens as input and generates new tokens as output.
Understand
What exactly is a token?
Humans think in words and sentences. Language models process smaller units selected by a tokenizer.
Human view
We naturally read a sentence as complete words carrying meaning.
Possible model view
A tokenizer may divide the same sentence into several separate units.
Visualize
How different content becomes tokens
Tokenisation depends on the vocabulary learned by the tokenizer, which means ordinary text, technical terms, and code may be split differently.
Original input
Hello, how are you?
Possible tokens
Common words and punctuation may become separate tokens. The exact split depends on the model's tokenizer.
Original input
unbelievable
Possible tokens
Longer or less common words may be divided into smaller reusable pieces.
Original input
Retrieval-Augmented Generation
Possible tokens
Technical phrases may be divided into full words, punctuation, and partial words.
Original input
console.log(userId);
Possible tokens
Code is also tokenised into identifiers, symbols, keywords, and fragments.
Process
How an LLM processes tokens
Tokens provide the bridge between human-readable language and the numerical calculations performed by a neural network.
Receive text
The application receives a prompt, document, message, or code sample.
Tokenise
A tokenizer divides the content into smaller units called tokens.
Convert
Each token is represented internally using a numerical identifier.
Process
The model analyses relationships between the tokens in the context.
Generate
The model predicts and produces output tokens one at a time.
Internal architecture
Prompt
Human-readable text
Tokenizer
Split text into units
Token IDs
Numerical representations
Neural network
Process and predict
Usage
Input tokens vs output tokens
A model request normally consumes tokens on both sides of the interaction.
Input tokens
Information sent to the model
- • System instructions
- • User prompt
- • Conversation history
- • Retrieved documents
- • Tool results and structured data
Output tokens
Information generated by the model
- • Text responses
- • Code
- • JSON or structured output
- • Summaries and explanations
- • Tool-call instructions
Simple usage example
500
Input tokens
300
Output tokens
800
Total processed tokens
Providers may price input and output tokens differently, so the actual cost depends on the selected model and service.
Context window
What fills an LLM context window?
The context window represents the total amount of tokenised information the model can consider during a request.
System instructions
Rules describing the model's role, behaviour, restrictions, and output expectations.
Conversation history
Earlier user and assistant messages included to maintain continuity.
Retrieved knowledge
Relevant document chunks, database results, or search information added by the application.
Current prompt
The user's latest question, instruction, or task.
Generated response
The output tokens being created also consume available context capacity.
Context-window boundary
When the combined content exceeds the available limit, the application must shorten, remove, summarise, or retrieve information more selectively.
Inside the current context
The model can use the information
Instructions, messages, and documents included in the current request can directly influence the generated response.
Outside the current context
The model cannot automatically see it
Older messages or external information must be stored and added again by the application when they are needed.
Cost and performance
Why tokens affect cost and speed
Token usage is not only a technical limit. It directly influences the economics and responsiveness of an AI application.
Smaller request
+ Selected context
+ Concise output
- ✓ Lower token consumption
- ✓ Often lower cost
- ✓ Often faster processing
- ✓ Less irrelevant context
Oversized request
+ Entire documents
+ Long unnecessary output
- • Higher token consumption
- • Greater cost
- • Longer processing time
- • More potential distraction
RAG architecture
Why RAG systems do not send every document
A large document collection may contain millions of tokens. A RAG system searches the collection and sends only the most relevant content to the model.
Documents
Large knowledge base
Chunking
Split into sections
Embeddings
Represent meaning
Retrieval
Find relevant chunks
Context
Add selected text
LLM
Generate answer
Without retrieval
The application may send too much irrelevant content.
With retrieval
Only the most useful document chunks enter the context.
Result
Lower token usage and more focused answers.
Use
Why tokens matter in real AI systems
API pricing
Many model providers calculate usage separately for input and output tokens.
Context management
Token limits determine how much conversation, documentation, and retrieved knowledge can be processed.
RAG design
Document chunks must fit within the available context together with prompts and generated answers.
Application latency
Larger prompts and longer outputs generally require more processing time.
Agent workflows
Every planning step, tool result, observation, and response may add additional tokens.
Scalability
Small inefficiencies become expensive when an application handles thousands or millions of requests.
Optimise
Practical token optimisation
Good token optimisation removes waste without removing information the model genuinely needs.
Remove repeated instructions
Avoid sending the same background information multiple times within one request.
Retrieve only relevant chunks
Use search, metadata filters, and reranking instead of sending an entire knowledge base.
Summarise old history
Compress older conversation messages while retaining important decisions and facts.
Set output limits
Ask for a response length appropriate to the task rather than generating unnecessary detail.
Use structured context
Clear headings, fields, and formats help reduce repetition and improve model interpretation.
Choose the right model
Not every request requires the largest or most expensive available model.
Avoid
Common token mistakes
Sending an entire document
Large documents consume context, increase cost, and may distract the model with irrelevant information.
Ignoring output tokens
A short prompt can still become expensive when the application generates very long answers.
Repeating conversation history
Sending unnecessary old messages consumes context without improving the current response.
Using oversized RAG chunks
Very large chunks can contain mixed topics and leave less room for other useful context.
Assuming words equal tokens
Token counts vary by language, punctuation, formatting, code, and the tokenizer used.
Optimising only for cost
Reducing too much context can damage answer quality, so cost and usefulness must be balanced.
AI for Real Work
Good AI engineers optimise the complete request
Token management is not simply about shortening prompts. It means selecting the right instructions, retrieving relevant knowledge, controlling output length, and balancing quality, cost, latency, and scalability.
Less waste
Remove duplicate and irrelevant information.
Lower cost
Reduce unnecessary input and output processing.
Better speed
Keep requests focused and responses appropriate.
Greater scale
Make every request more efficient before traffic grows.
Key takeaway
Words are for humans. Tokens are the units processed by language models.
Tokens influence how much information an LLM can consider, how long its answers can be, how much an API request costs, how quickly it responds, and how RAG and agent systems must be designed.
Remember the basic pattern:
Text becomes tokens → tokens fill the context window → the model processes them → output tokens become the final response.
Continue Learning
Related Lessons & Next Steps
Explore more practical AI guides from AIMates.
Stay in the loop
Get practical AI tutorials, frameworks, and real-work insights.