Current Section

Overview

0%

← Back to AI Foundations
AI Foundations · Chapter 6

What is an AI Agent?

Understand how AI agents combine language models, tools, memory, planning, and controlled workflows to complete multi-step tasks.

Beginner12–14 min readAgentic AI

What you will learn

✓What makes an AI agent different from a chatbot
✓How the Observe → Think → Act loop works
✓The main components inside an agent architecture
✓How agents use tools, APIs, memory, and RAG
✓The difference between single-agent and multi-agent systems
✓Why permissions, validation, and human approval matter

30-second explanation

An AI agent is a software system that can interpret a goal, decide what to do next, use tools, inspect the result, and continue until the task is complete.

A language model provides flexible reasoning and communication, while tools, memory, workflows, permissions, and validation allow the system to work with real information and applications.

Goal → Observe → Think → Plan → Act → Verify → Finish or repeat

Understand

Chatbot vs AI agent

A chatbot mainly generates a response. An agent can combine conversation with planning, tool use, and task execution.

Chatbot

Primarily answers

User question
↓
Language model
↓
Generated response
  • • Usually produces one response
  • • May not access external systems
  • • Relies on provided context
  • • Often stops after answering

AI agent

Works toward an outcome

User goal
↓
Plan + Tools + Memory
↓
Completed task
  • • Can complete multiple steps
  • • Can use tools and APIs
  • • Can inspect intermediate results
  • • Can continue until a stopping condition is met
Important: Many real applications combine both. The user sees a chat interface while an agentic workflow operates behind it.

Mental model

Observe → Think → Act

The most useful way to understand an agent is as a repeating loop. It observes the current state, decides what to do, performs an action, and then observes again.

Step 1

Observe

Read the user’s goal, available information, previous tool outputs, and current environment.

Step 2

Think

Decide whether the task is complete, which information is missing, and which action is most useful next.

Step 3

Act

Call a tool, retrieve information, update a system, ask for clarification, or produce the final result.

The loop

Observe

Read new information

Think

Select next action

Act

Use a tool or respond

Observe again

Inspect the outcome

Visualize

How an AI agent completes a task

A production agent typically moves through several controlled steps rather than taking one unrestricted action.

01

Observe

Read the user’s goal, available context, previous results, and current system state.

02

Think

Interpret the task, identify missing information, and decide what should happen next.

03

Plan

Break the goal into smaller steps and select the tools required to complete them.

04

Act

Call an API, search documents, query a database, run code, or update a system.

05

Verify

Inspect the result and decide whether to finish, correct the approach, or continue.

Stopping matters: The workflow needs a clear condition for success, failure, escalation, or maximum allowed attempts.

Architecture

Anatomy of an AI agent

An agent is not only an LLM. It is a complete software architecture made from several coordinated parts.

DirectionGoal+

The task or outcome the agent is expected to achieve. A clear goal helps constrain planning and tool use.

Example

Prepare a monthly sales report and draft an email for the regional manager.

Reasoning engineLLM+

The language model interprets instructions, evaluates context, plans steps, and decides which action may be useful.

Example

Decide that sales data must be retrieved before the report can be written.

RulesInstructions+

System instructions define the agent’s responsibilities, boundaries, preferred behaviour, and prohibited actions.

Example

Use approved company data only and never send an email without human confirmation.

CapabilitiesTools+

Tools allow the agent to interact with information and software beyond the language model itself.

Example

Query Salesforce, run Python analysis, create a chart, and prepare an email draft.

ContextMemory+

Memory helps the agent retain information about the current task, previous actions, user preferences, or earlier interactions.

Example

Remember the reporting format preferred by the manager.

ControlWorkflow+

The workflow defines how steps are ordered, repeated, validated, approved, or stopped.

Example

Retrieve data, validate totals, analyse trends, draft report, request approval, then distribute.

Capabilities

Tools an AI agent can use

Tools allow the agent to move beyond text generation and interact with real data, software, and business systems.

Search

Find current information from approved websites, internal indexes, or enterprise search.

Files

Read, summarise, compare, or extract information from documents and spreadsheets.

Databases

Query structured business data such as customers, orders, inventory, or transactions.

APIs

Connect to external services and internal applications using controlled interfaces.

Calculators

Perform reliable calculations instead of asking the LLM to estimate mathematical results.

Code execution

Run Python or other approved code for analysis, transformation, testing, or automation.

Email

Draft, classify, summarise, or—with appropriate approval—send messages.

Calendar

Check availability, propose meeting times, or create events under controlled permissions.

Cloud systems

Inspect logs, trigger workflows, manage approved resources, or retrieve operational information.

More tools are not always better. Give an agent only the capabilities required for its role.

Memory

What does agent memory mean?

Memory is normally implemented using application state, databases, retrieved documents, or conversation history outside the LLM itself.

Working memory

Information needed while completing the current task, including intermediate results and decisions.

Example

The sales totals already calculated for each region.

Conversation memory

Recent messages and instructions that help maintain continuity during an interaction.

Example

The user asked for a concise report written for senior leadership.

Long-term memory

Stored preferences or facts that may be reused in future tasks when appropriate.

Example

The user normally wants reports delivered as PDF and email drafts.

Knowledge memory

External documents or retrieved information made available through RAG or search systems.

Example

Sales policies, product documentation, or operational procedures.

Real-work example

Prepare the monthly sales report

A user asks:

“Prepare the monthly sales report, identify the main trend, and draft an email for my manager.”

1. Retrieve data

Query the approved sales system for the requested period.

2. Validate

Check missing values, totals, dates, and regional coverage.

3. Analyse

Calculate growth, compare regions, and identify unusual changes.

4. Generate

Prepare a report, key findings, and a professional email draft.

5. Verify

Compare the written conclusions with the calculated results.

6. Request approval

Show the draft to the user before any external communication.

7. Deliver

Create the final file and save or send it after confirmation.

8. Record

Store the execution outcome and relevant audit information.

The language model does not perform every step alone.

The application gives the model controlled access to data tools, calculations, document generation, and approval workflows.

Collaboration

Single-agent vs multi-agent systems

Some systems use one agent with several tools. Others divide work across multiple specialised agents.

Single agent

One agent manages the complete task

User
↓
General agent
↓
Search + Data + Code + Email

Simpler to build, monitor, and debug. Often the best starting point for an MVP.

Multi-agent system

Specialised agents divide the work

Manager agent
↓
Research + Analysis + Writing + Review

Can separate responsibilities, but adds orchestration, communication, cost, and debugging complexity.

Example multi-agent architecture

Manager agent

Assigns tasks and combines results

Research agent

Collect information

Analysis agent

Calculate and compare

Writer agent

Prepare the report

Reviewer agent

Check final quality

Enterprise architecture

A production agent needs controlled layers

Enterprise agents should not directly connect an unrestricted LLM to sensitive business systems.

User

Defines goal

Application

Authenticates user

Agent

Plans actions

Memory

Stores context

Knowledge

Retrieves facts

Tools

Perform actions

Approval

Controls impact

✓Allow only approved tools and actions.
✓Use least-privilege access for every integration.
✓Validate tool inputs and outputs.
✓Set iteration, time, and spending limits.
✓Require human approval for high-impact actions.
✓Record actions for monitoring and audit.

Use

Where AI agents are used

Research agent

Search approved sources, compare evidence, organise findings, and produce a referenced summary.

Customer-support agent

Read customer history, retrieve policies, suggest actions, and draft a response for review.

Software-engineering agent

Inspect code, generate changes, run tests, analyse failures, and prepare a proposed fix.

Data-analysis agent

Retrieve data, clean it, run calculations, create charts, and explain significant trends.

Operations agent

Monitor workflows, classify incidents, retrieve runbooks, and recommend or execute approved actions.

Meeting assistant

Prepare context, capture decisions, generate actions, assign owners, and draft follow-up communication.

Consider an agent when

  • ✓ The task requires several dependent steps.
  • ✓ The next action depends on previous results.
  • ✓ Tools or external systems must be used.
  • ✓ Some flexibility is genuinely valuable.
  • ✓ Progress and outcomes can be validated.

Use normal software when

  • • A fixed workflow already solves the problem.
  • • Only one predictable API call is needed.
  • • The task requires exact deterministic behaviour.
  • • The system cannot safely tolerate uncertainty.
  • • Agent cost and complexity exceed the benefit.

Avoid

Common AI agent mistakes

The biggest agent failures often come from system design rather than the language model itself.

Using an agent for a simple task

A fixed rule, search query, or normal API call may be faster, cheaper, and more reliable.

Giving access to too many tools

Every additional tool increases complexity, security exposure, and the number of ways the agent can fail.

Unclear instructions

A vague goal allows the agent to make inconsistent assumptions and choose inappropriate actions.

No stopping condition

The agent may repeat actions, consume unnecessary tokens, or continue without meaningful progress.

No validation

Tool outputs and generated conclusions must be checked before they affect important systems.

Excessive autonomy

High-impact operations should require limits, permissions, approvals, and clear accountability.

AI for Real Work

An LLM answers. An agent coordinates work.

The important shift is not from chat to unlimited autonomy. It is from isolated text generation to controlled systems that can understand goals, retrieve knowledge, use tools, verify progress, and involve people at the right moments.

Understand

Interpret the user’s goal and context.

Plan

Choose a safe sequence of actions.

Execute

Use approved tools and information.

Validate

Check results before creating impact.

Key takeaway

An AI agent combines reasoning with controlled action.

The LLM helps interpret goals and decide what to do, but a useful agent also requires tools, memory, workflows, permissions, validation, monitoring, and clear stopping conditions.

Remember the basic pattern:

Observe the current state → think about the next step → act using an approved tool → verify the result → finish or repeat.

Continue Learning

Related Lessons & Next Steps

Explore more practical AI guides from AIMates.

Stay in the loop

Get practical AI tutorials, frameworks, and real-work insights.