Current Section

Overview

0%

← Back to Tutorials
AIMates Tutorials · AI Agents

Build Your First AI Agent with Python

Build a small but genuine AI agent that can interpret a goal, select a Python tool, execute it safely, inspect the result, and return a final answer.

Beginner FriendlyPython18–22 min readTool Calling

What you will build

✓A basic Python assistant using the Responses API
✓A reliable calculator tool with validated arguments
✓A model-driven tool selection workflow
✓A complete tool execution and response loop
✓Error handling and maximum-iteration protection
✓A foundation you can extend with APIs, RAG, and memory

30-second explanation

An AI agent combines a language model with tools and an execution loop.

The model decides which approved capability may help. Your Python application executes that capability, validates the result, and sends the result back to the model so it can continue or finish.

User goal → Model decision → Tool call → Python execution → Tool result → Final response

Understand

Chatbot vs AI agent

A chatbot generates language. An agent can combine language with controlled actions.

Basic assistant

Produces an answer

User prompt
↓
Language model
↓
Text response
  • • Explains or generates content
  • • Uses the supplied context
  • • Usually returns one response
  • • Cannot automatically execute your Python code

Tool-using agent

Works toward an outcome

User goal
↓
Model selects tool
↓
Python executes tool
↓
Model explains result
  • ✓ Selects from approved tools
  • ✓ Uses real tool results
  • ✓ Can perform multiple controlled steps
  • ✓ Stops when the outcome is complete

Mental model

Observe → Decide → Act → Inspect

Agent behaviour is best understood as a loop rather than a single prompt.

01

Observe

Read the user’s goal, available context, previous tool results, and current task state.

02

Decide

Determine whether to answer directly, ask for clarification, or use an approved tool.

03

Act

Call a Python function, API, database, search service, file system, or another application.

04

Inspect

Review the tool result and decide whether the task is complete or another step is required.

05

Respond

Return a useful final answer containing the result, assumptions, and any necessary warnings.

Architecture

Components of the agent we will build

OutcomeGoal+

The concrete result the agent should achieve. Clear goals reduce unnecessary planning and tool calls.

Example

Calculate the estimated project cost and explain the result to the user.

BehaviourInstructions+

Rules describing the agent’s role, boundaries, response style, and when it should ask for help.

Example

Never invent missing values. Ask the user when required information is unavailable.

Decision engineModel+

The language model interprets the task, selects an appropriate tool, and explains the final result.

Example

Recognise that cost estimation requires the calculator tool rather than mental arithmetic.

CapabilitiesTools+

Python functions or external integrations that provide reliable information and controlled actions.

Example

Calculate a value, retrieve an order, search documentation, or create an email draft.

ContextMemory+

Application-managed information about the current workflow, earlier messages, or approved user preferences.

Example

Remember the hourly rate already supplied during the current task.

ControlValidation+

Checks that tool arguments, tool results, and final answers are safe, complete, and appropriate.

Example

Reject negative working hours and require approval before sending an email.

Build

Implementation roadmap

01

Install

Create a Python environment and install the official OpenAI package.

02

Configure

Store the API key safely in an environment variable.

03

Define tools

Write small Python functions with clear inputs and predictable outputs.

04

Describe tools

Provide JSON schemas so the model knows when and how each function may be called.

05

Run the loop

Execute requested functions, return results to the model, and continue until completion.

06

Add controls

Validate input, limit iterations, handle errors, and log important actions.

Step 1

Set up the Python project

Create a new project directory and virtual environment.

mkdir first-ai-agent
cd first-ai-agent

python -m venv .venv
source .venv/bin/activate

pip install --upgrade openai

Store the API key in your terminal environment rather than placing it directly in source code.

export OPENAI_API_KEY="your-api-key"

# Confirm that the variable exists without printing the secret.
python -c "import os; print(bool(os.getenv('OPENAI_API_KEY')))"
Never commit API keys. Keep secrets in environment variables or a managed secret store, and ensure local environment files are ignored by Git.

Step 2

Start with a basic AI assistant

This first example has instructions and a user goal, but it does not yet have tools or an execution loop.

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-5-mini",
    instructions=(
        "You are a practical AI learning assistant. "
        "Give structured, beginner-friendly answers. "
        "Do not invent missing information."
    ),
    input="Create a seven-day learning plan for Python AI development.",
)

print(response.output_text)

This is useful, but it is not yet a tool-using agent.

The model can generate a plan, but it cannot execute a Python function, retrieve live information, or update another system.

Step 3

Create a reliable Python tool

A tool is a normal function controlled by your application. Keep tool responsibilities narrow and validate all arguments.

def calculate_project_cost(
    hours: float,
    hourly_rate: float,
) -> dict[str, float]:
    """Calculate the estimated cost of a project."""

    if hours <= 0:
        raise ValueError("hours must be greater than zero")

    if hourly_rate <= 0:
        raise ValueError("hourly_rate must be greater than zero")

    total = round(hours * hourly_rate, 2)

    return {
        "hours": hours,
        "hourly_rate": hourly_rate,
        "total": total,
    }


result = calculate_project_cost(hours=10, hourly_rate=75)

print(result)
# {'hours': 10, 'hourly_rate': 75, 'total': 750}

Deterministic

The same valid inputs produce the same calculated result.

Validated

Invalid values are rejected before they create a misleading output.

Structured

The result can be serialised and returned to the language model.

Step 4

Describe the tool to the model

The tool schema tells the model what the function does and which arguments it accepts. It does not execute the function.

TOOLS = [
    {
        "type": "function",
        "name": "calculate_project_cost",
        "description": (
            "Calculate the estimated project cost from the number "
            "of working hours and the hourly rate."
        ),
        "parameters": {
            "type": "object",
            "properties": {
                "hours": {
                    "type": "number",
                    "description": "Estimated working hours.",
                },
                "hourly_rate": {
                    "type": "number",
                    "description": "Cost per working hour.",
                },
            },
            "required": ["hours", "hourly_rate"],
            "additionalProperties": False,
        },
        "strict": True,
    }
]
The model proposes a call. Your Python application still decides whether the request is valid and whether the tool may actually run.

Complete implementation

Build the tool execution loop

Save the following as agent.py.

import json
from typing import Any

from openai import OpenAI


MODEL = "gpt-5-mini"
MAX_TOOL_ROUNDS = 5

client = OpenAI()


def calculate_project_cost(
    hours: float,
    hourly_rate: float,
) -> dict[str, float]:
    """Calculate the estimated cost of a project."""

    if hours <= 0:
        raise ValueError("hours must be greater than zero")

    if hourly_rate <= 0:
        raise ValueError("hourly_rate must be greater than zero")

    return {
        "hours": hours,
        "hourly_rate": hourly_rate,
        "total": round(hours * hourly_rate, 2),
    }


TOOLS = [
    {
        "type": "function",
        "name": "calculate_project_cost",
        "description": (
            "Calculate the estimated project cost from working hours "
            "and an hourly rate."
        ),
        "parameters": {
            "type": "object",
            "properties": {
                "hours": {
                    "type": "number",
                    "description": "Estimated number of working hours.",
                },
                "hourly_rate": {
                    "type": "number",
                    "description": "Cost per working hour.",
                },
            },
            "required": ["hours", "hourly_rate"],
            "additionalProperties": False,
        },
        "strict": True,
    }
]


def execute_tool(
    name: str,
    arguments: dict[str, Any],
) -> dict[str, Any]:
    """Execute one approved tool."""

    if name == "calculate_project_cost":
        return calculate_project_cost(
            hours=float(arguments["hours"]),
            hourly_rate=float(arguments["hourly_rate"]),
        )

    raise ValueError(f"Unknown tool: {name}")


def run_agent(user_goal: str) -> str:
    """Run a small agent until it returns a final answer."""

    response = client.responses.create(
        model=MODEL,
        instructions=(
            "You are a project-estimation assistant. "
            "Use the calculator tool for cost calculations. "
            "Never estimate arithmetic mentally when the tool can be used. "
            "Ask for missing information instead of inventing it. "
            "Explain the final result clearly and concisely."
        ),
        input=user_goal,
        tools=TOOLS,
    )

    for _ in range(MAX_TOOL_ROUNDS):
        tool_calls = [
            item
            for item in response.output
            if item.type == "function_call"
        ]

        if not tool_calls:
            return response.output_text

        tool_outputs: list[dict[str, str]] = []

        for tool_call in tool_calls:
            try:
                arguments = json.loads(tool_call.arguments)

                result = execute_tool(
                    name=tool_call.name,
                    arguments=arguments,
                )

                output = json.dumps(
                    {
                        "ok": True,
                        "result": result,
                    }
                )

            except (KeyError, TypeError, ValueError, json.JSONDecodeError) as exc:
                output = json.dumps(
                    {
                        "ok": False,
                        "error": str(exc),
                    }
                )

            tool_outputs.append(
                {
                    "type": "function_call_output",
                    "call_id": tool_call.call_id,
                    "output": output,
                }
            )

        response = client.responses.create(
            model=MODEL,
            previous_response_id=response.id,
            input=tool_outputs,
            tools=TOOLS,
        )

    raise RuntimeError(
        "Agent stopped because the maximum number of tool rounds was reached."
    )


if __name__ == "__main__":
    goal = (
        "A project needs 42 hours of work at an hourly rate "
        "of 85 euros. Calculate the estimated cost and explain it."
    )

    try:
        answer = run_agent(goal)
        print(answer)
    except Exception as exc:
        print(f"Agent failed: {exc}")

Follow the execution

What happens when the program runs?

Goal

Calculate project cost

Model

Select calculator

Arguments

42 hours and €85

Python

Validate and execute

Result

Return €3,570

Answer

Explain to user

Expected calculation

42 hours × €85 = €3,570

The exact wording of the explanation may vary, but the numeric result comes from the validated Python function.

Test behaviour

Test missing information

A reliable agent should ask for missing information rather than inventing it.

answer = run_agent(
    "A project will take 30 hours. What will the total cost be?"
)

print(answer)

Wrong behaviour

The agent invents an hourly rate and produces an unsupported estimate.

Desired behaviour

The agent explains that the hourly rate is missing and asks the user to provide it.

Extend

Tools you can add next

Calculator

Perform exact calculations instead of relying on probabilistic language generation.

Search

Retrieve current or approved information from external or internal search services.

Documents

Read files, extract content, summarise sections, or retrieve knowledge through RAG.

Database

Read structured business information using controlled, parameterised queries.

Business API

Retrieve orders, create tickets, inspect inventory, or update approved systems.

Email

Prepare a draft or send a message only after appropriate validation and approval.

Add one tool at a time. Test tool selection, valid arguments, invalid arguments, errors, and stopping behaviour before adding another capability.

Memory

Memory usually belongs in your application

A model does not automatically provide durable application memory. Store useful state in your own data layer and include only relevant information in future requests.

Current-task state

Tool results, intermediate calculations, decisions, and remaining steps.

Conversation state

Relevant earlier messages needed to understand the current request.

Long-term preferences

Approved user preferences stored in a database and retrieved only when useful.

Production architecture

A real agent needs controlled application layers

Do not connect an unrestricted model directly to sensitive tools. Authentication, policy checks, validation, and monitoring belong around the agent.

User

Provides goal

API

Authenticates

Agent

Selects action

Policy

Checks permission

Tool

Executes safely

Validation

Checks result

Audit

Records outcome

Input validation

Validate every argument before passing it to a tool or external application.

Least privilege

Give the agent only the tools and permissions required for its specific responsibility.

Iteration limits

Stop runaway loops by setting maximum tool calls, time, and spending limits.

Human approval

Require confirmation before sending messages, spending money, deleting data, or creating external impact.

Observability

Record requests, tool calls, errors, latency, token usage, and final outcomes.

Evaluation

Test representative success cases, edge cases, missing inputs, and unsafe requests.

Build next

Practical agent projects

Research agent

Search approved sources, organise findings, compare evidence, and prepare a referenced summary.

Support agent

Retrieve customer and product context, suggest a resolution, and prepare a reply for review.

Data-analysis agent

Load data, run calculations, create summaries, and explain significant trends.

Software agent

Inspect code, propose changes, run tests, analyse failures, and prepare a patch for review.

Meeting agent

Convert notes into decisions, actions, owners, deadlines, and follow-up communication.

Cloud operations agent

Inspect alerts and logs, retrieve runbooks, recommend remediation, and escalate when required.

Avoid

Common beginner mistakes

Calling a chatbot an agent

A single LLM response is useful, but an agent normally includes tools, state, and an execution loop.

Starting with too much autonomy

Build a narrow, controlled workflow before attempting open-ended autonomous behaviour.

Giving too many tools

Every additional capability increases security risk, cost, ambiguity, and debugging complexity.

No argument validation

The model proposes tool arguments, but your application remains responsible for validating them.

No stopping condition

An uncontrolled loop can repeat actions, consume tokens, and continue without useful progress.

No human checkpoint

High-impact actions should not occur merely because a model generated a tool call.

Agent checklist

Before adding another tool

✓Does the tool solve one clear problem?
✓Are all arguments validated?
✓Can the requesting user access this capability?
✓Are errors returned in a safe structured form?
✓Is there a maximum number of execution rounds?
✓Can repeated calls create duplicate side effects?
✓Does a high-impact action require approval?
✓Are tool calls and outcomes logged?

AI for Real Work

The model proposes. Your application controls.

A production agent is not an unrestricted model with access to everything. It is a controlled software system in which the model selects from approved capabilities and the application validates, executes, monitors, and limits every action.

Understand

Interpret the goal

Choose

Select a tool

Validate

Check permission and input

Execute

Run application code

Verify

Inspect the outcome

Key takeaway

An AI agent is a model inside a controlled execution loop.

The model interprets the goal and proposes actions. Python functions and external systems perform the real work. Validation, permissions, iteration limits, monitoring, and human approval make that work safe and reliable.

Remember the complete pattern:

Observe the goal → select an approved tool → validate the arguments → execute in Python → inspect the result → finish or continue.

Continue learning

Add better instructions and trusted knowledge