Build a Tool-Calling Agent from Scratch
Build a tool-calling agent using nothing but the official OpenAI SDK — two real tools with JSON Schema, a bounded execution loop, and structured error handling for malformed model arguments. No agent framework.
PremiumProblem
Every agent framework — LangChain, LlamaIndex, CrewAI — implements the same core loop underneath: send the model a conversation plus a list of tools it's allowed to call, inspect the response for a request to call one of them, run that tool locally, feed the result back as part of the conversation, and repeat until the model is satisfied. Frameworks add convenience and structure on top of this, but the loop itself is something you can — and should — build yourself at least once with nothing but the provider's raw SDK, so you understand exactly what a framework is doing for you (and what it's hiding from you) later. In this lab you'll build that loop from scratch using only the official `openai` Python SDK: two real tools with proper JSON Schema parameter definitions, a dispatcher that executes the tool the model asked for and formats the result back correctly, a loop that repeats this until the model stops asking for tools (bounded so a confused model can't loop forever), and error handling for the case every agent eventually hits in production — the model sending arguments that don't parse or don't match what a tool expects. One of your two tools, `calculate`, evaluates arbitrary-looking arithmetic expressions. The tempting shortcut is Python's `eval()` — don't. `eval()` on model-generated text is effectively remote code execution as a feature: a model (or a prompt-injected document it read) could produce an expression like `__import__('os').system(...)`. You'll instead parse the expression into an AST with the `ast` module and only evaluate a small allow-listed set of arithmetic node types — the same "parse first, allow-list the grammar" principle that underlies every safe expression evaluator in production systems.
Objectives
By the end of this lab you will be able to:
- Define tools for an LLM using JSON Schema parameter definitions, the
same shape used by every major provider's function/tool-calling API. - Send a model a conversation with tools available and correctly parse
tool_callsout of its response. - Execute a requested tool locally and return its result to the model
as arole="tool"message, correctly linked viatool_call_id. - Implement a bounded agent loop that keeps calling the model until it
stops requesting tools, without risking an infinite loop. - Handle malformed or invalid tool arguments from the model by
returning a structured error the model can react to, instead of
crashing. - Explain why evaluating a math expression with
astparsing plus an
operator allow-list is safe whereeval()is not.
Prerequisites
To complete this lab you'll need:
- Python 3.10+ installed.
-
pip install openai(the official OpenAI Python SDK). - An OpenAI API key with a small amount of available credit, set as the
OPENAI_API_KEYenvironment variable (a model likegpt-4o-miniis
inexpensive enough for this lab's handful of calls). - Basic familiarity with Python functions, dictionaries, and JSON.
The shape of the loop
Tool-calling with a chat-completions-style API always follows the same
pattern:
1. Send: conversation so far + list of available tools (JSON Schema)
2. Model replies with EITHER:
a) a normal text answer → done, return it
b) one or more tool_calls → go to 3
3. For each tool_call: run the named function locally with its args
4. Append a role="tool" message per call, linked by tool_call_id
5. Go back to 1 with the updated conversation
Everything in this lab is building and hardening that five-step loop.
The two tools you'll implement are deliberately simple —
get_time(timezone) and calculate(expression) — so you can focus
entirely on the mechanics of the loop rather than the tools' own
logic.
Why ast, not eval(), for calculate
eval() executes arbitrary Python. If the string passed to it ever
comes from a model — and by definition, in a tool-calling agent, it
always does — you have handed remote code execution to anything that
can influence the model's output, including a malicious or
prompt-injected document the model read earlier in the conversation.
Parsing the expression into an abstract syntax tree with ast.parse()
and then walking that tree yourself, evaluating only a small
allow-listed set of node types (numbers, +, -, *, /, %, **,
unary minus), makes the "grammar" of what's acceptable explicit and
closed — anything not on the list raises an error instead of running.
How to work through this lab
Work through the five steps in order. Steps 1–2 build understanding
without executing anything; Steps 3–5 build the real loop. Test
throughout with real calls to the model — this lab is not mocked, since
understanding what the model's raw response actually looks like is
itself part of the learning goal.