AI Engineering · 60 min

Build a Resilient LLM Client

Build a typed, async LLM client in pure Python (`httpx` + `asyncio` only) with configurable timeouts, exponential backoff with jitter, response streaming, and structured logging — tested against a local mock provider that fails on purpose.

Premium

Problem

Every agent, RAG pipeline, and chat feature you'll ever build sits on top of one unglamorous piece of infrastructure: the code that actually talks to the LLM provider's HTTP API. It's tempting to treat that call as a one-liner — `response = requests.post(url, json=payload)` — and move on. In production, that one-liner is the single most common source of flaky agents: a provider having a slow minute, a network blip, a connection that resets mid-request. None of these mean your prompt was wrong. All of them will crash a naive client and, without retries, silently fail a real user's request. In this lab you will build `LLMClient`, a small async client with no framework dependency — just `httpx` for HTTP and the standard library's `asyncio` — that handles the four things every production LLM client needs: a **configurable timeout** so one slow request can't hang your whole pipeline, **retry with exponential backoff and jitter** so transient failures resolve themselves instead of surfacing to the user, **streaming** so long responses can be shown token-by-token instead of all at once, and **structured logging** of every attempt so when something *does* go wrong in production, you have the data to diagnose it. You'll test all of this against a small local mock provider (included in this lab, stdlib-only) that can simulate a slow response, a flaky one that fails twice before succeeding, and a streamed one — so you can prove your retry logic actually works without spending a single real API credit or depending on a real provider being unreliable on demand.

Objectives

By the end of this lab you will be able to:

Prerequisites

To complete this lab you'll need:

What you will build

A single file, llm_client.py, holding an LLMClient class that grows
across five steps: a plain call, a call with a timeout, a call with
retry + backoff + jitter, a streaming call, and structured logging tying
it all together. You'll test each stage against mock_provider.py, a
small local HTTP server (included, stdlib-only) that stands in for a
real LLM API and can be told to be slow, flaky, or streaming.

Why this matters beyond this lab

This is the client shape you'll reuse (or recognize) in every agent
framework you touch professionally — LangChain, the OpenAI SDK,
Anthropic's SDK all implement some version of exactly these four
concerns internally. Building it yourself once means you'll be able to
read, configure, and debug that machinery in any framework, instead of
treating it as a black box.

How to work through this lab

Work through the five steps in order — each stage adds to the same
llm_client.py file. Keep the mock provider running in its own terminal
for the whole lab; restart it (Ctrl+C then re-run) whenever a step
tells you to, since its flaky-mode failure counter is stateful and needs
a fresh start between tests.

Steps

Content exclusive to subscribers. See plans