Build a Resilient LLM Client
Build a typed, async LLM client in pure Python (`httpx` + `asyncio` only) with configurable timeouts, exponential backoff with jitter, response streaming, and structured logging — tested against a local mock provider that fails on purpose.
PremiumProblem
Every agent, RAG pipeline, and chat feature you'll ever build sits on top of one unglamorous piece of infrastructure: the code that actually talks to the LLM provider's HTTP API. It's tempting to treat that call as a one-liner — `response = requests.post(url, json=payload)` — and move on. In production, that one-liner is the single most common source of flaky agents: a provider having a slow minute, a network blip, a connection that resets mid-request. None of these mean your prompt was wrong. All of them will crash a naive client and, without retries, silently fail a real user's request. In this lab you will build `LLMClient`, a small async client with no framework dependency — just `httpx` for HTTP and the standard library's `asyncio` — that handles the four things every production LLM client needs: a **configurable timeout** so one slow request can't hang your whole pipeline, **retry with exponential backoff and jitter** so transient failures resolve themselves instead of surfacing to the user, **streaming** so long responses can be shown token-by-token instead of all at once, and **structured logging** of every attempt so when something *does* go wrong in production, you have the data to diagnose it. You'll test all of this against a small local mock provider (included in this lab, stdlib-only) that can simulate a slow response, a flaky one that fails twice before succeeding, and a streamed one — so you can prove your retry logic actually works without spending a single real API credit or depending on a real provider being unreliable on demand.
Objectives
By the end of this lab you will be able to:
- Write an async LLM client using
httpx.AsyncClientwith a clean,
typedchat(messages) -> strinterface. - Enforce a configurable request timeout with
asyncio.wait_forand
distinguish a timeout from a connection failure. - Implement retry with exponential backoff and jitter for transient
errors, and explain why jitter matters when multiple clients retry
at once. - Implement response streaming as an async generator over a chunked
HTTP response. - Add structured (JSON) logging of every attempt — latency,
success/failure, attempt number — so failures are diagnosable after
the fact. - Build and use a local mock HTTP provider to test retry and timeout
logic deterministically, without a real API key.
Prerequisites
To complete this lab you'll need:
- Python 3.10+ installed.
-
pip install httpx(no other third-party packages required — the
mock provider uses only the standard library). - A terminal capable of running two Python processes at once (the mock
provider in one, your client in another). - Basic familiarity with Python's
async/awaitsyntax. If you've
never written async Python before, you can still complete this lab —
every async construct used is explained inline.
What you will build
A single file, llm_client.py, holding an LLMClient class that grows
across five steps: a plain call, a call with a timeout, a call with
retry + backoff + jitter, a streaming call, and structured logging tying
it all together. You'll test each stage against mock_provider.py, a
small local HTTP server (included, stdlib-only) that stands in for a
real LLM API and can be told to be slow, flaky, or streaming.
Why this matters beyond this lab
This is the client shape you'll reuse (or recognize) in every agent
framework you touch professionally — LangChain, the OpenAI SDK,
Anthropic's SDK all implement some version of exactly these four
concerns internally. Building it yourself once means you'll be able to
read, configure, and debug that machinery in any framework, instead of
treating it as a black box.
How to work through this lab
Work through the five steps in order — each stage adds to the same
llm_client.py file. Keep the mock provider running in its own terminal
for the whole lab; restart it (Ctrl+C then re-run) whenever a step
tells you to, since its flaky-mode failure counter is stateful and needs
a fresh start between tests.