Skip to main content

Streaming

Get responses in real-time, token by token.

Basic streaming​

from pure_agents import Agent

agent = Agent()

async for event in agent.stream("Explain quantum computing"):
if event.type == "text":
print(event.content, end="", flush=True)

Output appears as it's generated, not all at once.

Events​

stream() yields StreamEvent objects. Each has a type:

typeWhenFields you'll use
textA chunk of the answer arrivedcontent
tool_callThe model asked for a toolname, arguments, id
tool_resultThat tool finishedname, content, id
doneThe run finishedcontent (the full answer)

With tools​

Tool calls are visible as they happen:

@tool
def get_weather(city: str) -> str:
"""Get the weather for a city."""
return f"Sunny, 22°C in {city}"


agent = Agent(tools=[get_weather])

async for event in agent.stream("What's the weather in Madrid?"):
if event.type == "text":
print(event.content, end="", flush=True)
elif event.type == "tool_call":
print(f"\n[{event.name}({event.arguments})]")
elif event.type == "tool_result":
print(f"[-> {event.content}]")

Output:

[get_weather({'city': 'Madrid'})]
[-> Sunny, 22°C in Madrid]
The weather in Madrid is sunny with a temperature of 22°C.

Collecting the full response​

The done event carries the complete answer:

async for event in agent.stream("Hello"):
if event.type == "done":
full_response = event.content

Images​

async for event in agent.stream("What's in this photo?", images=["photo.jpg"]):
...

Reliability​

Streaming shares run()'s reliability features: retries, timeout, provider fallback, and token accounting. Retries only apply before the first chunk reaches you, since a stream already in flight cannot be restarted cleanly.

agent = Agent(retries=3, timeout=120.0, fallback="anthropic")

async for event in agent.stream("Write a long essay"):
...

print(agent.usage.total_tokens)

Provider support​

Streaming works with all providers:

# Mistral
agent = Agent()

# OpenAI
agent = Agent(provider="openai")

# Anthropic
agent = Agent(provider="anthropic")