In ten steps, build a support agent that classifies online-store inquiries, looks up orders, and files refunds when asked. You add Mistral features one by one โ structured output, function calling, reasoning mode, OCR โ then move the sensitive classification step to local Ministral 3 and finish with evaluation and serving.
The finished agent handles each inquiry in the order below. Each step adds one file, and each file reuses the ones before it.
Customer inquiry
โ
โโ classify.py classify (category, urgency, order number) โ Ministral 3 ยท structured output
โโ router.py send only hard inquiries to reasoning mode โ reasoning_effort
โโ agent.py answer using order and refund tools โ Mistral Small 4 ยท function calling
โ โโ tools.py
โโ receipt.py attached receipt image โ structured data โ Mistral OCR
โโ app.py expose it over HTTP โ FastAPI
local.py run classification on local Ministral 3 (Ollama)
evaluate.py compare classification accuracy and latency, API vs local| Step | Mistral feature | Result |
|---|---|---|
| 1-2 | Python SDK 3.0, model aliases | Shared client llm.py, latency and tokens per model |
| 3 | Structured output (chat.parse) | Inquiry classifier |
| 4 | Function calling (tools) | Order lookup and refund agent |
| 5 | Reasoning mode (reasoning_effort) | Difficulty-based router |
| 6 | OCR (ocr.process) | Receipt extraction |
| 7-8 | Local Ministral 3 | Local classifier and accuracy comparison |
| 9 | โ | FastAPI service |
The model lineup, licensing and SDK basics are in the Mistral guide. This tutorial focuses on combining those features into one service.
uv init support-agent --python 3.12
cd support-agent
uv add "mistralai>=3" pydantic ollama fastapi "uvicorn[standard]"
export MISTRAL_API_KEY="your_key" # Windows PowerShell: $env:MISTRAL_API_KEY="..."Keep the client and model settings used by every step in one file.
import os
from mistralai.client import Mistral
from mistralai.client.models import TextChunk
# Split models by task โ a small fast model for short classification and extraction, a general model for chat and tools
FAST_MODEL = "ministral-3-8b-2512"
AGENT_MODEL = "mistral-small-latest"
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
def text_of(content) -> str:
"""Return only the final answer text from a response (skips thinking chunks in reasoning mode)"""
if isinstance(content, str):
return content
return "".join(chunk.text for chunk in content or [] if isinstance(chunk, TextChunk))Since SDK 3.0, from mistralai import Mistral no longer works. When you borrow older examples from the web, change it to from mistralai.client import Mistral.
Send the same question to a small model (Ministral 3 8B) and a general model (Mistral Small 4) and compare latency and token counts. These numbers are what you later use to decide which task goes to which model.
import time
from llm import AGENT_MODEL, FAST_MODEL, client, text_of
PROMPT = "I ordered three days ago and shipping still hasn't started. What should I do?"
for model in (FAST_MODEL, AGENT_MODEL):
started = time.perf_counter()
resp = client.chat.complete(
model=model,
messages=[{"role": "user", "content": PROMPT}],
temperature=0.1,
max_tokens=300,
)
elapsed = time.perf_counter() - started
usage = resp.usage
print(f"[{model}] {elapsed:.2f}s ยท input {usage.prompt_tokens} / output {usage.completion_tokens} tokens")
print(text_of(resp.choices[0].message.content)[:200], "\n")uv run first_call.pyUse dated names such as ministral-3-8b-2512 in production so behavior does not change when a new model ships. Aliases such as mistral-small-latest move to new versions automatically, which suits experiments.
Free-text classification results are fragile to parse. Pass a Pydantic model to chat.parse and you get JSON that matches the schema, directly as a Ticket object. Spell out the category and urgency criteria in the system prompt.
from typing import Literal
from pydantic import BaseModel, Field
from llm import FAST_MODEL, client
class Ticket(BaseModel):
category: Literal["billing", "delivery", "account", "bug", "other"]
urgency: Literal["low", "medium", "high"]
order_id: str | None = Field(default=None, description="The order number if the inquiry has one, otherwise null")
summary: str = Field(description="One-sentence summary")
SYSTEM = """You classify customer inquiries for an online store.
category: billing (payments, refunds, receipts) / delivery (shipping, returns) / account (sign-in, profile) / bug (app or web errors) / other
urgency: high (financial loss, service unusable, legal action mentioned) / medium (complaints about delays) / low (simple questions)"""
def classify(text: str, model: str = FAST_MODEL) -> Ticket:
resp = client.chat.parse(
model=model,
messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": text}],
response_format=Ticket,
temperature=0,
)
return resp.choices[0].message.parsed
if __name__ == "__main__":
for text in [
"I was charged twice. If it isn't refunded today, I'm blocking my card.",
"Order A-1024 has been stuck in shipping for three days",
"I'm not getting the password reset email",
]:
print(classify(text).model_dump()){'category': 'billing', 'urgency': 'high', 'order_id': None, 'summary': '...'}
{'category': 'delivery', 'urgency': 'medium', 'order_id': 'A-1024', 'summary': '...'}
{'category': 'account', 'urgency': 'low', 'order_id': None, 'summary': '...'}Restricting values with Literal stops the model from inventing labels that are not on the list. When the taxonomy changes, update the list and the system prompt together.
So the model never invents order status, real data comes only from tools. Start by defining the tool functions and their schemas.
import json
# In a real service these call your order DB and payment API
ORDERS = {
"A-1024": {"status": "in transit", "carrier": "UPS", "eta": "2026-10-07"},
"A-2048": {"status": "paid", "carrier": None, "eta": None},
}
def get_order_status(order_id: str) -> dict:
order = ORDERS.get(order_id.upper())
return {"order_id": order_id, **order} if order else {"error": f"Order {order_id} not found"}
def request_refund(order_id: str, reason: str) -> dict:
return {"order_id": order_id, "refund_ticket": f"RF-{order_id}", "status": "received", "reason": reason}
TOOLS = {"get_order_status": get_order_status, "request_refund": request_refund}
TOOL_SCHEMAS = [
{"type": "function", "function": {
"name": "get_order_status",
"description": "Get an order's status, carrier and estimated delivery date by order number",
"parameters": {"type": "object",
"properties": {"order_id": {"type": "string", "description": "Order number (e.g. A-1024)"}},
"required": ["order_id"]},
}},
{"type": "function", "function": {
"name": "request_refund",
"description": "File a refund only when the customer clearly asks for one",
"parameters": {"type": "object",
"properties": {"order_id": {"type": "string"}, "reason": {"type": "string", "description": "Short refund reason"}},
"required": ["order_id", "reason"]},
}},
]
def run_tool(name: str, raw_args) -> str:
"""Validate and run a call the model produced. The Mistral SDK may give arguments as a dict or a JSON string"""
if name not in TOOLS:
return json.dumps({"error": f"unknown tool: {name}"})
try:
args = raw_args if isinstance(raw_args, dict) else json.loads(raw_args or "{}")
return json.dumps(TOOLS[name](**args))
except (json.JSONDecodeError, TypeError) as e:
return json.dumps({"error": f"bad arguments: {e}"})The agent repeats "call โ run โ return the result" until the model stops calling tools and answers.
from llm import AGENT_MODEL, client, text_of
from tools import TOOL_SCHEMAS, run_tool
SYSTEM = """You are a customer-support agent for an online store.
- Always check order status with the tools, and never make up anything the tool results do not say.
- File a refund only when the customer clearly asks for one.
- Answer politely in three sentences or fewer."""
def answer(question: str, reasoning: str = "none", max_steps: int = 4) -> str:
messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": question}]
for _ in range(max_steps):
resp = client.chat.complete(
model=AGENT_MODEL,
messages=messages,
tools=TOOL_SCHEMAS,
tool_choice="auto",
parallel_tool_calls=False, # one at a time โ keeps logs and error handling simple
reasoning_effort=reasoning,
temperature=0.1,
)
msg = resp.choices[0].message
messages.append(msg)
if not msg.tool_calls:
return text_of(msg.content)
for tc in msg.tool_calls:
result = run_tool(tc.function.name, tc.function.arguments)
print(f"[tool] {tc.function.name}({tc.function.arguments}) -> {result}")
messages.append({"role": "tool", "name": tc.function.name, "content": result, "tool_call_id": tc.id})
return "This is taking longer than expected. A support agent will follow up shortly."
if __name__ == "__main__":
print(answer("When will order A-1024 arrive?"))[tool] get_order_status({"order_id": "A-1024"}) -> {"order_id": "A-1024", "status": "in transit", "carrier": "UPS", "eta": "2026-10-07"}
Your order A-1024 is in transit with UPS and should arrive on October 7. ...
(sample output โ the wording varies between runs)The Mistral SDK sometimes gives tool arguments as a dict and sometimes as a JSON string. Unless you handle both, as run_tool does, json.loads raises a TypeError. For tools that are hard to reverse, such as refunds, re-check amount and status rules in code before executing.
Mistral Small 4 can turn reasoning on and off within one model. Turning it on for every inquiry is slow and expensive, so use the step 3 classification to send only financial-loss (urgency high) or bug inquiries with reasoning_effort="high".
from agent import answer
from classify import Ticket, classify
def needs_reasoning(ticket: Ticket) -> bool:
"""Reasoning mode is expensive โ use it only for hard inquiries such as financial loss or error analysis"""
return ticket.urgency == "high" or ticket.category == "bug"
def handle(text: str) -> dict:
ticket = classify(text)
effort = "high" if needs_reasoning(ticket) else "none"
reply = answer(text, reasoning=effort)
return {"ticket": ticket.model_dump(), "reasoning": effort, "reply": reply}
if __name__ == "__main__":
for text in ["When will order A-1024 arrive?", "I was charged twice. Please refund order A-2048."]:
print(handle(text), "\n"){'ticket': {'category': 'delivery', 'urgency': 'medium', ...}, 'reasoning': 'none', 'reply': '...'}
{'ticket': {'category': 'billing', 'urgency': 'high', 'order_id': 'A-2048', ...}, 'reasoning': 'high', 'reply': '...'}
(sample output)With reasoning on, the response content becomes a list of thinking and text chunks. Step 1โs text_of collects only the text chunks, so the agent reads answers the same way whether reasoning is on or off.
Refund requests often come with a receipt photo. Mistral OCR turns the image into Markdown, and the same structured output as step 3 extracts the merchant, payment time and total. Send local images as base64 data URLs.
import base64
import mimetypes
import sys
from pydantic import BaseModel, Field
from llm import FAST_MODEL, client
class Receipt(BaseModel):
merchant: str
paid_at: str | None = Field(default=None, description="Payment date and time (YYYY-MM-DD HH:MM)")
total: float = Field(description="Total amount paid")
items: list[str]
def to_data_url(path: str) -> str:
mime = mimetypes.guess_type(path)[0] or "image/jpeg"
with open(path, "rb") as f:
return f"data:{mime};base64,{base64.b64encode(f.read()).decode()}"
def read_receipt(path: str) -> Receipt:
# 1) OCR: image โ Markdown
ocr = client.ocr.process(
model="mistral-ocr-latest",
document={"type": "image_url", "image_url": to_data_url(path)},
)
markdown = "\n\n".join(page.markdown for page in ocr.pages)
# 2) Structure: Markdown โ Receipt
resp = client.chat.parse(
model=FAST_MODEL,
messages=[
{"role": "system", "content": "Extract receipt details from OCR output. Use null for missing values."},
{"role": "user", "content": markdown},
],
response_format=Receipt,
temperature=0,
)
return resp.choices[0].message.parsed
if __name__ == "__main__":
print(read_receipt(sys.argv[1] if len(sys.argv) > 1 else "receipt.jpg").model_dump())uv run receipt.py ./receipt.jpg
# {'merchant': 'AI Mart Downtown', 'paid_at': '2026-10-03 14:22', 'total': 23.8, 'items': ['Coffee beans', 'Mug']}Receipts contain personal data such as partial card numbers and names. Confirm your company policy allows sending them to an external API, and do not store extracted values you do not need.
Classification is a short input with a fixed output format, so a local model is often good enough โ and inquiry text never leaves your machine. Ministral 3 is Apache 2.0 open-weight, so Ollama runs it directly.
ollama pull ministral-3 # 8B, about 6 GB
ollama run ministral-3 "hello"from ollama import chat
from classify import SYSTEM, Ticket
def classify_local(text: str, model: str = "ministral-3") -> Ticket:
"""Classify with the same Ticket schema on local Ministral 3 โ inquiry text never leaves your machine"""
resp = chat(
model=model,
messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": text}],
format=Ticket.model_json_schema(),
options={"temperature": 0},
)
return Ticket.model_validate_json(resp.message.content)
if __name__ == "__main__":
print(classify_local("Order A-1024 has been stuck in shipping for three days").model_dump())The key is reusing the same Ticket schema and system prompt as step 3. That keeps the comparison in the next step fair, and the rest of the code stays the same whichever classifier you use.
Compare accuracy and latency of both classifiers on labeled inquiries. In practice, take at least 100 past inquiries that agents have already handled.
{"text": "I was charged twice. If it isn't refunded today, I'm blocking my card.", "category": "billing", "urgency": "high"}
{"text": "Order A-1024 has been stuck in shipping for three days", "category": "delivery", "urgency": "medium"}
{"text": "I'm not getting the password reset email", "category": "account", "urgency": "low"}
{"text": "The app freezes on a white screen when I tap the pay button", "category": "bug", "urgency": "high"}import json
import statistics
import time
from classify import classify
from local import classify_local
cases = [json.loads(line) for line in open("data/labeled.jsonl", encoding="utf-8")]
def evaluate(name: str, fn) -> None:
hits_category = hits_urgency = 0
latencies = []
for case in cases:
started = time.perf_counter()
ticket = fn(case["text"])
latencies.append(time.perf_counter() - started)
hits_category += ticket.category == case["category"]
hits_urgency += ticket.urgency == case["urgency"]
if ticket.category != case["category"]:
print(f" โ [{name}] {case['text'][:40]}โฆ expected {case['category']} / got {ticket.category}")
n = len(cases)
print(f"{name}: category {hits_category}/{n}, urgency {hits_urgency}/{n}, "
f"latency p50 {statistics.median(latencies):.2f}s / max {max(latencies):.2f}s")
evaluate("API ministral-3-8b", classify)
evaluate("local ministral-3", classify_local)โ [local ministral-3] The app freezes on a white screen when I taโฆ expected bug / got billing
API ministral-3-8b: category 4/4, urgency 4/4, latency p50 0.62s / max 0.91s
local ministral-3: category 3/4, urgency 4/4, latency p50 0.48s / max 1.30s
(sample output โ accuracy and latency depend on hardware, data and model version)Collecting the misses (โ) shows where the system promptโs criteria are ambiguous. Fix the criteria and evaluate again.
Wrap the flow in an HTTP API. Limit input length, hand off to a human when AI processing fails, and log the classification, reasoning use and latency per inquiry.
import logging
import time
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
from router import handle
logging.basicConfig(level=logging.INFO)
log = logging.getLogger("support-agent")
app = FastAPI(title="Support Agent")
class Inquiry(BaseModel):
text: str = Field(min_length=1, max_length=2000)
@app.post("/inquiries")
def create_inquiry(inquiry: Inquiry) -> dict:
started = time.perf_counter()
try:
result = handle(inquiry.text)
except Exception:
log.exception("inquiry failed")
raise HTTPException(status_code=502, detail="AI processing failed. Connecting you to a human agent.")
ticket = result["ticket"]
log.info("category=%s urgency=%s reasoning=%s elapsed=%.2fs",
ticket["category"], ticket["urgency"], result["reasoning"], time.perf_counter() - started)
return resultuv run uvicorn app:app --port 8080
curl -X POST http://localhost:8080/inquiries \
-H "Content-Type: application/json" \
-d '{"text": "I was charged twice. Please refund order A-2048."}'As concurrent requests grow, prepare for the Mistral APIโs per-minute limits (429) with retries and a queue, and measure load with k6.
The Mistral API is billed by token usage, and each model has its own input and output price. Small models such as Ministral 3 are cheap, which suits repetitive work like classification and extraction. Check the pricing page in Mistral AI Studio for current prices and free-tier limits.
Classification is a short input with a fixed output format, so a small local model is often accurate enough, and inquiry text never leaves your machine. Conversational answers that use tools benefit much more from a larger model, so the API is the better fit there. Let your evaluation results decide where each step runs.
Reasoning mode generates thinking tokens before answering, which makes responses slower and more expensive. For simple inquiries such as delivery checks the quality difference is negligible, so routing only hard inquiries โ financial loss or error analysis โ to reasoning gives the best value.
Describe when each tool may be used, and for actions that are hard to reverse, such as refunds, re-check rules like amount limits and order status in code before executing. Send high-value or high-risk requests to a human approval queue instead of executing them directly.