During an incident you jump between the runbook and logs from several services to find the cause. In ten steps, put everything into DeepSeekβs 1M-token context, cut the cost of follow-up questions on the same material to about 1/50 with automatic caching, and build an agent that searches the logs itself.
Bundle one incidentβs runbook and logs into a single document and fix it at the start of every request. Because the prefix is identical, from the second question on that part hits the cache and input cost drops sharply. Question answering, timeline extraction and a log-search agent sit on top.
incident/ runbook + per-service logs
β
ββ bundle.py bundle in a fixed order as the system prompt β the key to cache hits
ββ ask_many.py many questions on the same material Β· cache and cost
ββ effort_compare.py thinking off / low / high
ββ timeline.py incident timeline as JSON β JSON output
ββ agent.py investigate with log-search and metric tools β thinking mode + tools
ββ redact.py mask personal data and secrets before sending
ββ local_triage.py first pass with a local R1 (when data cannot leave)
ββ offpeak.py detect off-peak hours
ββ evaluate.py evaluate against ground truth Β· total cost
ββ app.py FastAPI serviceThe model lineup, pricing and thinking-mode rules are in the DeepSeek guide. The old deepseek-chat and deepseek-reasoner were retired in July 2026, so this tutorial uses deepseek-flash.
uv init incident-analyst --python 3.12
cd incident-analyst
uv add openai ollama fastapi "uvicorn[standard]"
export DEEPSEEK_API_KEY="your_key" # platform.deepseek.comKeep the client, model name, price table and cost calculation in one file. Thinking mode is on by default, so every request states it with thinking().
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["DEEPSEEK_API_KEY"], base_url="https://api.deepseek.com")
MODEL = "deepseek-flash"
# USD per 1M tokens at peak hours (official prices as of October 2026 β check before relying on them)
PRICE = {"cache_hit": 0.006, "cache_miss": 0.30, "output": 1.20}
def thinking(enabled: bool) -> dict:
"""Thinking mode is on by default, so always state it explicitly"""
return {"thinking": {"type": "enabled" if enabled else "disabled"}}
def cost_usd(usage, off_peak: bool = False) -> float:
hit = getattr(usage, "prompt_cache_hit_tokens", 0) or 0
miss = getattr(usage, "prompt_cache_miss_tokens", usage.prompt_tokens - hit) or 0
total = (hit * PRICE["cache_hit"] + miss * PRICE["cache_miss"] + usage.completion_tokens * PRICE["output"]) / 1_000_000
return total / 2 if off_peak else totalPut the runbook and logs in incident/ and read them in a fixed order into one document, which becomes the system prompt of every request. It must be the exact same string every time to hit the cache, so read in sorted order and leave out changing values such as timestamps and request IDs.
incident/
βββ runbook.md # checkout incident runbook
βββ deploy.log # deploy history
βββ checkout.log # application log
βββ db.log # DB connection pool logfrom pathlib import Path
INCIDENT_DIR = Path("incident")
def load_bundle() -> str:
"""Combine the runbook and logs into one fixed document.
File order and content must be identical every time for the prefix to hit the cache β read in sorted order
and leave out changing values such as timestamps or request IDs."""
parts = []
for path in sorted(INCIDENT_DIR.glob("*")):
if path.suffix in {".md", ".log", ".txt"}:
parts.append(f"===== {path.name} =====\n{path.read_text(encoding='utf-8')}")
return "\n\n".join(parts)
SYSTEM_PREFIX = """You are an SRE helping with incident analysis. Answer only from the runbook and logs below,
and quote the log lines you rely on with their file name and timestamp. Say so when something is a guess not found in the logs.
"""
def system_message() -> dict:
return {"role": "system", "content": SYSTEM_PREFIX + load_bundle()}
if __name__ == "__main__":
text = load_bundle()
print(f"{len(text):,} characters β roughly {len(text) // 4:,} tokens (rough estimate; check usage in a response for the real value)")Fix the material at the front and ask three different questions. Check cache hits with prompt_cache_hit_tokens and cost with step 1βs cost_usd.
import time
from bundle import system_message
from ds import MODEL, client, cost_usd, thinking
QUESTIONS = [
"When did the incident start, and what is the first error log line?",
"Which service and endpoint have the most errors?",
"According to the runbook, what should we do now, in order?",
]
system = system_message() # identical prefix for every question β cache hits from the second one on
total = 0.0
for q in QUESTIONS:
started = time.perf_counter()
resp = client.chat.completions.create(
model=MODEL,
messages=[system, {"role": "user", "content": q}],
extra_body=thinking(False),
)
u = resp.usage
cost = cost_usd(u)
total += cost
print(f"Q: {q}\n {time.perf_counter() - started:.1f}s Β· cache hit {getattr(u, 'prompt_cache_hit_tokens', 0):,}"
f" / miss {getattr(u, 'prompt_cache_miss_tokens', 0):,} Β· ${cost:.5f}")
print(" " + resp.choices[0].message.content[:150].replace("\n", " "), "\n")
print(f"total ${total:.5f}")Q: When did the incident start, and what is the first error log line?
2.3s Β· cache hit 0 / miss 182,410 Β· $0.05479
checkout.log 09:13:10 β the first 500 error (NullPointerException) marks the start of the incident.
Q: Which service and endpoint have the most errors?
1.9s Β· cache hit 182,400 / miss 18 Β· $0.00117
...
(sample output β with 180K tokens of material, input cost drops to about 1/50 from the second question)Caches usually persist for hours to days. During an incident, questions keep coming on the same material, so only the first question is expensive and the rest cost little more than their output tokens.
On a hard question such as root-cause estimation, compare thinking off, low and high for latency, output tokens and cost. Output tokens drive cost the most.
import time
from bundle import system_message
from ds import MODEL, client, cost_usd, thinking
QUESTION = "From the logs alone, estimate the root cause and give the supporting evidence and an alternative hypothesis."
system = system_message()
settings = [
("thinking off", {"extra_body": thinking(False)}),
("thinking low", {"extra_body": thinking(True), "reasoning_effort": "low"}),
("thinking high", {"extra_body": thinking(True), "reasoning_effort": "high"}),
]
for name, params in settings:
started = time.perf_counter()
resp = client.chat.completions.create(model=MODEL, messages=[system, {"role": "user", "content": QUESTION}], **params)
msg = resp.choices[0].message
reasoning = getattr(msg, "reasoning_content", None) or ""
print(f"[{name}] {time.perf_counter() - started:.1f}s Β· output {resp.usage.completion_tokens:,} tokens "
f"(reasoning {len(reasoning):,} chars) Β· ${cost_usd(resp.usage):.5f}")
print(" " + msg.content[:160].replace("\n", " "), "\n")[thinking off] 3.1s Β· output 412 tokens (reasoning 0 chars) Β· $0.00172
[thinking low] 9.8s Β· output 1,935 tokens (reasoning 3,240 chars) Β· $0.00355
[thinking high] 24.6s Β· output 6,880 tokens (reasoning 13,900 chars) Β· $0.00948
(sample output β values vary widely with the question and material)Thinking mode ignores temperature. When comparing answer quality, change only thinking on/off and reasoning_effort.
A timeline has to be structured to go into a report or dashboard. JSON output requires the word "json" and a format example in the prompt, and empty responses can occasionally occur, so retry.
import json
from bundle import system_message
from ds import MODEL, client, thinking
INSTRUCTION = """Output the incident timeline from the logs above as json only. Example format:
{"events": [{"time": "09:12:04", "service": "checkout", "event": "deploy v2.4.1 completed", "source": "deploy.log"}],
"first_error_at": "09:13:10", "suspected_cause": "one sentence"}"""
def extract_timeline(retries: int = 2) -> dict:
for _ in range(retries + 1):
resp = client.chat.completions.create(
model=MODEL,
messages=[system_message(), {"role": "user", "content": INSTRUCTION}],
response_format={"type": "json_object"},
max_tokens=2000,
extra_body=thinking(False),
)
content = resp.choices[0].message.content
if content: # empty responses happen occasionally β retry
return json.loads(content)
raise RuntimeError("Could not extract the timeline")
if __name__ == "__main__":
timeline = extract_timeline()
for e in timeline["events"]:
print(f"{e['time']} {e['service']:<10} {e['event']} ({e['source']})")
print("first error:", timeline["first_error_at"], "| suspected cause:", timeline["suspected_cause"])09:12:58 checkout deploy v2.4.1 completed (deploy.log)
09:13:10 checkout POST /api/orders 500 errors begin (checkout.log)
first error: 09:13:10 | suspected cause: an OrderService defect in deploy v2.4.1
(sample output)Let tools fetch older logs that did not fit in the bundle and live metrics. When using tools in thinking mode, you must send back the reasoning_content of every previous turn, so store the whole response message with model_dump().
import json
import re
from bundle import INCIDENT_DIR, system_message
from ds import MODEL, client, cost_usd, thinking
def grep_logs(pattern: str, max_lines: int = 20) -> dict:
"""Find log lines matching a regex (can search older logs that are not in the bundle)"""
hits = []
for path in sorted(INCIDENT_DIR.glob("*.log")):
for line in path.read_text(encoding="utf-8").splitlines():
if re.search(pattern, line):
hits.append(f"{path.name}: {line}")
return {"pattern": pattern, "count": len(hits), "lines": hits[:max_lines]}
def get_metric(service: str, metric: str) -> dict:
"""Service metrics (query the Prometheus HTTP API in a real implementation)"""
sample = {("checkout", "error_rate"): 0.083, ("checkout", "p95_ms"): 1840, ("payment", "error_rate"): 0.002}
return {"service": service, "metric": metric, "value": sample.get((service, metric))}
TOOLS = {"grep_logs": grep_logs, "get_metric": get_metric}
TOOL_SCHEMAS = [
{"type": "function", "function": {
"name": "grep_logs", "description": "Search log files for lines matching a regex",
"parameters": {"type": "object", "properties": {
"pattern": {"type": "string", "description": "Python regex"},
"max_lines": {"type": "integer"}}, "required": ["pattern"]}}},
{"type": "function", "function": {
"name": "get_metric", "description": "Get a service's current metric (error_rate, p95_ms)",
"parameters": {"type": "object", "properties": {
"service": {"type": "string"},
"metric": {"type": "string", "enum": ["error_rate", "p95_ms"]}}, "required": ["service", "metric"]}}},
]
def investigate(question: str, max_steps: int = 6) -> str:
messages = [system_message(), {"role": "user", "content": question}]
spent = 0.0
for _ in range(max_steps):
resp = client.chat.completions.create(
model=MODEL, messages=messages, tools=TOOL_SCHEMAS,
reasoning_effort="high", extra_body=thinking(True),
)
spent += cost_usd(resp.usage)
msg = resp.choices[0].message
# In a thinking-mode conversation with tools, reasoning_content must be sent back on every turn
messages.append(msg.model_dump(exclude_none=True))
if not msg.tool_calls:
print(f"(cost ${spent:.5f})")
return msg.content
for tc in msg.tool_calls:
fn = TOOLS.get(tc.function.name)
try:
result = fn(**json.loads(tc.function.arguments)) if fn else {"error": "unknown tool"}
except (json.JSONDecodeError, TypeError, re.error) as e:
result = {"error": str(e)}
print(f"[tool] {tc.function.name}({tc.function.arguments}) -> count={result.get('count', result.get('value'))}")
messages.append({"role": "tool", "tool_call_id": tc.id, "content": json.dumps(result)})
return "Investigation step limit reached."
if __name__ == "__main__":
print(investigate("Check with logs and metrics whether the checkout errors come from the deploy or the DB"))[tool] grep_logs({"pattern": "ERROR|5\\d\\d"}) -> count=2
[tool] get_metric({"service": "checkout", "metric": "error_rate"}) -> count=0.083
(cost $0.01240)
Errors at OrderService.java:118 began right after deploy v2.4.1 while the DB pool stayed at 44%, so the deploy is the likely cause.
(sample output)Storing only some fields, as in {"role": "assistant", "content": msg.content}, drops reasoning_contentand the next request fails with a 400 error. The code in this tutorial was verified against a mock server that enforces this rule strictly.
Logs mix in emails, IPs and tokens. Mask them before sending to an external API, and for data that must not leave at all, do only a first-pass triage with a local R1 distill.
import re
import sys
from pathlib import Path
# Mask personal data and secrets in logs before sending them to an external API
PATTERNS = [
(re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"), "<EMAIL>"),
(re.compile(r"\b(?:\d{1,3}\.){3}\d{1,3}\b"), "<IP>"),
(re.compile(r"(?i)(authorization|api[_-]?key|token|password)([\"'=:\s]+)[^\s\"',]+"), r"\1\2<SECRET>"),
(re.compile(r"(?<!\w)(?:\+?\d{1,3}[-. ]?)?\(?\d{3}\)?[-. ]?\d{3}[-. ]?\d{4}\b"), "<PHONE>"),
]
def redact(text: str) -> str:
for pattern, repl in PATTERNS:
text = pattern.sub(repl, text)
return text
if __name__ == "__main__":
src, dst = Path(sys.argv[1]), Path(sys.argv[2])
dst.mkdir(exist_ok=True)
for path in src.glob("*"):
if path.is_file():
(dst / path.name).write_text(redact(path.read_text(encoding="utf-8")), encoding="utf-8")
print("redacted", path.name)uv run redact.py incident incident_redacted
# token=abc123secret β token=<SECRET>
# started by ops@example.com β started by <EMAIL>
# completed on 10.0.3.21 β completed on <IP>ollama pull deepseek-r1:14b # 16 GB GPU classfrom ollama import chat
from bundle import load_bundle
# When raw logs cannot leave: do a first-pass triage with a local R1 distill and send only the summary to the API
resp = chat(
model="deepseek-r1:14b",
messages=[{
"role": "user",
"content": "From these logs, pick only the incident-related lines and summarize them in time order in 10 lines or fewer. Leave out personal data.\n\n" + load_bundle()[-60_000:],
}],
think=True,
options={"num_ctx": 32768},
)
print(resp.message.content)Regex masking only covers known formats. Add patterns for in-house identifiers such as employee or account numbers, and have a person spot-check the masked output.
DeepSeek charges half outside peak hours (weekdays 01:00-04:00 and 06:00-10:00 UTC). Incident response cannot wait, but non-urgent work such as postmortem reports or summarizing past incidents in bulk can run off-peak.
from datetime import datetime, time, timezone
# Peak hours (weekdays, UTC): 01:00-04:00 and 06:00-10:00 β prices are half outside them
PEAK_WINDOWS = [(time(1), time(4)), (time(6), time(10))]
def is_off_peak(now: datetime | None = None) -> bool:
now = (now or datetime.now(timezone.utc)).astimezone(timezone.utc)
if now.weekday() >= 5: # Saturday and Sunday
return True
return not any(start <= now.time() < end for start, end in PEAK_WINDOWS)
if __name__ == "__main__":
now = datetime.now(timezone.utc)
print(f"Now (UTC {now:%a %H:%M}) is {'off-peak' if is_off_peak(now) else 'peak'}.")
# Run non-urgent batch jobs, such as postmortem reports, only off-peakCheck answers against ground truth confirmed by the people who handled the incident. Start with key terms each answer must contain, and move to human scoring as cases grow.
{"question": "When did the incident start?", "must_include": ["09:13"]}
{"question": "Which version was deployed just before?", "must_include": ["v2.4.1"]}
{"question": "What is the first action in the runbook?", "must_include": ["roll back"]}import json
from bundle import system_message
from ds import MODEL, client, cost_usd, thinking
# Ground truth confirmed by the people who handled the incident β key terms the answer must contain
cases = [json.loads(line) for line in open("data/qa.jsonl", encoding="utf-8")]
system = system_message()
total_cost, passed = 0.0, 0
for case in cases:
resp = client.chat.completions.create(
model=MODEL, messages=[system, {"role": "user", "content": case["question"]}], extra_body=thinking(False),
)
total_cost += cost_usd(resp.usage)
answer = resp.choices[0].message.content
missing = [k for k in case["must_include"] if k.lower() not in answer.lower()]
passed += not missing
print(("β" if not missing else "β"), case["question"], f"(missing: {', '.join(missing)})" if missing else "")
print(f"passed {passed}/{len(cases)} Β· total cost ${total_cost:.5f}")Finally, wrap it as an API for the team. Log cache hits and cost per request to keep an eye on how thinking mode affects spend.
import logging
from fastapi import FastAPI
from pydantic import BaseModel, Field
from bundle import system_message
from ds import MODEL, client, cost_usd, thinking
from offpeak import is_off_peak
logging.basicConfig(level=logging.INFO)
log = logging.getLogger("incident-analyst")
app = FastAPI(title="Incident Analyst")
class Question(BaseModel):
text: str = Field(min_length=1, max_length=2000)
deep: bool = False # True turns on thinking mode β slower, but suited to root-cause analysis
@app.post("/ask")
def ask(q: Question) -> dict:
resp = client.chat.completions.create(
model=MODEL,
messages=[system_message(), {"role": "user", "content": q.text}],
extra_body=thinking(q.deep),
)
u = resp.usage
cost = cost_usd(u, off_peak=is_off_peak())
log.info("deep=%s hit=%s miss=%s out=%s cost=$%.5f", q.deep,
getattr(u, "prompt_cache_hit_tokens", 0), getattr(u, "prompt_cache_miss_tokens", 0), u.completion_tokens, cost)
return {"answer": resp.choices[0].message.content, "cost_usd": round(cost, 6)}uv run uvicorn app:app --port 8080
# INFO:incident-analyst:deep=False hit=182400 miss=18 out=96 cost=$0.00122prompt_cache_hit_tokens)reasoning_content sent back on every turn?PRICE) kept in sync with the official pricing page?For material that fits within a few hundred thousand tokens β one incidentβs runbook and related logs β putting it in whole avoids retrieval misses and keeps things simple. DeepSeekβs automatic caching means asking many questions about the same material does not multiply input cost. When material exceeds millions of tokens or keeps growing, RAG to select only the relevant parts is better.
The cache hits only when the start of a request exactly matches an earlier one. Putting the current time or a request ID in the system prompt, or reading log files in a different order each time, prevents hits. Move changing values into the user message and read files in sorted order.
Logs often contain personal data and secrets such as emails, IPs and tokens. Mask them before sending, and for data your policy forbids sending, run a first pass with a local R1 distill and send only the summary. DeepSeek weights are MIT-licensed, so self-hosting is also an option under strict data policies.
Questions answered by reading the logs β first error time, deployed version β work fine, fast and cheap with thinking off. Turn it on only for root-cause analysis that compares hypotheses or investigations that use tools several times. Output tokens multiply, so compare cost directly as in step 4.