In ten steps, build an assistant that answers questions like "How many vacation days do I get?" or "What do I do with my laptop when I leave?" from company policy documents, with citations. Everything runs locally so documents never leave, unsupported answers are filtered out, and the retrieval threshold comes from evaluation data.
RAG (retrieval-augmented generation) first finds document chunks related to the question, then answers using only those chunks as sources. In this tutorial both embeddings and answers come from local Qwen models, so company documents never leave your machine.
docs/*.md company policy documents
โ
โโ chunk.py split by headings, re-split long sections by length
โโ store.py index with Qwen3 Embedding โ SQLite, re-index changed files only
โโ search.py embed the query (with instruction) โ top-k by cosine similarity
โโ answer.py Qwen3.5 answers from chunks above the threshold, citing [n]
โโ verify.py rule check of citations + model check of support
โโ evaluate.py retrieval and answer accuracy, choose MIN_SCORE
โโ app.py FastAPI โ answer, sources, needs-review flagModel choice, the embedding instruction rule and thinking mode are covered in the Qwen guide.
ollama pull qwen3-embedding:0.6b # embeddings (multilingual, up to 1024 dimensions)
ollama pull qwen3.5:9b # answers (about 6.6 GB) โ qwen3.8:27b on a 24 GB GPU
uv init docs-assistant --python 3.12
cd docs-assistant
uv add ollama pydantic fastapi "uvicorn[standard]"import os
EMBED_MODEL = "qwen3-embedding:0.6b"
CHAT_MODEL = os.environ.get("QWEN_CHAT_MODEL", "qwen3.5:9b") # qwen3.8:27b on a 24 GB GPU
DB_PATH = "rag.db"
# Search results scoring below this are not used as sources. Score distributions differ by embedding model and documents,
# so set it in step 7 between the lowest top score of answerable questions and the highest of unanswerable ones.
MIN_SCORE = float(os.environ.get("RAG_MIN_SCORE", "0.35"))
DOCS_DIR = "docs"
# Instruction added to search queries only (the format recommended on the Qwen3 Embedding model card)
QUERY_TASK = "Given a web search query, retrieve relevant passages that answer the query"Qwen3 Embedding expects an instruction in the form Instruct: task description\nQuery:question on search queries only, not on documents. Adding it to documents or leaving it off queries lowers retrieval accuracy.
Retrieval is accurate when each chunk covers one topic. Split by Markdown headings first, re-split only long sections by length, and overlap slightly so boundary sentences are not lost. Store the file name and heading with each chunk so answers can show where a source came from.
docs/
โโโ leave-policy.md
โโโ equipment.md
โโโ security.mdimport re
from dataclasses import dataclass
from pathlib import Path
@dataclass
class Chunk:
source: str # file name
heading: str # nearest heading โ shown as the source location in answers
text: str
def split_markdown(path: Path, max_chars: int = 500, overlap: int = 80) -> list[Chunk]:
"""Split by headings (#) first, then re-split only long sections by character count.
Character counts behave predictably across languages, unlike word counts."""
chunks: list[Chunk] = []
heading, buf = path.stem, []
def flush():
body = "\n".join(buf).strip()
start = 0
while body and start < len(body):
chunks.append(Chunk(path.name, heading, body[start:start + max_chars]))
start += max_chars - overlap
buf.clear()
for line in path.read_text(encoding="utf-8").splitlines():
m = re.match(r"^(#{1,3})\s+(.*)", line)
if m:
flush()
heading = m.group(2).strip()
else:
buf.append(line)
flush()
return chunks
if __name__ == "__main__":
for c in split_markdown(Path("docs/leave-policy.md")):
print(f"[{c.source} > {c.heading}] {c.text[:60]}")[leave-policy.md > Annual leave] Employees in their first year earn one day of leave per month, ...
[leave-policy.md > Requesting leave] Request leave in the HR portal at least three days in advance ...Embed the chunks and store them in SQLite. Record a hash per file, so rerunning re-embeds only changed files and cleans up deleted ones. Even with hundreds of documents you never need to re-index everything daily.
import hashlib
import json
import math
import sqlite3
from pathlib import Path
from ollama import embed
from chunk import split_markdown
from config import DB_PATH, DOCS_DIR, EMBED_MODEL
db = sqlite3.connect(DB_PATH, check_same_thread=False)
db.executescript("""
CREATE TABLE IF NOT EXISTS files (name TEXT PRIMARY KEY, sha256 TEXT);
CREATE TABLE IF NOT EXISTS chunks (id INTEGER PRIMARY KEY, source TEXT, heading TEXT, text TEXT, vec TEXT);
""")
def normalize(vec: list[float]) -> list[float]:
norm = math.sqrt(sum(x * x for x in vec)) or 1.0
return [x / norm for x in vec]
def index_file(path: Path) -> int:
chunks = split_markdown(path)
vecs = embed(model=EMBED_MODEL, input=[c.text for c in chunks]).embeddings # no instruction on documents
db.execute("DELETE FROM chunks WHERE source = ?", (path.name,))
db.executemany(
"INSERT INTO chunks (source, heading, text, vec) VALUES (?, ?, ?, ?)",
[(c.source, c.heading, c.text, json.dumps(normalize(v))) for c, v in zip(chunks, vecs)],
)
return len(chunks)
def sync() -> None:
"""Re-index only changed documents โ files with the same hash are not embedded again"""
current = {p.name: p for p in Path(DOCS_DIR).glob("*.md")}
known = dict(db.execute("SELECT name, sha256 FROM files"))
for name, path in current.items():
digest = hashlib.sha256(path.read_bytes()).hexdigest()
if known.get(name) == digest:
continue
n = index_file(path)
db.execute("INSERT OR REPLACE INTO files VALUES (?, ?)", (name, digest))
print(f"indexed {name}: {n} chunks")
for name in set(known) - set(current): # clean up deleted documents
db.execute("DELETE FROM chunks WHERE source = ?", (name,))
db.execute("DELETE FROM files WHERE name = ?", (name,))
print(f"removed {name}")
db.commit()
if __name__ == "__main__":
sync()
print("total chunks:", db.execute("SELECT COUNT(*) FROM chunks").fetchone()[0])uv run store.py # first run: indexes all three files
uv run store.py # again: does nothing if no file changedEmbed the query with the instruction, compute cosine similarity against stored chunks and pick the top k. Vectors are normalized in advance, so the dot product is the cosine similarity.
import json
import sys
from ollama import embed
from config import EMBED_MODEL, QUERY_TASK
from store import db, normalize
def search(question: str, k: int = 3) -> list[dict]:
q = normalize(embed(model=EMBED_MODEL, input=[f"Instruct: {QUERY_TASK}\nQuery:{question}"]).embeddings[0])
scored = []
for source, heading, text, vec in db.execute("SELECT source, heading, text, vec FROM chunks"):
score = sum(a * b for a, b in zip(q, json.loads(vec))) # normalized vectors, so dot product = cosine similarity
scored.append({"source": source, "heading": heading, "text": text, "score": round(score, 3)})
return sorted(scored, key=lambda r: -r["score"])[:k]
if __name__ == "__main__":
for r in search(sys.argv[1] if len(sys.argv) > 1 else "How many vacation days do I get?"):
print(f"{r['score']:.3f} [{r['source']} > {r['heading']}] {r['text'][:50]}")uv run search.py "How many vacation days do I get?"
uv run search.py "Where can I see the cafeteria menu?" # a question with no matching document should score lowPass only chunks scoring at or above MIN_SCORE, numbered, and require a source number on every sentence. If no chunk clears the threshold, skip the model and reply "could not find" โ that stops answers made up from unrelated material at the source.
import sys
from ollama import chat
from config import CHAT_MODEL, MIN_SCORE
from search import search
SYSTEM = """You are a company policy assistant. Always answer in the language of the question.
- Answer only from the numbered [n] sources below, and add the source number to every sentence, like [1].
- If the sources do not contain the answer, reply only "I could not find this in the company documents." and do not guess."""
def answer(question: str, min_score: float = MIN_SCORE) -> dict:
hits = [h for h in search(question) if h["score"] >= min_score]
if not hits: # no relevant document โ do not call the model at all, so it cannot make up an answer
return {"answer": "I could not find this in the company documents.", "sources": []}
context = "\n\n".join(f"[{i}] ({h['source']} > {h['heading']})\n{h['text']}" for i, h in enumerate(hits, 1))
resp = chat(
model=CHAT_MODEL,
messages=[{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"Sources:\n{context}\n\nQuestion: {question}"}],
think=False,
options={"temperature": 0.2, "num_ctx": 8192},
)
return {"answer": resp.message.content, "sources": hits}
if __name__ == "__main__":
result = answer(sys.argv[1] if len(sys.argv) > 1 else "I'm in my second year. How many vacation days do I have?")
print(result["answer"])
for i, s in enumerate(result["sources"], 1):
print(f" [{i}] {s['source']} > {s['heading']} ({s['score']})")When you leave, return your laptop and badge to IT by your last working day [1].
[1] equipment.md > Returning equipment (0.612)
(sample output โ scores depend on the model and documents)These are factual answers, so thinking mode is off (think=False) and temperature is low. If many questions need policy interpretation, try thinking mode and compare response times.
The model may ignore instructions, dropping source numbers or citing ones that do not exist. Run a fast, reliable rule check on every sentence first, then, when needed, a model check that judges whether each claim is actually in the sources.
import re
from ollama import chat
from pydantic import BaseModel
from config import CHAT_MODEL
class Verdict(BaseModel):
supported: bool
unsupported_claims: list[str]
def check_citations(answer: str, n_sources: int) -> list[str]:
"""Rule check: find sentences without a source number or citing a number that does not exist"""
problems = []
# Move citations that follow the punctuation ("15 days. [1]") in front of it ("15 days [1].") before splitting sentences
text = re.sub(r"([.!?])\s*((?:\[\d+\]\s*)+)", lambda m: " " + m.group(2).strip() + m.group(1), answer.strip())
for sentence in re.split(r"(?<=[.!?])\s+", text):
if not sentence or "could not find" in sentence:
continue
cited = [int(n) for n in re.findall(r"\[(\d+)\]", sentence)]
if not cited:
problems.append(f"no source: {sentence}")
elif any(n < 1 or n > n_sources for n in cited):
problems.append(f"unknown source number: {sentence}")
return problems
def check_faithfulness(answer: str, sources: list[dict]) -> Verdict:
"""Model check: judge again whether each claim in the answer is backed by the sources"""
context = "\n\n".join(f"[{i}] {s['text']}" for i, s in enumerate(sources, 1))
resp = chat(
model=CHAT_MODEL,
messages=[{"role": "user", "content": f"Sources:\n{context}\n\nAnswer:\n{answer}\n\nFind any claims in the answer not supported by the sources."}],
format=Verdict.model_json_schema(),
think=False,
options={"temperature": 0},
)
return Verdict.model_validate_json(resp.message.content)
if __name__ == "__main__":
from answer import answer as ask
result = ask("I'm in my second year. How many vacation days do I have?")
print("rule check:", check_citations(result["answer"], len(result["sources"])) or "passed")
print("model check:", check_faithfulness(result["answer"], result["sources"]))>>> check_citations("You get 15 days. Half days are allowed [5].", n_sources=3)
['no source: You get 15 days.', 'unknown source number: Half days are allowed [5].']
>>> check_citations("You get 15 days of paid leave. [1] Request it in the HR portal[2].", n_sources=3)
[]Citations often come after the punctuation, as in "15 days. [1]". Splitting sentences as is would push the number into the next sentence and misjudge both, so the code moves citations in front of the punctuation before splitting.
Put both answerable and unanswerable questions in the evaluation set. An unanswerable question is correct only when no search result clears the threshold. Set MIN_SCORE between the "lowest top score of answerable questions" and the "highest of unanswerable ones" that the evaluation prints.
{"question": "How many vacation days do I get after one year?", "source": "leave-policy.md", "must_include": ["15 days"]}
{"question": "What do I do with my laptop when I leave the company?", "source": "equipment.md", "must_include": ["return"]}
{"question": "Where do I request a VPN account?", "source": "security.md", "must_include": ["portal"]}
{"question": "Where can I see the cafeteria menu?", "source": "-", "must_include": ["could not find"]}import json
from answer import answer
from config import MIN_SCORE
from search import search
cases = [json.loads(line) for line in open("data/qa.jsonl", encoding="utf-8")]
retrieval_hits = answer_hits = 0
answerable_scores, unanswerable_scores = [], []
for case in cases:
results = search(case["question"], k=3)
top = results[0]["score"] if results else 0.0
if case["source"] == "-": # the documents do not answer this question
unanswerable_scores.append(top)
retrieved = top < MIN_SCORE # correct only if nothing clears the threshold
else:
answerable_scores.append(top)
retrieved = case["source"] in [r["source"] for r in results]
retrieval_hits += retrieved
reply = answer(case["question"])["answer"]
ok = all(k.lower() in reply.lower() for k in case["must_include"])
answer_hits += ok
print(f"{'โ' if retrieved and ok else 'โ'} {case['question']} top score {top:.3f} ยท retrieval {'O' if retrieved else 'X'} ยท answer {'O' if ok else 'X'}")
n = len(cases)
print(f"retrieval {retrieval_hits}/{n} ยท answers {answer_hits}/{n}")
if answerable_scores and unanswerable_scores:
print(f"lowest top score of answerable questions {min(answerable_scores):.3f} / highest of unanswerable {max(unanswerable_scores):.3f}"
" โ set MIN_SCORE between the two")โ How many vacation days do I get after one year? top score 0.701 ยท retrieval O ยท answer O
โ What do I do with my laptop when I leave the company? top score 0.488 ยท retrieval O ยท answer X
โ Where do I request a VPN account? top score 0.655 ยท retrieval O ยท answer O
โ Where can I see the cafeteria menu? top score 0.312 ยท retrieval O ยท answer O
retrieval 4/4 ยท answers 3/4
lowest top score of answerable questions 0.488 / highest of unanswerable 0.312 โ set MIN_SCORE between the two
(sample output โ run with RAG_MIN_SCORE=0.5, which cut the laptop questionโs source. 0.4, between the two, passes all)RAG_MIN_SCORE=0.4 uv run evaluate.pyBuild the evaluation set from at least 50-100 questions employees actually asked. Too high a threshold misses answerable questions; too low a threshold answers from unrelated documents.
On startup the server re-indexes only changed documents, and each answer comes with its sources and the rule-check result. When needs_review is true, the UI shows a needs-review flag.
from contextlib import asynccontextmanager
from fastapi import FastAPI
from pydantic import BaseModel, Field
from answer import answer
from store import sync
from verify import check_citations
@asynccontextmanager
async def lifespan(app: FastAPI):
sync() # on startup, re-index only changed documents
yield
app = FastAPI(title="Company Docs Assistant", lifespan=lifespan)
class Question(BaseModel):
text: str = Field(min_length=1, max_length=500)
@app.post("/ask")
def ask(q: Question) -> dict:
result = answer(q.text)
problems = check_citations(result["answer"], len(result["sources"]))
return {
"answer": result["answer"],
"sources": [{"source": s["source"], "heading": s["heading"], "score": s["score"]} for s in result["sources"]],
"needs_review": bool(problems), # flag answers with unsupported sentences as "needs review" in the UI
"problems": problems,
}RAG_MIN_SCORE=0.4 uv run uvicorn app:app --port 8080| Symptom | Improvement |
|---|---|
| Weak on proper nouns and code names | Hybrid search that adds keyword search (BM25) |
| Top results in the wrong order | Take the top 20 and re-sort them with a reranker model |
| Many tables and attached PDFs | Convert PDFs to Markdown to keep table structure before indexing |
| Hundreds of thousands of chunks | Move to a vector database such as pgvector or Qdrant (replace only store.py and search.py) |
| Access differs by department | Store permission tags on chunks and filter at search time |
MIN_SCORE with an evaluation set that includes unanswerable questions?Up to tens of thousands of chunks, storing vectors in SQLite and comparing against all of them is fast enough. Beyond hundreds of thousands of chunks, or with many servers searching at once, move to a vector database such as pgvector or Qdrant. Only store.py and search.py need to change.
Score distributions differ by embedding model and documents, so there is no universally right value. Put both answerable and unanswerable questions in your evaluation set and choose a value between the lowest top score of answerable questions and the highest of unanswerable ones. Recheck when documents change.
Qwen3 Embedding supports over 100 languages and Qwen3.5 around 200, so documents and questions in different languages can be searched together. The system prompt tells the model to answer in the language of the question. Add questions in each language to your evaluation set, because quality varies by language and model size.
This tutorial blocks it in three layers: if no search result clears the threshold the model is not called at all, every sentence must carry a source number checked by a rule, and a model can re-judge support when needed. Flag answers with weak support as needs-review in the UI.