Skip to main content
AIDevOps
  • Learn
  • Learning Paths
  • Practice
  • Open Source
  • Books
  • Engineering

    AI DevOpsThe full map of building and operating AI servicesLLMOpsLLM deployment ยท evaluation ยท observabilityHands-on ProjectsBuild AI Agent projects

    Knowledge

    DocsTechnical documentationBlogEngineering articlesPloggerDevelopment log feed

    Validate

    Certification3-level skills certification ยท coming soon
AI Models
LlamaTutorial|MistralTutorial|GemmaTutorial|DeepSeekTutorial|QwenTutorial
๐Ÿง  AI Core
AI Intro & RoadmapML FundamentalsLLM Fundamentals|Python AIC++|PyTorchTensorFlowJAX
๐Ÿค– AI Applied Development
Applied AI Intro & RoadmapHugging FaceLangChainLlamaIndexLLMOps|LangGraphMCPMulti-AgentAgent Evaluation
๐Ÿง  AI Agent Development
Finance AI AgentLLM API ServerStock Investing AgentAIOps AI AgentEducation AI AgentCoding AI Agent
๐ŸŒฑ Spring Cloud
Spring Intro & RoadmapSpring Cloud GatewaySpring BootJava|Spring AISpring SecuritySpring BatchSpring JPA
๐Ÿณ DevOps
DevOps Intro & RoadmapLinuxDockerCI/CD|Kubernetes BasicsK8s AdvancedPrometheusGrafana
๐Ÿงฑ Infrastructure
Infrastructure Intro & RoadmapNginxRedis
โ˜๏ธ Cloud
Cloud Intro & RoadmapAWSGCPAzureNCPCloudflare
๐ŸŽจ Frontend
Frontend Intro & RoadmapJavaScriptTypeScript|ReactNext.js|VueNuxt|Electron
๐Ÿ“ฑ Mobile
Mobile Intro & RoadmapKotlinAndroidFlutter
โš™๏ธ Backend
Backend Intro & RoadmapPython BasicsFastAPIDjangoFlask|CGoGinNode.js
๐Ÿ’พ Database
DB Intro & RoadmapCore SQLOracleMySQLPostgreSQL|MongoDBVector DB
๐Ÿงช Testing
k6JMeternGrinder
AIDevOps

Engineering AI. From Code to Production.
An engineering learning platform for building and operating AI and AI Agents

Learn

  • All Guides
  • Learning Paths
  • AI Tutorial
  • Practice
  • Books

Resources

  • AI DevOps
  • LLMOps
  • Hands-on Projects
  • Docs
  • Blog
  • Plogger
  • Open Source
  • Certification (coming soon)

Start Here

  • AI Core Roadmap
  • AI Applied Development Roadmap
  • Spring Cloud Roadmap
  • DevOps Roadmap
  • Infrastructure Roadmap

ย 

  • Cloud Roadmap
  • Frontend Roadmap
  • Mobile Roadmap
  • Backend Roadmap
  • Database Roadmap
ยฉ 2026 AIDevOps. All rights reserved.
Terms of ServicePrivacy PolicySitemaptestforge.kr
Qwen hands-on tutorial

๐Ÿค– Build a company-docs RAG assistant with Qwen

Visitors

In ten steps, build an assistant that answers questions like "How many vacation days do I get?" or "What do I do with my laptop when I leave?" from company policy documents, with citations. Everything runs locally so documents never leave, unsupported answers are filtered out, and the retrieval threshold comes from evaluation data.

โ† Qwen guide
๐Ÿค–

Contents

0 / 12
  1. Architecture
  2. 1. Models and settings
  3. 2. Split documents
  4. 3. Index and detect changes
  5. 4. Search
  6. 5. Cited answers
  7. 6. Verify citations
  8. 7. Evaluate and set the threshold
  9. 8. API service
  10. 9. Next improvements
  11. Pre-production checklist
  12. FAQ
Contents 12 sections
  1. Architecture
  2. 1. Models and settings
  3. 2. Split documents
  4. 3. Index and detect changes
  5. 4. Search
  6. 5. Cited answers
  7. 6. Verify citations
  8. 7. Evaluate and set the threshold
  9. 8. API service
  10. 9. Next improvements
  11. Pre-production checklist
  12. FAQ

Architecture

RAG (retrieval-augmented generation) first finds document chunks related to the question, then answers using only those chunks as sources. In this tutorial both embeddings and answers come from local Qwen models, so company documents never leave your machine.

TEXT
docs/*.md  company policy documents
  โ”‚
  โ”œโ”€ chunk.py    split by headings, re-split long sections by length
  โ”œโ”€ store.py    index with Qwen3 Embedding โ†’ SQLite, re-index changed files only
  โ”œโ”€ search.py   embed the query (with instruction) โ†’ top-k by cosine similarity
  โ”œโ”€ answer.py   Qwen3.5 answers from chunks above the threshold, citing [n]
  โ”œโ”€ verify.py   rule check of citations + model check of support
  โ”œโ”€ evaluate.py retrieval and answer accuracy, choose MIN_SCORE
  โ””โ”€ app.py      FastAPI โ€” answer, sources, needs-review flag
โ„น๏ธ

Model choice, the embedding instruction rule and thinking mode are covered in the Qwen guide.

Step 1 โ€” Models and settings

BASH
ollama pull qwen3-embedding:0.6b   # embeddings (multilingual, up to 1024 dimensions)
ollama pull qwen3.5:9b             # answers (about 6.6 GB) โ€” qwen3.8:27b on a 24 GB GPU

uv init docs-assistant --python 3.12
cd docs-assistant
uv add ollama pydantic fastapi "uvicorn[standard]"
config.pyPYTHON
import os

EMBED_MODEL = "qwen3-embedding:0.6b"
CHAT_MODEL = os.environ.get("QWEN_CHAT_MODEL", "qwen3.5:9b")   # qwen3.8:27b on a 24 GB GPU
DB_PATH = "rag.db"

# Search results scoring below this are not used as sources. Score distributions differ by embedding model and documents,
# so set it in step 7 between the lowest top score of answerable questions and the highest of unanswerable ones.
MIN_SCORE = float(os.environ.get("RAG_MIN_SCORE", "0.35"))
DOCS_DIR = "docs"

# Instruction added to search queries only (the format recommended on the Qwen3 Embedding model card)
QUERY_TASK = "Given a web search query, retrieve relevant passages that answer the query"
๐Ÿ’ก

Qwen3 Embedding expects an instruction in the form Instruct: task description\nQuery:question on search queries only, not on documents. Adding it to documents or leaving it off queries lowers retrieval accuracy.

Step 2 โ€” Split documents

Retrieval is accurate when each chunk covers one topic. Split by Markdown headings first, re-split only long sections by length, and overlap slightly so boundary sentences are not lost. Store the file name and heading with each chunk so answers can show where a source came from.

TEXT
docs/
โ”œโ”€โ”€ leave-policy.md
โ”œโ”€โ”€ equipment.md
โ””โ”€โ”€ security.md
chunk.pyPYTHON
import re
from dataclasses import dataclass
from pathlib import Path


@dataclass
class Chunk:
    source: str      # file name
    heading: str     # nearest heading โ€” shown as the source location in answers
    text: str


def split_markdown(path: Path, max_chars: int = 500, overlap: int = 80) -> list[Chunk]:
    """Split by headings (#) first, then re-split only long sections by character count.
    Character counts behave predictably across languages, unlike word counts."""
    chunks: list[Chunk] = []
    heading, buf = path.stem, []

    def flush():
        body = "\n".join(buf).strip()
        start = 0
        while body and start < len(body):
            chunks.append(Chunk(path.name, heading, body[start:start + max_chars]))
            start += max_chars - overlap
        buf.clear()

    for line in path.read_text(encoding="utf-8").splitlines():
        m = re.match(r"^(#{1,3})\s+(.*)", line)
        if m:
            flush()
            heading = m.group(2).strip()
        else:
            buf.append(line)
    flush()
    return chunks


if __name__ == "__main__":
    for c in split_markdown(Path("docs/leave-policy.md")):
        print(f"[{c.source} > {c.heading}] {c.text[:60]}")
TEXT
[leave-policy.md > Annual leave] Employees in their first year earn one day of leave per month, ...
[leave-policy.md > Requesting leave] Request leave in the HR portal at least three days in advance ...

Step 3 โ€” Index and detect changes

Embed the chunks and store them in SQLite. Record a hash per file, so rerunning re-embeds only changed files and cleans up deleted ones. Even with hundreds of documents you never need to re-index everything daily.

store.pyPYTHON
import hashlib
import json
import math
import sqlite3
from pathlib import Path
from ollama import embed
from chunk import split_markdown
from config import DB_PATH, DOCS_DIR, EMBED_MODEL

db = sqlite3.connect(DB_PATH, check_same_thread=False)
db.executescript("""
CREATE TABLE IF NOT EXISTS files (name TEXT PRIMARY KEY, sha256 TEXT);
CREATE TABLE IF NOT EXISTS chunks (id INTEGER PRIMARY KEY, source TEXT, heading TEXT, text TEXT, vec TEXT);
""")


def normalize(vec: list[float]) -> list[float]:
    norm = math.sqrt(sum(x * x for x in vec)) or 1.0
    return [x / norm for x in vec]


def index_file(path: Path) -> int:
    chunks = split_markdown(path)
    vecs = embed(model=EMBED_MODEL, input=[c.text for c in chunks]).embeddings   # no instruction on documents
    db.execute("DELETE FROM chunks WHERE source = ?", (path.name,))
    db.executemany(
        "INSERT INTO chunks (source, heading, text, vec) VALUES (?, ?, ?, ?)",
        [(c.source, c.heading, c.text, json.dumps(normalize(v))) for c, v in zip(chunks, vecs)],
    )
    return len(chunks)


def sync() -> None:
    """Re-index only changed documents โ€” files with the same hash are not embedded again"""
    current = {p.name: p for p in Path(DOCS_DIR).glob("*.md")}
    known = dict(db.execute("SELECT name, sha256 FROM files"))
    for name, path in current.items():
        digest = hashlib.sha256(path.read_bytes()).hexdigest()
        if known.get(name) == digest:
            continue
        n = index_file(path)
        db.execute("INSERT OR REPLACE INTO files VALUES (?, ?)", (name, digest))
        print(f"indexed {name}: {n} chunks")
    for name in set(known) - set(current):          # clean up deleted documents
        db.execute("DELETE FROM chunks WHERE source = ?", (name,))
        db.execute("DELETE FROM files WHERE name = ?", (name,))
        print(f"removed {name}")
    db.commit()


if __name__ == "__main__":
    sync()
    print("total chunks:", db.execute("SELECT COUNT(*) FROM chunks").fetchone()[0])
BASH
uv run store.py      # first run: indexes all three files
uv run store.py      # again: does nothing if no file changed

Step 4 โ€” Search

Embed the query with the instruction, compute cosine similarity against stored chunks and pick the top k. Vectors are normalized in advance, so the dot product is the cosine similarity.

search.pyPYTHON
import json
import sys
from ollama import embed
from config import EMBED_MODEL, QUERY_TASK
from store import db, normalize


def search(question: str, k: int = 3) -> list[dict]:
    q = normalize(embed(model=EMBED_MODEL, input=[f"Instruct: {QUERY_TASK}\nQuery:{question}"]).embeddings[0])
    scored = []
    for source, heading, text, vec in db.execute("SELECT source, heading, text, vec FROM chunks"):
        score = sum(a * b for a, b in zip(q, json.loads(vec)))     # normalized vectors, so dot product = cosine similarity
        scored.append({"source": source, "heading": heading, "text": text, "score": round(score, 3)})
    return sorted(scored, key=lambda r: -r["score"])[:k]


if __name__ == "__main__":
    for r in search(sys.argv[1] if len(sys.argv) > 1 else "How many vacation days do I get?"):
        print(f"{r['score']:.3f}  [{r['source']} > {r['heading']}] {r['text'][:50]}")
BASH
uv run search.py "How many vacation days do I get?"
uv run search.py "Where can I see the cafeteria menu?"   # a question with no matching document should score low

Step 5 โ€” Answers with citations

Pass only chunks scoring at or above MIN_SCORE, numbered, and require a source number on every sentence. If no chunk clears the threshold, skip the model and reply "could not find" โ€” that stops answers made up from unrelated material at the source.

answer.pyPYTHON
import sys
from ollama import chat
from config import CHAT_MODEL, MIN_SCORE
from search import search

SYSTEM = """You are a company policy assistant. Always answer in the language of the question.
- Answer only from the numbered [n] sources below, and add the source number to every sentence, like [1].
- If the sources do not contain the answer, reply only "I could not find this in the company documents." and do not guess."""


def answer(question: str, min_score: float = MIN_SCORE) -> dict:
    hits = [h for h in search(question) if h["score"] >= min_score]
    if not hits:   # no relevant document โ†’ do not call the model at all, so it cannot make up an answer
        return {"answer": "I could not find this in the company documents.", "sources": []}
    context = "\n\n".join(f"[{i}] ({h['source']} > {h['heading']})\n{h['text']}" for i, h in enumerate(hits, 1))
    resp = chat(
        model=CHAT_MODEL,
        messages=[{"role": "system", "content": SYSTEM},
                  {"role": "user", "content": f"Sources:\n{context}\n\nQuestion: {question}"}],
        think=False,
        options={"temperature": 0.2, "num_ctx": 8192},
    )
    return {"answer": resp.message.content, "sources": hits}


if __name__ == "__main__":
    result = answer(sys.argv[1] if len(sys.argv) > 1 else "I'm in my second year. How many vacation days do I have?")
    print(result["answer"])
    for i, s in enumerate(result["sources"], 1):
        print(f"  [{i}] {s['source']} > {s['heading']} ({s['score']})")
TEXT
When you leave, return your laptop and badge to IT by your last working day [1].
  [1] equipment.md > Returning equipment (0.612)

(sample output โ€” scores depend on the model and documents)
โ„น๏ธ

These are factual answers, so thinking mode is off (think=False) and temperature is low. If many questions need policy interpretation, try thinking mode and compare response times.

Step 6 โ€” Verify citations

The model may ignore instructions, dropping source numbers or citing ones that do not exist. Run a fast, reliable rule check on every sentence first, then, when needed, a model check that judges whether each claim is actually in the sources.

verify.pyPYTHON
import re
from ollama import chat
from pydantic import BaseModel
from config import CHAT_MODEL


class Verdict(BaseModel):
    supported: bool
    unsupported_claims: list[str]


def check_citations(answer: str, n_sources: int) -> list[str]:
    """Rule check: find sentences without a source number or citing a number that does not exist"""
    problems = []
    # Move citations that follow the punctuation ("15 days. [1]") in front of it ("15 days [1].") before splitting sentences
    text = re.sub(r"([.!?])\s*((?:\[\d+\]\s*)+)", lambda m: " " + m.group(2).strip() + m.group(1), answer.strip())
    for sentence in re.split(r"(?<=[.!?])\s+", text):
        if not sentence or "could not find" in sentence:
            continue
        cited = [int(n) for n in re.findall(r"\[(\d+)\]", sentence)]
        if not cited:
            problems.append(f"no source: {sentence}")
        elif any(n < 1 or n > n_sources for n in cited):
            problems.append(f"unknown source number: {sentence}")
    return problems


def check_faithfulness(answer: str, sources: list[dict]) -> Verdict:
    """Model check: judge again whether each claim in the answer is backed by the sources"""
    context = "\n\n".join(f"[{i}] {s['text']}" for i, s in enumerate(sources, 1))
    resp = chat(
        model=CHAT_MODEL,
        messages=[{"role": "user", "content": f"Sources:\n{context}\n\nAnswer:\n{answer}\n\nFind any claims in the answer not supported by the sources."}],
        format=Verdict.model_json_schema(),
        think=False,
        options={"temperature": 0},
    )
    return Verdict.model_validate_json(resp.message.content)


if __name__ == "__main__":
    from answer import answer as ask
    result = ask("I'm in my second year. How many vacation days do I have?")
    print("rule check:", check_citations(result["answer"], len(result["sources"])) or "passed")
    print("model check:", check_faithfulness(result["answer"], result["sources"]))
TEXT
>>> check_citations("You get 15 days. Half days are allowed [5].", n_sources=3)
['no source: You get 15 days.', 'unknown source number: Half days are allowed [5].']

>>> check_citations("You get 15 days of paid leave. [1] Request it in the HR portal[2].", n_sources=3)
[]
โš ๏ธ

Citations often come after the punctuation, as in "15 days. [1]". Splitting sentences as is would push the number into the next sentence and misjudge both, so the code moves citations in front of the punctuation before splitting.

Step 7 โ€” Evaluate and set the retrieval threshold

Put both answerable and unanswerable questions in the evaluation set. An unanswerable question is correct only when no search result clears the threshold. Set MIN_SCORE between the "lowest top score of answerable questions" and the "highest of unanswerable ones" that the evaluation prints.

data/qa.jsonlJSON
{"question": "How many vacation days do I get after one year?", "source": "leave-policy.md", "must_include": ["15 days"]}
{"question": "What do I do with my laptop when I leave the company?", "source": "equipment.md", "must_include": ["return"]}
{"question": "Where do I request a VPN account?", "source": "security.md", "must_include": ["portal"]}
{"question": "Where can I see the cafeteria menu?", "source": "-", "must_include": ["could not find"]}
evaluate.pyPYTHON
import json
from answer import answer
from config import MIN_SCORE
from search import search

cases = [json.loads(line) for line in open("data/qa.jsonl", encoding="utf-8")]

retrieval_hits = answer_hits = 0
answerable_scores, unanswerable_scores = [], []
for case in cases:
    results = search(case["question"], k=3)
    top = results[0]["score"] if results else 0.0
    if case["source"] == "-":                       # the documents do not answer this question
        unanswerable_scores.append(top)
        retrieved = top < MIN_SCORE                  # correct only if nothing clears the threshold
    else:
        answerable_scores.append(top)
        retrieved = case["source"] in [r["source"] for r in results]
    retrieval_hits += retrieved
    reply = answer(case["question"])["answer"]
    ok = all(k.lower() in reply.lower() for k in case["must_include"])
    answer_hits += ok
    print(f"{'โœ“' if retrieved and ok else 'โœ—'} {case['question']}  top score {top:.3f} ยท retrieval {'O' if retrieved else 'X'} ยท answer {'O' if ok else 'X'}")

n = len(cases)
print(f"retrieval {retrieval_hits}/{n} ยท answers {answer_hits}/{n}")
if answerable_scores and unanswerable_scores:
    print(f"lowest top score of answerable questions {min(answerable_scores):.3f} / highest of unanswerable {max(unanswerable_scores):.3f}"
          " โ†’ set MIN_SCORE between the two")
TEXT
โœ“ How many vacation days do I get after one year?  top score 0.701 ยท retrieval O ยท answer O
โœ— What do I do with my laptop when I leave the company?  top score 0.488 ยท retrieval O ยท answer X
โœ“ Where do I request a VPN account?  top score 0.655 ยท retrieval O ยท answer O
โœ“ Where can I see the cafeteria menu?  top score 0.312 ยท retrieval O ยท answer O
retrieval 4/4 ยท answers 3/4
lowest top score of answerable questions 0.488 / highest of unanswerable 0.312 โ†’ set MIN_SCORE between the two

(sample output โ€” run with RAG_MIN_SCORE=0.5, which cut the laptop questionโ€™s source. 0.4, between the two, passes all)
BASH
RAG_MIN_SCORE=0.4 uv run evaluate.py
๐Ÿ’ก

Build the evaluation set from at least 50-100 questions employees actually asked. Too high a threshold misses answerable questions; too low a threshold answers from unrelated documents.

Step 8 โ€” API service

On startup the server re-indexes only changed documents, and each answer comes with its sources and the rule-check result. When needs_review is true, the UI shows a needs-review flag.

app.pyPYTHON
from contextlib import asynccontextmanager
from fastapi import FastAPI
from pydantic import BaseModel, Field
from answer import answer
from store import sync
from verify import check_citations


@asynccontextmanager
async def lifespan(app: FastAPI):
    sync()                      # on startup, re-index only changed documents
    yield


app = FastAPI(title="Company Docs Assistant", lifespan=lifespan)


class Question(BaseModel):
    text: str = Field(min_length=1, max_length=500)


@app.post("/ask")
def ask(q: Question) -> dict:
    result = answer(q.text)
    problems = check_citations(result["answer"], len(result["sources"]))
    return {
        "answer": result["answer"],
        "sources": [{"source": s["source"], "heading": s["heading"], "score": s["score"]} for s in result["sources"]],
        "needs_review": bool(problems),    # flag answers with unsupported sentences as "needs review" in the UI
        "problems": problems,
    }
BASH
RAG_MIN_SCORE=0.4 uv run uvicorn app:app --port 8080

Step 9 โ€” Next improvements

SymptomImprovement
Weak on proper nouns and code namesHybrid search that adds keyword search (BM25)
Top results in the wrong orderTake the top 20 and re-sort them with a reranker model
Many tables and attached PDFsConvert PDFs to Markdown to keep table structure before indexing
Hundreds of thousands of chunksMove to a vector database such as pgvector or Qdrant (replace only store.py and search.py)
Access differs by departmentStore permission tags on chunks and filter at search time

Pre-production checklist

  • Did you set MIN_SCORE with an evaluation set that includes unanswerable questions?
  • Are documents re-indexed automatically when they change (on startup or on a schedule)?
  • Does the UI flag unsupported answers as needs-review?
  • Are documents that a department or role must not see kept out of search results?
  • Are questions and answers logged without personal data?
  • Do document owners review the list of unanswered questions and fill the gaps?

Frequently asked questions

Is SQLite enough, without a vector database?

Up to tens of thousands of chunks, storing vectors in SQLite and comparing against all of them is fast enough. Beyond hundreds of thousands of chunks, or with many servers searching at once, move to a vector database such as pgvector or Qdrant. Only store.py and search.py need to change.

How do I choose the retrieval threshold (MIN_SCORE)?

Score distributions differ by embedding model and documents, so there is no universally right value. Put both answerable and unanswerable questions in your evaluation set and choose a value between the lowest top score of answerable questions and the highest of unanswerable ones. Recheck when documents change.

Does it work with documents in several languages?

Qwen3 Embedding supports over 100 languages and Qwen3.5 around 200, so documents and questions in different languages can be searched together. The system prompt tells the model to answer in the language of the question. Add questions in each language to your evaluation set, because quality varies by language and model size.

What if the answer makes up something not in the documents?

This tutorial blocks it in three layers: if no search result clears the threshold the model is not called at all, every sentence must carry a source number checked by a rule, and a model can re-judge support when needed. Flag answers with weak support as needs-review in the UI.