Skip to main content
AIDevOps
  • Learn
  • Learning Paths
  • Practice
  • Open Source
  • Books
  • Engineering

    AI DevOpsThe full map of building and operating AI servicesLLMOpsLLM deployment Β· evaluation Β· observabilityHands-on ProjectsBuild AI Agent projects

    Knowledge

    DocsTechnical documentationBlogEngineering articlesPloggerDevelopment log feed

    Validate

    Certification3-level skills certification Β· coming soon
AI Models
LlamaTutorial|MistralTutorial|GemmaTutorial|DeepSeekTutorial|Qwen
🧠 AI Core
AI Intro & RoadmapML FundamentalsLLM Fundamentals|Python AIC++|PyTorchTensorFlowJAX
πŸ€– AI Applied Development
Applied AI Intro & RoadmapHugging FaceLangChainLlamaIndexLLMOps|LangGraphMCPMulti-AgentAgent Evaluation
🧠 AI Agent Development
Finance AI AgentLLM API ServerStock Investing AgentAIOps AI AgentEducation AI AgentCoding AI Agent
🌱 Spring Cloud
Spring Intro & RoadmapSpring Cloud GatewaySpring BootJava|Spring AISpring SecuritySpring BatchSpring JPA
🐳 DevOps
DevOps Intro & RoadmapLinuxDockerCI/CD|Kubernetes BasicsK8s AdvancedPrometheusGrafana
🧱 Infrastructure
Infrastructure Intro & RoadmapNginxRedis
☁️ Cloud
Cloud Intro & RoadmapAWSGCPAzureNCPCloudflare
🎨 Frontend
Frontend Intro & RoadmapJavaScriptTypeScript|ReactNext.js|VueNuxt
πŸ“± Mobile
Mobile Intro & RoadmapKotlinAndroidFlutter
βš™οΈ Backend
Backend Intro & RoadmapPython BasicsFastAPIDjangoFlask|CGoGinNode.js
πŸ’Ύ Database
DB Intro & RoadmapCore SQLOracleMySQLPostgreSQL|MongoDBVector DB
πŸ§ͺ Testing
k6JMeternGrinder
AIDevOps

Engineering AI. From Code to Production.
An engineering learning platform for building and operating AI and AI Agents

Learn

  • All Guides
  • Learning Paths
  • Practice
  • Books

Resources

  • AI DevOps
  • LLMOps
  • Hands-on Projects
  • Docs
  • Blog
  • Plogger
  • Open Source
  • Certification (coming soon)

Start Here

  • AI Core Roadmap
  • AI Applied Development Roadmap
  • Spring Cloud Roadmap
  • DevOps Roadmap
  • Infrastructure Roadmap

Β 

  • Cloud Roadmap
  • Frontend Roadmap
  • Mobile Roadmap
  • Backend Roadmap
  • Database Roadmap
Β© 2026 AIDevOps. All rights reserved.
Terms of ServicePrivacy PolicySitemaptestforge.kr
Llama hands-on tutorial

πŸ¦™ Llama from install to fine-tuning

Visitors

Start by running Llama on your own machine, then train it on your team’s data, import it into Ollama, and serve it behind an API β€” in ten steps. Every step comes with commands to verify it worked and the places people usually get stuck.

← Llama guide
πŸ¦™

Contents

0 / 13
  1. Roadmap
  2. 1. Check your hardware
  3. 2. Python environment
  4. 3. Install Ollama
  5. 4. First chat & parameters
  6. 5. Hugging Face model
  7. 6. Prepare training data
  8. 7. Train with QLoRA
  9. 8. Evaluate the result
  10. 9. GGUF & Ollama import
  11. 10. Serve it as an API
  12. Pre-production checklist
  13. FAQ
Contents 13 sections
  1. Roadmap
  2. 1. Check your hardware
  3. 2. Python environment
  4. 3. Install Ollama
  5. 4. First chat & parameters
  6. 5. Hugging Face model
  7. 6. Prepare training data
  8. 7. Train with QLoRA
  9. 8. Evaluate the result
  10. 9. GGUF & Ollama import
  11. 10. Serve it as an API
  12. Pre-production checklist
  13. FAQ

Roadmap

This tutorial follows one running example β€” an incident-response assistant for an engineering team β€” from installing Llama to training it on your data and serving it behind an API. Steps 1-4 work without a GPU; the training steps (5-9) need an NVIDIA GPU.

StepWhat you doResultGPU
1-2Check hardware, set up Python and PyTorchA working dev environmentOptional
3-4Install Ollama, first chat, generation parametersA local Llama API on port 11434Optional
5Download the original model from Hugging Face, understand the chat templateOriginal weights for trainingRecommended
6-7Prepare training data, fine-tune with QLoRAA LoRA adapterRequired
8Compare answers before and after trainingAn evaluation reportRequired
9-10Merge, convert to GGUF, import into Ollama, serve with FastAPIYour own model APIOptional
ℹ️

The hands-on model is Llama 3.2 3B Instruct, which can be trained on a laptop GPU. Change only the model ID and the same code scales to Llama 3.1 8B. For model differences and licensing, read the Llama guide first.

Step 1 β€” Check your hardware

GPU memory (VRAM) decides which model sizes you can run and train. Check your machine first.

BASH
# NVIDIA GPU name, driver version, VRAM usage
nvidia-smi

# macOS (Apple Silicon) β€” unified memory size
system_profiler SPHardwareDataType | grep -E "Chip|Memory"

# Linux β€” CPU and memory
lscpu | grep "Model name"; free -h
Your machineRun (inference)Train (QLoRA)
No GPU, 16 GB RAMLlama 3.2 3B, 3.1 8B (slow)Use a cloud GPU
Apple Silicon 16 GB+Llama 3.1 8B runs smoothlyThis tutorial’s 4-bit training (bitsandbytes) targets NVIDIA CUDA β€” use a cloud GPU
NVIDIA 8-12 GBLlama 3.1 8B (Q4)Llama 3.2 1B / 3B
NVIDIA 16-24 GBLlama 3.1 8B (Q8/FP16)Llama 3.1 8B
NVIDIA 48 GB+Llama 3.3 70B (Q4)Llama 3.3 70B
πŸ’‘

On Windows, run Ollama natively and do training and serving in WSL2 (Ubuntu). WSL2 uses the Windows NVIDIA driver for GPU access, so do not install a Linux driver inside WSL. If nvidia-smi works inside WSL, you are ready.

Step 2 β€” Set up a Python environment

Use a per-project virtual environment to avoid version conflicts. This tutorial uses the fast package manager uv.

BASH
# Install uv (macOS/Linux/WSL)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Create the project
uv init llama-lab --python 3.12
cd llama-lab

PyTorch has to match your GPU and CUDA version. Pick your OS, package manager, and CUDA version in the PyTorch install selector and use the command it gives you. Then confirm the GPU is visible:

check_gpu.pyPYTHON
import torch

print("PyTorch:", torch.__version__)
print("CUDA available:", torch.cuda.is_available())
if torch.cuda.is_available():
    props = torch.cuda.get_device_properties(0)
    print("GPU:", props.name)
    print(f"VRAM: {props.total_memory / 1024**3:.1f} GB")
    print("BF16 supported:", torch.cuda.is_bf16_supported())   # if False, train in fp16
BASH
uv run check_gpu.py

# Packages used in later steps
uv add "transformers>=4.56" "trl>=1.0" peft datasets accelerate bitsandbytes ollama "huggingface_hub[cli]"
⚠️

If you see CUDA available: False, the CPU build of PyTorch is installed. Reinstall the CUDA build before training β€” otherwise training runs on the CPU without any error, just extremely slowly.

Step 3 β€” Install and verify Ollama

Ollama handles model downloads, GPU placement, and the API server in one tool. Install it for your OS and confirm the server is up.

BASH
# Linux / WSL2
curl -fsSL https://ollama.com/install.sh | sh

# macOS β€” download the app from ollama.com, or
brew install ollama

# Windows β€” download and run the installer (OllamaSetup.exe) from ollama.com

# Verify
ollama --version
curl http://localhost:11434/api/version     # the server responds
BASH
# Pull the models used here
ollama pull llama3.2          # Llama 3.2 3B (about 2 GB)
ollama pull llama3.1          # Llama 3.1 8B (about 4.9 GB)
ollama list

Server settings you will change often are environment variables:

VariableDefaultUse
OLLAMA_MODELSmacOS ~/.ollama/models, Linux /usr/share/ollama/.ollama/modelsMove models to a larger disk
OLLAMA_HOST127.0.0.1:11434Allow other machines to connect (0.0.0.0)
OLLAMA_CONTEXT_LENGTH4K-256K depending on VRAMDefault context length
OLLAMA_KEEP_ALIVE5 minutesHow long a model stays loaded after a request
OLLAMA_NUM_PARALLEL1Concurrent requests per model (memory grows with it)
BASH
# Linux (systemd) β€” add variables to the service
sudo systemctl edit ollama.service
#   [Service]
#   Environment="OLLAMA_CONTEXT_LENGTH=16384"
#   Environment="OLLAMA_KEEP_ALIVE=30m"
sudo systemctl daemon-reload && sudo systemctl restart ollama

# macOS β€” then restart the Ollama app
launchctl setenv OLLAMA_CONTEXT_LENGTH 16384
⚠️

The Ollama API has no authentication, so with OLLAMA_HOST=0.0.0.0 anyone on the network can use your models. If you expose it, put authentication and IP restrictions in a reverse proxy such as Nginx.

Step 4 β€” First chat and generation parameters

Check that it works in the terminal, then call the REST API your application will actually use.

BASH
ollama run llama3.2
>>> What should I check, in order, when a Kubernetes pod is in CrashLoopBackOff?
>>> /set parameter temperature 0.2     # change a parameter mid-chat
>>> /bye
BASH
curl http://localhost:11434/api/chat -d '{
  "model": "llama3.2",
  "messages": [
    {"role": "system", "content": "You are an SRE. Answer step by step and keep it short."},
    {"role": "user", "content": "Five metrics to check when API latency suddenly jumps"}
  ],
  "stream": false,
  "options": {"temperature": 0.2, "num_ctx": 8192, "num_predict": 512}
}'
ParameterMeaningSuggested value
temperatureRandomness. Lower means the same question gets the same answer0-0.3 for facts and classification, ~0.7 for writing
top_pSample only from the top cumulative-probability candidates0.9 (tune this or temperature, not both)
num_ctxContext length for input plus output8K-32K, sized to your conversation or documents
num_predictMaximum tokens to generateCaps answer length (cost and latency)
seedRandom seedFix it for evaluation and tests to make results reproducible
first_chat.pyPYTHON
from ollama import chat

messages = [{"role": "system", "content": "You are an SRE. Answer step by step and keep it short."}]

while True:
    user = input("you> ").strip()
    if user in {"exit", "quit"}:
        break
    messages.append({"role": "user", "content": user})

    answer = ""
    for part in chat(model="llama3.2", messages=messages, stream=True, options={"temperature": 0.2}):
        print(part.message.content, end="", flush=True)
        answer += part.message.content
    print()

    messages.append({"role": "assistant", "content": answer})   # keep the conversation context
πŸ’‘

The model does not remember the conversation. You have to keep sending previous questions and answers in messages, and they take up context. Summarize or drop old messages in long conversations.

Step 5 β€” Download the original model from Hugging Face

Ollama models are quantized files for inference and cannot be used for training. Training uses the original weights (safetensors) from Hugging Face. Llama repositories are gated and require accepting the license.

  1. Sign in to Hugging Face, open meta-llama/Llama-3.2-3B-Instruct, accept the license, and wait for approval.
  2. Create a token with Read permission under Settings β†’ Access Tokens.
  3. Log in from the terminal and download the model.
BASH
uv run hf auth login                # paste the token
uv run hf download meta-llama/Llama-3.2-3B-Instruct --local-dir ./models/llama-3.2-3b-instruct

Llama was trained on conversations wrapped in special tokens β€” the chat template. Training data and inference input must follow it, so never build the string by hand; use the tokenizer’s apply_chat_template.

chat_template.pyPYTHON
import torch
from transformers import AutoTokenizer, pipeline

MODEL_DIR = "./models/llama-3.2-3b-instruct"
messages = [
    {"role": "system", "content": "You are an SRE."},
    {"role": "user", "content": "A disk usage alert just fired"},
]

# ── See the exact string the model receives ──────
tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
print(tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True))

# ── Run inference with Transformers ──────────────
pipe = pipeline("text-generation", model=MODEL_DIR, dtype=torch.bfloat16, device_map="auto")
out = pipe(messages, max_new_tokens=256, do_sample=False)
print(out[0]["generated_text"][-1]["content"])
TEXT
<|begin_of_text|><|start_header_id|>system<|end_header_id|>

Cutting Knowledge Date: December 2023
Today Date: ...

You are an SRE.<|eot_id|><|start_header_id|>user<|end_header_id|>

A disk usage alert just fired<|eot_id|><|start_header_id|>assistant<|end_header_id|>

(sample output β€” the date line depends on when you run it)

Step 6 β€” Prepare training data

Data quality decides almost everything about a fine-tune. You are showing the model example answers β€” "for this kind of question, answer like this" β€” so write the answers exactly in the format and tone you want, and make them correct.

PrincipleWhat it means
Consistent formatAnswer the same kind of question with the same structure (e.g. causes β†’ checks β†’ actions)
CorrectnessThe model learns wrong answers too. Have an expert review the data
VarietyInclude several phrasings of the same intent
Remove personal dataMask names, emails, tokens, and internal IPs before training
Hold out evaluation dataKeep about 10% out of training for evaluation

Store the data as JSONL in the conversational prompt-completion format TRL reads directly. With this format the loss is computed only on the completion, so the model does not spend capacity memorizing the questions.

data/incidents.jsonlJSON
{"prompt": [{"role": "system", "content": "Incident-response assistant"}, {"role": "user", "content": "API 5xx rate suddenly went above 10%"}], "completion": [{"role": "assistant", "content": "Likely causes: a recent deploy, an upstream outage, DB connection exhaustion\nCheck: 1) deploy history 2) upstream latency dashboard 3) DB connection pool usage\nAction: if it lines up with the last deploy, consider rolling back first."}]}
{"prompt": [{"role": "system", "content": "Incident-response assistant"}, {"role": "user", "content": "A node disk usage alert fired"}], "completion": [{"role": "assistant", "content": "Likely causes: container logs piling up, image cache, large temp files\nCheck: 1) df -h 2) du on the largest directories 3) log rotation settings\nAction: prune old images and shorten log retention."}]}
prepare_data.pyPYTHON
from datasets import load_dataset
from transformers import AutoTokenizer

MODEL_DIR = "./models/llama-3.2-3b-instruct"
MAX_LENGTH = 1024
tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)

ds = load_dataset("json", data_files="data/incidents.jsonl", split="train")

# ── 1) Drop duplicate questions ──────────────────
seen = set()
def first_time(example) -> bool:
    key = example["prompt"][-1]["content"].strip()
    if key in seen:
        return False
    seen.add(key)
    return True
ds = ds.filter(first_time)

# ── 2) Check token length β€” beyond MAX_LENGTH the end (the answer) is cut off ─
def count_tokens(example) -> dict:
    enc = tokenizer.apply_chat_template(example["prompt"] + example["completion"], tokenize=True, return_dict=True)
    return {"n_tokens": len(enc["input_ids"])}
ds = ds.map(count_tokens)
too_long = ds.filter(lambda e: e["n_tokens"] > MAX_LENGTH)
print(f"{len(ds)} examples, max {max(ds['n_tokens'])} tokens, {len(too_long)} over {MAX_LENGTH}")
ds = ds.filter(lambda e: e["n_tokens"] <= MAX_LENGTH).remove_columns("n_tokens")

# ── 3) Train / eval split ────────────────────────
split = ds.train_test_split(test_size=0.1, seed=42)
split["train"].to_json("data/train.jsonl", force_ascii=False)
split["test"].to_json("data/eval.jsonl", force_ascii=False)
print("train:", len(split["train"]), "eval:", len(split["test"]))
πŸ’‘

Start with 200-500 examples and run one full train-and-evaluate loop. Strengthening the weak question types based on the results is faster than producing a large dataset up front.

Step 7 β€” Train with QLoRA

QLoRA loads the original model in 4-bit and trains only small LoRA matrices added to each layer. Only about 1% of the parameters are trained, so Llama 3.2 3B fits on an 8-12 GB GPU.

train.pyPYTHON
import torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer

MODEL_DIR = "./models/llama-3.2-3b-instruct"   # scale up: meta-llama/Llama-3.1-8B-Instruct
OUTPUT_DIR = "./outputs/llama-3.2-3b-incident-lora"

data = load_dataset("json", data_files={"train": "data/train.jsonl", "eval": "data/eval.jsonl"})
bf16 = torch.cuda.is_bf16_supported()

trainer = SFTTrainer(
    model=MODEL_DIR,
    train_dataset=data["train"],
    eval_dataset=data["eval"],
    quantization_config=BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16 if bf16 else torch.float16,
        bnb_4bit_use_double_quant=True,
    ),
    peft_config=LoraConfig(
        r=16, lora_alpha=32, lora_dropout=0.05,
        target_modules="all-linear", task_type="CAUSAL_LM",
    ),
    args=SFTConfig(
        output_dir=OUTPUT_DIR,
        num_train_epochs=3,
        per_device_train_batch_size=4,
        gradient_accumulation_steps=4,       # effective batch = 4 Γ— 4 = 16
        learning_rate=1e-4,
        lr_scheduler_type="cosine",
        warmup_steps=10,
        max_length=1024,
        bf16=bf16,
        fp16=not bf16,
        logging_steps=10,
        eval_strategy="steps",
        eval_steps=50,
        save_strategy="steps",
        save_steps=50,
        save_total_limit=2,
        load_best_model_at_end=True,         # keep the checkpoint with the lowest eval_loss
        metric_for_best_model="eval_loss",
        model_init_kwargs={"dtype": torch.bfloat16 if bf16 else torch.float16},
    ),
)

trainer.train()
trainer.save_model()          # saves the LoRA adapter (tens of MB) to OUTPUT_DIR
print("saved to:", OUTPUT_DIR)
BASH
uv run train.py

# Watch GPU usage from another terminal
watch -n 2 nvidia-smi

Read loss (error on training data) together with eval_loss (error on held-out data):

Log patternWhat it meansWhat to do
loss and eval_loss both fallTraining normallyKeep going
loss keeps falling, eval_loss risesOverfitting β€” memorizing the training dataFewer epochs, more data, use the load_best_model_at_end checkpoint
Neither moves muchLearning rate too low or a data format problemRaise learning_rate, check the template from step 6
loss is NaNLearning rate too high or fp16 overflowLower learning_rate, use a GPU with bf16 support
CUDA out of memoryBatch or sequence too large for the VRAMBatch size 1-2 with more gradient accumulation, shorter max_length

Step 8 β€” Evaluate the result

A lower loss does not guarantee better answers. Compare answers from before (base model) and after training on questions that were not used for training. PEFT’s disable_adapter() lets you load the model once and produce both answers.

compare.pyPYTHON
import json
import torch
from datasets import load_dataset
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_DIR = "./models/llama-3.2-3b-instruct"
ADAPTER_DIR = "./outputs/llama-3.2-3b-incident-lora"

tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
base = AutoModelForCausalLM.from_pretrained(MODEL_DIR, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, ADAPTER_DIR)
model.eval()

def generate(messages: list[dict]) -> str:
    inputs = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_tensors="pt", return_dict=True,
    ).to(model.device)
    with torch.no_grad():
        out = model.generate(**inputs, max_new_tokens=300, do_sample=False)
    return tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()

REQUIRED = ["Likely causes", "Check", "Action"]      # the answer format we want

rows = []
for ex in load_dataset("json", data_files="data/eval.jsonl", split="train"):
    with model.disable_adapter():
        before = generate(ex["prompt"])
    after = generate(ex["prompt"])
    rows.append({
        "question": ex["prompt"][-1]["content"],
        "expected": ex["completion"][0]["content"],
        "before": before,
        "after": after,
        "format_ok_before": all(k in before for k in REQUIRED),
        "format_ok_after": all(k in after for k in REQUIRED),
    })

with open("eval_report.jsonl", "w", encoding="utf-8") as f:
    for r in rows:
        f.write(json.dumps(r, ensure_ascii=False) + "\n")

n = len(rows)
print(f"format compliance: before {sum(r['format_ok_before'] for r in rows)}/{n}, after {sum(r['format_ok_after'] for r in rows)}/{n}")
ℹ️

Alongside rule checks like format compliance, read eval_report.jsonl yourself and look for factual errors. With many evaluation questions you can add an LLM judge β€” ask a larger model (for example Llama 3.3 70B) to score each answer 1-5 against the expected one. See the Agent Evaluation guide for evaluation design.

Step 9 β€” Merge, convert to GGUF, import into Ollama

Training produces an adapter that sits on top of the base model. To use it in Ollama, merge the adapter into the base model, convert and quantize it to GGUF with llama.cpp, and import it into Ollama.

merge.pyPYTHON
import torch
from peft import AutoPeftModelForCausalLM
from transformers import AutoTokenizer

ADAPTER_DIR = "./outputs/llama-3.2-3b-incident-lora"
MERGED_DIR = "./outputs/llama-3.2-3b-incident-merged"

# Merge on top of the 16-bit original weights, not the 4-bit ones
model = AutoPeftModelForCausalLM.from_pretrained(ADAPTER_DIR, dtype=torch.bfloat16)
model.merge_and_unload().save_pretrained(MERGED_DIR)
AutoTokenizer.from_pretrained("./models/llama-3.2-3b-instruct").save_pretrained(MERGED_DIR)
BASH
uv run merge.py

# Get llama.cpp (conversion script + quantize tool)
git clone https://github.com/ggml-org/llama.cpp
uv pip install -r llama.cpp/requirements.txt
cmake -S llama.cpp -B llama.cpp/build && cmake --build llama.cpp/build --config Release -j

# 1) safetensors β†’ GGUF (16-bit)
uv run python llama.cpp/convert_hf_to_gguf.py ./outputs/llama-3.2-3b-incident-merged \
  --outfile ./outputs/incident-f16.gguf --outtype f16

# 2) Quantize to 4-bit (Ollama does not quantize GGUF files on import, so do it here)
./llama.cpp/build/bin/llama-quantize ./outputs/incident-f16.gguf ./outputs/incident-Q4_K_M.gguf Q4_K_M
ModelfileTEXT
FROM ./outputs/incident-Q4_K_M.gguf

PARAMETER temperature 0.2
PARAMETER num_ctx 8192

SYSTEM """Incident-response assistant"""
BASH
ollama create incident-llama -f Modelfile

# Check the chat template made it in β€” if TEMPLATE is empty, see the callout below
ollama show --modelfile incident-llama

ollama run incident-llama "API 5xx rate suddenly went above 10%"
⚠️

If answers never end or special tokens show up in the output, the chat template is missing. Run ollama show --modelfile llama3.2, copy the base Llama 3.2 model’s TEMPLATE and PARAMETER stop lines into your Modelfile, and run ollama create again.

Step 10 β€” Serve it as an API

Instead of calling the Ollama API from your app directly, put a thin backend in front of it so authentication, rate limits, logging, and prompt management live in one place. Here is a FastAPI backend that relays a streaming response.

BASH
uv add fastapi "uvicorn[standard]"
app.pyPYTHON
import time
import logging
from fastapi import FastAPI, Header, HTTPException
from fastapi.responses import StreamingResponse
from ollama import AsyncClient
from pydantic import BaseModel, Field

logging.basicConfig(level=logging.INFO)
log = logging.getLogger("incident-llama")

app = FastAPI(title="Incident Assistant")
ollama = AsyncClient(host="http://localhost:11434")
API_KEYS = {"team-sre-key"}               # in production, read these from env vars or a secret store

class ChatRequest(BaseModel):
    message: str = Field(min_length=1, max_length=4000)

@app.get("/health")
async def health() -> dict:
    await ollama.list()                    # checks the Ollama server too
    return {"status": "ok"}

@app.post("/chat")
async def chat(req: ChatRequest, x_api_key: str = Header()):
    if x_api_key not in API_KEYS:
        raise HTTPException(status_code=401, detail="invalid api key")

    async def stream():
        started = time.perf_counter()
        first_token_at = None
        async for part in await ollama.chat(
            model="incident-llama",
            messages=[{"role": "user", "content": req.message}],
            stream=True,
        ):
            if first_token_at is None:
                first_token_at = time.perf_counter()
            yield part.message.content
            if part.done:
                log.info(
                    "ttft=%.2fs total=%.2fs prompt_tokens=%s output_tokens=%s",
                    first_token_at - started, time.perf_counter() - started,
                    part.prompt_eval_count, part.eval_count,
                )

    return StreamingResponse(stream(), media_type="text/plain; charset=utf-8")
BASH
uv run uvicorn app:app --host 0.0.0.0 --port 8080

curl -N http://localhost:8080/chat \
  -H "Content-Type: application/json" -H "X-API-Key: team-sre-key" \
  -d '{"message": "A node disk usage alert fired"}'
πŸ’‘

When concurrent users grow, serve the same merged model (safetensors) with vLLM and only change what this backend calls. Export the logged time to first token (TTFT) and token counts as Prometheus metrics and track them on a dashboard.

Pre-production checklist

  • Have you met the Llama license terms ("Built with Llama" attribution, "Llama" at the start of a derivative model’s name)?
  • Did you remove personal data and secrets from the training data, and record its source and version?
  • Did you compare before and after training on an evaluation set and keep the report?
  • Are the model files (GGUF, adapter) and the Modelfile under version control?
  • Does the API have authentication, request size limits, and rate limits?
  • Do you monitor TTFT, total response time, token counts, and error rate?
  • Did you measure latency and error rate at the expected number of concurrent users with k6?

Frequently asked questions

How much data do I need to fine-tune Llama?

To match a tone or an output format, a few hundred high-quality examples already make a visible difference. To make the model follow a new workflow or domain conventions consistently, you often need thousands. Consistency and correctness matter more than volume, so start small and grow the dataset based on evaluation results.

Can I fine-tune Llama without a GPU?

You can run (infer) Llama on a CPU, but training on a CPU is not practical. QLoRA training effectively requires an NVIDIA GPU with CUDA: about 8-12 GB of VRAM for Llama 3.2 3B and around 16 GB for Llama 3.1 8B. Without a GPU, use a cloud GPU or a notebook environment such as Google Colab.

Should I fine-tune or use RAG first?

If the model needs to answer with facts that change often β€” documents, policies, prices β€” start with RAG. Fine-tuning is better at making answer format, tone, and procedures consistent than at adding knowledge. Most services start with RAG and add fine-tuning when format or behavior problems remain that prompting cannot fix.

Can I set up a Llama training environment on Windows?

Ollama runs on Windows directly. Training and vLLM serving are built around Linux, so WSL2 (Ubuntu) causes the fewest problems. In WSL2 the Windows NVIDIA driver is all you need for GPU access β€” do not install a separate Linux driver inside WSL.

How do I use my fine-tuned model in Ollama?

Merge the LoRA adapter into the base model, convert it to a GGUF file with llama.cpp’s conversion script, and quantize it to the size you want. Then point a Modelfile at the GGUF path with FROM and run ollama create; the model is available through ollama run and the API right away.