Start by running Llama on your own machine, then train it on your teamβs data, import it into Ollama, and serve it behind an API β in ten steps. Every step comes with commands to verify it worked and the places people usually get stuck.
This tutorial follows one running example β an incident-response assistant for an engineering team β from installing Llama to training it on your data and serving it behind an API. Steps 1-4 work without a GPU; the training steps (5-9) need an NVIDIA GPU.
| Step | What you do | Result | GPU |
|---|---|---|---|
| 1-2 | Check hardware, set up Python and PyTorch | A working dev environment | Optional |
| 3-4 | Install Ollama, first chat, generation parameters | A local Llama API on port 11434 | Optional |
| 5 | Download the original model from Hugging Face, understand the chat template | Original weights for training | Recommended |
| 6-7 | Prepare training data, fine-tune with QLoRA | A LoRA adapter | Required |
| 8 | Compare answers before and after training | An evaluation report | Required |
| 9-10 | Merge, convert to GGUF, import into Ollama, serve with FastAPI | Your own model API | Optional |
The hands-on model is Llama 3.2 3B Instruct, which can be trained on a laptop GPU. Change only the model ID and the same code scales to Llama 3.1 8B. For model differences and licensing, read the Llama guide first.
GPU memory (VRAM) decides which model sizes you can run and train. Check your machine first.
# NVIDIA GPU name, driver version, VRAM usage
nvidia-smi
# macOS (Apple Silicon) β unified memory size
system_profiler SPHardwareDataType | grep -E "Chip|Memory"
# Linux β CPU and memory
lscpu | grep "Model name"; free -h| Your machine | Run (inference) | Train (QLoRA) |
|---|---|---|
| No GPU, 16 GB RAM | Llama 3.2 3B, 3.1 8B (slow) | Use a cloud GPU |
| Apple Silicon 16 GB+ | Llama 3.1 8B runs smoothly | This tutorialβs 4-bit training (bitsandbytes) targets NVIDIA CUDA β use a cloud GPU |
| NVIDIA 8-12 GB | Llama 3.1 8B (Q4) | Llama 3.2 1B / 3B |
| NVIDIA 16-24 GB | Llama 3.1 8B (Q8/FP16) | Llama 3.1 8B |
| NVIDIA 48 GB+ | Llama 3.3 70B (Q4) | Llama 3.3 70B |
On Windows, run Ollama natively and do training and serving in WSL2 (Ubuntu). WSL2 uses the Windows NVIDIA driver for GPU access, so do not install a Linux driver inside WSL. If nvidia-smi works inside WSL, you are ready.
Use a per-project virtual environment to avoid version conflicts. This tutorial uses the fast package manager uv.
# Install uv (macOS/Linux/WSL)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Create the project
uv init llama-lab --python 3.12
cd llama-labPyTorch has to match your GPU and CUDA version. Pick your OS, package manager, and CUDA version in the PyTorch install selector and use the command it gives you. Then confirm the GPU is visible:
import torch
print("PyTorch:", torch.__version__)
print("CUDA available:", torch.cuda.is_available())
if torch.cuda.is_available():
props = torch.cuda.get_device_properties(0)
print("GPU:", props.name)
print(f"VRAM: {props.total_memory / 1024**3:.1f} GB")
print("BF16 supported:", torch.cuda.is_bf16_supported()) # if False, train in fp16uv run check_gpu.py
# Packages used in later steps
uv add "transformers>=4.56" "trl>=1.0" peft datasets accelerate bitsandbytes ollama "huggingface_hub[cli]"If you see CUDA available: False, the CPU build of PyTorch is installed. Reinstall the CUDA build before training β otherwise training runs on the CPU without any error, just extremely slowly.
Ollama handles model downloads, GPU placement, and the API server in one tool. Install it for your OS and confirm the server is up.
# Linux / WSL2
curl -fsSL https://ollama.com/install.sh | sh
# macOS β download the app from ollama.com, or
brew install ollama
# Windows β download and run the installer (OllamaSetup.exe) from ollama.com
# Verify
ollama --version
curl http://localhost:11434/api/version # the server responds# Pull the models used here
ollama pull llama3.2 # Llama 3.2 3B (about 2 GB)
ollama pull llama3.1 # Llama 3.1 8B (about 4.9 GB)
ollama listServer settings you will change often are environment variables:
| Variable | Default | Use |
|---|---|---|
OLLAMA_MODELS | macOS ~/.ollama/models, Linux /usr/share/ollama/.ollama/models | Move models to a larger disk |
OLLAMA_HOST | 127.0.0.1:11434 | Allow other machines to connect (0.0.0.0) |
OLLAMA_CONTEXT_LENGTH | 4K-256K depending on VRAM | Default context length |
OLLAMA_KEEP_ALIVE | 5 minutes | How long a model stays loaded after a request |
OLLAMA_NUM_PARALLEL | 1 | Concurrent requests per model (memory grows with it) |
# Linux (systemd) β add variables to the service
sudo systemctl edit ollama.service
# [Service]
# Environment="OLLAMA_CONTEXT_LENGTH=16384"
# Environment="OLLAMA_KEEP_ALIVE=30m"
sudo systemctl daemon-reload && sudo systemctl restart ollama
# macOS β then restart the Ollama app
launchctl setenv OLLAMA_CONTEXT_LENGTH 16384The Ollama API has no authentication, so with OLLAMA_HOST=0.0.0.0 anyone on the network can use your models. If you expose it, put authentication and IP restrictions in a reverse proxy such as Nginx.
Check that it works in the terminal, then call the REST API your application will actually use.
ollama run llama3.2
>>> What should I check, in order, when a Kubernetes pod is in CrashLoopBackOff?
>>> /set parameter temperature 0.2 # change a parameter mid-chat
>>> /byecurl http://localhost:11434/api/chat -d '{
"model": "llama3.2",
"messages": [
{"role": "system", "content": "You are an SRE. Answer step by step and keep it short."},
{"role": "user", "content": "Five metrics to check when API latency suddenly jumps"}
],
"stream": false,
"options": {"temperature": 0.2, "num_ctx": 8192, "num_predict": 512}
}'| Parameter | Meaning | Suggested value |
|---|---|---|
temperature | Randomness. Lower means the same question gets the same answer | 0-0.3 for facts and classification, ~0.7 for writing |
top_p | Sample only from the top cumulative-probability candidates | 0.9 (tune this or temperature, not both) |
num_ctx | Context length for input plus output | 8K-32K, sized to your conversation or documents |
num_predict | Maximum tokens to generate | Caps answer length (cost and latency) |
seed | Random seed | Fix it for evaluation and tests to make results reproducible |
from ollama import chat
messages = [{"role": "system", "content": "You are an SRE. Answer step by step and keep it short."}]
while True:
user = input("you> ").strip()
if user in {"exit", "quit"}:
break
messages.append({"role": "user", "content": user})
answer = ""
for part in chat(model="llama3.2", messages=messages, stream=True, options={"temperature": 0.2}):
print(part.message.content, end="", flush=True)
answer += part.message.content
print()
messages.append({"role": "assistant", "content": answer}) # keep the conversation contextThe model does not remember the conversation. You have to keep sending previous questions and answers in messages, and they take up context. Summarize or drop old messages in long conversations.
Ollama models are quantized files for inference and cannot be used for training. Training uses the original weights (safetensors) from Hugging Face. Llama repositories are gated and require accepting the license.
meta-llama/Llama-3.2-3B-Instruct, accept the license, and wait for approval.uv run hf auth login # paste the token
uv run hf download meta-llama/Llama-3.2-3B-Instruct --local-dir ./models/llama-3.2-3b-instructLlama was trained on conversations wrapped in special tokens β the chat template. Training data and inference input must follow it, so never build the string by hand; use the tokenizerβs apply_chat_template.
import torch
from transformers import AutoTokenizer, pipeline
MODEL_DIR = "./models/llama-3.2-3b-instruct"
messages = [
{"role": "system", "content": "You are an SRE."},
{"role": "user", "content": "A disk usage alert just fired"},
]
# ββ See the exact string the model receives ββββββ
tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
print(tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True))
# ββ Run inference with Transformers ββββββββββββββ
pipe = pipeline("text-generation", model=MODEL_DIR, dtype=torch.bfloat16, device_map="auto")
out = pipe(messages, max_new_tokens=256, do_sample=False)
print(out[0]["generated_text"][-1]["content"])<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: ...
You are an SRE.<|eot_id|><|start_header_id|>user<|end_header_id|>
A disk usage alert just fired<|eot_id|><|start_header_id|>assistant<|end_header_id|>
(sample output β the date line depends on when you run it)Data quality decides almost everything about a fine-tune. You are showing the model example answers β "for this kind of question, answer like this" β so write the answers exactly in the format and tone you want, and make them correct.
| Principle | What it means |
|---|---|
| Consistent format | Answer the same kind of question with the same structure (e.g. causes β checks β actions) |
| Correctness | The model learns wrong answers too. Have an expert review the data |
| Variety | Include several phrasings of the same intent |
| Remove personal data | Mask names, emails, tokens, and internal IPs before training |
| Hold out evaluation data | Keep about 10% out of training for evaluation |
Store the data as JSONL in the conversational prompt-completion format TRL reads directly. With this format the loss is computed only on the completion, so the model does not spend capacity memorizing the questions.
{"prompt": [{"role": "system", "content": "Incident-response assistant"}, {"role": "user", "content": "API 5xx rate suddenly went above 10%"}], "completion": [{"role": "assistant", "content": "Likely causes: a recent deploy, an upstream outage, DB connection exhaustion\nCheck: 1) deploy history 2) upstream latency dashboard 3) DB connection pool usage\nAction: if it lines up with the last deploy, consider rolling back first."}]}
{"prompt": [{"role": "system", "content": "Incident-response assistant"}, {"role": "user", "content": "A node disk usage alert fired"}], "completion": [{"role": "assistant", "content": "Likely causes: container logs piling up, image cache, large temp files\nCheck: 1) df -h 2) du on the largest directories 3) log rotation settings\nAction: prune old images and shorten log retention."}]}from datasets import load_dataset
from transformers import AutoTokenizer
MODEL_DIR = "./models/llama-3.2-3b-instruct"
MAX_LENGTH = 1024
tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
ds = load_dataset("json", data_files="data/incidents.jsonl", split="train")
# ββ 1) Drop duplicate questions ββββββββββββββββββ
seen = set()
def first_time(example) -> bool:
key = example["prompt"][-1]["content"].strip()
if key in seen:
return False
seen.add(key)
return True
ds = ds.filter(first_time)
# ββ 2) Check token length β beyond MAX_LENGTH the end (the answer) is cut off β
def count_tokens(example) -> dict:
enc = tokenizer.apply_chat_template(example["prompt"] + example["completion"], tokenize=True, return_dict=True)
return {"n_tokens": len(enc["input_ids"])}
ds = ds.map(count_tokens)
too_long = ds.filter(lambda e: e["n_tokens"] > MAX_LENGTH)
print(f"{len(ds)} examples, max {max(ds['n_tokens'])} tokens, {len(too_long)} over {MAX_LENGTH}")
ds = ds.filter(lambda e: e["n_tokens"] <= MAX_LENGTH).remove_columns("n_tokens")
# ββ 3) Train / eval split ββββββββββββββββββββββββ
split = ds.train_test_split(test_size=0.1, seed=42)
split["train"].to_json("data/train.jsonl", force_ascii=False)
split["test"].to_json("data/eval.jsonl", force_ascii=False)
print("train:", len(split["train"]), "eval:", len(split["test"]))Start with 200-500 examples and run one full train-and-evaluate loop. Strengthening the weak question types based on the results is faster than producing a large dataset up front.
QLoRA loads the original model in 4-bit and trains only small LoRA matrices added to each layer. Only about 1% of the parameters are trained, so Llama 3.2 3B fits on an 8-12 GB GPU.
import torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer
MODEL_DIR = "./models/llama-3.2-3b-instruct" # scale up: meta-llama/Llama-3.1-8B-Instruct
OUTPUT_DIR = "./outputs/llama-3.2-3b-incident-lora"
data = load_dataset("json", data_files={"train": "data/train.jsonl", "eval": "data/eval.jsonl"})
bf16 = torch.cuda.is_bf16_supported()
trainer = SFTTrainer(
model=MODEL_DIR,
train_dataset=data["train"],
eval_dataset=data["eval"],
quantization_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16 if bf16 else torch.float16,
bnb_4bit_use_double_quant=True,
),
peft_config=LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules="all-linear", task_type="CAUSAL_LM",
),
args=SFTConfig(
output_dir=OUTPUT_DIR,
num_train_epochs=3,
per_device_train_batch_size=4,
gradient_accumulation_steps=4, # effective batch = 4 Γ 4 = 16
learning_rate=1e-4,
lr_scheduler_type="cosine",
warmup_steps=10,
max_length=1024,
bf16=bf16,
fp16=not bf16,
logging_steps=10,
eval_strategy="steps",
eval_steps=50,
save_strategy="steps",
save_steps=50,
save_total_limit=2,
load_best_model_at_end=True, # keep the checkpoint with the lowest eval_loss
metric_for_best_model="eval_loss",
model_init_kwargs={"dtype": torch.bfloat16 if bf16 else torch.float16},
),
)
trainer.train()
trainer.save_model() # saves the LoRA adapter (tens of MB) to OUTPUT_DIR
print("saved to:", OUTPUT_DIR)uv run train.py
# Watch GPU usage from another terminal
watch -n 2 nvidia-smiRead loss (error on training data) together with eval_loss (error on held-out data):
| Log pattern | What it means | What to do |
|---|---|---|
| loss and eval_loss both fall | Training normally | Keep going |
| loss keeps falling, eval_loss rises | Overfitting β memorizing the training data | Fewer epochs, more data, use the load_best_model_at_end checkpoint |
| Neither moves much | Learning rate too low or a data format problem | Raise learning_rate, check the template from step 6 |
| loss is NaN | Learning rate too high or fp16 overflow | Lower learning_rate, use a GPU with bf16 support |
| CUDA out of memory | Batch or sequence too large for the VRAM | Batch size 1-2 with more gradient accumulation, shorter max_length |
A lower loss does not guarantee better answers. Compare answers from before (base model) and after training on questions that were not used for training. PEFTβs disable_adapter() lets you load the model once and produce both answers.
import json
import torch
from datasets import load_dataset
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_DIR = "./models/llama-3.2-3b-instruct"
ADAPTER_DIR = "./outputs/llama-3.2-3b-incident-lora"
tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
base = AutoModelForCausalLM.from_pretrained(MODEL_DIR, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, ADAPTER_DIR)
model.eval()
def generate(messages: list[dict]) -> str:
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt", return_dict=True,
).to(model.device)
with torch.no_grad():
out = model.generate(**inputs, max_new_tokens=300, do_sample=False)
return tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()
REQUIRED = ["Likely causes", "Check", "Action"] # the answer format we want
rows = []
for ex in load_dataset("json", data_files="data/eval.jsonl", split="train"):
with model.disable_adapter():
before = generate(ex["prompt"])
after = generate(ex["prompt"])
rows.append({
"question": ex["prompt"][-1]["content"],
"expected": ex["completion"][0]["content"],
"before": before,
"after": after,
"format_ok_before": all(k in before for k in REQUIRED),
"format_ok_after": all(k in after for k in REQUIRED),
})
with open("eval_report.jsonl", "w", encoding="utf-8") as f:
for r in rows:
f.write(json.dumps(r, ensure_ascii=False) + "\n")
n = len(rows)
print(f"format compliance: before {sum(r['format_ok_before'] for r in rows)}/{n}, after {sum(r['format_ok_after'] for r in rows)}/{n}")Alongside rule checks like format compliance, read eval_report.jsonl yourself and look for factual errors. With many evaluation questions you can add an LLM judge β ask a larger model (for example Llama 3.3 70B) to score each answer 1-5 against the expected one. See the Agent Evaluation guide for evaluation design.
Training produces an adapter that sits on top of the base model. To use it in Ollama, merge the adapter into the base model, convert and quantize it to GGUF with llama.cpp, and import it into Ollama.
import torch
from peft import AutoPeftModelForCausalLM
from transformers import AutoTokenizer
ADAPTER_DIR = "./outputs/llama-3.2-3b-incident-lora"
MERGED_DIR = "./outputs/llama-3.2-3b-incident-merged"
# Merge on top of the 16-bit original weights, not the 4-bit ones
model = AutoPeftModelForCausalLM.from_pretrained(ADAPTER_DIR, dtype=torch.bfloat16)
model.merge_and_unload().save_pretrained(MERGED_DIR)
AutoTokenizer.from_pretrained("./models/llama-3.2-3b-instruct").save_pretrained(MERGED_DIR)uv run merge.py
# Get llama.cpp (conversion script + quantize tool)
git clone https://github.com/ggml-org/llama.cpp
uv pip install -r llama.cpp/requirements.txt
cmake -S llama.cpp -B llama.cpp/build && cmake --build llama.cpp/build --config Release -j
# 1) safetensors β GGUF (16-bit)
uv run python llama.cpp/convert_hf_to_gguf.py ./outputs/llama-3.2-3b-incident-merged \
--outfile ./outputs/incident-f16.gguf --outtype f16
# 2) Quantize to 4-bit (Ollama does not quantize GGUF files on import, so do it here)
./llama.cpp/build/bin/llama-quantize ./outputs/incident-f16.gguf ./outputs/incident-Q4_K_M.gguf Q4_K_MFROM ./outputs/incident-Q4_K_M.gguf
PARAMETER temperature 0.2
PARAMETER num_ctx 8192
SYSTEM """Incident-response assistant"""ollama create incident-llama -f Modelfile
# Check the chat template made it in β if TEMPLATE is empty, see the callout below
ollama show --modelfile incident-llama
ollama run incident-llama "API 5xx rate suddenly went above 10%"If answers never end or special tokens show up in the output, the chat template is missing. Run ollama show --modelfile llama3.2, copy the base Llama 3.2 modelβs TEMPLATE and PARAMETER stop lines into your Modelfile, and run ollama create again.
Instead of calling the Ollama API from your app directly, put a thin backend in front of it so authentication, rate limits, logging, and prompt management live in one place. Here is a FastAPI backend that relays a streaming response.
uv add fastapi "uvicorn[standard]"import time
import logging
from fastapi import FastAPI, Header, HTTPException
from fastapi.responses import StreamingResponse
from ollama import AsyncClient
from pydantic import BaseModel, Field
logging.basicConfig(level=logging.INFO)
log = logging.getLogger("incident-llama")
app = FastAPI(title="Incident Assistant")
ollama = AsyncClient(host="http://localhost:11434")
API_KEYS = {"team-sre-key"} # in production, read these from env vars or a secret store
class ChatRequest(BaseModel):
message: str = Field(min_length=1, max_length=4000)
@app.get("/health")
async def health() -> dict:
await ollama.list() # checks the Ollama server too
return {"status": "ok"}
@app.post("/chat")
async def chat(req: ChatRequest, x_api_key: str = Header()):
if x_api_key not in API_KEYS:
raise HTTPException(status_code=401, detail="invalid api key")
async def stream():
started = time.perf_counter()
first_token_at = None
async for part in await ollama.chat(
model="incident-llama",
messages=[{"role": "user", "content": req.message}],
stream=True,
):
if first_token_at is None:
first_token_at = time.perf_counter()
yield part.message.content
if part.done:
log.info(
"ttft=%.2fs total=%.2fs prompt_tokens=%s output_tokens=%s",
first_token_at - started, time.perf_counter() - started,
part.prompt_eval_count, part.eval_count,
)
return StreamingResponse(stream(), media_type="text/plain; charset=utf-8")uv run uvicorn app:app --host 0.0.0.0 --port 8080
curl -N http://localhost:8080/chat \
-H "Content-Type: application/json" -H "X-API-Key: team-sre-key" \
-d '{"message": "A node disk usage alert fired"}'When concurrent users grow, serve the same merged model (safetensors) with vLLM and only change what this backend calls. Export the logged time to first token (TTFT) and token counts as Prometheus metrics and track them on a dashboard.
To match a tone or an output format, a few hundred high-quality examples already make a visible difference. To make the model follow a new workflow or domain conventions consistently, you often need thousands. Consistency and correctness matter more than volume, so start small and grow the dataset based on evaluation results.
You can run (infer) Llama on a CPU, but training on a CPU is not practical. QLoRA training effectively requires an NVIDIA GPU with CUDA: about 8-12 GB of VRAM for Llama 3.2 3B and around 16 GB for Llama 3.1 8B. Without a GPU, use a cloud GPU or a notebook environment such as Google Colab.
If the model needs to answer with facts that change often β documents, policies, prices β start with RAG. Fine-tuning is better at making answer format, tone, and procedures consistent than at adding knowledge. Most services start with RAG and add fine-tuning when format or behavior problems remain that prompting cannot fix.
Ollama runs on Windows directly. Training and vLLM serving are built around Linux, so WSL2 (Ubuntu) causes the fewest problems. In WSL2 the Windows NVIDIA driver is all you need for GPU access β do not install a separate Linux driver inside WSL.
Merge the LoRA adapter into the base model, convert it to a GGUF file with llama.cppβs conversion script, and quantize it to the size you want. Then point a Modelfile at the GGUF path with FROM and run ollama create; the model is available through ollama run and the API right away.