Even in a basement plant room or a factory with no internet, one laptop can analyze equipment photos, write inspection reports and create maintenance work orders. In ten steps you add Gemma 4 E4Bβs image and audio input, thinking mode and function calling one at a time.
The inspector takes an equipment photo and optionally leaves a voice memo; the assistant writes a report and, if maintenance is needed, files a work order. Everything runs on the laptop in the field without the internet.
Equipment photo (+ voice memo)
β
ββ inspection.py photo β inspection report (JSON) β Gemma 4 image input Β· structured output
ββ review.py re-check maintenance-level findings β thinking mode (think=True)
ββ agent.py asset lookup Β· work order creation β function calling
β ββ assets.py local SQLite asset register
ββ voice_memo.py voice memo β text β Gemma 4 audio input (Transformers)
ββ app.py LAN API for field tablets β FastAPI
batch.py process a day's photos (resumable)
evaluate.py compare with inspector ratings, count under/over-ratedModel sizes, licensing and Ollama/Transformers basics are in the Gemma guide.
| Field hardware | Model | Ollama size (4-bit) |
|---|---|---|
| 16 GB laptop or mini PC | gemma4:e4b | from about 6.6 GB |
| 8 GB-class device | gemma4:e2b | from about 4.6 GB |
| Apple Silicon MacBook | gemma4:e4b-mlx | about 9.5 GB |
# Download in advance at the office (with internet)
ollama pull gemma4:e4b
uv init field-inspector --python 3.12
cd field-inspector
uv add ollama pydantic fastapi "uvicorn[standard]" python-multipartKeep the model name and generation options in one place so a hardware change is a one-line edit.
import os
# Pick by field hardware β e4b for laptops and mini PCs, e2b for tablet-class devices
MODEL = os.environ.get("GEMMA_MODEL", "gemma4:e4b")
# Inspection reports need consistent judgments, so use a low temperature (the model card default is 1.0)
REPORT_OPTIONS = {"temperature": 0.2, "num_ctx": 8192}Start with a free-form question to see how the model reads your photos. In the CLI, include the image path in the prompt.
ollama run gemma4:e4b "What looks wrong in this equipment photo? ./photos/pump-07.jpg"If gauge readings matter, a close, straight-on photo without glare has the biggest effect on accuracy. Give field inspectors a photo guide.
Free-text reports are hard to aggregate or connect to work orders. Define the report with Pydantic and pass its schema to formatto always get the same JSON structure. Keep the ratings conservative: "when unsure, go one level higher".
import sys
from typing import Literal
from ollama import chat
from pydantic import BaseModel, Field
from config import MODEL, REPORT_OPTIONS
class Finding(BaseModel):
component: str = Field(description="Component name (e.g. pipe joint, pressure gauge)")
issue: str = Field(description="The observed problem in one sentence")
severity: Literal["ok", "watch", "action", "stop"]
class InspectionReport(BaseModel):
asset_id: str | None = Field(default=None, description="The asset tag in the photo if visible, otherwise null")
readings: list[str] = Field(description="Values read from gauges and displays")
findings: list[Finding]
overall: Literal["ok", "watch", "action", "stop"]
SYSTEM = """You are an equipment inspection assistant. Base the report only on what is visible in the photo.
severity: ok (no issue) / watch (recheck next inspection) / action (maintenance request needed) / stop (consider shutting down now)
Leaks, smoke, sparks or heavy deformation are action or higher; when unsure, go one level higher.
overall is the highest severity among the findings."""
def inspect_photo(path: str, note: str = "") -> InspectionReport:
resp = chat(
model=MODEL,
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"Inspection photo. {note}".strip(), "images": [path]},
],
format=InspectionReport.model_json_schema(),
options=REPORT_OPTIONS,
)
return InspectionReport.model_validate_json(resp.message.content)
if __name__ == "__main__":
report = inspect_photo(sys.argv[1] if len(sys.argv) > 1 else "photos/pump-07.jpg")
print(report.model_dump_json(indent=2))uv run inspection.py photos/pump-07.jpg{
"asset_id": "PUMP-07",
"readings": ["pressure 7.8 bar"],
"findings": [{"component": "pipe joint", "issue": "water is leaking from the joint", "severity": "action"}],
"overall": "action"
}Thinking mode on every photo would slow processing a lot. Re-check only action and stop findings β where a wrong call is costly β with think=True, and have the model list the immediate steps for the field worker.
from ollama import chat
from config import MODEL
from inspection import InspectionReport
SERIOUS = {"action", "stop"}
def review(report: InspectionReport, photo: str) -> str | None:
"""Re-check only maintenance-level findings in thinking mode β spend time only where a wrong call is costly"""
if report.overall not in SERIOUS:
return None
resp = chat(
model=MODEL,
messages=[{
"role": "user",
"content": (
"Check whether this inspection finding matches the photo, and list three steps the field worker should take right now.\n"
f"Finding: {report.model_dump_json()}"
),
"images": [photo],
}],
think=True,
)
print("[thinking]", (resp.message.thinking or "")[:200], "...")
return resp.message.contentThe reasoning arrives in message.thinking and the final answer in message.content. Show workers only the answer; keep the reasoning in logs for checking how a finding was reached.
Keep the asset register and work orders in SQLite on the field laptop and sync them back at the office, so no network is needed on site.
import sqlite3
from datetime import date
DB = sqlite3.connect("plant.db", check_same_thread=False)
DB.executescript("""
CREATE TABLE IF NOT EXISTS equipment (asset_id TEXT PRIMARY KEY, type TEXT, location TEXT, last_inspection TEXT);
CREATE TABLE IF NOT EXISTS work_orders (id INTEGER PRIMARY KEY AUTOINCREMENT, asset_id TEXT, severity TEXT, summary TEXT, created TEXT);
INSERT OR IGNORE INTO equipment VALUES
('PUMP-07', 'centrifugal pump', 'Building B, level B2', '2026-08-14'),
('VALVE-12', 'gate valve', 'Building A, rooftop', '2026-09-02');
""")
def lookup_equipment(asset_id: str) -> dict:
"""Look up equipment type, location and last inspection date by asset ID.
Args:
asset_id: Asset ID (e.g. PUMP-07)
"""
row = DB.execute("SELECT type, location, last_inspection FROM equipment WHERE asset_id = ?", (asset_id.upper(),)).fetchone()
if not row:
return {"error": f"Equipment {asset_id} not found"}
return {"asset_id": asset_id.upper(), "type": row[0], "location": row[1], "last_inspection": row[2]}
def create_work_order(asset_id: str, severity: str, summary: str) -> dict:
"""Create a maintenance work order. Call it only when severity is action or stop.
Args:
asset_id: Asset ID
severity: action or stop
summary: One sentence on why maintenance is needed
"""
if severity not in {"action", "stop"}:
return {"error": "Work orders can only be created for action or stop findings"}
cur = DB.execute(
"INSERT INTO work_orders (asset_id, severity, summary, created) VALUES (?, ?, ?, ?)",
(asset_id.upper(), severity, summary, date.today().isoformat()),
)
DB.commit()
return {"work_order_id": cur.lastrowid, "asset_id": asset_id.upper(), "severity": severity}
TOOLS = {"lookup_equipment": lookup_equipment, "create_work_order": create_work_order}Pass functions directly in tools and the Ollama Python library builds the schema from type hints and the docstringβs Args. So writing clearly in the docstring when a tool should be called is the prompt.
from ollama import chat
from assets import TOOLS, create_work_order, lookup_equipment
from config import MODEL
from inspection import InspectionReport
SYSTEM = """You are an equipment inspection assistant.
1) If the report has an asset ID, check its location and last inspection date with lookup_equipment.
2) If overall is action or stop, create a work order with create_work_order. Do not create one for ok or watch.
3) Finish with a three-line summary for the field worker."""
def process(report: InspectionReport, max_steps: int = 4) -> str:
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"Inspection report: {report.model_dump_json()}"},
]
for _ in range(max_steps):
resp = chat(model=MODEL, messages=messages, tools=[lookup_equipment, create_work_order])
messages.append(resp.message)
if not resp.message.tool_calls:
return resp.message.content
for call in resp.message.tool_calls:
fn = TOOLS.get(call.function.name)
result = fn(**call.function.arguments) if fn else {"error": f"unknown tool: {call.function.name}"}
print(f"[tool] {call.function.name}({call.function.arguments}) -> {result}")
messages.append({"role": "tool", "tool_name": call.function.name, "content": str(result)})
return "Automatic processing did not finish. Please review the report."[tool] lookup_equipment({'asset_id': 'PUMP-07'}) -> {'asset_id': 'PUMP-07', 'type': 'centrifugal pump', 'location': 'Building B, level B2', ...}
[tool] create_work_order({'asset_id': 'PUMP-07', 'severity': 'action', 'summary': 'water is leaking from the joint'}) -> {'work_order_id': 1, ...}
(sample output)As create_work_order re-checks severity, enforce rules inside tools so a wrong model call cannot cause harm. Instructions in the prompt alone are not enough.
Wearing gloves in the field, voice memos beat typing. Gemma 4 E2B, E4B and 12B accept audio directly, so no separate speech recognition model is needed. Audio is processed with Hugging Face Transformers, up to 30 seconds per input.
uv add -U transformers torch accelerate
# Download at the office in advance (after accepting the license)
uv run hf download google/gemma-4-E2B-itimport sys
from transformers import AutoModelForMultimodalLM, AutoProcessor
# Audio input is supported only by E2B, E4B and 12B. Up to 30 seconds per input.
MODEL_ID = "google/gemma-4-E2B-it"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(MODEL_ID, dtype="auto", device_map="auto")
def transcribe(wav_path: str) -> str:
messages = [{
"role": "user",
"content": [
{"type": "text", "text": "Transcribe this field voice memo in its original language. Output only the transcription, and write numbers as digits."},
{"type": "audio", "audio": wav_path},
],
}]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
enable_thinking=False,
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
outputs = model.generate(**inputs, max_new_tokens=256)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
parsed = processor.parse_response(response, prefix=inputs["input_ids"])
return parsed["content"] if isinstance(parsed, dict) else str(parsed)
if __name__ == "__main__":
print(transcribe(sys.argv[1] if len(sys.argv) > 1 else "memos/pump-07.wav"))Pass the transcript as the note of step 3βs inspect_photo to get a report that considers both the photo and the memo. Split memos longer than 30 seconds by sentence.
After the inspection round, process the photo folder in one go. If the laptop battery dies or the run stops, rerunning continues with the unprocessed photos only.
import json
import sys
import time
from pathlib import Path
from inspection import inspect_photo
photo_dir = Path(sys.argv[1] if len(sys.argv) > 1 else "photos")
out_path = Path("data/reports.jsonl")
out_path.parent.mkdir(exist_ok=True)
# Skip photos already processed β if the battery dies or the run stops, rerunning picks up where it left off
done = set()
if out_path.exists():
done = {json.loads(line)["photo"] for line in out_path.read_text(encoding="utf-8").splitlines() if line}
photos = sorted(p for p in photo_dir.iterdir() if p.suffix.lower() in {".jpg", ".jpeg", ".png"})
with out_path.open("a", encoding="utf-8") as out:
for photo in photos:
if photo.name in done:
continue
started = time.perf_counter()
try:
report = inspect_photo(str(photo))
row = {"photo": photo.name, "ok": True, "report": report.model_dump()}
except Exception as e: # one failed photo does not stop the rest
row = {"photo": photo.name, "ok": False, "error": str(e)}
row["seconds"] = round(time.perf_counter() - started, 2)
out.write(json.dumps(row) + "\n")
out.flush()
print(f"{photo.name}: {row.get('report', {}).get('overall', 'ERROR')} ({row['seconds']}s)")gauge-03.jpg: watch (4.8s)
pump-07.jpg: action (5.3s)
valve-12.jpg: ok (4.6s)
(sample output β processing time depends on hardware)Use ratings from experienced inspectors on the same photos as ground truth. The number of under-rated findings β risk rated too low β matters far more than the match rate. If any appear, make the system promptβs criteria more conservative and evaluate again.
photo,severity
pump-07.jpg,action
valve-12.jpg,ok
gauge-03.jpg,watchimport csv
import json
from collections import Counter
LEVEL = {"ok": 0, "watch": 1, "action": 2, "stop": 3}
labels = {row["photo"]: row["severity"] for row in csv.DictReader(open("data/labels.csv", encoding="utf-8"))}
reports = [json.loads(line) for line in open("data/reports.jsonl", encoding="utf-8") if line.strip()]
result = Counter()
for r in reports:
expected = labels.get(r["photo"])
if expected is None or not r["ok"]:
result["skipped"] += 1
continue
got = r["report"]["overall"]
if got == expected:
result["match"] += 1
elif LEVEL[got] < LEVEL[expected]:
result["under"] += 1 # risk rated too low β the most dangerous error
print(f" β under-rated {r['photo']}: inspector {expected} / model {got}")
else:
result["over"] += 1 # rated too high β the cost of an unnecessary dispatch
judged = result["match"] + result["under"] + result["over"]
print(f"match {result['match']}/{judged}, under-rated {result['under']}, over-rated {result['over']}, skipped {result['skipped']}")β under-rated gauge-11.jpg: inspector action / model watch
match 41/48, under-rated 1, over-rated 6, skipped 2
(sample output)Inspectors upload photos from a tablet and the laptop processes them. Put both on the same private Wi-Fi (or tethering) and open the API on the LAN only.
import shutil
import tempfile
from pathlib import Path
from fastapi import FastAPI, File, Form, HTTPException, UploadFile
from agent import process
from inspection import inspect_photo
from review import review
app = FastAPI(title="Field Inspection Assistant")
MAX_BYTES = 10 * 1024 * 1024
@app.post("/inspections")
def create_inspection(photo: UploadFile = File(...), note: str = Form("")) -> dict:
if photo.content_type not in {"image/jpeg", "image/png"}:
raise HTTPException(status_code=415, detail="Only JPEG or PNG photos are accepted")
with tempfile.TemporaryDirectory() as tmp:
path = Path(tmp) / (photo.filename or "photo.jpg")
with path.open("wb") as f:
shutil.copyfileobj(photo.file, f)
if path.stat().st_size > MAX_BYTES:
raise HTTPException(status_code=413, detail="Photos must be 10 MB or smaller")
report = inspect_photo(str(path), note)
return {
"report": report.model_dump(),
"review": review(report, str(path)),
"summary": process(report),
}# Keep the model in memory to avoid first-request latency
OLLAMA_KEEP_ALIVE=-1 ollama serve &
curl http://localhost:11434/api/generate -d '{"model": "gemma4:e4b", "keep_alive": -1}' # preload with an empty request
uv run uvicorn app:app --host 0.0.0.0 --port 8080This API has no authentication, so never open it on a network connected to the outside. Before heading out, disconnect the laptop and run the whole flow from photo upload to work order to confirm it works offline.
stop (shutdown) findings?Once the model files are downloaded, Ollama and Gemma 4 run entirely locally and need no internet connection. The Transformers model used for voice transcription also works offline if you download it in advance. Before going to the field, disconnect the network and run the whole flow once.
No. The assistant helps inspectors record findings; decisions such as shutting equipment down must be confirmed by a person. Under-rating risk can lead to accidents, so check the under-rated share in the evaluation step and keep the criteria conservative.
On a laptop or mini PC, E4B reads photos better. On tablet-class devices with 8 GB of memory or less, use E2B and check quality in the evaluation step. Both models accept images and audio.
The model can only judge what it sees. Have the system prompt rate hard-to-read photos at watch or higher, and have the app ask for a retake. For gauges that must be read, give inspectors a photo guide: close up, straight on, no glare.