본문으로 건너뛰기
AIDevOps
  • Learn
  • Learning Paths
  • Practice
  • Open Source
  • Books
  • Engineering

    AI DevOpsAI 서비스 개발·운영 전체 지도LLMOpsLLM 배포·평가·관측실전 프로젝트AI Agent 프로젝트 실습

    Knowledge

    Docs기술 문서 모음Blog엔지니어링 아티클Plogger개발 기록 피드

    Validate

    Certification3단계 역량 인증 · 준비 중
AI Models
LlamaMistralGemmaDeepSeekQwen
🧠 AI Core
AI 입문 & 로드맵ML FundamentalsLLM Fundamentals|Python AIC++|PyTorchTensorFlowJAX
🤖 AI 실전 개발
AI 실전 입문 & 로드맵Hugging FaceLangChainLlamaIndexLLMOps|LangGraphMCPMulti-AgentAgent Evaluation
🧠 AI Agent 개발
금융 AI AgentLLM API 서버주식 투자 AgentAIOps AI Agent교육 AI Agent코딩 AI Agent
🌱 Spring Cloud
Spring 입문 & 로드맵Spring Cloud GatewaySpring BootJava|Spring AISpring SecuritySpring BatchSpring JPA
🐳 DevOps
DevOps 입문 & 로드맵LinuxDockerCI/CD|Kubernetes 기본K8s 심화/실무PrometheusGrafana
🧱 인프라
인프라 입문 & 로드맵NginxRedis
☁️ 클라우드
클라우드 입문 & 로드맵AWSGCPAzureNCPCloudflare
🎨 Frontend
Frontend 입문 & 로드맵JavaScriptTypeScript|ReactNext.js|VueNuxt
📱 Mobile
Mobile 입문 & 로드맵KotlinAndroidFlutter
⚙️ Backend
Backend 입문 & 로드맵Python 기본FastAPIDjangoFlask|CGoGinNode.js
💾 Database
DB 입문 & 로드맵공통 SQLOracleMySQLPostgreSQL|MongoDB벡터 DB
🧪 검증
k6JMeternGrinder
AIDevOps

Engineering AI. From Code to Production.
AI와 AI Agent를 개발하고 운영하기 위한 엔지니어링 학습 플랫폼

Learn

  • 전체 가이드
  • Learning Paths
  • Practice
  • Books

Resources

  • AI DevOps
  • LLMOps
  • 실전 프로젝트
  • Docs
  • Blog
  • Plogger
  • Open Source
  • Certification (준비 중)

Start Here

  • AI Core 로드맵
  • AI 실전 개발 로드맵
  • Spring Cloud 로드맵
  • DevOps 로드맵
  • 인프라 로드맵

 

  • 클라우드 로드맵
  • Frontend 로드맵
  • Mobile 로드맵
  • Backend 로드맵
  • Database 로드맵
© 2026 AI DevOps Korea. All rights reserved.
이용약관개인정보처리방침Sitemaptestforge.kr
AI Agent 추론 모델 가이드

🔍 DeepSeek 완전 가이드

Visitors

DeepSeek R1의 Chain-of-Thought 추론은 복잡한 계획 수립이 필요한 AI Agent 개발에 최적입니다. MIT 라이선스, R1 추론 파싱, Function Calling, 추론-실행 분리 아키텍처, vLLM 프로덕션 배포까지 다룹니다.

← AI DevOps 홈으로
🔍

DeepSeek 실전 예제

R1 CoT 파싱, Function Calling Agent, 추론-실행 분리 아키텍처, vLLM 배포 예제 모음.

실전 예제 보기 →

목차

0 / 8
  1. 모델 라인업 & 선택 기준
  2. R1 Chain-of-Thought 파싱
  3. Ollama 로컬 실행
  4. DeepSeek API & Function Calling
  5. 추론-실행 분리 Agent
  6. CoT 스트리밍
  7. vLLM 프로덕션 서버
  8. 벤치마크 & 모델 선택
목차 8개 섹션
  1. 모델 라인업 & 선택 기준
  2. R1 Chain-of-Thought 파싱
  3. Ollama 로컬 실행
  4. DeepSeek API & Function Calling
  5. 추론-실행 분리 Agent
  6. CoT 스트리밍
  7. vLLM 프로덕션 서버
  8. 벤치마크 & 모델 선택

모델 라인업 & 선택 기준

DeepSeek는 추론 특화(R1)와 범용(V3) 두 계열로 나뉩니다. Agent 설계 시 계획·추론에는 R1, 도구 호출·실행에는 V3를 조합하면 효율적입니다.

모델파라미터특징VRAM추천 용도
R1-Distill-Qwen-7B7B경량 추론, CoT8GB로컬 추론 Agent
R1-Distill-Llama-8B8BLlama 기반 Distill8GBLangChain 통합
R1-Distill-Qwen-32B32B고성능 추론24GB복잡한 계획 수립
R1-Distill-Llama-70B70B최고 추론 품질48GBo1-mini 수준
DeepSeek-V3671B MoEFunction CallingAPI도구 호출 실행
DeepSeek-R1671B MoE최강 추론APIo1 수준 추론
💡

로컬 시작점: R1-Distill-Qwen-7B. 8GB VRAM으로도 수학·코딩 추론에서 GPT-4 수준의 성능을 보입니다.

R1 Chain-of-Thought 파싱

R1은 최종 답변 전에 <think>...</think> 블록에서 단계별 추론 과정을 생성합니다. 이 추론 과정을 파싱하면 모델의 판단 근거를 분석하거나 추론 품질을 평가할 수 있습니다.

r1_parser.pyPYTHON
import re
from dataclasses import dataclass
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_DEEPSEEK_KEY",
    base_url="https://api.deepseek.com",
)

@dataclass
class R1Response:
    thinking:   str          # <think> 내부 추론 과정
    answer:     str          # 최종 답변
    think_len:  int          # 추론 토큰 수 (복잡도 지표)

def call_r1(prompt: str, system: str = "") -> R1Response:
    messages = []
    if system:
        messages.append({"role": "system", "content": system})
    messages.append({"role": "user", "content": prompt})

    resp = client.chat.completions.create(
        model="deepseek-reasoner",
        messages=messages,
    )
    msg = resp.choices[0].message

    # DeepSeek API는 reasoning_content를 별도 필드로 제공
    thinking = getattr(msg, "reasoning_content", "") or ""

    # Ollama 등 로컬 모델은 <think> 태그로 반환
    if not thinking:
        match = re.search(r"<think>(.*?)</think>", msg.content, re.DOTALL)
        thinking = match.group(1).strip() if match else ""

    answer = re.sub(r"<think>.*?</think>", "", msg.content, flags=re.DOTALL).strip()
    return R1Response(thinking=thinking, answer=answer, think_len=len(thinking.split()))

# ── 사용 예시 ─────────────────────────────────────
result = call_r1(
    prompt="피보나치 수열의 40번째 항을 동적 프로그래밍으로 계산하는 Python 코드를 작성해줘",
    system="당신은 알고리즘 전문가입니다. 시간·공간 복잡도를 반드시 분석하세요.",
)
print(f"추론 길이: {result.think_len}개 단어")
print(f"추론 과정 일부: {result.thinking[:300]}...")
print(f"\n최종 답변:\n{result.answer}")

Ollama 로컬 실행

BASH
# R1 Distill 시리즈 (로컬 추천)
ollama run deepseek-r1:7b    # 7B  — 4.7GB,  8GB VRAM
ollama run deepseek-r1:14b   # 14B — 9.0GB, 16GB VRAM
ollama run deepseek-r1:32b   # 32B — 20GB,  24GB VRAM
ollama run deepseek-r1:70b   # 70B — 43GB,  48GB VRAM

# 실행 확인
ollama list
ollama ps
local_r1.pyPYTHON
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

def ask_r1(question: str) -> tuple[str, str]:
    import re
    resp = client.chat.completions.create(
        model="deepseek-r1:7b",
        messages=[{"role": "user", "content": question}],
        temperature=0.6,   # R1 추천 온도
    )
    content = resp.choices[0].message.content
    think_m = re.search(r"<think>(.*?)</think>", content, re.DOTALL)
    thinking = think_m.group(1).strip() if think_m else ""
    answer   = re.sub(r"<think>.*?</think>", "", content, flags=re.DOTALL).strip()
    return thinking, answer

thinking, answer = ask_r1("하노이 탑 문제를 재귀와 반복 두 가지 방법으로 풀어줘")
print(f"[추론] {thinking[:200]}...")
print(f"[답변] {answer}")

DeepSeek API & Function Calling

DeepSeek-V3는 OpenAI와 동일한 Function Calling 인터페이스를 지원합니다. API 비용은 GPT-4o의 약 1/30 수준입니다.

BASH
uv add openai  # DeepSeek은 OpenAI SDK 그대로 사용
deepseek_function.pyPYTHON
import json
from openai import OpenAI

# V3 — Function Calling (코드 실행, 도구 호출)
client = OpenAI(api_key="YOUR_KEY", base_url="https://api.deepseek.com")

TOOLS = [
    {"type": "function", "function": {
        "name": "run_python",
        "description": "Python 코드를 실행하고 결과를 반환합니다",
        "parameters": {"type": "object",
            "properties": {
                "code": {"type": "string", "description": "실행할 Python 코드"},
                "timeout": {"type": "integer", "default": 10},
            }, "required": ["code"]},
    }},
    {"type": "function", "function": {
        "name": "search_arxiv",
        "description": "arXiv에서 논문을 검색합니다",
        "parameters": {"type": "object",
            "properties": {
                "query": {"type": "string"},
                "max_results": {"type": "integer", "default": 5},
            }, "required": ["query"]},
    }},
]

def run_python(code: str, timeout: int = 10) -> str:
    # 실제 구현 시 subprocess + 샌드박스 사용
    return f"실행 완료:\n{code}\n→ [결과]"

def search_arxiv(query: str, max_results: int = 5) -> str:
    return f"arXiv 검색 '{query}': {max_results}개 논문 발견..."

TOOL_FNS = {"run_python": run_python, "search_arxiv": search_arxiv}

def chat_with_tools(user_msg: str) -> str:
    messages = [{"role": "user", "content": user_msg}]
    for _ in range(5):
        resp = client.chat.completions.create(
            model="deepseek-chat",   # V3 — Function Calling
            messages=messages, tools=TOOLS, tool_choice="auto",
        )
        msg = resp.choices[0].message
        messages.append(msg)
        if not msg.tool_calls:
            return msg.content
        for tc in msg.tool_calls:
            args   = json.loads(tc.function.arguments)
            result = TOOL_FNS[tc.function.name](**args)
            messages.append({"role": "tool", "tool_call_id": tc.id, "content": result})
    return "스텝 초과"

print(chat_with_tools("피보나치 수열 20번째 항을 Python으로 계산하고 실행 결과를 보여줘"))

추론-실행 분리 Agent

R1(추론)과 V3(실행)를 조합하는 아키텍처입니다. R1이 복잡한 계획을 수립하고, V3가 도구를 실행합니다. 각 모델의 강점을 최대한 활용하는 프로덕션 패턴입니다.

planner_executor.pyPYTHON
import json, re
from openai import OpenAI
from pydantic import BaseModel

# R1: 계획 수립 (CoT 추론)
r1_client = OpenAI(api_key="YOUR_KEY", base_url="https://api.deepseek.com")

# V3: 도구 실행 (Function Calling)
v3_client = OpenAI(api_key="YOUR_KEY", base_url="https://api.deepseek.com")

class Plan(BaseModel):
    steps:      list[str]   # 단계별 실행 계획
    tools_needed: list[str] # 필요한 도구 목록
    reasoning:  str         # 계획 수립 근거

def make_plan(user_goal: str) -> Plan:
    """R1으로 실행 계획 수립"""
    resp = r1_client.chat.completions.create(
        model="deepseek-reasoner",
        messages=[{
            "role": "user",
            "content": f"""다음 목표를 달성하기 위한 단계별 계획을 JSON으로 작성해줘.
목표: {user_goal}

반드시 JSON 형식으로:
{{"steps": [...], "tools_needed": [...], "reasoning": "..."}}"""
        }],
    )
    content = resp.choices[0].message.content
    # <think> 태그 제거 후 JSON 추출
    clean = re.sub(r"<think>.*?</think>", "", content, flags=re.DOTALL).strip()
    json_m = re.search(r"\{.*\}", clean, re.DOTALL)
    return Plan(**json.loads(json_m.group())) if json_m else Plan(
        steps=[user_goal], tools_needed=[], reasoning="단순 작업"
    )

def execute_step(step: str, context: str) -> str:
    """V3로 단계 실행 (도구 호출 가능)"""
    resp = v3_client.chat.completions.create(
        model="deepseek-chat",
        messages=[
            {"role": "system", "content": f"이전 컨텍스트: {context}"},
            {"role": "user",   "content": f"다음 단계를 실행해줘: {step}"},
        ],
    )
    return resp.choices[0].message.content

def run_planner_executor(goal: str) -> str:
    # 1. R1으로 계획 수립
    plan = make_plan(goal)
    print(f"계획 수립 완료: {len(plan.steps)}단계")
    print(f"근거: {plan.reasoning[:200]}")

    # 2. V3로 단계별 실행
    context, results = "", []
    for i, step in enumerate(plan.steps, 1):
        print(f"[{i}/{len(plan.steps)}] {step}")
        result = execute_step(step, context)
        results.append(result)
        context = f"단계 {i} 완료: {result[:200]}"

    return "\n".join(results)

output = run_planner_executor("Python으로 REST API 서버를 만들고 Docker로 배포하는 전체 과정을 안내해줘")

CoT 스트리밍

R1의 추론 과정을 스트리밍으로 받으면 사용자에게 "생각하는 중..." 피드백을 실시간으로 제공할 수 있습니다.

r1_stream.pyPYTHON
from openai import OpenAI
import sys

client = OpenAI(api_key="YOUR_KEY", base_url="https://api.deepseek.com")

def stream_r1(prompt: str):
    """추론 과정과 최종 답변을 실시간 스트리밍"""
    stream = client.chat.completions.create(
        model="deepseek-reasoner",
        messages=[{"role": "user", "content": prompt}],
        stream=True,
    )

    in_think = False
    for chunk in stream:
        delta = chunk.choices[0].delta

        # reasoning_content: 추론 과정 (R1 전용 필드)
        if hasattr(delta, "reasoning_content") and delta.reasoning_content:
            if not in_think:
                print("\n💭 [추론 중...]", flush=True)
                in_think = True
            print(delta.reasoning_content, end="", flush=True)

        # content: 최종 답변
        if delta.content:
            if in_think:
                print("\n\n✅ [최종 답변]", flush=True)
                in_think = False
            print(delta.content, end="", flush=True)

    print()  # 줄바꿈

stream_r1("피타고라스 정리를 증명하고 Python으로 직각삼각형 판별기를 만들어줘")

vLLM 프로덕션 서버

DeepSeek Distill 모델을 vLLM으로 서빙하면 Ollama 대비 최대 10배 높은 처리량을 얻을 수 있습니다. 특히 배치 추론이 많은 Agent 환경에 적합합니다.

BASH
pip install vllm

# R1-Distill-7B 서버 실행
python -m vllm.entrypoints.openai.api_server \
  --model deepseek-ai/DeepSeek-R1-Distill-Qwen-7B \
  --host 0.0.0.0 --port 8000 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90 \
  --enable-chunked-prefill          # 긴 추론 처리량 최적화

# 32B 멀티 GPU
python -m vllm.entrypoints.openai.api_server \
  --model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B \
  --tensor-parallel-size 2 \
  --port 8000
vllm_r1.pyPYTHON
import asyncio, re
from openai import AsyncOpenAI

client = AsyncOpenAI(base_url="http://localhost:8000/v1", api_key="none")

async def batch_reason(problems: list[str]) -> list[dict]:
    """여러 추론 문제를 동시에 처리"""
    async def solve(problem: str) -> dict:
        resp = await client.chat.completions.create(
            model="deepseek-ai/DeepSeek-R1-Distill-Qwen-7B",
            messages=[{"role": "user", "content": problem}],
            max_tokens=4096,
            temperature=0.6,
        )
        content = resp.choices[0].message.content
        think_m = re.search(r"<think>(.*?)</think>", content, re.DOTALL)
        return {
            "problem": problem,
            "thinking": think_m.group(1)[:200] if think_m else "",
            "answer":   re.sub(r"<think>.*?</think>", "", content, flags=re.DOTALL).strip(),
        }

    return await asyncio.gather(*[solve(p) for p in problems])

problems = [
    "피보나치 40번째 항을 메모이제이션으로 계산하는 Python 코드",
    "이진 탐색 트리의 삽입과 검색을 구현해줘",
    "동적 프로그래밍으로 최장 공통 부분 수열(LCS)을 구현해줘",
]
results = asyncio.run(batch_reason(problems))
for r in results:
    print(f"Q: {r['problem'][:50]}...")
    print(f"A: {r['answer'][:200]}\n")

벤치마크 & 모델 선택

벤치마크R1-7BR1-70BDeepSeek-R1GPT-4oo1
MATH-50083.0%94.5%97.3%74.6%96.4%
AIME 202413.3%60.0%79.8%9.3%74.4%
HumanEval72.0%87.2%92.0%90.2%92.4%
MMLU63.8%86.0%90.8%88.7%92.0%
ℹ️

Agent 용도별 추천: 수학·코딩 추론 → R1-Distill. 도구 호출·일반 작업 → DeepSeek-V3. 복잡한 다단계 계획 → R1 Full(API). 비용 절감 목적 → R1-7B 로컬.