LLM API 응답 지연 최적화부터 추론 서버 튜닝, RAG 파이프라인 성능 개선, k6 부하 테스트까지. 프로덕션 AI 서비스의 병목을 진단하고 해결하는 실전 기술을 다룹니다.
AI 서비스의 성능은 일반 웹 API와 다른 지표로 측정합니다. 단순 응답 시간 외에 토큰 처리 속도와 첫 토큰까지의 지연이 사용자 경험을 결정합니다.
| 지표 | 설명 | 목표치 (프로덕션) |
|---|---|---|
| TTFT (Time To First Token) | 첫 번째 토큰이 도착하기까지의 시간 | < 500ms |
| TPS (Tokens Per Second) | 초당 생성 토큰 수 | > 30 tok/s |
| E2E Latency | 요청 시작~응답 완료까지 전체 시간 | p95 < 10s |
| Throughput | 단위 시간당 처리 요청 수 | 모델/GPU에 따라 상이 |
| Queue Depth | 대기 중인 추론 요청 수 | 0 유지 목표 |
TTFT를 우선 최적화하세요. 스트리밍 UI에서 사용자는 첫 토큰이 도착하는 순간부터 응답이 왔다고 인식합니다. E2E latency가 길어도 TTFT가 짧으면 체감 성능이 크게 개선됩니다.
LLM 호출의 지연은 크게 세 구간에서 발생합니다. 프롬프트 처리(prefill), 토큰 생성(decode), 네트워크 전송입니다. 각 구간을 독립적으로 최적화할 수 있습니다.
전체 응답을 기다리는 대신 생성되는 토큰을 즉시 전송합니다. TTFT가 체감 성능의 핵심이므로 스트리밍은 사실상 필수입니다.
import { NextRequest } from 'next/server';
export async function POST(req: NextRequest) {
const { messages } = await req.json();
const upstream = await fetch('https://api.openai.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
},
body: JSON.stringify({
model: 'gpt-4o-mini',
messages,
stream: true, // 스트리밍 활성화
max_tokens: 1024,
}),
});
// 업스트림 SSE를 클라이언트에 그대로 파이프
return new Response(upstream.body, {
headers: {
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
Connection: 'keep-alive',
},
});
}export async function streamChat(messages: Message[], onChunk: (text: string) => void) {
const res = await fetch('/api/chat', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ messages }),
});
const reader = res.body!.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const lines = decoder.decode(value).split('\n');
for (const line of lines) {
if (!line.startsWith('data: ') || line === 'data: [DONE]') continue;
const delta = JSON.parse(line.slice(6))
?.choices?.[0]?.delta?.content;
if (delta) onChunk(delta);
}
}
}동일한 시스템 프롬프트나 컨텍스트가 반복되는 경우, KV cache를 재사용하면 prefill 비용과 TTFT를 크게 줄일 수 있습니다. Anthropic, OpenAI 모두 프롬프트 캐싱을 지원합니다.
import anthropic
client = anthropic.Anthropic()
# 긴 시스템 프롬프트에 cache_control 적용
# 첫 호출 이후 cache hit 시 TTFT 최대 85% 감소
response = client.messages.create(
model="claude-opus-4-8",
max_tokens=1024,
system=[
{
"type": "text",
"text": LONG_SYSTEM_PROMPT, # 수천 토큰의 컨텍스트
"cache_control": {"type": "ephemeral"} # 캐시 마킹
}
],
messages=[{"role": "user", "content": user_query}],
)
# 캐시 적중 확인
usage = response.usage
print(f"Cache read tokens: {usage.cache_read_input_tokens}")
print(f"Cache creation tokens: {usage.cache_creation_input_tokens}")캐시 TTL: Anthropic 프롬프트 캐시는 5분 유지됩니다. RAG 시스템에서 동일 문서를 반복 참조하거나, 멀티턴 대화에서 누적 컨텍스트가 길어질 때 효과가 큽니다.
실시간 응답이 불필요한 작업(분류, 요약, 임베딩)은 배치로 묶어 처리하면 GPU 활용률을 높이고 단위 비용을 낮출 수 있습니다.
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI()
async def classify_batch(texts: list[str]) -> list[str]:
"""여러 텍스트를 동시 분류 — 개별 호출 대비 3~5배 처리량"""
tasks = [
client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "텍스트를 긍정/부정/중립으로 분류하세요."},
{"role": "user", "content": text},
],
max_tokens=10,
)
for text in texts
]
responses = await asyncio.gather(*tasks)
return [r.choices[0].message.content for r in responses]
# 실행
results = asyncio.run(classify_batch(["좋아요", "별로예요", "평범해요"]))자체 호스팅 LLM(Llama, Mistral, Gemma 등)을 서빙할 때는 추론 엔진 선택이 성능에 직접적인 영향을 미칩니다.
| 엔진 | 특징 | 적합한 상황 |
|---|---|---|
| vLLM | PagedAttention, 연속 배칭, OpenAI API 호환 | 프로덕션 고처리량 |
| Ollama | GGUF 지원, 설치 간단, CPU/GPU 혼용 | 로컬 개발, 소규모 |
| TGI (HuggingFace) | Flash Attention, 텐서 병렬, 다양한 모델 | HuggingFace 생태계 |
| llama.cpp | GGUF 양자화, 최소 자원 | 엣지, 저사양 환경 |
FP16 모델을 INT4/INT8로 양자화하면 VRAM 사용량을 최대 75% 줄이고 추론 속도를 높일 수 있습니다. 품질 손실은 INT8 기준 1~3% 수준입니다.
# AWQ 양자화 (INT4, 품질 손실 최소화)
pip install autoawq
python -c "
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model_path = 'meta-llama/Llama-3.1-8B-Instruct'
quant_path = 'llama-3.1-8b-awq-int4'
model = AutoAWQForCausalLM.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path)
quant_config = {'zero_point': True, 'q_group_size': 128, 'w_bit': 4, 'version': 'GEMM'}
model.quantize(tokenizer, quant_config=quant_config)
model.save_quantized(quant_path)
"| 방식 | VRAM (8B 모델) | 속도 | 품질 |
|---|---|---|---|
| FP16 | 16 GB | 기준 | 100% |
| INT8 (GPTQ) | 8 GB | +20% | ~99% |
| INT4 (AWQ) | 4.5 GB | +60% | ~97% |
| GGUF Q4_K_M | 4.8 GB | CPU 가능 | ~96% |
services:
vllm:
image: vllm/vllm-openai:latest
runtime: nvidia
environment:
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN}
command: >
--model meta-llama/Llama-3.1-8B-Instruct
--quantization awq # AWQ INT4 양자화
--max-model-len 8192 # 컨텍스트 길이
--max-num-seqs 256 # 최대 동시 시퀀스 (연속 배칭)
--gpu-memory-utilization 0.90 # GPU 메모리 90% 활용
--enable-chunked-prefill # 긴 프롬프트 청크 처리
ports:
- "8000:8000"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]--max-num-seqs를 높이면 처리량이 증가하지만 개별 요청의 TTFT가 늘어납니다. 실시간 챗봇은 낮게(32~64), 배치 처리는 높게(256~512) 설정하세요.
RAG(Retrieval-Augmented Generation)의 성능 병목은 대부분 검색(retrieval) 단계에 있습니다. LLM 응답 시간보다 벡터 검색이 더 느린 경우가 많습니다.
import asyncio
from functools import lru_cache
from qdrant_client import AsyncQdrantClient
from openai import AsyncOpenAI
openai_client = AsyncOpenAI()
qdrant = AsyncQdrantClient(url="http://localhost:6333")
# 임베딩 캐싱 — 동일 쿼리 재호출 방지
@lru_cache(maxsize=1024)
def cache_key(query: str) -> str:
return query.lower().strip()
embedding_cache: dict[str, list[float]] = {}
async def embed_with_cache(text: str) -> list[float]:
key = cache_key(text)
if key not in embedding_cache:
res = await openai_client.embeddings.create(
model="text-embedding-3-small",
input=text,
)
embedding_cache[key] = res.data[0].embedding
return embedding_cache[key]
async def retrieve(query: str, top_k: int = 5) -> list[dict]:
# 임베딩 + 검색 병렬화
query_vec = await embed_with_cache(query)
results = await qdrant.search(
collection_name="docs",
query_vector=query_vec,
limit=top_k,
with_payload=True,
# HNSW 파라미터: ef 높을수록 정확하지만 느림 (기본 128)
search_params={"hnsw_ef": 64},
)
return [r.payload for r in results]| 최적화 | 효과 | 트레이드오프 |
|---|---|---|
| 임베딩 캐싱 | 반복 쿼리 ~0ms | 메모리 사용량 증가 |
| HNSW ef 감소 | 검색 속도 2~3배 | 재현율 소폭 감소 |
| 청크 크기 최적화 | 검색 정확도 향상 | 청크 수 증가 |
| 하이브리드 검색 | 키워드+의미 혼합 | BM25 인덱스 추가 필요 |
AI 서비스는 일반 API와 다르게 응답 시간이 길고 스트리밍 응답을 처리해야 합니다. k6로 이 특성을 반영한 현실적인 부하 시나리오를 구성합니다.
import http from 'k6/http';
import { check, sleep } from 'k6';
import { Trend, Counter } from 'k6/metrics';
// AI 전용 커스텀 메트릭
const ttft = new Trend('ai_time_to_first_token', true);
const totalTok = new Counter('ai_total_tokens');
export const options = {
stages: [
{ duration: '1m', target: 10 }, // 워밍업
{ duration: '3m', target: 50 }, // 정상 부하
{ duration: '2m', target: 100 }, // 최대 부하
{ duration: '1m', target: 0 }, // 쿨다운
],
thresholds: {
'http_req_duration': ['p(95)<15000'], // AI는 15s 이내
'ai_time_to_first_token': ['p(95)<1000'], // TTFT p95 < 1s
'http_req_failed': ['rate<0.01'],
},
};
const PROMPTS = [
"Redis 캐싱 전략을 3가지 설명해주세요.",
"API 응답 시간이 느릴 때 점검할 항목은?",
"Kubernetes HPA 설정 방법을 알려주세요.",
];
export default function () {
const prompt = PROMPTS[Math.floor(Math.random() * PROMPTS.length)];
const start = Date.now();
const res = http.post(
'http://localhost:8000/v1/chat/completions',
JSON.stringify({
model: 'meta-llama/Llama-3.1-8B-Instruct',
messages: [{ role: 'user', content: prompt }],
max_tokens: 256,
stream: false,
}),
{ headers: { 'Content-Type': 'application/json' } },
);
check(res, {
'status 200': (r) => r.status === 200,
'has content': (r) => r.json()?.choices?.[0]?.message?.content?.length > 0,
});
const body = res.json();
if (body?.usage) {
totalTok.add(body.usage.completion_tokens);
// non-streaming: TTFT ≈ 전체 응답 시간 (보수적 측정)
ttft.add(Date.now() - start);
}
sleep(Math.random() * 2 + 1); // 1~3초 think time
}스트리밍 테스트: k6는 SSE 스트림을 기본 지원하지 않습니다. 스트리밍 TTFT를 정확히 측정하려면 stream: false로 전체 응답을 받거나, 별도의 WebSocket 테스트를 구성하세요.
# 테스트 실행
k6 run ai-load-test.js
# 결과 예시
# ai_time_to_first_token p(95)=842ms ✓
# http_req_duration p(95)=6.2s ✓
# ai_total_tokens 12,840 tokens
# http_req_failed 0.00% ✓AI 서비스는 일반 서비스보다 더 세밀한 모니터링이 필요합니다. 토큰 사용량, GPU 활용률, 큐 대기 시간이 핵심 관측 지표입니다.
# vLLM Prometheus 메트릭 수집 설정
scrape_configs:
- job_name: 'vllm'
static_configs:
- targets: ['vllm:8000']
metrics_path: '/metrics'
# 핵심 PromQL 알림 규칙
groups:
- name: ai-service
rules:
- alert: HighTTFT
expr: histogram_quantile(0.95, vllm:time_to_first_token_seconds_bucket) > 1
for: 5m
annotations:
summary: "TTFT p95가 1초를 초과했습니다"
- alert: QueueBacklog
expr: vllm:num_requests_waiting > 20
for: 2m
annotations:
summary: "추론 큐 대기 요청이 20개를 초과했습니다"
- alert: LowGPUUtilization
expr: vllm:gpu_cache_usage_perc < 0.3
for: 10m
annotations:
summary: "GPU 캐시 활용률이 30% 미만입니다 — 스케일 다운 고려"| 메트릭 | 설명 | 임계값 (예시) |
|---|---|---|
| vllm:time_to_first_token_seconds | TTFT 히스토그램 | p95 < 1s |
| vllm:num_requests_running | 현재 처리 중인 요청 수 | max-num-seqs 이하 |
| vllm:num_requests_waiting | 큐 대기 요청 수 | < 20 |
| vllm:gpu_cache_usage_perc | KV 캐시 GPU 사용률 | > 0.7 |
| vllm:generation_tokens_total | 총 생성 토큰 수 | 비용 추적용 |
비용 대비 성능 최적화: GPU 캐시 사용률이 지속적으로 90% 이상이면 컨텍스트 길이를 줄이거나 GPU를 추가하세요. 30% 이하면 과잉 프로비저닝 상태로 스케일 다운을 고려합니다.