본문으로 건너뛰기
AIDevOps
  • Learn
  • Learning Paths
  • Practice
  • Open Source
  • Books
  • Engineering

    AI DevOpsAI 서비스 개발·운영 전체 지도LLMOpsLLM 배포·평가·관측실전 프로젝트AI Agent 프로젝트 실습

    Knowledge

    Docs기술 문서 모음Blog엔지니어링 아티클Plogger개발 기록 피드

    Validate

    Certification3단계 역량 인증 · 준비 중
AI Models
LlamaMistralGemmaDeepSeekQwen
🧠 AI Core
AI 입문 & 로드맵ML FundamentalsLLM Fundamentals|Python AIC++|PyTorchTensorFlowJAX
🤖 AI 실전 개발
AI 실전 입문 & 로드맵Hugging FaceLangChainLlamaIndexLLMOps|LangGraphMCPMulti-AgentAgent Evaluation
🧠 AI Agent 개발
금융 AI AgentLLM API 서버주식 투자 AgentAIOps AI Agent교육 AI Agent코딩 AI Agent
🌱 Spring Cloud
Spring 입문 & 로드맵Spring Cloud GatewaySpring BootJava|Spring AISpring SecuritySpring BatchSpring JPA
🐳 DevOps
DevOps 입문 & 로드맵LinuxDockerCI/CD|Kubernetes 기본K8s 심화/실무PrometheusGrafana
🧱 인프라
인프라 입문 & 로드맵NginxRedis
☁️ 클라우드
클라우드 입문 & 로드맵AWSGCPAzureNCPCloudflare
🎨 Frontend
Frontend 입문 & 로드맵JavaScriptTypeScript|ReactNext.js|VueNuxt
📱 Mobile
Mobile 입문 & 로드맵KotlinAndroidFlutter
⚙️ Backend
Backend 입문 & 로드맵Python 기본FastAPIDjangoFlask|CGoGinNode.js
💾 Database
DB 입문 & 로드맵공통 SQLOracleMySQLPostgreSQL|MongoDB벡터 DB
🧪 검증
k6JMeternGrinder
AIDevOps

Engineering AI. From Code to Production.
AI와 AI Agent를 개발하고 운영하기 위한 엔지니어링 학습 플랫폼

Learn

  • 전체 가이드
  • Learning Paths
  • Practice
  • Books

Resources

  • AI DevOps
  • LLMOps
  • 실전 프로젝트
  • Docs
  • Blog
  • Plogger
  • Open Source
  • Certification (준비 중)

Start Here

  • AI Core 로드맵
  • AI 실전 개발 로드맵
  • Spring Cloud 로드맵
  • DevOps 로드맵
  • 인프라 로드맵

 

  • 클라우드 로드맵
  • Frontend 로드맵
  • Mobile 로드맵
  • Backend 로드맵
  • Database 로드맵
© 2026 AI DevOps Korea. All rights reserved.
이용약관개인정보처리방침Sitemaptestforge.kr
AIDevOps Cloud Native Series · Book 01

K8s Kubernetes Production Engineering by John Bae

Visitors

Designing, Deploying, Scaling, Securing, Troubleshooting, and Operating Kubernetes in the Cloud

← Back to all books
Kubernetes Production Engineering cover

목차

0 / 5
  1. Overview
  2. Chapters (10)
  3. Companion Materials
  4. Publication Status
  5. Next in Series
목차 5개 섹션
  1. Overview
  2. Chapters (10)
  3. Companion Materials
  4. Publication Status
  5. Next in Series

Overview

About this book

A practitioner-focused book on designing, deploying, scaling, securing, troubleshooting, and operating Kubernetes in production. It follows a fictional commerce system, CloudShop, from a single containerized service to a full enterprise production platform, connecting failure-aware scheduling, traffic control, data protection, security, observability, troubleshooting, GitOps, and managed-cloud integration into one continuous story.

Target version

Kubernetes 1.35+ (currently maintained releases: 1.35 / 1.36 / 1.37)

Length

58,917 words (10 chapters)

Status

Content complete — final technical review and proofreading in progress before publication

Fictional system — CloudShop

A fictional e-commerce platform that evolves throughout the book. It starts as a single containerized deployment and grows into a seven-namespace production system covering networking, storage, security, observability, troubleshooting, delivery, and enterprise governance.

About the author — John Bae

An enterprise cloud architect with more than 20 years of experience in software development, application architecture, and technical architecture. Over the past five years he has designed, built, and operated Kubernetes-based platforms in real enterprise environments. He is the founder and operator of AIDevOps.kr.

Download companion files (.zip)

Table of Contents

10 chapters

Chapters 1-3 lay the foundation — containers, workloads, and cluster design. Chapters 4-10 connect networking, storage, security, observability, troubleshooting, delivery, and enterprise governance through one continuous system, CloudShop.

Chapter 13,869w

From Containers to Production Kubernetes

What containers solved and what they did not, Kubernetes’ declarative control-loop model, and CloudShop’s first container architecture.

Chapter 23,826w

Kubernetes Architecture and Workloads

The request path from API server to kubelet, the Deployment/ReplicaSet/Pod relationship, and a first workload contract with probes and graceful shutdown.

Chapter 34,604w

Designing the Production Cluster

Cluster design that starts from requirements, failure domains, resource requests/limits, and the HPA/node-autoscaling chain.

Chapter 46,502w

Production Networking and Traffic Management

Gateway API as an organizational interface, per-team HTTPRoute ownership, and default-deny NetworkPolicy with explicit allows.

Chapter 58,197w

Storage, State, and Data Protection

One chain from data classification through StatefulSet, the StorageClass promise, the backup system, to a verified restore.

Chapter 65,987w

Kubernetes Security for Production

RBAC, Pod Security Standards, and NetworkPolicy built around one question: what can an attacker reach in the first sixty seconds?

Chapter 75,495w

Observability and Reliability Engineering

Metrics, logs, and traces; SLIs, SLOs, and error budgets; alerting that pages only on actionable conditions.

Chapter 87,169w

Troubleshooting Production Kubernetes

Hands-on diagnosis using CloudShop’s real failure-scenario manifests: Pending, ImagePullBackOff, CrashLoopBackOff, OOMKilled, and more.

Chapter 95,995w

Delivery: From Commit to Cluster

Helm packaging, Git as desired state, the Argo CD pattern, and rollback as a designed capability rather than an afterthought.

Chapter 107,273w

Managed Kubernetes and Enterprise Production Architecture

Comparing EKS/AKS/GKE operating models, encoding governance as automation, and CloudShop’s final architecture and production readiness review.

Companion Materials

Companion manifests & lab files

Every companion/... path referenced throughout the book maps to the same path inside the zip downloadable from this page. Every companion path referenced throughout the book (companion/manifests/..., companion/labs/..., companion/charts/...) maps to the same path inside this zip. This page is the official distribution point for the companion materials.

Chapter 1-3

manifests/chapter01/product-service-deployment.yaml

CloudShop’s first Deployment — a minimal workload with no probes, security, or resource settings yet.

Chapter 2

manifests/chapter02/cloudshop-workload-contract.yaml

A complete workload contract with three-stage probes and graceful shutdown.

Chapter 3

manifests/chapter03/product-service-production-controls.yaml

Production controls: PriorityClass, topology spread, PodDisruptionBudget, and HPA.

Chapter 4

manifests/chapter04/cloudshop-network.yaml

Gateway, three team-owned HTTPRoutes, and default-deny NetworkPolicy with explicit allows.

Chapter 5

manifests/chapter05/cloudshop-storage-recovery.yaml

Redis StatefulSet, volumeClaimTemplates, and a backup-verification CronJob.

Chapter 6

manifests/chapter06/cloudshop-security.yaml

RBAC, a restricted Pod Security Standard, and a NetworkPolicy that allows only DNS egress.

Chapter 7

manifests/chapter07/cloudshop-observability.yaml

OTel/OTLP instrumentation, Prometheus scraping, and an SLO policy ConfigMap.

Chapter 8

manifests/chapter08/troubleshooting-scenarios.yaml

Five deliberately broken scenarios: Pending, ImagePullBackOff, CrashLoopBackOff, OOMKilled, and a probe-path mismatch.

Chapter 9

charts/cloudshop-service/

A Helm chart with a restricted-PSS-compliant securityContext and readiness/liveness probes.

Chapter 9

manifests/chapter09/gitops-application.yaml

An Argo CD Application — selfHeal true, prune false.

Chapter 10

manifests/chapter10/enterprise-baseline.yaml

ResourceQuota, LimitRange, restricted Pod Security Admission, and a production-readiness checklist ConfigMap.

All

labs/

Ten chapter-by-chapter lab guides.

Publication Status

Publication status

This book is in pre-publication draft. It will be formally published once every gate below is complete. The EPUB available for download today reflects current progress as-is.

Done

  • All 10 chapters drafted
  • Every diagram illustrated
  • Kindle EPUB built
  • Author name and bio finalized
  • Kubernetes 1.35+ version audit
  • Paperback layout (English + Korean)

Remaining

  • Technical verification (Codex)
  • Human technical review
  • YAML/CLI executable verification
  • Final proofreading

AIDevOps Cloud Native Series

Next in the series

This book is designed as the series foundation — a mental model and an architecture map. The books below go far deeper into each individual topic.

AIDevOps Cloud Native Series roadmap

  • Book 02 — Kubernetes Troubleshooting Handbook: a much larger scenario catalog and complete forensic command sequences
  • Book 03 — Kubernetes Security in Production: RBAC design patterns, Pod Security Standards, and supply-chain security
  • Book 07 — Kubernetes Networking: CNI internals, full Gateway API conformance, and multi-cluster routing
  • Book 08 — Kubernetes Observability: Production Prometheus/Grafana deployment and complete alerting design
  • Book 10 — GitOps in Production: Argo CD/Flux operating patterns and progressive delivery

← Back to all books