Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #203Intermediate 10 min read

Kubernetes Core Concepts: Pods, Nodes, Services, & Ingress

💡
Core Architecture Summary

Master container orchestration: Control Plane (API Server, etcd, Scheduler, Controller Manager), Worker Nodes, Pods, ClusterIP vs NodePort vs LoadBalancer Services, and Ingress Controllers.

Key Glossary Concepts in this TopicAll Glossary Terms

Kubernetes Cluster Architecture & Traffic Ingress ☸️

Control Plane consensus and scheduling managing Worker Nodes, Services (ClusterIP VIPs), and Layer 7 Ingress routing.

Kubernetes Cluster Architecture & Traffic Ingress ☸️
100%
Rendering visual architecture flowchart...

01.1. What is Kubernetes (K8s) & Cluster Architecture

Kubernetes is an open-source, declarative container orchestration platform originally designed by Google (Borg) and maintained by the Cloud Native Computing Foundation (CNCF). It automates the provisioning, scheduling, horizontal scaling, networking, and self-healing of containerized applications across clusters of physical or virtual machines.

A production Kubernetes cluster is bifurcated into two primary planes:

  1. The Control Plane (Master Nodes): Makes global decisions about the cluster (e.g., scheduling pods, detecting node failures).
  2. Worker Nodes: Provide compute, memory, disk, and network resources to execute user container workloads.

02.2. Control Plane Internals: API Server, etcd, Scheduler, and Controllers

The Control Plane maintains the declarative truth of the entire cluster:

  • kube-apiserver (The Hub): The single stateless entrypoint for all cluster management. It exposes the Kubernetes REST API, handles authentication (X.509 certs, OIDC tokens), authorization (RBAC), and runs Admission Webhooks (Validating and Mutating). No component touches etcd directly except kube-apiserver.
  • etcd (Distributed Consensus Store): A strongly consistent, distributed key-value store implementing the Raft consensus algorithm. It persists the desired and actual state of all cluster resources. A standard HA deployment runs 3 or 5 etcd replicas to tolerate 1 or 2 node failures.
  • kube-scheduler: Watches for newly created pods without an assigned node. It evaluates candidates through two phases:
    1. Filtering (Predicates): Eliminates nodes lacking sufficient CPU/RAM requests, lacking required node labels, or matching node taints.
    2. Scoring (Priorities): Ranks qualifying nodes based on topology spread (spreading pods across Availability Zones) and resource balance, assigning the pod to the highest-scoring node.
  • kube-controller-manager: Runs continuous reconciliation control loops (e.g., NodeController, ReplicaSetController, EndpointSliceController, ServiceAccountController). It continuously drives the actual cluster state toward the declared desired state.

03.3. Worker Node Architecture: kubelet, kube-proxy, and CNI

Worker nodes execute container workloads and report health back to the control plane:

  • kubelet: The primary node agent running as a system daemon. It registers the node with the API server, watches for PodSpecs assigned to its node, and instructs the container runtime via the CRI (Container Runtime Interface) (e.g., containerd, CRI-O) to launch or stop containers. It also executes periodic liveness, readiness, and startup health probes.
  • kube-proxy: A network proxy running on each node that implements the Kubernetes Service abstraction:
    • iptables mode: Generates Linux netfilter NAT rules to randomly distribute traffic arriving at a Service Virtual IP (VIP) to healthy backend Pod IPs.
    • IPVS mode (IP Virtual Server): Utilizes Linux kernel IPVS hash tables for O(1) lookups, scaling cleanly to clusters with tens of thousands of services where iptables O(N) rule evaluation causes CPU latency.
  • CNI (Container Network Interface): Implements flat, non-NAT pod-to-pod networking (e.g., AWS VPC CNI, Calico, Cilium with eBPF). Every pod receives a unique, routable IP address in the cluster CIDR.

04.4. The Pod Primitive & Multi-Container Patterns

A Pod is the smallest deployable compute unit in Kubernetes. A pod represents a single instance of a running process:

  • A pod encapsulates one or more closely coupled containers sharing the same Network Namespace (same IP address, port space, and reachable via localhost) and shared Storage Volumes.
  • The Sidecar Pattern: Running a helper container alongside the main application container within the same pod (e.g., an Envoy proxy for service mesh mTLS, a FluentBit agent for log shipping, or an OpenTelemetry collector).
  • Init Containers: Specialized containers that execute sequentially to completion before the application containers start (e.g., executing database schema migrations or waiting for a cache to become reachable).

05.5. Services & Ingress: Solving Ephemeral IP Discovery

Pods are ephemeral; when a pod crashes or scales, it is replaced with a new pod with a completely new IP address. Kubernetes Services provide a permanent, stable network identity:

  1. ClusterIP (Default): Assigns a permanent internal Virtual IP (VIP) and internal cluster DNS name (my-service.my-namespace.svc.cluster.local). Traffic sent to this VIP is load-balanced across healthy pods tracked by EndpointSlices.
  2. NodePort: Exposes the service on a dedicated static port (range 30000-32767) across every worker node's physical IP address.
  3. LoadBalancer: Integrates with cloud providers (AWS NLB, GCP Cloud LB) to automatically provision an external hardware/cloud load balancer that forwards public internet traffic to the cluster NodePorts.
  4. Ingress & Ingress Controllers (Layer 7): An API object that manages external HTTP/HTTPS routing rules (Host-based api.example.com and Path-based /orders vs /users), TLS termination via cert-manager (Let's Encrypt), and rate limiting before routing traffic to internal ClusterIP Services.

⚖️Architectural Trade-offs & Production Realities

Architectural Advantages

  • Declarative state management eliminates imperative manual scripting and ClickOps human errors
  • Unified abstraction layer across public clouds (AWS EKS, GCP GKE, Azure AKS) and on-premise bare-metal
  • Automated service discovery, load balancing, health monitoring, and self-healing restarts

Trade-offs & Constraints

  • High architectural complexity and steep learning curve (RBAC, CNI, CSI, CRDs, Ingress)
  • Operational overhead of managing control plane upgrades, etcd snapshots, and node pool maintenance
Production Implementation in Big Tech
Spotify & OpenAI• Massive Kubernetes Orchestration

Spotify migrated their entire backend infrastructure of over 1,500 microservices to Kubernetes clusters across multiple regions to standardize deployment pipelines. Similarly, OpenAI manages specialized multi-thousand-node Kubernetes clusters orchestrating distributed GPU clusters for training large language models like GPT-4.

🎯 Staff+ Engineering Takeaways

  • The Control Plane (kube-apiserver, etcd, scheduler, controller-manager) coordinates the cluster declarative state.
  • Worker nodes run kubelet (agent), kube-proxy (networking), and the container runtime.
  • Pods are ephemeral and share network namespaces (localhost communication between sidecars).
  • Services provide stable VIPs and DNS endpoints backed by EndpointSlices.
  • Ingress controllers provide Layer 7 HTTP path/host routing and TLS termination.

Topic Knowledge Assessment 🧠

Step through 3 scenario questions to test your staff-level grasp.

Question 1 of 30 answered
#1

Why do Kubernetes applications connect to a Service rather than connecting directly to backend Pod IP addresses?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?