Kubernetes Interview Questions for AI Training Work
AI training platforms often test Kubernetes directly, through a live coding round or a technical screen, rather than just taking a resume's word for it. These questions cover the parts of Kubernetes that actually come up under that kind of scrutiny: Pods, Deployments & Scaling, Networking & Services and Configuration & Secrets.
Below are 10 questions split into technical, scenario, and behavioral rounds, each with a full written answer so you can see what a strong response sounds like.
Technical (5)
What's the difference between a Pod and a Deployment?
A Pod is the smallest deployable unit, one or more containers that share networking and storage, but a Pod created directly has no self-healing: if it dies, it's gone. A Deployment manages a set of identical Pods via a ReplicaSet, continuously reconciling the actual state toward the desired replica count, so if a Pod crashes or its node fails, the Deployment creates a replacement automatically.
Explain the difference between a Service of type ClusterIP, NodePort, and LoadBalancer.
ClusterIP exposes a Service only inside the cluster, the default and most common choice for internal communication between components. NodePort opens a static port on every node, making the Service reachable from outside the cluster via any node's IP, but it's fairly manual to use directly. LoadBalancer provisions an external cloud load balancer, typically the choice for exposing a Service to the internet on a cloud provider.
What's the difference between a liveness probe and a readiness probe?
A liveness probe tells Kubernetes whether a container is still functioning; if it fails, Kubernetes restarts the container, since it's assumed to be stuck or broken. A readiness probe tells Kubernetes whether a container is ready to receive traffic; if it fails, the Pod is removed from a Service's endpoints without being restarted, which matters for things like a container still warming up a cache on startup that shouldn't be killed, just temporarily skipped.
How does a Kubernetes Deployment perform a rolling update, and how do you control how disruptive it is?
A rolling update gradually replaces old Pods with new ones, controlled by `maxSurge`, how many extra Pods can be created above the desired count during the rollout, and `maxUnavailable`, how many existing Pods can be down at once. Tuning these lets you trade rollout speed against risk: a lower `maxUnavailable` keeps more old Pods serving traffic during the transition, at the cost of a slower rollout.
What's the difference between a ConfigMap and a Secret, and why doesn't using a Secret automatically make data secure?
Both store configuration data injectable into Pods as environment variables or mounted files, but a Secret is intended for sensitive values, like API keys, while a ConfigMap is for non-sensitive configuration. By default, Secrets are only base64-encoded, not encrypted, in etcd, so anyone with API access to read Secrets can trivially decode them. Real protection requires enabling encryption at rest and tightly scoping RBAC permissions on who can read Secrets.
Scenario (3)
A Pod keeps entering CrashLoopBackOff. How do you diagnose it?
I'd start with `kubectl describe pod` to see recent events, like an OOMKilled status or a failed liveness probe, and `kubectl logs <pod> --previous` to see the logs from the last crashed instance, since logs from a fresh restart won't show why the previous one died. Common causes are a missing config value the app fails to start without, insufficient memory limits causing OOM kills, or a liveness probe misconfigured to fail before the app finishes starting up.
Traffic to a service is being routed to a Pod that isn't actually ready to handle requests yet. What's likely misconfigured?
This usually points to a missing or too-permissive readiness probe, if there's no readiness probe at all, Kubernetes considers a Pod ready as soon as its containers start, even before the application inside has finished initializing. I'd add a readiness probe that checks something meaningful, like an HTTP health endpoint that only returns success once startup tasks, like warming a cache or connecting to a database, are complete.
A namespace's Pods are being evicted under memory pressure even though the cluster has capacity elsewhere. What would you check?
I'd check whether the Pods have resource requests and limits set at all, since Pods without requests are scheduled less predictably and evicted first under pressure on their current node. I'd also check whether a ResourceQuota or LimitRange on the namespace is constraining things more tightly than the cluster's actual available capacity, since eviction is a per-node decision and doesn't automatically account for capacity on other nodes.
Behavioral (2)
Tell me about a production incident caused by a Kubernetes resource limit being set incorrectly.
A service had no memory limit set, and a slow leak in one Pod eventually consumed enough node memory to get it OOM-killed by the kernel along with unrelated Pods on the same node, turning one service's bug into a wider outage. I added memory limits and requests to every Deployment in that cluster afterward, and it became a required field in our deployment review checklist rather than an optional one.
Describe a time you had to decide between debugging a Kubernetes issue further versus just rolling back a deployment.
During a rollout that was causing intermittent errors, I gave myself a fixed 10-minute window to identify the cause from logs and metrics. When that window passed without a clear answer, I rolled back rather than continuing to investigate live in production, since restoring service took priority over root-causing it immediately. I debugged the actual cause afterward from the rolled-back environment's logs, without the pressure of live user impact.
Knowing the answer and saying it out loud under pressure are different skills.
The Academy has free modules and mock exams to build the second one.