Skip to content

VOL 07 / CH 07 / LESSON 01

7.1 Container Runtime Mechanisms: Images, Isolation, Resources, and Lifecycle ​

As services and instances multiply, manually copying directories, setting environment variables, and launching processes no longer produces consistent results. Teams need a repeatable way to deliver processes and their dependencies, but does “putting it in a container” create a lightweight virtual machine?

This lesson covers ordinary Linux containers: processes constrained by namespaces, cgroups, capabilities, and filesystem views that share the kernel of their Linux host. With Docker Desktop, that Linux host may itself be a virtual machine. Containers and VMs can be used together.

Images Are Immutable Content Graphs ​

OCI images are composed of a manifest, configuration, and read-only layers. Layers are addressed by content digests and can be shared across images.

This runtime-image example requires an executable app.jar in the build context. It demonstrates packaging and does not include the project’s compilation steps.

dockerfile
FROM eclipse-temurin:21-jre
WORKDIR /app
COPY --chown=10001:10001 app.jar /app/app.jar
USER 10001:10001
ENTRYPOINT ["java", "-jar", "/app/app.jar"]

The example uses a base-image tag for readability. Production builds should pin a reviewed base-image digest and releases should pin the final image digest. The ellipsis below is only a format placeholder, not a pullable image reference:

text
registry.example/atlas@sha256:...

When compilation tools are needed, a multi-stage build can leave them in the build stage. Control base-image contents, lock dependencies, and record SBOMs and signatures. Verify trusted signers and policy; merely finding a signature is insufficient. Deleting a secret does not erase it from an earlier layer. Exclude secrets from the build context. Credentials needed during a build can use temporary BuildKit secret mounts, avoiding ARG, ENV, or output files.

Namespace Isolation View ​

Namespaces such as PID, mount, network, UTS, IPC, and user allow processes to see different views of system resources. Isolation is not the entirety of security: kernel vulnerabilities, overly permissive capabilities, host-mounted filesystems, and privileged containers can still breach these boundaries.

By default, remove unnecessary Linux capabilities, enforce restrictions using seccomp, AppArmor, or SELinux, use a read-only root filesystem, and run processes as non-root users. The presence of a process inside a container does not imply trust in arbitrary code.

Cgroups Resource Measurement and Limitation ​

  • Orchestrator CPU requests inform scheduling and may map to cgroup competition weights. CPU quotas/limits cap CPU time within a period and may trigger throttling;
  • In cgroup v2, memory.max is a hard limit and can lead to an OOM kill if reclaim cannot satisfy allocation. memory.high can apply reclaim and allocation pressure, but excess memory cannot be time-shared like CPU;
  • PID, IO, and other resources can also be controlled.

Applications must know the available CPU and memory within a container, not the total host resources. Thread pool sizes, JVM heap limits, and cache boundaries should be aligned with the cgroup budget, leaving room for native memory and thread stacks within that container’s budget. Sidecars usually have separate container budgets; Pod and node planning must include all containers.

Writable Layer Is Transient ​

A writable layer belongs to the container object. Docker stop/start or restart normally retains it; removing and recreating the container replaces it. Kubernetes container replacement likewise cannot rely on files in the old layer. Put persistent business data in volumes with explicit lifecycles or external storage, and normally send logs to stdout/stderr.

A Pod’s emptyDir survives container restarts but disappears with Pod deletion. Persistent-volume retention depends on PVC/PV lifecycle and reclaim policy. Local sessions, uploads, and queues may be lost during replacement or migration and do not automatically become shared when scaling out. Applications with a read-only root filesystem should get a separate, capacity-limited temporary mount if they need temporary files.

PID 1 and Graceful Termination ​

The main process must handle termination signals and, if it spawns children, reap them correctly, using an appropriate init process when needed. The exec-form ENTRYPOINT above lets Java receive signals directly, avoiding a shell that may not forward them. Normal Kubernetes termination works roughly as follows:

  1. Record the Pod’s terminating state and update endpoints asynchronously; existing connections and cached routes do not disappear immediately;
  2. Run preStop if configured, counting its time against the termination grace period;
  3. Send the container’s main process its stop signal, usually SIGTERM, though runtime configuration can change it;
  4. Let the application drain and exit within the remaining grace period;
  5. Forcibly terminate processes still running after the grace period expires.

Applications should stop accepting new work, drain in-flight requests, and either complete or abandon tasks with clear semantic meaning before exiting. The grace period is not an infinite amount of time for processing.

Image Startup Success Does Not Mean It's Runnable ​

Pre-deployment validation:

  • Runs as a non-root user and on a read-only filesystem;
  • Behavior under CPU throttling, memory nearing capacity, and OOM (out-of-memory) conditions;
  • SIGTERM, connection draining, and duplicate message handling;
  • Vulnerabilities in the base image, digital signatures, and provenance;
  • Runtime dependencies such as temporary directories, certificates, and time zones;
  • Compatibility with host architecture and kernel capabilities.

The next lesson will integrate containers into Kubernetes' declarative control loop.

References ​

Built with VitePress | Software Systems Atlas