The highest-value container security controls are unglamorous: do not run as root, keep the final image minimal, pin base images by digest, drop capabilities, use a read-only root filesystem, and scan continuously rather than only at build. Most breaches involving containers exploit ordinary misconfiguration rather than exotic escapes.
What are the most important container security practices?
Run containers as a non-root user, keep the final image minimal so it contains no shell or build tooling, pin base images by digest rather than by floating tag, drop all Linux capabilities and add back only what is needed, mount the root filesystem read-only, and rescan deployed images continuously as new advisories are published.
Build-time: make the image small and unprivileged
Everything in the final image is attack surface. A multi-stage build that copies only the compiled artifact into a slim base removes compilers, package managers, shells and source code, all of which are useful to an attacker and none of which are needed at run time.
FROM golang:1.22-bookworm@sha256:9a1c3f... AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -o /out/app ./cmd/app
# Distroless: no shell, no package manager, no coreutils.
FROM gcr.io/distroless/static-debian12:nonroot@sha256:6cd6f4...
COPY --from=build /out/app /app
USER 65532:65532
ENTRYPOINT ["/app"]- Pin the base image by digest. A tag can be repointed; a digest cannot, which is what makes the build reproducible and the supply chain reviewable.
- Prefer a distroless or minimal base. If there is no shell in the image, a great many post-exploitation techniques simply do not work.
- Never bake secrets into layers. A deleted file remains in the layer it was added to and is recoverable from the image.
- Generate an SBOM per image, so the next advisory is answered by a query rather than a rescan of the estate.
Runtime: restrict what the container can do
A hardened image running with a permissive pod specification is not hardened. The Kubernetes-side controls matter as much as the image-side ones.
securityContext:
runAsNonRoot: true
runAsUser: 65532
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"] # add back only what is genuinely needed
seccompProfile:
type: RuntimeDefault
# Writable paths, if the application needs them, are explicit and bounded.
volumeMounts:
- name: tmp
mountPath: /tmp
volumes:
- name: tmp
emptyDir: {}| Control | What it prevents |
|---|---|
| runAsNonRoot | Root inside the container, which makes escapes far more productive |
| allowPrivilegeEscalation: false | setuid binaries regaining privilege inside the container |
| readOnlyRootFilesystem | Attacker tooling being written into the container |
| drop: ALL capabilities | Kernel operations the workload never needs |
| seccompProfile: RuntimeDefault | Access to rarely needed syscalls used in escape chains |
| Network policy | Lateral movement from a compromised pod to unrelated services |
| No hostPath, hostNetwork or hostPID | Direct access to the node from inside a container |
Supply chain: know what you shipped
The container supply chain runs from base image, through dependencies, through build, to registry, to cluster. Each link is a point where something can be substituted.
- Deploy by digest, not by tag, so what was scanned is provably what runs.
- Restrict which registries a cluster may pull from, via admission policy.
- Sign images and verify signatures at admission, so an image that did not come from your pipeline cannot be deployed.
- Keep provenance: commit, build inputs, digest, scan verdict, approver. This is what answers "where did this container come from" during an incident.
- Maintain a curated set of approved base images, patched centrally, with an inventory of which service uses which.
Scan continuously, not once
A build-time scan checks an image against the advisories known at that moment. An image that was clean in March is not clean in October, and nothing about the image changed. The world did.
Continuous re-evaluation of deployed artifacts against new advisories is what catches this. Combined with an SBOM index, it also turns the question that follows a major disclosure (which of our services ship this library) into a query answered in seconds rather than a week of scanning.
What to do first
- Stop running containers as root. Highest risk reduction per unit of effort by a wide margin.
- Move to minimal base images. Removes a large amount of attack surface in one change.
- Enforce a pod security baseline at admission, in report mode first.
- Pin digests and deploy by digest.
- Turn on continuous rescanning of deployed images.
- Add network policies, starting with default deny in the most sensitive namespaces.
Key takeaways
- Everything in the final image is attack surface. Multi-stage builds remove most of it.
- A hardened image with a permissive pod spec is not hardened; both sides matter.
- Pin and deploy by digest so what was scanned is provably what runs.
- Build-time scanning does not protect against advisories published after the build.
- Not running as root is the single highest-value change available.
Frequently asked questions
Because root inside a container is a much better starting position for an attacker. Many container escape techniques require root in the container, and a compromised root process can install tooling, modify the filesystem and interact with the runtime in ways an unprivileged one cannot.
A container image containing only the application and its runtime dependencies: no shell, no package manager, no standard utilities. It reduces attack surface substantially and makes post-exploitation harder, at the cost of being harder to debug interactively.
At build, and then continuously against new advisories. Vulnerabilities are disclosed after images are built, so a single build-time scan gives a false sense of currency.
A software bill of materials: a machine-readable list of every component in an artifact. Its practical value is that when a new vulnerability is disclosed, you can answer which of your services ship the affected component from an index rather than by rescanning everything.
For most applications, yes. Where temporary files are needed, mount an emptyDir at the specific path. The exercise of enumerating which paths genuinely need to be writable is itself useful.
They answer different questions. Image scanning tells you what vulnerabilities exist in the artifact; runtime detection tells you whether something unexpected is happening now. Start with hardening and scanning, which prevent more than detection catches.