Most Dockerfiles in the wild are single-stage, run as root, invalidate their dependency cache on every source change and ship a compiler to production. Fixing that per repository does not scale. Generate the definition from repository analysis, apply hardening as organisation policy, order layers so dependencies cache independently of source, and track base images centrally so patching is one operation instead of forty.
How do you automate Docker image builds?
Automate Docker builds by generating the container definition from repository analysis rather than writing it by hand: detect the language and framework from manifest files, emit a multi-stage Dockerfile with an approved base image and a non-root user, order instructions so dependency layers cache independently of source changes, and build with BuildKit cache mounts. Scan the resulting image before it is tagged as releasable.
The four problems with a typical Dockerfile
- It is single-stage, so the compiler, build tools and source all ship to production inside the final image.
- It copies the whole source tree before installing dependencies, so every source change invalidates the dependency layer and reinstalls everything.
- It runs as root, because that was the fastest way to make it start.
- It uses a floating tag such as node:20, so the image built today and the image built next month are different in ways nobody recorded.
Each is easy to fix in one repository and hard to fix in eighty, which is why the standard drifts. The scalable answer is to generate the definition rather than to review it.
Multi-stage builds, and why layer order matters
A multi-stage build compiles in one stage and copies only the result into a clean final stage, so build tooling never reaches production. Combined with correct instruction ordering, it also fixes the caching problem: dependency manifests are copied and installed before the source is copied, so a source change reuses the cached dependency layer.
# syntax=docker/dockerfile:1.7
# ---- build stage ----------------------------------------------------------
FROM node:20.14.0-bookworm-slim@sha256:aa1c0e... AS build
WORKDIR /app
# Manifests first: this layer only changes when dependencies change.
COPY package.json package-lock.json ./
RUN --mount=type=cache,target=/root/.npm \\
npm ci --omit=dev
# Source second: a code change invalidates from here down, not above.
COPY . .
RUN npm run build
# ---- runtime stage --------------------------------------------------------
FROM node:20.14.0-bookworm-slim@sha256:aa1c0e...
ENV NODE_ENV=production
WORKDIR /app
# Non-root by default. The image will not start as root.
RUN useradd --system --uid 10001 --no-create-home app
COPY --from=build --chown=10001:10001 /app/node_modules ./node_modules
COPY --from=build --chown=10001:10001 /app/dist ./dist
USER 10001
EXPOSE 8080
CMD ["node", "dist/server.js"]Caching that survives real workflows
- Copy dependency manifests and lock files before the rest of the source, so the install layer is independent of code changes.
- Use BuildKit cache mounts for the package manager cache directory, so even a dependency change does not redownload everything.
- Export and import a build cache in CI, because ephemeral runners start with nothing.
- Keep .dockerignore accurate. Copying .git or node_modules into the build context invalidates caches and inflates images.
The measurable outcome is the difference between a cold and warm build. If a one-line source change triggers a full dependency install, the caching is not working regardless of what the configuration says.
Managing this across eighty services
The per-repository approach has a well-known failure mode: the standard is agreed, applied to the services that are actively worked on, and never reaches the rest. A year later there are three generations of Dockerfile in the estate and no inventory of which is which.
The alternative is to generate the definition from analysis of the repository and treat the container standard as organisation policy: approved base images, required user handling, forbidden final-layer contents. Because the policy is applied at generation rather than at review, a change to the standard can be rolled out as a rebuild rather than as eighty pull requests.
It also makes base image lifecycle tractable. When a base image picks up a critical vulnerability, the platform knows which services use it and can raise one rebuild plan covering all of them, instead of the security team asking each team to check.
How ArkBuilder does it
ArkBuilder reads the repository, identifies the stack from its manifest and lock files, and generates a multi-stage definition that it commits to your repository so the result stays reviewable in Git rather than hidden in a platform. It builds with BuildKit and cache mounts, reports the size delta of every layer against the previous build, scans the image before it can be promoted, produces an SBOM, and publishes by immutable digest.
Base image lineage is tracked across every service, so a patched base image produces a single rebuild plan listing everything affected.
Key takeaways
- Multi-stage builds keep compilers and build tooling out of the production image.
- Copy dependency manifests before source so the install layer caches independently of code changes.
- Pin base image digests, not floating tags, if you want reproducible builds.
- Run as a non-root user by default, enforced at generation rather than at review.
- Generate container definitions from repository analysis so the standard reaches every service, not just the active ones.
- Track base image lineage centrally so patching is one plan rather than forty conversations.
Frequently asked questions
Almost always because the source is copied before dependencies are installed, so any code change invalidates the dependency layer. The other common causes are an inaccurate .dockerignore that lets changing files into the build context, and CI runners that start fresh without importing a build cache.
A build with more than one FROM instruction, where earlier stages compile or prepare artifacts and the final stage copies only what is needed to run. The result is a smaller image containing no compilers, build tools or source.
No. A container running as root that is compromised gives the attacker root inside the container, and a container escape then has a much better starting position. Create a non-root user in the image and set USER before the entrypoint.
Use a multi-stage build so build tooling never reaches the final image, choose a slim maintained base, avoid installing recommended packages, and remove package manager caches in the same layer that created them. Then measure per-layer size on each build so growth is visible.
Yes, if you want reproducible builds. A tag can be repointed at a new image; a digest cannot. Pinning by digest means updates are a deliberate, reviewable change rather than something that happens on the next build.
You still have one: ArkBuilder commits the generated definition to your repository so it stays visible and reviewable. What changes is that you are not writing or maintaining it by hand.