DevOps

How to automate Docker builds without hand-writing Dockerfiles

Multi-stage builds, layer caching that actually works, hardening defaults, and how to generate and maintain container definitions across a large service estate.

DevOpsArk EngineeringEngineering team, DevOpsArkPublished 19 March 2026 · Updated 22 July 20269 min read
TL;DR

Most Dockerfiles in the wild are single-stage, run as root, invalidate their dependency cache on every source change and ship a compiler to production. Fixing that per repository does not scale. Generate the definition from repository analysis, apply hardening as organisation policy, order layers so dependencies cache independently of source, and track base images centrally so patching is one operation instead of forty.

Short answer

How do you automate Docker image builds?

Automate Docker builds by generating the container definition from repository analysis rather than writing it by hand: detect the language and framework from manifest files, emit a multi-stage Dockerfile with an approved base image and a non-root user, order instructions so dependency layers cache independently of source changes, and build with BuildKit cache mounts. Scan the resulting image before it is tagged as releasable.

The four problems with a typical Dockerfile

  1. It is single-stage, so the compiler, build tools and source all ship to production inside the final image.
  2. It copies the whole source tree before installing dependencies, so every source change invalidates the dependency layer and reinstalls everything.
  3. It runs as root, because that was the fastest way to make it start.
  4. It uses a floating tag such as node:20, so the image built today and the image built next month are different in ways nobody recorded.

Each is easy to fix in one repository and hard to fix in eighty, which is why the standard drifts. The scalable answer is to generate the definition rather than to review it.

Multi-stage builds, and why layer order matters

A multi-stage build compiles in one stage and copies only the result into a clean final stage, so build tooling never reaches production. Combined with correct instruction ordering, it also fixes the caching problem: dependency manifests are copied and installed before the source is copied, so a source change reuses the cached dependency layer.

# syntax=docker/dockerfile:1.7

# ---- build stage ----------------------------------------------------------
FROM node:20.14.0-bookworm-slim@sha256:aa1c0e... AS build
WORKDIR /app

# Manifests first: this layer only changes when dependencies change.
COPY package.json package-lock.json ./
RUN --mount=type=cache,target=/root/.npm \\
    npm ci --omit=dev

# Source second: a code change invalidates from here down, not above.
COPY . .
RUN npm run build

# ---- runtime stage --------------------------------------------------------
FROM node:20.14.0-bookworm-slim@sha256:aa1c0e...
ENV NODE_ENV=production
WORKDIR /app

# Non-root by default. The image will not start as root.
RUN useradd --system --uid 10001 --no-create-home app
COPY --from=build --chown=10001:10001 /app/node_modules ./node_modules
COPY --from=build --chown=10001:10001 /app/dist ./dist

USER 10001
EXPOSE 8080
CMD ["node", "dist/server.js"]
Pin the digest, not just the tag. node:20.14.0 can be rebuilt and republished. The digest cannot change. Pinning the digest is what makes a build reproducible; pinning only the tag makes it approximately reproducible, which is not a property you can rely on during an incident.

Caching that survives real workflows

  • Copy dependency manifests and lock files before the rest of the source, so the install layer is independent of code changes.
  • Use BuildKit cache mounts for the package manager cache directory, so even a dependency change does not redownload everything.
  • Export and import a build cache in CI, because ephemeral runners start with nothing.
  • Keep .dockerignore accurate. Copying .git or node_modules into the build context invalidates caches and inflates images.

The measurable outcome is the difference between a cold and warm build. If a one-line source change triggers a full dependency install, the caching is not working regardless of what the configuration says.

Managing this across eighty services

The per-repository approach has a well-known failure mode: the standard is agreed, applied to the services that are actively worked on, and never reaches the rest. A year later there are three generations of Dockerfile in the estate and no inventory of which is which.

The alternative is to generate the definition from analysis of the repository and treat the container standard as organisation policy: approved base images, required user handling, forbidden final-layer contents. Because the policy is applied at generation rather than at review, a change to the standard can be rolled out as a rebuild rather than as eighty pull requests.

It also makes base image lifecycle tractable. When a base image picks up a critical vulnerability, the platform knows which services use it and can raise one rebuild plan covering all of them, instead of the security team asking each team to check.

How ArkBuilder does it

ArkBuilder reads the repository, identifies the stack from its manifest and lock files, and generates a multi-stage definition that it commits to your repository so the result stays reviewable in Git rather than hidden in a platform. It builds with BuildKit and cache mounts, reports the size delta of every layer against the previous build, scans the image before it can be promoted, produces an SBOM, and publishes by immutable digest.

Base image lineage is tracked across every service, so a patched base image produces a single rebuild plan listing everything affected.

Key takeaways

  • Multi-stage builds keep compilers and build tooling out of the production image.
  • Copy dependency manifests before source so the install layer caches independently of code changes.
  • Pin base image digests, not floating tags, if you want reproducible builds.
  • Run as a non-root user by default, enforced at generation rather than at review.
  • Generate container definitions from repository analysis so the standard reaches every service, not just the active ones.
  • Track base image lineage centrally so patching is one plan rather than forty conversations.

Frequently asked questions

DockerContainersBuild automationCI/CD

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.