Agentic AI

Agents that show their working

Ark investigates your environment the way an engineer would, and presents the path it took. An answer you cannot check is not usable for an operational decision, however good it is on average.

Short answer

What is agentic DevOps?

Agentic DevOps is the use of software agents that observe an infrastructure environment, decide at run time which evidence to gather, reason about what they find, and then propose or execute remediation, as distinct from automation that executes a predetermined sequence.

The distinction

Automation executes a sequence. An agent chooses one.

This is not a marketing distinction: it changes what the system can handle and what safeguards it needs.

Scripted automationAIOpsAgentic
Decides what to do nextNo, fixed sequenceNo, fixed modelsYes, at run time
Handles unanticipated situationsNoDetection onlyYes, within its access
Typical outputAn actionA detection or correlationA conclusion and a proposed plan
Main riskActing wrongly in an unforeseen caseFalse positivesConfident, wrong reasoning
Required safeguardTestingTuningLegibility, scope and approval

A script that restarts a pod when memory exceeds a threshold is automation. Something that notices the memory pattern, checks whether it correlates with a recent deployment, reads the logs for an out-of-memory signature, compares the configured limit against a month of observed usage, and concludes that the limit is too low rather than that the pod needs restarting. That is an agent. Nobody wrote that sequence of checks.

Grounding

A general model knows Kubernetes. It does not know your cluster.

This is the difference between a plausible list of causes and an answer.

Asked why a payment service is restarting, a general language model can only produce a generic list: memory limits, failing probes, image problems. That is a starting point you already had.

Ark resolves the same question into queries against your data: the workload configuration, its recent events, its memory usage against its limit, the deployments in that window, the log patterns that appeared. The model interprets the question and plans the investigation. The data provides the answer.

This is also why the agent layer sits on the same inventory as the rest of the platform rather than beside it. Grounding requires the data to be there.

Question or signal
Alert firedEngineer asksAnomaly detected
Ark agent
Plan investigationQuery platform dataCorrelateForm conclusion
Grounding data
InventoryMetricsLogsDeploymentsCostFindings
Output
Execution traceEvidenceProposed plan
Gate
Human approvalScoped executionAudit record
Ark reasons over the same inventory every other module writes to.
Legibility

The execution trace is the interface

Why is payments-api restarting?
  • Resolved service payments-api / prod-eu-west
  • Read pod events 14 restarts in 2h
  • Exit code 137 (OOMKilled)
  • Memory limit vs 30d peak 1Gi limit, 1.04Gi peak
  • Deployments in window 1 at 09:12 today
  • Comparing memory profile across revisions
The 09:12 release raised steady-state memory above the 1Gi limit. Recommended: set the limit to 1.5Gi. Six other services share this base image and show the same pattern.

Why this matters more than accuracy

Every system is wrong sometimes. The failure mode of an opaque agent is not that it is occasionally wrong; it is that being wrong is indistinguishable from being right until the consequences arrive.

A trace turns an unverifiable assertion into something an engineer can evaluate in seconds. It shows which queries ran, against which sources, what they returned, and how that produced the conclusion.

Trace, then trust. The trace is not a debug view for when something goes wrong. It is the mechanism by which an agent conclusion becomes usable for an operational decision.
Autonomy

Where to actually operate

Most of the value is in the first three levels, which is worth saying because the marketing usually points at the last.

1 · Observe

The agent gathers and correlates evidence, and reports. No actions. Almost pure benefit, because this is the phase that consumes most incident time.

2 · Recommend

The agent proposes a specific remediation with its reasoning. A person decides and acts.

3 · Approve and execute

The agent proposes a scoped plan, a person approves, the agent executes and reports.

4 · Bounded autonomy

Specific vetted playbooks run without approval within declared limits. Every execution is recorded.

5 · Full autonomy

The agent acts on its own judgement in unfamiliar situations. Very few organisations should be here, and none should start here.

FAQ

Questions about agents

Watch Ark investigate your own environment

Connect a cluster read-only and ask it something you would normally spend twenty minutes on.