Agents that show their working
Ark investigates your environment the way an engineer would, and presents the path it took. An answer you cannot check is not usable for an operational decision, however good it is on average.
What is agentic DevOps?
Agentic DevOps is the use of software agents that observe an infrastructure environment, decide at run time which evidence to gather, reason about what they find, and then propose or execute remediation, as distinct from automation that executes a predetermined sequence.
Automation executes a sequence. An agent chooses one.
This is not a marketing distinction: it changes what the system can handle and what safeguards it needs.
| Scripted automation | AIOps | Agentic | |
|---|---|---|---|
| Decides what to do next | No, fixed sequence | No, fixed models | Yes, at run time |
| Handles unanticipated situations | No | Detection only | Yes, within its access |
| Typical output | An action | A detection or correlation | A conclusion and a proposed plan |
| Main risk | Acting wrongly in an unforeseen case | False positives | Confident, wrong reasoning |
| Required safeguard | Testing | Tuning | Legibility, scope and approval |
A script that restarts a pod when memory exceeds a threshold is automation. Something that notices the memory pattern, checks whether it correlates with a recent deployment, reads the logs for an out-of-memory signature, compares the configured limit against a month of observed usage, and concludes that the limit is too low rather than that the pod needs restarting. That is an agent. Nobody wrote that sequence of checks.
A general model knows Kubernetes. It does not know your cluster.
This is the difference between a plausible list of causes and an answer.
Asked why a payment service is restarting, a general language model can only produce a generic list: memory limits, failing probes, image problems. That is a starting point you already had.
Ark resolves the same question into queries against your data: the workload configuration, its recent events, its memory usage against its limit, the deployments in that window, the log patterns that appeared. The model interprets the question and plans the investigation. The data provides the answer.
This is also why the agent layer sits on the same inventory as the rest of the platform rather than beside it. Grounding requires the data to be there.
The execution trace is the interface
- Resolved service payments-api / prod-eu-west
- Read pod events 14 restarts in 2h
- Exit code 137 (OOMKilled)
- Memory limit vs 30d peak 1Gi limit, 1.04Gi peak
- Deployments in window 1 at 09:12 today
- Comparing memory profile across revisions
Why this matters more than accuracy
Every system is wrong sometimes. The failure mode of an opaque agent is not that it is occasionally wrong; it is that being wrong is indistinguishable from being right until the consequences arrive.
A trace turns an unverifiable assertion into something an engineer can evaluate in seconds. It shows which queries ran, against which sources, what they returned, and how that produced the conclusion.
Where to actually operate
Most of the value is in the first three levels, which is worth saying because the marketing usually points at the last.
The agent gathers and correlates evidence, and reports. No actions. Almost pure benefit, because this is the phase that consumes most incident time.
The agent proposes a specific remediation with its reasoning. A person decides and acts.
The agent proposes a scoped plan, a person approves, the agent executes and reports.
Specific vetted playbooks run without approval within declared limits. Every execution is recorded.
The agent acts on its own judgement in unfamiliar situations. Very few organisations should be here, and none should start here.
Questions about agents
Agentic DevOps is the use of software agents that observe an infrastructure environment, decide at run time which evidence to gather, reason about what they find, and then propose or execute remediation, as distinct from automation that executes a predetermined sequence.
Ark is the DevOpsArk agent layer. An Ark agent resolves a question or a signal into queries against your live platform data, correlates what it finds across metrics, logs, deployments and configuration, and reports a conclusion together with the execution trace that produced it.
Not by default. Ark proposes a scoped plan and requires approval from someone with the necessary permissions. Specific vetted playbooks can be allowed to run within declared bounds where you decide the risk is acceptable, and every execution is recorded and attributable.
You check the trace. Every conclusion is presented with the queries it ran, the sources it read and what they returned. This is why the trace is a primary part of the interface rather than a debug view: an operational decision needs a verifiable answer.
No. Ark answers within the asking user own permissions. If you cannot see a namespace in the console, Ark will not report on it for you.
No. Your environment data is used to answer your questions and to build your own baselines. It is not used as training data.
AIOps generally means applying statistical and machine learning techniques to operational data for detection, correlation and noise reduction. Agentic systems include those techniques and add the ability to plan and carry out multi-step investigation and propose remediation.
No. It removes the evidence-gathering phase, which is where a large share of incident time goes. The judgement about what to do stays with the engineer, which is also why approval gates exist rather than being a temporary limitation.
The mitigations are structural rather than accuracy claims: it proposes rather than executes, it can only invoke a fixed set of tested operations within a declared scope, its reasoning is visible for checking, and every action is recorded so it can be reviewed and reversed.
More on agentic DevOps
What is agentic DevOps?
A precise definition of agentic DevOps, how it differs from scripted automation and from AIOps, and the conditions under which an agent is safe to give real access.
AI agents in DevOps: where they help and where they do not
A practical assessment of where AI agents genuinely improve infrastructure work, where they are oversold, and how to introduce them without creating a new class of incident.
AI-powered incident response: compressing the first ten minutes
Where incident time actually goes, which parts an agent can take over safely, and how to structure incident response so the automation helps rather than adds noise.
Watch Ark investigate your own environment
Connect a cluster read-only and ask it something you would normally spend twenty minutes on.