AutoOps
A CLI and API-driven agentic harness for AIOps. Governed autonomous SRE for AI infrastructure, Kubernetes platforms, and private AI operations. It turns telemetry into diagnosis, remediation plans, and auditable operational evidence.
For platform & AI infra teams
Designed specifically for the engineers maintaining AI workloads, focusing on raw telemetry, terminal interfaces, and un-abstracted API access rather than point-and-click dashboards.
Read-only by default
AutoOps operates strictly in shadow mode out of the box. It analyzes live production data to propose root causes without ever mutating infrastructure state without explicit approval.
Evidence first
Every remediation proposal is backed by a cryptographically verifiable ledger of correlated logs, traces, and policy evaluations. No black-box AI decisions.
Policy governed
Execution is gated by strict Cedar policies. Destructive operations (like cluster rollbacks) require human-in-the-loop escalation.
What AutoOps adds to Setloop
Setloop already helps teams design and run GPU workloads, AI platforms, private AI, and agent observability. AutoOps closes the operational loop: diagnose, govern, verify, and document what happened.
Incident diagnosis
Reads logs, metrics, traces, Kubernetes state, deployments, and service bindings to localize faults.
Governed automation
Cedar policy, destructive-command checks, and short-lived authority keep operations controlled.
Private telemetry handling
Classifies and redacts secrets, tokens, public IPs, and PII before data reaches the model.
Evidence ledger
Every run can produce structured events, result rows, manifests, and an audit trail for review.
Model portability
Uses a clean LLM client port with Anthropic and OpenRouter adapters, ready for private model routes.
Platform fit
Designed around Kubernetes, observability systems, GPU infrastructure, and AI service operations.
The autonomous operations copilot.
AutoOps connects infrastructure signals to operational decisions.
Autonomous operations copilot
For teams running inference platforms, private AI systems, managed GPU clusters, and agent workloads.
Complements GPU Cloud, FinOps & LLMTrace
AutoOps connects infrastructure signals to operational decisions and customer-facing reliability.
Diagnose first, act when authorized
The product defaults to read-only investigation, then escalates or executes governed remediation.