PortCon: The Agentic SDLC Summit

Resolve AI Alternatives: How to Choose Your Incident AI Layer

What does Resolve AI do? It runs AI agents that investigate production incidents. A buyer's guide to its features, limits, pricing, and alternatives.

Kevin Wolf
Kevin Wolf
September 28, 2026
Kevin Wolf
Kevin Wolf&
September 28, 2026
Kevin Wolf
Kevin Wolf&&
September 28, 2026
Resolve AI Alternatives: How to Choose Your Incident AI Layer

Resolve AI runs AI agents that investigate production incidents, tracing root causes across your code, infrastructure, telemetry, and internal knowledge. Investigation is the strong part. On actions, their documentation is narrower, listing silencing and snoozing alerts as the only mitigation Resolve performs today, so carrying out a fix still falls to a person.

This guide is for the engineering leader or platform owner asking: what does Resolve AI do, and what do I need around it? By the end you will know how Resolve AI builds its picture of your systems, where it sits in a wider stack, and how to choose between the specialist and the layer it runs on.

What is Resolve AI?

Resolve AI is a production operations company whose agents investigate incidents on your behalf. They describe the product as "AI for prod," software that "works across your code, infrastructure, telemetry, and knowledge to help engineering teams run production more reliably and efficiently."

Spiros Xanthos and Mayank Agarwal founded it, and it came out of stealth in late 2024. Both co-created OpenTelemetry and ran Splunk's observability business before starting the company.

Three capabilities anchor it, in their own framing:  

  1. Agents investigate alerts automatically as they fire, gather evidence across your systems, and surface a root cause hypothesis. 
  2. Agents join incident channels and work remediation alongside the people already there. 
  3. Teams encode their own practices as repeatable workflows that run health checks, generate reports, and execute multi-step investigations on demand. Their documentation stresses that investigations pursue several hypotheses in parallel, and that the product works against your tools without asking the engineer to know tool specific queries or unfamiliar code.

They connect to roughly 45 tools, including Datadog, Grafana, Splunk, PagerDuty, ServiceNow, and the major clouds. For companies whose telemetry cannot leave their own network, Resolve ships Satellite, a component you run inside your VPC or on ECS so the product reaches on-prem data without that data leaving. Their customer page names DoorDash, Coinbase, Snowflake, and Robinhood, among others.

What problem does Resolve AI solve?

Production knowledge concentrates on a few people who have been there longest, and every incident re-derives something someone already knew. Resolve AI aims to address that gap.

An alert fires at 2am and the engineer on call spends the first twenty minutes working out which team owns the failing service. MTTR is accumulating on routing alone before diagnosis has even started.

The same failure comes back a quarter later and gets investigated from scratch, because last time's root cause lives in a Slack thread nobody can find. The organization pays twice for one lesson, and the repeat-incident rate, the share of alerts recurring within ninety days, stays flat no matter how good any single investigation was.

The agent works out the fix in minutes, and an engineer then spends an hour carrying it out by hand. Proposing and doing are separate steps, so a fast diagnosis does not reach the outcome, and what moves here is the share of remediations that finish without a person typing commands.

How does Resolve AI solve it?

Building the picture: Resolve pulls your Kubernetes, code repositories, and observability tools (Datadog, Grafana, Splunk, and 40+ others). When you enable Infer Applications, it scans your Kubernetes resource manifests, extracts application names using JSON paths you configure, and groups related resources together. On top of that, it builds causal timelines, linking code changes, infrastructure events, and metric anomalies, and reasons across them during an investigation.

Surfacing your knowledge: You upload documentation, runbooks, and playbooks (Resolve calls them Skills: written procedures for triage, rollback, or postmortem). Resolve stores these at two levels, organization-wide and team-specific, and surfaces the ones that match during an investigation. When a Skill's description fits the incident, Resolve activates it automatically.

Resolve proposes and a person approves: Their documentation is direct about it: "Resolve AI only proposes actions. Every action requires explicit human approval before execution." The model doing the reasoning cannot call write APIs at all, and a separate execution engine holds the credentials and runs only once someone has approved. Today that covers silencing and snoozing alerts across Grafana, Kloudfuse, Datadog, AlertManager, and Sumo Logic, which they class as low risk because you can revoke a silence at any time.

Exposing findings: Resolve also exposes an MCP server and a REST API. Coding agents such as Claude Code and Cursor connect over MCP to open and steer investigations, pull production context before a risky change, and turn findings into pull requests. From outside the product, an investigation is how you reach any of this. The entity model, Team Knowledge, and the action layer each appear in the documentation as part of how an investigation runs.

Approaches to building your Agentic SDLC stack, and where Resolve AI fits

Resolve AI: buy the incident layer

This is the fastest path to a working investigation. The product ships domain reasoning that usually takes a dedicated team to build, an on-call experience engineers can use without training, and, through Satellite, an answer for organizations whose telemetry cannot leave the network. Resolve's systems picture and action controls are scoped to investigations, not to a shared platform. The catalog it reads from, the action framework it calls, and governance for other agents sit outside this product, so your team buys or builds them alongside it.

Port: buy the layer underneath it

Port keeps one catalog of your services, teams, ownership, deployments, incidents, and agents, and it updates itself from the systems it connects to instead of going stale between reviews. Everything reads that same catalog, whether it is a developer in the UI, a workflow, a scorecard, or an agent over MCP. Approvals and audit attach to the action rather than to whoever calls it, so they hold for a person and an agent alike. Incidents are one workload on that layer, sitting alongside ticket resolution, resource management, and code review. The scope is wider than an incident tool's, so there are more of your systems to map.

Where the shared layer pays off

Our position is that you should buy the specialist for the job it does best, and keep the model of your systems and the authority to act one layer below it, where every agent can reach them. That means mapping more of your systems than an incident tool needs, and the return shows up in each of the three situations above.

Ownership at 2am resolves from a catalog you defined, with service-to-team relationships you set, so it holds for the things that were never in a Kubernetes cluster: a managed database, a vendor API, a scheduled job, a data pipeline. Port auto-discovers your infrastructure too, so the catalog covers the cluster and everything sitting outside it.

When the same failure returns, the next alert starts from the service's history, and repeat rate and MTTR move with it. The incident attaches to the service entity, so any agent can read it. A scorecard turns what the team learned into the standard that measures the service.

The gap between proposing and doing closes because the action is a first-class object. You define what it does, who can run it, and what approval it needs, and the audit trail follows on its own. An agent calls that action the same way a person does, through the same approval, and the run history records who asked and who approved. Resolve's documented mitigations stop at silencing and snoozing alerts.

Port (Agentic SDLC Platform) vs Resolve AI

Resolve's picture of your production comes out of investigating it. What ends up in that picture, what can read it, and what it can act on all follow from that.

Capability Port Resolve AI
Acting on what you find
Custom workflow framework ✅ Define any workflow (script, API call, workflow) with typed inputs ❌ Docs list alert silencing and snoozing only, no framework for user-defined actions
What the gate covers ✅ Any action in the catalog. An agent calling run_action over MCP reaches only actions its token has permission to run ◐ The mitigations Resolve itself proposes
Human approval on execution ✅ Manual approval per action, with dynamic RBAC on who runs and who approves ✅ Every proposal approved by a human in their UI or Slack before execution
Reach of the model
What can read the model ✅ MCP server and REST API expose catalog entities, blueprints, and scorecards to any client ◐ Eleven MCP tools and three REST endpoints, all scoped to investigations, chats, and alerts
Same model feeds scorecards and workflows ✅ One set of entities drives scorecards, automations, and self-service ◐ Knowledge and entities are retrieved inside investigations
Span of the model
Auto-discovery of infrastructure ✅ Kubernetes stack integration plus catalog auto-discovery ✅ Infer Applications extracts names from Kubernetes via JSON paths
Entity types and relationships you define ✅ Blueprints for any type (service, team, domain, agent, incident, cost center) with typed relations between them ◐ Application entities grouped from cluster resources, with the grouping built by the product
Investigation depth
Causal timeline across code, infra, and telemetry ◐ Port AI reasons over the catalog and connected MCP servers ✅ Purpose-built causal timelines linking code changes, infrastructure events, and telemetry
Time to a first useful investigation ◐ You configure the agent and the context it reads ✅ Auto-investigation runs on the alert with no per-incident setup
Agents you didn't build
Inventory of external agents ✅ Built-in AI Agent blueprint, extended with platform, status, model, region, and owning team ❌ Their documented surface treats outside agents as MCP clients that consume investigations
What an outside agent can reach ✅ Agents query the catalog through the Port MCP server, which exposes only what the calling token's permissions allow ◐ Their MCP scopes outside agents to investigations, chats, and alerts
Commercial
Published tiers ✅ Free, Basic, and Standard published with seat, entity, and run limits ❌ Contact sales only
Free tier ✅ 15 seats, no time limit, no credit card ❌ None disclosed

Port and Resolve disclose very different amounts about what they cost. Port publishes its tiers with seat counts, entity and run limits, and what each tier gates, plus a free tier with no time limit and no credit card, so you can model most of your cost before you talk to anyone. Resolve publishes a contact form inviting you to "talk with our team about pricing, deployment guidance, and integration planning," with no tiers, no per-unit rates, and no free tier disclosed. Port's Enterprise tier is quoted on request, so the published numbers cover Free, Basic, and Standard.

Whatever number you are quoted covers the investigation layer. Everything listed above as yours to buy or build is a separate line item on top of it.

Bottom line: Resolve investigates well, but what it learns and what it may do stay inside it, so you still have to own the layer underneath.

How to choose, and where to start

The teams that get the most out of agents put the context and the controls in place first, then point agents at them. dLocal is the clearest public example. They run Port as their Agentic SDLC Platform, with the catalog and governance established before the agent went to work, and they now resolve 45% of tickets autonomously with MTTR cut in half. dLocal did not evaluate Resolve AI, so this is not a displacement story. The catalog and the governance came first, and the 45% came after.

Choose Resolve AI if production incidents are your sharpest pain, you want a deep investigation agent running in weeks rather than quarters, your telemetry needs to stay inside your network, and you are content for the context and controls it builds to live inside it.

Choose Port if you expect to run agents from several vendors, you want one model of your systems that every agent, workflow, and scorecard reads, and you want the approval and audit on an action to apply no matter who or what calls it.

Book a demo and bring a real incident from last quarter, and we will walk it against your own stack. If you would rather look around first, sign up for the free tier, which has no time limit and no credit card.

FAQ

What does Resolve AI do?

Resolve AI runs AI agents that investigate production incidents. It connects to your observability, cloud, and git tools, builds causal timelines across code changes, infrastructure events, and telemetry, and surfaces root cause hypotheses automatically when an alert fires. It also proposes mitigations, currently silencing and snoozing alerts, which a human approves before anything executes.

What is an Agentic SDLC Platform?

An Agentic SDLC Platform holds one model of your engineering organization, its services, teams, ownership, incidents, deployments, and agents, and serves that model identically to your UI, your APIs, your MCP clients, and every agent you run, with governance on what any of them may do. The agents are consumers of the platform rather than owners of the context.

Is Resolve AI an Agentic SDLC Platform?

No, and it does not claim to be. Resolve AI positions itself as AI agents for production operations, which is a solution built for one part of the lifecycle. Its model of your systems and its action controls are scoped to its own investigations, so it works as a layer inside a wider stack rather than as the platform underneath it.

What are the alternatives to Resolve AI?

The direct comparisons are other AI SRE and incident investigation products, including incident.io's agent, PagerDuty's SRE Agent, and AWS DevOps Agent, all of which Resolve compares itself against publicly. The different choice is an Agentic SDLC Platform such as Port, where incident work is one agentic workload running on a model and a governance layer you own.

Port vs Resolve AI, which should I pick?

They answer different questions. Pick Resolve AI if a deep incident investigation agent is what you need next and you are comfortable with its context and controls staying inside it. Pick Port if you will run several agents and want one model of your systems and one set of approvals covering all of them. Plenty of organizations run both.

When should I choose Resolve AI over an Agentic SDLC Platform?

When incidents are your most expensive problem right now, you need something working in weeks, you lack the platform engineering capacity to own a shared layer, or your telemetry cannot leave your network and Satellite solves that

Tags:
{{survey-buttons}}

Get your survey template today

By clicking this button, you agree to our Terms of Use and Privacy Policy
{{stay_tuned}}

Stay tuned for our upcoming tutorial

With a step-by-step guide walking you through how to implement and scale Anthropic’s playbook in Port’s free tier. Register to be notified once the guide is published:

By clicking this button, you agree to our Terms of Use and Privacy Policy
Thank you!You’ll be notified when the guide hits!
{{survey}}

Download your survey template today

By clicking this button, you agree to our Terms of Use and Privacy Policy
{{roadmap}}

Free Roadmap planner for Platform Engineering teams

  • Set Clear Goals for Your Portal

  • Define Features and Milestones

  • Stay Aligned and Keep Moving Forward

{{rfp}}

Free RFP template for Internal Developer Portal

Creating an RFP for an internal developer portal doesn’t have to be complex. Our template gives you a streamlined path to start strong and ensure you’re covering all the key details.

{{ai_jq}}

Leverage AI to generate optimized JQ commands

test them in real-time, and refine your approach instantly. This powerful tool lets you experiment, troubleshoot, and fine-tune your queries—taking your development workflow to the next level.

{{cta_1}}

Check out Port's pre-populated demo and see what it's all about.

Check live demo

No email required

{{cta_webinar_aug_18}}

LIVE WEBINAR, Aug 18, 2026:

Context-aware Vibe Coding for Platform Engineering

{{cta_webinar_oct_22}}

Thursday, October 22 12:00pm EDT⋅6:00pm CET

To learn more about the new capability and to see a live demo, join our upcoming community session with GitHub

{{cta_explore_port}}

Move fast while staying in control

Build governed agentic workflows on one central platform.

{{public_demo}}

See it in action:

Watch this video on generating Terraform with Port, or explore our public demo.

{{cta_survey}}

Check out the 2025 State of Internal Developer Portals report

See the full report

No email required

{{cta_2}}

Minimize engineering chaos. Port serves as one central platform for all your needs.

Explore Port
{{cta_3}}

Act on every part of your SDLC in Port.

Schedule a demo
{{cta_4}}

Your team needs the right info at the right time. With Port's software catalog, they'll have it.

{{cta_5}}

Learn more about Port's agentic engineering platform

Read the launch blog

Let’s start
{{cta_6}}

Contact sales for a technical walkthrough of Port

Let’s start
{{cta_7}}

Every team is different. Port lets you design a developer experience that truly fits your org.

{{cta_8}}

As your org grows, so does complexity. Port scales your catalog, orchestration, and workflows seamlessly.

{{cta_n8n}}

Port × n8n Boost AI Workflows with Context, Guardrails, and Control

{{port_builders_session}}

Port Builders Session: A Single, Governed Interface for All MCP Servers

{{cta-demo}}
{{read_case}}
{{n8n-template-gallery}}

n8n + Port templates you can use today

walkthrough of ready-to-use workflows you can clone

Template gallery
{{from_manual_to_autonomous_engineering}}

From manual to autonomous engineering

One platform to build, govern, and operate the Agentic SDLC.

Explore Port
{{port_is_open_for_you_to_try_it}}

Port is open for you to try it

build your first agentic workflow today

Sign up
{{reading-box-backstage-vs-port}}
{{cta-backstage-docs-button}}

Starting with Port is simple, fast, and free.