Co-work with agents and engineers to resolve incidents
Agent Teams investigate incidents in parallel. Engineers steer and remediate through Workbench
Specialized agents for harder investigations
A team of domain-specialized agents investigates in parallel and verifies findings against production evidence, the way a war room of senior engineers would
pgdb-orders-instance-3 — migrations 0335 and 0337 were silently skipped weeks ago.event_outcome column) was merged 2026-04-24, but Drizzle's migrator skipped it due to timestamp ordering (PR #27504).column "event_outcome" does not exist and relation "order_doc_state" does not exist errors across 5 services.Work alongside your agents in Workbench
Interrogate every finding, evidence, or theory by just interacting with the report
What happened
Alert fired at Wed May 6 1:19am for PostgreSQL High Rollback Rate on pgdb-orders-instance-3 (database: orders, cluster: orders-db-cluster). Rollback ratio was ~2.2% at alert time, with peaks up to 4.3%. Alert cleared at Wed May 6 1:29am (~10 minutes after firing).
Chronic schema drift on orders-db-cluster — missing event_outcome column and order_doc_state table. The order_doc_state escalation was the acute trigger that pushed rollback ratio above threshold.
Errors stopped at ~Wed May 6 7:33am, indicating remediation (migrations) was applied.
Root cause
Missing database migrations on orders-db-cluster
Drizzle migrations 0335_event_outcome_column.sql and 0337_quick_argent.sql were never applied to the orders-db-cluster orders database. Application code deployed 8–13 days earlier references schema objects that don't exist, causing every transaction touching them to roll back.
Causal chain:
Investigate with your team, agents, or on your own
Pursue parallel hypotheses, redirect agents, and add context as new evidence emerges
Main Investigation
Follow the agent's reasoning and steer it directly
Search for 5xx HTTP errors during the deadlock window
Verify PostgreSQL deadlock issue resolution status
What triggered the scaling event on October 23?
What happened
PostgreSQL deadlock alert fired on Thu October 23, 2025 4:55pm for the orders database on pgdb-orders-instance-2. The order-events-ingest service's OrderReconciler experienced 25 deadlocks over 14 minutes when attempting to batch update order events with resolution timestamps.
Impact
- 100% error rate on
updateOrderEventAPI for 2 minutes - 456-second latency spike on
getOrderEventCountsAPI - 68% error rate on
spanRequestAPI - 12+ downstream services affected including investigation system and analytics APIs
- Self-resolved when load normalized; no data loss
Root cause
Confirmed (HIGH confidence): Lack of deterministic lock ordering in batchUpdateOrderEvents() method. The SQL VALUES clause processes order IDs in arbitrary iteration order, allowing concurrent requests to acquire row locks in different sequences. A scaling event deployed ~27 new pods (vs baseline 1–2), dramatically increasing concurrency and triggering circular lock dependencies.
Gets you to verified root cause
Every finding is backed by production evidence for you to verify or explore further
Missing schema migrations on orders-db causing transaction rollbacks
Two schema migrations — adding the event_outcome column and the order_doc_state table — were merged but never applied to orders-db. The migration runner is invoked manually rather than through the deploy pipeline, so the cluster was silently overlooked. When a morning deployment triggered a traffic surge, transactions hit the missing schema and the rollback ratio crossed the 2% alert threshold. Applying both migrations cleared the errors.
event_outcome column never applied to orders-dborder_doc_state table never applied to orders-dborders-db after dependent code shippedorder-fulfillment, checkout-router, inventory-sync, catalog-service, payments-apiRemediate from the same surface
Trigger commit reverts, GitHub Actions, and alert silencing without leaving the context or interface
Revert: disable checkout-v2-routing in production
Revert recent enablement of enableCheckoutV2Routing, identified as the trigger for elevated p95 latency on checkout-router. Restores the previous routing path.
This change:
- Sets
enableCheckoutV2Routing: falseinhelm/values/production/values.yaml
Used and loved by engineers
Removing the toil of investigations, war rooms, and on-call.
“Resolve AI allowed us to move from hours to minutes for investigations in many incidents. We pull fewer engineers into war rooms, on-call is materially better, and that translates directly to advertiser trust and revenue protection for a billion-dollar ads business.”

“Resolve AI proved it could deliver real results in a constrained environment. It identified dependencies, surfaced accurate root causes 72% faster than our teams, all while integrating cleanly into our existing stack.”

“Resolve AI has changed how our teams work through production incidents. What used to take hours of manual investigation and coordination across teams now gets resolved in a fraction of the time. Our engineers aren't only faster, they're focused on the work that actually drives impact.”

“We’ve seen the value of AI in development, and now we’re applying that same approach to production. We started by partnering with Resolve AI for alert triage, incident investigation, and root cause analysis. We’ve seen positive signs of improvement in mean time to resolve for our critical incidents, and the North Star is self-healing systems.”

“What excites me most about Resolve AI's background agents is that I’m no longer starting from zero. The alerts are already investigated. The deployment summaries are already written. The findings are verified and the next steps are waiting for me. A lot of the operational work I used to handle manually is now happening continuously in the background with my oversight. I’m still making the important calls, but I can operate at a scale that just wasn’t possible before.”

“Resolve AI feels like a teammate who’s already done half the work. It tells me immediately if something’s critical or can wait, saving countless hours and frustration.”

“Resolve AI helps my team navigate incidents by correlating signals across logs, metrics, traces, and code automatically. Instead of switching between multiple platforms hunting for clues, we get immediate context. At our deployment velocity, that speed makes all the difference.”

“Incident response at our scale isn't about collecting more signals. It's about understanding why something is failing, quickly enough to limit customer impact.”

Recent updates.
- May 2026
Agent Teams
Specialized agents investigating in parallel with verified findings.
- May 2026
Workbench
Shared workspace with real-time visibility and engineer steering.
- May 2026
Closed-loop actions
Commit reverts, GitHub Actions, and alert silencing from investigation findings.
- April 2026
Adaptive knowledge
Every investigation makes the platform smarter.
Frequently asked questions
Single agents investigate sequentially and commit to one path early. Agent Teams pursue multiple hypotheses in parallel across different domains, with a verifier checking findings against production evidence before delivery.
Yes. Workbench shows what each agent is doing in real time. You can redirect agents, add context, or pursue a side hypothesis without interrupting the main investigation.
You take action from the same surface. Trigger a commit revert, execute a GitHub Actions workflow, or silence related alerts. Engineers approve the action. No context switching.
A dedicated reviewer validates findings against production data: timing, deployment history, metric patterns, service dependencies. It checks whether the evidence actually supports the conclusion, not just whether the conclusion sounds plausible.
It depends on the complexity. Resolve dynamically assembles the right team for the problem. A straightforward issue might need fewer agents. A cross-domain incident spanning multiple services will have more investigators working in parallel.
Yes. Resolve integrates with PagerDuty for incident lifecycle, Slack for communication, and your observability stack (Datadog, Grafana, Prometheus) for evidence. Findings and actions flow through the tools your team already uses.


