SRE Agent
Investigates active PagerDuty incidents by retrieving incident, service and dependency data, querying connected log, metric and knowledge-base tools, and recommending diagnostic, remediation and workflow steps; it can add notes to the incident it is working on and, on a responder's click, run a recommended incident workflow.
Recorded characteristics
- Function
- PagerDuty describes the SRE Agent as "a continuously learning agent to help rapidly diagnose, troubleshoot, and remediate issues" that "works alongside you… ingesting event data, runbooks, and logs to build an understanding of the incident's scope and likely cause, then recommending remediation actions". Documented behaviour: it loads an incident summary and suggests next steps in the Operations Console, the incident details page, Slack, and Microsoft Teams (Early Access); answers responder questions about the incident; surfaces likely root causes; prioritises actions by urgency and impact; recommends diagnostic and remediation steps; recalls similar past incidents and past resolutions; generates and saves playbooks; and presents "nudges" (buttons) such as Upload Runbook, Update Runbook, Analyze Past Incidents, Analyze Related Incidents, Generate a Playbook, Check Change Events, Search Logs and Update Memory. PagerDuty lists the agent's own capabilities as the tools get_incident_details, list_incidents, add_incident_note, get_service_details and get_related_services — that is, four read operations and one write of a note to the current incident. Through connectors the agent can query logs, metrics, traces, code and knowledge-base tools in third-party systems. For response execution, PagerDuty documents that the agent "analyz[es] configured incident workflows and recommend[s] the most relevant one", presents a ranked recommendation with reasoning and a "Run recommended workflow" button, and that the responder clicks that button "to execute the automation"; each recommended workflow can run only once per incident and the agent confirms successful execution. As a virtual responder (Early Access) the agent can be attached to an incident workflow or to an escalation policy level so that it begins investigating automatically when an incident triggers, gathering context and posting findings to Slack or the PagerDuty web app "in parallel with your human responders — not in place of them".
- Data access
- PagerDuty states the agent analyses event and alert payload information, historical and related incidents, change events, and user-provided data (runbooks, logs, documentation), and notes as a current limitation that it "has limited access to incident timeline details, incident workflows, and alert grouping data". Its documented PagerDuty tools read incident details including recent notes, changes and status updates; incident lists filtered by service, team or user; service details; and upstream and downstream service dependencies. Responders can attach files (.txt, .pdf, .md, .jpg, .png; up to 25 files per conversation, 100 Kb each). The agent fetches runbook URLs from the event payload (runbook_url.confluence, runbook_url.github, runbook_url.servicenow in custom_details) and can fetch documents it judges relevant. Through administrator-configured connectors it reads logs, metrics or traces from Arize, AWS CloudWatch, Azure Monitor, Coralogix, Datadog, Dynatrace, Elasticsearch, GCP, Grafana, Honeycomb.io, Logz.io, New Relic, Observe Inc., Sentry, Splunk and Sumo Logic, knowledge bases from Confluence, GitHub and ServiceNow, and GitHub code, with each connector's individual tools selected by an administrator under "Agent May Use These Tools". It also analyses new notes posted during active incidents, and in Slack it can read history of channels associated with the incident or service where optional Slack scopes are granted. The agent maintains per-service memory artifacts (incident playbook or "scratchpad", customer service runbook, incident summarisation, service profile), which PagerDuty exposes for viewing, updating and redaction through an SRE Agent Memory API. PagerDuty states the agent analyses custom_details and notes only up to the first 2,000 characters of each.
- Actions
- Can take actions
- External actions
- Conditional
- Human confirmation
- Conditional
- Permission basis
- Mixed
- Administrative control
- Documented controls. Entitlement: PagerDuty Advance is an add-on on the Professional, Business and Enterprise plans; the SRE Agent in the Operations Console additionally requires PagerDuty AIOps, while Slack and incident-details access requires PagerDuty Advance only. Enablement: Account Owners and Global Admins can enable or disable AI agents at AI > AI Settings > Assistant and AI Agents Configuration, where each agent has its own toggle; other users see a "Request to Admin" button that emails admins. Setting up PagerDuty Advance requires the Admin, Global Admin or Account Owner base role. Access scope: an Access tab lets admins grant PagerDuty Advance per team, enforced per user on team membership across the web app, chat integrations and the API (all teams are enabled by default). Chat surfaces: Slack and Microsoft Teams integrations are toggled on or off, with separate Configure options for Proactive Incident Insights and Proactive Incident Summarization; the SRE Agent requires additional Slack scopes that a PagerDuty admin may need to reauthorise. Tooling: connectors are configured at AI Settings > SRE Agent Configuration > Connectors by users with Admin or Account Owner permissions, who supply the third-party credentials or API keys and tick the individual tools under "Agent May Use These Tools"; connectors show Active, Unhealthy or Not Connected. Skills (Early Access) let an administrator or user add custom instructions that tell the agent when to act and what steps to follow, and can be scoped to one user or to everyone. Data sources: a GitHub personal access token and an Amazon Q Business data accessor are each toggled on and configured in AI Settings. Proactive note messages from the agent can be disabled on the AI Settings page. Virtual responder behaviour is controlled by an incident workflow trigger (conditional or manual) or by a per-escalation-policy "Trigger SRE Agent for incidents in this escalation policy" toggle, which requires a Manager or Admin role. Usage is metered in AI Actions, viewable in PagerDuty Advance Analytics and AI > Usage Analytics.
- Default state
- Disabled
- Availability
- PagerDuty Advance is documented as General Access, with certain features in Early Access as noted in the PagerDuty Advance AI Disclosure. The SRE Agent is documented in the current PagerDuty knowledge base under AI > PagerDuty Advance alongside the Scribe, Shift and Insights agents (SRE Agent page updated 20 August 2026). Generally available surfaces are the Operations Console (requires AIOps and Advance), the incident details page and Slack (require Advance). Documented as Early Access: the SRE Agent in Microsoft Teams, the SRE Agent virtual responder via incident workflow, triggering the SRE Agent via escalation policy, and Skills. Where an account does not have AIOps, PagerDuty states certain AIOps information (related incidents, past incidents, change events, outlier incidents) is still reachable, but only inside the agent chat interface in Slack. PagerDuty Advance for Automation Digest, an adjacent feature, is documented as available in the US and EU service regions; no capability-specific regional restriction is documented for the SRE Agent itself.
- Licensing
- PagerDuty Advance is sold as an add-on to the Professional, Business and Enterprise pricing plans, or through one-time AI Actions; PagerDuty states that if an account has neither Advance nor AIOps it begins a trial to give access to the SRE Agent. Usage is metered: each request submitted to the SRE Agent through Slack or the Operations Console, each nudge-button click, and each trigger of the agent by an incident workflow or escalation policy consumes four AI Actions. PagerDuty states the number of AI Actions consumed by action type may change at its sole discretion.
- External model or provider
- Not established at capability level. PagerDuty states that "PagerDuty Advance generates outputs with assistance from certain LLM providers" and directs readers to the PagerDuty Advance AI Disclosure, obtained from PagerDuty's Assurance Profile, for details of any third-party models used; no model or provider is named in the public documentation for the SRE Agent. PagerDuty states it does not use customer data to train the models used for PagerDuty Advance.
- Limitations and uncertainty
- Not established: the model or provider used; whether any PagerDuty incident-state change other than adding a note (for example acknowledging, reassigning, escalating, changing urgency or resolving) is performed by the agent itself — PagerDuty's list of SRE Agent capabilities contains only add_incident_note as a write, while a summary table and a feature list state the agent "marks the incident resolved" and "summarize[s] conversations and inputs, mark[s] resolved", and PagerDuty does not reconcile these statements; whether the agent can invoke PagerDuty Automation Actions or Runbook Automation jobs (no such invocation is documented — the documented Automation relationship is PagerDuty Advance for Automation Digest, which summarises Automation Actions Log output, and AI Generated Runbooks in Rundeck, an Early Access feature outside this capability); the operational detail, approval requirements and failure handling of the GitHub "Code (Revert PR)" tool, which PagerDuty lists among the connector tools an administrator can permit the agent to use but does not otherwise document; whether the agent can run a workflow without a responder click when acting as a virtual responder; retention periods for agent memory artifacts and conversations; regional processing for this capability; and whether execution history is immutable. PagerDuty states generative AI output can be misleading or false and that output must be fact-checked. Boundaries: the PagerDuty Advance Assistant, the Scribe, Shift and Insights agents, PagerDuty AIOps noise reduction, correlation and probable-origin features, PagerDuty Automation Actions, PagerDuty Runbook Automation, Incident Workflows themselves and the PagerDuty MCP Server are separate and are not recorded here.
Evidence
- SRE Agent
Supports: Function · Actions · Data access · Human confirmation · Availability · Licensing · Admin controls · Limitations · General · Primary source
"A continuously learning agent to help rapidly diagnose, troubleshoot, and remediate issues"; it ingests event data, runbooks and logs "then recommending remediation actions". Key features include surfacing likely root causes, recommending diagnostic and remediation steps, recalling similar incidents, generating and saving playbooks, and prioritising actions by urgency and impact.
"SRE Agent Capabilities": get_incident_details, list_incidents, add_incident_note ("Add a note to the current incident"), get_service_details, get_related_services. Supported actions are presented as nudges: Upload Runbook, Update Runbook, Analyze Past Incidents, Analyze Related Incidents, Generate a Playbook, Check Change Events, Search Logs, Update Memory.
FAQ: the agent analyses event and alert payload information, historical and related incidents, change events, and user-provided data (runbooks, logs, documentation), with "limited access to incident timeline details, incident workflows, and alert grouping data". It analyses new notes during active incidents and reads custom_details and notes only to the first 2,000 characters of each.
Responders drive the agent by asking questions or clicking nudges; the documented wrap-up behaviour requires incidents to be resolved before memory is written, with an Update Memory button for later additions.
Operations Console access requires AIOps and PagerDuty Advance; Slack and incident-details access require PagerDuty Advance. MS Teams and the virtual responder are Early Access.
PagerDuty Advance is available through one-time AI Actions or as an add-on with Enterprise, Business and Professional plans; the SRE Agent consumes four AI Actions per request or nudge click.
Proactive note messages can be disabled on the AI Settings page; the SRE Agent requires additional Slack scopes that a PagerDuty admin may need to reauthorise.
Memory artifacts are scoped per service and exposed through an SRE Agent Memory API for viewing, updating and redaction "at human speed, not machine speed"; a summary table states the agent "marks the incident resolved", which the documented tool list does not include.
Surfaces: Operations Console, incident details page, Slack, and MS Teams (Early Access); file uploads limited to .txt, .pdf, .md, .jpg, .png, 25 files per conversation at 100 Kb each.
- SRE Agent Recommended Workflows
Supports: Actions · Human confirmation · Admin controls · General · Primary source
"Review the agent's logic and click the Run recommended workflow prompt to execute the automation." The agent evaluates all configured incident workflows and returns a ranked recommendation with reasoning.
Workflow execution is initiated by the responder clicking "Run recommended workflow"; recommendations only appear where workflows are configured and the agent identifies a high-confidence match.
Execution limit: each recommended workflow can run only once per incident, with a re-run notification if attempted again.
"Timeline Audit: The system automatically logs all workflow executions in the Timeline tab of the Incident Details page for auditing and visibility", and a confirmation message verifies successful execution.
- Connectors, Tools, and Skills
Supports: Data access · External actions · Permission basis · Admin controls · Primary source
Connector table: logs, metrics, traces, knowledge base and code tools across Arize, AWS CloudWatch, Azure Monitor, Confluence, Coralogix, Datadog, Dynatrace, Elasticsearch, GCP, GitHub, Grafana, Honeycomb.io, Logz.io, New Relic, Observe Inc., Sentry, ServiceNow, Splunk and Sumo Logic, by API or MCP.
GitHub connector tools include "Code (Revert PR)" alongside Knowledge Base and Code; connector setup instructs the administrator to check the tools to enable "Under Agent May Use These Tools", and "Once a connector is active, the agent can use its associated tools… automatically or via user request". No further operational detail for the Revert PR tool is published.
"You have Admin or Account Owner permissions in PagerDuty" and "You have credentials or API keys for the third-party integrations you plan to connect" are prerequisites; each connector holds its own stored connection.
Connectors are managed at AI Settings > SRE Agent Configuration > Connectors with per-tool checkboxes and Active / Unhealthy / Not Connected statuses; Skills (Early Access) supply custom instructions with when-to-use conditions, execution steps, error handling and success criteria, scoped to one user or everyone.
- Add SRE Agent as Virtual Responder
Supports: Default state · Actions · Availability · Primary source
The virtual responder is enabled through existing Incident Workflows using two PagerDuty templates; once configured, "the SRE Agent activates automatically when an incident triggers—no manual invocation required".
As virtual responder the agent reviews past incidents and notes, gathers information from configured integrations and connectors, and surfaces recommended workflows; it posts findings to a dedicated Slack channel or the PagerDuty web app.
Marked Early Access.
- Trigger SRE Agent via Escalation Policy
Supports: Human confirmation · Admin controls · Availability · Primary source
"Add the SRE Agent to an escalation level to put it on call as a virtual responder, in parallel with your human responders — not in place of them"; it investigates, recommends a workflow with reasoning, and neither party blocks the other.
A per-escalation-policy toggle, "Trigger SRE Agent for incidents in this escalation policy", requires a Manager or Admin role; output destinations (PagerDuty side panel, Slack) are selectable.
Marked Early Access.
- PagerDuty Advance
Supports: Permission basis · Default state · Admin controls · External model · Licensing · Availability · Limitations · Primary source
Account Owners and Global Admins can enable or disable AI agents; setting up PagerDuty Advance requires the Admin, Global Admin or Account Owner base role; team-level access "is enforced per user, based on PagerDuty team membership… This applies to every access method: the web app, chat integrations (Slack and Microsoft Teams), and the API."
AI agents are turned on by toggling each agent to Enabled at AI > AI Settings > Assistant and AI Agents Configuration; users without permission click "Request to Admin". Data sources (GitHub token, Amazon Q Business) are separately toggled on.
Chat integrations are toggled per app with Configure options for Proactive Incident Insights and Proactive Incident Summarization; an Access tab grants PagerDuty Advance per team (all teams enabled by default).
"PagerDuty Advance generates outputs with assistance from certain LLM providers", with details in the PagerDuty Advance AI Disclosure; "PagerDuty does not use your data to train models used for PagerDuty Advance."
PagerDuty Advance is an add-on on Professional, Business and Enterprise plans; AI Actions usage is visible in PagerDuty Advance Analytics and AI > Usage Analytics.
PagerDuty Advance is General Access, with certain features in Early Access as noted in the AI Disclosure.
"Generative AI is a predictive technology, and sometimes the information it creates is misleading or false. You must fact-check the output of generative AI before you use it."
- PagerDuty Advance User Guide
Supports: General · Primary source
Boundary: PagerDuty Advance also includes the PagerDuty Advance Assistant, the Amazon Q integration, post-incident reviews and status updates, and the Insights, Shift and Scribe agents; none of these are recorded as this capability.
- PagerDuty Automation Actions
Supports: General · Primary source
Boundary: PagerDuty Advance for Automation Digest "synthesizes log output generated by Automation Action invocations" into a digest; no AI invocation of Automation Actions is documented.
- Incident Workflows
Supports: General · Primary source
Boundary: incident workflows are administrator-configured deterministic automation with conditional or manual triggers; the SRE Agent recommends one and a responder runs it.
- SRE Agent (PagerDuty documentation source)
Supports: Function · Primary source
Monitored PagerDuty-published source of the SRE Agent page, used for weekly change detection.