Skip to content

🚨 Service Incident ​

Definition ​

A Service Incident is an unplanned event that disrupts or degrades the quality of a service.

How to define it properly ​

  • Describe the user or system impact
  • Classify severity (Critical, High, Medium, Low)
  • Include:
    • Start time and detection time
    • Affected services or workloads
  • Document:
    • Root cause (when identified)
    • Mitigation actions
    • Resolution steps
  • Link to related Workloads and PRDs if relevant
  • Provide a post-mortem with:
    • Learnings
    • Preventive actions

Get an overview of service incidents ​

The incidents view offers both a list and a calendar layout, with severity and status at a glance. Use it to track ongoing and past incidents over time.

service-incidents

Declare a service incident ​

Use the Add button to open the incident form. Capture the impact, severity, start date, and affected workloads, then save to start tracking it.

create-service-incident

Update a service incident ​

Open an incident to follow up: update its status, document the root cause, mitigation, and resolution, and capture the post-mortem.

update-service-incident

Made from Lyon - France with ❤️