🚨 Service Incident ​
Definition ​
A Service Incident is an unplanned event that disrupts or degrades the quality of a service.
How to define it properly ​
- Describe the user or system impact
- Classify severity (Critical, High, Medium, Low)
- Include:
- Start time and detection time
- Affected services or workloads
- Document:
- Root cause (when identified)
- Mitigation actions
- Resolution steps
- Link to related Workloads and PRDs if relevant
- Provide a post-mortem with:
- Learnings
- Preventive actions
Get an overview of service incidents ​
The incidents view offers both a list and a calendar layout, with severity and status at a glance. Use it to track ongoing and past incidents over time.

Declare a service incident ​
Use the Add button to open the incident form. Capture the impact, severity, start date, and affected workloads, then save to start tracking it.

Update a service incident ​
Open an incident to follow up: update its status, document the root cause, mitigation, and resolution, and capture the post-mortem.
