ALAI Client Incident Response — Runbook
ALAI Client Incident Response — Runbook
Purpose:Version: What1.0
toCreated: do2026-07-28
whenOwner: John (AI Director) / Alem Basic (CEO)
Source process: processes/incident-management.md (full spec), processes/quick-ref/incident-response-quickref.md
Related: processes/templates/incident-report.md, processes/communication-protocol.md
1. Purpose
Client-facing runbook for handling any incident that affects a client-facingclient's incidentservice: happensproduction —outages, classify,bugs, respond,security communicate,incidents, escalate.SLA breaches. This is the condensed operational runbook;version — the full processgovernance document (RCAclassification methodology, post-mortem template,rationale, GDPR breachappendix, templates)KPIs, audit trail) lives in .~/ALAI/processes/incident-management.md For diagnosing ALAI's own infrastructure alerts (tunnel/daemon/Docker), see ~/system/docs/runbooks/incident-response-playbook.md.
Audience: John (primary), Alem (P1 escalation), FlowForge/CodeCraft/Securion (delegated fixes)
Last updated: 2026-07-28
Step 1 — Classify2. Severity ImmediatelyMatrix
| Severity | Response Time | Resolution Target | ||
|---|---|---|---|---|
| P1 CRITICAL | Service completely down, active data breach, revenue impact, legal/regulatory violation | 15 min | 4 hours | Alem + Client + Datatilsynet (if GDPR) |
| P2 HIGH | Major feature broken, client blocked, |
1 hour | 8 hours | John + Tech Lead + Client |
| P3 MEDIUM | Minor feature broken, workaround |
4 hours | 24 hours | John + relevant agent |
| P4 LOW | 24 hours | 5 days | Assigned agent |
When in doubt, escalate upup. — betterBetter to downgrade a P1 to P2 than to underestimate a critical incident.
P1
3. include:Response anyFlow
Detection service-> unavailableTriage & Classify -> Investigation -> Resolution -> Verification -> Communication -> Post-Mortem -> Closure
Step 1 — Detection
| Source | Action |
|---|---|
| Automated monitoring (Mission Control, health checks) | Auto-alert to |
| Client |
Logged via support-ticket.js, |
| Internal discovery ( |
Report immediately (Mattermost #ai-ops) |
| Third-party/vendor | Forward to John + Alem |
Step 2 — Triage & OpenClassify the Incident Record
Owner: (John, within theSLA response-response time window above.
- Acknowledge receipt to the reporter
(client email, Mattermost, monitoring alert). - Classify severity
(Stepusing1).the matrix in Section 2 - Assign
anowner—(agent/roleroutepertoincidentthe right specialist company (see~/system/agents/specialist-mapping.json; infra → FlowForge, security → Securion, backend bug → CodeCraft).type) - Create
theincidentrecordrecord:comms/incidents/INC-YYYY-MM-DD-NNN.mdfrom~/ALAI/processes/templates/OPERATIONS/incident-report.md, - Notify
formatstakeholdersINC-YYYY-MM-DD-NNN.per Escalation Matrix (Section 4) - Open
aMattermost thread in#incidents.
Step 3 — Investigation
- Reproduce the issue in the affected environment
- Gather evidence: logs, metrics, screenshots
- Root cause analysis (5 Whys / Fishbone / Timeline) — mandatory for P1/P2
- Assess impact: users, data, services affected
- Fix plan: immediate mitigation vs. long-term fix
Step 4 — Resolution
- Implement fix (code, config, infra, or process change)
- Validate: staging for P2-P4, production hotfix + 1h monitoring for P1
- Deploy per deployment protocol
- Monitor metrics/logs/error rates for 1 hour post-deploy
Step 5 — Verification
- Fix deployed to production
- Smoke tests passing
- No related errors in the last hour
- Monitoring shows normal metrics
- Client confirms resolution (if client-reported)
- Incident record updated to "Resolved"
Step 6 — Communication cadence
| Severity | Initial Update | Progress Updates | Resolution Notice |
|---|---|---|---|
| P1 | Within 15 min | Every 30 min | Immediately + 24h follow-up |
| P2 | Within 1 hour | Every 2 hours | Within 1h of resolution |
| P3 | Within 4 hours | Daily if >24h | Next business day OK |
| P4 | N/A unless client-reported | N/A | Next sprint review |
Step 7 — Post-Mortem (mandatory P1/P2, optional P3)
- Scheduled within 2 business days of resolution, conducted within 5
- Blameless, fact-based, action-oriented
- Every post-mortem produces at least one action item, tracked in Mission Control
- Template and full 10-section structure: see
processes/incident-management.md§7.3
Step 8 — Closure
- Incident resolved and verified
- Post-mortem completed (P1/P2)
- Action items created and assigned
- Documentation/runbooks updated
- Client notified (final resolution email)
- Incident record archived to
comms/incidents/archive/YYYY/
4. Escalation Matrix
P4: Developer -> Tech Lead (review) -> Archive
P3: Developer -> Tech Lead (notified) -> Archive
P2: Tech Lead (owns) -> John (notified) -> Client notified -> Archive
-> if SLA breach: Alem involved
P1: John coordinates -> Alem notified IMMEDIATELY -> Client notified IMMEDIATELY
-> Datatilsynet if GDPR breach (72h)
Escalate to Alem when: any P1; P2 affecting a major client or revenue; client relationship risk; legal/regulatory implication (GDPR, lawsuit threat); media/PR risk; financial impact >50,000 NOK; requires external assistance.
| Authority | When | Deadline | Contact |
|---|---|---|---|
| Datatilsynet | Personal data breach (GDPR Art. 33) | 72h from detection | [email protected] |
| Client | Breach of data processing agreement | 24h per DPA | Client contact per contract |
| Politiet (Kripos) | Criminal activity | Immediately | 02800, [email protected] |
| NSM | Critical infrastructure attack | Immediately | [email protected] |
5. Client Communication Templates
Initial notification (email, within SLA)
Subject: [Action Required / Informational] Service Incident Notification — [Client Name]
We are writing to inform you of a service incident affecting [affected service/feature].
Incident Summary:
- What: [impact — what users experience]
- When Detected: [date, time]
- Current Status: [Investigating / Mitigating / Resolved]
- Impact to Your Users: [specific — e.g. "Login unavailable"]
What We Are Doing: [response actions underway]
Next Steps: [what client should expect, workaround if any, next update time]
Estimated Resolution: [if known, or "Investigating"]
We will provide updates every [timeframe] until resolved.
Resolution confirmation
Subject: RESOLVED: Service Incident — [Client Name]
Incident Summary:
- Issue: [what happened]
- Root Cause: [brief, non-technical]
- Resolution: [what was fixed]
- Resolved At / Total Duration: [...]
Preventive Actions: [what we are doing to prevent recurrence]
Impact to Your Data: [confirm no data loss, or describe]
Full template set (internal declared/update/resolved, GDPR Art. 33/34 regulatory notices): processes/incident-management.md §6.
6. On-Call
- Primary: John (AI Director) — 24/7 via automated monitoring
- Human backup: Alem (CEO) — for P1 decisions requiring a human
- Business hours: Mon-Fri 08:00-18:00 CET; after-hours P1 = John responds immediately + alerts Alem by phone, P2 = John within 1h + email to Alem (reviewed within 4h)
- Contact priority: Mattermost #ai-ops -> email [email protected] -> phone (P1 only)
7. Tools
| Tool | Purpose | Command | ||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| support-ticket.js | Create incident record from client report | node ~/system/tools/support-ticket.js create --title "[Title]" --severity [P1-P4] |
||||||||||||||||||||||||||||||||
| Mission Control | Track incident as task | node ~/system/tools/mc.js add "[INC-
|
node | |||
IfMission JohnControl isDashboard: offline, all P0/P1 alerts route to Alem via email ([email protected]).http://localhost:3030
Step8. 7Hard — Post-Mortem & ClosureRules
P1/P2:Blameless post-mortemrequired, using the template in~/ALAI/processes/incident-management.md§7 (or~/ALAI/processes/templates/OPERATIONS/post-mortem.md). P3: optional.Closure requires: fix verified live, client notified (if applicable), post-mortem filed, lessons logged.Log the lesson:discover.js memoryentry or feedback memory if the root cause reveals a repeatable gap — don't let the same incident recur silently.
Hard Rules
Classify before you actmortems —no fix attempt without a severity label; the label sets the response clock.No claim without evidence— verify fixes with live curl/screenshot/enumeration, not a green build (ZAKON NULA).No deploy without ZAKON PI2's 6 checks, even mid-incident.P1 always reaches Alem— phone first. Never sitfocus onasystems/process,P1neverto "confirm root cause first."GDPR clock starts at detection— 72h to Datatilsynet, not at RCA completion.Client comms use the approved templates— don't freelance the wording on data-loss or breach notices; legal risk.individuals.- Every P1/P2 gets a post-mortem
—within 5 business days, noexceptions, no "fixed already, skip write-up."exceptions.
comms/incidents/.References9. Related Documents
Fullprocesses/incident-management.mdincident—processfull 13-section governance document (classificationseveritydetail,rationale, KPIs, RCAmethodology, all 8 communication templates,appendix, GDPRforms):breach assessment tree, quarterly review process)- processes/quick-ref/incident-response-quickref.md — 1-page field version
- processes/templates/incident-report.md
- ALAI-CLIENT-ONBOARDING.md
- ALAI-CLIENT-INVOICE-CYCLE.md
Document Location: ~/ALAI/processes/
incident-management.runbooks/ALAI-CLIENT-INCIDENT-RESPONSE.md
~/system/docs/runbooks/incident-response-playbook.md~/ALAI/processes/templates/OPERATIONS/incident-report.md~/ALAI/processes/templates/OPERATIONS/post-mortem.md~/ALAI/processes/templates/OPERATIONS/sla-report.md