OpsSwarmIncident Simulation Lab
GitHub checking
API checking
INCIDENT SIMULATION CONTROL PLANE

Operations overview

Inject controlled faults, observe service degradation, collect evidence, and verify recovery from one console.

Monitored services◉
— Waiting for API
Healthy now✓
— Live health checks
Active incidents!
— Faults awaiting recovery
Evidence runs≡
— Recorded simulation runs
LIVE TOPOLOGY

Service health

ACTIVITY

Latest runs

Loading run history…
DEMO WORKFLOW

Incident lifecycle

Live labExternal orchestration
01Fault injectionIncident Simulation Lab
→
02Monitoring eventPrometheus / Alertmanager
→
03GitHub IssueSystem of record
→
04InvestigationOpsSwarm agents
→
05Recovery policyAUTO / HUMAN / DENY
→
06VerificationPASS / FAIL evidence
EXPERIMENT CATALOG

Incident scenarios

Repeatable failure contracts for service, database, resource, network, and telemetry conditions.

⌕
CONTROLLED INJECTION

Fault runner

Select a scenario and keep the fault active until OpsSwarm or an operator performs recovery. Timed auto-reset is optional for short demos.

1 Select scenario
2 Configure run
RUNTIME TELEMETRY

Services

Current health and metrics from the five DemoMart simulator services.

TRACEABILITY

Runs & evidence

Inspect fault state, timestamps, GitHub linkage, and recovery evidence for each simulation run.

⌕
RunScenarioServiceFaultStateCreatedGitHub
Loading evidence…
CONNECTIONS

Monitoring & integrations

Runtime endpoints that feed the demo incident lifecycle.

P

Prometheus

Metrics scraping and alert rule evaluation.

localhost:9090
A

Alertmanager

Alert routing into the monitoring webhook contract.

localhost:9093
API

IncidentLab API

Fault injection, recovery, services, and evidence.

/docs
OS

OpsSwarm + OpenClaw

Checking orchestration runtime…

—
CHECKING
GH

GitHub Issues

Checking repository configuration…

—
CHECKING
MONITORING CONTRACT

Ingress path

POST /hooks/monitoring
Service metrics→Prometheus→Alertmanager→/hooks/monitoring→GitHub / OpsSwarm