Operations overview
Inject controlled faults, observe service degradation, collect evidence, and verify recovery from one console.
Service health
Latest runs
Incident lifecycle
Incident scenarios
Repeatable failure contracts for service, database, resource, network, and telemetry conditions.
Fault runner
Select a scenario and keep the fault active until OpsSwarm or an operator performs recovery. Timed auto-reset is optional for short demos.
Services
Current health and metrics from the five DemoMart simulator services.
Runs & evidence
Inspect fault state, timestamps, GitHub linkage, and recovery evidence for each simulation run.
| Run | Scenario | Service | Fault | State | Created | GitHub | |
|---|---|---|---|---|---|---|---|
Loading evidence… | |||||||
Monitoring & integrations
Runtime endpoints that feed the demo incident lifecycle.
Prometheus
Metrics scraping and alert rule evaluation.
localhost:9090Alertmanager
Alert routing into the monitoring webhook contract.
localhost:9093IncidentLab API
Fault injection, recovery, services, and evidence.
/docsOpsSwarm + OpenClaw
Checking orchestration runtime…
—GitHub Issues
Checking repository configuration…
—Ingress path
POST /hooks/monitoring