Incident History

Investigating issues in US Central (prod-us-central-0, prod-us-central-5)

We have identified the cause as an issue in our cloud provider’s US Central region. A mitigation is in place and recovery is in progress. Customers in prod-us-central-0 and prod-us-central-5 may still see errors or delays with metrics, logs, or Grafana while services come back. We will update this page as recovery continues.

1788276354 Ongoing

Partial Logs Write Outage

This incident has been resolved.

1787937162 - 1787942529 Resolved

Some Grafana UI features may be unavailable or reverting to legacy behaviour

This incident has been resolved.

1787878683 - 1787884812 Resolved

Elevated error rates affecting metrics writes in prod-us-central-0

This incident has been resolved.

1787832597 - 1787835995 Resolved

Mimir Writes Incident in prod-us-central-0

Mimir writes in the prod-us-central-0 region had elevated error rates for approximately 15 minutes from 1:05 to 1:20 UTC. The issue has been resolved and we are monitoring.

1787795648 - 1787795648 Resolved

Incident Management unavailable in US Central

This incident has been resolved. Incident Management in US Central is operating as normal. Customers can create, view, and query Incidents, and the Incident public API is fully available. Grafana OnCall was not affected at any point during this incident.

1787654389 - 1787655779 Resolved

Cloud Logs read path outage on eu-west-2

We observed an issue which impacted the following service: reads of Cloud Logs in the eu-west-2 region.

Affected services would include Explore, dashboards, alert evaluation when querying Logs.

The time of impact lasted from approximately 22h15 and 22h23 UTC, on loki-prod-012. This incident has since been resolved.

1787093607 - 1787093607 Resolved

K6 Test Outage

This incident has been resolved.

1786634703 - 1786639638 Resolved

Metrics: Elevated Error Rates Reads/Writes

Between 12:20 and 13:15 UTC, we experienced elevated error rates for metrics reads and writes due to a networking issue. Impact was limited to a subset of deployments in the prod-us-central-0 region. A fix has been applied and error rates have since recovered.

1786368609 - 1786368609 Resolved

Alerting expressions pipeline failing when recovery settings

Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.

1786218488 - 1786218488 Resolved
⮜ Previous