Incident History

Incident Management unavailable in US Central

This incident has been resolved. Incident Management in US Central is operating as normal. Customers can create, view, and query Incidents, and the Incident public API is fully available. Grafana OnCall was not affected at any point during this incident.

1787654389 - 1787655779 Resolved

Cloud Logs read path outage on eu-west-2

We observed an issue which impacted the following service: reads of Cloud Logs in the eu-west-2 region.

Affected services would include Explore, dashboards, alert evaluation when querying Logs.

The time of impact lasted from approximately 22h15 and 22h23 UTC, on loki-prod-012. This incident has since been resolved.

1787093607 - 1787093607 Resolved

K6 Test Outage

This incident has been resolved.

1786634703 - 1786639638 Resolved

Metrics: Elevated Error Rates Reads/Writes

Between 12:20 and 13:15 UTC, we experienced elevated error rates for metrics reads and writes due to a networking issue. Impact was limited to a subset of deployments in the prod-us-central-0 region. A fix has been applied and error rates have since recovered.

1786368609 - 1786368609 Resolved

Alerting expressions pipeline failing when recovery settings

Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.

1786218488 - 1786218488 Resolved

Some Cloud Test Runs Terminated

Between approximately 13:40 and 15:05 UTC today, a subset of cloud test runs were unexpectedly terminated and marked Aborted (by system). This affected runs that were in progress during three short windows in that period; test results and metrics for completed runs were not impacted.

We have identified the cause and no further occurrences have been observed since 15:05 UTC. A fix is in place. The platform is currently operating normally and test runs are executing as expected

1785952635 - 1785952635 Resolved

o11y requests too large for nods in the prod-eu-west-2 region.

This incident has been resolved.

1785930095 - 1785930275 Resolved

Some Grafana Instances Unavailable

We are observing a continued period of stability and at this point, we are marking the incident resolved.

1785870669 - 1785893521 Resolved

Logs latency increase within prod-eu-west-3

This incident has been resolved.

1785839653 - 1785851340 Resolved

K6 - Cloud test-run issues

This incident has been resolved.

1785569138 - 1785573097 Resolved
⮜ Previous