Incident History

Grafana rulers crash-looping on prometheus

This incident has been resolved.

1783513036 - 1783531416 Resolved

Cannot access dashboard set up as a home page

We have found that the issues are related to the grafana.unifiedHomepage feature rollout active 16:50 UTC yesterday, to 10:20 UTC today. We have rolled this back and systems are now working as expected.

1783506507 - 1783508443 Resolved

Delayed usage and billing data

We identified an issue with the internal job that calculates month-to-date usage and cost data, which caused usage attribution and billing dashboards to display stale information. The root cause has been identified and resolved.

Your billing dashboard may show a sudden jump in usage. This is expected. It reflects several days of accumulated usage that hadn't been showing up while the issue was ongoing, not a sudden change in your actual usage.

This issue does not affect your end-of-month invoice. Your bill will be calculated based on actual usage, not the numbers displayed during this period.

1783431847 - 1783431847 Resolved

Grafana Cloud IRM alert groups failing to produce alerts in us-east-3 region.

We haven't noticed any further issues in this region for alert group processing since yesterday. This incident i fully resolved.

1783251461 - 1783323425 Resolved

Partial Write Outage for Grafana Cloud Logs in prod-eu-north-0.

Grafana Cloud Logs in prod-eu-north-0 experienced a 10-minute partial write outage between 13:45 and 13:54 UTC. Impacted users may have experienced 5xx errors during this time.

1783088865 - 1783088865 Resolved

Some Queries Failing

This incident has been resolved. Thank you for your patience.

1783015830 - 1783028148 Resolved

Loki and Frontend Observability - Major Outage in prod-us-central-0 region

This incident has been resolved by restarting the affected services.

1782968152 - 1782974984 Resolved

Elevated Loki Query Bytes Reporting

We’ve implemented a fix and can confirm the issue is fully resolved as of 20:25 UTC.

Thank you for your patience.

1782934998 - 1782943244 Resolved

Confluent API Outage

This incident has been resolved.

1782736649 - 1782743233 Resolved

Mimir read errors and high latency in prod-eu-west-0

Since the mitigation has been applied, we have not seen the errors return. At this point, we are considering the incident resolved.

1782733060 - 1782807653 Resolved
⮜ Previous Next ⮞