Incident History

Test Run Browser Screenshot Upload Failing

Test run browser screenshot upload experienced failures from 13:12 to 14:51 UTC.

The issue has been resolved

1772562955 - 1772562955 Resolved

Grafana Cloud Logs - Write degradation in Azure Netherlands (eu-west-3)

This incident has been resolved.

1772539643 - 1772735479 Resolved

Write outage for logs in prod-eu-west-3

This incident has been resolved.

1772437055 - 1772466536 Resolved

Complete outage in prod-me-central-1

Following our ongoing communications regarding the complete outage in prod-me-central-1, we are now closing this incident. As noted in the latest AWS update, the Middle East (UAE) region (ME-CENTRAL-1) has suffered significant damage and restoration is expected to take several months. We strongly recommend all affected customers migrate workloads to an alternate Grafana Cloud region as soon as possible.

If you have not already done so, please follow the migration steps outlined in our previous updates:

  1. Create a Grafana Cloud stack in an alternate region
  2. Update clients to send telemetry to the new region, if using Grafana Alloy then you can use Fleet Management https://grafana.com/docs/grafana-cloud/send-data/fleet-management/introduction/
  3. If your instance remains available and you have not configured your dashboards as code, then you may be able to use grafanactl to migrate dashboards https://grafana.com/docs/grafana/latest/as-code/observability-as-code/grafana-cli/grafanacli-workflows/ https://grafana.github.io/grafanactl/

For further details, please refer to the AWS incident communication directly: https://health.aws.amazon.com/health/status Please reach out to our Support team if you need any assistance with the above - https://grafana.com/profile/org#support We will continue to monitor the situation and update the incident once circumstances change.

1772433809 - 1781690293 Resolved

Increased Latency for Small Subset of Customers

A recent rollout caused the AuthZ (RBAC) service to perform many redundant folder-tree fetches for each authorization check. For a small number of tenants in the prod-us-east-0 and prod-eu-west-2 regions with very large folder trees. This added a few milliseconds to every check, which increased request latency.

The approximate timeframe of the impact is:

2026-02-26 17:24:43 UTC to 2026-02-27 14:33:53 UTC.

This has now been resolved.

1772209516 - 1772209516 Resolved

Trace querying issue in all Tempo clusters

This incident has been resolved.

1772199975 - 1772235481 Resolved

Incorrect pipeline assignment after custom attributes are assigned

This incident has been resolved.

1772197067 - 1772205879 Resolved

Grafana Cloud Faro slowness of listing and uploading sourcemaps in all regions.

This incident has been resolved.

1772110806 - 1772160545 Resolved

Grafana Cloud Metrics - Intermittent Write Latency in prod-us-central, prod-us-central-5, and prod-eu-west-0

This incident is now resolved.

During the incident the Cloud Metrics platform experienced intermittent latency spikes communicating with a backend cloud service in the prod-us-central-0 and prod-us-central-5 regions. During the incident the internal CSP-facing issue was escalated to a P1. After determining the scope of the latency spikes was limited to only one availability zone, the team mitigated the situation by migrating all write traffic from to the single nearly unaffected availability zone.

As the CSP service team attempted to remedy the situation, the situation became worse and began affecting the previously unaffected zone. Given this, another mitigation path was needed. Changing the connection strategy employed by Cloud Metrics to a different method was deployed to all environments, stabilizing the write path once again as we found the different connection method was more reliable and not affected by these increases in latency.

We have migrated all tenants back to multi-zone write paths and are happy with and confident in the current method of connectivity to the backend cloud service, which is the one we migrated to during the course of the incident. We have no immediate plans to use the previous problematic connectivity method for the foreseeable future.

1772049255 - 1773771739 Resolved

Issues Loading Dashboards and Alert Folders in Hosted Grafana

This incident has been resolved.

1772041482 - 1772049066 Resolved
⮜ Previous Next ⮞