Investigating issues in US Central (prod-us-central-0, prod-us-central-5)
This incident has been resolved.
This incident has been resolved.
This incident has been resolved.
This incident has been resolved.
This incident has been resolved.
Mimir writes in the prod-us-central-0 region had elevated error rates for approximately 15 minutes from 1:05 to 1:20 UTC. The issue has been resolved and we are monitoring.
This incident has been resolved. Incident Management in US Central is operating as normal. Customers can create, view, and query Incidents, and the Incident public API is fully available. Grafana OnCall was not affected at any point during this incident.
We observed an issue which impacted the following service: reads of Cloud Logs in the eu-west-2 region.
Affected services would include Explore, dashboards, alert evaluation when querying Logs.
The time of impact lasted from approximately 22h15 and 22h23 UTC, on loki-prod-012. This incident has since been resolved.
This incident has been resolved.
Between 12:20 and 13:15 UTC, we experienced elevated error rates for metrics reads and writes due to a networking issue. Impact was limited to a subset of deployments in the prod-us-central-0 region. A fix has been applied and error rates have since recovered.
Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.