Between 12:20 and 13:15 UTC, we experienced elevated error rates for metrics reads and writes due to a networking issue. Impact was limited to a subset of deployments in the prod-us-central-0 region. A fix has been applied and error rates have since recovered.
Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.
Between approximately 13:40 and 15:05 UTC today, a subset of cloud test runs were unexpectedly terminated and marked Aborted (by system). This affected runs that were in progress during three short windows in that period; test results and metrics for completed runs were not impacted.
We have identified the cause and no further occurrences have been observed since 15:05 UTC. A fix is in place. The platform is currently operating normally and test runs are executing as expected
We are investigating a partial read outage affecting Loki in prod-us-east-4. Between 19:41 UTC and 19:50 UTC, a significant portion of read queries may have failed or returned errors.
The issue has been identified and service has been restored. We are continuing to investigate the underlying cause and will provide additional information as it becomes available.