Incident History

AWS Service Disruption

This incident has been resolved.

1784891552 - 1784899132 Resolved

MQTT service disruption preventing users from starting prints (recurring)

Summary

On 2026-07-20, our MQTT device-connectivity service experienced two connection disruptions originating from the same root cause. A single node in the service cluster developed intermittent network failures, which disrupted communication with the rest of the cluster and led to a backlog in connection handling. As a result, some devices were disconnected, reconnections were delayed, and the number of online connections dropped.

The first phase fully recovered by 16:00 UTC. The issue recurred around 17:00 and was fully resolved by 18:30, after the faulty node was removed and replaced.

Customer impact

Root cause

A single cluster node developed intermittent network failures, disrupting its communication with the other nodes in the cluster. This blocked data synchronization across the cluster and led to a backlog in connection handling. The issue recurred because the faulty node remained in the cluster; it was fully resolved once the node was removed and replaced (17:53).

Why it happened twice

The first round of mitigation cleared the accumulated backlog but did not eliminate the underlying network fault on the node itself, so the cluster appeared healthy after 16:00. When the same node faulted again around 17:00, the identical failure pattern was re-triggered. Restarting nodes at 17:27 released the backlog a second time, but with the faulty node still in the cluster the problem returned around 17:47. Isolating the fault source by replacing the node was what ultimately brought the incident to an end. These were therefore two recurrences of the same intermittent fault, not two unrelated incidents.

Follow-up actions

  1. Strengthen monitoring, alerting, and automatic isolation for inter-node communication anomalies;
  2. Improve timeout handling, graceful degradation, and backlog protection during abnormal conditions;
  3. Add dedicated failure drills covering intermittent network faults, partially unreachable nodes, reconnection surges, and faulty-node replacement;
  4. Over the coming months, roll out multi-region deployment of the service to reduce the blast radius of any single point of failure.
1784568432 - 1784602728 Resolved

MQTT service disruption preventing users from starting prints

This incident has been resolved.

1784562091 - 1784565752 Resolved

ISP Network Disruption in Parts of Europe Impacting Cloud Service Access

This incident has been resolved.

1782980733 - 1782993057 Resolved

Europe & South Africa → AWS Network Issues

The issue has been fully resolved. It has been confirmed as a network configuration issue on the carrier side. All affected regions have returned to normal. We apologize for any inconvenience caused.

1779176448 - 1779191662 Resolved

Network Issues Affecting Bambu Lab Cloud Services in Parts of Europe

This incident has been resolved.

1778421767 - 1778465242 Resolved

Network Outage

The incident has been resolved.

1764925189 - 1764928058 Resolved

Partial Connectivity Issues for Users in Some Regions

This incident has been resolved.

1763467315 - 1763515047 Resolved

Partial Service Outage

Service Disruption Caused by ElastiCache Automated Certificate Rotation

Summary

Based on our initial analysis with AWS, this incident was caused by an automated certificate update for the ElastiCache middleware.

Timeline

At 02:58 AM on October 29 (Beijing Time), AWS initiated an automated certificate update for our ElastiCache instances. During this process, the primary and replica nodes of the ElastiCache cluster experienced issues, preventing backend services from accessing the component.

Next Steps & Action Items

We have raised two critical issues with AWS Support:

AWS Support has escalated these issues to their internal engineering team for a detailed root cause analysis. We will provide further updates as soon as we receive more information from AWS.

Updated on November 13, 2025

1761679745 - 1761686678 Resolved

Network Connectivity Issues Resolved and Mitigated

Between 08:02 and 11:52 UTC+8, some users experienced intermittent issues accessing our cloud services.

After a joint investigation with our cloud provider, AWS, we have confirmed the root cause was network instability from the carrier, Cogent. Access requests routed through the Cogent network were subject to timeouts and packet loss.

Due to several recent incidents involving this provider, AWS has proactively rerouted traffic away from Cogent to alternative network paths. This action significantly mitigates the risk of similar disruptions in the future.

Cogent Network Status: https://ecogent.cogentco.com/network-status

1759476975 - 1759476975 Resolved
⮜ Previous