“Global Outage Hits Microsoft Azure Due to DDoS Attack”

On July 30, 2024, Microsoft faced a major disruption affecting its Azure cloud services and Microsoft 365 suite. This outage, which persisted for nearly 10 hours, was instigated by a Distributed Denial-of-Service (DDoS) attack and had global repercussions.

The issue started around 11:45 UTC and was resolved by 19:43 UTC. Throughout this period, users encountered problems with various services, including Azure App Services, Application Insights, Azure IoT Central, Azure Log Search Alerts, Azure Policy, the Azure portal, and several Microsoft 365 and Microsoft Purview tools.

Microsoft identified the cause as a DDoS attack, which resulted in a sudden surge in traffic that overwhelmed Azure Front Door (AFD) and Azure Content Delivery Network (CDN). This led to intermittent errors, timeouts, and increased latency.

Adding to the problem, a flaw in Microsoft’s defense mechanisms exacerbated the situation. The company explained, “The initial incident was due to a Distributed Denial-of-Service (DDoS) attack, but early findings indicate that a deficiency in our defensive measures amplified the attack’s effects rather than neutralizing them.”

To counter the issue, Microsoft adjusted its network configurations and executed failovers to alternative pathways. By 14:10 UTC, these preliminary measures had alleviated most of the impact, though some users still faced reduced availability until around 18:00 UTC.

Subsequently, Microsoft deployed an updated mitigation strategy, starting with regions in Asia Pacific and Europe before extending it to the Americas. By 19:43 UTC, failure rates had normalized to pre-incident levels, with full resolution confirmed by 20:48 UTC.

This outage follows a series of recent service disruptions for Microsoft. Just two weeks earlier, an update from CrowdStrike’s Falcon agent caused Windows virtual machines to experience BSOD errors. These repeated issues have sparked concerns about the resilience of cloud infrastructure and the risks associated with centralized services.

The impact was felt across various sectors, with businesses like Starbucks in the US forced to disable their mobile ordering system for several hours due to the Azure disruption.

Microsoft has pledged to conduct an internal review to gain insights into the incident. A Preliminary Post-Incident Review will be released within 72 hours, followed by a Final Post-Incident Review within 14 days, detailing further findings and lessons learned.

More Articles & Posts