Microsoft Outage: Automated Maintenance Bug Crippled Azure and Microsoft 365 Services
A recent widespread outage impacting **Microsoft Azure** and **Microsoft 365** services was traced back to a critical bug in an automated network maintenance system. The flaw mistakenly removed IP routes from more devices than intended, disrupting connectivity for numerous cloud services and users, particularly in the West US Azure region.

**Microsoft** has disclosed that a bug within its automated network maintenance request system was the root cause of a significant outage on Thursday, July 23. The incident led to the erroneous removal of IP routes from an excessive number of devices, severely impacting **Azure** and **Microsoft 365** services.
### Outage Details and Impact
The disruption commenced at 10:44 AM ET, primarily affecting customers accessing **Microsoft 365** services via network infrastructure linked to Microsoftβs West US Azure region. By 11:11 AM ET, **Downdetector** recorded 2,403 outage reports, a stark increase from its typical baseline of 29.
The most reported service disruptions included **SharePoint** (78% of complaints), **Excel** (11%), and the **Microsoft 365 Admin Center** (6%).
Microsoft tracked the incident under ID **MO1437424**, confirming a broad impact across various **Microsoft 365** services:
* **Microsoft OneDrive**: Intermittent access.
* **SharePoint Online**: Users encountered "Something went wrong" errors.
* **Microsoft Teams**: Degraded chat functionality, including image loading issues.
* **Microsoft 365 Admin Center**: Slow loading or complete unresponsiveness.
* **Power Automate**: Flows failed to load.
* **Copilot Chat**: Intermittent delays or failures for actions and queries.
* **Microsoft Loop**: Inability to open or load Loop pages.
Other affected services included **Fabric**, **Power BI**, **Power Apps**, **Copilot Studio**, **Windows 365**, and **Microsoft Defender**. Some **Microsoft Defender** customers experienced delays in responses from **Microsoft Defender Experts**, with investigations and remediation actions via **Threat Explorer** and **Advanced Hunting** also failing.
Initially, Microsoft attempted to mitigate the outage by rerouting traffic, which offered some relief but did not fully resolve the issues. Before identifying the cause, Microsoft advised customers to review their business continuity and disaster recovery plans.
### The Root Cause: A Maintenance Bug
A preliminary Post Incident Review for the **Azure** incident revealed that the outage was triggered during routine device maintenance in the West US Azure region, where specific network paths were being isolated.
Microsoft's maintenance process typically converts these requests into system-readable instructions, incorporating checks to ensure at least one of two redundant paths remains healthy before any work commences. However, a critical bug in the request conversion system erroneously flagged additional network devices as part of the maintenance event.
This led to the unintended removal of IP routes from more devices than planned, specifically between Microsoft's West US datacenter and its wide-area network. The removed routes disrupted network traffic entering or leaving the West US region, though traffic remaining entirely within the region was unaffected.
### Azure Services Affected
The **Azure** incident resulted in connectivity failures, increased latency, and access problems for numerous cloud services, including:
* **Azure App Service**
* **Application Gateway**
* **Azure AD B2C**
* **Azure AI Search**
* **Azure API Management**
* **Azure Cosmos DB**
* **Azure Databricks**
* **Azure Firewall**
* **Azure Kubernetes Service**
* **Azure Monitor**
* **Azure Virtual Desktop**
* **ExpressRoute**
* **Log Analytics**
* **Microsoft Graph**
* **Microsoft Sentinel**
* **Power BI Embedded**
* **Virtual WAN**
* **VPN Gateway**
### Resolution and Future Steps
Microsoft engineers began investigating the issues immediately after the outage began. The problem initially manifested as large-scale route churn in Microsoft's WAN, which was later traced to a datacenter in the West US region and linked to recent maintenance activity.
Microsoft initiated a rollback of the maintenance change at 1:45 PM ET, completing it by 2:26 PM ET. This rollback restored the affected network infrastructure, allowing **Microsoft 365** services to recover. Some **Azure** services continued their recovery, with all affected services fully operational by 3:41 PM ET.
Microsoft is now conducting a comprehensive internal review, focusing on the safety checks and automated processes involved in executing maintenance requests. "We will be performing a full analysis focusing on safety checks, automated maintenance request change process, and more as we progress through our post mitigation internal retrospective," Microsoft stated. A final Post Incident Review is expected within 14 days following the completion of their investigation.