
Microsoft says a bug in its automated network maintenance request system caused Thursday's massive outage by mistakenly removing IP routes from more devices than intended, disrupting Azure and Microsoft 365 services.
The outage began at 10:44 AM ET on Thursday, July 23, and mostly affected customers accessing Microsoft 365 services through network infrastructure connected to Microsoft’s West US Azure region.
At 11:11 AM ET, Downdetector had recorded 2,403 outage reports, sharply above its normal baseline of 29. SharePoint accounted for 78% of the complaints, followed by Excel at 11% and the Microsoft 365 Admin Center at 6%.
Microsoft tracked the Microsoft 365 outage under incident ID MO1437424 and confirmed that multiple Microsoft 365 services were impacted:
- Microsoft OneDrive - Access to OneDrive was intermittent.
- SharePoint Online - Users received "Something went wrong" errors.
- Microsoft Teams - Chat functionality was degraded, including images not loading.
- Microsoft 365 Admin Center - The Admin Center loaded slowly or not at all.
- Power Automate - Automate flows did not load.
- Copilot Chat - Users experienced intermittent delays or failures when performing actions and queries.
- Microsoft Loop - Users were unable to open or load Loop pages.
Other affected services included Fabric and Power BI, Power Apps, Copilot Studio, Windows 365, and Microsoft Defender.
Some Defender customers experienced delays receiving responses from Microsoft Defender Experts, while investigations, workflows, and remediation actions triggered through Threat Explorer and Advanced Hunting could fail.
Microsoft initially attempted to mitigate the outage by rerouting traffic through alternate network paths, which helped customers, but many services continued to be affected.
Before determining what caused the outage, Microsoft warned customers that they might need to review their business continuity and disaster recovery plans and take actions appropriate for their environments.
The company later identified a recent networking change as the cause and began reverting it.
Microsoft completed the reversion at 2:26 PM ET and confirmed through service telemetry and customer reports that the Microsoft 365 incident had been resolved.
Maintenance bug caused outage
In a preliminary Post Incident Review for the Azure incident, Microsoft said the outage was triggered during routine device maintenance in its West US Azure region, where specific network paths were being isolated.
Microsoft says its maintenance process converts these types of requests into system-readable instructions and checks that at least one of two redundant paths remains healthy before the work begins.
However, a bug in the request conversion system incorrectly marked additional network devices as part of the maintenance event.
As a result, IP routes were removed from more devices than intended between Microsoft's West US datacenter and its wide-area network.
The removed routes disrupted network traffic entering or leaving the West US region. However, Microsoft said traffic remaining entirely within the region was not affected.
The Azure incident caused connectivity failures, increased latency, and problems accessing numerous cloud services, including Azure App Service, Application Gateway, Azure AD B2C, Azure AI Search, Azure API Management, Azure Cosmos DB, Azure Databricks, Azure Firewall, Azure Kubernetes Service, Azure Monitor, Azure Virtual Desktop, ExpressRoute, Log Analytics, Microsoft Graph, Microsoft Sentinel, Power BI Embedded, Virtual WAN, and VPN Gateway.
Microsoft said its engineers began investigating the issues immediately after the outage began at 10:44 AM ET.
The problem initially presented itself as large-scale route churn in Microsoft's WAN. Engineers later traced the route removals to a datacenter in the West US region and linked them with the recent maintenance activity.
Microsoft initiated a rollback of the maintenance change at 1:45 PM ET, which was completed at 2:26 PM ET.
The rollback restored the affected network infrastructure and allowed Microsoft 365 services to recover. Some Azure services continued recovering after the fix was put in place, with Microsoft reporting that all affected services had fully recovered by 3:41 PM ET.
Microsoft is now conducting a full internal review focused on the safety checks and automated processes used to execute maintenance requests.
"We will be preforming a full analysis focusing on safety checks, automated maintenance request change process, and more as we progress through our post mitigation internal retrospective," explained Microsoft.
The company said it will publish a final Post Incident Review after completing its investigation, which is usually within 14 days.
Test every layer before attackers do
Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.
The Picus whitepaper shows how breach and attack simulation tests your SIEM and EDR rules so threats stop slipping by detection.










English (US) ·