On July 23, 2026, a bug in Microsoft’s automated network maintenance system triggered a nearly five-hour outage across Microsoft 365 and Azure, cutting off a West US datacenter from Microsoft’s global network.

Reading our article? Try our product:

Full Protection for Windows Servers - Zero-Risk Trial

What happened

The outage started at 10:44 AM ET during routine maintenance meant to isolate specific network paths in Microsoft’s West US Azure region. By 11:11 AM, Downdetector had logged 2,403 reports — 83 times its normal baseline. SharePoint made up 78% of complaints, followed by Excel and the Microsoft 365 Admin Center.

Affected services included:

  • SharePoint Online – “Something went wrong” errors
  • Microsoft Teams – degraded chat, broken image loading
  • OneDrive – intermittent access
  • Microsoft 365 Admin Center – slow or unresponsive
  • Power Automate, Copilot Chat, Microsoft Loop – failures and delays
  • Microsoft Defender – delayed expert responses, failed threat-hunting workflows

The disruption also spread into Azure infrastructure, hitting App Service, Cosmos DB, Kubernetes Service, ExpressRoute, Microsoft Sentinel, VPN Gateway, and more than a dozen other services.

The root cause

Microsoft’s maintenance process converts human-issued requests into machine-readable instructions, then runs a safety check confirming at least one of two redundant network paths stays healthy before work begins.

A bug in that conversion step incorrectly marked additional devices as part of the maintenance scope. Because the safety check validated the buggy output rather than the engineers’ original request, it passed a change that removed IP routes from far more devices than intended — severing the West US datacenter from Microsoft’s wide-area network (WAN). Traffic inside the region kept working; anything needing to enter or leave hit a dead end.

The downstream ripple effect

At least 19 third-party services with infrastructure in Azure’s West US region reported cascading outages, per monitoring firm IncidentHub. Sixteen of those were still showing active incidents after Microsoft declared recovery. Named examples included delayed Azure metrics in Datadog, cluster connectivity failures in MongoDB Cloud, and portal degradation in Wiz.

This is the third major 2025–2026 outage where a well-designed safety check failed because a bug elsewhere in the automation pipeline corrupted its input — following similar patterns in a February 2026 Cloudflare BGP incident and a May 2026 Azure OpenAI outage. The recurring lesson: a safety check that validates a system’s output rather than the original human intent has a structural blind spot.

Microsoft is conducting a full internal review of its safety checks and maintenance request process, with a final Post Incident Review expected within 14 days.

Takeaways for IT teams

  • Have a fallback communication channel. Organizations with backup messaging or file-sharing tools weathered the outage better.
  • Consider multi-region deployments. Microsoft itself recommended this in its PIR for mission-critical workloads.
  • Don’t mistake SLA credits for risk coverage. The 99.9% Microsoft 365 SLA pays 25% of monthly fees on a breach — a fraction of what an extended outage actually costs most businesses.

Fortify Your Server with Messageware Security

Data breaches have increased by 72%, servers are compromised in under 90 minutes. Ensure you have multiple layers of security software protecting your Windows Servers.

Server Threat Guard (STG) for All Windows Servers: Next-gen server protection, providing detection, alerting, and response (MDR) to zero-day and server penetration cyber-attacks. No need to research complicated deployments and no learning curve to install and manage.

EPG Guard for Exchange Servers: Real-time security. Stop AD account lockouts, eliminate password attacks, intelligent GEO blocking, and prevent Exchange Server vulnerability probing.

Don’t leave your critical infrastructure vulnerable, be proactive and stay ahead of evolving threats.