Featured image

Table of Contents Link to heading

On September 18, 2025, Optus suffered a major network outage that disrupted Triple Zero (000) emergency services across several Australian states. The outage lasted over 13 hours and resulted in four confirmed deaths and hundreds of failed emergency calls — making it not just an infrastructure failure, but a public safety crisis with real human consequences.

What Went Wrong in 2025 Link to heading

The September 2025 outage was triggered during a routine firewall upgrade. A misconfiguration blocked emergency call traffic at the network layer, and the backup mechanism — the “camp-on” system, which allows mobile devices to route emergency calls through an alternative carrier when the primary network is unavailable — also failed.

The dual failure of the primary path and its backup is the critical detail here. A single point of failure is a known risk; a simultaneous failure of a backup designed to cover exactly that scenario points to either a testing gap or an architectural dependency between the two systems that should have been isolated. Among the victims were an eight-week-old baby, a 68-year-old woman, a 74-year-old man, and a 49-year-old man. Delayed incident notification to emergency services compounded the impact.

The 2023 Outage: A Missed Warning Link to heading

The 2023 Optus outage — also a 12+ hour nationwide event — was caused by a BGP routing error during a software upgrade. That incident affected millions of users, generated over 2,000 failed Triple Zero calls, and prompted a government review recommending 18 specific reforms to strengthen emergency call reliability.

By September 2025, only 12 of those 18 reforms had been implemented. The recurrence of a structurally similar failure — again during a planned maintenance window, again affecting emergency services — makes it difficult to interpret the 2025 outage as an unforeseeable event. Unimplemented reforms are unmitigated risks.

Systemic Issues and Regulatory Gaps Link to heading

The pattern across both outages points to systemic issues rather than isolated incidents:

  • Emergency call routing was not functionally isolated from general network operations, meaning a change that should only affect data traffic could cascade into 000 disruption.
  • Redundancy mechanisms existed on paper but were either misconfigured or never validated under realistic failure conditions.
  • Pre-deployment testing for network changes was insufficient — particularly for changes touching infrastructure adjacent to emergency services.
  • Outage notification to emergency services was delayed, degrading the ability of police and ambulance services to adjust their operations during the outage window.
  • Regulatory enforcement post-2023 was not sufficient to drive full compliance with the recommended reforms before the next major incident.

What Needs to Change Link to heading

The regulatory and industry response to the 2025 outage has centred on several structural reforms:

  • Mandatory real-time outage notifications to emergency services, with defined SLA thresholds.
  • Stricter technical standards for emergency call systems, including mandatory isolation from general network changes.
  • Independent auditing of Triple Zero infrastructure, separate from the telco’s internal quality assurance.
  • Automated pre- and post-change testing protocols for any network modification that could affect emergency services paths.
  • Federal investment in resilient emergency communications infrastructure, potentially under a dedicated national custodian for Triple Zero services.

Pattern of Failure: Why the 2025 Outage Was Not Unforeseeable Link to heading

The September 2025 Optus outage did not emerge from a vacuum. The technical failure modes — upgrade-induced misconfiguration, untested redundancy, delayed escalation — are identical in character to what happened in 2023. The difference is consequence: in 2025, people died.

From a systems engineering perspective, the lesson is not that network upgrades are inherently dangerous. It is that safety-critical systems — and emergency communications infrastructure qualifies — require a different operational discipline: independent failure domains, validated failover, mandatory pre-change testing against emergency service paths, and regulatory teeth strong enough to enforce compliance between incidents, not only after them.

Optus now faces legal exposure, regulatory scrutiny, and a serious erosion of public trust. More significantly, the 2025 outage has opened a national conversation about whether critical emergency communications infrastructure can remain in the hands of commercial operators without stronger mandatory standards and independent oversight.

Key Lessons Learned Link to heading

Emergency Systems Must Be Isolated and Redundant Link to heading

Both outages demonstrated that emergency call routing shared failure domains with general network operations. The camp-on mechanism’s failure in 2025 revealed that the backup system was architecturally dependent on components of the primary system it was supposed to replace.

Lesson: Emergency call systems must operate in independently managed failure domains, with automated failover tested against real failure scenarios — not just nominal configurations — on a defined regular cadence.

Real-Time Outage Reporting Is Critical Link to heading

In both incidents, Optus’s delayed notification to emergency services limited their ability to activate contingency procedures. Every minute of delay in notification translates directly to uncoordinated response in the field.

Lesson: Mandatory real-time outage alerts to emergency services must be a regulatory requirement with defined response-time SLAs, not a best-effort practice.

Infrastructure Upgrades Require Rigorous Testing Link to heading

The 2023 outage stemmed from a BGP routing change during a software upgrade. The 2025 outage stemmed from a firewall misconfiguration during a routine change. Both were planned maintenance events — not unexpected failures.

Lesson: Any network change within scope of emergency services infrastructure must undergo automated pre- and post-deployment validation, with defined rollback criteria and verified rollback procedures before the change window opens.

Regulatory Oversight Needs Strengthening Link to heading

Eighteen reforms were recommended after 2023. Six remained unimplemented by the time the next major incident occurred. This is a compliance and enforcement failure, not just a telco failure.

Lesson: Regulatory bodies need the authority and mandate to audit compliance with reform commitments on a continuous basis, with meaningful penalties for non-compliance that create real urgency rather than treated as a cost of doing business.

Public Trust Depends on Transparency and Accountability Link to heading

Optus’s response in both incidents was characterised by slow communication, limited proactive outreach, and insufficient support for affected users during the outage window. In a crisis, silence is the worst possible communication strategy.

Lesson: Telcos operating critical infrastructure must maintain crisis communication playbooks with pre-approved messaging, proactive outbound contact capabilities, and defined escalation paths to government and emergency services.

Emergency Communications Should Be a National Priority Link to heading

Both outages exposed the fragility of relying on commercial network infrastructure — subject to normal commercial pressures and upgrade cycles — for life-safety communications.

Lesson: Emergency communications infrastructure must be classified and governed as critical national infrastructure, with dedicated funding, independent technical oversight, and standards that are set and enforced by government rather than self-regulated by industry.