NTA - We are aware of an incident and are currently working towards a resolution. – Incident details

We are aware of an incident and are currently working towards a resolution.

Resolved
Major outage
Started 10 days agoLasted 12 hours 36 minutes

Affected

Inbound Calls

Major outage from 5:57 PM to 6:34 AM

Outbound Calls

Major outage from 5:57 PM to 6:34 AM

Busy Lamps

Major outage from 5:57 PM to 6:34 AM

Updates
  • Postmortem
    UTC
    Postmortem

    RFO on 23.09.26

    Wednesday 23rd September 2026 17:51
    Upstream Network Outage: A third-party provider (Cogent) suffered an outage, breaking connectivity between data centers (THN and ENF).

    Wednesday 23rd September 2026 18:20
        All ENF servers stopped and calls were diverted to THN, which enabled calls to be made, but BLF's were disabled

    Database Split Dependency: Call handling and Busy Lamp Fields (BLF) rely on sqlsrv2 (ENF), while remaining systems rely on sqlsrv (THN). Because database connections were intentionally split across sites to balance load, the network outage severed cross-site DB communication and disrupted calls and BLF updates.

     SIP Server Bug: A  bug in the  SIP server caused registration drops during the extended outage, and failed to recover automatically once connectivity was restored.

    Thursday 24th September 2026 04:00

         Cogent resolved the primary network link issue, and the Data links were restored at ENF

    Thursday 24th September 2026 06:15

        BLF's re-enabled

    SIP Service Recovery: Manually rebooted the SIP server and executed a service failover, restoring registration levels to normal

    Immediate / Short-Term Actions

    Consolidate Databases (Target: Nov 15): Following the Oct 6 web portal release, migrate call handling and BLF back to a single primary database on new hardware.
    Enable Inter-DC Link: Complete the ILF–ENF and ENF-THN interconnect to provide resilient cross-site database failover.

    Medium-Term Improvements


    Database Proxy (Contingency): Implement and test a DB proxy to automatically fail over call handling and BLF traffic if the primary database becomes unavailable.
    Asterisk Infrastructure: We will add more Asterisk servers to each DB to handle the traffic increase when running on a single DC in case of an outage.

  • Resolved
    UTC
    Resolved

    This incident has been resolved.
    The datacentre link was restored shortly before 05:00. Registrations have since returned to normal, and services at the affected site have been brought back online and monitored to confirm stability.



  • Monitoring
    UTC
    Monitoring

    Inbound, outbound and BLF have been working for some time now, however we are continuing to monitor.
    Investigations confirmed there has been a major outage with one of our datacenter's links which required some services to be manually switched over.
    We are still awaiting further information and resolution on this link so this incident will remain in a monitoring state until all connections resume and have been confirmed working.

  • Investigating
    UTC
    Investigating

    Inbound and outbound calls have been working for some time but we are still monitoring the service.
    BLF has not been brought back yet while we investigate.