Banking disruption: MAS orders Citibank, DBS to open probe, says it will take ‘supervisory action’ after
THE Monetary Authority of Singapore (MAS) has instructed Citibank and DBS to conduct a “thorough investigation” into a recent disruption of numerous online services, and said it will take “appropriate supervisory actions after gathering the necessary facts”.
This comes after a technical issue at an Equinix data centre in Singapore led to the disruption of services last Saturday (Oct 14), including those of Citibank and DBS.
An MAS statement on Thursday night said that both DBS and Citibank activated their back-up data centres when their primary data centres failed to perform normally.
However, the banks were unable to fully recover their systems within the required four-hour timeframe, MAS noted.
In its statement, MAS said that it requires all banks to ensure that their critical systems and services are resilient: “Banks are required to have in place back-up data centres and systems and test them periodically.
“In addition, the unscheduled downtime for a critical system affecting a bank’s operations or service to customers must not exceed four hours within any 12-month period.”
MAS said that it expects banks to establish contractual agreements with data-centre providers that incorporate the regulator’s requirements. “No IT system is infallible. Banks and customers should have contingency measures in the event of service disruptions caused by IT outages,” it added.
Indeed, market watchers The Business Times spoke to said even the best data-centre operators do not guarantee 100 per cent uptime.
Companies providing critical services therefore need to prepare redundancies to counter data centre disruptions.
Equinix has since said that it had a technical issue with the chilled-water system at its data centre following a planned system upgrade. This raised the temperature in some of the data centre’s halls and impacted its tenants’ operations.
The company advertises 99.9999 per cent uptime for its data centres worldwide on its website. It did not respond to queries on the standards it promises its tenants, or how it would compensate affected tenants.
ST Telemedia Global Data Centres group chief technology officer Dan Pointon said that data centres may advertise anything from 99.99 per cent to 99.9999 per cent uptime, but no data centre can promise 100 per cent uptime.
A 99.99 per cent uptime promise works out to up to 52.6 minutes of downtime each year.
The actual disruption could be longer, however, as the downtime figure does not take into account time taken by tenants, such as banks, to restore their services.
“If somebody does want to provide 100 per cent uptime to their customers, they’re going to have to think about multiple data centres and duplicating the technologies,” Pointon said.
He added that data centres typically sign service-level agreements with tenants that specify uptime requirements. Failure to meet these commitments can lead to onerous financial penalties.
“We want to be able to see ourselves as a long-term stable provider of critical infrastructure, and so I think the industry... typically does a fairly decent job at self-regulating,” he noted.
National University of Singapore Associate Professor Lee Poh Seng said that financial institutions (FIs) with mission-critical workloads must also be able to recover quickly.
“Those who were affected should have moved over to their back-up systems,” he said. “We didn’t expect such an extended down period for their services.”
Prof Lee said that the Equinix data centre could have suffered a failure in a segment of its chilled-water architecture that could not easily have redundancies set up.
Equinix said on its website that its data centres feature “N+1” cooling redundancy. This means they have a single additional component in their architecture to support failure or required maintenance.
Whatever the case, Prof Lee hopes that regulators can get them to share key information when outages happen. “This can be a good learning point for other operators.”
Teresa Tan, research analyst at data centre market intelligence company DC Byte, said that some data centres provide for “2N” redundancy – a fully redundant, mirrored system with two independent distribution systems.
“This needs to be balanced against the exponential rise in cost stemming from the greater layers in redundancy, as well as the impact this has on the environment,” she added.
In August this year, a lightning strike disrupted the power supply to a data centre in Sydney, Australia, resulting in an almost 16-hour disruption for customers of the Bank of Queensland.
A post-outage analysis by Microsoft, whose Azure services were also hosted in the data centre, attributed the extended outage to insufficient staffing at night to manually restart the chillers that failed to do so automatically.
In June last year, MAS released a set of business-continuity management guidelines. FIs are encouraged to consider the separation of infrastructure, such as data centres, into different zones to mitigate wide-area disruption.
The central bank also directed FIs to consider engaging alternate service providers for redundancy when the primary provider is unavailable.
In response to queries, a DBS spokesperson said that the lender has robust business recovery plans in place, as well as data centres across Singapore.
“In this instance, the rapid overheating of the data centre triggered an abrupt shutdown of our systems, which delayed the full recovery process,” the spokesperson added.
A Citi spokesperson said: “Citi places great importance on the resilience of our infrastructure, and we will use the lessons from this incident to continually improve.”