Software

Slack's Nine-Hour Outage Exposes Database Architecture Challenges

Slack's widespread service failure Wednesday revealed vulnerabilities in its sharded MySQL infrastructure, prompting industry experts to question whether the company's legacy database approach can scale reliably.

7 min read
Slack: Takeaways From This Week’s Service Outage

Slack has offered limited public disclosure regarding the root cause of Wednesday's service disruption. However, Spencer Kimball, co-founder and CEO of Cockroach Labs, has offered analysis of what may have triggered the incident.

The messaging platform experienced a major outage beginning at 10:47 a.m. U.S. Eastern time, with engineers requiring nearly nine hours to restore functionality, according to incident details posted on Slack's status page. The disruption affected users across more than 150 countries, including 77 of the Fortune 100 companies. Organizations relying on Slack include Airbnb, Target, Uber, and the U.S. Department of Veterans Affairs. Slack is owned by Salesforce.

Slack has not publicly disclosed what caused the outage beyond status updates issued during the incident itself. The company declined to respond to inquiries about the service failure.

Remediation work involves repairing affected database shards, which are causing feature degradation issues. This has become a diligent process to ensure we're prioritizing the database replicas with the most impact.

Slack status update, 4:04 p.m. EST Wednesday

By 7:42 p.m. EST, Slack announced restoration of full functionality across affected features including message sending, workflows, threads, and API-related capabilities. Recovery efforts continued into Thursday, when the company posted an update at 10:17 a.m. EST noting that events generated during the outage remained queued and paused, with hourly progress updates promised.

The Problem With MySQL

Without fuller disclosure from Slack, industry observers must piece together what may have failed. Slack launched in 2013 using MySQL as its data storage engine from inception. Beginning in 2017, the company initiated a migration to Vitess, an open source database scaling system built on MySQL.

According to a December 2020 post on Slack's engineering blog, the company faced mounting challenges as it expanded. The post noted that "our application performance teams were regularly running into scaling and performance problems and having to design workarounds for the limitations of the workspace sharded architecture."

The shift toward Vitess reflected a dilemma many organizations encounter during growth: the disruption required to abandon entrenched legacy systems. Slack remained deeply committed to MySQL infrastructure.

At the time there were thousands of distinct queries in the application, some of which used MySQL-specific constructs. And at the same time we had years of built-up operational practices for deployment, data durability, backups, data warehouse ETL, compliance, and more, all of which were written for MySQL.

Slack engineering blog post by Rafael Chacón, Arka Ganguli, Guido Iaquinti, and Maggie Zhou

The blog post continued: "This meant that moving away from the relational paradigm (and even from MySQL specifically) would have been a much more disruptive change, which meant we pretty much ruled out NoSQL datastores like DynamoDB or Cassandra, as well as NewSQL like Spanner or CockroachDB."

Maintaining a sharding model may have created conditions for incidents like Wednesday's, according to Kimball's assessment. During the outage, Slack's status communications referenced "problems with corrupted database shards."

Kimball explained the mechanics of database sharding to The New Stack: "basically what you do is you have a lot of customers, a lot of data, way too much to put into one single, monolithic database. So you create lots and lots of databases, and you call them shards. And so you say, 'OK, well, customers one through 100 are on shard one, and 100 through 200 are in shard two,' and so forth, right? Problem is, you're kind of in the position of managing 100 databases. So you don't just have one database. You've got 100 of them, and all of them are separate. They're siloed, which is actually a huge problem, because you might have a customer that's too big to fit on one shard."

This architecture presents inherent tradeoffs. Kimball noted that "when you lose a shard, you only lose a subset of your customers. But the problem is, you've got 100 things to manage, and they all do have their own unique kind of weirdness, because different customers have different ways of using them."

Testing What Went Wrong

Cockroach Labs, Kimball's company, develops CockroachDB, a distributed SQL database system. While Kimball has commercial incentive to promote alternative architectures to Slack's approach, he also provided perspective on why legacy sharded systems may struggle with complexity and reliability.

I don't know how many shards Slack has, but I imagine there's quite a few at this point. They have become essentially a database company as well as a corporate messaging company, because of the thing they've built that accommodates this idea of shards and the resilience on each one of those shards. When you have one of these things, every piece of code you write, every new feature you write, it has to also deal with the underlying, exposed reality of this complex architecture that you've cobbled together, that, by the way, doesn't work together. It's not one integrated, holistic whole.

Spencer Kimball

For organizations like Slack that have invested heavily in MySQL infrastructure across hundreds or thousands of shards, what preventative measures exist against future service disruptions?

Whatever just happened to them, this should be part of their standard testing process.

Spencer Kimball

Kimball recommended that testing teams focus on improving Recovery Time Objective (RTO)—reducing it from the approximately nine hours required for Slack to restore most services following Wednesday's incident.

It's not like Slack is anywhere unique. Everywhere saw these outages that have been happening this year. It's insane. From the Federal Aviation Administration, to Barclays and Capital One, everyone has outages. But the question is, OK, whatever just happened, let's routinely test that. And when we have our runbooks and we apply them, what can we optimize our RTO to?

Spencer Kimball

Organizations conducting quarterly or biannual database infrastructure testing, combined with team capability to resolve outages and visibility into resolution timelines, gain valuable insight. According to Kimball, "then you at least know when you're regressing a little bit, and you know what the cost is going to be when this inevitably happens again."

Deploying backup databases on different cloud providers than primary systems represents another recommended practice, Kimball suggested.

The Cost of Resilience

Resilience standards themselves continue evolving, Kimball observed. When Cockroach Labs was established a decade ago, the baseline expectation was "let's survive a data center going away. Because that's pretty common: Like, when Google started off, it's like, hey, let's survive a node or a physical machine failing or something. Then it's, 'Let's survive data centers going away. Then it's, 'Hey, people want to survive whole regions going away.' And it's like, now people want to survive whole cloud providers going away."

Few organizations conduct resilience testing twice yearly or quarterly, and Kimball identified a practical reason: "It's expensive to do these things." Beyond direct costs, staff time and resources required for testing and re-architecting legacy systems often compete with priorities around new feature development.

Some companies decide, let's just keep our fingers crossed. It might not be a terrible answer, if they're really under some serious constraints. It's like, we'll take the egg on our face if things go wrong and we'll hope for the best. It just depends on what your use case is and what you think the cost will be from the downtime. For financial services, that's not an option anymore, especially with regulator scrutiny. For Slack, you know, maybe they're going to stay on their thing because it's just too hard to move. Eventually, you have to modernize things, and they'll just find the right time.

Spencer Kimball

Source: The New Stack

Source: The New Stack · Reporting supplemented by The Silicon Ledger staff.