Home » T-SQL Tuesday #202 – Two weeks of ransomware and my most memorable outage

T-SQL Tuesday #202 – Two weeks of ransomware and my most memorable outage

by Vlad Drumea
0 comments 8 minutes read

This month’s T-SQL Tuesday invitation comes from Marlon Ribunal, and it’s titled “That One SQL Server Outage You’ll Never Forget”.

I’ve been meaning to write something about dealing with the biggest/most memorable outage of my career for a while, and this month’s invitation gives me an excuse to do just that.

This is about the outage caused by a ransomware attack in October 2020 on my then employer, Steelcase, and the 2+ weeks required to get everything back online.

You can read more about the two-week halt of global order management, manufacturing, and distribution, as well as the fact that forensic investigation found no evidence of exfiltration, in this BleepingComputer article.

Disclaimer: I mention and link some products and companies in this blog post. This isn’t a sponsored post, and none of these companies have had a say in what I’m writing here.

What is ransomware?

In short, ransomware is a type of malicious software (aka malware) that holds a victim’s data hostage until a ransom is paid.
The files are generally encrypted with an algorithm and key that only the attacker knows.

Note: I haven’t signed any NDAs related to this, so I do have some flexibility in the amount of details I can include, but I’ll try to keep some things on the vaguer side.
But, since this happened 6 years ago, I might be a bit fuzzy on some things where I wasn’t directly involved plus at the end of those 2+ weeks I was severely burned out.

So how does an outage caused by ransomware, specifically RYUK in this case, look like from a DBA’s point of view?

Detection and network shutdown

It all started on October 22nd 2020 at around noon EET, when we started noticing multiple servers, including SQL Server hosts, becoming unresponsive.

RDP-ing into a couple of SQL Server hosts revealed the following three signs confirming a Ryuk attack:

  • User profile corruption errors on logon.
  • The .ryk extension appended to all encrypted files.
  • The usual ransom note.

For anyone wondering how that ransom note looks:

Side-note: Malwarebytes has a really good analysis on the Ryuk ransomware.

Since it was already affecting some production systems, a Critical Incident call was initiated (as was the MO for any production impacting issue)
During the call we’ve made the decision to shutdown network traffic for the entire environment until further notice.

At this point we didn’t really know how bad the damage was, but we were suspecting things will get worse before they start getting better.

Assessing the damage and putting together a recovery plan

The Security team reached out to Mandiant (now part of Google) for incident response consulting.
And, under Mandiant’s guidance, the following steps were outlined:

  1. Deploy FireEye and a new VPN on all unaffected laptops.
    So that we can safely do our part in bringing production back up.
  2. Identify VM snapshots taken before the attack.
  3. All Windows servers should be treated as potentially infected and have the VMs restored from snapshots.
  4. Confirmed infected Windows servers should have a snapshot made for cyber insurance purposes.

So, pretty much all Windows servers were going to have to be restored from snapshots.
And, depending on what servers lived on them, more restore and/or config work might be required.

Since we did regular DR tests, we already had a prioritized list of which servers should be brought up first.

The SQL Server perspective

One thing I’d like to note here is that, while we did keep a local copy of the current week’s backups, all of the backups were safe and sound because we were using NetBackup as our long-term backup solution.
Unfortunately, this also meant that Netbackup would become the main bottleneck since everyone was pulling backups from it for everything (VMs, file shares, databases, etc).
All this while ongoing backups were being pushed to it.

From a SQL Server ransomware recovery perspective, this meant 60+ instances spread over 30+ VMs.
And, since most non-prod VMs as well as some prod VMs were clown cars*, this meant that ~20 out of those 30+ VMs were hosting production SQL Server instances, while the rest were non-prod.

*my go-to derogatory term for a VM with multiple SQL Server instances.

The plan was pretty straight-forward on this side:

  1. Have the Sysadmin team restore the VM from snapshot, harden it and install FireEye.
  2. Attach the now formatted database disks to the VM.
  3. Re-create the instance-related directories.
  4. Have the Ops team provide the necessary files from Netbackup.
  5. Restore the system databases (master, msdb, model).
  6. Start the SQL Server instance.
  7. Restore the user databases.
  8. Kick off a fresh full backup for the entire instance.
  9. Rinse and repeat for the next VM.

To take some of the load off of the Sysadmin team, I could handle point 2 myself since I already knew how to add disks to VMs.
And to speed things up for the entire DBA team, I’ve automated points 3 and 5 through 8 with PowerShell, T-SQL and dbatools, and documented all the steps.

Luckily, the environment didn’t include Availability Groups, so the restore process was pretty straight-forward.
Or was it?

Reinfection, broken LSN chains, and VMs without viable snapshots

During the recovery process we did run into a few hiccups.

Reinfection

A couple of the VMs got reinfected shortly after they were restored.
At this point, I’ve decided to make a change in our backup process so that we could still rely on the local copies of the database backups even if the host VM gets reinfected.

A bit of context: when Ryuk encrypts files it excludes files with the following extensions:

  • .exe
  • .dll
  • .ini
  • .lnk
  • .hrmlog

The explanation for this is simple: as a ransomware developer, the purpose of your software is to encrypt the user’s files, in order to do that you need a working operating system, which you can’t have if you also encrypt binaries. And, another reason to not nuke the OS is that you might still need it for other post-exploitation work.

Going back to the aforementioned change: I’ve changed the backup file extensions from .bak to .exe for full and diff backups, and .dll for transaction log backups.
This way I could leverage Ryuk’s own file extension filtering rules to ensure that the local copies of the new backups are safe and we don’t have to pull them from Netbackup again.

Note that I didn’t do this based on a wild hunch or a guess, I had confirmation of this behavior from the code of that version of RYUK that hit us.
And this also isn’t a guarantee that absolutely all ransomware behaves the same.

I’ve actually talked about this before, on a GitHub issue for Brent Ozar’s First Responder Kit, and in a couple of LinkedIn comments which I’m too lazy to dig for at the moment.

Side-note: if you’re going to ask “but how did it manage to encrypt .mdf, .ndf, and .ldf files since they were opened by SQL Server?”, the answer is that Ryuk stops SQL Server related processes as part of its attack chain before starting the encryption process.

Broken LSN chains

This happened on two non-prod instances where the host was poorly configured and VSS snapshots ended up invalidating our differential backups.
The solution here was simple and, at the request of the app owner, we restored a fresh backup from prod on the two non-prod instances.

VMs without viable snapshots

If I recall correctly, only one VM didn’t have a viable snapshot, so it had to be rebuilt from scratch.

But, after SQL Server was installed with the exact patch level as the original instance, both system and user database backups were restored without any issues.

Split across 3 teams and 16 hour work days

Due to my experience, I’ve also volunteered to help the Sysadmin team with some of the post-restore VM tasks, as well as the Security team with sifting through weeks worth of connection logs to get an idea of how the attack spread across the environment, and also validating infected VMs.

For a little over 2 weeks, including weekends, my schedule looked something like this:

  • 8AM wake up.
  • Join calls, work on restores, go through status updates.
  • 12AM (aka midnight) call it a day and try to get some rest.

The queen of all burnouts

To absolutely no one’s surprise, this whole thing ended up taking a toll on me and, when all systems were back online with no sign of infection, I was just exhausted.
I took some days off to just sleep and avoid anything that made me think of stuff more complex than “food goes in mouth”.

Lessons learned and what I’d do differently

I guess the main thing would be to manage my effort, and avoid spreading myself too thin to make up for someone else’s poor staffing decisions.

Another thing I’ve learned was that the environment could have been better hardened, which I’ve handled afterwards.

The main factors that played an important part in bringing everything back online:

  • Recurrent DR tests, which helped us have an updated view of the environment as well as a prioritized list of servers and services.
  • Backup test restores.
    When we knew we had to restore everything, we were confident that our production backups were viable.

This experience also pushed me more into the cybersecurity side of things, leading to me getting my OSCP in August of the following year.

What I’d tell myself in 2020

If I could get ahold of September 2020 Vlad, I’d tell him to do that backup file extension change ASAP, so that any backup files stored locally would have the .exe/.dll extensions a month before the attack.

You may also like

Leave a Comment

* By using this form you agree with the storage and handling of your data by this website.

This site uses Akismet to reduce spam. Learn how your comment data is processed.