SecurityMetrics Podcast | 14
6 Phases of an Incident Response Plan
“Something has happened.” Your company has experienced the worst: a data breach. You’ll need to answer questions. You’ll need to implement emergency operations and plans, run backup, and talk to investigators. Not a convenient time to start your Incident Response Plan.
According to Dave Ellis, SecurityMetrics VP of Investigations (GCIH, PFI, QSA, CISSP), an Incident Response Plan (IRP) is, in short, “What you do ahead of time, in preparation for an event that you hope never happens.”
Dave Ellis sits down with Host and Principal Security Analyst Jen Stone (MCIS, CISSP, CISA, QSA) to discuss in detail the phases of an IRP, along with the circumstances, variables, and options surrounding this “worst case scenario.”
- Emergency-Mode Operations, contingency planning, and the recovery phase
- How to get initial buy-in from your executives, C-suites, and decision makers
- Case studies and examples from the field: the practical realities involved in maintaining a current Incident Response Plan
- Tips to avoid, handle, and learn from data breaches, ransomware, and other types of malware
Resources:
Download our Guide to PCI Compliance! - https://www.securitymetrics.com/lp/pci/pci-guide
Download our Guide to HIPAA Compliance! - https://www.securitymetrics.com/lp/hipaa/hipaa-guide
[Disclaimer] Before implementing any policies or procedures you hear about on this or any other episodes, make sure to talk to your legal department, IT department, and any other department assisting with your data security and compliance efforts.
6 Phases of an Incident Response Plan Transcript
Hello, and welcome back to the Security Metrics podcast. I'm Jen Stone. I'm a principal security analyst here at Security Metrics, and I have with me again today, Dave Ellis. Dave, will you please reintroduce? I know we've had you on on in the past but some people might not have seen that one yet. Tell people who you are.
Oh, sure. Thanks. I'm Dave Ellis. I've been with Security Metrics for, thirteen years. I'm the vice president of investigations.
And prior to coming here, I was in law enforcement.
Spent twenty years with the Oakland Police Department.
Graduated from the FBI National Academy.
I was one of the SWAT team commanders. And a lot of stuff that's to my kids, a whole lot more interesting than when I talk about doing computer forensics.
And yet, it seems like it's a really good, kind of career path to go from from the one thing to the other.
Yeah. That's kind of how it it did begin is, SecurityMetrics was looking for somebody who had law enforcement background because in the earlier days, the nature of the investigations, you interfaced a lot with local law enforcement, FBI, secret service. And so it was that background plus I I did oversee the the cyber crimes division while I was there. So that's kinda what, you know, led me to, finding Security Metrics and finding a career and a home here and I've I've loved it.
That's yeah. Security Metrics is a great place to work. I gotta Good. I gotta say. So, we have you here today to talk about the six phases of an incident response plan. And before we kind of dive into the phases, maybe it would help if we talked a little bit about what's the difference between an incident response plan and a disaster recovery plan and a contingency plan and, emergency operations mode for for HIPAA people, you know, those types of things.
Help me frame that.
Well, you know, the the truth is is all four of those elements that you just mentioned are are kind of woven together.
Some of them take place at different points in time. Like, for example, the incident response plan is something that you're going to develop before anything has hit the fan.
Hopefully what you're doing ahead of time in preparation for an event that you hope never happens. Sure. Disaster recovery on the other hand is how you're going to get your systems back online and returning to normal business operations after a breach has occurred.
Right. And then emergency mode operations is when you're still in the thick of things and they're not normal yet. That's the emergency mode. And then, oh, contingency planning is like the whole set of things that are planned for. What's likely to happen, what incidents are likely to happen, and then the planning that goes so these are all subsets of the overarching contingency planning.
Yeah. So, you know, if you look at the different phases of an insert incident response plan, they will encompass each of those other aspects.
Right. So let's dive into those phases. First phase is to prepare. Right?
Right. And that's that's truly the workhorse of of it all. This is if you do your homework and and you, apply yourself and and develop a thorough comprehensive plan, You train all of your employees on what their security responsibilities in the company are.
You train the members of the incident response team, on what their roles are when that day comes, that you hope never comes.
Right.
You get buy in from the c suite. That's one of the important parts of preparation is that, you get executive buy in on your plan because sometimes, they're going to have to be purchasing equipment or paying for training or things along those lines. So you you want to get their advance buy in, before an incident occurs so that you don't have somebody hemming and hawing over having to write a check. So, you know, something along that line. So that's all of the the stuff that you're gonna do beforehand.
And then you're going to be training annually, ideally.
If if you're in a ecommerce or ecommerce kind of a role, PCI does require annual training on your, on your plan.
The more you train, the better and more naturally everyone will respond when an event does occur. One of the services that we do is, you know, we will go out, and meet with companies and put on mock drills for them when they have to actually invoke their incident response plan and everybody has to respond to the this mock drill the way they would had it been a real event. Right. And by doing that, it kinda helps you figure out where the holes in your plan might be.
Right. So testing it is is a real critical part of actually developing the plan because if you can't test it, if you well, testing, it helps you find out where where you've kind of fallen down in in certain areas. Right?
Absolutely. Yeah.
One of the the interesting things that I encounter, occasionally is I I work with a lot of universities. And so sometimes an incident response response plan development and test makes sense to a group of IT guys or who, you know, people who are already in the frame of mind, well, a SOC or things that they're already used to it. And I'll go out to some of these smaller merchants that are that make up a university and they'll say, we don't know how to test this. There is no test for this.
Right.
And I'll say, well, if you I want you to to tell me what what is the one thing that could prevent you from, usually, it's in relation to PCI. Tell me the what's the one thing that could prevent you from taking credit card payments? And they'll say, oh, well well, what if somebody comes and steals the machine? Like, that would be an incident.
Right? So what would you do? And so helping them understand, in their environment, what does an incident response look like? What are the likely things to happen?
So I don't know if you've run into that with people where they've had a struggle even coming up with what are the incidents that are are possible for them.
Right. You know, and and sometimes they confuse, an incident and an event.
And, you know, an incident is, let's say your systems pick up that you've you've been under attack. You hit like a, a DDoS attack or something like that. But your systems held up to it. So that was an incident but it didn't turn into an event.
Right.
The event is where, you know, like the lady said, somebody comes in and steals their POS machine. It it's quantitative, it shut them down, or a hacker has, breached their system and, obtained, you know, classified or HIPAA data, credit cards or whatever it might be.
So yeah, and that actually leads right into the second phase which is identification.
Right.
So something has happened, whether it's an incident or an event.
You you have to be able to answer the following questions.
You know, when did it happen? When when did it start?
Now, a lot of times people will confuse when we found it with when it started.
Right.
Yeah. And and how was it discovered?
In in a lot of cases, it is brought to your attention by an outside party. If you're in the the, you know, credit card processing side, one of the credit card brands might notify your bank that they think there's been a breach. We've seen cases where the FBI have found credit cards on the, you know, for sale on the dark web and they trace it back and they determine that they came from your location or HIPAA in information, same exact scenario. They find it for sale on the dark web and they trace it back to, you know, to your facility.
But that's probably the worst way to find out.
Well, okay. I I'm glad you said it. Second to worst.
Okay. Okay.
But starting with the best way to find out that you've been breached is when you discover it internally.
Yeah.
You know, you have file integrity monitoring or an IDS or an IPS that alerts someone and that person is actually paying attention to the alerts. I I I say that one because I'll throw in a quick horror story. We had a case, it was a large credit card breach, nine hundred stores in, you know, in a large chain.
All of them affected by it.
They had a wonderful, IDS system in place, intrusion detection system in place. And it was doing its job. It was throwing alerts that, hey guys, there's a problem going on and and nobody was watching.
Mhmm. Yeah.
And almost a year went by, where credit cards were being harvested the whole time. Had somebody been watching, they would have killed that breach on day one.
Right. And and that's I I'm glad you brought that up because a lot of the the people that I talk to not a lot. But sometimes I'll I'll especially with a new customer, I'll talk to them and say, alright. How are you monitoring network activity?
How are you monitoring your antivirus? Tell me how what kind of logging do you have? What kind of alerting do you have? And they'll say, well, we don't really have that.
We don't we're small. We don't feel the need for it.
And I think that you need to assume that you're being breached if you don't have instrumentation in place to show that you're not.
You know? If you can't demonstrate that the integrity of your data in your systems, then, you need to assume that that you don't have integrity around it.
And and a follow-up to that when you ask, you know, what monitoring capabilities are are are taking place regularly.
Excuse me. The follow-up to that is whose job is it to, you know, review those logs? Who who's who's has the role or the responsibility to receive and respond to those alerts?
Because, you know, as I had mentioned, you can have systems that are doing their job and doing a good job at it.
But if nobody's watching Yeah.
Yeah. That that can be a real frustration for your business.
Yeah. And and tuning tuning those logs dedicating someone or a team depending on the size of your organization to do that, It's hard to explain to senior management sometimes why that's important to them.
Right. Yeah. And I I think there's been enough stuff going out in the media in the last couple of years where you can you can bring the potential of a horrific and, you know, experience to the attention of of the c suite at your company. Yeah. Let me circle back though where we started was how, your breach might have been discovered. So I mentioned that, you know, the best way to find it is that you do it internally.
Yes.
That you, you know, somebody is watching. Then after that, typically, it will be an external like like I mentioned, a bank or the FBI or something along that line. Mhmm. And now and the worst one, great place, it's a little more rare, but, you know, you read read a Krebs article and he mentions your company. Yeah. And that's like Krebs is your IDS, you know, so.
That we don't want that. You know, what else is really bad is if the those who have breached your company reach out and say, hey. We've breached you and we have your information. And also, we want you to pay all this money to either not make us reveal it or to get it back or, you know, ransomware. That's worse. The whole extortion path, that's probably pretty bad too.
You know, and it's it's one of those that a couple years ago we thought was gonna kinda subside.
But it it's just like that bad penny. It just keeps coming back and they continue to make money at it.
And the the healthcare industry seems to be, you know, one of their their favorite targets because they recognize that in, you know, in the healthcare industry, these these, facilities don't have the option of being down for very long.
Right.
And and so they're probably more likely to to pay. And so that that's what a ransomware attacker is looking for is somebody who, is going to pay upfront instead of taking the the time to bring their systems back online.
Yeah.
You know, through through other, you know, means. You know, reimaging and restoring systems, things like that.
Unfortunately, true. Okay. So we've prepared our plan. We we've identified that there is an incident happening. What do we do?
K. You know, the state phase three is containment.
And and this is one where you have to, resist then what might be a knee jerk reaction that says, okay. I'm just gonna nuke our entire system. We're gonna reimage everything and, you know, and come back online.
While that might contain you, it might not.
Mhmm.
Because you're you're also going to be eliminating, information that that that's essential for you. The kind of information that will tell you how they got in, whose credentials were, were compromised if that's how they got in.
It it will, you know, that that nuclear option this early in the scenario, would prevent you from learning from what just happened and being able to prevent it from happening again.
Right.
And in some cases, that nuclear option doesn't solve it because they actually come in through, you know, third party credentials. Mhmm. You know, third party users that have credentials to get into your site. Things along those lines.
Right. So in containment, what you're looking to do is to stop the bleeding right now. This isn't the investigative part where you go back to find out necessarily how they got in. So an an element of containment could be, you know, hey I just unplugged from the internet.
Right.
You've achieved containment, but you it's gonna be a little tough to start moving back toward natural operations. So you're you're looking at, both, you know, what do I have to do to contain short term? What do I have to do to contain long term? Containment would involve things like discovering the malware and quarantining it.
Going through and looking at all the access credentials and hardening those.
Verifying that all the credentials are legitimate. A lot of times if if an attacker gets in by by compromising somebody's credentials, they will then go on to create their own set of credentials, you know, separate from the one that they had used before. Right. In case that person then changes there. So so you're gonna, you know, do a thorough review of everyone's credentials.
You're going to look to see if if you've applied all the recent, patches and updates, security releases, things like that. You're gonna review your room remote access to make sure that true multi factor authentication is in place.
You know, we've seen a couple of cases where they said, yeah, we have multi factor authentication because I have to put in both the username and a password.
That's not that's not multifactor authentication.
Yeah. Kinda missing it there. But so that that's what you're gonna go through. The the most important part of of containment on there is finding the the malware if you can and getting that quarantined, and then going through and, you know, like so, looking at all the access controls.
Once you believe that you've achieved containment, that's when then you circle back and you're looking at the deeper or taking the deeper dive and looking to eradicate, you know, all the remnants of of the attack.
Sure. So eradication is phase four. Right?
Yeah. Okay. And so in here, you're going to be looking at, you know, have you found all the artifacts?
How you And in keeping with that, a lot of times, I see people where they will say, okay, we found the malware and we've removed it and so they think we've eradicated it. Well, by finding the malware, you're actually finding one of the later stages that occurred Right. In the breach. Because what what happened long before they installed malware on the system.
How did it get there?
Yeah. Exactly. How did they get there in the first place? So so you've gotta be able to trace back and, you know, and close those doors.
And sometimes this is also gonna involve, you know, having to update your systems In in the the credit card side, there's Magento is very popular Sure. E commerce platform.
And the, the Magento one was the popular one being used for years. And, we're still bumping into people that are using Magento one and it's no longer supported. It's rife with holes and and so, you you know, during the eradication phase, you you need to look at your system saying, okay, is everything up to date? Right.
And then you're at this point, you can now consider, alright, we've we've hardened everything, we've patched, we've applied updates.
Okay. Can we reimage the system? Can or or do we have a a backup that we can restore from? And and this is kind of starting to slide into the the next area. But if if you're going to apply a nuclear option, sort of at the end of phase four would be the place where you'd consider doing that.
Okay. Because it it you you've eradicated and you know that you can do the nuclear option. So we we all we always called it Nuke and Pave.
Yeah. So you can do that. So that's entirely reimaging your system so that you can restore because you feel like that's the point where where you know that you can do that without reinfection happening. So so that moves us into the recover phase.
Right. Yeah. And so, and this is the one where you ask yourself, you know, can we safely return to to production, to normal operations? And and in order really to be able to answer that question, you have to be able to say, alright, has everything been patched, hardened and tested? Mhmm.
You know, can we reimage or do we have a trusted backup that that predates the breach?
Right.
That, you know, that we can go back to or roll back to?
And then you've gotta ask, you know, how will the affected systems be monitored in the future? You know, and and you kind of addressed that before.
The you know, the importance of keeping an eye on your system.
Right.
And and and creating those processes that make sure all the patching, all the updating, all of the things that you did to make sure that your your systems are now robust, that those processes continue. So a lot of times they'll say, oh, yeah. We patched our systems last year. Yeah.
Yeah. You know, a good analogy to that is if you could look at your company's, security as like a living, breathing organism.
You know, that if it's healthy today, to keep it healthy, you have to keep feeding it. You have to keep taking care of it. And that's, you know, continually nurturing, you know, and and taking care of of things because, you know, what, you know well, look at SSL. SSL was great right up until it wasn't. Right.
And and and our security can be the same way. You know, one day to the next, everything can change.
Yeah. And that and that I think that's a hard concept sometimes for people who are kind of, not on the technical side of things.
Working with customers that I've had over the years and telling them you have to be at TLS one point one right now. And and the the heavy lift that that was for some of them. And for some groups, it's not a heavy lift. It's just, you know, applying the new the new standard in in their servers.
But for for others, especially when you have maybe older systems that have to talk to that system and are not prepared for that upgrade, it was a harder thing. And so then the next year, you come along and say, sorry. Now you have to be TLS one point two. And they say, will the will this ever end?
Yeah.
No.
No. It won't. Yeah.
That's right. Yeah. Yeah. Yeah. Sadly, yeah. No.
No. Sorry. This is the life you live now. So, so we've taken a a trusted backup.
We've restored, that we've recovered, so that the systems are working properly and tested that. So, are we good then? Is that are we done? Is that the last phase?
No. Not really. The the next phase is, you know, after you've done all the mechanical stuff to get your your, you know, company back on its feet, now you wanna have an after action meeting where you get together and you discuss the this event and say, you know, what have we learned from it? Mhmm. You know, where did things go right in our in our system?
Where did it fall a little bit short? Did we notify everybody? You know, and that was one of the things I didn't mention a little little bit earlier is, you know, when you after you identify an event, there's going to be notifications certainly internally in your company but there's probably gonna be some notifications that you need to make outside. Well, during this end, this last phase, lessons learned, you need to ask have all the notifications been made? Have we made, you know, legal notifications to the state, to consumers, to whomever?
Then we've talked about, you know, what worked for us. Now what didn't work? What's what caused us to stumble during this event?
And, Yeah.
And to to improve the to improve the plan itself, and and react better the next time, I wanted to to step back just a second and say the notification piece. One of the things that I've seen is companies can get themselves into worse situation, than before if they try to not notify without a plan for notification. So just like every other action, who is allowed to communicate with the various, state organizations or the people whose information was breached? Not just allowed, but who is responsible for that. Because those are some some decisions that have to be made on what's the timing, what's the messaging, and and, what rises to the level of notification.
And understanding that is not always an easy question. Because like you said earlier, is it an incident? Is it an event? Is it reportable? Is it not reportable? So so that has to be very clearly laid out, and and I've seen that be a struggle for some people.
Yeah. Jen, you made a perfect point there.
Those those notifications oftentimes are going to be dictated by law or sometimes by, you know, other, outside entities. And a lot of times what will happen is the pool of notifications that you have to make, that you think you have to make at the beginning, that pool will look very differently at the end of the event.
Mhmm.
And so and and then a lot of these notifications have time frames attached to it. You know, forty five days or sixty days or ninety days or or whatever. And so you you really want to make those notifications when you have as much information as possible. Mhmm. We can give the people that you're notifying, you know, instructions on this is what you can expect as we move forward.
Makes them feel a whole lot better than, oh, hey, your data might have been compromised.
Yeah. We don't know. We'll let you know. Maybe we'll let you know.
Great. So those are some so six phases of, incident response.
How often does a group need to check its plan? Like, how often should they be doing those tests and reviews?
You know, minimally, they should be doing it annually.
And in in some, you know, spheres, there are, things that dictate that.
And and, you know, the payment card industry space for example, PCI DSS twelve point ten point two, indicates that they need to test their incident response plan annually.
Know, and HIP HIPAA, I shouldn't really think that they would be any different.
Yeah. They say regularly, but, HIPAA HIPAA is a little harder to, to come up with specifics because it's not prescriptive. But on the other hand, if you start with the risk assessment that that HIPAA says start with that risk assessment, then that kind of dictates. Some groups are going to need to to to perform this plan, review it more more frequently, maybe quarterly. And some groups maybe maybe every other year would be okay.
Yeah. And you know, and another thing to take into account is if you have gone through and made some major system changes, you want to then review your instant response plan to say, did any of these changes that we made in our, you know, network environment, will they have any impact on our instant response plan?
Do we need to alter the plan now a little bit to put somebody else in in charge of or responsible for doing you know x y z Right.
Related to this new bank of servers or whatever.
Sure.
So before we wrap up, I was wondering what's your what's your thoughts on the best way to test a plan? Is it is it a simulation? Is it a tabletop exercise? How do you how do you test it?
Great great question. I I there's a couple of different types of tests. And I think the first one, that you should do is a a tabletop. It's the easiest training to pull off. It's where people aren't aren't going to actually, you know, pick up the phone and make a notification. They're not going to have to do a lot of the physical things that you would do in a, more directed mock drill.
And they're a little bit easier to pull off. You know, the IT manager who's responsible for putting on the training can find a lot of, you know, good sample, tabletop exercises online.
Or you can, you know, hire a company to come in and do it as well. But it's easier to stage than it would be a full mock drill. Now, once you've done a tabletop or two, I would highly recommend you move into an actual mock, drill. And the difference there is is you actually have people logging into systems and going through the motions that they would otherwise do.
It just adds just a little bit more realism Yeah. You know, to it. You can't quite do this in real time because, you know, the truth is is these play out over, you know, several days if not weeks.
But the the the more realistic you can make the training, the better it is for everyone and and the better, report card it will give you of the quality of of your current incident response plan.
And I think it's important even for small small so back again to the very small merchants that make up, a lot of universities.
One of the groups that I went and talked to said, we have this incident response plan that IT gave us. All we do is call their number. And I said, great. What what was your test for that?
Well, we don't have to test it because it's just that simple. I said, okay. Did you call the number? Nope.
Let's call the number. Let's find out. And so they called the number, followed all of the instructions on this sheet of paper. Well, since getting that sheet of paper from the the central IT, they had changed the the NOC to a SOC.
And so they called and said, I need to talk to the NOC about this issue. And they said, you you have the wrong number. And and the the person who answered was new, wasn't familiar with with why they would be calling. So the whole test was it it it caused a problem.
And it it created a new procedure because all they did honestly, all they had to do was call that number. And the number was very confused about why they were calling, which is why you test even the smallest groups.
Right.
Just whatever your plan is, give it a test.
Yep.
Yep. Yep. And and and it's amazing how much you learn from that. So you like that last phase that we talked about, in in the phases of the incident response lessons learned. I've seen a lot of companies come back and say that was the most important step for them is because they could then go back and and refine their plan.
Yeah.
And, you know, they can say, okay, this is what we need to do to prevent this from happening again.
And it doesn't have to be boring either.
Yeah. Yeah. Sure.
It can be fun. Like, the you can use there's a lot of companies that they use the, D and D dice, you know, the twenty sided dice and all of those. And then so they have different scenarios and then they have to roll to see what happens. One of the groups that I talked to, I was reading their their instant response plan test documentation.
And at the end, it said, things went so wrong, they blew up the data center and everybody died.
That is a series of bad roles, guys.
That's great. That's awesome.
We've we've actually, done some of these, you know, mock drills for companies. And we'll usually go in with one or two different scenarios.
Or excuse me, two or three different scenarios. We'll kinda give them a softball warm up and then get a little bit, deeper.
But we try to think of things that they might not think of in a in a, you know, the natural course because people who design environments know what that environment is supposed to do.
Mhmm.
And sometimes it takes someone from the outside to come in and challenge that and like in one case, you know, they said, oh, yeah. We're hosted. And but in this scenario, their hosting facility experienced a fire. Mhmm. And and, you know, and they lost access to, you know, those systems.
So now what do you do? Right.
And they step back and went, I don't know.
I don't know.
And something else they gotta learn. So another one, that I helped with was, who's in charge? He gets called away to something very important, and cannot help you. Because I didn't wanna say he died because that's just rude. He he has called away and lost his phone and cannot help you.
So and then He's in Borneo, and he doesn't have Internet.
Exactly. It's time for his vacation. We gotta let him go. So, yeah, just just, shuffling things so that once you start feeling like your team is good at it, they're getting bored, then do something different.
Change it up. Take people out. Add people in. Yeah.
So well, this has been super. I really appreciate you talking to me about incident response today. It's it's it's critical. Everybody needs to do it, and, it was super fun talking to you about it.
Oh, it's always a pleasure.
Is there anything that I missed before we close?
No. I think, I think you covered every possible aspect of incident response to this.
Well, I might have had a little help with the talking points. Thanks, Dave. Alright. Well, good to talk to you.
Thank you for joining us here at the Security Metrics podcast. I'm Jen Stone. I've been talking with Dave Ellis, and I hope you join us again in the future. Thanks for watching.
To watch more episodes of Security Metrics podcast, click on the box on the right. If you prefer to listen to this podcast, it's available on all your favorite podcast platforms. See you on the slopes.
