Whether they’re caused by a cyber-attack, a human mistake or failing technology, organizations of any size will have to deal with technology incidents. While these incidents are a given, what you as an organization won’t know is when they will happen or what kind of incident they will be.
An incident response plan allows you to respond to such incidents using a controlled process. You can think of a response plan as a playbook to help you continue operating while you resolve an incident. This helps you avoid being caught in a situation where you don’t know what steps to take, who will make the decisions and whether there is money available to fix things.
This document guides you through creating and, when necessary, using such a plan. It was created by the Cybersecurity Advisory Team (CAT) at the Ford Foundation’s BUILD program. The document follows the six phases outlined by the SANS Institute.
Phase one: preparation
The first phase is the preparation phase. This is what you do before an incident occurs, to provide your future selves with well-written procedures to follow during the stressful time after an incident has occurred. Good news: by reading this document you have started this phase!
Preparation starts by sitting down with the security team, IT team or daily management team, whichever is applicable to your organization.
Together, imagine an incident has occurred. This could include spyware being found on your organization’s servers, or someone impersonating your organization on social media. Who would you inform about this incident? Your IT provider? All staff members? The board? Your funders?
Then there are the people who could help you in case of an incident. Make sure you have all the right contact details for them. Do you already have an IT or security company you can contact? Good, write down their contact details. If you don’t, now is a good time to find one. Is your website important to your organization? Write down the contact details for your web hosting company. Do you rely on social media and are you concerned about abuse or harassment against you or your staff? Write down how to report this.
Decide on a chain of command in case of an incident: who will be responsible for executing which steps of the incident response plan. And make sure the people who will be leading the incident response efforts, whether they are internal or external, have the appropriate access to tools and accounts.
Don’t forget to ensure these groups also have the appropriate management approval to handle incidents. And just as important: make sure a budget has been set aside to deal with incidents.
Finally, make sure that your plan is regularly updated and easily accessible to all relevant staff members, even when your main communication channels are down.
Phase two: identification
The next phase is the identification phase. This is where you identify that an incident has occurred, where you determine its scope and, more generally, learn as much about the incident as possible.
Not all incidents will be obvious and it is important that your organization is set up to detect unusual activity that could be the sign of an incident or an attack.
Security products can play an important role here: many of these don’t just aim to stop bad things from happening, they also give you insight into what is happening. Intrusion detection and prevention systems (IDS/IPS) report on unusual activity on the network, while a security information and event management (SIEM) product collects security logs at a central location and gives IT administrators an easy way to look for suspicious activity.
Many smaller organizations don’t run their own network and instead rely solely on cloud services, but they may still benefit from running a centrally managed endpoint security solution. In addition to including some kind of antivirus, these solutions alert administrators about unusual activity happening on individual devices, or “endpoints.”
Even organizations that aren’t able to centrally manage security on endpoints may still use a collaboration suite like Google Workspace or Office 365. These allow administrators to monitor and be notified about unusual activity regarding accounts.
Identification isn’t just about using tools and logs, though. It’s just as important to have your staff members, volunteers and even clients report unusual activity. This requires a clear tool or platform through which to report such activity, but also an internal culture where people know to look out for “strange things happening” and aren’t afraid to report things, even if they believe it was “their own fault.”
Sometimes, incidents are discovered by people from outside your organization and you want to make sure your organization is easily approachable by people reporting incidents. You may want to consider creating a security@ email alias and advertising that prominently on your website. Note that sometimes people use direct messages on social media to report possible incidents. Make sure you regularly check these even if the account isn’t actively used.
When you discover an incident, you should also consider whether there is a legal requirement for the incident to be reported to a regulator. For example, this is often the case in many jurisdictions when the incident involves personal data, such as required by GDPR.
Finally, as part of the identification phase, you need to make a decision whether unusual or unwanted activity actually warrants an investigation. It is important that you create clear guidance for making this call, but also that you adopt this guidance as and when needed.
Phase three: containment
The third phase is the containment phase. You enter this phase after you have fully identified the incident and have begun the process of resolving it.
Just like people infected with a virus may be quarantined to prevent further spread, you will want to isolate a device that is possibly infected with malware by disconnecting it from the network.
If the incident concerns a compromised user account, change its password to remove access for this account from devices and services, so that someone with unauthorized access cannot use it to cause further harm.
During the containment phase, the incident response team may also take forensic images of affected devices for later analysis and save logs of affected accounts.
During the containment phase, you may apply temporary fixes, such as creating a new account for an accepted user or setting them up with a temporary device, so that they can continue their daily work while you handle the incident. Temporary fixes might require some flexible thinking. For example, an organization that has had an incident with its website might temporarily switch to social media to communicate its message.
Phase four: eradication
The containment phase is followed by the eradication phase. This occurs when you leave the “crisis mode” following an incident and prepare to return to normal operations.
During this phase, you remove any malware from affected systems, remove accounts that may have been created by an intruder and ensure there is no access to compromised accounts any longer.
Completely eradicating malware from a system is no trivial task and this is a phase where smaller and medium-sized organizations can really benefit from help from a third party.
You may also decide to rebuild affected machines, which in almost all cases will get rid of malware. Do beware, though, that a compromise may actually be caused not by malware but by a compromised account, which won’t be fixed by rebuilding a machine.
Phase five: recovery
Once the incident has been eradicated, you enter the fifth phase: recovery. That occurs when you bring affected systems back online and make sure they work properly.
The latter is the most important part of this phase: always consider the possibility of the incident not having been fully eradicated in the previous phase, so make sure you properly test the affected systems and accounts and thoroughly monitor their activity.
Recovery isn’t just about systems. While some incidents are nothing but computer glitches, others may have a serious psychological impact on members of your staff. Always make sure you ensure your staff’s emotional and psychological needs are met following the incident.
Phase six: learning
It might seem that the incident handling is finished once the affected systems have been recovered, but there is one more phase, perhaps the most important one of all: that of looking at lessons learned.
In this phase, you will complete any documentation of the incident and its handling that wasn’t written up during the incident handling itself. Do this very soon after the incident has been handled, while your memory is still fresh.
Once that is done, the team handling the incident should sit down and discuss what worked and what didn’t, both when it comes to the handling of the incident and when it comes to the general running of the organization. Maybe the incident points to a weakness that needs to be resolved if similar incidents are to be avoided.
This learning phase may include recommendations for better tools to detect and handle incidents, clearer roles for staff members, better contacts with third parties who could help the handling of the incident or simply more budget for handling incidents.
Outsourcing your incident response
Larger organizations will be able to handle the majority of incidents in-house, but smaller organizations often only have a small IT team, if they have such a team at all. They will likely outsource most incident response handling. However, even larger organizations may occasionally find themselves having to outsource the handling of more complex incidents.
When preparing for incidents, you will need to decide whether it works best for your organization to use a one-time engagement for each incident or have an incident response company on retainer. Even if you decide one-time engagements work best, you will want to decide who to call during the preparation phase to avoid losing critical time.
When engaging a third party, it is important to agree on the responsibilities and capabilities of both your organization and the third party, such as who is responsible for reporting the incident in the first place and within what timeframe. It is possible you may decide to engage more than one third party, for example, a lawyer, a communications consultant and a security company.
Outsourcing the handling of an incident can be costly, which is why it could be worth speaking to your insurance broker (if you have one) to see if there is a policy for incident handling in general and cybersecurity incidents in particular.