Pipeline cybersecurity incident response: IT and OT considerations

  • TSA Pipeline Security Directives
  • Incident response
  • OT security

An incident response plan written once, for the network as a whole, tends to read fine and fail in practice. The failure shows up at the same step every time: containment. Pulling a compromised server off the corporate network is a known, low-risk action. Doing the equivalent to a PLC mid-cycle on a live line is a different decision with different consequences, and a plan that doesn’t say so out loud leaves whoever is on shift to work that out during the incident itself.

What SD-02G actually requires

SD Pipeline-2021-02G requires a Cybersecurity Incident Response Plan built around four objectives: prompt containment of an infected server or device; segregating the infected network or devices so malicious code doesn’t spread; keeping backups secure, separate from the system, and verified free of malicious code; and, distinctly from the first three, established capability and governance for isolating the Information and Operational Technology systems from each other during an incident. That fourth objective isn’t a restatement of segregation in general. It’s the directive naming the IT/OT boundary specifically as something the plan has to address on its own terms, not as a side effect of a broader containment step.

The plan also has to name who, by position, is responsible for each measure, and has to be exercised at least annually, testing at least two of the four objectives with the named position-holders actually participating. A plan that’s never been exercised is a document, not a capability, and TSA’s own structure around this requirement treats that distinction as the point.

Why containment isn’t the same action twice

The segregation objective spells out what containing a compromised device can require: removing it from the network, removing anything that shared a network with it, and, notably, preserving a forensic memory image before powering off or moving the device. That last step matters more on the OT side than it sounds. An IT server can usually be taken offline, imaged, and restored on a timeline measured in minutes without consequence beyond the outage itself. A controller running a live process often can’t be power-cycled the same way without a safety or operational review first, and the SD-02G language anticipates that: preserve the evidence, but don’t assume powering off is the default first move on every device the way it might be on the IT side.

This is the same distinction active vs. passive vulnerability assessment draws for testing: not every device on the network can absorb the same action safely, and an incident response plan that doesn’t split its containment steps by what a device can actually tolerate is making the same mistake in a higher-stakes moment.

Reporting is a separate clock

Incident response and incident reporting run on separate tracks, and conflating them is a common gap. SD Pipeline-2021-01G requires reporting cybersecurity incidents to CISA as soon as practicable, but no later than 72 hours after the operator identifies the incident, covering unauthorized access, discovered malware, denial of service, physical attacks on network infrastructure, and any other incident with the potential to disrupt IT or OT systems. The report has to include what’s known at the time: earliest known date of compromise, date of detection, who’s been notified, what’s been observed, and a description of the incident’s actual or potential impact. The 72-hour clock starts at identification, not at containment, so a response plan that treats reporting as something to handle once the incident is fully resolved is already behind the deadline the directive sets.

The Cybersecurity Coordinator’s 24/7 accessibility requirement exists specifically so that clock can be met: someone has to be reachable to make the reporting call regardless of when the incident starts.

What this means for building the plan

A plan that treats “isolate IT from OT” as its own named objective, rather than folding it into general segregation language, needs the network architecture to actually support that isolation before an incident, not during one. That’s the same question a segmentation review answers directly: whether the conduits between zones, laid out in the segmentation reference architecture, are deliberately defined and restricted enough that isolating OT from IT during an incident is a switch configuration someone already tested, rather than a scramble to figure out which connections need to be cut and what breaks if they are.

For what a Cybersecurity Incident Response Plan and its exercise program need to cover end to end, see the IT/OT segmentation design and review and OT/ICS security assessment service pages.

All posts