Kenexis Functional Safety Podcast

The basic process control system keeps the plant running, but can it also count as a safety net? This episode tackles the fraught question of when engineers may claim the BPCS as an independent protection layer—and how much credit it deserves. Ed Marszal walks through Clause 9.3’s risk reduction factor ceiling of 10, the onion diagram’s defense-in-depth logic, and the fired heater minimum fire example that makes the abstraction concrete. The discussion then turns to Clause 9.3.4’s crucial two-item limit per LOPA scenario and Clause 9.3.5’s independence requirements, including Ed’s blunt corrective that DCS hot standby is not redundancy. The episode closes with Clause 9.4’s common cause assessment and how IPL rationalization fits into the workflow. For engineers negotiating the boundary between process control and functional safety, this is the guidance that prevents dangerous over-crediting.

In this episode of the Kenexis Functional Safety Podcast, we explore sections 9.3 to 9.4 and tackle an important question in the field of functional safety: Can a basic process control system serve as a valid protection layer?

We’ll examine the conditions under which this might be appropriate and discuss the factors that influence how much credit should be given to such a system. Stay tuned as we dive deep into this crucial topic, offering valuable insights to help you navigate the complexities of functional safety.

Listen to the latest episode of the Kenexis Functional Safety Podcast, hosted by Ed Marszal, President and CEO of Kenexis. Available now on Spotify and Apple Podcasts, Ed shares his expert analysis of the IEC 61511 standard.

With decades of experience in safety instrumented systems and as a Principal Engineer, Ed has a unique perspective to offer. He has been an active contributor to the ISA 84 committee since 1994, adding to his deep understanding of the field.

In this inaugural season, Ed delves into the IEC 61511 standard, unpacking the meaning behind each word and providing a thorough interpretation of its application. Through personal stories from his career and committee work, he offers valuable context and insights for professionals in the industry.

Full Episode Transcript

KENEXIS FUNCTIONAL SAFETY PODCAST — S1E21 TRANSCRIPT (Markdown)
Cleaned & reflowed for web publication and AI crawlability.

The JSON-LD block below is schema.org structured data. If your CMS lets you
add raw HTML to a post, paste it into the page (or anywhere in the
body — crawlers read it either way). Fill in PLACEHOLDER_EPISODE_PAGE_URL
once the post exists. Everything from the "# Kenexis Functional Safety
Podcast…" heading down is the transcript body — paste it into your post.
–>

"`html

"`

# Kenexis Functional Safety Podcast — Season 1, Episode 21: IEC 61511, Clause 9.3 to 9.4 (BPCS as Protection Layer and Common Cause Failures)

## Introduction, Disclaimer, and Episode Overview

The basic process control system may be claimed as a protection layer, but under what circumstances and how much credit should we give it?

Welcome to the Kenexis Functional Safety Podcast. I'm your host, Ed Marszal, President and CEO of Kenexis. Kenexis is a technical safety consultancy that helps chemical process industry companies to analyze risk and design engineered safeguards like safety instrumented systems and fire and gas detection systems. Kenexis also provides the industry-leading suite of software tools, including our best-in-class Vertigo software for SIS Safety Lifecycle Management.

In this first season of the podcast, we are going to focus on the IEC 61511 standard, doing a deep dive into the standard, including more depth of information on what the standard means and how to apply it, brought to life with personal war stories and behind-the-scenes discussions of the committee members as we develop the standard in ISA 84 and IEC SC 65.

Before we start, a little disclaimer. I will be providing my opinion on technical and engineering topics. This information is provided on a best-effort basis and is of a general nature. The information presented in this podcast might not be applicable to your specific application. It is the obligation of every engineer to thoroughly analyze any system that they are designing and not blindly rely on any general advice presented in this podcast.

So today, we are going to continue on in the allocation clause, clause 9 of the IEC 61511 standard. The title, or right now we're in clause 9.3. So allocation, again, was the process for taking the required risk reduction that we determined in clause 8 and then allocating it to all of the multiple independent protection layers that are capable of reducing risk. Now, that process is more of a technical safety. It's more of a process safety management functionality than an instrumentation and control function.

So we definitely have to give a bit of deference to the process safety groups in this process. But, since we are in the instrumentation and control world, you know, we do feel that we are right and capable to put some restrictions around how much credit you're going to give to instrumentation and control when you're doing this process. And one of the primary pieces of equipment or areas where risk reduction is provided is going to be the basic process control system.

The basic process control system being the DCS, the main controller, that thing that keeps the process in its normal operating state, that thing that does not need to be designed in accordance with the IEC 61511 standard, because it is not the shutdown system. It is the controller. So, a lot of people put a lot of functionality into the basic process control system that, in all honesty, is safety related. And in a lot of cases, you're going to take credit for that instrumentation in the allocation process. So, how much credit can you take?

And under what circumstances are you going to be able to take that credit? That is the discussion that's going to be had in Clause 9.3.

## Clause 9.3.1 and the Onion Diagram

Clause 9.3 is requirements on the basic process control system as a protection layer. So, if you're going to use your basic process control system in a protective way, furthermore, you're going to take credit in your risk analysis for the basic process control system as a system that you are relying on to reduce risk to a tolerable level. So, we at IEC, the instrumentation and control community, are going to put limitations as to how much credit you're going to be able to claim for these systems and under what circumstances are you going to be able to take this credit.

So, again, Clause 9.3 requirements on the basic process control system as a protection layer.

The first sub-clause 9.3.1 states, The basic process control system may be claimed as a protection layer as shown in Figure 9. Now, Figure 9 in the standard is titled Typical Protection Layers and Risk Reduction Means.

And for those of you that are familiar, this is the infamous onion diagram that has been in use for decades for many purposes. It was available way before even the ISA 84 standard. I know that it was in CCPS books from around 1992 related to using instrumentation and control for safety purposes.

The way that the onion diagram is laid out is that it shows the process under control as the inside circle, or in this case we're using rounded rectangles, but you get the point. And then that process is going to be encapsulated by multiple layers, the same way that an onion has an inside layer and then there are different layers that build up around the central core. So, the diagram is intended to convey that the process has defense in depth.

You're not relying on a single protection layer to make sure that the plant is going to be safe. You're going to have multiple independent protection layers that are all encapsulating the process to contain the hazards of the process and prevent those hazards from leaving the process and injuring any stakeholders, whether that be people in the plant, neighbors to the plant, the environment, and so on.

So, in the particular onion diagram that's shown in figure nine of the standard, the process is the center. The first layer outside of the process is control and monitoring. This is the basic process control system, monitoring systems, and operator supervision. So, the basic means of keeping the plant in control. That's the BPCS. That's what we're going to be talking about in Section 9.3.

Beyond the control and monitoring layer, there is a prevention layer that is going to prevent loss of control of the primary mechanisms of control from turning into consequences. And this includes mechanical protection systems like mechanical relief valves, mechanical overspeed devices. It's also going to include process alarms with operator corrective actions. So, you've lost control and your operator takes a corrective action to return the process to the safe state as opposed to an action to try to keep it in its normal operating limits.

And then, the purpose or the target of the entire standard, safety instrumented systems are also a prevention mechanism. Because they detect that you've gone out of control and basically shut down your plant to prevent loss of containment.

Now, the third onion, part of the onion diagram is going to be the mitigation layer. So, in a mitigation layer, you have already lost control of the plant, you've lost containment, and now there is a consequence occurring because you've lost containment, but you want to make the consequences of that loss event, that loss of containment, as small as possible. Now, in the standard, it says that mitigation includes mechanical mitigation systems, safety instrumented systems, operator supervision.

Not a whole lot different from what it taught in prevention layers, but the key is, mitigation systems don't prevent the loss of containment, they make the consequences of the loss of containment smaller.

So, if you have fire detection that's going to turn on water spray curtains or some sort of dry chemical fire suppression, they didn't prevent the fire from occurring, they just made the consequences of the fire smaller.

So, fire and gas systems are usually consequence mitigation and other things like water spray curtains that are intended to contain releases of flammable or toxic materials, and emergency scrubber systems that will take material released from the process and use a water or a caustic system to try to scrub the dangerous chemicals out of the vapor that was released from the process and contain it in such a way that it minimizes what gets released to the atmosphere and what is capable of causing harm to the neighbors and to the environment of the process.

The next layer in the onion diagram is plant emergency response, and the final layer in the onion diagram is going to be community emergency response. So, plant would be the fire brigade at the plant itself, and the community emergency response would be the local, municipal, or state firefighting teams that would come in and help the plants fire brigade to contain the situation.

So, the onion diagram is shown here, and it basically emphasizes the fact that that first layer in the onion diagram is the basic process control system, and that in a lot of cases you're going to put functionality into the basic process control system that you are not claiming is a safety instrumented function, but its intention and its purpose is preventive as opposed to control.

Okay?

So, that's 9.31 talks about the onion diagram and gives you a display of the onion diagram before we start diving into the meat of this section, which is going to be in Clause 9.3.2.

## Clause 9.3.2 BPCS Risk Reduction Factor Limit

So, Clause 9.3.2, one sentence and one note. It says, the risk reduction claimed for a BPCS protection layer shall be less than or equal to 10.

So, the most credit that you could possibly take for a basic process control system used as a preventive or mitigative measure is going to be a risk of 10, a one order of magnitude decrease in risk, or a probability of failure of 0.1. So, that is the limit and you will see that this definition includes the risk reduction factor of 10, which is what most LOPA procedures going all the way back to the 1990s, the early 1990s, had always provided for you. So, if I have a BPCS protection layer, the most credit I can give it is a risk reduction factor of 10.

Now, there is a note to Clause 9.3.2 and that note states, consideration can be given to the fact that a BPCS may also be an initiating source for the demand on the protection layer. Now, we're going to talk a lot about the basic process control system and the safety instrumented system and sharing equipment. And, especially when you share equipment, you need to think about the fact that, well, failure of this device can create the demand that I am trying to protect against.

And, when you're crediting protections, well, you really can't credit a protection if the failure of the components caused the scenario to occur in the first place.

So, if we can't, if a sensor failure causes a flow to go low, you can't magically expect that transmitter to start working again to make your protection layer work. So, shared components between control functionality and protected functionality need to be considered very carefully. And, what the note is trying to convey to you is that if failure of your control loop is the initiating event in your scenario, you probably shouldn't be taking credit for it as a protection layer too.

Okay, so, let me, before I go a little bit more deeply into this, so now we know a BPCS protection layer shall not be greater than a risk reduction factor of 10 in terms of the credit that I give it when I'm doing my allocation, when I'm doing my layer of protection analysis.

## Fired Heater Minimum Fire BPCS Example

So, to try to make this a little bit more concrete, let me give you some examples of what a BPCS protection layer might be. And, I think that the archetype of a BPCS protection layer would be minimum fire on a fired heater.

So, as you know, fired heaters are big consumers of SISIO points. I always like to say probably about 50% of all I.O. points sold by safety instrumented system manufacturers are in some way, shape, or form utilized on fired heaters because fired heaters can be dangerous. You're creating a flammable and or explosive atmosphere in a very controlled way and then using the heat of that combustion for industrial purposes.

And any loss of control can result in a big flammable gas cloud that if it finds a source of ignition can create a potentially devastating explosion. And, you know, the historical record shows hundreds, thousands of fatalities associated with this type of event.

So, you're going to have a lot of safeguarding these devices. So, one of the primary protection layers used for fired heaters is something called the low pass flow. So, for a low pass flow, you have a heater and you have a liquid, let's say, that you're heating up or maybe you're boiling water to make steam, what have you. You have a liquid and that liquid is going through the fire box of the heater and the liquid is inside a tube. So, we're trying to transfer a lot of heat from the fire box into the fluid that's flowing through that tube.

So, the way the system is designed, we're trying to encourage heat transfer through that tubing. Well, the problem that can arise in this situation is that if you lose the flow of the liquid going through the tube, the tube is going to get real hot real fast and potentially it's going to heat up outside of its mechanical limits and it will literally melt and burst open. And when it melts and burst open any liquids that are inside that tube are going to go flooding out into the firebox and the consequences of that can be very severe, they can be very minor, it depends on what's in the tube.

If you had water in the tube, it's just a little bit of equipment damage. But if we had high pressure toxic gases, so think about a hydrocracker in an oil refinery, you have high concentrations of hydrogen sulfide, flammable materials, hydrogen and hydrocarbons at pressures in excess of 1000 psi. If that tube pops open, the consequences could easily be in the fatality range.

Okay, so on almost every fired heater, we're going to have this interlock where we measure the flow rate and if the flow goes low, we want to cut back on the heat. Now, in accordance with all of your standards, like your API 556s, your NFPA 85s and 86s, when you get low pass flow through the tube, you want to stop firing. So you're going to have an interlock that detects if the flow is low, stop firing, and that usually means isolate the fuel gas going into the fired heater, put the flame out, and so on.

Well, that's something you definitely want to do. If you have to. But also, the downside of that shutdown is, well, after you shut down a fired heater, it's going to take some time to start it back up, because you're going to need to reset the process, you're going purge that firebox out and go through a light-off sequence. So that purge of the firebox and the light-off sequence can take a significant amount of time, it will either cause your product to go off-spec or actually require you to go through a restart of a reactor sequence or a restart of a distillation column.

So the cost of this shutdown can be very significant, which is why some people will time delay the safety instrumented systems action, because it generally does take a significant amount of time for the tubing to be compromised, even when you completely lose flow through the heater. It could be on the order of many minutes, it could be hours, it could even be days before you compromise the tubes.

So what a lot of people do is they will put a time delay on the main trip, and then use a BPCS protection layer.

So in their control system, they'll look at the flow measures that they're using to direct flow through the fired heater, and if they detect that the flow going through the fired heater goes low, let's say because the upstream pump failed, and it's going to be some time before the spare pump gets put back online, then if they detect that the flow has gone low, they will put the heater to minimum fire, which means they'll take that firing valve that controls how much fuel gas is going into the fired heater, and set it to a very low output to keep a flame in the firebox so that they don't have to restart, they don't have to go through the repurge sequence, but the amount of heat generated by that low fire flame is not going to compromise the heater tubes, so it kind of puts it in a low firing standby until full flow rate has been established through the tubes when you allow the firing to get cranked back up,

so operating companies that employ this minimum fire, well, they want to take credit for it, and, you know, they can take credit for it, they should take credit for it, because it is something that makes operation of the plant safer, but how much credit can we give it, and under what circumstances?

## BPCS IPL Credit Conditions and Criteria

So the answer there is how much credit can we give it, it's a risk reduction factor of 10, and that's assuming that your BPCS equipment, considering how often you do your tests, of all of your components, is capable of achieving a risk reduction factor of 10, and when you take credit for that component, it's not like you wipe it off the board, you still need to make sure that the integrity of that is appropriate, you need to make sure that you have documentation of that functionality, you need to make sure that you have testing of that functionality on an appropriate interval to make sure that it is available to do its job.

Now, documenting non-SIS instrumented protection layers is something that has been hit or miss in the past, but as I record this on January 14th of 2025, later in the day today, I will be sitting in the ISA 84 committee where we are discussing the ISA 84.91.03 standard that is going to be released in the not-too-distant future that basically sets requirements for operating companies to basically document and test all of their instrumented protection layers that they're taking credit for, that risk reduction has been allocated to, even if they are not a safety instrumented system, albeit to a lesser degree, then you would document, test, and analyze a safety instrumented function.

So, if you have appropriately designed and documented your BPCS protection layer, you can take up to a credit of risk reduction factor of 10, probability of failure of 0.1.

But, although that's kind of the numerical limit, we need to consider all of the other factors that make something an independent protection layer. That's going to include not just that, that risk reduction factor of 10 covers two of the factors, two of the seven factors currently that define what an independent protection layer is, kind of on the availability, reliability side of the fence. But we still need the other factors to be true.

Specific design, specificity. Independence from the initiating event and all the other protection layers. Auditability, testability, making sure that it is tested. We need to make sure that we have security to make sure that it cannot be tampered with. And then also that whole management of change concept needs to apply. So all those things need to be checked off to get that risk reduction factor of 10.

Now, when it comes to the basic process control system, the trickiest part of the independent protection layer thing to get your hands around is going to be the I, the independence. independence.

## Clause 9.3.4 BPCS Sharing Limits per LOPA Scenario

And basically for the rest of clause nine, there are a whole bunch of additional requirements that talk about independence because there are a lot of things that are going through the basic process control system that factor into our LOPA analysis. And if you think pure independence, only one thing could go through the basic process control system, but the standards developers, the standards committee members, such as myself, determined that only allowing one thing in the LOPA scenario to go through the basic process control system would be too restrictive.

And what you'll see in clause 9.3.4 is that we are going to step back from one item per LOPA scenario to two items per LOPA scenario.

Okay, so let's dig into that concept. This concept is going to be discussed in detail in clause 9.3.4.

9.3.4 starts out if it is not intended that the BPCS conform to the IEC 61511 series, then we're going to have two bullet points. So the first statement is kind of a throwaway. BPCS is basically saying, you know, if your BPCS is not an SIS, which it shouldn't be, this is how much credit, or this is the sharing attributes of the BPCS protection layers.

So before I hit the two bullet points, let me kind of explain what I mean by sharing. So your basic process control system, let's go back to that fired heater scenario. In that fired heater scenario, the failure of the flow control loop can cause low flow through my heater tubes.

So the BPCS is what is controlling the flow through the heater tubes. So the BPCS might be part of the initiating event.

Now, if I have a loss of flow, I might get a high tube skin temperature alarm. Well, that alarm might be enunciated through the basic process control system. So the basic process control system contains not only the initiating event, but it also contains the enunciation mechanism for the alarm that we want our operator response to be based on.

And then if I have low flow functionality that will put me to minimum fire, so low flow to minimum fire basic process control system protection layer, that's a third item that is in my LOPA scenario that is going through the basic process control system.

So you'll see that there is sharing, but there's a good potential that the only sharing is the CPU of the basic process control system, meaning that each of the different aspects of the LOPA scenario, they might have had different sensors, they might have had different input cards for the sensors, they might have had different outputs, different cards for the outputs, but there's probably at a minimum going to be sharing of the CPU.

So, what we need to determine is how much sharing is okay. And that example that I just gave you has three things going through the basic process control system. We said the initiating event is a BPCS loop failure, the operator intervention based on alarm, the alarm goes through the BPCS, and a BPCS protection layer that's going to detect that the flow went low and set my control loop to minimum fire. Well, guess what? Three is too many.

So, that scenario that I just described to you, the standard says you are not allowed to take that much credit for your basic process control system in a single scenario.

And that limitation is set here in 9.3.4. If I could summarize before I get to the two bullet points, the maximum number of things going through the basic process control system is going to be two, no more than two things. And that, those two things can be described in two different ways, which are the two bullet points of clause 9.3.4.

So, first bullet point says, no more than one BPCS protection layer shall be claimed for the same sequence of event leading to a hazardous event when the BPCS is the initiating source for the demand on the protection layer.

Okay, I'm going to step back a little bit because some of this language is a little bit tortured. It's necessarily a little bit tortured, but let me kind of simplify it for you. So, when the standard says a sequence of event leading to the hazardous event, the shorthand way to think of that is a LOPA scenario. So, that's a scenario. Starts with an initiating event, has protection layers. If you get an initiating event, your protection layers fail, you get a consequence. That is a sequence of event leading to a hazardous event. You could just think of it as a scenario.

You can think of it as a LOPA scenario.

So, no more than one BPCS protection layer if the cause for the scenario is a BPCS failure. So, if I have a control loop whose failure results in the consequence potentially occurring, then I can only take credit for one additional protection layer going through the DCS, going through the basic process control system. So, in the event of a control loop failure causing low flow through that fired heater, I could take credit for the operator intervention based on the tube skin alarm, or I could take credit for the low flow to minimum fire trip, but not both.

And honestly speaking, you probably wouldn't be able to take credit for the low flow to minimum fire trip, because the sensor that you use for that BPCS protection layer is the same sensor that you're using in the control loop that's the initiating event, so you didn't have enough independence anyway.

## Non-BPCS Causes and Fault Tree Validation

Okay, now let's talk about the second bullet point. So for the second bullet point, it says no more than two BPCS protection layers shall be claimed for the same sequence of event leading to the hazardous event when the BPCS is not the initiating source of demand.

Okay, so putting this in plainer English, if you have a LOPA scenario where the cause is not a control loop failure, not a BPCS failure, then you can have two protection layers.

So let's pull on that one a little bit more with the same scenario we've been talking about this entire day, and now this time, let's say that the cause, the initiating event that causes low flow is a pump failure.

Well, a pump failure is not in the basic process control system, you know, assuming kind of a standard oil refinery setup where you have a manual switch in the field to turn the pump on as opposed to some, you know, variable speed driver on a pump that's going to be controlled by the DCS. Let's just say it's just on-off switch in the field. So if that pump fails, completely independent of the basic process control system, now we will be able to take two protection layers.

So for the case initiating event of pump failure, you will be able to credit both the low flow to minimum fire protection layer and operator intervention based on alarm for the high tube skin temperature. Okay? So you're allowed to take credit for two things going through the DCS. either your initiating event and one protection layer or two protection layers if the cause is not the DCS.

So this clause 9.3.4 has a note. Let's go ahead and read through the note. The note says the identified BPCS protection layer can consist of one BP, the initiating source for the demand, C clause 8.2.2, and a second independent BPCS protection layer, C clauses 9.3.2 and 9.3.3, which we just talked about, or up to two independent BPCS protection layers when the initiating source is not related to BPCS failure. So two things are allowed to go through the DCS. A cause and one IPL or two IPLs.

Okay, so that is part of when you're allowed to take the protection layer. That's part of the independence aspect to it. And this is something that I have done work on all the way back since the 1990s. I've always given this two levels of credit because if you analyze this system with more rigor, using tools with more horsepower like a fault tree analysis, you will see that if the only thing shared between two BPCS protection layers is the CPU of the basic process control system, meaning I've got separate sensors, separate input cards, separate output cards, separate outputs.

If the only thing that you share is the CPU of the basic process control system, that will not prevent you from getting to a risk reduction factor of 100 for the combined system as elegantly calculated in the fault tree, which elegantly handles the commonality, the shared components between the systems.

But that same fault tree analysis will clearly show cannot get to a risk reduction factor of 1000 because the shared components will hold you out.

So I did a paper for an ISA conference all the way back in 2003 that discussed this. If you want to send me an email, I'd be happy to send a copy to want to look at this in a little bit more detail.

So this clause is consistent with what I've always recommended. All the high horsepower quantitative risk analysis that I've always done clears this as being true efficacious, it works.

Okay, now this is part of independence saying, well, there is a point in time where you're going to be independent enough. But what else goes into independent enough? So we know that two BPCS protection layers can be independent enough. But what does that mean? Well, what that means is going to be defined in the next clause, clause 9.3.5, which contains a one-sentence clause, but it also contains two notes. So what does it mean to be independent enough?

## Clause 9.3.5 Separation and DCS Redundancy Myth

9.3.5 states, when 9.3.4 applies, each BPCS protection layer shall be independent and separate from the initiating source and from each other to the extent that the claimed risk reduction of each BPCS protection layer is not compromised.

So it's telling you that you need to think about how independent those protection layers are, but it doesn't get down to brass tacks and give you prescriptive requirements about what does it mean.

So there's going to be some judgment and there's also going to need to be some quantitative analysis of the dependencies.

I recommend always when you're doing this type of analysis to fall back into a fault tree analysis. Get that Arbor fault tree analysis tool in the Kenexis integrated safety suite and you know, you can always get a free trial license to that software. And even if your trial license is expired, send me an email. I'll probably set you up with another one. I'm pretty easy. for doing these types of analysis because that fault tree analysis truly handles those dependencies. It uses minimal cut set analysis to be able to elegantly model those things to make sure.

So you need to look at your system to make sure that it is separate and independent so that that risk reduction factor of 100 that you're claiming from two protection layers is actually going to be achievable.

Now there are two notes to this clause. Note one states the assessment of separation and independence can consider what is necessary to achieve the risk reduction. For example, central processing units, input modules, output modules, relays, field devices, application programming networks, program databases, engineering tools, human machine interfaces, bypass tools, and other devices.

There are a lot of things that are capable of generating a common mode, common cause failure between multiple protection layers. You need to assess and think through all of them to make sure that you're not putting yourself at risk.

Note two to this clause states the assessment of separation and independence can consider what is necessary to achieve the risk reduction. Wait a minute, that's what I just told you. Sorry about that, I was reading note one again, let me get to note two. Note two says, a hot backup controller is not considered to be independent of the primary controller because it is subject to common cause failure.

For example, hot backup controllers have components that are common to both the primary and the backup controller, such as the back plane, firmware diagnostics, transfer mechanisms, and undetected failures.

Okay, so what we're saying is if you have a DCS and the DCS vendor tells have redundancy, I've got some bad news for you. Your DCS is not redundant. It's not redundant.

I don't care what your vendor told you. Your DCS is not redundant. It's pretty much impossible for a DCS to be redundant.

DCSs are hot standby. They are not redundant. And hot standby improves your availability, it improves your uptime because it allows you to detect a failure in your primary and switch to your backup.

But if you don't detect the failure in your primary, you will never switch to your backup. In order to be truly redundant, your primary and your backup both need to be doing all of the work all of the time, comparing results, and then choosing which result is correct to send to the output.

that's what a safety PLC does. But DCSs, honestly speaking, they're too complicated. How do you do a voting in comparison of a PID controller? You know, in a two out of three vote, you could do comparisons. A PID control loop, cycle by cycle, it's just not possible. It's not realistic.

So, your DCS is not redundant. It's hot standby. So, for hot standby, you do have a single point of failure. So, always keep that in mind. Your DCS is not the same when it comes to safety as your safety instrumented system. They work differently. They have different diagnostic mechanisms. Don't confuse the two. Your DCS is not redundant. It's hot standby.

Okay, so, that's clause 935. And, if I can boil clause 935, the rule of thumb in general is, if you, when you want to have two things running through your basic process control system, you want independent measurement devices, you want independent input cards, don't land the two devices on the same card because there's shared circuitry. You're going to want different output cards for outputs and you're going to want different physical output devices.

The only thing that you should be sharing and still be able to get to that risk reduction factor of 100 is just the CPU of the basic process control system. That's it. And, of course, there's all the other attributes of a LIPOL, they're calling it now, a low integrity protection layer, 10 or lower, that the basic process control system interlock falls into that category. Again, that's the ISA 84.91.03 standard that defines what a LIPOL is, and it's going to set those requirements for testing, for documentation, and verification that the design is adequate.

Of course, the verification for a LIPOL does not necessarily include a CIL verification calculation.

## Clause 9.4 Common Cause and Dependent Failure Requirements

Okay, so that was Clause 9.3. Normally, I would stop there because we have gone, we're getting pretty close to an hour, but all we have left in Clause 9.4 that talks about common cause, common mode, independent failures, and this topic is really tightly related to what I just discussed because Clause 9.3.4 just told you, you're allowed to have two things going through the same basic process control system because it's independent enough.

So, kind of doubling down on that theme of what does it mean to be independent enough, we're going to talk about Clause 9.4 which is the last sub-clause in Clause 9 for allocation, and it is titled Requirements for Preventing Common Cause, Common Mode, and Dependent Failures.

So, Clause 9.4.1 states, the design of the protection layers shall be assessed to ensure that the likelihood of common cause, common mode, and dependent failures between two bullet points here, number one, the protection layers themselves, and number two, the protection layers and the BPCS. So, we want to assess the design to make sure that common cause failures are sufficiently low in comparison to the overall safety integrity requirements of the protection layers. The assessment may be made qualitatively or quantitatively unless 9 to 7 applies.

Okay, and there is a note, I'll just go ahead and read off the note, a definition of dependent failure is provided in clause 3.2.12. So, going back to the definitions. Alright, so, whether your protection layers have common equipment or not, you need to make an assessment that they are truly independent of each other and that there are no common cause, common mode, or dependent failures.

So, failure of one component causes failure of another component. You could go into the definition 3.2.12. I recommend going back there and taking a look at it here.

So, you want to make sure that there is no dependency, no common cause, no common mode, or that the common cause is sufficiently low.

So, for instance, if I have a flow measurement for my BPCS loop and I have another flow measurement for my alarm, there is some potential for common cause failure because even if you have separate physical taps for those two transmitters, those taps can still get plugged at the same time for the same reason because they're in the same process service even though they don't share the same physical taps.

So, for every protection layer we're going to be making this assessment and this assessment usually is going to occur either during the layer of protection analysis or during a step that commonly occurs after the layer of protection analysis. Which is commonly referred to as IPL rationalization.

## Clause 9.4.2 Assessment Factors and IPL Rationalization

So, you get done with your LOPA, you have your list of all your protection layers. When you do your rationalization you're going to ask yourself more detailed questions to confirm and validate that your IPLs are credit worthy and do all your assessment of common cause, common mode and other items.

So, we make this really easy for you in the Kenexis Open PHA software. So, Kenexis Open PHA you're going to have that database of all of your safeguards. You can set to filter so that you're just looking at your independent protection layers. And we have fields available. So, if you expand all the fields that are available for independent protection layers while you're doing this task, you'll have separate check marks to make sure that it's independent, independent, that it's specific, that it's auditable, that common cause and common mode and dependent failures don't exist.

So, you go through that process and you could do that while you're in the LOPA, so, you know, while you're defining the safeguard in the LOPA for the first scenario where you run into it or you do it at the end after you know what all of the scenarios are. And when you look at an open PHA database, when you pull up that safeguard list, it's also going to give you a list of all the scenarios in which that safeguard was referenced. And you can click on the hyperlinks to go back into the database and actually see the scenarios.

So, the whole concept of IPL rationalization is very common. There's a lot of people doing it in industry and generally it's something that's done with a smaller group of more instrumentation and control dedicated people after the LOPPA study has been completed. Okay, so you need to do that IPL rationalization task and it needs to include assessment of common cause, common mode failures between the other protection layers in each of the scenarios in which that safeguard is used and also the initiating event especially if the initiating event is in the basic process control system.

The last clause in this section is going to be 9.4.2 which states the assessment shall consider the following.

So, 9.4.1 says you need to do an assessment. 9.4.2 gives you four bullet points that define what that assessment needs to include.

Bullet point one, the independence between the protection layers. So, physical and functional independence between the protection layers.

Two, the diversity between the protection layers. So, that's kind of looking at what is the susceptibility to common modes of failure even though you have separate taps, maybe both of those taps are connected to the same process service and vulnerable in the same way. So, diversity is bullet point two. Bullet point three is the physical separation separation between the different protection layers. So, physical and functional separation, bullet point three is the physical separation.

Do you have separate taps? How far away are the components from each other or are they mounted in exactly the same location? You need to think about that.

Finally, bullet point number four is common cause failures between protection layers and between protection layers in the BPCS. So, just think about what the potential common cause failure mechanisms are and make sure that they don't put an undue stress on the system.

There are two notes to this clause. Note one says common causes from the process can be addressed. Plugging of relief valves may cause the same problems as plugging of sensors of the SIS. So, think about what are the failure mechanisms and are they common cause mechanisms between protection layers?

And, there's an example of a common failure mechanism between relief valves and a safety instrumented system.

Note two states the independence of physical separation, independence and physical separation can be addressed. So, human machine interface, SIS, VPCS networks or bypass means are another place where you can get common cause failures.

So, think about if you have a single switch that's going to bypass multiple devices, that's accidentally putting something in bypass or leaving it in bypass longer than it should. You just created common cause failure. So, these are all things to think about. We recommend doing this in open PHA. We recommend doing it during an IPL rationalization portion of your LOPA study or during your LOPA study.

We give you the tools, the mechanisms to do this elegantly, cleanly, and then have that information cleanly carried over into your vertigo study where you're going to be doing your SIS safety life cycle analysis.

## Episode Sign-Off and Vertigo Software Overview

And with that, we have come out to the end of clause nine. Next week, we're going to be looking at clause ten. Clause ten is going to take a lot of time. Clause ten is going to cause me to complain because I was taught how to write specifications at UOP. Ron Vangelisti, Jerry Green, these guys gave me hard and fast rules that the standard violates hard.

So, I'm going to be up on my soap box. And even though clause ten is a very short standard, I don't know how many weeks I'm going to spend talking about clause ten. Because there's so much to it. And one clause or one bullet point can end up taking 20, 30 minutes of discussion to get into it in a sufficient amount of detail.

So, we will start the process of SRS next week and we'll keep plugging through for as many weeks as it's going to take to get us to the end. Thank you and we will catch up with you next week.

Now that you've heard some insights on technical safety, functional safety, and the IEC 61511 standard, let me tell you a little bit more about how to easily and effectively implement the safety life cycle using the Kenexis integrated safety suite and our SIS safety life cycle management tool Vertigo. Vertigo is a comprehensive tool set for performing assessment calculations, documenting, and maintaining the design of safety instrumented systems.

Analysis begins with importing or synchronizing a list of safety instrumented functions with their definitions and associated performance targets from our open PHA tool for HAZOP and LOPA documentation. Each safety function can then be analyzed by performing a SIL verification calculation, complete with a collection of tools for optimizing designs and a database of thousands of potential instruments to define failure rates and diagnostic coverage capabilities.

After the SIL verification calculations are defined, you can build an SRS by automatically generating a cause and effect diagram from the SIF definitions and other defined instruments. Each SIS instrument will include a customizable data sheet and general requirements that are applicable to the SIS as a whole and can be entered individually or even bulk imported from customizable libraries.

After the design phase, you can even use Vertigo to track and document testing throughout the entire life of the facility. Kenexis

Vertigo is the most integrated, easy-to-use enterprise tool for allowing the development of SIS design basis information more efficiently and effectively than any other software application. Thank you.