In this episode of the Functional Safety Podcast:

How should a safety instrumented system respond when it detects its own failure? The answer, Ed Marszal explains, is however you design it to respond — and that design choice carries far more engineering weight than the standard’s brevity suggests. This episode dissects Clause 11.3’s four normative sentences, tracing the arc from automatic compensating measures through redundant voting arrangements to the elaborate human-centered alternate protection plans required when no redundancy exists. The centerpiece is a detailed sulfur recovery unit war story that finally illuminates the standard’s most puzzling sentence: when a bad-PV alarm forces an operator to take process action, that alarm itself becomes part of the SIS. Ed closes with the easily overlooked proof-testing mandate of Clause 11.3.2 and a candid assessment of how often diagnostics are claimed in calculations yet left unimplemented in practice. For engineers who have ever stared at 11.3 and wondered what it actually demands, this is the episode that connects the text to the plant floor.

How should the safety instrumented system respond to a detected failure?  Well, it will respond however you want it to respond…. Join Ed as he breaks down clause 11.3 and provides his thoughts and expert analysis of this topic in more detail.

Tune in to the latest episode of the Kenexis Functional Safety Podcast, hosted by Ed Marszal, President and CEO of Kenexis. Now available on Spotify and Apple Podcasts, Ed offers his expert insights on the IEC 61511 standard.

With decades of experience in safety instrumented systems and as a Principal Engineer, Ed has a unique perspective to offer. He has been an active contributor to the ISA 84 committee since 1994, adding to his deep understanding of the field.

In this inaugural season, Ed delves into the IEC 61511 standard, unpacking the meaning behind each word and providing a thorough interpretation of its application. Through personal stories from his career and committee work, he offers valuable context and insights for professionals in the industry.

Full Episode Transcript

Episode Teaser and Introduction

How should the safety instrumented system respond to a detected failure? Well, it will respond however you want it to respond.

Welcome to the Kenexis Functional Safety Podcast. I’m your host, Ed Marszal, President and CEO of Kenexis. Kenexis is a technical safety consultancy that helps chemical process industry companies to analyze risk and design engineered safeguards like safety instrumented systems and fire and gas detection systems. Kenexis also provides the industry-leading suite of software tools, including our best-in-class Vertigo software for SIS Safety Lifecycle Management.

In this first season of the podcast, we are going to focus on the IEC 61511 standard, doing a deep dive into the standard, including more depth of information on what the standard means and how to apply it, brought to life with personal war stories and behind-the-scenes discussions of the committee members as we develop the standard in ISA 84 and IEC SC 65.

Before we start, a little disclaimer, I will be providing my opinion on technical and engineering topics. This information is provided on a best effort basis and is of a general nature. The information presented in this podcast might not be applicable to your specific application. It is the obligation of every engineer to thoroughly analyze any system that they are designing and not blindly rely on any general advice presented in this podcast.

All right.

Clause 11.3 Overview and Diagnostic Context

So what are we talking about today? Today we are talking about requirements for system behavior on detection of a fault. We are hitting Clause 11.3. We’re going to go all the way through Clause 11.3, which is dramatically shorter today than it was when the standard was originally written. So the 2003 version of the IEC 61511 standard, this is like about two pages. Now it is two clauses. The normative clauses total up to three sentences. And then you’re going to get two informative notes.

So we were able to condense all of the information in the previous version of the standard in just a few requirements. And still, the 2016 version still conveys exactly the same information as the 2003 version. It just does it in a much more concise way.

So starting off, Clause 11.3, the title, Requirements for System Behavior on Detection of a Fault. So with safety instrumented systems, they are, I don’t want to say unique, but an attribute of a safety instrumented system is that there are going to be diagnostics. That continuous, high-speed, automatic, online testing that’s occurring to make sure that your safety system equipment is functioning properly. And if it’s not functioning properly, you know that it’s not functioning properly and are able to do something about it.

Now, Clause 11.3 talks about the system behavior. And usually when I hear the word system behavior, I’m thinking about hardware. You know, what are the specs for how the hardware is supposed to behave when this type of action happens? But honestly, in this case, a lot of it has to do with how the human beings are supposed to respond as opposed to how the safety instrumented system is supposed to respond. But there is a big portion of how the safety instrumented system is supposed to respond also.

Now, even though there are only, you know, three sentences here in these clauses, there’s a lot to talk about because we need to talk about kind of what are the diagnostics? How do they work? How do you know that a diagnostic has been activated? Where is the test actually occurring? So there’s a whole lot of complexity in the engineering of diagnostics for the end user.

Now, obviously, there’s a lot of effort required from the equipment vendor who needs to actually get their equipment to perform the test. But on the integration side, how do you know that a component is sick? How do you know that a component is trying to tell you that it’s not working anymore? And what do you do with that information? That’s kind of the configuration on the end user part that’s going to have a lot to do with our discussion today.

Also, on top of that, there’s a lot of human actions that are required. So a device can put up an alarm. Maybe it’s sending smoke signals, which is generally not the mechanism that we’d prefer for communication of diagnostic failures. But it may be a natural outcome of whatever happened to the device. But when that communication happens, what do you do? What do you do from the system side? What do you do from the configuration side? What do you do from the maintenance side? What do you do from the operation side? There’s a lot. There’s a lot going on.

And we’re going to dig into all of it over the course of this podcast. So let’s begin.

Clause 11.3.1 Text and Compensating Measures Definition

So there are two clauses, 11.3.1 and 11.3.2. And we’ll start off with 11.3.1, which says, when a dangerous fault in the SIS has been detected. Then we’re going to go off into a parenthetical statement that says, you know, it gives a list of the ways that you can detect a dangerous fault. So that could be through diagnostic tests. But it could be also happening through manual proof tests. The standard just says proof tests. But we take it, you know, when we hear proof tests, we mean, we’re thinking that that’s going to be the manually executed proof tests. Or by any other means.

So any other way for you to know that an SIS component is in the failed state. So we start out by saying, when a dangerous fault in the SIS has been detected, then compensating measures shall be taken to maintain safe operation. So right there is our first sentence. Actually, it’s going to total up to four. I missed a period as I was doing my counting earlier. When a dangerous fault in the SIS has been detected, compensating measures shall be taken to maintain safe operation.

Okay.

So right out of the gate, we’re going to be digging back into some definitions. We’re going to be focusing on some terms. So a fault would be any situation where a SIS component will not be able to perform its function as designed.

Now, that’s not really too dramatic. The thing here that is a little bit more nuanced and a little bit more niche is going to be compensating measures. So compensating measures is a term that is used in the standards. Way back in the day, we at Kenexis started off using a term called an alternate protection plan. So an alternate protection plan could be part of compensating measures, shall we say. Because when we say an alternate protection plan, that’s going to assume that there’s going to be some sort of human involvement in how you do things or what those compensating measures are.

We’re partially going to be using humans as those compensating measures. But compensating measures don’t necessarily require human beings to be involved.

Types of Compensating Measures

So what is a compensating measure? A compensating measure is this. The safety instrumented function is unavailable to perform its action. It’s not okay just to say, oh, well, it’s not going to work. We better go fix it. And to fly blind for the time period when that device is in the unavailable state.

Instead, when a safety instrumented function becomes available, and we know that it’s unavailable because there was some sort of enunciation of a fault, it is incumbent upon us to replace that failed safety instrumented function with other equipment, other humans, other activities that will take the place of that component for the duration of its unavailability.

Now, this can be done many different ways. Now, if, for instance, I have a two out of three voting arrangements and one of the devices fails, well, the compensating measure is I’ve got two other devices that are doing exactly the same thing. And if one of those two devices fails, then I will have my majority vote and I will take my safety action. So in a lot of cases, these compensating measures are going to be equipment.

Now, something that is much more rare. So using the redundant component is very common. And you can use this in a variety of voting arrangements. Even a two out of two vote, or I’m sorry, a one out of two vote, if you know that one of the two devices is in the failed state, instead of just tripping out, you can degrade to a one out of one voting arrangement. It’s a very common, perfectly acceptable option that you should keep in your bag of tricks, if you will. So that’s a very standard compensating measure.

Now, another automatic compensating measure that is a little bit more esoteric, it’s a little bit more out there, would be to communicate a basic process control system variable to the SIS to replace a failed SIS measurement for the duration of its inactivity. And similarly, you can, you know, maybe send an output to a basic process control valve when we know that an SIS valve is unavailable. So those are all automatic compensating measures that don’t require any human intervention.

But, and the ones that I just presented to you, the, you know, the redundancy is kind of a slam dunk done by everybody all the time. Trying to replace the failed SIS component with something that’s existing in the BPCS is, it’s so uncommon that I would call it not just rare, but very rare. But I have heard it done. And, you know, being in the business as a consultant for 30 years, I’ve heard one of just about everything.

So that’s kind of a, another option to think about maybe, but anyway, the other type of compensating measure, if you have no redundancy in your SIS to fall back on and you can’t do the compensation automatically, is manual compensation. So if I have a slow moving process variable, like a level, and my level device, uh, diagnosis that it’s in the failed state, would I actually want to shut down or would I want to use compensating measures? So that one level transmitter and a one out of one vote is in the failed state. Uh, I want to compensate for it using a human.

Manual Compensating Measures and Alternate Protection Plans

So in that case, you’re going to need to put together your alternate protection plan. Now, if you’re using really, really good, uh, SIS safety life cycle software like Vertigo from Kenexis, then when you go into your, the, the bypass authorization section of Vertigo, if you’re doing a bypass of a component that where there is no redundancy, you’re going to be asked to fill out a alternate protection plan form or compensating measures form. And you’re going to need to think through a lot of things as you’re going through that process.

You’re going to need to number one, uh, explain, well, what measurement am I going to look at in lieu of the failed measurement? Who is going to be looking at that measurement? Uh, at what measured value of that alternate measurement are you going to need to take action?

So since it’s a human being, you’re probably going to want to give them a little bit more time to react than an automatic system would. So the set point’s probably going to get a little bit further or, uh, further away from the dangerous condition. Um, and then you’re going to need to say, okay, that person who’s doing that activity, are they sufficiently independent from normal work activity? So I don’t want somebody to keep an extra eye on it. Nobody has an extra eye. Well, maybe somebody does, but that would be kind of a freaky thing.

Um, so we’re basically want to have, going to have like a dedicated operator who’s going to be looking at that process variable.

We’re going to want to ask ourselves, can they respond in the process safety time? Otherwise you probably ought to just be shutting down because then at that point in time, your compensating measures are not really expected to work. So why are you even bothering with this? Um, and going beyond that, um, so that’s kind of on the measurement side. Then you’re also going to need to say, okay, well, if that measurement is violated, what action is going to be taken? So, and who is going to take that action and how is it going to be communicated?

So maybe that outside operator is going to walk over to a manual valve and manually turn it with their hands. Maybe they need to call to another operator over the radio and get them to go to a manual valve. Maybe they will call back into the control room and have the control room put a valve into 0% from, from the board because it’s a basic process control valve. All of that needs to be documented. And then all of that documentation needs to be kind of signed off and approved by the people that are implementing the bypass.

So it sounds like a lot of work. It is a lot of work. But, uh, if you’re using Vertigo, if you’re one of the lucky ones using Vertigo, once you enter all that information in, in terms of compensating measures for a device, it stays in the database forever. So next time you need to put that, uh, device in the bypass, all that information has already been filled out and you’re going to be good to go. So those are the compensating measures that you need to take to maintain safe operation.

Shutdown Option and Clause 11.3.1 Second Sentence

Now, after reading this first sentence, it sounds to me like the standard is trying to tell you that you should just keep going. You should just continue operating, continue letting the plant go. Well, you don’t have to do that. Um, which is kind of in the second sentence. Um, and in the second sentence, it kind of makes it seem like it’s a, a second choice or a suboptimal option, but shutting down is always an option.

If you detect that that level transmitter in the previous example failed, you always have the option to just shut your plant down, bring it to a safe state. That’s going to be the second sentence in the clause, which says, if safe, safe operation cannot be maintained. Now that’s the first clause. Um, I would kind of change that to say if safe operation cannot be maintained or you don’t want to maintain it. It’s not that you can only shut down if there’s no way for you to continue to operate. You always have the choice to shut down.

Um, and for a lot of operating facilities, a lot of plants, that is going to be the preferred, that’s going to be the preferential option, uh, is to take that, that type of action.

So if safe operation cannot be maintained or you don’t want to, a specified action to achieve or maintain a safe state of the process shall be taken. Okay. So continue to operate, but make sure that you have compensating measures in place to maintain the same level of safety as you had when the safety instrumented function was operational.

If you don’t want to do that or you can’t do that because, you know, maybe there is no effective compensating measure, then you shut the plant down. Well, shut the plant down is the colloquialism. The standard says, take an action to achieve or maintain a safe state.

Third Sentence – Alarm as Part of the SIS

Now, the third sentence here is the weirdest of them all. And it takes, uh, the most, it, it, it, it involves the most head scratching. It involves the most looking at the standard and wondering what the heck are these guys talking about? Why are they even stating this? Okay. So everything that we talked about so far in the clause says we are going to invoke compensating measures.

And if you can hear weird things going on in the background, Sammy, Sammy, my dog was outside and I just had to let him back inside my 16 and a half year old puppy, uh, who I, you know, take extra special care of because he’s such a delicate senior citizen.

Um, anyway, so, so you continue to operate with compensating measures or you shut down the plant. Now, if you’re going to continue to operate with compensating measures, I’ve already explained quite elaborately that that’s, that’s a big deal.

There’s, there’s a lot to the compensating measures, but there’s even more to it with this last clause. All right, let me go ahead and begin by just reading this clause and, and, and, uh, getting into it. So where the compensating measures depend on an operator taking specific action in response to an alarm, for example, opening or closing a valve, then the alarm shall be considered part of the SIS. Okay. That’s confusing. It seems kind of like, I don’t know, there’s just words out there that they’re saying for no reason. I don’t actually understand what they’re saying.

Well, no, I, I do, but you might not understand what they’re saying, uh, with regards to this clause. So what are the compensating measures that you’re going to need to take? Now, this last sentence says the compensating measures require you to take a specific action. And then it gives the examples of opening or closing a valve. So if the bad PV alarm, so that alarm that says my level transmitter has faulted, my temperature transmitter has faulted, whatever the measurement is, it has faulted. What is the normal actions that you will take?

The normal action that you will take is you will contact maintenance to get the device fixed. No big deal. That seems fairly easy. That calling maintenance to get the device fixed is not opening or closing a valve. It’s not starting or stopping a pump. It’s not some sort of action that you need to take with respect to the process.

So why would a bad PV alarm ever require you to take some sort of action on the process? Uh, so this actually, um, some of you who are listening to the podcast may know one of my long time colleagues and just, you know, straight up friends, Dennis Zetterberg from, uh, formerly from Chevron. Now he is still doing work on the IEC 61511 committee, uh, in an, uh, early, I dare say retirement from Chevron, but he’s still very active in standards committees. Now,

Dennis was part of the team with me that helped to write the ISA, uh, certification exams. And as we were going through the certification exams, Dennis looked at this clause and said, I don’t understand why you would ever do this. I don’t know of any example of where this is applicable. And everyone around the table, we couldn’t figure out why this clause was here either. So, you know, what, when does a bad PV alarm require you to take some sort of action? And then a couple of years passed and I finally came up with an example of, oh, that’s what this clause means.

Sulfur Recovery Unit Example – Process Background

So let me give you an example that’s going to fire this clause and tell you what the ramifications are. Many of you are in or familiar with the oil and gas business, specifically oil refining. In oil refining, you are generally cleaning up crude oil, uh, on the way to selling valuable fuels. And due to various regulatory environmental requirements, we need to get every last little tiny bit of sulfur, uh, out of that fuel. And as a result, um, there’s a lot of sulfur. I’m going to call it waste. Let’s call it a brought a byproduct. Sulfur becomes a very common byproduct of the refining process.

Now, while we’re removing the sulfur from the fuels, there’s kind of an intermediate step where it gets turned into hydrogen sulfide, which is a wildly toxic gas. And, uh, we will take that hydrogen sulfide and basically carry it all to one portion of the plant called the sulfur recovery unit.

Now the sulfur recovery unit operates by burning the hydrogen sulfide to make, uh, sulfur dioxide and water. And while sulfur dioxide’s not nearly as toxic as hydrogen sulfide, but it’s what we don’t want to avoid to the atmosphere. So we can’t just release it to atmosphere. Otherwise, what, what was the purpose? Now in the sulfur recovery unit, after you burn the H2S and create the SO2, it’s going to go to a reaction section where the sulfur is the, the oxygen will be removed from the sulfur and you end up with elemental sulfur and oxygen.

And this is because we took our hot SO2 and passed it over a catalyst.

Now, as many of you know, this process is very commonly the bottleneck for most refineries because, you know, we built our plants, you know, 50, 100 years ago, and we keep expanding them. But, you know, the sulfur recovery units don’t make money. Uh, so we try to squeeze as much out of them as we can. And, uh, you know, we, we usually end up building new ones, but there, it’s always a bottleneck. So

I’m trying to shove more H2S through my Klaus reactor system, my, my sulfur recovery unit. How can we do this? So I’ve got, you know, my pots and pans are the size that they are. How can we squeeze more H2S through this system? Well, the answer is instead of just burning the H2S in air, why don’t we spike our air with oxygen? And so instead of 21% oxygen, maybe we’re at 30, 35% oxygen. I can get a lot more gas through my sulfur recovery unit without having to buy more pots and pans. Everybody is happy.

SRU Catalyst Hazard and High Temperature Safety Function

Well, all right, let’s talk about one of the hazards in the sulfur recovery unit. One of the hazards in the sulfur recovery unit is the catalyst on the back end of the Klaus reactor can catch on fire. And the reason that it will catch on fire is well, there’s sufficient free oxygen to allow it to catch on fire. And if it catches on fire, you may have a hazardous destruction of your SRU. It’s kind of a stretch. The reality is what we’re really worried about is when it catches on fire, I’m destroying my catalyst.

If I destroy my catalyst, I can’t push any more oil through my refinery until I fix my SRU unit. And this is a ginormous business interruption. Okay. So in order to deal with that, I have a high temperature shutdown. So if the temperature is significantly higher than it’s supposed to be, I’m going to shut my unit down. Okay. So I’ve got a safety function, more of a financial protection, but let’s call it a safety instrumented function that will shut down the plant on high temperature. My measurement is usually a thermocouple, which are not the most reliable pieces of equipment.

And it’s at insanely high temperatures, which is going to make it even less reliable. And when this safety function activates, I shut down myself a recovery unit. So at a minimum, I need to reduce rates through the refinery. I might need to shut my whole refinery off. So I don’t want to shut this unit down spuriously for any reason. So when I detect that the thermocouple has failed, and you know what, it’s really easy to detect the thermocouple has failed because you’re going to get an open circuit and your voltage is going to go to zero.

So the diagnostic coverage is between 90 to 100% on a thermocouple failure. Can you get a coverage that’s more than 100%? This is one of those diagnostics that is slam dunk. So if I detect that my thermocouple went open circuit, do I want to shut down my plant?

Absolutely not. So what am I going to do? I’m going to put in compensating measures, or maybe I could have just done a two out of three votes. But trying to put an extra two thermal wells into a reactor is a monumental cost undertaking also. So if you didn’t already have three thermal wells for a two out of three vote, probably not going to want to do that. And the SIL target on this is relatively low. So generally, when we detect that this failure occurs, and it’s going to occur, you know, it might occur like once a year, once every other year, you don’t want to shut the plant down.

You want to continue operating, you want to put in place compensating measures. And this is another one of those places where the hazard takes a while to develop. There are other measurements in the basic process control system you can use. It’s a pretty good case for saying I’m going to put in a manual alternate protection plan. But that pesky oxygen that we’re putting in really dramatically increases the risk of this operation.

So for those people that are spiking their air with extra oxygen, for the time period where that high temperature shutdown is in place, they’re going to want to go and cut off the oxygen spike and just go back to regular air.

Oxygen Spike Action and Alarm Classification as SIS

Here we go. There it is. Finally, we found that situation where the bad PV causes me to take a process action. So in this application, when the temperature in the reactor showed a bad PV, we’re going to manually close off the oxygen spike. So what did the standard say?

Where the compensating measures depend on an operator taking a specific response to an alarm. That’s exactly what’s happening here. Okay. When that is the case, the alarm shall be considered part of the SIS. What does that mean? Well, kind of make it a long story short, and it’s kind of written in a vague, nondescript kind of way that can get a lot of different people to interpret it different ways. But basically, if the alarm is part of the SIS, that means it’s not in the BPCS, or at least it’s not in the BPCS alone. And this action should be possible even if the BPCS has failed.

So we’re not going to rely on the basic process control system to enunciate this alarm for the operator to take the compensating measures. So the bad PV alarm from my temperature transmitter, I’m not going to say it can’t go through the basic process control system, but it can’t go through the basic process control system alone. You can repeat it, but the BPCS cannot be the only way that it shows up. So what would I recommend that you do? Well, different people do different things.

There are some operating companies that have an operator interface that is dedicated to the safety instrumented system. If you want to do that, that will meet the requirements of this clause. And that activation is all part of the safety instrumented function all the way up through the human machine interface that you’re using for that system. A little bit of an easier way to do it. And those of you who know me have run into me on a variety of my speaking engagements know I am a fan of the good old-fashioned panel arm, the old light box in the control room. I’ve got all kinds.

I could probably, you know, do a two-hour presentation about the value of light boxes because they are so much better tuned to human beings, pattern recognition. I’ve got all kinds of good war stories. But so basically, in this case, what we would want to do for that bad PV, so if I get a thermocouple burnout on the temperature measurement, the transmitter should indicate that I’ve got thermocouple burnout and be able to communicate that to the SIS. Now, how does that communication occur? It’s a configuration in your sensor.

But probably the most common best way to do it is for the transmitter to set itself to a fault output. And the most common fault output made popular by the NAMUR standards out of Germany is to set it to 3.7 milliamps. So a valid signal is 0 is 4, 100% is 20 milliamps. Most transmitters are going to have a saturation above and below to where they’ll actually work between 3.9 and 20.8. So they could go a little bit negative and a little bit above 100% and still give you a valid signal.

If you go down to 3.7, there’s no reason that you should be at 3.7 milliamps other than the transmitter is trying to tell you that it has failed.

So if the SIS logic solver reads a fault condition from the transmitter, maybe it’s 3.7 milliamps, maybe you’re using digital communications, maybe you’re piggybacking a heart signal on your 4 to 20. However, the SIS logic solver knows that you have a broad bad process variable, that should then go to a SIL rated effectively enunciation system. So you’re going to, and it would be basically at the same level as the safety instrumented function that you’re replacing.

In this case, it’s either SIL 1 or might even be SIL 0 because it’s not necessarily safety related, but let’s just say it was SIL 1. And that is going to meet all of the requirements. And you would have to understand that that light box is part, it’s basically effectively kind of like a final element of your safety instrumented system that needs to be tested and maintained, uh, meeting all of the other requirements of the IEC 61511 standard. So that last sentence, let me repeat it again.

Um, where the compensating measures depend on an operator taking a specific action in response to an alarm, e.g. opening or closing the valve, then the alarm shall be considered part of the SIS. So if you were ever wondering what that clause meant, well, I just gave you like a 10 minute explanation with an example of where that becomes the case. All right.

Note 1 and Note 2 – Factors Affecting Compensating Measures

Now there are a couple of additional notes that are part and parcel of, uh, this clause also, which, uh, Sammy, my dog keeps going in and out. So I keep having to stand up and open and close the door for him. I apologize for that. Uh, note one, uh, to clause 11.3.1 says the specified action, fault reaction required to achieve or maintain a safe state of the process can be specified in the SRS. And that’s going to be in clause 10.3.1, 10.3.2. It can consist of the safe shutdown of the process or of that part of the process, which relies on the faulty SIS for risk reduction. So that note is,

I would kind of argue a little bit of motherhood and apple pie. Uh, yeah, you’re, you’re generally going to trigger the safety function that the failed component is part of as though, uh, there were a vote to trip in place is generally what you’re going to do. But since it’s a bad PV, you know, you’ve got all kinds of flexibility to do all kinds of other things. So, you know, as is common for the SRS section, you just decide what you’re going to do, think it through, write it down. And that’s, that’s basically it. It’s part of your application logic.

Note 2 is a little bit more elaborate than note 1. Note 2 says, the compensating measures required for continued safe operations can depend on the safety integrity requirements, the tolerable risk associated with the event, the hardware fault tolerance of the SIS, the anticipated MRT, and the availability of any other layers of protection. So that first sentence of the note is kind of telling you what attributes you need to think about when you’re doing your risk analysis of and writing your alternate protection plan for those compensating measures. Let’s kind of hit them one at a time.

So compensating measures are going to depend on safety integrity requirements. Replacing a high-speed SIL-3 safety instrumented function with a human is probably not realistic. You should probably just go ahead and shut your plant down if that’s the case, not even bother with these compensating measures. So the higher the SIL, the harder it’s going to be to replace. SIL-1, slow-moving process might be something where you can replace it manually.

SIL-3, maybe not so much. So SIL, how much risk reduction is required factors into it. What else factors into it? The tolerable risk associated with the hazardous event. Now, so now instead of just looking at the safety function itself, we’re looking at, well, what is the consequence if the hazard that’s protected against by the safety instrumented function occurs? And even maybe just looking at it from a consequence-only perspective, what is the outcome that you’re going to be doing?

But of course it says tolerable risk, so you’re going to look at kind of all attributes to it, and we’ll get to that.

Okay, continuing on. Compensating measures depend on hardware fault tolerance. So one out of one voting arrangement is going to require, you know, human alternate protection plans or compensating measures, where if you have a two out of three vote, you’ve got compensating measures effectively built in. So the hardware fault tolerance is going to tell you whether or not you have built in compensating measures depending on how you do your configuration.

Next item we need to think about is the anticipated MRT. So asking an operator to look at a sight glass for an hour is a pain, but it’s doable. Asking an operator to watch a sight glass for 72 hours hours is absolutely unrealistic. Maybe it’s probably even unrealistic with multiple shifts. So the longer it takes to repair something, the more of a case you have to say, I don’t believe human compensating measures are going to be reasonable here. And then finally, the last thing that the note talks about for making these consideration is what are the other layers of protection?

So if your safety function is backed up by redundant relief valves, then maybe not having the safety function in place for a little while, maybe that’s not that big of a deal. So what are all those other IPLs that are in the LOPA scenarios that the safety function that you’re taking out of services in, that all needs to get factored into what compensating measures look like?

Okay, moving along, there’s another sentence in the note here that says, in some cases, it can be adequate to ensure action is taken to ensure repair of the dangerous failure within the assumed MPRT in the calculation and the PFD average. Excuse me. But in other cases, it can be judged necessary to provide other measures to compensate for the reduced risk reduction until the SIS is fully restored.

All right. So this note has said MRT and MPRT. Let me begin by clarifying what those items are. So MRT is mean repair time. And in actuality, it is how long you think it will actually take for this repair to happen, which is going to be different from MPRT, which is a maximum limit, kind of conservative thing that honestly, this is what your calculations are based on. So when you say repair time is 72 hours, you don’t actually think it’s going to take 72 hours. You’re giving yourself a maximum limit.

So, and I could probably get this repair done in an hour and a half, but I’m going to give myself 72 hours because if it happens at five o’clock on Friday, I’ll let day shift on Monday, deal with it. That kind of thought process.

So MRT, how long you actually think it’s going to take to fix MPRT maximum permitted repair time is the absolute do not exceed limit for the time duration making that repair. So what the second sentence of this note says is in some cases, just making sure that you repair the component inside the MPRT is going to allow risk to be satisfied. And if you have a two out of three voting arrangement, that is absolutely the case. No additional compensating measures, no human alternate protection plan is going to be required.

But sometimes you do need to take other additional actions and that’s all part of that risk analysis workflow. Okay, so that was clause 11.3.1. We’ve got another sentence to close out clause 11.3.

Clause 11.3.2 – Bad PV Alarm Proof Testing Requirements

One more sentence, which is clause 11.3.2. And it reads, where any dangerous fault in an SIS is brought to the attention of an operator by an alarm, then the alarm shall be subject to appropriate proof testing and management of change. Huh? So the bad PV alarm needs to have it. Now, this is different. So, so in clause 11.3.1, we talked about when the bad PV alarm needed to be part of the SIS. Okay, this clause 11.3.2 is broader than that.

It is, if a bad PV alarm, and that would be dangerous fault in an SIS brought to the attention of an operator by an alarm. So anytime I have a bad PV alarm, it needs to be subject to appropriate proof testing and management of change. So your alarm, master alarm databases should have all these bad PV alarms. Uh, the response to the bad PV alarms should be documented. Uh, the response to the bad PV alarms should be documented. And when you are performing your periodic proof tests of these alarms, you should generate a bad PV signal and verify that the alarm comes through the operator.

So, uh, kind of a, a, a subtle item that, you know, if you go and look at a lot of proof tests, proof test templates, this clause isn’t in there because it takes, uh, a weird action like lifting a wire to be able to actually create a bad PV. Um, that might not happen during the course of a test that you would normally do. If generation of that bad

PV alarm does not happen during the normal course of testing, then you’re going to need to write additional steps into your test plan of all of your SIS components where bad PV, uh, alarms are critical to the operation of the plants, meaning you don’t shut down automatically based on the bad PV. Um, because if you do shut down automatically based on the bad PV, then there’s going to be a shutdown alarm. It’s not going to be a mystery to anyone, uh, what happened.

Uh, but in this case, uh, yeah, if you’re not shutting down based on a bad PV, then that bad PV alarm needs to be tested as part of the manual proof test process.

Summary and Hardware Fault Tolerance Preview

All right. So a lot of information there, a lot of good information there. We went through clause 11.3. It was only four sentences. Uh, but we’re, you know, we’re on our way to an hour. We’re a little bit over 51 minutes, uh, into the podcast right now. So, uh, a lot was talked about.

Uh, it’s honestly response to detection of a fault is one of those things that is, doesn’t get the attention that it should. A lot of times diagnostics just don’t get implemented because engineers don’t understand that they need to manage the communication of faults from the sensors through the logic solvers out to operations and maintenance. And sometimes, you know, those fault alarms just, they don’t go anywhere. Nobody knows about them.

And even more sadly, most people took credit for those alarms in their calculation, even though they’re not actually doing anything with them, which is a dangerous behavior that we’re going to need to, uh, put the kibosh on going forward.

Okay. That is all that I have, uh, for this episode. The next episode, we are going to get into the curious topic of hardware fault tolerance. And when we do this, I don’t know if I’m going to be able to get through all of hardware fault tolerance in one episode or whether I’m going to need to break it into two. Cause we’re going to need to talk about the new way. We’re going to need to talk about the old way. We’re going to need to talk about the IEC 61508 way. And the IEC 61508 way is the 1-H way and the 2-H way.

Okay. So there’s a lot to discuss. There’s a lot of tables. There’s a lot of exclusions. There’s a lot of clauses that get out and get you out of jail for free, but I don’t think I’ve seen anybody ever use them. And then you get the weird clause 11-4-8 that seems like it should get a lot more attention, but it’s tucked away. It’s hidden so that nobody really ever even sees it. And I’m sure people have gotten themselves into trouble. And then clause 11-4-9, oh my gosh. We’re going to talk about the chi-squared factor. We’re going to talk about confidence limits.

So that might, I might be able to squeeze all of this into one episode. It might get broken into two, but regardless, next episode, get ready. We’re going to be talking about hardware fault tolerance, why it’s there, and what to do about it. See you next week.

Kenexis Vertigo Software Overview

Now that you’ve heard some insights on technical safety, functional safety, and the IEC 61511 standard, let me tell you a little bit more about how to easily and effectively implement the safety life cycle using the Kenexis integrated safety suite and our SIS safety life cycle management tool, Vertigo. Vertigo is a comprehensive tool set for performing assessment calculations, documenting, and maintaining the design of safety instrumented systems.

Analysis begins with importing or synchronizing a list of safety instrumented functions with their definitions and associated performance targets from our open PHA tool for HAZOP and LOPA documentation. Each safety function can then be analyzed by performing a SIL verification calculation, complete with a collection of tools for optimizing designs and a database of thousands of potential instruments to define failure rates and diagnostic coverage capabilities.

After the SIL verification calculations are defined, you can build an SRS by automatically generating a cause and effect diagram from the SIF definitions and other defined instruments. Each SIS instrument will include a customizable data sheet and general requirements that are applicable to the SIS as a whole and can be entered individually or even bulk imported from customizable libraries. After the design phase, you can even use Vertigo to track and document testing throughout the entire life of the facility.

Kenexis Vertigo is the most integrated, easy to use enterprise tool for allowing the development of SIS design basis information more efficiently and effectively than any other software application.

Thank you.