In this episode of the Functional Safety Podcast:
The most debated topic in functional safety gets barely a page in the standard itself. Ed Marszal walks through Clause 11.9.1 and 11.9.2, the twin requirements that every SIF achieve its calculated failure measure by calculation and account for twelve specific factors — from architecture and voting to diagnostic coverage, common cause failures, and the notorious operator-response clause that committee politics will soon erase. Along the way he explains why human failures must stay out of these calculations entirely, why the “calibrated 2×4” test creates dangerous never-detected failures, and which subclauses the next revision of IEC 61511 will quietly eliminate. For engineers tired of black-box software and seeking the intent behind the bullet points, this is the episode that connects standard to practice.
When it comes to SIL verification calculations, we look at Clause 11.9. It’s appropriately titled ‘Quantification of Random Failures’…putting the focus on ‘random’ right there in the name.
Listen in for more information and thoughts on this important topic as this section is discussed in more detail.
Tune in to the latest episode of the Kenexis Functional Safety Podcast, hosted by Ed Marszal, President and CEO of Kenexis. Now available on Spotify and Apple Podcasts, Ed offers his expert insights on the IEC 61511 standard.
With decades of experience in safety instrumented systems and as a Principal Engineer, Ed has a unique perspective to offer. He has been an active contributor to the ISA 84 committee since 1994, adding to his deep understanding of the field.
In this inaugural season, Ed delves into the IEC 61511 standard, unpacking the meaning behind each word and providing a thorough interpretation of its application. Through personal stories from his career and committee work, he offers valuable context and insights for professionals in the industry.
Full Episode Transcript
Introduction and Episode Overview
Clause 11.9 is the clause that talks about how we do SIL verification calculations. The title is Quantification of Random Failure. It’s got random right in the name.
Welcome to the Kenexis Functional Safety Podcast. I’m your host, Ed Marszal, President and CEO of Kenexis. Kenexis is a technical safety consultancy that helps chemical process industry companies to analyze risk and design engineered safeguards like safety instrumented systems and fire and gas detection systems. Kenexis also provides the industry-leading suite of software tools, including our best-in-class Vertigo software for SIS safety lifecycle management.
In this first season of the podcast, we are going to focus on the IEC 61511 standard, doing a deep dive into the standard, including more depth of information on what the standard means and how to apply it, brought to life with personal war stories and behind-the-scenes discussions of the committee members as we develop the standard in ISA 84 and IEC SC 65.
Before we start, a little disclaimer. I will be providing my opinion on technical and engineering topics. This information is provided on a best-effort basis and is of a general nature. The information presented in this podcast might not be applicable to your specific application. It is the obligation of every engineer to thoroughly analyze any system that they are designing and not blindly rely on any general advice presented in this podcast.
Okay, so today is the day where we spend the most time that we’re going to about a topic that has had more ink spilled over it than any other topic in functional safety, and that is SIL verification calculations.
You would think that with the attention that SIL verification calculations gets, that there would be hundreds of pages or dozens of pages at a minimum on what actually occurs, what is necessary to do a good SIL verification calculation.
But there is about a page. Yeah, when you put it all together, there’s about a page. And the standard actually never tells you how to do a SIL verification calculation. It tells you that you need to achieve your target. It tells you what you need to consider when you run your calculation. But how to actually perform the calculation is not part of the standard anywhere.
So we’re going to break clause 11.9 into a couple pieces. And I’m going to, because there’s a lot to unpack. But we’re going to, in this session, I’m going to want to talk about clause 11.9.1 and 11.9.2, which are basically what does the calculation need to consider. And then when we get toward the end of clause 11.9, we’re going to talk about some things that I would argue are not normative requirements. They’re most certainly informative. And then we’ll also talk a little bit about how to handle some of the uncertainty that shows up in these calculations.
And what some people do versus what I tell you you should do. And, of course, what I tell you to do is the most correct thing to do. All right. Let’s get into it.
Clause 11.9 Name and Random Failures
Clause 11.9 is called quantification of random failure. It doesn’t say SIL verification calculations. It doesn’t say PFD calculations. It says quantification of random failure. And I’m going to keep harping on that random failure piece of this.
Because we see a lot of value in running calculations that predict how well our systems are going to perform. But we also understand the limitations of those calculations, when they’re going to be effective, when they’re not going to be effective. And we don’t want to take excessive credit when we’re pretty confident that the calculation is not going to be effective. And where the calculations are not going to be effective is systematic failures of any kind.
And when we’re talking about systematic failures, the thing that we’re the most concerned about is going to be human failures.
So I’ll start out right now telling you we don’t want to include the effect of humans in this calculation. Period. End of discussion. Believe me now. Hear me later. Or no. Hear me now. Believe me later. We don’t want to do that. And I will explain to you a little bit later on in the podcast why if you include human failures in your calculations, not only are you wasting time. It’s worse than wasting time. Considering human failures in your SIL verification calculations is actually going to convince you that you should be doing something that is wrong.
That will not improve your PFD, but actually even make it worse. All right. So clause 11.9 says we need to quantify random failure of our safety instrumented functions.
Clause 11.9.1 Calculated Failure Measure Requirement
That’s going to be made apparent by clause 11.9.1, which is two sentences. It says the calculated failure measure of each SIF shall be equal to or better than the target failure measure related to the SIL as specified in the SRS. This shall be determined by calculation. So two brief sentences, three lines on the page, but very loaded with a lot of meaning and a lot of requirement in terms of what you’re doing with regards to these calculations.
So let’s start out. It doesn’t say that the SIL shall be equal to or better than. It says the calculated failure measure. This is the standards maker’s way of reinforcing to you that you have a target that you need to achieve. SIL may be everything, but it may only be part of what you’re looking at. So that target measure might be a risk reduction factor in the low demand mode of operation. It might be a frequency of failure or probability of failure per hour for high demand mode or continuous mode safety instrumented functions.
So there’s a kind of a number that you’re looking at that you’re geared toward when you’re doing all of this work.
Furthermore, it doesn’t specifically say SIL because sometimes your calculated failure measure, especially if you’re doing an explicit version of LOPA, is going to have more precision than just the order of magnitude. You might have done a LOPA that yields a risk reduction factor of 250. So 250 sits into the range between SIL 2 and SIL 3. But if all you say is SIL 2, then a risk reduction factor of 101 achieves SIL 2. So is that good enough? And the answer is no. You need to get all the way to 250.
So I mentioned this earlier when I was in Clause 9 that you really want to focus on the specific measure, whether it’s a PFD, it’s a risk reduction factor, or it’s a maximum failure frequency if you’re in continuous or high demand mode of operation.
If you’re calculating out the requirement explicitly, you should list it explicitly for the implementers of the system to actually achieve. So SIL 2 with a minimum risk reduction factor of 375 is a completely valid target. And in that case, you’re going to need to meet a risk reduction factor target of 375. Similarly, if you’re in that continuous mode or high demand mode, you might have a failure frequency of 3.5 times 10 to the minus 7 is the maximum frequency of failure that should be allowable for that type of system.
Okay, so we’re looking for a calculated failure measure, which is more than just what is the SIL.
If you’re using more qualitative approaches, implicit LOPAs, matrix-based methods, risk-graph-based methods, then your calculated failure measure is going to be the bottom of your safety integrity level target. But that’s the calculated failure measure that you need to achieve. Now, it furthermore says that this target failure measure is going to be related to the SIL of a SIF. So it’s going to be related to the SIL of the SIF. It might be the bottom of the range. It might be somewhere in the middle of the range. The Clause 1191 also says that it’s the calculated failure measure of each SIF.
So we’re not doing this for the safety instrumented system as a whole. We’re not doing it for a piece of equipment. We’re doing it for a SIF. A specific SIF needs to achieve a specific failure measure. And it says related to the SIL as specified in the SRS.
So this target failure measure, you know, as we talked about back in Clause 10, we need to specify what these calculated failure measures, what these performance targets are on a SIF-by-SIF basis. And that is going to be listed out in the SRS. Or as I preached back when we were talking about the SRS, it’s actually going to show up in a SIL selection, oftentimes done with layer of protection analysis. And in the SRS, you will refer back to where the SIL was assigned.
Now, if you are using a very good relational database type system to document your safety requirement specifications, you can report that one piece of knowledge out in many different locations.
Okay.
Continuing on, the second sentence says, this shall be determined by calculation. So for each SIF, I mean, there’s no ambiguity here. Each SIF needs a calculation. That means you can’t just look at the SIL targets on the boxes of the equipment that you bought and say, oh, it says SIL 2, it must be good enough. No, you have to run a calculation. And you can’t run a typical application that you hope to apply to everything. You need to have that calculation done for every safety instrumented function.
Now, sometimes you might want to shortcut and say, okay, low pass flow on my fired heater is going to be exactly the same equipment, tested exactly the same way. And I’ve got basically four or five or six passes on my heater. Each of those calculations is identical.
So that’s an area where you might want to, you know, document that, you know, all these equations are going to be identical to each other and do the equation once. But what is frowned upon is something like, okay, well, in my fleet of hundreds of fired heaters, they all have low fuel gas pressure shutdowns. And here is a typical calculation for a low fuel gas pressure shutdown for a fired heater. Now, the issue there is the test intervals might not be the same. The equipment probably is not going to be the same.
So the actual achieved SIL, the achieved PFD, the achieved risk reduction factor are not always going to be the same. So we’re going to want to have separate calculations for all of them. All right. That’s clause 1191.
Complex Applications and Fault Tree Analysis
And clause 1191 also has a note to it, which is one very long sentence with a parenthetical in it. So let’s go ahead and read the note to clause 11.9.1, which is, in complex applications, the hazardous event frequency can be used as an alternative to the target failure measures. For example, where different demand causes have different safety integrity requirements or where non-independent SISs act in sequence. Okay. This note is basically defining what the Kenexis-focused QRA service is.
And that’s where your safety instrumented function is part of a complex system where you might be sharing equipment between initiating events and protection layers and the safety instrumented systems.
And you really can’t excise and separate that safety instrumented function from everything else that’s going on in the scenario. So instead of doing a LOPA that assumes your SIS is independent, and as a matter of fact, it assumes also that everything in that LOPA scenario is independent, we’re going to use a tool that has more flexibility, it’s more elegant, it allows you to model the situation more rigorously.
Specifically, I’m talking about fault trees is what we’re normally going to do in this situation, using a really good fault tree analysis tool like Arbor, which is part of the Kenexis Integrated Safety Suite, and very tightly integrated into the Vertigo software application.
So in those situations, you’re going to have a target maximum event likelihood for the scenario as a whole that’s based on the magnitude of the consequence. And what this note is saying is that, well, a sensor might be used as an initiating event and as the sensor for the safety instrumented function. And because they’re not separate from each other, you need to use a fault tree to model the overall situation. There we’re not trying specifically to achieve a SIL for the safety function.
We simply need to know that given all of the equipment in that scenario, that we were able to achieve the frequency, the hazardous event frequency, that is less than the target maximum event likelihood of the scenario. And in that case, the PFD of the safety instrumented function is kind of buried inside the fault tree and doesn’t really have a whole lot of meaning and significance by itself because everything in the scenario is so intertwined.
So that note on 11.9.1 basically says, I’ve got a very complex situation with a lot of interdependencies between functions and initiating events, and I’m going to use a more sophisticated tool that understands the commonality. And all I really care about is that the estimated frequency of the consequence is less than the maximum allowable or TMEL, target maximum event likelihood for that scenario. So that type of analysis and assessment is obviously, it’s right here in the standard as an informative note that is something that you are allowed to do.
Clause 11.9.2 Random Hardware Failure Scope
Okay, so continuing on, we’re going to move into clause 11.9.2, and that’s probably all the further that we’re going to make it in today’s podcast because things do get very complicated and that there are a lot of bullet points. So clause 11.9.2 is only one sentence, but then you have A, B, C, D, E, F, G, H, I, J, K, L for bullet points under that one sentence, and then you have a fairly long note that we’re going to need to work our way through. So let’s get into it.
Clause 11.9.2 right out of the gate states that the calculated failure measure of each SIF shall be equal to or better than the target failure measure related to the SIL as specified in the SRS. This shall be determined by calculation. Huh, wait a minute. I just read you 11.9.1 again. Okay, I was going to say this text actually appears back in 11.9.1. No, I just read 11.9.1 again.
Let’s try this again, and this time I’m going to read you 11.9.2, which says the calculated failure measure of each SIF due to random failures shall take into account all contributing factors including the following, and then after that you’re going to get a bullet point of all of the items.
So both Clause 11.9.1 and 11.9.2 start with the phrase the calculated failure measure of each SIF. But now we’re specifically saying the next subclause here or portion of this clause is the calculated failure measure due to random failures.
So we’re only looking at random failures. We’re not looking at systematic failures, and we’re not looking at human failures in the design, testing, and operation of the safety instrument and function. Now, I’m going to hedge on this a little bit when we get to bullet point, is it I? It’s K. Bullet point K talks about operator response. That’s a little bit of an oddball that I’m going to kind of push back on.
And actually, we in the IEC 61511 standards committee are going to make this subclause go away in the next version of the standard because it just leads to confusion, and nobody does it the way that it was written, which I will get to.
So 11.9.2, we’re going to calculate a failure measure. We’re only going to look at random failures. Specifically, we’re going to look at random hardware failures. And then it says that that calculation of random hardware failures shall take into account these factors, and then it gives you a list of what you need to take into account. Now, you’ll notice it doesn’t tell you how to take them into account. It’s not giving you equations. It’s just telling you what you need to consider. Okay. So now we’re going to go through sub-items A through K. I’m sorry, A through L, one at a time.
Sub-item A: SIS Architecture and Voting
So item A says the calculated failure measure shall take into account the architecture of the SIS and its SIS subsystems where relevant as they relate to each SIF under consideration.
So what do we mean by architecture? What we mean by architecture is the voting arrangement, the redundancy. So if you have a simplex one out of one system, or you have a one out of two voting, or two out of two voting, two out of three voting, 17 out of 45, whatever it is, you need to consider that architecture when you’re running the calculations.
Now, you’ll notice that it says architecture of the SIS and its subsystems. And that’s because each subsystem, and as a matter of fact, individual components inside a subsystem may not have the same voting arrangements. It’s extremely common to have a two out of three transmitter system connected to a high diagnostics one out of one logic solver that is then connected to a one out of two voting valve system. Now, that valve system, each one of those valves in the one out of two might have two out of two solenoids associated with it.
So you need to consider the architecture not at the SIS level, but at the subsystem level.
So sensor subsystems, logic solver subsystems, and final element subsystems individually. And if you have multiple subsystems, so let’s say I measure temperature and pressure, those two measurements are inputs to my safety instrumented function. I might measure my pressure with a two out of three and my temperature with a one out of two. So there are different things. And all of that needs to be considered when you’re doing your SIL verification calculations. They can be different, which adds a lot of complexity to your calculations.
And this is something that your better SIL verification software, like Vertigo from Kenexis, is going to be able to allow you to group your inputs together, group your outputs together with different voting arrangements between groups, and then also allow you to build up legs of a sensor.
So that would be things like a sensor by itself is not, I don’t want to say it’s not relevant, but it’s not everything. You will need to consider your process connection. You will need to consider the sensor itself. You’ll need to consider interface devices, like input cards to PLCs or intrinsic safety barriers. All those things, anything in the loop that can contribute to a dangerous failure need to be considered, or to a failure in general. I’ll just kind of push back on that. All right.
So we need to consider the architecture, and that architecture, at a minimum, drops down to the subsystem level, and it might even drop further to components of a subsystem. All right. That’s A.
Sub-items B-D: Failure Categories
Next item is B. So the calculated failure measure shall take into account B, the estimated failure rate related to each failure mode due to random hardware failures.
So I want a failure rate, and the first failure rate I’m concerned about are failures due to random hardware failures, and that random hardware failures is going to show up in clause B, C, and D. So give me a failure rate due to random hardware failures, which would contribute to a dangerous failure rate of the SIS, but which are detected by diagnostic tests.
So I’ve got a full sentence to explain a topic that you will probably know better as Lambda DD. So item B is, if I wanted to do a shortcut, I would say subclause B says you need to consider Lambda DD failures. So Lambda DD failures are failures that are dangerous, that will inhibit the safety function from being able to perform its action, but that subset of the failures will be detected by automatic diagnostics. So you need to handle dangerous detected failures, DD failures.
That’s going to be clause B. And you’ll note, it doesn’t tell you that you need to consider that as an unavailability or an unreliability. It doesn’t tell you how to actually run the calculation. It just says you need to think about dangerous detected failures.
Clause C is very similar. Let me read it all the way through and then give you the simple version. So reading it all the way through, subclause C says you need to consider the estimated failure rate related to each failure mode due to random hardware failures, which would contribute to a dangerous failure of the SIS, which are undetected by the diagnostic tests, but which are detected by proof tests.
That’s a long, complicated, wordy sentence that says you need to consider DU failures, dangerous undetected failures. So subclause C is you need to consider DU failures. So DU failures are dangerous. And furthermore, the automatic diagnostics are not going to tell you that they’re there. So DU is kind of the subset after the first pass where we determine what fraction of failures are detected. Then the next, the balance are going to be undetected.
But that DU failure rate should only include the failures that are not detected by automatic diagnostics, but are going to be detected by the manual proof tests. So our DU, not detected by diagnostics, but will be detected by the automatic proof tests.
And we want that DU, if we want DU and DD to be as close to 100% as possible. And we are going to write our test procedures to make sure that they are as close to 100% as possible. But if not, we have a new category that was inserted in the 2016 version of the standard.
So subclause D. So we need to take into account D, which is the estimated failure rate related to each failure mode, due to random hardware failure, which would contribute to a dangerous failure of the SIS, which is undetected by the diagnostics and undetected by proof tests. So clause D, if I could simplify it, is your DN failures, lambda DN. What is the rate of dangerous failures that are never detected? Why are they never detected? Well, they’re never detected because the diagnostics, the automatic high-frequency diagnostics can’t detect that that failure mode is present.
And the manual proof tests can’t detect that that failure mode is present either.
So Lambda DN, when it was first introduced, got a lot of pushback in industry. It’s like, how can I possibly perform a function test of my SIS and see that it works, but have it not work?
So the example I’ll give you of a DN, dangerous never detected failure, is a, I’m going to give you an example that is a float level switch. So in your mind, close your eyes, not if you’re driving, but imagine a scenario where I have a tank and I have a float that is actually inside the tank. Now, how do I test a float that is inside the tank? Well, I could fill the tank up beyond the trip point of the level switch. Well, that is an operational nightmare, and it may be dangerous.
So you’re going to basically fill that tank up while the plant is shut down for a turnaround to make that float switch move. Most people are not going to do that. So instead, they’re going to probably open up a man way, and they’re going to use their calibrated 2×4 to actually lift that float switch up beyond the trip point and make sure that the switch changes position. Then they’re going to remove their calibrated 2×4 and let it drop back down to its end of travel and go back into the unsafe state, if you will, or the state where it’s not floated. Okay, so what is the problem with that test?
Is that test 100% effective? No, it’s not. There’s a failure mode for a float that is the float has a crack or a hole in it. If a float has a hole in it, it ain’t going to float because it’s going to fill full of your process fluid and stay down at the bottom of its range. Did you test when you picked that float up with the 2×4, did you test to see whether there was a hole in the float? No, you absolutely did not. As a matter of fact, that test might have actually put a hole in the float, which we’ll come back to that in when we’re talking about clause E, subclause E here in a second.
So there are some pieces of equipment that are very difficult to test while the plant is offline and running. Now, you could have installed that float type transmitter in a separate bridle that you could actually fill it and empty it during this testing process to get a much more comprehensive, better test, which is what we would recommend. So only if you can’t physically test for a failure mode would you say, okay, well, that’s going to be a DN failure.
Because if you have DN failures when you start running these calculations, you know that they accumulate over the entire life of the device, which makes a very small failure rate result in a very large PFD contribution if it accumulates for 10, 20, 30, 40 years. Okay, so that is subclause D. Subclause C says that you need to take into account the susceptibility of the SIS to failures caused by the proof test themselves.
Sub-items E-G: CCF and Diagnostics
So my proof test actually poked a hole in my device and now it doesn’t work. Now, you’re probably asking yourself, how do I model this? What number, is there some sort of beta factor that I use that is a function of the type of test that I do that says, well, when I perform the test, I’m actually going to cause the device to fail by the test.
Well, I don’t know how this subclause got into the standard originally, but in the next version of the 1511 standard, this clause is going away because it’s kind of impossible to put a number on the frequency at which a test would cause the device to fail. And it’s kind of counterproductive to analyze this.
And the causing a failure with a proof test is a systematic failure. It’s not a random hardware failure. So we don’t want to look at it anyway. So for a wide variety of reasons, I’ve explained to you what clause E is. And even though that subclause is in there, nobody, nobody, nobody really ever did anything with it. And in the next version of the standard, it’s going away. So I give you absolution to go ahead and start ignoring it today.
All right, subclause F says that you need to consider the susceptibility of the SIS to common cause failures. That common cause failure, again, being a single stressor, which causes all of the multiple components in a redundant voting arrangement to fail at the same time for the same reason.
That susceptibility is usually modeled with a separate term called the common cause term, an unreliability. And we’re typically going to use a beta factor, which is going to be a fraction of the overall failure rate. So a beta factor of 10% would basically be saying when a single device fails, 10% of the time, all of the other redundance components failed at the same time for the same reason due to the same stress. And then we kind of throw that into our calculation.
So common cause failures are nature’s limit to the effectiveness of redundancy. So if I have one device, that’s good. If I have two devices, that’s better. If I have three devices in a one out of three vote, it’s really not significantly better than one out of two. A fourth device is going to buy you nothing if your common cause failure factor is anything significant. So you need to consider common cause. That’s what clause 11.9 says. It doesn’t tell you how to do it. It doesn’t tell you what numbers you should use, but it just says that you do need to consider it.
And the beta factor, once you start getting into the informative notes, the informative annexes, informative technical reports like ISA 84.00.02, we’ll go into more detail on that.
All right, subclause G states that you need to consider the diagnostic coverage of any periodic diagnostic tests, the associated diagnostic test interval, and the probability of failure of the diagnostic facilities.
Okay, so most of your instruments are going to be performing high-frequency diagnostic tests. So an example of a high-frequency diagnostic test for a logic solver would be something like a watchdog timer, where you’re trying to determine whether or not the PLC has stopped operating. And you know this because the pulses to reset the watchdog timer stopped because the logic solver hung up. And once the pulse stops and the timer times out, you know that the logic solver has stopped. So you de-energize all of your outputs to bring your plant to a safe state. That is a diagnostic.
It’s actually a very good diagnostic to detect the failure mode of my logic solver froze up. It froze in place. Blue screen of death from Microsoft.
Okay, so you need to consider, so for those periodic proof tests, you need to consider, number one, what is the diagnostic coverage? So that means what percentage of the failure rate, what percentage of the failures, is going to be detected by that type of diagnostic. So kind of going back to that watchdog timer scenario, you’re going to detect that the logic solver hung up. But if a bit of memory in the logic solver gets stuck, that is not going to be detected by this particular diagnostic. So you need to look at all the failure modes that are possible.
You need to look at what tests you’re running. And for each failure mode, determine whether or not that failure mode will be detected by a diagnostic test to determine the overall diagnostic coverage.
Okay, so once you know the coverage, that tells you something. The other thing that you need to consider is the associated diagnostic test interval. Excuse me. Now, the diagnostic test interval is going to be important because that factors into, if a failure occurs, how quickly am I going to be able to repair out of it? And not only how quickly can I repair out of it, that duration between when you get a failure and when you repair out of it, that tells you what your unavailability is that’s going to contribute to that DD term. So how quickly you do diagnostic tests is a factor.
Now, most of the time it’s irrelevant. So the mean time to restore is going to include how long does it take between when a failure happens and when you know about it. And then, once you know about it, how long does it take to repair? So most of you assume that that repair time, mean repair time, the MRT, is going to be 72 hours. Now, if you’re running a diagnostic every 100 milliseconds and it takes 72 hours to repair something, you need to add those numbers together.
Or technically, you need to add half of the diagnostic test interval in because the failure might have occurred at the beginning of the time period or at the end of the time period. So on average, it happens in the middle. So we divide it by two.
But if your diagnostic test interval is 100 milliseconds and your mean repair time is 72 hours, when you add 100 milliseconds on the 72 hours, you still have 72 hours. So we ignore it for everything other than very long interval diagnostics, something like a partial stroke test that might not occur but once a week or once a month. Now that test interval becomes very significant in the calculations.
And then, the last item that honestly not a lot of people factor into their calculations is the failure of the diagnostic facilities. So if my device that is performing my diagnostic test fails, then the diagnostics are not going to occur. And all of my DD failures just became DU failures because my test’s not happening. So sometimes if you look at some certificates, this might show up like an enunciation failure rate. So the rate at which you’re not going to know that your diagnostics have diagnosed a failure in occurring.
This number is generally on the extremely low side. A lot of people are going to ignore them in their calculations. But it’s something that you can consider when you’re running your calculations. And basically, if the diagnostics fail, now all of your DDs have become DUs for the balance of time before you run your manual proof test to determine that the device has failed and repair out of that failure mode.
Sub-items H, I, L: Proof Tests and Utilities
Okay, item H is the next item. So now we need to take into account H, the coverage of any periodic proof tests, the associated proof test procedure, and the reliability for the proof test facilities and the procedure. So everything that I just said for diagnostics, you also need to do for that manual proof test.
Now, back in the day when we first started doing SIL verification calculations, we assumed that if we tested it and it worked, it worked. The DN failure rate was zero. There are no failures. But then we kind of over time said, yeah, maybe the test that we’re doing is not that good. And it’s not going to give me 100% proof test coverage because there are failure modes that I don’t detect.
So if you’re running a test by connecting a HART communicator to the transmitter and ramping its output up and down, you might be testing the electronics, but you’re not actually testing the measurement cell or the connection of the process. And if you look at some certification reports for that type of test, you might see a diagnostic coverage that’s in the 60, 50, 60, 70% range. It’s pretty low, actually. And the balance of the failures are not going to be detected by the test. And we need to consider that when we’re determining how effective our system is or, in this case, really is not.
Now, so we need to estimate what that coverage of our manual proof test is. And honestly, we want to write our test procedures so that it is effectively 100%. Any failure mode that I know about, I want to test and make sure it’s not there. And that’s going to be a function of the proof test procedure.
And that’s why the clause says the associated proof test procedure needs to be part of this assessment, as do the proof test facilities. What equipment am I using to perform this proof test? So think about how you’re doing the test, what procedure you’re using to perform the test, when you’re making an assessment of how good your periodic proof tests are. And whereas some of the more academic types might want to say, well, what we’re really concerned about is making sure that, you know, we’ve accurately calculated what that periodic proof test coverage is.
I would argue, no, no, no, that’s not what we’re really worried about. If my periodic proof test coverage is not 100%, I need to be looking at why. And I need to change my proof test procedures and my proof test facilities so that I can get to 100% or as close to 100% as possible. So more of a tool for making sure that your test procedure is well written, as opposed to some administrative certification paper pushing. All right.
Next subclause is I. So I, when I’m running my calculations, I need to take into account I, the repair times for detected failures and the state of the SIS during repairs online or offline. Okay.
So if I test a device and it has failed, it is unavailable to perform its safety action. So when that, when I know about that failure occurring, I will kind of start the clock on my mean repair time.
And hopefully I can report repair it in significantly less than my mean repair time. That is almost always 72 hours. So most of the time, your mean repair time is actually more of an MPRT. I’ve already discussed this, a maximum permitted repair time, as opposed to the actual amount of time that you think it’s going to take. But for that time duration of repair, you are unavailable. And that is going to go into that PFD calculation, specifically with dangerous detected failures.
Or if you have a voting arrangement, like two out of three, you’re going to be able to detect the failure in one out of those three transmitters by a simple comparison vote, which at the end of the day is actually a kind of a diagnostic. Okay. So that time duration that it takes to repair is going to factor in.
It’s going to be a contributor to the PFD. Now you’ll notice there’s a little parenthetical in this statement. Or, well, the back end of the sentence says that you need to consider the state of the SIS during repairs. And then you have a parenthetical online or offline. So if my plant is online and running when I detect the failure, and I continue to run the plant in the presence of that failure, that component is unavailable to perform its safety action, and it will contribute to the PFD.
But if I did my test while the plant is offline, or when I detect the failure, I shut my plant down, now it’s not going to contribute to the PFD because you can only contribute to the PFD when the plant is online and running.
So when you detect a failure when the plant is offline, it’s incumbent upon you to repair the device before you put the plant back online, and then it’s not going to contribute to the PFD of the safety instrumented function.
Okay, now we’re at clause K, but I’m going to come back to clause K. It’s the last thing I’m going to want to talk about because it’s a little bit of a complex topic. So before I get to K, let me finish the last item, which is L, which says we need to take into account L, the reliability of any utility necessary for the SIS. This subclause is going to go away in the next version of the standard because it’s simply a subset of clauses B, C, and D.
So if I have an energized to trip system and I lose my power to the energized to trip system, that is a dangerous failure of the SIS that I need to include in my calculations.
So it was kind of an unnecessary additional statement, which we in the standards committee are going to move in the next version of the standard back into the informative notes because it’s just a subset of the failure rates is all we’re talking about here, and that is the subset of the failure rates that are associated with utilities, which would include things like electricity, instrument air, pneumatics, communication utilities that communicate signals between one component and another component.
So all of those things, if they can cause a dangerous failure, need to get included into the dangerous failure rate that you’re going to put in your equations when you’re calculating the PFD. Okay, so yeah, whether it’s instrument air, whether it’s steam, whether it’s electricity, if that failure can be dangerous because you’re energized to trip, now we need to include that failure rate into the calculations. Okay, last item. Actually, before I leave that item L there, a lot of people say, well, how do you include that in your calculations?
Well, if you’re using Vertigo, you can include the utilities of energized to trip systems as interface devices in your subsystems. So if you’re using an energized to trip switch out in the field, you could put the frequency of failure and the duration of failure into the calculation and just put the electrical supply as an interface device when you’re running your calculations. So that’s where and how you would handle it.
Sub-item K: Operator Response in SIL Calculations
Okay, so now we’re going to go back to the item that I skipped, which is item K. So let me go ahead and read item K. And it says that in your calculations, you need to consider K, What the heck does this mean?
And you go up and you look. Everything in your calculations is based on failure rates. DD failures, dangerous detected. DU failures, dangerous undetected. DN failures, dangerous never detected. And for all those failure rates, the subclause screams due to random hardware failure. So where the hell did this estimated likelihood that operator response would cause a dangerous failure of the SIS come into this? Well, let’s unpack this.
Right out of the gate, it says operator response. Okay, so operator response is a trigger. And so basically, what is not operator response?
So if an engineer performs the calculation incorrectly and undersizes the shutoff valve and it doesn’t allow it to go to a safe state, that’s not an operator response. If a technician miscalibrates a safety instrumented system, that’s not an operator response. So what is an operator response?
Well, that is the operator taking an action in response to some sort of trigger like an alarm. So, well, why the heck would we ever want to include that in a SIL verification calculation? And I’m going to go up on my soapbox here and tell you, you know what? You probably don’t want to go and include those in your SIL verification calculations.
But if you go back to some of the discussions we had a long time ago about the definitions of the safety instrumented system, you might be reminded that SIL verification calculations can include a this as a SIF.
A alarm, a safety critical alarm comes in, which is received by an operator, and then an operator pushes a button, which closes a valve. Now, I would never consider that to be a safety instrumented function. I would never assign that a risk reduction factor greater than 10. But some people want to treat that as a safety instrumented function using the same hardware, using the same software. And what just happened? In that situation, my human operator has become the logic solver.
And if that is the case, well, then I need to include that human operator as the logic solver in my SIL verification calculations. And there are some operating companies that do exactly that. Now, Kenexis does not recommend that you do that. We don’t recommend that you would ever give an operator enough credit to get into the SIL 1 territory in the first place. So we would have stopped you at risk reduction factor of 10, which obviates you from the need to follow this standard in the first place.
But if you do, if you’re in some sort of organization that says a safety critical alarm where an operator pushes a button is a SIL rated safety function, then Clause K kicks in. Now, that being said, Clause K is going to disappear in the next version of the IEC 61511 standard because if an operator is part of the loop, you shouldn’t have been able to get in the SIL 1 in the first place.
And also, operator response is, again, going back to Clause B, C, and D, it is a subset of failure rate. So we’re basically kind of calling out a specific type of failure when we really needn’t be. So that’s going to go away in the next version of the standard. You shouldn’t be doing it because you shouldn’t be allowing the operator to be that good in your SIL verification calculations. But hey, there’s a possibility that you might have done that. And some people will include an operator’s logic solver in their safety instrumented functions, at least for the time being.
Episode Wrap-Up and Vertigo Advertisement
All right. So with that, that has given us all of the clauses, sub-clauses, in Clause 11.9.2. And that is where I am going to bag it for the day, even though I haven’t talked about the informative note yet. I will talk about the informative note in Clause 11.9.2. When I talk about the balance of Clause 11.9 that talks about reliability data, it talks about uncertainties, and it talks about ways to achieve your target failure measure.
The note is, kind of could, you know, stand on its own in one of those because it talks about the approaches that you might want to use for running your SIL verification calculations. But I will get into all of that next week.
Now that you’ve heard some insights on technical safety, functional safety, and the IEC 61511 standard, let me tell you a little bit more about how to easily and effectively implement the safety lifecycle using the Kenexis Integrated Safety Suite and our SIS safety lifecycle management tool, Vertigo. Vertigo is a comprehensive tool set for performing assessment calculations, documenting, and maintaining the design of safety instrumented systems.
Analysis begins with importing or synchronizing a list of safety instrumented functions with their definitions and associated performance targets from our open PHA tool for HAZOP and LOPA documentation.
Each safety function can then be analyzed by performing a SIL verification calculation, complete with a collection of tools for optimizing designs and a database of thousands of potential instruments to define failure rates and diagnostic coverage capabilities. After the SIL verification calculations are defined, you can build an SRS by automatically generating a cause and effect diagram from the SIF definitions and other defined instruments.
Each SIS instrument will include a customizable data sheet and general requirements that are applicable to the SIS as a whole and can be entered individually or even bulk imported from customizable libraries.
After the design phase, you can even use Vertigo to track and document testing throughout the entire life of the facility. Kenexis Vertigo is the most integrated, easy-to-use enterprise tool for allowing the development of SIS design basis information more efficiently and effectively than any other software application. Thank you.