Kenexis Functional Safety Podcast
The standards committee doesn’t trust your calculations — and with good reason. In this episode, Ed Marszal unpacks Clause 11.4’s minimum hardware fault tolerance requirements, the committee’s skepticism toward vendor failure rate data, and why redundancy against dangerous failures became the non-negotiable floor beneath every SIL claim. He works through Table 6’s SIL-by-SIL requirements, the rare cases where IEC 61508’s Route 1H or Route 2H become necessary, and the baffling exception clauses that let engineers ignore the rules entirely — clauses he admits he cannot find a single real-world example for. The episode closes with two stealth requirements hidden at the end of 11.4: the 60% diagnostic coverage mandate for programmable devices and the 70% statistical confidence limit on failure rates, both of which belong elsewhere in the standard but carry enormous practical weight. For anyone who has ever wondered why SIL 3 demands redundancy regardless of what the math says, this is the explanation from inside the committee room.
You achieve a given SIL target by running a calculation that verifies you achieved the proper probability of failure on demand, yet the standards committee doesn’t trust to do that properly…. Join Ed as he breaks down clause 11.4 covering hardware fault tolerance and provides his thoughts and expert analysis of this topic in more detail.
Tune in to the latest episode of the Kenexis Functional Safety Podcast, hosted by Ed Marszal, President and CEO of Kenexis. Now available on Spotify and Apple Podcasts, Ed offers his expert insights on the IEC 61511 standard.
With decades of experience in safety instrumented systems and as a Principal Engineer, Ed has a unique perspective to offer. He has been an active contributor to the ISA 84 committee since 1994, adding to his deep understanding of the field.
In this inaugural season, Ed delves into the IEC 61511 standard, unpacking the meaning behind each word and providing a thorough interpretation of its application. Through personal stories from his career and committee work, he offers valuable context and insights for professionals in the industry.
Full Episode Transcript
KENEXIS FUNCTIONAL SAFETY PODCAST — S1E37 TRANSCRIPT (Markdown)
Cleaned & reflowed for web publication and AI crawlability.
The JSON-LD block below is schema.org structured data. If your CMS lets you
add raw HTML to a post, paste it into the page
body — crawlers read it either way). Fill in PLACEHOLDER_EPISODE_PAGE_URL
once the post exists. Everything from the "# Kenexis Functional Safety
Podcast…" heading down is the transcript body — paste it into your post.
–>
"`html
"`
# Kenexis Functional Safety Podcast — Season 1, Episode 37: IEC 61511, Clause 11.4 (Hardware Fault Tolerance)
—
## Introduction and Episode Overview
You achieve a given SIL target by running a calculation that verifies that you've achieved the proper probability of failure on demand. Yet, the Standards Committee doesn't trust you to be able to do that properly.
Welcome to the Kenexis Functional Safety Podcast. I'm your host, Ed Marszal, President and CEO of Kenexis. Kenexis is a technical safety consultancy that helps chemical process industry companies to analyze risk and design engineered safeguards like safety instrumented systems and fire and gas detection systems. Kenexis also provides the industry-leading suite of software tools, including our best-in-class Vertigo software for SIS safety lifecycle management.
In this first season of the podcast, we are going to focus on the IEC 61511 standard, doing a deep dive into the standard, including more depth of information on what the standard means and how to apply it, brought to life with personal war stories and behind-the-scenes discussions of the committee members as we develop the standard in ISA 84 and IEC SC 65.
Before we start, a little disclaimer. I will be providing my opinion on technical and engineering topics. This information is provided on a best-effort basis and is of a general nature. The information presented in this podcast might not be applicable to your specific application. It is the obligation of every engineer to thoroughly analyze any system that they are designing and not blindly rely on any general advice presented in this podcast. All right.
## Why Minimum Hardware Fault Tolerance Exists
So what we're talking about today, today we are working our way through Section 11. Again, my favorite section. And we're talking today about hardware fault tolerance. So hardware fault tolerance is redundancy. But there are two different kinds of hardware fault tolerance. You can be tolerant to a dangerous failure or you can be tolerant to a spurious trip. This particular clause doesn't specifically say so. But what we're really interested in is the tolerance to a dangerous failure.
So sometimes I'll refer to it as HFT sub D, meaning the ability to tolerate dangerous failures and still be able to execute the action for a subsystem.
Now, why do we have minimum hardware fault tolerance requirements? Let's get straight to the facts. The standards committee doesn't believe that you can do a good calculation or doesn't believe that everyone will be able to do a good calculation. And it also believes that there's a lot of really crappy failure rate data out there. There are a lot of equipment vendors who are kind of playing fast and loose with the books, might be overstating the quality of the equipment that they're providing, using questionable techniques to accumulate failure rate data.
So, for instance, failure rate data based on field returns is junk. So a vendor says, OK, my failure rate is one times 10 to the minus nine per year, which immediately should throw up red flags as kind of ridiculous. And how did they come to this conclusion? They say, well, we manufactured a million of these devices last year. And, you know, in the following year or the following two years or whatever, we only received this tiny number of devices back from the users as failed. Well, this is guaranteed to underreport the number of failures that occur.
Number one, there's channel stuffing. So you push a bunch of equipment out to your distributors and it's sitting on their shelves. And that's counted as devices that haven't failed. They're not in service, though, so they shouldn't count. But, you know, the equipment vendor has no way to really know that that's the case. So it's not a failure. So things that are sitting on the shelf are counted as working. Then you get potential failures. And some of the failures are going to get returned.
Now, those failures that get returned, the equipment vendor can look at the device, look at the stresses on the device and make the judgment that, well, this device was used out of its tolerances. And I'm not going to count it as a failure. And I'm not going to count it as a failure because the end user applied it in a way that they're not supposed to. Somewhat reasonable, but also totally open to abuse.
Beyond that, if you have a $100 device, are you going to go through, are you going to spend several thousand dollars of man hours trying to get a replacement device? Or are you just going to grab another one off the shelf, install it and throw the one that you have in the trash? Okay, so there are a lot of pieces of equipment that aren't worth the trouble of letting the vendor know that it's failed. And then finally, if you're, what is the warranty period? A year? Two years? Maybe? And if the device is out of warranty, you're not going to get anything for reporting that it's failed.
So that's just a simple explanation of why a lot of failure rate data that's out there is really bad.
So the Standards Committee wanted to have a backstop that said if a person uses bad data that was provided to them by an equipment vendor or they didn't collect the data properly in the field, whatever reason, they're not using good failure rate data. The Standards Committee said, you know what? We're going to apply a hardware fault tolerance that forces people to put in redundancy for safety, not against spurious trips, but for safety, for under certain circumstances. And that's basically higher SIL targets are going to require more redundancy.
And that really matches up with if you use reasonable data, what your calculations will show you.
## Clause 11.4.1 – SIF-Level HFT Requirement
So 11.4 is titled hardware fault tolerance. The first clause in 11.4, 11.41 says, The SIS shall have a minimum hardware fault tolerance with respect to each SIF it implements. Okay. So the SIS shall have a minimum hardware fault tolerance with respect to each SIF.
So basically each SIF is going to have a hardware fault tolerance. And that hardware fault tolerance is going to change on a SIF by SIF basis. So it's something that you need to consider at the SIF level for what the overall target is. But what we're going to learn a few clauses down from here is that the actual implementation, how you achieve that hardware fault tolerance, is going to change based on the subsystem level. So we'll get to that right directly in the next clause.
But before we dig into the next clause, let's go ahead and talk about that note, which says, This does not exclude the possibility that the hardware fault tolerance may be reduced below the minimum requirement at certain times during operation of the system following the occurrence of faults.
So what this says is a clause 11.41 that says, Every SIF is going to have a hardware fault tolerance you need to achieve. It says that that hardware fault tolerance applies to the normal state of operation where all the equipment is working as designed. If a device fails, so let's say you have a one out of two vote, you detect a failure in one of those two devices and you degrade to a one out of one voting. That one out of one voting has zero hardware fault tolerance where the one out of two vote had one degree of hardware fault tolerance.
So you're no longer achieving the required degree of hardware fault tolerance during that point in time where you continue to operate with one of those devices in the failed state.
So that's okay. All of the requirements from last week, 11.3, are going to apply. Meaning that you need to have some sort of compensating measures in place to maintain the effectiveness of the safety instrumented function. Even though we have a failed component. And in this case, we're not achieving the hardware fault tolerance that is required for the safety function. But only for the duration of the bypass of that device that includes the use of the compensating measures. Okay.
Now, when we say that we need a minimum hardware fault tolerance, that is going to be achieved by using redundant components. And those redundant components need to be wired up, installed, designed in such a way that they are going to provide tolerance to dangerous failures. So everything in Clause 11.4 is to prevent dangerous failures accumulating as opposed to spurious trick. Standard, with regards to minimum hardware fault tolerance, doesn't care, doesn't give a hoot about spurious trips. That's not going into the equation. So how do you know what your minimum hardware fault tolerance is?
So if you kind of look at definitions, if I have a hardware fault tolerance of X, then that means X plus one failures will cause the system to not be able to respond. That's one way of thinking about it. The other way of thinking about it is if I have a hardware tolerance of X, then X failures can occur, but the system will still continue to operate. So one out of two voting has one degree of tolerance. Because if the A device fails, the B device can still shut you down. See, we got that with redundancy because there were two devices and one of them can fail.
Now, mathematically speaking, if you want to know what the hardware fault tolerance against dangerous failure is, if you have an M out of N voting arrangement, then the hardware fault tolerance to dangerous failures is going to be N minus M. So if you have a two out of three voting arrangement, M is two, N is three. So the hardware tolerance is three minus two or one. Yes, that's right. Even though a two out of three voting arrangement contains three devices, it only has one degree of tolerance to a dangerous failure.
It is a unique kind of arrangement because it also has one degree of tolerance to a safe failure, but tolerance to safe failures doesn't come into play when we're looking at clause 11.4. So we're trying to achieve tolerance to failures by adding redundant components, but putting them, wiring them up in a fashion, putting the logic in a fashion that it will prevent dangerous failures.
Now, if I have a two out of two voting arrangement, that means that both of the devices need to agree that the shutdown needs to happen before the shutdown will happen. So if I use my equation, the hardware fault tolerance is going to be N minus M or two minus two or zero. So a two out of two voting arrangement, even though it has two devices, has no hardware fault tolerance to a dangerous failure. It does have fault tolerance to a spurious trip, but not to a dangerous failure. All right.
So we're trying to set a floor for what redundancy we need in such a fashion that it provides tolerance to dangerous failures because the standards committee doesn't necessarily trust or believe that people are going to be able to run their calculations properly. Specifically, they're going to use underreported failure rate data. Okay. So that's clause 11.4.1.
## Subsystem-Level HFT and Route 1H
Let's move on to clause 11.4.2, which gets into a little bit more nitty gritty on achieving the hardware fault tolerance. So 11.4.2 says, when the Sys can be split into independent subsystems, for example, sensors, logic solvers, and final elements, then the hardware fault tolerance can be assigned at the subsystem level.
What that means is that the way that you achieve your hardware fault tolerance can change on a subsystem by subsystem level. So you might use two out of three voting sensors, which have one degree of hardware fault tolerance. You might achieve your SIL target with a one out of one D logic solver, which has no fault tolerance, but it makes up for that by having increased diagnostics. That's going to be route one H in IEC 61508, which we'll get into in the very next clause. And then maybe you can use a one out of two voting arrangement for your valve.
So your voting arrangement doesn't need to be consistent through your SIF as a whole. You might look at it at the subsystem level, and you might even look at it at the component or subsub system level, because you may have redundancy on your solenoid valves, but not on your process valve.
So you have infinite flexibility, and you need to just do your confirmation of adequate hardware fault tolerance on the subsystem level, not at the SIF as a whole. You don't need everything to be consistent. All right, the next clause explains the three, one, two, three ways that you can confirm that you have an appropriate degree of hardware fault tolerance.
Those three ways are, number one, you can comply with clauses 1145 through 1149 of clause 11 in the IEC 61511 standard. So you can use the tables and notes and requirements in 1511. And 99.9% of the time, that's going to be perfectly acceptable, especially in the 2016 version of the standard versus the 2003 version of the standard. It is very easy, and it's very reasonable to achieve appropriate fault tolerance structures with table six that we're going to get into, a couple clauses here from now that define how much fault tolerance you need for the different voting arrangements.
But, you also have the ability to jump out into the mother standard, 6508, also known as the equipment vendor standard, because that standard has a lot more detail, and it allows you to look at a lot more different factors, because table six in the IEC 61511 standard that you're going to use 99.9% of the time, it's basically for field devices. Yeah, that's it. And logic solver devices get a little bit more complicated.
So, whereas table six only looks at what is the SIL target, and a little bit about mode of operation, but primarily the SIL target tells you what the hardware fault tolerance is. When you get into 6508, you're going to have a lot more flexibility to look at not just the required SIL, but also the type of device that you have, the amount of diagnostics the device has. You can look a little bit more rigorously at your failure rate data, and apply less hardware fault tolerance with higher confidence of failure rate. So, we'll get into that in just a second.
So, most of the time, you're going to be using a logic solver where the equipment vendor told you what SIL it achieves, and you're going to be applying table six in clause 1145 to the field devices, and you're good to go. Literally, 99.9% of the time or greater.
But, occasionally, you're going to run into some oddball situations where you're going to need to jump over into the 1508 standard to look at things in a little bit more detail. So, bullet points two and three for the methods to determine whether or not your hardware fault tolerance is appropriate. Item number two says you could follow the requirements of clause 7442 in 6508. That is called the Route 1H. So, in Route 1H 7442, what you have is a series of tables that will tell you what SIL you're capable of achieving.
So, if you're looking at the 6508 standard part two on page 27 of the PDF that I am looking at that we received or that we purchased, there is a table that looks at what hardware fault tolerance do you have, what is the safe failure fraction of that device, and what type of device do you have. And it will look at those three parameters to tell you what SIL you're capable of achieving with that equipment arrangement, with that subsystem architecture.
So, the two things that are in 1508 Route 1H that are not in the IEC 61511 tables are, number one, what type of device is it? So, devices that have microprocessors in them are complicated and the failure modes are not very well understood. That requires you to add some more hardware fault tolerance. Those are type B devices. But if you have a type A device, your board-on-tube pressure switch, your solenoid valve, your process valve, it doesn't have any microprocessors in it. That's one big key. But it's a well-known, well-established device with well-known and well-understood failure modes.
So, I don't need to have as much fault tolerance because I understand how the device fails a lot better. You know, reducing the risk. And basically, whether it's a type A or a type B device is going to increase or decrease the amount of hardware fault tolerance that you need by one degree.
Now, the dirty little secret in table six, I don't want to call it a dirty little secret, is that it is a subset of the tables in Route 1H of the 6508 standard. Basically, we are assuming that we have a type B device with a safe failure fraction between 60 and 90%. That's going to give you the numbers that you're going to be looking at. And it's a little bit more flexible. No need to get into the details. All right.
## Route 2H and Rigorous Failure Rate Data
There's also another route that is going to be the third bullet point under 1143, which is that you can comply with the requirements of 7443 or Route 2H of IEC 61508. Now, Route 2H, making a long story short, requires you to understand and develop the failure rates with a lot more rigor. So, you need a lot better tracking of your failure data, your failures that occur during testing, your failures that occur during operation, and you need to calculate the failure rate to a much higher degree of rigor.
So, whereas failure rate data as per IEC 61511 is going to require 70% confidence, if you're going to claim Route 2H, you're going to need a much, much higher degree of confidence in your failure data, which adds conservatism in the number.
So, why would you use either of these two approaches in IEC 61508 instead of just following the tables in IEC 61511? Well, here's the answer. If you're trying to achieve SIL 3 with no hardware fault tolerance, you're going to need to go Route 2H.
And, the biggest consumer of this is going to be if you need to achieve SIL 3 in a safety function where you are de-energizing a high power, high voltage motor, this is a route that you're probably going to look at because the motor starter systems are extremely or can be extremely good. Internally, they provide all kinds of redundancy and, well, a lot of the time the equipment vendor doesn't get themselves certified to a certain SIL level. But, internally, they're using multiple coils and multiple latches to be able to latch and unlatch the motor starter.
And, it's exceedingly uncommon for a high voltage motor starter that receives a valid command to stop to not actually disconnect. Most of the failures that you'll see with high voltage motor starters are related to the signaling circuit, not the motor starter itself. So, that's kind
> *[inaudible, 25:22 – 25:25]*
those examples where you're actually trying SIL 3 with no hardware fault tolerance is going to require a very meticulous, accurate assessment of what your actual failure rate data are. And, motor starters are pieces of equipment that you generally, at a single plant, have hundreds of those devices where you can collect data and potentially make this case.
Alright, so that's one reason you would go Route 2H. Route 1H is basically I am using commercial off-the-shelf programmable hardware for my logic solver. So, I'm using a non-certified equipment vendor can't spell IEC 61508. and I'm using that as my logic solver for my safety functions. I'm doing my safety configuration and the things that you need to do to allow a general purpose PLC to be used in a safety application.
The table, table 6 in the IEC 61511 standard is not going to be applicable. So, you're going to need to go back to 1508 and look at the 1508 tables anytime you're looking at a programmable device in a safety application. And by that I mean a LVL programmable. You're not going to be able to use an FVL full variability language programming for other reasons. But LVL meaning ladder logic function block diagram type programming.
Okay, so that's 1143. Now we know 99.9% of the time we're going to follow what's in 1511. Occasionally if I'm trying to hit SIL 3 with a single device I might want to go route 2H. Good luck collecting the data. Good luck making that justification. And if I want to use a commercial off-the-shelf PLC I'm going to go route 1H for that PLC while I'm doing my assessment. Okay, so that was clause 1143.
## Clause 11.4.4 – Fault Exclusions
What are the appropriate degree of hardware fault tolerance is. Clause 1143 does have a note that says the route developed in IEC 61511 is derived from route 2H of 1508. And it's kind of a little bit of both. I would argue that it's more route 1H. But it's kind of a mishmash of 1H and 2H. Doesn't really matter. You get what you get. What's in 1511 is there and it's easy for all to see.
Okay. 11.4.4 states when determining the achieved hardware fault tolerance certain faults may be excluded provided that the likelihood of them occurring is very low in relation to the safety integrity requirements. Any such fault exclusions shall be justified and documented. Okay.
So we're going to get a whole bunch a series of get out of jail free clauses that I don't understand why they needed to put these clauses in. I don't understand who or why any of these clauses has ever been used. You know coming from somebody that sits on the standards committee that's saying a lot.
But what clause 11.4.4.4 said is when I am determining whether or not I achieved a certain hardware fault tolerance. You have the ability to exclude certain faults and say okay that type of failure that failure mode I don't care about it. And you would do that if that failure mode is its likelihood of occurrence is going to be very low in relation to the safety requirements. So I don't even know how that type of failure mode would have factored into the analysis in the first place.
But you could basically say oh that mode of failure I'm going to ignore it because it's unrealistically low. It's not going to affect anything. And when you do that you need to justify and document the exclusion. So somebody somewhere at some point in time wanted to ignore a certain failure mode of the equipment. And they have a clause now that allows them to do that so anybody listening edward.marszal at Kenexis.com I'm sure everybody knows how to get a hold of me. It's no secret. If you know of an example of where this is important please send it to me.
I would love to hear about this because I I just can't figure any reason why you'd want to do this. And furthermore this clause has an informative note which is note. Further information about fault exclusion can be found in ISO 13849-1 and ISO 13849-2. Huh if it wasn't so early in the morning for me here I might actually go try to dig up those standards. But I'll put that on my to do list for another day. 13849. Note to self. Okay let's move on to the next clause which is the meat clause 1145
## Table 6 – Minimum HFT Requirements by SIL
contains table 6 which is going to tell you what your minimum hardware fault tolerance requirements are. Clause 11.4.5 is one sentence which simply says the minimum HFT the minimum hardware fault tolerance for a SIS or its SIS subsystems implementing a SIF of a specified SIL shall be in accordance with table 6. So
on a SIF by SIF basis you're going to determine what the hardware fault tolerance requirements are because different SIFs are going to have different SIL targets and ultimately the hardware fault tolerance is a function of SIL which we're going to see when we look at table 6 and so you're going to need to meet the requirements that are in table 6. And then also there's exceptions more exceptions in clause 11.4.6 and 11.4.7. And yeah we'll get to those in just a second. So let's start.
Before we get to table 6 let's take a look at the note which precedes table 6 that note says the HFT requirements in table 6 represent the same minimum system or where relevant the subsystem redundancy. So hardware fault tolerance requirements represent a minimum redundancy. But again and I just kind of broke the note halfway in the middle here. They really don't because we're only looking at redundancy that is designed to prevent dangerous failures. Redundancy that's designed to prevent spurious failures is ignored.
So how you connect and wire and program your devices together is an essential part of that hardware fault tolerance not just redundancy.
Two out of two voting has two devices but it has zero ability to prevent a dangerous failure from causing your entire subsystem to fail. Okay the next sentence in the note says depending on the application device failure rate and proof testing interval additional redundancy can be required to satisfy the failure measure for the SIL of the SIF according to 11.9. So the other note that says even though you meet the redundancy requirements the hardware fault tolerance requirements you still might not be able to achieve your PFD. So just looking at this table is not enough.
You also need to run your SIL verification calculation in accordance with clause 11.9 to make sure that you met your SIL target. So this table is not enough. Go to clause 11.9 run your SIL calx to make sure you achieved your PFD target. Okay so let's dive into table 6. Table 6 is titled minimum hardware fault tolerance requirements according to SIL. Now this table has two columns. The first column is SIL and the second column is minimum required HFT or minimum required hardware fault tolerance.
So if you know what the SIL is the table will tell you what the hardware fault tolerance requirement is. And it seems simple but it was made ever so slightly complicated for SIL 2. Because SIL 2 shows up twice. Okay let me I will come back to SIL 2 in just a second but the table says
SIL 1 for any mode of operation the minimum hardware fault tolerance requirement is zero. So one out of one voting simplex voting is fine for for SIL 1 for SIL 3 you need one degree of hardware fault tolerance requirements. So according to 61511 a single simplex device is not acceptable in SIL 3 no way no how ain't gonna happen. You need one degree of hardware fault tolerance which you're gonna get looking at the big four standard architectures. You will get that with one out of two voting or two out of three voting.
You will not get it with one out of one you will not get it with two out of two. SIL 4 in any mode of operation requires you to have a minimum hardware fault tolerance of two none of them. You need something like a one out of three vote or a two out of four vote which are not very common. Most of the resources out there don't even give you equations to calculate those voting arrangements because they're kind of on the rare and kind of odd side but then again SIL 4 in and of itself is on the kind of rare and kind of odd side.
I've mentioned several times now that I've been doing this for about 30 years and I have never actually implemented a SIL 4 system we had a requirement for SIL 4 in one instance in an offshore oil production application and we decided to go with a SIL 2 pneumatic system and a SIL 2 electronic system that were separated to the absolute maximum degree conceivable. Granted still there's you know all the valves are sitting on the same pipe so there's still going to be some degree of common cause between those systems. But I digress. So
SIL 1 zero hardware fault tolerance. SIL 3 1. SIL 4 2. SIL 2 gets a little bit strange because there are two entries. Now you'll notice that for 1 3 and 4 I said any mode of operation. What did I mean by mode of operation? I mean is it a low demand mode or is it a high demand mode? And when I say high demand mode I also mean continuous mode because high demand mode and continuous mode is a distinction without a difference. You treat them the same way mathematically when you're doing calculations for the most part.
Some of you also watch are on the Kenexis mailing list if you're not on the Kenexis mailing list you should get on the mailing list to find out at a minimum what the monthly topic of our webinar is.
So the webinar last month was running calculations for high demand mode and continuous mode safety functions where you ignore your test interval and focus on the frequencies because if you fail you're going to have a consequence almost immediately before you run your test. So the chances of you your test actually finding a failure are essentially zero because if a failure would have happened you would have basically had an explosion. Now things get a little bit tricky when you start looking at redundancy and diagnostics.
So a diagnostic might tell you about a failure before the plant explodes and it might tell you about it quick enough that you could do something about it. But anyway as a result of the basically complete ineffectiveness of testing in the high demand mode and continuous mode. For SIL2 if you're in the low demand mode you don't need any hardware fault tolerance. HFT is zero.
But if you're in high demand mode or continuous mode you need one degree of hardware fault tolerance. And when you start running your calculations both ways you're going to see that oh because I can't take any credit for testing in achieving my performance goals I actually do need some redundancy even in SIL2 to be able to get to that tolerable frequency of the event. Because when you're in the high demand mode again I recommend the webinar if you want more details but basically failure of your safety function is your initiating event that causes the consequence.
So you need to limit that frequency. Alright so that's basically the big picture. Most of the time so generally let's throw out SIL4 because that's so unusual to happen. And let's throw out high demand mode because that's so unusual to happen. Which I cover why that's so unusual to happen in that webinar.
Basically your basic process control needs to be fixed if you're in high demand mode as opposed to putting in high demand mode safety function. But I will put that on the table. So if you kind of assume that everything is a low demand mode and ignore SIL4 what you're left with is SIL1 doesn't require any hardware fault tolerance SIL2 doesn't require any hardware fault tolerance SIL3 requires one degree of hardware fault tolerance. So I just made it very simple for you the only time you need to think about
## Clauses 11.4.6 and 11.4.7 – HFT Exception Clauses
hardware fault tolerance is if you're in SIL3 and you're going to need at least one degree of hardware fault tolerance there. And I bet you when you run your SIL verification calculations all got a lot more text here in clause 11.4 but most of this text is the two get out of jail free clauses. So the next two clauses are 11.46 and 11.47 now these two clauses when you read them when you look at them you can understand the words that they're saying. But like me probably never understand why they put these clauses in who would use them under what circumstances they would use them.
But you know we put in some basically 11.47 and 11.46 are two get out of jail free cards that say eh you know what? If you don't like that hardware fault tolerance forget about it. Just decide whatever you want. No it doesn't say that exactly come on oh what what does it actually say okay so for clause 11.46 it says for a SIS or SIS subsystem that does not use FVL or FVL programmable devices and if the minimum hardware fault tolerance as specified in table six would result in additional failures and lead to overall decreased process safety then the HFT may be reduced.
This shall be justified and documented. The justification shall provide evidence that the proposed architecture is suitable for its intended purpose and meets the safety integrity requirements. So
1146 basically says that if adding redundancy is going to make your system less safe then you don't need to do it. Hmm. Okay. It says you can decrease your hardware fault tolerance. Can you decrease it all the way down to zero? Yes you can. That's an 1147 which we'll get to next. And unless you're at SIL basically basically kind of just boiling it all down how can redundancy make your plant less safe? Well you're going to necessarily increase your spurious trip rate when you put in redundancy.
If your spurious trips are as dangerous or more dangerous than the consequence you're trying to prevent in the first place then you might have an argument here that says okay that's the case. So but bottom line if adding redundancy makes the overall plant less safe the overall safety is lower then you don't have to do it that's what clause 1147 says once again I cannot give you an example of this. I cannot. If somebody has a good example please contact me. I would love to know. Next so that's the clause. The note is huge.
So let's dig into the note for clause 1146. Fault tolerance is the preferred solution to achieve the required confidence that a robust architecture has been achieved. So yeah if we want to be confident that we have a good safety system fault tolerance is a great qualitative ! attribute when 1146 applies the purpose of the justification is to demonstrate that the proposed alternative architecture is equivalent or better. All right let's continue on in the note. This may vary depending on the application and or technology in use. Examples include backup arrangements.
For example analytical redundancy replacing a failed sensor output by a physical calculation result from other sensors. Using more reliable items of the same technology. Changing for a more reliable technology. Decreasing common cause failure of impact by using diverse technology increasing design margins constraining the environmental conditions for example for electronic components decreasing reliability uncertainty by gathering more field feedback or expert judgment.
So you just read through a laundry list of ways to improve your probability of failure on demand numbers and kind of tolerate failures in ways that are not just a pure switch to an identical redundant component. Okay
I would argue that if you fail to a non-identical redundant component that is still hardware fault tolerance. But I digress a little bit. We've got clause 1146 that lets you dance around. If you want to use it I don't know anybody who's ever used it but if you want to use it please document what you're trying to achieve and what you're trying to justify. All right.
## Diagnostic Coverage and Failure Rate Confidence
Next clause up is 1147 which states if a fault tolerance equal to zero results from applying clause 1146 the justification required by 1146 shall provide evidence that the related dangerous failure modes can be excluded in accordance with 1144 including consideration of the potential for systematic failures. So especially if you determine that your hardware fault tolerance is zero then you're going to need a justification. I mean you need a justification in clause 1146. 1147 says you really need a justification now buster.
Okay so 1146 1147. If you think you're going to be in a more dangerous condition because you applied fault tolerance then you don't need to justify it. There you go all right couple more clauses. And both these clauses are stealth mega requirements that are probably in the wrong place. So clause 1148 says FVL and LVL programmable devices. So a PLC or computer of any kind which you would typically use as a logic solver.
So programmable devices shall have diagnostic coverage is not less than 60% in my mind this should be in clause 11. 5 where it talks about what are the requirements for you to be allowed to use a device. That is big. That is big time. That clause right there probably excludes your ability to use any type of commercial off the shelf PLC. Not because it can't achieve a diagnostic coverage of 60% but because there's no way for you to know what diagnostic coverage it achieves.
You're guessing at best maybe you're going into a database that says oh a commercial off the shelf PLC has a safe failure fraction of 75% that is a I don't want to call it a wild guess but it's an amalgamation of a bunch of different equipment types at best. So if you're using a commercial off the shelf PLC do you know that it achieves a diagnostic coverage greater than 60% if you don't then this clause right here 1148 says you shouldn't be using it at all. But at the same time it's kind of hidden at the end of minimum hardware fault tolerance glossed over repeatedly.
Everyone kind of looks at that and goes okay yeah that seems like a reasonable number we probably achieve. Well a reasonable number we probably achieve is not confirmation of achievement of a requirement. Well stew on that one. Stew on that one. And we'll probably come back to it in clause 11.5 when we start talking about what equipment is acceptable and what equipment is not acceptable. All right. The last clause in 11.4 is 1149 which is also completely out of place. Clause 11.4.9 belongs in clause 11.9 where we talk about how to do SIL verification calculations.
And in clause 11.9 this statement is kind of danced around as opposed to being definitive. So when I read clause 11.4.9 it says reliability data used in the calculation of the failure measure shall be determined by an upper bound statistical confidence limit of no less than 70% this is precise. This is definitive. You have one and only one safety margin that you are required to put in your calculations.
That is that the failure rate data you use for your analysis shall be determined with an upper bound statistical confidence of 70% single sided confidence of 70% that's a safety margin that's going to gross up your failure rate over the average number. So if you say if you have one failure every 10 years you could say that your failure rate is 0.1 per year. But that is with a 50% confidence. So 50% of the time the actual failure measure is going to be higher. 50% of the time it's going to be lower.
We can use a statistical technique to gross up that failure rate to say 70% of the time the actual is going to be lower but 30% of the time it's going to be higher. So that's the conservatism that you're going to be built in the chi squared test is the way to do it.
I would recommend going to a book from David Smith called Reliability Maintainability and Risk. Old Smithy. That's another long story. But he provides an excellent explanation of how to calculate confidence limits on your failure rate data specifically single sided confidence limits that we're going to use in our SIL verification calculations. So pick up a copy of Reliability Maintainability and Risk from David Smith.
It should be on everybody's bookshelf. I have edition three is where I first learned how to do it. Edition seven is the most recent one that I have and there might actually be a new one since then. Also talking about Smithy Faradip is his failure rate data from his company Contexas Technus which is a great resource when you're trying to identify failure rates of equipment actually in operation. And that database is particularly good because it's for real.
And you might be surprised at how high your failure rates actually are in comparison with what maybe your equipment vendor might tell you that it is estimated to be in a somewhat idealized world as opposed to that real world where your equipment is actually installed alright so we'll talk more about confidence limits when we get to clause 11.9 which is the clause that talks about how to do the calculations. But with that that is all that we have for today.
And we were able to crunch through clause 11.4 hardware fault tolerance which is I mean if you look at the standard it's a page and two sentences and the page includes that table which takes up a lot of space. Next week we're going to get into requirements for device selection.
Clause 11.5 talks about what devices are you allowed to use in a safety application what devices are you not allowed to use in safety applications and what are the special things you need to do to allow a device to be used in a safety application that you might not necessarily have to do if that device was used in a basic process control application. But we will get to all of that next week. Talk to you then.
## Kenexis Vertigo Software Overview
Now that you've heard some insights on technical safety functional safety and the IEC 61511 standard let me tell you a little bit more about how to easily and effectively implement the safety life cycle using the Kenexis integrated safety suite and our SIS safety life cycle management tool Vertigo Vertigo is a comprehensive tool set for performing assessment calculations documenting and maintaining the design of safety instrumented systems.
Analysis begins with importing or synchronizing a list of safety instrumented functions with their definitions and associated performance targets from our open PHA tool for HAZOP and LOPA documentation. Each safety function can then be analyzed by performing a SIL verification calculation complete with a collection of tools for optimizing designs and a database of thousands of potential instruments to define failure rates and diagnostic coverage capabilities.
After the SIL verification calculations are defined you can build an SRS by automatically generating a cause and effect diagram from the SIF definitions and other defined instruments. Each SIS instrument will include a customizable data sheet and general requirements that are applicable to the SIS as a whole and can be entered individually or even bulk imported from customizable libraries.
After the design phase you can even use Vertigo to track and document testing throughout the entire life of the facility Kenexis Vertigo is the most integrated easy to use enterprise tool for allowing the development of SIS design basis information more efficiently and effectively than any other software application.
Thank you.