In this episode of the Functional Safety Podcast: Safety is freedom from risk which is not tolerable — but what exactly is risk, and who decides when it becomes tolerable? In this episode, Ed Marszal works through definitions 3.2.57 through 3.2.64, the cluster of terms that underpin every risk analysis and SIL calculation in process safety. He explains why the standard’s definition of protection layer deliberately outruns the narrower independent protection layer used in LOPA, why the word ‘probability’ in the risk definition (3.2.61) can lead to dangerously wrong answers in high-demand scenarios, and why ‘safe failure’ (3.2.62) is terminology he has never once used in thirty years of teaching. Ed also traces the committee’s choice of ‘tolerable’ over ‘acceptable’ risk (3.2.64), the distinction between safe states and mere shutdowns (3.2.63), and the ISO 9000 pedigree of a quality definition (3.2.58) that leaves him scratching his head. Anyone who writes SIL verification reports, argues over spurious trip rates, or sits in hazard review meetings will find an hour of clarifying context that the standard itself never supplies.

Safety is the freedom from risk that is not tolerable, but what is risk and how do you know it is tolerable?

Please join Ed Marszal, President and CEO of Kenexis, for the latest episode of the inaugural season of our new Functional Safety Podcast on Spotify and Apple Podcasts where he continues his discussion of the IEC 61511 standard. Clause 3.2.57 through 3.2.64 is covered in this episode.

As a Principal Engineer (PE) himself with decades of experience in safety instrumented systems, Ed brings a unique perspective to this podcast, having actively contributed to the ISA 84 committee since 1994.

In this inaugural season, Ed will delve into the IEC 61511 standard, examining each word’s significance. He provides detailed insights into the standard’s interpretation and application, complemented by personal stories from his career and committee discussions.

Full Episode Transcript

Introduction and Episode Overview

Safety is the freedom from risk which is not tolerable. But what is risk? And how do you know whether or not it’s tolerable? More definitions coming up.

Welcome to the Kenexis Functional Safety Podcast. I’m your host, Ed Marszal, President and CEO of Kenexis. Kenexis is a technical safety consultancy that helps chemical process industry companies to analyze risk and design engineered safeguards like safety instrumented systems and fire and gas detection systems. Kenexis also provides the industry-leading suite of software tools, including our best-in-class Vertigo software for SIS Safety Lifecycle Management.

In this first season of the podcast, we are going to focus on the IEC 61511 standard, doing a deep dive into the standard, including more depth of information on what the standard means and how to apply it, brought to life with personal war stories and behind-the-scenes discussions of the committee members as we develop the standard in ISA 84 and IEC SC 65.

Before we start, a little disclaimer. I will be providing my opinion on technical and engineering topics. This information is provided on a best-effort basis and is of a general nature. The information presented in this podcast might not be applicable to your specific application. It is the obligation of every engineer to thoroughly analyze any system that they are designing and not blindly rely on any general advice presented in this podcast.

We left off at 3.2.57, which was the definition of protection layer. So we’re going to start back up at 3.2.58. And it’s going to be an interesting episode because we’re going to be talking a lot about the definitions that are related to risk analysis, kind of the Clause 8 and Clause 9 layer of protection analysis. So we’ll be talking about what it means to be safe. What is risk, which is a component of safety? What does it mean for something to be tolerable? So lots of good definitions. Let’s go ahead and get into it.

Protection Layer Definition

Kind of starting off in this vein, 3.2.57 is a protection layer.

Now, we all know about protection layers. We all know about independent protection layers. I wrote an entire book way back in the year 2000 about independent protection layers and layer of protection analysis. So protection layer is a concept that’s really ingrained in process safety in general and in the risk analysis that’s used to determine requirements for safety instrumented systems.

So what is a protection layer? It is any independent mechanism that reduces risk by control, prevention, or mitigation.

Well, that’s really interesting because most of the time when you’re looking at independent protection layers, you’re specifically looking at prevention. And the layer of protection analysis technique pretty much requires that it be preventive, that you detect that you’ve gone out of control and take an action to return yourself to a safe state before you have a loss of containment event. So mitigation is a little bit different in that it requires or it acts after the loss of containment has already occurred to make your consequences smaller.

And control is going to prevent you from getting into a dangerous state in the first place.

So the definition of protection layer is necessarily broad in the standard. But in terms of independent protection layer, as we would use it in layer of protection analysis, it’s a lot more narrow than this.

Now, there is a note to the entry for protection layer. Note one to the entry says, It can be a process engineering mechanism, such as the size of vessels containing hazardous chemicals, a mechanical mechanism, such as a relief valve, a SIS or an administrative procedure, such as an emergency plan against an imminent hazard. These response may be automated or initiated by human action.

So in the definition of protection layer, we’re giving it a very broad definition. There are a lot of ways to reduce risk that we need to consider when we’re assigning performance targets to our safety instrumented system. And all of those techniques are covered in the definition of protection layer, even though the definition of protection layer here and the definition of independent protection layer, as you would use in a layer of protection analysis, are not the same.

What we have here is a little bit broader, whereas when we get into layer of protection analysis and we say independent protection layer, we are definitely tightening up on that definition.

Quality Definition

Okay, moving on to 3.2.58. Another one of those definitions that I’m kind of left scratching my head as to why they needed to put it in here, it is the definition for quality.

And let me tell you, though, IEC and ISO are big on quality. The ISO 9000 standard is huge across all kinds of industries, including the process industries. And as a matter of fact, before I even get into the definition of quality, I’ll read the note first because the note goes right to the ISO standards. And it says, note one to the entry, see ISO 9000 for more details.

So definitely there are a lot of big proponents pushing ISO 9000 in a lot of applications. It is valuable. It does provide a lot of benefit. But I don’t know. I’m a little bit surprised sometimes by the degree to which it is given so much importance. Anyway, so 3.258, quality.

Definition is totality of characteristics of an entity that bear on its ability to satisfy stated and implied needs. Wow. Yeah. Okay.

So that’s a definition that is kind of very good from just kind of a general perspective. It doesn’t give me a lot of help in what things I need to do to achieve quality, but I guess that’s what the entire ISO 9000 standard is for. And with a definition that’s this kind of esoteric, I don’t know what benefit it provides to the reader of IEC 61511. But hey, there you go. The definition of quality. All right.

Random Hardware Failure Definition

Now let’s move into something that’s got a little bit more teeth with regards to SIS and a little bit more. It’s relevant and it’s also at the same time a little bit of a cause for dissension or disagreement between standards committee members. And that is random hardware failure.

So as we go through this series of podcasts, you’re going to hear me talk many, many times about random hardware failures versus systematic failures and human failures and how when we’re running SIL verification calculations, you really need to focus on the random hardware failures because trying to include systematic failures and human failures into that process is it’s quite literally going to lead you down the wrong road and get you to make bad choices about your design instead of making good choices.

So let’s hit the definition. Random hardware failure is failure occurring at a random time, which results in one or more of the possible degradation mechanisms in the hardware.

So it’s one of those things that the piece of equipment isn’t born with this failure. It’s going to happen later and it’s going to happen at a random time. And we are definitely talking about the hardware itself as opposed to the engineering that went into it, how it was implemented or how it was maintained.

Now there are a large series of notes, associated with this, or actually the notes are large. There’s only two of them.

All right, the first one, note one to entry. There are many degradation mechanisms occurring at different rates in different components and since manufacturing tolerances cause components to fail due to these mechanisms after different times in operation. Failures of a total equipment comprising many components occur at predictable rates, but at unpredictable times.

Okay, let me unpack the note a little bit. Basically, when we start out with saying there are many degradation mechanisms, that means that there are a lot of different ways that a device can fail. Those are also known as the failure modes. So what specifically happened to cause the device to no longer be able to operate? So we’re acknowledging that there are many different failure modes. And the rates at which these different modes happen is going to vary based on the manufacturing process and the tolerances or degree of over-design that’s built into the manufacturing process.

So while we know that failures are going to be possible, there’s really no good way to predict precisely when they’re going to happen. But at the same time, we can look at the overall trends in terms of rates of failure, even though we don’t know exactly when that failure is going to happen. So we can predict rates, but knowing the exact timing is not realistic.

Okay, note number two to the entry. Two major differences distinguish random hardware failures from systematic failures. So we’re getting already into that discussion of there are random hardware failures, there are systematic failures, they’re different, it’s important to understand that they’re different, and it’s important to treat them differently.

Okay, so the first bullet point is a random hardware failure involves only the system itself, while a systematic failure involves both the system itself and a particular condition, and then it’s going to refer you to clause 3281. Let’s go ahead and scroll down to 3281 and see what that definition is. The definition at 3281 is the definition for a systematic failure.

So we will get to systematic failures, we’re just not getting to them yet, but this definition is referring to the definition of systematic failures. So we’re still in the second bullet of the note here, and we’re still talking about random hardware failures. That was just the first sentence.

The second sentence now says, then a random hardware failure is characterized by a single reliability parameter, for instance, the failure rate, while systematic failures are characterized by two reliability parameters, the probability of a pre-existing fault and the hazard rate of the particular condition.

So systematic failures, you’re kind of looking at failures in human actions that generate what is referred to here as a pre-existing fault. So some sort of action, whether it’s maintenance or design, has caused a device to not be able to operate.

Okay, the second bullet point says a systematic failure can be eliminated after being detected, while random hardware failures cannot.

So if, for instance, you miscalibrated a transmitter, once you know that you’ve miscalibrated the transmitter, you can calibrate it correctly and that will make that failure go away. Whereas random hardware failures, such as corrosion in a float type level measurement, you really can’t make them go away. That possibility that it’s going to corrode always exists. All right. Continuing on in the note, it says, this implies that the reliability parameters of random hardware failures can be estimated from field feedback, while it’s very difficult to do the same for systematic failures.

Okay, so right here in the definitions, they’re saying we can quantify based on real data, the failure rate of random hardware failures. But trying to quantify systematic failures is really hard to do. And that’s why the last sentence here in the note says, a qualitative approach is preferred for systematic failures.

So even in the definition section, they’re already leading you down the road of quantifying systematic failures, human interactions is a fool’s errand. You don’t have proper data to do it. It’s not something that’s repeatable. It’s not something that you can really quantify correctly or well based on experience. Instead, we’re going to use clause five, six, and seven to address systematic failures. We’re going to make sure our people are competent. We’re going to do verification of every step to make sure that failures are detected.

We’re going to do validation testing to make sure the system works. We’re going to do functional safety assessments and audits. That’s how we’re addressing systematic failures, not by putting some numbers into a calculation, doing some hand waving and saying it’s okay.

All right.

So this information, these notes are sourced from IEC 61508 part four, uh, in clause 3.6.5. So, uh, the, the definition and the notes came from that standard, although it says that the notes have been modified to more appropriately address the process industries. Okay.

Redundancy Definition

3.2.60 is redundancy. So important to understand what we mean by redundancy, uh, important to understand what we mean by fault tolerance, which is something that we’ve already talked about. Uh, and redundancy is generally the mechanism that we’re going to use to address fault tolerance, uh, hardware fault tolerance.

So what is the definition of redundancy? It is the existence of more than one means for performing a required function or for representing information. Okay.

Notes to the entry. Note one, examples are the use of duplicate devices and the addition of parity bits. And note two to the entry is redundancy is used primarily to improve reliability or availability. Okay.

So basically we have more than one way to execute the same function. That’s what redundancy is. And also the source for that definition is IEC 61508 part four, uh, clause 3.4.6.

Risk Definition and Probability vs. Frequency

All right, let’s move on to one of the most important, most loaded definitions in the standard. That is 3.2.61. And it is the definition of risk.

Everything that we’re doing in this standard is related to risk. We’re calculating the risk of the process plant. We’re determining how much risk is tolerable. We are determining how much risk reduction is required for our safety instrumented systems to perform. So knowing, uh, very clearly, very, very precisely what risk is, is going to be very important.

So the definition of risk is a combination of the probability of occurrence of harm and the severity of that harm. Probability of occurrence of the harm and severity of the harm.

Hmm. Boy, I said it was important to have a precise definition, but I’m going to take umbrage with the word probability. Um, we’ll come back to that in just a second. Um, note one to the entry is that the probability of occurrence includes the exposure to a hazardous situation, the occurrence of a hazardous event, and the possibility to avoid or limit the harm. And you know what? Hazardous situation, hazardous event, and harm are all definitions that we’ve already talked about.

Um, the source for this definition, and in ISO IEC world, there’s a lot of cross pollination of definitions to try to minimize the same word being used different ways in different applications. This, uh, definition is sourced from ISO IEC guide 51 and it’s clause 3.8.

So this is such a ubiquitously used term that they wanted to source it kind of higher in the IEC ISO, uh, hierarchy. Now, how, where, where am I taking umbrage with this definition and, and, and why is, why are my complaints relevant? Um, how do they impact, uh, your workflow and what you’re going to do with the standard?

Now, we say here that risk is a combination of probability of occurrence of harm and severity of the harm. Now, uh, on a topic that I have already brought up, um, there are three words that I will use to describe, uh, how often, if you will, an event happens. Those terms are probability, frequency, and likelihood.

Now, in my definition of risk, I would prefer to use the word likelihood because it is deliberately ambiguous. And I will continue to use it deliberately, ambiguously, whereas there are some busy bodies in the IEC and ISA standards committee world that want to, to tell you, want to convince you that likelihood and frequency are exactly the same thing. They are not. Um, you know, definitions be damned. I am a strong proponent of reserving use of the word likelihood to be, uh, definition, definitionally ambiguous.

I want likelihood to be maybe a probability, maybe a frequency, depending on the context, because the words likelihood are the words frequency and probability are very precisely defined.

So probability is for a given event, uh, a, a numerical representation on what the outcome of an event will be. So if you flip a coin, the probability of it landing heads is 0.5. The probability of it landing tails is also 0.5. Frequency looks at how often an event happens in time.

So probability and frequency are completely different terms, and you are going to use them for different applications, and you are going to use them to define risk in different applications.

Now, as we mentioned in the mode of operation discussion, a couple, uh, podcasts ago, months ago, there is low demand mode and high demand mode. And when you’re in high demand mode, you deal with frequencies. In low demand mode, you deal with probabilities. And it’s all related to the, uh, the likelihood that a failure is detected by a test or a failure is detected by an actual demand.

So the same way risk, whether you use frequency to calculate risk or you use probability to calculate risk, uh, depends on the rare event approximation. So probability makes sense. So probability makes sense. If the event that you’re looking at, uh, is extremely rare, but if the event that you’re talking about happens very frequently, you’re going to need to use frequency when you’re calculating risk, as opposed to using probability.

So I am unhappy, uh, with the definition of risk as it’s presented here, because there are a lot of situations where if you calculate risk using probability, your calculation is going to be wrong. It’s going to be dramatically dangerously wrong.

So let me give you an example. If you’re an insurance adjuster and you have, you’re selling life insurance and the amount of premium that a customer needs to pay for a life insurance policy is based on the risk of the, them dying. Now, when you’re an insurance underwriter and you want to calculate, uh, the, the, the, the premium that you need to charge someone, you can basically say, well, if the insurance policy is a million dollars, uh, I’m going to multiply that million dollars by a likelihood, you know, mark it up to make a profit.

And then that’s what I’m going to charge for the life insurance premium. Well, if you basically say, well, what is the, if I have a pool of a million people, what is the probability of one of them dying this year? It’s going to be a number that’s pretty close to one. So you multiply one by a million dollars. And that’s kind of, uh, you know, what, what you would charge for the policy. But the, the issue with that is, well, what do you mean in terms of the pool for the probability? Is it the probability for one person dying or is the probability for anybody in the pool dying?

But then if you look at the probability of anybody in the pool dying, your numbers are very high. Um, also, if you look at that million person pool in a given year, it’s very likely that more than one person will die. So you, it’s really more of a frequency calculation. That’s going to give you a legitimate expected value of loss, uh, which is, you know, frequency multiplied by severity. And that’s what the insurance company is going to use to calculate what your premium is.

So 3.261, we have to go with what IEC says. We have to go with ISO says, but in the back of your mind, you should always remember that strictly saying probability is not precisely correct. And depending on the frequency at which that harm is occurring, it may be much more appropriate to use frequency instead of probability when you’re calculating the risk. Okay.

Safe Failure Definition and Notes

Moving on. There is clause 3.2.62, which is going to define a safe failure.

So, uh, a safe failure is a failure which get, which favors a given safety function. Um, wow.

Uh, again, you know, I, I, I never, I never failed to be amazed by looking at some of these definitions and going, who thought that this definition was a good idea? Uh, a failure which favors a given safety function favors is not really an engineering term. Now, is it, uh, so a failure which favors a given safety function.

I’ve been teaching the difference between safe failures and dangerous failures for about 30 years now. I can guarantee you, I have never used this definition to explain what a safe failure is.

A safe failure, uh, if you ask me, is a failure which results in the safety action being taken, even though there was no hazardous condition present, which would have necessitate, necessitate, uh, the action to occur. So, the safety function activated, even though the process parameter that you’re measuring never went into an out of control, uh, position or an out of control value. Uh, the safety function activated when we didn’t want it to activate.

Um, any one of those would have been a better definition than the definition we have here in the standard. But, it is what it is. So, uh, in addition to giving you a relatively poor, uh, definition, we’re going to go ahead and lather on five notes into this definition to just kind of let, just turn it up a notch. Okay.

Note one to the entry says, a failure is safe only with regard to a given safety action. So, um, they’re basically trying to say here, uh, that, uh, again, even the note isn’t very clear.

Uh, what you’re trying to get to here is the fact that when we, we say safe failure, that means that the safety function activated. We didn’t want it to activate that it activated. So we moved to a, you know, air quotes, safe state, but the standards committee wanted to impress upon people that, well, we’re calling it a safe failure, but it might not, might not actually be safe. So, you know, shutting down your plant is not necessarily a safe thing to do.

Uh, so we’re basically saying, okay, well it activated when you didn’t want it to, and it went to the, uh, proverbial safe state, but that safe state might not actually be safe.

So, you know, that’s another, um, vote for calling it a spurious failure as opposed to safe failure, but we don’t, the definitions, the standards terminology, all use safe failure. Note two to the entry.

when fault tolerance is implemented, a safe failure can lead to either a operation where safety action is available, but with a higher probability of success on demand or lower likelihood to cause a hazardous event, or B spurious operation where the safety action is initiated. Okay.

So, what clause two is trying to convey is that a safe failure might not actually result in a trip. So, like, kind of the second bullet point says, yeah, a spurious activation of the safety function, you cause the safety function to operate. That is a safe failure. But, when you have fault tolerance against spurious trips, like in a two out of two voting arrangement, a single spurious failure, single safe failure does not result in a trip, because you need both of the devices to agree to trip before you have a trip in a two out of two voting arrangement.

Um, so, uh, specifically they’re talking about safe fault tolerance. And, uh, when you’re in a two out of two vote, and you have a spurious trip of one of them, you’re more likely now to actually have a spurious trip overall, because only one more device is required to go into a trip state.

Note three to the entry. When no fault tolerance is implemented, safe failures result in the initiation of the safety action, regardless of the process condition. This is also known as a spurious trip.

Okay.

Well, note three, uh, is actually kind of a better definition than the definition itself. Uh, but, uh, well, hey, at least they got it into the notes, if not into the definition itself.

Note four to the entry. A spurious trip may be safe with regard to a given safety function, but may be dangerous with regard to another safety function.

So, uh, uh, a spurious trip might result in a, another condition, which is even more dangerous possibly than the trip that you just activated.

So, uh, analysis of spurious shutdowns is something that should always be considered, always be implemented. When you’re divine, defining a safety instrumented system, look at the activation of the safety function and see what new hazards you’re creating.

Note five to the entry. Spurious trips may also have detrimental effects on the production availability of the process. Okay.

That would seem very obvious that if you spuriously shut down your plant, your plant is shut down. And when your plant is shut down, you can’t make your product. Uh, so if that wasn’t completely obvious to you, hopefully after you read note five, it will now be obvious to you.

Safe State Definition and Notes

So let’s take that, uh, discussion of safe failure and go to a related definition, three to 63 for safe state. Safe state is the state of the process. When safety is achieved.

Now, the objective of all safety instrumented functions is to, uh, achieve a safe state when safe operating condition parameters have been violated.

Now, what is a safe state? Uh, the kind of easy colloquial way to put it is a shutdown, but sometimes a shutdown is not the safe state. You need to take some other action. Like you need to put steam into the process unit. You need to nitrogen blanket your, uh, your storage vessel. There are maybe some active things, active conditions that need to be present in order for you to have a safe state. So safe state is broader than just a shutdown.

It is whatever the collection of actions are that are required to make sure that a hazardous event that causes harm does not occur.

Now, safe state itself has four notes to the entry. Note one is some states are safer than others. And in going from a hazardous condition to the final safe state, or in going from the nominal safe condition to a hazardous condition, the process may have to go through a number of intermediate safe states.

So, um, very, uh, interesting, uh, note to the entry that, um, getting to a safe state can be a complex series of steps. Um, I, you know, the, the preponderance of your safety functions are simply going to stop things. But if stopping a process, for instance, can result in a flammable condition occurring in your vessel, you may need to kind of step down the, the concentration of flammable materials in your vessel by injecting nitrogen before you can vent it to atmosphere, potentially letting air in.

So, uh, that would be for those of you familiar with the phrase walking around the nose on the flammability curve might be a multiple step process. That’s kind of what’s described in note one.

Note two, for some situations, a safe state exists only so long as the process is continuously controlled. Uh, our safety function is not a shutdown function. Our safety function is a controller. And if that controller stops working, that failure of control results in the dangerous condition. Hmm. Interesting stuff. Uh, continuing on with that note, such continuous control may be for a short or for an indefinite period of time.

Note three to the entry, a safe state, which is safe with regard to a given safety function may increase the probability of a hazardous event with regard to another safety function. In this case, the maximum allowable spurious trip frequency for the first function can consider the potential increased risk associated with the other function.

All right. This note, um, we’re going to tackle, um, we’re going to tackle when we get to, uh, the safety requirements specification section in clause 10. Um, in terms of minimizing the spurious trip rate, there are situations where your shutdown action creates a new, actually creates a consequence. So it, instead of hazardous event, a occurring hazardous event, B is going to occur because you activated your safety function spurious.

So, uh, uh, that is something that needs to be considered and it might be completely considered by minimizing the spurious trip rate and setting a performance target in your safety requirements specifications for the maximum allowable frequency of a spurious trip.

All right. Note four to the entry. This definition deviates from the definition in IEC 61508 part four. to reflect differences in process sector terminology.

So yes, a safe state in a process industry application, uh, is going to be dramatically different from a safe state in a railway or a self-driving car. So, hey, it makes sense that we’re going to have our own more precise definitions. All right.

Safety Definition and Tolerable Risk

Another key definition is, uh, 3.2.64. It’s kind of, uh, the, one of the ones that started it all. It’s actually in the title of this standard. And that is the definition of safety.

So 3.2.64 says safety is freedom from risk, which is not tolerable. Very short definition. Very precise definition. Very loaded. And kind of starts referring you all over the place.

So basically, it starts by saying it’s the freedom from risk. So we have to know what risk is. And you know what? Just a couple minutes ago, we talked about what risk is. Uh, so, uh, that’s, uh, been defined. Uh, so we know what risk is. And it’s freedom from a certain degree of risk. And that degree of risk is a risk, which is not tolerable. So we know what risk is.

Tolerable is another kind of, uh, liberal arts squishy term. It’s, uh, more of an opinion than a fact. Um, so that’s going to require a little bit more discussion.

So let’s look at the note to the entry. Note one to the entry says, according to ISO IEC guide 51, where did that come from? We just talked about guide 51, uh, in the definition of risk itself. Uh, so this is, you know, safety and risk are just higher level terms that are going to be used in so many applications in so many industries. We’re going to try to tighten up that definition as much as possible. So we are referring to a, a higher level ISO document for this definition. Okay.

So according to ISO guide 51, the terms acceptable risk and tolerable risk are considered to be, synonymous. Hmm. Hmm. Okay. Before I, uh, pick at that thread, uh, the, the, the document here also says that the source of this standard is I, ISO IEC guide 51, and it will be clause 3.14, uh, uh, in that guide guidance document. If you look at it. Okay. Um, acceptable risk and tolerable risk.

Thus it’s, it’s again, one of those things where you’ve got a squishy, um, liberal arts kind of term that is subjective and open to interpretation. And engineers don’t like things that are squishy, subjective and open to interpretation. So they define things basically in defiance of human nature and language as it’s been used for thousands of years. Uh, hoping that their adamancy in creating a definition is going to cause all human beings to stop being human and behave like robots. Um, good luck with that.

Okay.

So acceptable risk and tolerable risk in just normal human speech are going to have different meanings and different connotations. Um, acceptable is more accepting. It’s more acceptable. Acceptable is more acceptable than tolerable. How do you like that for a circular definition? Um, we are willing to allow a situation that we call acceptable, uh, ease more easily than we’re willing to allow a situation that’s tolerable.

Acceptable means I can live with this. It’s okay. I don’t need to do anything about it. I’m going to accept it. Tolerable has an implication that, well, I am not happy about this. I don’t want to allow this, but I have no other options. So I’m going to tolerate it.

So in the actual real world, uh, busy body engineers trying to, uh, define away the English language aside, acceptable and tolerable are different words that have different meanings. I mean, if, if, if they meant exactly the same thing, why would we have two different words? Uh, so, be careful how you use this.

Uh, it’s much more acceptable to use the phrase tolerable risk than it is to use the word acceptable risk because of the implication that tolerable means that I really have no choice in this as opposed to acceptable meaning, yeah, yeah, you know, I looked at a bunch of options and this is what I want. So, um,

I prefer to use the word tolerable and I’m very happy, uh, that our definition uses the word tolerable. Now, what is a tolerable risk? Ooh, that is a whole other book chapter in and of itself. I would, uh, recommend that you read my book on layer of protection analysis, which is available from ISA. There’s an entire section on, uh, acceptable risk, the different philosophical mechanisms for determining what amount of risk is acceptable or tolerable, uh, and kind of numerical methodologies and benchmarking techniques that you use to make the determination of what is tolerable.

But that’s kind of outside the scope, uh, of what we’re going to talk about here, uh, today, although when we get to clause eight and we’re talking about hazard and risk analysis, I think I might even designate an entire, uh, episode of this podcast just to the concept of tolerable. Tolerable risk because it’s something that’s not well understood and very much underappreciated. All right.

So with that, we have gone all the way through, uh, clause 3.2.64 and we’re going to pick it up again in the next episode. When we start talking about safety functions, safety instrumented functions, safety instrumented systems, and all the confusion as to why I need like 10 different definitions to know what a safety instrumented function is. Lots more on that. Next time we will see you then.

Now that you’ve heard some insights on technical safety, functional safety, and the IEC 61511 standard, let me tell you a little bit more about how to easily and effectively implement the safety lifecycle using the Kenexis integrated safety suite and our SIS safety lifecycle management tool, Vertigo. Vertigo is a comprehensive tool set for performing assessment calculations, documenting, and maintaining the design of safety instrumented systems.

Analysis begins with importing or synchronizing a list of safety instrumented functions with their definitions and associated performance targets from our Open PHA tool for HAZOP and LOPA documentation. Each safety function can then be analyzed by performing a SIL verification calculation, complete with a collection of tools for optimizing designs and a database of thousands of potential instruments to define failure rates and diagnostic coverage capabilities.

After the SIL verification calculations are defined, you can build an SRS by automatically generating a cause and effect diagram from the SIF definitions and other defined instruments. Each SIS instrument will include a customizable data sheet and general requirements that are applicable to the SIS as a whole and can be entered individually or even bulk imported from customizable libraries. After the design phase, you can even use Vertigo to track and document testing throughout the entire life of the facility.

Kenexis Vertigo is the most integrated, easy-to-use enterprise tool for allowing the development of SIS design basis information more efficiently and effectively than any other software application.