Kenexis Functional Safety Podcast
Including human failure rates in SIL verification seems rigorous, yet Ed Marszal argues it is actively counterproductive. This episode unpacks why systematic human errors defy time-based probability models and why common-cause effects render redundancy useless against miscalibration or misdesign. The discussion then moves through the reliability data requirements of Clause 11.9.3, the uncertainty provisions of 11.9.4, and the iterative design improvement guidance of 11.9.5 that Ed regards as unenforceable suggestion dressed as requirement. Engineers wrestling with whether their failure rates are credible, their confidence limits conservative enough, or their improvement strategies sound will find the committee perspective and practical warnings essential listening.
Do not incorporate human failure into your SIL verification calculations, it will actually lead you to do the wrong thing…
Listen in for more information and thoughts on this important topic as it is discussed in more detail by an industry expert.
Tune in to the latest episode of the Kenexis Functional Safety Podcast, hosted by Ed Marszal, President and CEO of Kenexis. Now available on Spotify and Apple Podcasts, Ed offers his expert insights on the IEC 61511 standard.
With decades of experience in safety instrumented systems and as a Principal Engineer, Ed has a unique perspective to offer. He has been an active contributor to the ISA 84 committee since 1994, adding to his deep understanding of the field.
In this inaugural season, Ed delves into the IEC 61511 standard, unpacking the meaning behind each word and providing a thorough interpretation of its application. Through personal stories from his career and committee work, he offers valuable context and insights for professionals in the industry.
Full Episode Transcript
KENEXIS FUNCTIONAL SAFETY PODCAST — S1E47 TRANSCRIPT (Markdown)
Cleaned & reflowed for web publication and AI crawlability.
The JSON-LD block below is schema.org structured data. If your CMS lets you
add raw HTML to a post, paste it into the page
body — crawlers read it either way). Fill in PLACEHOLDER_EPISODE_PAGE_URL
once the post exists. Everything from the "# Kenexis Functional Safety
Podcast…" heading down is the transcript body — paste it into your post.
–>
"`html
"`
# Kenexis Functional Safety Podcast — Season 1, Episode 47: IEC 61511, Clauses 11.9.3 to 11.9.5 (SIL Verification Data, Uncertainty, and Design Improvement)
—
## Episode Teaser and Introduction
Do not incorporate human failure into your SIL verification calculations. It will actually lead you to do the wrong thing.
Welcome to the Kenexis Functional Safety Podcast. I'm your host, Ed Marszal, President and CEO of Kenexis. Kenexis is a technical safety consultancy that helps chemical process industry companies to analyze risk and design engineered safeguards like safety instrumented systems and fire and gas detection systems. Kenexis also provides the industry-leading suite of software tools, including our best-in-class Vertigo software for SIS safety lifecycle management.
In this first season of the podcast, we are going to focus on the IEC 61511 standard, doing a deep dive into the standard, including more depth of information on what the standard means and how to apply it, brought to life with personal war stories and behind-the-scenes discussions of the committee members as we develop the standard in ISA 84 and IEC SC 65.
Before we start, a little disclaimer, I will be providing my opinion on technical and engineering topics. This information is provided on a best-effort basis and is of a general nature. The information presented in this podcast might not be applicable to your specific application. It is the obligation of every engineer to thoroughly analyze any system that they are designing and not blindly rely on any general advice presented in this podcast.
## Human Failures and SIL Verification Scope
All right, so we're still in Clause 11.9, which is the clause that talks about SIL verification calculations, what you need to consider, what you don't need to consider. In the last episode, we went through Clause 11.9.1 and 11.9.2, and we're moving on to Clause 11.9, which is three more sub-clauses, but I'm going to backtrack, not necessarily backtrack, but in the last episode, I promised that I would provide a little bit more depth, a little bit more commentary on human failure in the SIL verification calculations.
And what is proposed by the Standards Committee, or it's not proposed, what is required in the standard is that you consider random hardware failures. So it looks at the dangerous undetected, the DU failures, the DD failures, dangerous detected, and then the dangerous never detected, and wants you to quantify the random hardware failures.
It doesn't talk about systematic failures. It doesn't talk about systematic human failures. So when you're looking at those factors, do they contribute to the PFD?
Absolutely. Human failure is a big deal. Humans mess things up all the time. Quantifying how often they mess things up is a little bit more difficult to do. And incorporating those failures into your calculations is difficult, number one. Number two, it will lead you to do the wrong thing more often than it will lead you to do the right thing. And this is going to tie into Clause 11.95 that we're going to get into.
In Clause 11.95, what it's basically doing is it's telling you, if you don't achieve your SIL target, what do you do about it? And it turns out that if the reason that you don't achieve your SIL is because of human interaction with the safety instrument and function, then it is most certainly going to tell you to do the wrong thing. It's going to tell you to do stuff that will actually make the situation worse, not better. And that might not be reflected in your SIL verification calculations. All right.
So going back, right from Clause 11.1, it talks about calculated failure measures. And then going into Clause 11.92, it talks about random hardware failures. So the calculated failure measure of each SIF due to random failures, not systematic failures, not human failures, needs to be quantified.
And when we're looking at the failure rates to put into our calculations, Clause 11.9.2, sub clauses B, C, and D all say random hardware failures. It doesn't say human failures. The only reference to human failures is in Clause, sub clause K, which says the estimated likelihood that operator response would cause a dangerous failure of the SIS. But we clarified last week that what we're talking about right there is in the rare strange case that, you know, most SIS practitioners disagree with, that you include a human being as your logic solver.
Well, then you need to consider their probability of failure in the overall calculus. Of course, I would never consider anything that has a human being as the logic solver to be a safety instrumented function. That's something that can't achieve SIL 1 for the most part. Okay, that's debatable in mathematical terms. But in risk analysis and standards compliance, we generally wouldn't give you anything better than a risk reduction factor of 10 for something that includes a human being. So you're not even SIL 1, so why did you bust this standard out in the first place?
## Why Human Failures Must Be Excluded
So, now, why are we so adamant about not including random hardware failures?
The answer is that we want you to address human failures in other ways.
Basically, Clause 5, Clause 7, Clause 15, 14. Wow, my mind is… Well, we will get to it shortly. Human factors are dealt with administrative controls that are going to be discussed in Clause 5. So, functional safety planning, who does what when. Competency, making sure people know what they're doing, they're trained, they can demonstrate that they're able to perform their work tasks. Verification, where one person checks the work of another person to make sure that it was done correctly.
Validation, which is that final physical test on the safety instrumented system to make sure that it is working properly before you introduce hazards into the process or as us oil and gas guys go, before you put oil in. And then also functional safety assessments to make sure you're following your rules for your projects. Functional safety audits, which periodically look at the entire SIS program for the entire facility. So, that's the mechanism that we're using to achieve functional safety for human factors issues, human interactions with the system in terms of maintenance, design, etc.
Because maintenance and design can lead to a significant amount of failures. As a matter of fact, if you go back and look at overall failure statistics, most failures related to safety instrumented systems occur because the specification of what the safety system is supposed to do was wrong. So, the most common thing to get wrong in a safety instrumented system is 100% generated by human beings.
Now, that being said, when you're in the operational phase or even in the design phase, what is the benefit versus what is the liability of trying to include human failures into your calculations?
Well, number one, human failures are kind of task-based, not time-based. So, the probability that a technician miscalibrates a transmitter doesn't increase with time. They either got it right or they got it wrong when they did the calibration. So, this whole concept of unreliability, unavailability, where, you know, probability of failure increases over time, that kind of goes out the window with human interaction with the system in the first place.
So, the concept of adding an additional failure rate that is human interaction with the safety instrumented system, just mathematically, it doesn't check out.
But let's say we were to do it. And the example of this would be, let's say that we can get a dangerous failure of a safety instrumented system because a technician miscalibrates the sensor subsystem. Let's say I've got a one-out-of-one sensor subsystem, and I'm going to include the failure rate, some sort of failure rate, into the total failure rate of the system that's associated with a technician miscalibration of the device.
So, that's going to make the overall failure rate go higher. So, let's say we needed to achieve SIL 2, and we were able to meet it with a single device because we're testing very frequently, maybe, you know, once a year, once a quarter. And we achieved our SIL 1 target, and then we add in the human failure associated with the technician miscalibrating the device. And now we want to look at, well, what SIL are we achieving now?
And let's say that that additional failure rate associated with technician miscalibration caused us to not be able to achieve SIL 1. So, we've got a one-out-of-one device, we're testing, let's say, once a year, and now we're not meeting our SIL target. So, we need to improve the performance of the subsystem. Well, how do you typically do that? Well, you either increase testing, or you increase fault tolerance. So, if I increase fault tolerance, that could mean that instead of a one-out-of-one simplex system, I'm going to put in a two-out-of-three system. Did that fix the problem?
Mathematically speaking, it will, because you've got this comparison diagnostics, you have multiple devices, you have fault tolerance. Theoretically, you're going to achieve better performance.
You will achieve better performance, but not with relation to human factors, because the technician, if the technician miscalibrates device A, it's systematic. And systematic failures have 100% common cause associated with them. So, if they miscalibrate device A, they're going to miscalibrate device B, and device C.
So, putting in three transmitters instead of one transmitter to try to address human failures is basically a gigantic waste of money, because if they make the error, you're finished, and there's nothing that redundancy is going to be able to do to help you, because human failures have effectively 100% common cause. And to model that accurately, you would need to model that common cause in your SIS design.
Okay, so let's look at the other factor. Well, let's increase testing. So, I tested once a year, but now I'm going to test every six months. I'm going to test every three months in order to increase my SIL. Well, if the reason I am not achieving my target is because my technician miscalibrates, I'm making things worse because I am giving the technician more opportunities to make a mistake.
So, the techniques that you use to decrease your probability of failure on demand and increase your achieved SIL, when you look at systematic human failures, those techniques either don't do anything to make the situation better, or they will actually make them worse by giving the human more opportunities to screw up.
That is why we don't include human failures in your SIL verification calculations, period. No probability of misdesign, no probability of miscalibration, et cetera, et cetera, et cetera, et cetera. We just, you don't do it. It is counterproductive. It is worse than nothing. It is actually counterproductive. It will lead you to do things that will make your PFD in reality worse.
But if you're not intelligent, or if you're not sophisticated enough to know that you have 100% common cause related to systematic human failures, you might think that you're making the situation better, but in fact you are not.
So once again, on the SIL verification side of the fence, leave your calculations to be 100% random hardware failures. I do not want you including human factors in those calculations for any reason. All right. So that's kind of a little bit of an addition to Clause 1193, which is part and parcel of 1195 when we get to it. But we're not there yet.
## SIL Verification Modeling Approaches
Let's talk about the last thing that we didn't talk about last week in this section was the note to Clause 1192. And the note to Clause 11.9.2 talks about modeling approaches.
So the note says, several modeling approaches are available and the most appropriate approach is a matter for the analyst and can depend on the circumstances. Available means include, and then they're going to give you a bullet list, but they're also giving you a reference to IEC 61508 Part 6 Annex B. So more information on the approaches that you can use to run your calculations in 61508 Part 6 Annex B. And there is a list of the modeling approaches that it talks about.
There is five. First one is Cause-Consequence Analysis. Second is Reliability Block Diagrams, or technically you would be using Unreliability Block Diagrams to calculate failures. The third one is Fault Tree Analysis. The fourth one is Markov Models. And the fifth one is Petri Nets. Petri Nets, if you're not familiar, are kind of most similar to Markov Models, generally Petri Nets are not used for quantitative calculations, but you can jack them up, if you will, to make them quantitative. The same way you can make Bowtie Diagrams quantitative, but I am starting to digress a little bit.
All right, so the last sentence in the note says that the probabilistic calculations can be performed analytically or by numerical simulation. So numerical simulation, and the example they give is Monte Carlo simulation, is kind of, you don't have a precise answer. You're doing a whole bunch of numerical simulations and kind of mushing them together to get a result that has a failure distribution associated with it. So if all of your inputs have a failure rate distribution, you can calculate an output with a failure rate distribution using Monte Carlo simulation.
Most people are not going to look at a full range of error bands around their calculations. So ultimately, in the process industry, 6, 15, 11, most of the time you're going to follow the analytical solution equations. I deliberately don't say simplified equations. They are not simplified. They provide a comprehensive analytical solution. So anything that you do in a fault tree or anything you do in a Markov model, if you solve the Markov model with a, symbolically instead of numerically, you will get an analytical solution equation.
So the analytical solution equations are not some sort of shortcut. They are 100% rigorous and they will give you a 100% accurate solution, but they are not as flexible. shall we say. So if I perform an analytical solution for a two out of three voting arrangement where you have identical redundancy and you assume that your probability of failure when you start is zero, well, it's only good in that situation. So if the probability of failure when you start is not zero, you'll need to have a different equation for that.
Or if you're using diverse voting, you're going to need a different equation from that. So the equations are not less accurate. Don't let anybody tell you that. They're just less flexible.
Whereas if you go closer down to the bare metal, if you will, you're going to be able to be a lot more flexible in what you model using things like Markov models and fault tree analysis. So fault tree analysis is used very frequently in the process industries because it's kind of easier to understand, easier to work with. The tools are a lot more available. Arbor, for instance, the Kenexis integrated safety suite contains our Arbor model, which allows you to run fault tree analyses for SIL verification and a plethora of other different applications.
And Markov models are things that, you know, I know how to solve Markov models. Not that I would ever actually use one in the real world, but just to prove to the other people who are trying to flex with their ability to run Markov models, well, yeah, I could do a Markov model too, but why would you? The answer to why would you is if you have really complicated partial repairs with indeterminate time intervals. That's real difficult to model with a fault tree, whereas you can model that with a Markov model. But then again, most people are not using the Markov models to do that.
So, I mean, if you're looking at, you know, your typical software application that, you know, under the hood is using Markov models to do SIL verification calculations, they have a canned Markov model. And if you have a canned Markov model, you can analytically solve that to get an analytical solution equation. Why are you still running the numerical simulation on the Markov model when you can easily do an analytical solution that will run faster and be more accurate? Well, I will tell you the answer, marketing. It sounds more impressive, but it's really not any more accurate.
Anyway, I digress. Those are the options that are in the note to Clause 1192. Most of the time, you're going to end up with analytical solution equations.
In Vertigo, we use a bit of a mix. We have a series of analytical solution equations that we use. But in a lot of situations, we're going to basically have, under the hood, we will build a fault tree of your equipment and then solve the fault tree using the Arbor calculation engine for doing fault tree analysis, which provides us a lot more flexibility that we wouldn't have if we tried to do analytical solution equations for everything. Okay. So that is the note to Clause 11.9.2.
## Clause 11.9.3 Reliability Data Requirements
Let's move on to 11.9.3. 11.9.3 states, the reliability data used when quantifying the effect of random failures shall be credible, traceable, documented, justified, and shall be based on field feedback from similar devices used in a similar operating environment. Okay.
So this is, it's kind of loaded, a statement. Basically, your numbers that you're using in SIL verification calculations should be realistic. Traceable is key. Documented is key. Credible is key. Justified. All of these are really important. Now, how do you comply with this clause? Well, when you run your SIL verification calculations, you should always have a source for where that data came from.
Whether that data came from a corporate standard, whether it came from a data bank like NPRD or RITA, you need traceability as to where that data came from, and someone needs to provide some sort of justification to say that that data is acceptable in your calculation. Now, the easiest, easiest, easiest way to do this is to use Kenexis Vertigo software. Or, you know, there are other software vendors out there where there is a database available to you, and the numbers in that database are traceable, credible, documented, and justified, and shall be based on field feedback.
Especially us at Kenexis, we are constantly tracking the failure rates of our vendors that they are achieving in their, in the actual performance of their systems. We're constantly reviewing that and updating our generic data, and then also, you know, things that show up in the database in Vertigo will be reviewed. So, even if we have a third-party certificate, we don't always publish it in our database because a lot of times we don't trust or believe the person that issued the certificate. We don't believe that they use the right techniques.
We don't believe that they use credible data, and we will throw it out. So, everything that is in the Kenexis database has been curated, and it's going to meet all the requirements of 1193.
But you, the end user, still need to make sure that the number that you pick out of the database is appropriate for the specific device that you're looking at, number one. And number two, it's incumbent upon you to always track the performance of those devices and compare them against what's in our database to make sure that they're credible.
Maybe you have a high failure rate application that you need to adjust those failure rates for whatever reason. Okay, so make sure you know where your data came from and that it's legit and that in your reporting, you have a traceability back to where the data came from.
## Notes on Reliability Data Sources and Vendor Data
Now, there are three notes here in Clause 1193. Note one states, this includes user collected data, vendor provider user data derived from data collected on devices, data from general field feedback reliability databases. In some cases, engineering judgment can be used to assess missing reliability data or evaluate the impact on reliability data collected in a different operating environment.
So, note one says, you know, this is data driven. This is a data driven standard. We want to make sure that we have data that is going to be collected by end users, collected by general field feedback databases.
And the second part, the second sentence in there is interesting and noteworthy in that it says that sometimes you're going to need to use engineering judgment to look at your data because the data set is not rich enough. So, while you might have a lot of data on how often a device fails, you might not have enough actual failures to be able to give a good breakdown between what percent of the failures are safe and what percent of the failures are dangerous.
Because failures don't happen that frequently and getting a statistically valid number of failures to be able to do a safe dangerous split is much, much, much, much, much, much harder than collecting data to determine what the overall failure rate is. So, in one case you need a lot of experience. In the other case you need a lot of failures which hopefully you're not getting.
So, things like safe dangerous splits you might want to add additional judgment, you might want to look at vendor data and general reliability databases in addition to the failure rate that you're collecting from your own plant.
Alright, note two states, the lack of reliability data reflective of the operating environment is a recurrent shortcoming of probabilistic calculations. operations. So, what that first sentence is saying is that in a lot of cases you might not have enough data in the plant to, or you might have data for a device used in something like purchased natural gas service, but now you're applying it into a process service.
So, making sure that you have data that's relevant to your operating environment is something that's difficult to do. It's just a note saying, yeah, something to think about, keep in the back of your mind that the data that you're using might not reflect the installation and that's whether it's the service, the process service, or the ambient environment around the device.
Okay, the second sentence in note 2 says, end users can organize relevant device reliability data collections in accordance with IEC 60300 Part 3 or ISO 14224 to improve the implementation of the IEC 61511 series. So, specifically ISO 14224 is basically industry's attempt to standardize failure modes and the way that you collect data, the way you assign failure modes to data. If you are familiar with the Offshore Reliability Databank or RITA, the failure modes that you see in or RITA are going to be consistent with ISO 14224. So, that's kind of an example of what you can look at.
So, we're just trying to standardize what those failure modes are.
Note 3 states that vendor data based on returns can be restricted to a population where there is full knowledge of the operational environment and fully recorded in accordance with IEC 60300 or ISO 1425. That's the first sentence. So, basically, vendors have the ability to restrict the population to where there is full knowledge.
So, here's the rub with vendor data and why a lot of the times it's hopelessly optimistic and not representative of reality.
So, let's say an equipment vendor makes a thousand devices and they say, okay, I made a thousand devices and one year later only one of those devices failed. Gee, isn't my failure rate great? Well, not necessarily. Of the thousand devices that they made, there's a good chance that 500 of them are still sitting in the factory or a distributor shelf and it hasn't been installed yet. So, we really shouldn't be taking credit for those. The ones that have been sold might be sitting in inventory at a process plant. We shouldn't be counting those.
When a failure occurs, if it's a low cost device, there's a very good chance that the device will just get thrown in the trash and its failure never reported back to the vendor. So, we shouldn't be counting those. And even if it is reported back to the vendor, the vendor might say, oh, well, you used it outside of the operating environment that we specify, so we're not going to take credit for that failure because you misused it. So, we're not going to count those failures.
So, the number of devices that you can really thoroughly watch and by that it means that we're looking at the quality of the data or that we know what the service is is a very small fraction. So, when you're looking at vendor return data as a source of your failure rates, well, don't. Don't. Just don't because it's not good unless the vendor went to a lot of trouble to track how many devices are actually in service, when were they put in service, and there are very few vendors that do that with a high enough degree of rigor that you can use it in your SIL verification calculations.
All right,
1193 Note 3 is, I'll go ahead and read that note. Oh, wait, no, I did just read the note. That was the first sentence of Note 3. The second sentence of Note 3 says, the user can also record the operational environment for the SIF and be able to demonstrate that the vendor's operational environment data matches the environment of the SIF. So, kind of the second sentence there is, if I'm getting vendor data, I want to know what applications that that data is based on and make sure that my data is in a consistent application with that.
Okay, so,
1193, a little bit of traceability, validity, documentation of the numbers that you're using in your SIL verification calculations. If you're using really good sophisticated software, like Vertigo from Kenexis, that'll be kind of built into the system for you.
## Clause 11.9.4 Reliability Data Uncertainty
Okay, 1194 is one sentence, and there's more notes than there are requirements, but the requirement for 1194 states, the reliability data uncertainties shall be assessed and taken into account when calculating the failure measure.
Shall. So, we need to consider the reliability data uncertainties. Well, we already talked about this in clause 11.4, I believe clause 1149. Now,
I'm going to need to back up in the standard and double check. But the official way to consider the uncertainty of your failure rate data is clause 11.4.9, which again states, reliability data used in the calculation of the failure measure shall be determined by an upper bound statistical confidence limit of no less than 70%. So, you can't just divide number of failures by time because that's going to give you a 50% confidence. 50% of the time your actual failure rate is going to be higher, 50% of the time it's going to be lower.
We want to make sure that only 30% of the time is the failure rate higher and 70% of the time the failure rate is lower and we do that by a statistical technique called the chi-squared test, which if you want to learn all about that, that is in I would recommend the Smith book, Reliability, Maintainability, and Risk from David Smith, or you can always take the ISA EC54 Advanced SIL verification to learn how to do that technique.
So, that is the official required method for assessing the uncertainty of your failure rates. There are some notes here, so let's go ahead and hit those notes.
Note one says the reliability data uncertainties can be evaluated according to the amount of field feedback and or exercise of expert judgment. And there's also a parenthetical in there that says less field feedback results in more uncertainty, uncertainty, which is going to be definitely the case when you do the chi-squared test mathematics. You'll see that the less data accumulation you have, the more uncertainty you're going to have, the higher that failure rate that you're going to use in your calculations ends up being.
Okay, the second sentence in note one states, published standards, such as IEC 60604, part four, Bayesian approaches, engineering judgment techniques, et cetera, can be used to estimate the reliability data uncertainties.
Okay, note two, states, give some techniques, note two, the following techniques can be used for calculating failure measures. More information can be found in IEC 61511 part two, in the 11.9.4 portion of part two. Again, part two is that the informative guidance that's associated with the standard. And two sub-bullet points under note two, the first one says, use an upper bound confidence of 70% for each input reliability parameter instead of its mean or 50% confidence in order to obtain conservative point estimates of the failure measures. confidence.
Um, so this is a bullet point in a note that says to use 70% confidence, but if you go back up to clause 1149, it is an actual requirement. It's not a something to think about, something to consider. It's an actual requirement. So, uh, kind of a little bit of inconsistency in the way that the standard presents it. Here in 1194, it seems like it's a suggestion, but if you go back to clause 1149, it is not a suggestion. It is a requirement.
Uh, the second bullet point on note two states that the use of probabilistic distribution functions of input reliability parameters. So, use probabilistic distribution functions of reliability parameters, and then perform Monte Carlo simulations to obtain a histogram representing the distribution of the failure measure and assess a conservative value from this distribution.
Okay, I have seen exactly zero people do Monte Carlo distributions in their SIL verification calculations. Exactly zero. It's a whole lot of work. Now, some software tools have Monte Carlo distribution techniques built into them where you can run this, but you would need a failure rate distribution of your input data to calculate a failure rate distribution of your output data, and I can provide you with exactly zero sources of failure rate distributions for inputs to your SIL verification calculations. So, yep, it's something you could do. Never seen it done. Don't plan on doing it ever.
I might, you know, just for edification purposes, but not for a real live project.
Basically, when it comes to uncertainty, 70% confidence in your failure rates is what the standard requires, and that goes back to Clause 1149.
## Clause 11.9.5 Improving SIF Design to Meet SIL
Okay, last item in Clause 11.9 is 1195, which states, if for a particular design, the target failure measure for the relevant SIF is not achieved. So, Clause
1195 is technically, it's not a requirement. it's some guidance, some rules of thumb, some suggestions on how to achieve your SIL target if your first calculation failed. So, if your proposed design fails, how do you change your design in order to achieve the SIL target? This is in no way, shape, or form, or requirement. I don't know why the standard lists it as a requirement, but I digress.
So, there are four bullet points for the things that you can do if you don't achieve the SIL target.
Number one, A, identify the devices or parameters contributing most to the failure measure with a note that says fault tree cut set analysis can be useful here.
Okay, so most, so if you go into Arbor, if you run your SIL verification calculation, you're going to get minimal cut sets, and you'll know which cut sets have the most importance because they contribute the most to the PFD, and you'll know what events are associated with those cut sets.
But let's get real here. Most people are not going to use fault trees or Markov models in their SIL verification calculations. They're going to use CAN software tools. And those CAN software tools, like Vertigo from Kenexis, is going to calculate little pie charts, little percentages that will tell you for each subsystem, sensor, logic solver, final element, to what degree, what percentage of the overall PFD is that subsystem contributing? contributing? What's the percent contribution? And you're generally going
*[inaudible]*
reduce the PFD of the item that is contributing most heavily to the PFD. Of course, there are cost implications and all kinds of other things that real people in the real world think about to try to achieve the SIL target at a lower price. But we generally want to focus on the thing that is the biggest problem.
Item B says evaluate the effect of possible improvement measures on the identified devices or parameters.
For example, more reliable devices, so things that have a lower failure rate. Additional defenses against common mode failures, so if your common cause is the biggest contributor contributor to not achieving the PFD. Think about things that can reduce your common cause failure and change the design.
Next item would be increased diagnostics or proof test coverage. That allows us to detect a failure early and repair out so you're unavailable for a shorter period of time. So increased diagnostics, increased redundancy, so go for a simplex device to a 2 out of 3 vote, 1 out of 2 vote. Reduce the proof test intervals and stagger tests.
That's an interesting one. It's on the list, so if you have a 1 out of 2 vote and your test interval is one year, if you test one device at time zero and another device at time equals 6 months, that will give you a different PFD than if you test both the devices at the same time. Believe me now and hear me later. You can run some calculations. Bust out your Excel spreadsheet to prove that to yourself.
Now, do I recommend doing that? Absolutely not. Now you're getting into human factors. You're making your plant so difficult to maintain that you've lost any numerical benefit. So that's one of those things that only a PhD number cruncher who has no idea about what it takes to actually run a plant would recommend.
But mathematically speaking, it will actually work.
Okay, item C is select and implement improvement measures to establish a new result. So A, identify the problem. B, muck around with the parameters. And then C, determine which of those parameters you want to mess with. And D, compare the new result to the target failure measure and repeat steps A through D until the target failure measure is achieved in a conservative manner.
Okay, so that's Clause 1195, which shouldn't be a clause, honestly. It should be in Part 2 because it is literally a suggestion and there's no way to actually enforce this requirement. So 1195, you can kind of toss that onto the pile of suggestions and informative guidance because no one is going to verify how you achieved the SIL target if you didn't achieve it with the preliminary design. So there's nothing you can audit against to verify, to prove that you actually met Clause 1195.
But the guidance is correct. Find your biggest problem and figure out what parameter you can diddle with on your biggest problem and then make that change, rerun it, and see if you meet your SIL target.
And that's it. That is Clause 1195.
## Summary and Clause 12 Software Preview
And with this episode and the previous episode, I told you absolutely everything that the standard says about SIL verification, which ain't a whole lot, which is why there are so many books that talk about SIL verification calculations, so many webinars, so many training classes on how to do SIL verification calculations. Because we didn't talk about, for instance, when do you use availability? When do you use reliability? What are these different fault propagation models that you can use? What are their strengths and limitations?
There's a whole lot to it, but the standard just doesn't get into it. It says you need to consider these DU, DD, DN, common cause, test intervals, proof test coverage, diagnostic coverage, diagnostic intervals, everything that we talked about in Clause 1192 are the things that you need to consider, and then you calculate your failure measure compared against your target. If your calculated is better than your target, you are good to go, and you can move on to the balance of the safety life cycle.
All right, so with that, at this stage in the game, we are all the way through Clause 11, which is hands down my favorite clause. It's all kind of the prescriptive part of the SIS process. Got into a lot of detail, a lot of good stuff there.
Now, in next week's episode, we are going to start in the Clause 12, and Clause 12 is all about software, and software is the neglected stepchild of the SIS design, and maybe for good reason. A lot of the stuff in Clause 12 is over the top.
It assumes that you're going to be doing something crazy like writing your code in JavaScript, and it's a little bit overblown. But we're going to go through it.
We're going to kind of go through one clause at a time, and I will give you my opinions, kind of beginning with the underpinning that most people are going to take the cause and effect diagrams and turn the cause and effect diagrams into function block diagrams using their equipment vendors' maintenance and engineering interface software without a whole series of specifications for what the software is supposed to do, a bunch of V diagrams, a multi-step staged approach to software development.
*[inaudible]*
no, that's not what I see. Most of the time I see a programmer being given a cause and effect diagram and they go, is that good? Is that bad? Well, I guess you're going to have to tune in next week when we start talking about software and clause 12.
## Kenexis Vertigo Software Overview
Now that you've heard some insights on technical safety, functional safety, and the IEC 61511 standard, let me tell you a little bit more about how to easily and effectively implement the safety lifecycle using the Kenexis integrated safety suite and our SIS safety lifecycle management tool Vertigo. Vertigo is a comprehensive tool set for performing assessment calculations, documenting, and maintaining the design of safety instrumented systems.
Analysis begins with importing or synchronizing a list of safety instrumented functions with their definitions and associated performance targets from our open PHA tool for HAZOP and LOPA documentation.
Each safety function can then be analyzed by performing a SIL verification calculation, complete with a collection of tools for optimizing designs and a database of thousands of potential instruments to define failure rates and diagnostic coverage capabilities.
After the SIL verification calculations are defined, you can build an SRS by automatically generating a cause and effect diagram from the SIF definitions and other defined instruments. each SIS instrument will include a customizable data sheet and general requirements that are applicable to the SIS as a whole and can be entered individually or even bulk imported from customizable libraries.
After the design phase, you can even use Vertigo to track and document testing throughout the entire life of the facility.
Kenexis Vertigo is the most integrated, easy-to-use enterprise tool for allowing the development of SIS design basis information more efficiently and effectively than any other software application.