Kenexis Functional Safety Podcast

A bypass left open too long quietly erodes the very risk reduction the SIS was built to deliver. In this episode, Ed Marszal tackles the final three sub-clauses of 11.8, beginning with why maximum bypass duration must be grounded in the mean repair time or online test interval already baked into SIL verification calculations. He walks through a complete Intelligent MOC workflow for authorizing, tracking, and closing out bypasses with documented compensating measures. The episode closes with a pragmatic look at forcing inputs and outputs in safety PLCs—once anathema, now accepted for maintenance only under strict procedural and access-security controls. For engineers wrestling with the gap between standard text and daily practice, this is the connective tissue that makes compliance operable.

A device should never remain in bypass indefinitely. Bypass duration must be time-limited, and appropriate compensating measures must be implemented.

Listen in for more information and thoughts on this important topic as sections 11.8.4 to 11.8.6 are discussed in more detail.

Tune in to the latest episode of the Kenexis Functional Safety Podcast, hosted by Ed Marszal, President and CEO of Kenexis. Now available on Spotify and Apple Podcasts, Ed offers his expert insights on the IEC 61511 standard.

With decades of experience in safety instrumented systems and as a Principal Engineer, Ed has a unique perspective to offer. He has been an active contributor to the ISA 84 committee since 1994, adding to his deep understanding of the field.

In this inaugural season, Ed delves into the IEC 61511 standard, unpacking the meaning behind each word and providing a thorough interpretation of its application. Through personal stories from his career and committee work, he offers valuable context and insights for professionals in the industry.

Full Episode Transcript

KENEXIS FUNCTIONAL SAFETY PODCAST — S1E45 TRANSCRIPT (Markdown)
Cleaned & reflowed for web publication and AI crawlability.

The JSON-LD block below is schema.org structured data. If your CMS lets you
add raw HTML to a post, paste it into the page (or anywhere in the
body — crawlers read it either way). Fill in PLACEHOLDER_EPISODE_PAGE_URL
once the post exists. Everything from the "# Kenexis Functional Safety
Podcast…" heading down is the transcript body — paste it into your post.
–>

"`html

"`

# Kenexis Functional Safety Podcast — Season 1, Episode 45: IEC 61511, Clause 11.8.4 to 11.8.6 (Bypass Time Limits, Compensating Measures, and Forcing)

## Episode Teaser: Bypass Time Limits

You cannot leave a device in bypass forever. You need to limit the time duration for your bypass and provide compensating measures too.

## Podcast Introduction and Disclaimer

Welcome to the Kenexis Functional Safety Podcast. I'm your host, Ed Marszal, President and CEO of Kenexis. Kenexis is a technical safety consultancy that helps chemical process industry companies to analyze risk and design engineered safeguards like safety instrumented systems and fire and gas detection systems. Kenexis also provides the industry-leading suite of software tools, including our best-in-class Vertigo software for SIS Safety Lifecycle Management.

In this first season of the podcast, we are going to focus on the IEC 61511 standard, doing a deep dive into the standard, including more depth of information on what the standard means and how to apply it, brought to life with personal war stories and behind-the-scenes discussions of the committee members as we develop the standard in ISA 84 and IEC SC 65.

Before we start, a little disclaimer. I will be providing my opinion on technical and engineering topics. This information is provided on a best-effort basis and is of a general nature. The information presented in this podcast might not be applicable to your specific application. It is the obligation of every engineer to thoroughly analyze any system that they are designing and not blindly rely on any general advice presented in this podcast. All right.

## Episode Overview and Prior Testing Recap

We are following up on Clause 11.8. I thought I would be able to get through it in one episode, but an hour later and we only got through Clause 11.8.3. There's a lot to discuss about testing. I'm still going to say that testing is probably the thing that I worry the most about in SIS design. I mean, it's rare, almost non-existent to find an accident that happened even though there was a well-designed, well-maintained simplex system. It's always improper design, improper testing. Anyway, I digress a little bit.

We talked about last week the fact that you could do your testing either end-to-end or in segments, in parts. You could test your sensors today, logic solvers tomorrow, and your final elements some other time beyond that. And then we also spent a lot of time talking about online proof testing, which is 11.8.2 and 11.8.3. Usually when you're doing online proof testing, you are doing it in parts because the plant is online and running. Now, some of these are partial tests. In a lot of cases, they are complete tests. So when you're testing a sensor online, it's generally a complete test.

And even as we discussed last week, when you perform an online test of a valve, it's possible to do a complete test of that valve while the plant is online.

In addition to what you would do while the plant is online and running, so we talked about test facilities a lot. And the test facilities are the equipment and the features of the process that are going to allow you to do this test while the plant is online and running. Very important stuff. But that's not it. Because we need to limit the amount of time that things are in bypass. In general, but for Clause 11.8.4, we're specifically talking about testing. And actually, it does mention repair.

So that whole repair part maybe should be moved out to Clause 16 that talks about operation and maintenance. But, you know, with respect to the design process, making sure that you have the systems in place are going to be important.

So time duration, that's going to be 11.4. Then we're going to get to Clause 11.8.5, which talks a little bit more about compensating measures. Now, I did talk about compensating measures in Clause 11.3, which is the first place that this topic came up. We're going to revisit some of those issues. We're going to talk about different best practices for what compensating measures are and how to implement them. Maybe even using intelligent MOC and OpenPHA in combination with the checklisting functionality and the authorization functionality that's provided inside those tools.

That reminds me, I need to set up a webinar where I talk all about how to use management of change tools to properly document, properly authorize, properly manage defeat of safety critical equipment, also known as bypassing. So that's in Clause 11.8.5. And then 11.8.6 is kind of a fun topic to talk about. Well, it was a lot more controversial back in the olden days when I first started programming PLCs, but it's kind of a non-event at this point in time. And that's going to be forcing of inputs and outputs.

So using a program, programmatically, using the maintenance and engineering interface to hold a logical value in its position regardless of what the inputs and outputs and the process logic might dictate that should happen when that occurs.

So we'll talk about when you're allowed to use forcing, when you're not allowed to use forcing. You know, back in the day, it was absolutely verboten. Today, it's kind of a non-event. But we talked about forcing a couple, three episodes ago when we were talking about the maintenance and engineering interface and, you know, the fact that you really shouldn't abuse it. And I don't want to belabor it more than I already did. I think in that episode of the podcast, I went into a very long history of forces. But I'll kind of recap some of the high points on forcing of inputs and outputs.

## Clause 11.8.4: Maximum Bypass Time Duration

All right, so now that we reviewed where we're at, let's get back into where we're going with this topic, which is clauses 1184, 1185, and 1186. Let's start with clause 1184. One simple sentence. It states, the maximum time the SIS is allowed to be in bypass, and then parenthetically, for repair or testing. So right there, it's not just a limitation on testing. It's also a limitation on repair. Okay, the maximum time the SIS shall be allowed into bypass while safe operation of the process is continued shall be defined. So, three items that we need to think about separately.

The first item is there is a maximum amount of time I'm allowed to be in bypass. The second item is that safe operation needs to be maintained during that time duration. And then the unspoken part of that sentence is, what if I get to the end of that maximum time duration and I'm not done and I need to leave it into bypass? What do I do now? All right, so let's talk about all these issues.

So why does there need to be a maximum time duration? Well, we assumed in our calculations, actually, let me step back even further than that. Management has defined a tolerable level of risk. And in our design process, we are doing everything in our power under our control to make sure that that value for tolerability of risk is never exceeded.

And that value for tolerability of risks is going to work its way into our calculation in the form of the mean repair time, the MRT. Formerly called MTTR, mean time to repair. We don't do that anymore because MTTR, the term, was abused so often that nobody knows really what it means. Does it mean time to restore? Is it mean time to repair? Do I mean just the amount of time to actually fix the device? Does that include the duration between when the device failed and when I know that it failed? That's the realization time.

In the current version of IEC 61511, MTTR is mean time to restore, which is going to include the mean repair time plus the time to detect. So how much time elapsed between when the failure occurred and when I know that the failure has occurred and I can trigger the repair process. So MRT is the term that we use for the time that it takes to perform the repair. So we definitely want to repair within the mean repair time.

Now, actually, there's another term that's used in the standard that is more appropriate, which is the MPRT, or the maximum permitted repair time. Now, 99 times out of 100, when we perform a calculation or we talk about mean repair time, we don't actually mean the mean repair time. We actually mean the maximum permitted repair time. So when I put into my calculation 72 hours, I don't actually expect that my repair is going to take exactly 72 hours.

Normally, that 72 hours, giving you a little historical reference, means that if something fails at 5.30 p.m. on a Friday, we can wait for day shift on Monday to fix it. That's where the 72 hours came from. God's honest truth, I swear. That's where it came from. So we're expecting the actual duration of the repair to be less than eight hours because day shift, or 12 hours, depending on how you do your math, day shift on Monday is going to fix the device. So it doesn't need to be fixed over the weekend.

So that's the MRT that we're using in our calculations, which technically speaking is not the mean repair time, but the maximum permitted repair time. That's what's used in our calculations. And we use that duration in our calculations to ensure that tolerable risk has been achieved.

So we shouldn't be violating that in our testing. Also, in our testing or our repairs. Furthermore, if you are doing online testing, you should be calculating an online testing term. So if you go into a really good SIL verification software tool like Vertigo from Kenexis, part of the Kenexis integrated safety suite, you will see that there is a checkbox that you would check if you're going to perform online testing. And if you're going to perform online testing, there is a time duration.

And what our assumption is, is that for the duration of the test, that subsystem will be unavailable to perform its safety function. And that unavailability is going to contribute to the overall PFD of the safety instrumented function. Therefore, when you're performing a test, you should not be in bypass longer than the number that you put in for the test duration in the online testing term.

So if you're using the greatest SIL verification software out there, Vertigo, then your maximum duration should either be the MRT or MPRT in reality for a repair and the test duration for a test. Those should be the limits of how long you are in bypass. They should be defined and they will be defined in your calculations. Just go to the calculation for the device you're testing and those numbers will be in that detailed SIL verification report that you're going to print out of Vertigo. Okay, so those numbers are already there. They're already documented.

## Bypass Authorization and Intelligent MOC Workflow

But when you're putting something into bypass for a test, you need to follow your bypass authorization procedure. And my recommendation for you is to follow a bypass authorization procedure that includes a documentation step and an authorization step. Now, in our Vertigo tool, we do have a light bypass authorization form to help you document what is in bypass. It is light though. The authorizations are a little bit on the simplified side and the reporting is a little bit on the simplified side.

What I would recommend that you do instead is do this in your management of change software. So management of change software, I highly, highly recommend Intelligent MOC, of course, from Kenexis, which is part of the Kenexis integrated safety suite of software. So anytime you do a management of change and putting something into bypass is a change that you should document through a management of change process. It is, specifically speaking, it is a temporary change. Okay?

So you go into Intelligent MOC, you create a new MOC ticket, you check that it is a temporary bypass, you put the description of what you're going to be doing, the time duration, all that stuff in the request for change, and then that is going to trigger a workflow.

So you're going to need an Intelligent MOC to build out a bypass workflow. That bypass workflow is going to include a risk analysis. So risk analysis can be as simple as incorporating a checklist. So the way Intelligent MOC works is that you can have a checklist in OpenPHA that you can attach to your MOC ticket for performing your bypass risk assessment with any form of PHA. So if you wanted to do a HAZOP, you could do that. Just do the HAZOP in the MOC or a what if in, I'm sorry, in OpenPHA.

Once you've done that, you can attach it to your MOC file so that all things are synchronized and linked automatically.

The next thing you're going to want to do is you're going to want to put together another checklist that is your compensating measures. And that compensating measures is going to include things like who needs to be informed, who is going to replace the device that's out of service, what set point, if you will, do they need to act on, when that safe operation limit is violated, what action does that person take? Is it automatic? Do they call to the control room? Do they move a manual valve?

And then considerations for whether or not there's enough time to respond to that alarm is the thing that you're going to want to document.

So after you've done your rationale, your bypass risk assessment, your compensating measures checklist, then you're going to need to get a series of approvals. So you're going to need to possibly get an approval from instrumentation and control, an approval from operations, an approval from management. And after all those people have checked all of the information, they will do the approval, you can activate the bypass.

And then after you're done with the bypass, you would close it out in the MOC system. And until you close it out, the MOC system is going to flag that that bypass is currently present, whether it's authorized and so on. So you're going to be able to go into intelligent MOC, you're going to be able to sort for bypass being the MOC type, and see what bypasses are active in the MOC system. So elegantly tracking your bypasses is a super critical thing.

We do have a light version of this directly in our Vertigo software, but the more elegant system that includes multiple authorizations, tracks the authorizations, provides email notices to a series of people saying, hey, you're beyond your test interval and you haven't closed out your bypass yet. Things like that are going to be more elegantly performed in a management of change system. And hey, with Kenexis integrated safety suite, that management of change system is extremely cost effective. Look into it.

And also, keep your eye out. I think maybe February, maybe March, I'm probably going to want to do a full webinar on doing bypass management using a management of change system and the Kenexis intelligent MOC software. Okay, so that's 1184. Maximum time the SIS is allowed to be in bypass for either repair testing while safe operation of the process is continued shall be defined. It will show up in your MOC documentation.

It will show up in your SIL verification documentation as either the bypass duration or the maximum permitted repair time, often referred to as the mean repair time, somewhat incorrectly.

## Clause 11.8.5: Compensating Measures

All right, next clause up is 1185. This says, compensating measures that ensure continued safe operation shall be provided in accordance with 11.3 when the SIS is in bypass and parenthetically repair or testing. So I need these compensating measures when the device is being repaired. I also need these compensating measures when the device is being tested. Basically, I need compensating measures anytime the device is not in service.

We had a long discussion of compensating measures when we talked about 11.3 several weeks ago.

Now, you could go ahead and review that webinar for some of these details. But we like to put together that alternate protection plan. And again, if you go into Open PHA, you should be able to find a checklist in the, or a form would be a better way to describe it, in the, in the Kenexis library for doing the compensating measures. So, take a, take, take a look at that.

Some of the things that you're going to want to consider in the compensating measures. Number one, do you have a device that does the same functionality that's not in bypass? That's a great one. But if you don't have any redundancy, generally, you need to replace that device with a human and you're going to need to give them instructions. instructions.

And those instructions are going to need to include things like what variable are they looking at? What measurement are they looking at? What is the activation point? What action do they need to take? And then, ultimately, the big deal is, is it reasonable to expect a human being to be able to do this considering the process safety time? All right.

## Clause 11.8.6: Forcing of Inputs and Outputs

Last item in this clause, which I, you know, now that we're, we're only about halfway through the standard now. That's not true. We're definitely more than halfway through the standard. And a lot of these topics I've talked about several times now. So I hate to keep repeating myself, but, you know, then again, this is kind of a grab bag. So anybody listening to this podcast might not listen from beginning all the way through to the end. They might just kind of look for the clause that they're interested in, which is why I do repeat myself a little bit on these podcasts.

Hopefully, that's not infuriating to the loyal followers who listen to every podcast segment immediately when it was released. But the forcing of inputs and outputs is something that I talked about when I talked about the maintenance and engineering interface, which is the location where you actually perform the forcing.

So what is forcing? So if I have my maintenance and engineering interface computer that's running the software that I use to program the PLC, I can generally run that software in an online mode where I can see the status of everything that's going on in the plant. I know the input signals that are going in. I know the output signals that are going out. I know the state of all the logic. And in that paradigm, I can click on a variable and I can force it.

So when I force it, what that means is that it's going to ignore all of the logic that is trying to set the value of that variable and simply give it a number and that's what the number is. So I can take a level transmitter and instead of using the A to D conversion, I could just set that transmitter's value to 50% and it will stay at 50% regardless of what the A to D converter is trying to do.

Or I can set the output of a two out of three vote to true, meaning that it will never shut down regardless of what the two out of three inputs say to do and so on. So that's what I mean by forcing.

## Forcing Rules: Clause Text and Permitted Uses

Now, let's read the clause three sentences and discuss what the concerns with forcing are and what the rules are related to forcing in relation to tests. It says, forcing of inputs and outputs in PESIS, so Programmable Electronic Safety Instrumented Systems, shall not be used, period. No, of course not. A lot of people wanted a period there, but it's not. So forcing shall not be used as part of application programs, operating procedures, and maintenance. Period? No, no, no, no, no, not a period. Parenthetical, except as noted below.

So, the standard is saying that you shouldn't use forcing as part of the application program, which again, that also violates the rule for using your maintenance and engineering interface as an operator interface. So you shouldn't have to force something to make the process go. So if I'm in batch step number two, I don't want to force a value inside the PLC to get the batch number three. We should write better software than that. or have manual switches that are wired into the PLC instead of forcing. So we don't want to use forces as part of the application program.

We don't want to use them as part of operating procedures, and we don't want to use them as part of maintenance either.

Now, what's a little bit tricky here is the parenthetical except as noted below. The way I read the standard, it says that the except as noted below means that you can only use the forcing for maintenance, and you can't use forcing for application program and operating procedures, which kind of is consistent with the fact that you're not allowed to use your maintenance and engineering interface as your operator interface.

So really, the first clause says no, no, no forcing for application program, no forcing for operating procedures. yes, you can use forcing of inputs and outputs for maintenance, but under these two sentences that limit what you're allowed to do.

Okay, so the second two sentences, forcing of inputs and outputs without taking the SIS out of service shall not be allowed unless supplemented by procedures and access security. So, if you are going to force anything in the PLC, you can't just do it ad hoc. You need to be following a procedure that tells you what to force, how to force it, when to force it, how long it's going to be in force, and when you're going to take it out of force and how you're going to take it out of force.

So, you need to be following a procedure and also the forcing needs to be secured in terms of access, access security. I have talked several times now about that access security. So, key cards to get into the plant, key cards to get into the room, Azure AD passwords to get onto the computer, another password to get into the application program that is the maintenance and engineering interface. Okay? So, we got security, but we need to be following procedures. So, it's not a, well, I need to test this transmitter.

So, I got a work order to test the transmitter and I'm going to go into the safety PLC and put stuff into force, not following procedures. This is no good. We need to be following a procedure. Okay, second sentence.

And again, these are the exceptions that allow forcing of inputs and outputs only during maintenance. And, I guess testing would be considered part of maintenance. Okay, the second sentence does any such forcing shall be announced or set off an alarm as appropriate. So, what you would, a best practice here, which is very reasonable and very doable, is when you put something into force, that's going to trigger a value in your safety PLC that should be communicated into your alarm system and activate an enunciated alarm that will forever more be shown in your alarm log.

So, when you force something, it's going to show up in the alarm log so that when I do my bypass management, when I do my auditing and tracking of my bypassing system, I will know even if somebody forced something that it was put into bypass and I'm going to want to go into my intelligent MOC system and confirm that there is a bypass authorization ticket that allowed that to happen. setting off an alarm would be great but you don't have to. You're going to announce it or set off an alarm as appropriate.

So, how do I announce it? Okay, well, this can be as loosey-goosey as just kind of going into the control room and saying, hey, Mrs. Operator, I'm going to put this sensor into bypass. and if she says it's cool, you can go ahead and do it. Now, let's be a little bit more reasonable and a little bit more realistic. If you do that, there should have been a work order and that work order should have been authorized by the operator. So, this is not that loosey-goosey in reality if you're following, and there's no reason for you not to be following a formal permit-to-work system to do all of this.

Okay, so that is kind of the big picture on forcing.

## Episode Summary and Next Episode Preview

So, a little bit of an abbreviated episode today, only a little bit over a half hour. We rounded out the discussion of testing, and specifically we rounded out the discussion of Clause 11.8, Maintenance or Testing Design Requirements. Last week we talked about end-to-end or in parts. We talked about, especially if you're doing it in parts and online design of all the equipment that you're going to need to allow that bypass to occur.

Then, in today's episode, we talked about limiting time durations, using management of change to document the time duration, the authorization, and being able to track what's in bypass, what's not in bypass, and also the documentation and approval of the compensating measures that you're going to use while that bypass is in place. We rounded out the discussion with our third or fourth, at least third, maybe fourth discussion of forcing inputs and outputs in the PLC.

When you can do it, when you can't do it, and if you do do it, set off an alarm or announce it, preferably through a permit-to-work system where there is an official authorization from the operator.

That's all I got for this week. Next week, we're going to get into clause 11.9. Clause 11.9 is quantification of random failure. Now, I don't know how far we're going to get into it.

We're probably only going to talk about clause 11.9.1 and 11.9.2. But those are some big dogs because I'm going to need to discuss what are all the things that the standard requires you to consider when you're doing your SIL verification calculations, number one, and then another very long discussion of what's going to change in the next version of the standard because pretty much universally the standards committee doesn't like a lot of the statements that are currently contained in this section.

So I'll talk a little bit about how we've gotten our scissors out and we're going to cut things that are redundant or just don't make any sense. And then also, I will go into another preach mode discussion of how SIL verification calculations are for random hardware failures. Do not include probability of human failure or frequency of human failure in these calculations because not only do they not benefit you, they will cause you to do the wrong thing and think you improved the situation when you actually made it worse. We'll talk to you about this again next week.

## Vertigo Software and KISS Suite Advertisement

Now that you've heard some insights on technical safety, functional safety, and the IEC 61511 standard, let me tell you a little bit more about how to easily and effectively implement the safety lifecycle using the Kenexis integrated safety suite and our SIS safety lifecycle management tool Vertigo. Vertigo is a comprehensive tool set for performing assessment calculations, documenting, and maintaining the design of safety instrumented systems.

Analysis begins with importing or synchronizing a list of safety instrumented functions with their definitions and associated performance targets from our open PHA tool for HAZOP and LOPA documentation.

Each safety function can then be analyzed by performing a SIL verification calculation, complete with a collection of tools for optimizing designs and a database of thousands of potential instruments to define failure rates and diagnostic coverage capabilities. After the SIL verification calculations are defined, you can build an SRS by automatically generating a cause and effect diagram from the SIF definitions and other defined instruments.

Each SIS instrument will include a customizable data sheet and general requirements that are applicable to the SIS as a whole and can be entered individually or even bulk imported from customizable libraries. After the design phase, you can even use Vertigo to track and document testing throughout the entire life of the facility.

Kenexis Vertigo is the most integrated, easy-to-use enterprise tool for allowing the development of SIS design basis information more efficiently and effectively than any other software application.