Butterworth-Heinemann Linacre House, Jordan Hill,Oxford OX2 8DP 225 Wildwood Avenue, Woburn, MA 01801-2041 A division of Reed Educational and Professional Publishing ltd
-&A member of the Reed Elsevier pie group OXFORD
BOSTON
MELBOURNE
JOHANNESBURG
NEW DELHI
SINGAPORE
First published 1991 Reprinted 1992, 1993,1994 (twice). 1995,1996 Second edition 1997 Reprinted 1997 John Moubray 1991, 1997 All rights reserved. No part of this publication may be reproduced in any material fOIDl (including photocopying or storing in any medium by electronic means and whether or not transiently or incidentally to some other use of this publication) without the written permission of the copyright holder except in accordance with the provisions of the Copyright, Designs and Patents Act 1988 or under the terms of a licence issued by theCopyright Licensing Agency ltd, 90 Tottenham Court Road, London, England WIP 9HE. Applications for the copyright holder's written permission to reproduce any part of this publication should be addressed to the publishers
British Library Cataloguing in Publication Data A catalogue record for this book is available from the British Library ISBN 0 7506 3358 I
Typeset by the author Printed and bound in Great Britain by Biddies ltd, Guildford and King's Lynn
For Edith
Contents
/ Prefaee Acknowledgements
1
Introduction to Reliability-centred Maintenance
1
1.1
The changing world of maintenance
1
1.2
Maintenance and RCM
6
1.3
RCM: The seven basic questions
1.4
Applying the RCM process
16
1.5
What RCM achieves
18
7
2
Functions
21
2.1
Describing functions
22
2.2
Performance standards
22
2.3
The operating context
28
2.4
Different types of functions
35
2.5
How functions should be listed
44
3
Functional Failures
45
3.1
Failure
45
3.2
Functional failures
46
4
Failure Modes and Effects Analysis
53
4.1
What is a failure mode?
53
4.2
Why analyse failure modes?
55
4.3
Categories of failure modes
58
4.4
How much detail?
64
4.5
Failure effects
73
4.6
Sources of information about modes and effects
77
4.7
Levels of analysis and the information worksheet
80
5
:Failure Consequences
90
5.1
Technically feasible and worth doing
90
5.2
Hidden and evident functions
92
5.3
Safety and environmental consequences
94
viii
Contents
Reliability-centred Maintenance
IX
5.4
Operational consequences
103
11
Implementing ReM Recommendations
212
5.5
Non-operational consequences
108
ILl
Implementation - the key steps
212
5.6
Hidden failure consequences
III
11.2 The RCM audit
Conclusion
127
214
5.7
11.3 Task descriptions
218
6
Proactive Maintenance 1: Preventive Tasks
129
6.1
Technical feasibility and proactive tasks
129
6.2
Age and deterioration
130
6.3
Age-related failures and preventive maintenance
133
11.4 Implementing once-off changes
220
11.5 Work packages
221
11.6 Maintenance planning and control systems
224
11.7 Reporting defects
233 235
6.4
Scheduled restoration tasks
134
12
6.5
Scheduled discard tasks
137
12.1 The six failure patterns
235
6.6
Failures which are not age-related
140
12.2 Technical history data
250
7
Proactive Maintenance 2: Predictive Tasks
144
13
7.1
Potential failures and on-condition maintenance
144
13.1 Who knows?
7.2
The P-F interval
145
13.2 RCM review groups
266
7.3
The technical feasibility of on-condition tasks
149
13.3 Facilitators
269
7.4
Categories of on··condition techniques
149
13.4 Implementation strategies
277
7.5
On-condition tasks: some of the pitfalls
155
13.5 RCM in perpetuity
284
7.6
Linear and non-linear P-F curves
157
13.6 How RCM should not be applied
286
7.7
How to determine the P-F interval
163
13.7 Building skills in ReM
291
7.8
When on-condition tasks are wOlth doing
166
7.9
Selecting proactive tasks
167
14
14.1 Measuring maintenance performance
292
8
Default Actions 1: Failure-finding
170
14.2 Maintenance effectiveness
293
8.1
Default actions
170
14.3 Maintenance efficiency
304
8.2
Pailure-finding
171
14.4 What RCM achieves
307
15
318
Actuarial Analysis and Failure Data
Applying the ReM Process
What ReM Achieves
261 261
292
8.3
Failure-finding task intervals
175
8.5
The technically feasibility of failure-finding
185
9
Other Default Actions
187
15.2 RCM in other sectors
321
9.1
No scheduled maintenance
187
15.3 Why RCM 2?
323
9.2
Redesign
188
9.3
Walk-around checks
197
10
The ReM Decision Diagram
A Brief History of ReM
15.1 The experience of the airlines
318
Appendix 1: Asset hierarchies and functional block diagrams
327
Appendix 2: Human error
335
198
Appendix 3: A continuum of risk
343
10.1 Integrating consequences and tasks
198
Appendix 4: Condition monitoring
348
10.2 The RCM decision process
198
Bibliography
412
10.3 Completing the decision worksheet
209
Index
415
10.4 Computers and RCM
211
Preface
Humanity continues to depend to an ever-increasing extent on the wealth generated by highlY mechanised and automated businesses. We also depend more and more on services such as the unintermpted supply of electricity or trains which nm on time. More than ever, these depend in turn on the continued integrity of physical assets. Yet when these assets fail, not only is this wealth eroded and not only are these services intermpted, but our very survival is threatened. Equip ment failure has played a part in some of the worst accidents and environ mental incidents in industrial history - incidents which have become by words, such as Amoco Cadiz, Chernobyl, Bhopal and Piper Alpha. As a result, the processes by which these failures occur and what must be done to manage them are rapidly becoming very high priorities indeed, especi ally as it becomes steadily more apparent just how many of these failures are caused by the very activities which are supposed to prevent them. The first industry to confront these issues was the international civil aviation industry. On the basis of research which challenges many of our most firmly and widely-held beliefs about maintenance, this industry evolved a completely new strategic framework for ensuring that any asset continues to perform as its users want it to perform. This framework is known within the aviation industry as MSG3, and outside it as Reliabil ity-centred Maintenance, or RCM. Reliability-centred Maintenance was developed over a period of thirty years. One of the principal milestones in its development was a report commissioned by the United States Department of Defense from United Airlines and prepared by Stanley Nowlan and the late Howard Heap in 1978. The report provided a comprehensive description of the develop ment and application of RCM by the civil aviation industry. It forms the basis of both editions of this book and of much of the work done in this field outside the airline industry in the last fifteen years. Since the early 1980's, the author and his associates have helped com panies to apply RCM in hundreds of industrial locations around the world work which led to the development of RCM2 for industries other than aviation in 1990.
xii
Reliability-centred Maintenance
The first edition of this book (published in the UKin 1991 and the USA in 1992) provided a comprehensive introduction to RCM2. Since then, the RCM philosophy has continued to evolve, to the extent that it became necessary to revise the first edition to incorporate the new developments. Several new chapters have been added, while others have been revised and extended. Foremost among the changes are: •
a more comprehensive review of the role of functional analysis and the definition of failed states in Chapters 2 and 3
•
a much broader and deeper look at failure modes and effects analysis
Preface
The book is intended for maintenance, production and operations managers who wish to learn what RCM is, what it achieves and how it is applied. It will also provide students on business or management studies
courses with a comprehensive introduction to the formulation of strate gies for the management of physical (as opposed to financial) assets. Finally, the book will be invaluable for any students of any branch of engineering who seek a thorough understanding of the state-of-the-art in
modern maintenance. It is designed to be read at three levels: •
•
•
new material on how to establish acceptable levels of risk in Chapter 5
the addition of more rigorous approaches to the determination of failure finding task intervals in Chapter 8
•
more about the implementation of RCM recommendations in Chapter 1 1, with extra emphasis on the ReM auditing process
•
more information on how RCM should - and should not - be applied in Chapter 13, including a more comprehensive look at the role of the RCM facilitator
•
new material on the measurement of the overall performance of the maintenance function in Chapter 14
•
Chapters 2 to 10 describe the main elements of the technology of RCM, and will be of most value to those who seek no more than a reasonable
technical grasp of the subject.
and Appendix 3 •
Chapter 1 is written for those who only wish to review the key elements of Reliability-centred Maintenance.
in the context of RCM, with special emphasis on the question of levels of analysis and the degree of detail required in Chapter 4
xiii
•
The remaining chapters are for those who wish to explore ReM in more detail. Chapter 1 1 provide a brief summary of the key steps which must
be taken to implement the recommendations arising from RCM analy ses. Chapter 12 takes an in-depth look at the sometimes contentious subject of the relationship between age and failure. Chapter 13 consid ers how RCM should be applied, with emphasis on thc role of the people
involved. After reviewing ways in which maintenance effectiveness and efficiency should be measured, Chapter 14 describes what ReM
achieves. Chapter 15 provides a brief history of RCM.
a brief review of asset hierarchies in Appendix 1, together with a sum mary of the (often overstated) role played by functional hierarchies and JOHN MOUB RAY
functional block diagrams in the application of RCM •
a review of different types of human error in Appendix 2, together with a look at the part they play in the failure of physical assets
•
the addition of no fewer than 50 new techniques to the appendix on condition monitoring (now Appendix 4).
In the second impression of the second edition, the word 'tolerable' has been substituted for 'acceptable' in discussions about risk in Chapters 5 and 8 and in Appendix 3, in order to align this book more with standard terminology used in the world of risk. It also includes further material on the practicality of failure-finding task intervals in Chapter 8, and slightly revised material on RCM implementation strategies in Chapter 13.
Lutterworth Leicestershire Serptember 1997
f1cknowledgelllents
Acknowledgements
It has only been possible to write both editions of this book with the help of a great many people around the world. In particular, I would like to record my continuing gratitude to every one of the hundreds of people with whom I have been privileged to work over last ten years, each of whom has contributed something to the material in these pages. In addition, I would like to pay special tribute to a number of people who played a major role in helping to develop and refine the RCM philosophy to the point discussed in this edition of this book. Firstly, special thanks are due to the late Stan Nowlan for laying the foundations for both editions of this book so thoroughly, both through his own writings and in person, and to all his colleagues in the civil aviation industry for their pioneering work in this field. Special thanks are also due to Dr Mark Horton, for his help in develop ing many of the concepts e mbodied in Chapters 5 and 8, and to Peter Stock for researching and helping to co-author Appendix 4. I a m also indebted to all members of the Aladon network for their help in applying the concepts and for their continuous feedback about what works and what doesn't work, much of which is also reflected in these pages. Foremost among these are my colleagues Joel Black, Chris James, Hugh Colman and Ian Hipkin, and my associates Alan Katchmar, Sandy Dunn, Tony Geraghty, Frat A marra, Phil Clarke, Michael Hawdon, Brian Oxenham, Ray Peden, Simon Deakin and Theuns Koekemoer . A mong the many clients who have proved and are continuing to prove that RCM is a viable force in industry, I am especially indebted to the following: Gino Palarchio and Ron Thomas of Dofasco Steel Mike Hopcraft, Terry Belton and Barry Camina of Ford of Europe Joe Campbell of the British Steel Corporation Vincent Ryan and Frank O'Connor of the Irish Electricity Supply Board Francis Cheng of Hong Kong Electric Bill Seeland of the New Venture Gear Company Denis Udy, Roger Crouch, Kevin Weedon and Malcolm RegIer of the Royal Navy
xv
Don Turner and Trevor Ferrer of China Light & Power John Pearce of the Mars Corporation Dick Pettigrew of Rohm & Haas Pat McRory of BP Exploration Al Weber and Jerry Haggerty of Eastman Kodak Derek Burley of Opal Engineering . The roles played by Don Humphrey, Richard Hall, Brian Davies, Tom Edwards, David Willson and the late Joe Versteeg in helping to develop or to propagate the concepts discussed in this book are also acknowl edged with gratitude. Finally, a special word o f thanks to my family for creatinG an environ b ment in which it was possible to write both editions of this book, and to Aladon Ltd for permission to reproduce the RCM Information and Deci sion Worksheets and the RCM 2 Decision Diagram.
1
Introduction to Reliability-centred Maintenance
1.1 The Changing World of Maintenance Over the past twenty years, maintenance has changed, perhaps more so than any other management discipline. The changes are due to a huge increase in the number and variety o f physical assets (plant, equipment and buildings) which must be maintained throughout the world, much more complex designs, new maintenance techniques and changing views on maintenance organisation and responsibilities. Maintenance is also responding to changing expectations. These include a rapidly growing awareness o f the extent to which equipment failure a ffects safety and the environment, a growing awareness o f the connec tion between maintenance and product quality, and increasing pressure to achieve high plant availability and to contain costs. The changes are testing attitudes and skills in all branches o f industry to the limit. Maintenance people are having to adopt completely new ways o f thinking and acting, as engineers and as managers. At the same time the limitations o f maintenance systems are becoming increasingly apparent, no matter how much they are computerised. In the face o f this avalanche o f change, managers everywhere are looking for a new approach to maintenance. They want to avoid the false starts and dead ends which always accompany major upheavals. instead they seek a strategic framework which synthesises the new developments into a coherent pattern, so that they can evaluate them sensibly and apply those likely to be of most value to them and their companies. This book describes a philosophy which provides just such a frame work. It is called Reliability-centred Maintenance, or RCM. If it is applied correctly, RCM transfor ms the relationships between the undertakings which use it, their existing physical assets amI the people who operate and maintain those assets. It also enables new assets tn be put into effective service with great speed, confidence and precision. This chapter provides a brief introduction to ReM, stal1ing with a look at how maintenance has evolved over the past fifty years.
2
Introduction to Reliability-centred Maintenance
Reliability-centred Maintenance
Since the 1930's, the evolution of maintenance can be traced through three generations. ReM is rapidly becoming a cornerstone of the Third Generation, but this generation can only be viewed in perspective in the light of the First and Second Generations . The First Generation
The First Generation covers the period up to World War II. In those days industry was not very highly mechanised, so downtime did not matter much. This meant that the prevention of equipment failure was not a very high priority in the minds of most managers. At the same time, most equipment was simple and much of it was over-designed. This made it reliable and easy to repair. As a result, there was no need for systematic maintenance of any sort beyond simple cleaning, servicing and lubrica tion routines. The need for skills was also lower than it is today. The Second Generation
Things changed dramatically during World War II. Wartime pressures increased the demand for goods of all kinds while the supply of industrial manpower dropped sharply. This led to increased mechanisation . By the 1950's machines of all types were more numerous and more complex. Industry was beginning to depend on them. As this dependence grew, downtime came into sharper focus. This led to the idea that equipment failures could and should be prevented, which led in tum to the concept of preventive maintenance. In the 1960's, this consisted mainly of equipment overhauls done at fixed intervals. The cost of maintenance also started to rise sharply relative to other operating costs. This led to the growth of maintenance planning and control systems. These have helped greatly to bring maintenance under control, �nd are now an established part of the practice of maintenance. Finally, the amount of capital tied up in fixed assets together with a sharp increase in the cost of that capital led people to start seeking ways in which they could maximise the life of the assets. The Third Generation
Since the mid-seventies, the process of change in industry has gathered even greater momentum. The changes can be classified under the head ings of new expectations, new research and new techniques.
3
New expectations Figure 1 . 1 shows how expectations of maintenance have evolved. Third Generation:
Figure 1.1 Growing expectations of m aintenance
•
�------l ......J "'-P;-jr,-s-t - - - - - - -· G e n er, a tlon: •
Fix it when it broke
1 940
1 950
Second Gen13ration:
•
• •
•
Greater
•
Better product quality
•
Higher plant availability
Longer equipment life
Lower costs
1 960
1 970
Higher plant availability and reliability
No damage to the environment
•
longer equipment life
•
Greater cost effectiveness
1 980
1 990
2000
Downtime has always affected the productive capability of physical assets by reducing output, increasing operating costs and interfering with customer service. By the 1960's and 1970's, this was already a major concern in the mining, manufacturing and transport sectors. In manufac turing, the effects of downtime are being aggravated by the worldwide move towards just-in-time systems, where reduced stocks of work-in progress mean that quite small breakdowns are now much more likely to stop a whole plant. In recent times, the growth of mechanisation and automation has meant that reliability and availahility have now also become key issues in sectors as diverse as health care, data processing, telecommunications and building management. Greater automation also means that more and more failures affect our ability to sustain satisfactory quality standards. This applies as much to standards of service as it does to product quality. Forinstance, equipment failures can affect climate control in buildings and the punctuality of transport networks as much as they can inteIi"ere with the consistent achievement of specified tolerances in manufacturing. More and more failures have serious safe ty or environmental conse quences, at a time when standards in these areas are rising rapidly. In some parts of the world, the point is approaching where organisations either conform to society'S safety and environmental expectations, or they cease to operate. This adds an order of magnitude to our dependence on the integrity of our physical assets one which goes beyond cost and which becomes a simple matter of organisational survival.
Introduction to Reliability-centred Maintenance
Reliability-centred Maintenance
4
At the same time as our dependence on physical assets is growing, so too is their cost - to operate and to own. To secure the maximum return on the investment which they represent, they must be kept working effici ently for as long as we want them to. Finally, the cost of maintenance itself is still rising, in absolute terms and as a proportion of total expenditure. In some industries, it is now the second highest or even the highest element of operating costs. As a result, in only thirty years it has moved from almost nowhere to the top of the league as a cost control priority. New research Quite apart from greater expectations, new research is changing many of our most basic beliefs about age and failure. In particular, it is apparent that there is less and less connection between the operating age of most assets and how likely they are to fail. Figure 1.2 shows how the earliest view Third Generation of failure was simply that as things got older, they were more likely to fail. A growing awareness of 'infant mortality' led to widespread Second Generation belief in the 'bathtub' curve. Figure 1.2: Changing views on equipment failure
1 940
1 950
1 960
1 970
1 980
1 990
2000
However, Third Generation research has revealed that not one or two but six failure patterns actually occur in practice. This is discussed in detail later, but it too is having a profound effect on maintenance.
Figure 1.3 shows how the classical emphasis on overhauls and administra tive systems has grown to include many new developments in a number of dif· ferent fields .
r----------1 Second Generation: •
First Generation: •
1 940
•
Scheduled overhauls Systems for planning and controlling work Big, slow computers
1 950 Figure 1.3:
1 960
1 970
Third Generation: Con dition monitoring
•
•
Design for reliability and main tain abi lity
•
Hazard studies
•
S m al l , fast
•
Fail ure modes and effects a nalyses
•
•
M ultiskillin g and teamwork
1 980
1 990
2000
Changing maintenance techniques
The new developments include: •
•
•
•
decision support tools, such as hazard studies, failure modes and effects analyses and expert systems new maintenance techniques, such as condition monitorinu "" designing equipment with a much greater emphasis on reliability and maintainability a major sh(ft in organisational thinking towards participation, teamworking and flexibility.
A major challenge facing maintenance people nowadays is not only to learn what these techniques are, but to decide which are worthwhile and which are not in their own organisaiions. If we make the right choices, it is possible to improve asset performance and at the same time contain and even reduce the cost of maintenance. If we make the wrong choices, new problems are created while existing problems only get worse. The challenges facing maintenance In a nutshell, the key challenges facing modern maintenance managers can be summarised as follows: to select the most appropriate techniques to deal with each type of failure process in order to fulfil all the expectations of the owners of the assets, the users of the assets and of society as a whole in the most cost-effective and enduring fashion with the active support and co-operation of all the people involved . •
•
New techniques There has been explosive growth in new maintenance concepts and tech niques. Hundreds have been developed over the past fifteen years, and more are emerging every week .
Fix it when it broke
•
5
•
•
•
6
Reliability-centred Maintenance
Introduction to Reliability-centred Maintenance
ReM provides a framework which enables users to respond to these chal lenges, quickly and simply. It does so because it never loses sight of the fact that maintenance is about physical assets. Uthese assets did not exist, the maintenance function itself would not exist. So ReM starts with a comprehensive, zero-based review of the maintenance requirements of each asset in its operating context. All too often, these requirements are taken for granted. This results in the development of organisation structures, the deployment of resources and the implementation of systems on the basis of incomplete or incorrect assumptions about the real needs of the assets. On the other hand, if these requirements are defined correctly in the light of modern thinking, it is possible to achieve quite remarkable step changes in maintenance effi ciency and effectiveness. The rest of this chapter introduces ReM in more detail. It begins by exploring the meaning of 'maintenance' itself. lt goes on to define ReM and to describe the seven key steps involved in applying this process.
1.2 Maintenance and RCM From the engineering viewpoint, there are two elements to the manage ment of any physical asset. It must be maintained and from time to time it may also need to be modified. The major dictionaries define maintain as cause to continue (Oxford) or keep in an existing state ( Webster). This suggests that maintenance means preserving something . On the other hand, they agree that to modify something means to change it in some way. This distinction between maintain and modify has profound implications which are discussed at length in later chapters. However, we focus on maintenance at this point. When we set out to maintain something, what is it that we wish to cause to continue? What is the existing state that we wish to preserve? The answer to these questions can be found in the fact that every phys ical asset is put into service because someone wants it to do something. In other words, they expect it to fulfil a specific function or functions. So it follows that when we maintain an asset, the state we wish to preserve must be one in which it continues to do whatever its users want it to do.
Maintenance: Ensuring that physical assets continue to do what their users want them to do
7
What the users want will depend on exactly where and how the asset is being used (the operating context). This leads to the following formal de finition of Reliability-centred Maintenance:
Reliability-centred Maintenance: a process used to determine the maintenance requirements of any physical asset in its operating context In the light of the earlier definition of maintenance, a fuller definition of ReM could be 'a process used to determine what must be done to ensure that any physical asset continues to do whatever its lIsers want it to do in its present operating context' .
1.3 RCM: The seven basic questions The ReM process entails asking seven questions about the asset or sys tem under review, as follows: •
what are the functions and associated peljormallce standards of the asset in its present operating context?
•
in what ways does it fail to fulfil its functions?
•
what causes each functional failure?
•
what happens when each failure occurs?
•
in what way does each failure matter?
•
what can be done to predict or prevent each failure?
•
what should be done if a suitable proactive task cannot be found?
These questions are introduced brietly in the following paragraphs, and then considered in detail in Chapters 2 to 10. Functions and Performance Standards
Before it is possible to apply a process used to determine what must be done to ensure that any physical asset continues to do whatever its users want it to do in its present operating context, we need to do two things: •
determine what its users want it to do
•
ensure that it is capable of doing what its users want to start with.
8
Reliability-centred Maintenance
Introduction to Reliability-centred Maintenance
This is why the first step in the RCM process is to define the functions of each asset in its operating context, together with the associated desircd standards of performance. What users expect assets to be able to do can be split into two categories: •
•
In addition to the total inability to function, this definition encompas ses partial failures, where the asset still functions but at an unacceptable level of performance (including situations where the asset cannot sustain acceptable levels of quality or accuracy). Clearly these can only be identified after the functions and performance standards of the asset have been defined. Functional failures are discussed at greater length in Chapler 3.
primary functions, which summarise why the asset was acquired in the first place. This category of functions covers issues such as speed, out put, carrying or storage capacity, product quality and customer service. secondaryfunctions, which recognise that every asset is expected to do more than simply fulfil its primary functions. Users also have expecta tions in areas such as safety, control, containment, comfort, structural integrity, economy, protection, efficiency of operation, compliance with environmcntal regulations and even the appearance of the asset,
The users of the assets are usually in the best position by far to know exactly what contribution each asset makes to the physical and financial well-being of the organisation as a whole, sO it is essential that they are involved in the RCM process from the outset. Done properly, this step alone usually takes up about a third of the time involved in an entire RCM analysis. It also usually causes the team doing the analysis to learn a remarkable a mount often a frightening a mount about how the equipment actually works. Functions are explored in more detail in Chapter 2. Functional Failures
The objectives of maintenance are defined by the functions and associ ated performance expectations of the asset under consideration. But how does maintenance achieve these objectives? The only occurrence which is likely to stop any asset performing to the standard required by its users is some kind of failure. This suggests that maintenance achieves its objectives by adopting a suitable approach to the management of failure. However, before we can apply a suitable blend of failure management tools, we need to identify what failures can occur. The RCM process does this at two levels:
Failure Modes
As mentioned in the previous paragraph, once each functional failure has been identified, the next step is to try to identify all the events which are reasonably likely to cause each failed state. These events are known as failure modes. 'Reasonably likely' failure modes include those which have occurred on the same or similar equipment operating in the same context, failures which are currently being prevented by existing main tenance regimes, and failures which have not happened yet but which are considered to be real possibilities in the context in question. Most traditional lists of failure modes incorporate failures caused by deterioration or normal wear and tear. However, the list should include failures caused by human errors (on the part of operators and maintainers) and design flaws so that all reasonably likely causes of equipment failure can be identified and dealt with approptiately. It is also i mportant to identify the cause of each failure in enough detail to ensure that time and effort are not wasted trying to treat ;;ymptoms instead of causes. On the other hand, it is equally important to ensure that time is not wasted on the analysis itself by going into too much detail. Failure Effects
The fourth step in the RCM process entails listing failure effects, which describe what happens when each failure mode occurs . These descrip tions should include all the information needed to support the evaluation of the consequences of the failure, such as:
•
firstly, by identifying what circumstances a mount to a failed state
•
•
then by asking what events can cause the asset to get into a failed state.
•
In the world of R CM, failed states are known asfunctional failures be cause they occur when an asset is unable tofulfil afunction to a standard ofperformance which is acceptable to the user.
9
what evidence (if any) that the failure has occurred in what ways (if any) it poses a threat to safety or the environment
•
in what ways (if any) it affects production or operations
•
what physical da mage (if any) is caused by the failure
•
what must be done to repair the failure.
10
Failure modes and effects are discussed at greater length in Chapter 4. The process ofident�fvingfunctions,functionalfailures,failure modes andfailure effects yields surprising and often very exciting opportunities for improving performance and safety, and also for eliminating waste Failure Consequences
A detailed analysis of an average industrial undertaking is likely to yield between three and ten thousand possible failure modes. Each of these failures affects the organisation in some way, but in each case, the effects are different. They may affect operations. They may also affect product quality, customer service, safety or the environment. They will all take time and cost money to repair. It is these consequences which most strongly influence the extent to which we try to prevent each failure. In other words, if a failure has seri ous consequences, we are likely to go to great lengths to try to avoid it. On the other hand, if it has little or no effect, then we may decide to do no routine maintenancc beyond basic cleaning and lubrication. A great strength of RCM is that it recognises that the consequences of failures are far more important than their technical characteristics. In fact, it recognises that the only reason for doing any kind of proactive main tenance is not to avoid failures per se, but to avoid or at least to reduce the consequences of failure. The RCM process classifies these consequences into four groups, as follows: •
•
•
•
Introduction to Reliability-ceillred Maintenance
Reliability-centred Maintenance
Hidden failure consequences: Hidden failures have no direct impact, but they expose the organisation to multiple failures with serious, often catastrophic, consequences. (Most of these failures are associated with protective devices which are not fail-safe.) Safety and environmental consequences: A failure has safety conse quences if it could hurt or kill someone. It has environmental conse quences if it could lead to a breach of any corporate, regional, national or international environmental standard.
We will see later how the RCM process uses these categories as the basis of a strategic framework for maintenance decision-making. By forcing a stmctured review of the consequences of each failure mode in terms of the above categories, it integrates the operational, environmental and safety objectives of the maintenance function. This helps to bring safety and the environment into the mainstream of maintenance management. The consequence evaluation process also shifts emphasis away from the idea that allfailures are bad and must be prevented. In so doing, it focuses attention on the maintenance activities which have most effect 011 the per formance of the organisation, and diverts energy away from those which have little or no effect. It also encourages us to think more broadly about different ways of managing failure, rather than to concentrate only on failure prevention. Failure management techniques m'e divided into two categories: •
•
proactive tasks: these are tasks undertaken before a failure occurs, in order to prevent the item from getting into a failed state. They embrace what is traditionally known as 'predictive' and 'preventive' maintenance, although we will see later that RCM uses the terms scheduled restora tion, scheduled discard and on-condition maintenance default actions: these deal with the failed state, and are chosen when it is not possible to identify an effective proactive task. Defaultactions in clude failurefinding, redeSign and run-to-failure.
The consequence evaluation process is discussed again brie fly later in this chapter, and in much more detail in Chapter 5. The next section of this chapter looks at proactive tasks in more detail Proactive Tasks
Many people still believe that the best way to optimise plant availability is to do some kind of proactive maintenance on a routine basis. Second Generation wisdom suggested that this should consist of overhauls or componentreplacements at fixed intervals. Figure 1.4 illustrates the fixed interval view of failure.
Operational consequences: A failure has operational consequences if it affects production (output, product quality, customer service Of oper ating costs in addition to the direct cost of repair) Non-operational consequences: Evident failures which fall into this category affect neither safety nor production, so they involve only the direct cost of repair .
11
t
ro >c::t= (]) 0== :.-
Figure 1.4:
The traditional view of failure
+-------.-
"
Efa� "O..oal
00.. 0 § e::
LI FE
" _____
t=�===========1i1l1 Age ------
12
Reliability-centred Maintenance
Figure 1.4 is based on the assumption that most items operate reliably for a period 'X' , and then wear out. Classical thinking suggests that extensive records about failure will enable us to detennine this life and so make plans to take preventive action shortly before the item is due to fail in future. This model is true for certain types of simple equipment, and for some complex items with dominant failure modes. In particular,wear-out char acteristics are often found where equipment comes into direct contact with the product. Age-related failures are also often associated with fatigue, corrosion, abrasion and evaporation. However, equipment in general is far more complex than it was twenty years ago. This has led to startling changes in the patterns of failure, as shown in Figure 1.5 . The graphs show conditional probability of failure against operating age for a variety of electrical and mechanical items. Pattern A is the well-known bathtub curve. It begins with a high incidence of failure (known as infant mortality) followed by a constant or gradually increasing conditional probability of failure, then by a wear-out zone. Pattern B shows constant or slowly increasing conditional prob ability of failure, ending in a wear-out zone (the same as Figure 1.4).
o
Figure 1.5:
Six patterns of failure
F
Introduction to Reliability-centred Maintenance
13
Pattern C shows slowly increasing conditional probability of failure, but there is no identifiable wear-out age . Pattern D shows low conditional probability of failure when the item is new or just out of the shop, then a rapid increase to a constant level, while pattern E shows a constant con ditional probability of failure at all ages (random failure) . Pattern F starts with high infant mortality, which drops eventually to a constant or very slowly increasing conditional probability of failure. Studies done on civil aircraft showed that 4% of the items conformed to pattern A, 2% to B, 5% to C, 7% to D, 14o/t.l to E and no fewer than 68% to patte rn F. (The number of times these patterns occur in aircraft is not necessarily the same as in industry. But there is no doubt that as assets become more complex, we see more and more of patterns E and F .) These findings contradict the belief that there is always a connection between reliability and operating age . This belief led to the idea that the more often an item is overhauled, the less likely it is to fail . Nowadays, this is seldom true. Unless there is a dominant age-related failure mode, age limits do little or nothing to improve the reliability of complex items. In fact scheduled overhauls can actually increase overall failure rates by introducing infant mortality into otherwise stable systems . An awareness of these facts has led some organisations to abandon the idea of proactive maintenance altogether. In fact, this can be the right thing to do for failures with minor consequences. But when the failure consequenees are significant, something must be done to prevent or pre dict the failures, or at least to reduc.e the consequences. This brings us back to the question of proactive tasks. As mentioned earlier, RCM divides proactive tasks into three categories, as follows: • scheduled restoration tasks • scheduled discard tasks • scheduled on-condition tasks. Scheduled restoration and scheduled discard tasks Scheduled restoration entails remanufacturing a component or overhaul ing an assembly at or before a specified age limit, regardless of its con dition at the time . Similarly, scheduled discard entails disearding an item at or before a speeified life limit, regardless of its condition at the time. Collectively, these two types of tasks are now generally known as pre ventive maintenance. They used to be by far the most widely used form of proactive maintenance. However for the reasons discussed above, they are much less widely used than they were twenty years ago.
14
On-condition tash The continuing need to prevent certain types of failure, and the growing inability of classical techniques to do so, are behind the growth of new types of failure management. The majority of these techniques rely on the fact that most failures give some warning of the fact that they are about to occur. These warnings are known as potential failures, and are defined as identifiable physical conditions which indicate that afunctionalfail ure is about to occur or is in the process of occurring. The new techniques are used to detect potential failures so that action can be taken to avoid the consequences which could occur if they degen erate intQ functional failures. They are called on-condition tasks because items are left in service on the condition that they continue to meet desired performance standards. (On-condition maintenance includes predictive maintenance, condition-based maintenance and condition monitoring.) Used appropriately, on-condition tasks are a very good way of managing failures, but they can also be an expensive waste of time. RCM enables decisions in this area to be made with particular confidence. Default Actions
Whether or not a proactive task is technically feasible is governed by the technical characteristics of the task and of the failure which it is meant to prevent. Whether it is worth doing is governed by how well it deals with the consequences of the failure. If a proactive task cannot be found which is both technically feasible and worth doing, then suitable default action must be taken. The essence of the task selection process is as follows: •
•
•
•
•
failure-finding: Failure-finding tasks entail checking hidden functions periodically to determine whether they have failed (whereas condition based tasks entail checking if something is failing). redesign: redesign entails making any one-off change to the built-in capability of a system. This includes modifications to the hardware and also covers once-off changes to procedures. no scheduled maintenance: as the name implies, this default entails mak ing no effort to anticipate or prevent failure modes to which it is applied, and so those failures are simply allowed to occur and then repaired. This default is also called run-to-failure,
The ReM Task Selection Process A great strength ofRCM is the way it provides simple, precise and easily understood criteria for deciding which (if any) of the proactive tasks is technicallyfeasible in any context, and if so for deciding how often they should be done and who should do them. These criteria are discussed in more detail in Chapters 6 and 7 .
for hiddenfailures, a proactive task is worth doing if it reduces the risk of the multiple failure associated with that function to an acceptably low level. If such a task cannot be found then a scheduled failure finding task must be performed. If a suitable failure-finding task cannot be found, then the secondary default decision is that the item may have to be re designed (depending on the consequences of the multiple failure) . -
RCM recognises three major categories of default actions, as follows: •
15
Introduction to Reliability-centred Maintenance
Reliability-centred Maintenance
•
for failures with safety or environmental consequences, a proactive task is only worth doing if it reduces the risk of that failure on its own to a very low level indeed, if it does not eliminate it altogether. If a task can not be found which reduces the risk of the failure to an acceptably low level, the item must be redesigned or the process must be changed. if the failure has operational consequences, a proactive task is only worth doing if the total cost of doing it over a period oftime is less than the cost of the operational consequences and the cost of repair over the same period. In other words, the task must be justified 011 economic grounds. If it is not justified, the initial default decision is no scheduled maintenance. (If this occurs and the operational consequences are still unacceptable then the secondary default decision is again redesign). if a failure has non-operational consequences a proactive task is only worth doing if the cost of the task over a period of time is less than the cost of repair over the same period. So these tasks must also be justified on economic grounds. If it is not justified, the initial default decision is again no scheduled maintenance, and if the repair costs are too high, the secondary default decision is once again redesign .
This approach means that proactive tasks are only specified for failures which really need them, which in turn leads to substantial reductions in routine workloads . Less routine work also means that the remaining tasks are more likely to be done properly. This together with the elimination of counterproductive tasks leads to more effective maintenance.
16
Reliability-centred Maintenance
Compare this with the traditional approach to the development of maintenance policies. Traditionally, the maintenance requirements of each asset are assessed in terms of its real or assumed technical charactelistics, without considering the consequences of failure. The resulting schedules are used for all similar assets, again without considering that different consequences apply in different operating contexts. This results in large numbers of schedules which are wasted, not because they are 'wrong' in the technical sense, but because they achieve nothing. Note also that the R CM process considers the maintenance require ments of each asset before asking whether it is necessary to reconsider the design. This is simply because the maintenance engineer who is on duty today has to maintain the equipment as it exists today, not what should be there or what might be there at some stage in the future.
1.4 Applying the ReM process Before setting out to analyse the maintenance requirements of the assets in any organisation, we need to know what these assets are and to decide which of them are to be subjected to the RCM review process. This means that a plant register must be prepared if one does not exist already. In fact, the vast majority of industrial organisations nowadays already possess plant registers which are adequate for this purpose, so this book only touches on the most desirable attributes of such registers in Appendix 1. Planning If it is correctly applied, R CM leads to remarkable improvements in main tenance effectiveness, and often does so surprisingly quickly. However, the successful application of R CM depends on meticulous planning and preparation. The key elements of the planning process are as follows: •
• •
•
decide which assets arc most likely to benefit from the ReM process, and if so, exactly how they will benefit assess the resources required to apply the process to the selected assets in cases where the likely benefits justify the investment, decide in detail who is to perform and who is to audit each analysis, when and where, and arrange f()r them to receive appropriate training ensure that the operating context of the asset is clearly understood.
Introduction to Reliability-centred Maintenance
17
Review groups We have seen how the ReM process embodies seven basic questions. In practice, maintenance people simply cannot answer all these questions on their own. This is because many (if not most) of the answers can only be supplied by production or operations people. This applies especially to questions concerning functions, desired performance, failure effects and failure consequences. For this reason, a review of the maintenance requirements of any asset should be done by small teams which include at least one person from the maintenance function and one from the operations function. The senior ity of the group members is less important than the fact that they should have a thorough knowledge of the asset under review. Each group mem ber should also have been trained in R CM. The make-up of a typical RCM review group is shown in Figure 1.6: The use of these groups not Facilitator only enables management to gain access to the Engineering Operations knowledge and Supervisor Supervisor expertise of each member of the group on a systematic basis, Craftsman but the members (M and/or Operator themselves gain a greatly enhanced under standing of the asset in External Specialist (if needed) (Technical or Process) its operating context. Figure 1.6: A typical ReM review group
Facilitators ReM review groups work under the guidance of highly trained special ists in ReM, known as facilitators. The facilitators are the most impor tant people in the ReM review process. Their role is to ensure that: •
• •
the ReM analysis is carried out at the right level, that system bounda ries are clearly defined, that no important items are overlooked and that the results of the analysis are properly recorded RCM is correctly understood and applied by the group members the group reaches consensus in a brisk and orderly fashion, while retain ing the enthusiasm and commitment of individual members
18
Introduction to Reliability-centred Maintenance
Reliability-centred Maintenance
finishes on time. the analysis progresses reasonably quickly and or sponsors to ensure Facilitators also work with RCM project managers . es appropnate managethat each analysis is properly planned and receiv ' . rial and logistic support. . . m more detatl m Facili tators and R CM review groups are dIscussed
•
•
19
Greater safety and environmental integrity: ReM considers the safety and environmental i mplications of every failure mode before consider ing its effect on operations. This means that steps are taken to mini mise all identifiable equipment-related safety and environ mental hazards, if not eliminate them altogether. By integrating safety into the mainstream of maintenance decision-making, RCM also i mproves atti tudes to safety.
Chapter 13. •
The outcomes of an ReM analysis . ReM analysls results If it is applied in the manner suggested above, an le outco mes, as follows: 1· n three tanaib b to be done by the maintenance depar t ment • maintenance schedules dures for the operators of the asset • revised operating proce ffchanges must be made to the design of the • a list of areas where one-o situations where the . asset or the way in which it is operated to deal with t configuratIon. asset canno t deliver the desired performance in its curren the process learn a great Two less tangible outco mes are that participants in on better as teams. deal about how the asset works, and also tend to functi . Auditing and implementation for each as �et, semor I mmediately after the review has been completed I�ent must satIsf � them managers with overall responsibility for the equip ble and defensIble. senSI are selves that decisions made by the group . are I mple �nented ations mend After each review is approved, the recom main tenance plann �ng and by incorporating maintenance schedules into dure c ?anges mto the control systems, by incorporating operating proce by hand :ng recom men standard opera ting procedures for the asset, and n authonty. Key aspec ts dations for design changes to the appropriate desig in Chapter 1 1. of auditing and implementation are discussed
•
1.5 What ReM Achieves should ?nly be seen as Desirable as they are, the outcomes listed above the mamtena ?ce func e . a means to an end. Specifically, they should enabl e 1. 1 at the begmnmg of tion to fulfil all the expec tations listed in Figur the following paragraphs, this chap ter . How they do so is summarised in 14. and discussed again in more detail in Chapter
•
•
Improved operating performance (output, product quality and custo mer service): ReM recognises that all types of maintenance have some value, and provides rules for deciding which is most suitable in every situation. By doing so, it helps ensure that only the most effective forms of maintenance are chosen for each asset, and that suitable action is taken in cases where maintenance cannot help. This much more tightly focused maintenance effort leads to qnantum jumps in the performance of existing assets where these arc sought. ReM was developed to help airlines draw up maintenance programs for new types of aircraft before they en ter service . As a result, it is an ideal way to develop such programs for new assets, especially complex equipment for which no historical information is available. This saves much of the trial and error which is so often part of the development of new maintenance programs trial which is ti me-consuming and frus trating, and error which can be very costly. Greater maintenance cost-effectiveness: ReM continually focuses a ttention on the maintenance activities which have most effect on the performance of the plant. This helps to ensure that everything spent on maintenance is spent where it will do the most good. In addition, if R CM is correctly applied to existing maintenance sys tems, it reduces the a mount of routine work (in other words, mainte nance tasks to be undertaken on a cyclic basis) issued in each period, usually by 40% to 70%. On the other hand, if ReM is used to develop a new maintenance program, the resulting scheduled workload is much lower than if the program is developed by traditional methods. Longer useful life of expensive items, due to a carefully focused e m phasis on the use of on-condition maintenance techniques. A comprehensive database: An ReM review ends with a comprehen sive and fully documented record of the maintenance requirements of all the significant assets Llsed by the organisation . This makes it possible
20
Reliability-centred Maintenance
to adapt to changing circumstances (such as changing shiftpatterns or new technology) without having to reconsider all maintenance policies from scratch. It also enables equipment users to demonstrate that their mainte nance programs are built on rational foundations (the audit trail required by more and more regulators). Finally, the information stored on RCM worksheets reduces the effects of staff turnover with its attendant loss of experience and expertise. An RCM review of the maintenance requirements of each asset also provides a much clearer view of the skills required to maintain each asset, and for deciding what spares should be held in stock. A valuable by-product is also improved drawings and manuals. •
•
Greater motivation of illdividuals, especially people who are involved in the review process. This leads to greatly improved general under standing of the equipment in its operating context, together with wider 'ownership' of maintenance problems and their solutions. It also means that solutions are more likely to endure. Better teamwork: ReM provides a common, easily understood techni cal language for everyone who has anything to do with maintenance. This gives maintenance and operations people a better understanding of what maintenance can (and cannot) achieve and what must be done to achieve it.
All of these issues are part of the mainstream of maintenance manage ment, and many are already the target of improvement programs. A major feature of RCM is that it provides an effective step-by-step framework for tackling all of them at once, and for involving everyone who has any thing to do with the equipment in the process. RCM yields results very quickly. In fact, if they are correctly focused and correctly applied, RCM reviews can pay for themselves in a matter of months and sometimes even a matter of weeks, as discussed in Chapter 14. The reviews transform both the perceived maintenance requirements of the physical assets used by the organisation and the way in which the maintenance function as a whole is perceived. The result is more eost effective, more harmonious and much more successful maintenance.
2
Functions
Most people become engineers because they feel at least some affinity for things, be they mechanical, electrical or structural. This affinity leads them to derive pleasure from assets in good condition, but to feel offended by assets in poor condition. These reflexes have always been at the heart of the concept of preven tive maintenance. They have given rise to concepts such as 'asset care', which as the name implies, seeks to care for assets per se. They have also led some maintenance strategists to believe that maintenance is all about preserving the inherent reliability or built-in capability of any asset. In fact, this is not so. As we gain a deeper understanding of the role of assets in business, we begin to appreciate the significance of the fact that any physical asset is put into service because someone wants it to do something. So it follows that when we maintain an asset, the state which we wish to preserve must be one in which it continues to do whatever its llsers want it to do. Later in this chapter, we will see that this state what the users want is funda mentally different from the built-in capability of the asset. This emphasis on what the asset does rather than what it is provides a whole new way of defining the objectives of maintenance for any asset - one which focuses on what the user wants. This is the most important single feature of the ReM process, and is why many people regard ReM as 'TQM applied to physical assets'. Clearly, in order to define the objectives of maintenance in terms of user requirements, we must gain a crystal clear understanding of the functions of each asset together with the associated performance standards. This is why the RCM process starts by asking: •
what are the functions and associated performance standards of the asset in its present operating context?
This chapter considers this question in more detail . It describes how functions should be defined, explores the two main types of perfomlanee standards, reviews different categories of functions and shows how func tions should be listed.
22
Flinctions
Reliability-centred Maintenance
2. 1 Describing functions It is a well established principle of value engineering that a function state ment should consist of a verb and an object. It is also helpful to start such statement'> with the word 'to' ( 'to pump water', 'to transport people', etc). However, as explained at length in the next part of this chapter, users not only expect an asset to fulfil a function. They also expect it to do so to an acceptable level of performance. So a function definition and by implication the definition of the objectives of maintenance for the asset of is not complete unless it specifies as precisely as possible the level ity). capabil performance desired by the user (as opposed to the built-in _
be listed as: For instance, the primary function of the pump in Figure 2 . 1 would per minute. litres 800 than less not at Y Tank to X Tank To pump water from
•
verb, This example shows that a complete function statement consists of a user. the by desired ance perform an object and the standard of A jimction statement should consist of a verb,
an object and a desired standard ofpetiormance
2.2 Performance standards what The objective of maintenance is to ensure that assets continue to do asset any wants user any their users want them to clo. The extent to which If ance. perform of to do anything can be defined by a minimum standard ance we could build an asset which could deliver that minimum perform without deteriorating in any way, then that would be the end of the matter. The machine would run continuously with no need for maintenance. However, in the real world, things are not that simple. d The laws of physics tell us that any organised system which is expose total is ation deterior this of result to the real world will deteriorate. The end are disorganisation (also known as 'chaos' or 'entropy'), unless steps rate. deterio to taken to arrest whatever process is causing the system For instance, the pump in Figure 2 . 1 is pumping water into a tank from which the water is drawn at a rate of 800 litres/minute. One process that causes the pump to deteriorate (failure mode) is impeller wear. This happens regardless of whether it is pumping acid or lubricating oil, and regardless of whether the impeller is made of titanium or mild steel. The only question is how fast it will wear to the point that it can no longer deliver 800 litres/minute.
Figure 1:
Initial capability vs desired performance
23
- · _-� ·_J--r-�� �(.�_- ?i] I�Y -rJG _
water per mll1ute
_
Offtake from tank: 800 litres/minute
1
Margin for deterioration
w () z « � a: o u..
a: w a..
So if deterioration is inevitable, it must be allowed for. This means that when any asset is put into service, it must be able to deliver more than the minimum standard of performance desired by the user. What the asset is able to deliver is known as its initial capability (or in her ent reliability). Figure 2.2 illustrates the right relationship between this capabi lity and desired performanee. Figure 2.2:
Allowing for deterioration
For instance, in order to ensure that the pump shown in Figure 2. 1 does what its users want and to allow for deterioration, the system designers must specify a pump which has an initial built-in capability of something greater than 800 litresl minute. In the example shown, this initial capability is 1 000 litres per minute.
This means that performance can be defined in two ways, as follows: •
desired performance (what the user wants the asset to do)
•
built-in capability (what it can do)
Later chapters look at how maintenance helps ensure that assets continue to fulfil their intended functions, either by ensuring that their capability remains above the minimum standard desired by the user or by restoring something approaching the initial capability if it drops below this point. When considering the question of restoration, bear in mind that: •
•
the initial capability of any asset is established by its design and by how it is made maintenance can only restore the asset to this initial level of capability - it cannot go beyond it.
In practice, most assets are adequately designed and built, so it is usuallv possible to develop maintenance programs which ensure that such asse ;s continue to do what their users want.
24
Functions
Reliability-centred Maintenance
25
This underlines the i mportance of identifying precisely what the lIsers want when starting to d evelop a maintenance program. The following paragraphs explore k ey aspects o f performance standards in more d etail.
INITIAL CAPA B I LITY
Multiple performance standards Many function statements incorporate more than one and sometimes several performance standards.
w o z « ::2: n:: o u n:: w
For example, one function of a chemical reactor in a batch-type chemical plant might be listed as: To heat up to 500 kg of product X from ambient temperature to boiling point ( 1 25°C) in one hour. In this case, the weight of product, the temperature range and the time all present different performance expectations. Similarly, the primary function of a motor car might be defined as: To transport up to 5 people along made roads at speeds of up to 1 40 km/h Here the performance expectations relate to speed and number of passengers. •
0..
Figure 2.3: A m aintainable asset
In short, such assets are maintainable, as shown in Figure 2 .3. On the other hand, if the desired performance exceeds the initial cap ability, no a mount of maintenance can d eliver the d esired perform ance. 2.4. In other words, sllch assets are not maintainable, as shown in Figure For instance, if the pump shown in Fig ure 2 . 1 had an initial capability of 750 litres/minute, it would not be able to keep the tank full. Since the maintenance pro gram does not exist which makes pumps bigger, maintenance cannot deliver the desired performance in this context. Sim ilarly, if we make a habit of trying to draw 1 5 kW (desired pertormance) from a 1 0 kW electric motor (initial capability), the motor will keep tripping out and will even tually burn out prematurely. No amount of maintenance will make this motor big enough. It may be perfectly adequately designed and built in its own right - it just cannot deliver the desired performance in the context in which it is being used.
1
w o z « ::2: n:: o u n:: w
CL
Figure 2.4:
A non-maintainable situation
Two conclusions which can be drawn from the above exampl es are that: for any asset to be maintainable, the d esired performance of the asset must fall within the envelope of its initial capability is so, we not only need to know the • in order to determi ne whether this initial capability of the asset, but we also need to know exactly what in minimu m performance the user is prepared to accept in the context which the asset is being used .
•
Quantitative peJlormance standards Performance standards should b e quantified where possible, because quantitative standards are inherently much more precise than qualitative standards. Special care should be taken to avoid qualitative statements like 'to produce as many widgets as required by production ', or 'to go as fast as possible ' . Function statements of this type are meaningless, if only because they make it impossible to d efine exactly when the item is failed. In reality, it can be extraordinarily difficult to d efine precisely what is required, but just because it is difficult does not mean that it cannot or should not be done. One major user of ReM sum med up this point by saying ' If the users of an asset cannot specify precisely what p erformance they want from an asset, they cannot hold the maintainers accountable for sustaining that p erformance.' Qualitative standards In spite of the need to be precise, it is sometimes impossibl e to specify quantitative performance standards so we have to live with qualitative statements . For instance, the primary function of a painting is usually 'to look acceptable' (if not 'attractive'). What is meant by 'acceptable' varies hugely from person to person and is impossible to quantify. As a result, user and maintainer need to take care to ensure that they share a common understanding of what is meant by words like 'acceptable' before setting up a system intended to preserve that acceptability.
26
Reliahility-centred Maintenance
Functions
Absolute perjbrmance standards A function statement which contains no performance standard at all usually implies an absolute. For instance, the concept of containment is associated with nearly all enclosed systems. Function statements covering containment are often written as follows: To contain liquid X The absence of a performance standard suggests that the system must contain all the liquid, and that any leakage at all amounts to a failed state. In cases where an enclosed system can tolerate some leakage, the amount which can be toler ated should be incorporated as a performance standard in the function statement. •
Variahle performance standards Per formance expectations (or applied stress) sometimes vary infinitely between two extremes. Consider for example a truck used to deliver loads of assorted goods to urban retailers. Assume that the actual loads vary between (say) 0 (empty) and 5 tons, with an average of 2.5 tons, and the distribution of loads is as shown in Figure 2.5. To allow for deterio ration, the initial capability of the truck must be more than the 'worst case' load, which in this example is 5 tons. The maintenance program in turn must ensure that the cap ability does not drop below this level, in which case it would automatically satisfy the full range of performance expectations.
Figure 2.5:
Variable performance standards
Upper and lower limits In contrast to variable performance expectations, some systems exhibit variable capability. These are systems which simply cannot be set up to function to exactly the same standard every time they operate. For example, a grinding machine used to finish grind a crankshaft will not produce exactly the same finished diameter on every journal. The diameters will vary, if only by a few microns. Similarly, a filling machine in a food factory will not fill two successive containers with exactly the same weight of food. The weights will vary, if only by a few milligrams.
Figure 2 . 6 indicates that capability variations o f this nature usually vary about a mean. In order to a ccommodate this variability, the associated desired standards of per formance incorporate an upper and lower limit.
27
For instance, the primary function of a sweet-packing machine might be: To pack 250±1 gm of sweets into bags at a minimum rate of 75 bags per minute. . The pnmary function of the grinding ma chine might be: To finish grind main bearing journals in a cycle time of 3.00 ±0.03 minutes to a diameter of 75 ±0. 1 mm with a surface finish of RaO.2. •
•
t
- - - -
(In practice, this kind o fvariability is usually unwelcome for a number o f reasons. Ideally, processes should be so stable that there is no variation at Figure 2.6: all and hence no need for two limits. Upper and lower limits In pursuit o f this ideal, many indus tries are spending a great deal o ftime and energy on designing processes that vary as little as possible . How ever, th �s aspect of design and development is beyond the scope o f this book. RIght now we are concerned purely with variability from the view point of maintenance.) �ow much variability can be tolerated in the specification o fany prod u ct lS usually governed by external factors. For instance, the lower limit which can be tolerated on the crankshaft journal diameter IS governed by factors such as noise, vibration and harshness, and the upper limit by the clearances needed to provide adequate lubrication. The lower limit of the weight of the bag of sweets ( relative to the advertised weight) is usually governed by trading standards legislation, while the upper limit is governed by the amount of product which the company can afford to give away.
In cases like these, the desired performance limits are known as the upper and l ?wer specification limits. The limits o f capability (usually de fined as bemg three standard deviations either side o f the mean) are known as the �pper and lower control limits. Quality management theory suggests that m a well managed process, the difference between the control limits should ideally be half the difference between the specification limits. This multiple should allow a more than adequate margin for deterioration from a maintenance viewpoint. Upper and lower limits not only apply to product quality. They also apply to ot�er functional speci fications such as the a ccuracy of gauges and the settmgs of control systems and protective devices. This issue is discussed further in Chapter 3 .
28
Reliability-centred Maintenance
2.3 The Operati ng Context In Chapter 1, ReM was defined as 'a process used to determine the main tenance requirements of any physical asset in its operating context'. This context pervades the entire maintenance strategy formulation process, statting with the definition of functions.
For example, consider a situation where a maintenance program is being devel oped for a truck used to transport material from Startsville to Endburg. Before the functions and associated performance standards of this vehicle can be defined, the people developing the program need to ensure that they thoroughly under stand the operating context. For instance, how far is Startsville from Endburg? Over what sort of roads and what sort of terrain? What are the 'typical worst case' weather and traffic condi tions on this route? What load is the truck carrying (fragile? corroSive? abrasive? explosive?) What speed limits and other regulatory constraints apply to the route? What fuel facilities exist along the way? The answers to these questions might lead us to define the primary function of this vehicle as follows: 'To transport up to 40 tonnes of steel slabs at speeds of up to 60 mph (average 45 mph) from Startsville to Endburg on one tank of fuel'. The operating context also profoundly influences the requirements for secondary functions. In the case of the truck, the climate may demand air conditioning, regulations may demand special lighting, the remoteness of Endburg may demand that special spares be carried on board, and so on. Not only does the context drastically affect functions and performance expectations, but it also affects the nature of the failure modes which cou Id occur, their effects and consequences, how often they happen and what must be done to manage them.
For instance, consider again the pump shown in Figure 2.1. If it were moved to a location where it pumps mildly abraSive slurry into a Tank B from which the slurry is being drawn at a rate of 900 litres per minute, the primary function would be: To pump slurry into Tank B at not less than 900 litres per minute. This is a higher performance standard than in the previous location, so the stand ard to which it has to be maintained rises accordingly. Because it is now pumping slurry instead of water, the nature, frequency and severity of the failure modes also change. As a result, although the pump itself is unchanged, it is likely to end up with a completely different maintenance program in the new context •
All this means that anyone setting out to apply RCM to any asset or process must ensure that they have a crystal clear understanding of the operating context before they start. Some of the most important factors which need to be considered are discussed in the following paragraphs.
Functions
29
Batch amlflow processes In manufacturing plants, the most important feature of the operating con text is the type of process. This ranges from flow process operations where nearly all the equipment is interconnected, to jobbing operations where most of the machines are independent. In flow processes, the failure of a single asset can either stop the entire plant or significantly reduce output, unless surge capacit y Of stand-by plant is available. On the other hand, in batch or johbing plants, most fail ures only cUltail the output of a single machine or line. The consequences of such failures are determined mainly by the duratio n of the stoppage and the amount of work-in-process queuing in front of subseq uent operations. These ditIerenees mean that the maintenance strateg y applied to an asset which is part of a flow process could be radical ly different from the strategy applied to an identical asset in a batch environ ment.
Redundancy The presence of redundancy or alternative means of production is a feature of the operating context which must be conside red in detail when defining the functions of any asset
The importance of redundancy is illustrated by the three identical pumps shown 111 Figure 2.7. Pump B has a stand-by , while pump A does not.
Figure 2.7: Different operating contexts
Duty
Stand-by
00
This means that the primary function of pump A is to transfer liquid from one pOint to �nother on its own, and that of pump B to do it in the presence of a stand-by. ThiS difference means that the maintenance requirements of these pumps will be different (just how different we see later), even though the pumps are identical.
Quality standards Quality standards and standards of customer service are
two more aspects of the operating context which can lead to differences between the de scriptions of the functions of otherwise identical machine s.
For example, identical milling stations on two transfer machines might have the same basic function to mill a workpiece. However, deptll of cut, cycle time, flat ness tolerance and surface finish specifications might all be different. This could lead to quite different conclusions about their maintenance requirements.
30
Functions
Reliability-centred Maintenance
Environmental standards
An increasingly important aspect of the operating context of any asset is the impact which it has (or could have) on the environment. Growing worldwide interest in environmental issues means that when
we maintain any asset, we actually have to satisfy two sets of 'users'. The first is the people who operate the asset itself. The second is soeiety as a whole, which wants both the asset and the process of which it forms part not to calise undue harm to the environment.
31
In a single shift plant, production lost due t o failures can usually be made up by working overtime. This overtime leads to increased production costs, so maintenance strategies are evaluated in the light of these costs. On the other hand, if an asset is working 24 hours per day, seven days per week, it is seldom possible to make up for lost time, so downtime causes lost sales. This costs a great deal more than extra overtime, so it is worth trying much harder to prevent failures under these circumstances. However, it is also more difficult to make equipment available for maintenance in
What society wants is expressed in the form of increasingly stringent environmental standards and regulations. These are international, national,
a fully-loaded plant, so maintenance strategies need to be formulated with
narily wide range of issues, from the biodegradability of detergents to the
change, organisations can move from one end of this spectrum to the
regional, municipal or even corporate standards. They cover an extraordi
special care. As products move through their life cycles or as economic conditions
content of exhaust gases. In the case of processes, they tend to concentrate on unwanted liquid, solid and gaseous by-products.
other surprisingly quickly. For this reason, it is wise to review maintenance
ards. However, it is not enough simply to ensure that a plant or process is environmentally sound at the moment it is commissioned. Steps also have to be taken to ensure that it remains in compliance throughout its life.
Work-in-process refers to any material which has not yet been through all
Most industries are responding to society's environmental expectations by ensuring that equipment is designed to comply with the associated stand
Taking the right steps is becoming a matter of urgency, because all over the world, more and more incidents which seriously affect the environ ment are occurring because some physical asset did not behave as it should
in other words, because something failed. The associated penalties are becoming very harsh indeed, so long-term environmental integrity is now
a partieularly important issue for maintenance people.
Safety hazards
An increasing number of organisations have either developed themselves or subscribe to formal standards concerning acceptable levels of risk. In
some cases, these apply at corporate level, in others to individual sites and in yet others to individual processes or assets. Clearly, wherever such standards exist, they are an important part of the operating context.
policies every time this aspect of the operating context changes.
Work-in-process the steps of the manufacturing process. It may be stored in tanks, in bins, in hoppers, on pallets, on conveyors or in special stores. The consequen ces of the failure of any machine are greatly influenced by the amount of this work-in-process between it and the next machines in the process.
Consider an example where the volume of work in the queue is sufficient to keep the next operation working for six hours and it only takes four hours to repair the failure mode under consideration. In this case, the failure would be unlikely to affect overall output. Conversely, if it took eight hours to repair, it could affect overall output because the next operation would come to a halt. The severity of these consequences in turn depends on the amount of work-in-process between that operation and the next and so on down the line, and the extent to which any of the operations affected is a bottleneck operation (in other words an operation which governs the output of the whole line). •
•
Although plant stoppages cost money, it also costs money to hold stocks of work-in-process. Nowadays stock-holding costs of any kind are so high that reducing them to an absolute minimum is a top priority. This is
Sh({t arrangements
a major objective of 'just-in-time' systems and their derivatives.
Shift arrangements profoundly affect the operating context. Some plants operate for eight hours per day five days a week (and even less in bad
stocks provided against failure is rapidly disappearing. This is a vicious
times). Others operate continuously for seven days a week, and yet others
circle, because the pressure on maintenance departments to reduce failures
somewhere in between.
in order to make it possible to do without the cushion is also increasing.
These systems reduce work-in-process stocks, so the cushion that the
32
Reliability-centred Maintenance So from the maintenance viewpoint, a balance has to be struck between
Functions •
the economic implications of operational failures, and: •
the cost of holding work-in-process stock'> in order to mitigate the effeets of those failures, or
•
the cost of doing proactive maintenance tasks with a view to anticipat ing or preventing the failures.
To strike this balance successfully, this aspect of the operating context must
33
use ReM to develop a maintenance strategy based on existing spares
holding policies, •
review the failure modes associated with key spares on an exeeption basis,
by establishing what impact (if any) a change in the present stockholding policy would have on the initial maintenance strategy, and then picking
the most cost-effective maintenance strategy/spares holding policy.
be particularly clearly understood in manufacturing operations.
If this approach is adopted, then the existing spares holding policy can be seen as part of the (initial) operating context.
Repair time
Market demand
Repair times are influenced by the speed of response to the failure, which
The operating context sometimes features cyclic variations in demand for the products or services provided by the organisation.
is a function of failure reporting systems and staffing levels, and the ,speed
of repair itself, which is a function of the availability of spares and appro priate tools and of the capability of the person doing the repairs. These factors heavily influence the effects and the consequences of failures, and they vary widely from one organisation to ,mother. As a result, this aspect of the operating context also needs to be clearly understood.
Spares It is possible to use a derivative of the ReM process to optimise spares stocks and the associated failure management policies. This derivative is based on the fact that the only reason for keeping a stock of spare parts is to avoid or reduce the consequences of failure. The relationship between spares and failure consequences hinges on the time it takes to procure spares from suppliers. If it could be done instantly there would be no need to stock any spares at all. But in the real world procuring spares takes time. This is known as the lead time, and it ranges from a matter of minutes to several months or years. If the spare is not a stock item, the lead time often dictates how long it takes to repair the failure, and hence the severity of its consequences. On the other hand, holding spares in stock also costs money, so a balance needs to be struck, on a case-by-case basis, between the cost of holding a spare in stock and the total cost of not holding it. In some cases, the weight and/or dimen sions of the spares also need to be taken into account because of load and space restrictions, especially in facilities like oil platforms and ships. This spares optimization process is beyond the scope of this book. How ever, when applying ReM to an existing facility, one has to start somewhere. In most cases, the best way to deal with spares is as follows:
For example, soft drink companies experience greater demand for their products in summer than in winter, while urban transport companies experience peak de mand during rush hours. In cases like these, the operational consequences of failure arc much more serious at the times of peak demand, so in this type 0 f industry, this aspect
of the operating context needs to be especially clearly understood when defining functions and assessing failure consequences.
Raw material supply Sometimes the operating context is influenced by cyclic fluctuations in the supply of raw materials. Food manufacturers often experience peri
ods of intense activity during harvest times and periods of little or no activity at other times. This applies especially to fruit processors and sugar
mills. During peak periods, operational failures not only affect output, but can lead to the loss of large quantities of raw materials if these cannot be processed before they deteriorate.
Documenting the operating context For all the above reasons, it is essential to ensure that everyone involved in the development of a maintenance program for any asset fully understands the operating context of that asset. The best way to do so is to document
the operating context, if necessary up to and including the overall mission statement of the entire organisation, as part of the ReM process.
Figure 2.8 overleaf shows a hypothetical operating context statement for the grinding machine mentioned earlier. The crankshaft is used in a type of engine used in motor car model X.
34
Reliability-centred Maintenance Make car model X
(Corresponding asset: Model X Car Division)
Functions
Model X division employs 4 000 people to produce 220 000 cars this year. Sales forecasts indicate that this could rise to 320 000 per year within 3 years. We are now number 18 in national customer satisfaction rankings, and intend to reach 15th place next year and 10th place the following year. The target for lost time injuries throughout the division is one per 500 000 paid hours. The probability of a fatality occurring anywhere in the division should be less than one in 50 years. The division plans to conform to all known environmental standards.
Make engines
(Corresponding
The Motown Engine Plant produces all the engines for model X cars. 140 000 Type 1 and 80 000 Type 2 engines are produced per year. In order to achieve the customer satisfaction targets for the entire vehicle, warranty claims for engines
asset: Motown
must drop from the present level of 20 per 1000 to 5 per 1000. The plant suffered
Engine Plant)
three reportable environmental excursions last year
our target is not more
3S
The hierarchy starts with the division of the corporation which produces this model, but it could have gone up one level further to include the entire corporation. Note also that a context statement at any level should apply to aI/the assets below it in the hierarchy, not just the asset under review. The context statements at the higher levels in this hierarch y are simply broad function statements. Performance standards at the highest levels quantify expectations from the viewpoint of the overall business. At lower levels, performance standards become steadily more specific until one reaches the asset under review. The primary and secondary function s of
the asset at this level are defined as described in the rest of this chapter.
than one in the next three years. The plant shuts down for two weeks per year to allow production workers to take their main annual vacations.
Make Type 2 engines
The Type 2 engine line presently works 110 hours per week (2 x 10 hr shifts 5 days per week and one 10 hour shift on Saturdays). The assembly line could
(Corresponding
produce 140 000 engines per year in these hours if it ran continuously with no
assetType2
defects, but overall output of engines is limited by the speed of the crankshaft
Engine Line)
manufacturing line. The company would like as much maintenance as possible to be done during normal hours without interfering with production.
Machine crankshafts
The crankshaft line consists of 25 operations, and is nominally able to produce 20 crankshafts per hour (2 200 per week, 110 000 per 50 week year). It currently
(Corresponding
sometimes fails to produce the requirement 011 600 per week in normal time.
asset Crank
When this happens, the line has to work overtime at an additional cost of £800
shaft machining line2)
per hour. (Since most of the forecast growth will be for Type 2 engines, stop
pages on this line could eventually lead to lost sales of model X cars unless the performance is improved.) There should be no crankshafts stored between the end of the crankshaft line and the engine assembly line, but operations in fact keep a pallet of about 60 crankshafts to provide some 'insurance' against stop pages. This enables the crankshaft line to stop for up to 3 hours without stopping assembly. Crankshaft machining defects have not caused any warranty claims, but the scrap rate on this line is 4%. The initial target is 1.5%.
Finish grind crankshaft main and big end journals
The finish grinding machine grinds 5 main and 4 big end journals. It is the bottle neck operation on the crankshaft line, and the cycle time is 3.0 minutes. The finished diameter of the main journals is 75mm ± 0.1mm, and of the big ends
53mm ± 0.1 mm. Both journals have a surface finish of RaO.2. The grinding
(Corresponding
wheels are dressed every cycle, a process which takes 0.3 minutes out of each
asset: Ajax
3 minute cycle. The wheels need to be replaced after 3 500 crankshafts, and replacement takes 1.8 hours. There are usually about ten crankshafts on the
Mark 5 grinding machine)
conveyor between this machine and the next operation, so a stoppage of 25 minutes can be tolerated without interfering with the next operation. Total buffer stocks on the conveyors between this machine and the end of the line mean that this machine can stop for about 45 minutes before the line as a whole stops. Finish grinding contributes 0.4% to the present overall scrap rate.
Figure 2.8: An operating context statement
2.4 Different Types o f Fu nctions Every physical asset has more than one - often several
functions. If the objective of maintenance is to ensure that the asset can continue to fulfil these functions, then they must all be identified together with their cun-ent
desired standards of performance. At first glance, this may seem to be a fairly straightforward exercise. However in practice it nearly always turns out to be the single most challenging and time-consuming aspect of the maintenance strategy formulation process, This is especially true of older facilities. Products change, plant config
urations change, people change, technology changes and performa nce expectations change - but still we find assets in service that have been there since the plant was built. Defining precisely what they are supposed to be doing now requires very close cooperation between maintainers and users. It is also usually a profound learning experience for everyone involved . Functions are divided into two main categories (primary and
second ary functions) and then further divided into various sUb-categOIies. These are reviewed on the following pages, starting with primary function s.
Primary functions Organisations acquire physical assets for one, possibly two, seldom more than three main reasons. These 'reasons' are defined by suitably worded function statements. Because they are the 'main' reasons why the asset is acquired, they are known asprimaryjunctions. They are the reasons why the asset exists at all, so care should be taken to define them as precisely
as possible.
36
Reliability-centred Maintenance
Functions
Primary functions are usually fairly easy to recognise. In fact, the names
37
I n cases like this, one could list a separate function statement for each pro
of most industrial assets are based on their primary functions.
duct. This would logically lead to three separate maintenance programs
For instance the primary function of a packing machine is to pack things, of a crusher to crush something and so on.
for the sanle asset. Three programs may be feasible - perhaps even desirable if each product rnns continuously for very long periods. However, if the interval between long-tenn maintemmce tasks is longer
As mentioned earlier, the real challenge lies in defining the current per formance expectations associated with these functions. For most types of
than the change-over intervals, then it is impractical to change the tasks
equipment, the performance standards associated with primary functions
every time the machine is changed over to a different product. One way around this problem is to combine the 'worst case' standards
concern speeds, volumes and storage capacities. Product quality also usually
associated with each product into one function statement.
need to be considered at this stage. Chapter 1 mentioned that our ability to achieve and sustain satisfactory quality standards depends increasingly on the capability and condition of
In the above example, a combined function statement could be 'to reflux up to 750 litres of product at temperatures up to 180°C and pressures up to 10 bar.'
the assets which produce the goods. These standards are usually associ
This will lead to a maintenance program which might embody
ated with primary functions. As a result, take care to incorporate product quality criteria into primary function statements where relevant. These include dimensions for machining, fonning or assembly operations,
some over maintenance some of the time, but which will ensure that the asset can handle the worst stresses to which it will be exposed.
purity
standards for food, chemicals and pharmaceuticals, hardness in the case levels or weights for packaging, and so on.
of heat treatment, filling
Functional block diagrams If an asset is very complex or if the interaction between different systems is poorly understood, it is sometimes helpful to clarify the operating con text by drawing up functional block diagrams. These are simply diagrams showing all the primary functions of an enterprise at any given level. They m'e discussed in more detail in Appendix 1 .
Multiple independent primalY functions An asset can have more than one primary function. For instance, the very nmne of a military fighterlbomber suggests that it has two primary functions. In such cases, both should be listed in the functional specification.
A similar situation is often found in manufacturing, where the same asset may be used to perform different functions at different times. For instance, a single reactor vessel in a chemical plant might be used at different times to reflux (boil continu ously) three different products under three different sets of conditions, as follows: 3 2 1 Product Pressure 2 bar 10 bar 6 bar 140°C Temperature 180°C 120°C Batch size 500 litres 600 litres 750 litres (It could be said that this vessel is not performing three different functions, but that it is performing the same function to different standards of performance. In fact, the distinction does not matter because we arrive at the same conclusion either way.)
Serial or dependent primalY functions One often encounters assets which must perform two or more primary functions in series. These are known as serial functions.
For instance, the primary functions of a machine in a food factory may be 'to fill cans with food per minute' and then 'to seal 300 cans per minute'.
300
The distinction between multiple primary functions and serial primary functions is that in the fonner case, each function can be performe d inde pendently of the other, while in the latter, one function must be performed before the other. In other words, for the canning machine to work properly
it must fill the cans before it seals them. Secondary Functions
Most assets are expected to fulfil one or more functions in addition to their primary functions. These are known as secondary junctions.
For example, the primary function of a motor car might be described as follows: to transport up to 5 people at speeds of up to 140 km/h along made roads If this was the only function of the vehicle, then the only objective of the mainte nance program for this car would be to preserve its ability to carry up to 5 people at speeds of up 140 km/h along made roads. However, this is only part of the story, because most car owners expect far more from their vehicles, ranging from the ability to carry luggage to the ability to indicate how much fuel is in the fuel tank. •
To help ensure that none of these functions are overlooked, they are divi ded into seven categories as follows:
Reliability-centred Maintenance
38
•
environmental integrity
•
safety/structural integrity
•
control/containment/comfort
•
appearance
•
protection
•
•
Functions
39
A further subset o f safety-related fUllctiolls are those which deal with product contamination and hygiene. These are most often found in the food and pharmaceutical industries. The associated performance standards are usually tightly specified, and lead to rigorous and comprehensive main tenance routines (cleaning and testing/validation).
economy/efficiency
Structural integrity
superl1uous functions.
The first letters of each line in this list form the word ESCAPES. Although secondary functions are usually less obvious than primary functions, the loss of a secondary function can still have serious conse quences sometimes more serious than the loss of a primary function. As
a result, secondary functions often need as much if not more maintenance than primary functions, so they too must be clearly identified. The follow
Many assets have a structural secondary function. This usually involves supporting some other asset, sub-system or component.
For example, the primary function of the wall of a building might be to protect people and equipment from the weather, but it might also be expected to support the roof (and bear the weight of shelves and pictures).
ing pages explore the main categories of these functions in more detail.
Large, complex structures with multiple load bearing paths and high levels of redundancy need to be analysed using a specialised version of
Environmental integrity
and thc structural elements of offshore oil platforms.
RCM. Typical examples of such stmctures are airframes, the hulls of ships Part 2 of this chapter explained how society'S environmental expectations have become a critical feature of the operating context of many assets. RCM begins the process of compliance with the associated standards by incorporating them in appropriately worded function statements.
as any other function described in this chapter.
For instance, one function of a car exhaust or a factory smoke stack might be 'to contain no more than X micrograms of a specified chemical per cubic meter'. The car exhaust system might also be the subject of environmental restrictions deal ing with noise, and the associated functional specification might be 'to emit no more than X dB measured at a distance of Y metres behind the exhaust outlet'
�ty
.
Most users want to be reasonably sure that their assets will not hurt or kIll them. In practice, most safety hazards emerge later in the RCM process
as failure modes. However, in some cases it is necessary to write function statements which deal with specific threats to safety.
For instance, two safety-related functions of a toaster are 'to prevent users from touching electrically live components' and 'not to burn the users'. Many processes and components are unable to fulfil the safety expecta tions of users on their own. This has given rise to additional functions in
the form of protective devices. These devices pose some of the most diffi cult and complex challenges facing the maintainers of modern industrial plant. As a result, they are dealt with separately below.
Structures of this type are rare in industry in general, so the relevant analytical teehniques are not covered in this book. However, straightfor ward, single-celled structural elements can be analysed in the same way
Control In many cases, users not only want assets to fulfil functions to a given standard of performance, but they al�o want to be able to regulate the per fonnance. This expectation is summarised in separate function statements.
For instance, the primary function of a car as suggested earlier was 'to transport up to 5 people at speeds of up to 140 km/h along made roads'. One control function associated with this function could be 'to enable driver to regulate speed at will between -15 km/h (reverse) and +140 km/h'. Indication or feedback forms an important subset of the control category of functions. This includes functions which provide operators with real time information about the process (gauges, indicators, telltales, VDU's and control panels), or which record such information for later analysis
(digital or analog recording devices, cockpit voice recorders in aircraft, etc). Performance standards associated with these functions not only re late to the ease with which it should be possible to read and assimilate or to playback the information, but also cover its accuracy.
For instance, the function of the speedometer of a car might be described as 'to indicate the road speed to the driver to within +5 -0% of the actual
40
Reliability-centred Maintenance
FUllctions
Containment
41
Protective devices
In the case of assets used to store things, a primary function is to contain
As physical assets become more complex, the number of ways they can
whatever is being stored. However, containment should also be acknowl
fail is growing almost exponentially. This has led to corresponding growth
transfer material of
in the variety and severity of failure consequences. In an attempt to elimi
any sort -especially fluids. This includes pipes, pumps, conveyors, chutes,
nate (or at least to reduce) these consequences, increasing use is being
hoppers and pneumatic and hydraulic systems.
made of automatic protective devices. These work in one of five ways:
edged as a secondary function of all devices used to
Containment is also an important secondary function of items like gear boxes and transformers. (In this context, note again the remarks on Page 26
•
about performance standards and containment.)
Comf(Jrt
Most people expect their assets not to cause them anxiety, grief or pain.
These expectations are listed under the heading of 'comfort' because the major English dictionaries define comfort as being freedom from anxi
•
ety, pain, grief and so on. (These expectations can also be classified under the heading of 'ergonomics'.)
•
to draw the attention of the operators to abnormal conditions (warning lights and audible alarms which respond totailure effects. The ejlects are monitored by a variety of sensors including level switches, load cells, overload or overspeed devices, vibration or proximity sensors, temperature or pressure switches, etc) to shut down the equipment in the event of a failure (these devices also respond tofailure effects, using the same types of sensors and often the same circuits as alarms, but with different settings) to eliminate or relieve abnormal conditions which follow a failure and
Too much discomfort affects morale, so it is undesirable from a human
which might otherwise cause much more serious damage (fire-fIghting
viewpoint. It is also bad business because people who are anxious or in pain
equipment, safety valves, rupture discs or bursting discs, emergency medical equipment)
are more likely to make incorrect decisions. Anxiety is caused by poorly explained, unreliable or unintelligible control systems, be they for domestic appliances or for oil refineries. Pain is caused by assets - especially clothing and furniture which are incompatible with the people using them. The best time to deal with these problems is of course at the design stage. However, deterioration and/or changing expectations can cause this cat egory of functions to fail like any other. The best way to set about ensuring that this doesn't happen is to define appropriate functional specifications. For example. one function of a control panel might be 'To indicate clearly to a colour-blind operator up to five feet away whether pump A is operating or shut down'. A control-room chair might be expected 'To allow operators to sit com fortably for up to one hour at a time without inducing drowsiness'
•
to take over from a function which has failed (stand�hyplant olany sort,
redundant structural components) •
to prevent dangerous situations from arising in the first place (guards).
The purpose of these devices is to protect people from failures or to pro tect machines or to protect products.
in some cases all three.
Protective devices ensure that the failure of the function being protec ted is much less serious than it would be if there were no protection. The presence of protection also means that the maintenance requirements of a protected function are often less stringent than they would be otherwise. Consider a milling machine whose milling cutter is driven by a toothed belt. If the belt were to break in the absence of any protection, the feed mechanism would
Appearance The appearance of many items embodies a specific secondary function. For instance, the primary function of the paintwork on most industrial equipment is to protect it from corrosion, but a bright colour might be lIsed to enhance its visibility for safety reasons. Similarly, the main func tion of a sign outside a factory is to show the name of the company which occupies the premises, but a secondary [unction is to project an image.
drive the stationary cutter into the workpiece (or vice versa) and cause serious secondary damage. This can be avoided in two ways: •
by implementing a comprehensive proactive maintenance routine designed to prevent the failure of the belt
•
by providing protection such as a broken belt detector to shut down the machine as soon as the belt breaks. In this case, the only consequence of a broken belt is a brief stoppage while it is replaced, so the most cost-effective maintenance policy might simply be to let the belt fail. But
belt detector is working,
this policy is only valid if the broken
and steps must be taken to ensure that this is so.
42
Reliability-centred Maintenance
Functions
43
The maintenance of protective devices - especially devices which are not
For instance, a car might be expected 'to consume not more than 6 litres of fuel
fail-safe - is discussed in much more detail in Chapters 5 and 8. However,
per 100 km at a constant 120 km/h, and not more than 4 litres of fuel per 100 km
this example demonstrates two fundamental points: •
that protective devices often need more routine maintenance attention
at 60 km/h.'
A fossil fuel power station might be expected 'to export at least 45% A plant using an expensive
of the latent energy in the fuel as electrical power.'
solvent might want 'to lose no more than 0.5% of solvent X per month'.
than the devices they are protecting •
that we cannot develop a sensible maintenance program for a protected function without also considering the maintenance requirements of the protective device.
It is only possible to consider the maintenance requirements of protective devices if we understand their functions. So when listing the functions of any asset, we must list the functions of
all protective devices.
A final point about proteeti ve devices concerns the way their functions should be described. These devices act by exception (in other words when something else goes wrong), so it is important to describe them correctly.
Superfluous Functions Items or components are sometimes encountered which arc completely superfluous. This usually happens when equipment has been modified frequently over a period of years, or when new equipment has been over
specified. (These comments do not apply to redundant components built in for safety reasons, but to items which serve no purpose at all in the con
text under consideration.)
For example, a pressure reducing valve was built into tile supply line between a gas manifold and a gas turbine. The original function of the valve was to reduce the gas pressure from 120 psi to 80 psi. The system was later modified to reduce
In particular, protective function statements should include the words 'if
the manifold pressure to 80 psi, after which the valve served no useful purpose.
or'in the event of, followed by a very brief summary of the circumstan
It is sometimes argued that items like these do no harm and it costs money
ces or the event which would activate the protection. For instance, if we were to describe the function of a tripwire as being 'to stop the machine', anyone reading this description could be forgiven for thinking that the tripwire is the normal stop/start device. To remove any ambiguity, the function of a tripwire should be described as follows: •
to be capable ofstopping the machine in the event ofan emergency at any pOint along its length
The function of a safety valve may be described as follows: •
to be capable of relieving the pressure in the boiler
ifit exceeds 250 psi.
Economy/efficiency Anyone who uses assets of any sort only has finite financial resources.
to remove them, so the simplest solution may be to leave them alone until the whole plant is decommissioned. Unfortunately, this is seldom true in practice. Although these items have no positive function, they can still fail
and so reduce overall system reliability. To avoid this, they need mainten ance, which means that they still consume resources. It is not unusual to find that between 5% and 20% of the components
of complex systems are superfluous. in the sense described above. If they are eliminated, it stands to reason that the same percentage of mainten ance problems and costs wi II also be eliminated. However, before this can
be done with confidence, the functions of these components first need to
be identified and clearly understood.
This leads them to put a limit on what they are prepared to spend on operating and maintaining it. How much they are prepared to spend is
A Note on Reliability
governed by a combination of three factors:
There is often a temptation to write 'reliability' function statements such as 'to operate 7 days a week, 24 hours per day'. In fact, reliability is not
•
•
•
the actual extent of their financial resources how much they want whatever the asset will do for them the availability and cost of competitive ways of achieving the same end.
At the operating context level, functional expectations concerning costs are usually spelled out in the form of expenditure budgets. At the asset level, economic issues can be addressed directly by function statements which define what users expect in areas such as fuel economy and loss of process materials.
a function in its own right. It is a performance expectation which pervades all the other functions. It is properly dealt with by dealing appropriately with each of the failure modes which could cause each loss of function. This issue is discussed further in Chapter 13.
44
Reliability-centred Maintenance
Using the ESCAPES Categories There will often be doubt about which of the ESCAPES categories some functions belong to. For instance, should the function of a seat reclining
3
Functional Failures
mechanism be classified under the heading of 'control' or 'comfort'? In practice the precise classification does not matter. What does matter is that wc identify and define al\ the functions which are likely to be ex pected by the user. The list of categories merely serves as an aide memoire to help ensure that none of thesc expectations are overlooked.
Chapter I explained that the RCM process entails asking seven questions about selected assets, as follows: •
2.5 How Functions should be lAsted
what are the functions and associated performance standards of the asset in its present operating context?
•
in what ways does it fail to fu?fil its functiolls?
•
what causes each fUllctional failure?
•
what happens when each failure occurs?
ensures that maintenance activities remain focused on the real needs of
•
in what way does each failure matter?
the users (or 'customers'). It also makes it easier to absorb changes trig
•
what call be dOlle to predict or prevent each failure?
•
what should be done (f a suitable proactive task cannot be found?
A properly written functional specification - especially one which is fully quantified
precisely defines the objectives of the enterprise. This en
sures that everyone involved knows exactly what is wanted, which in turn
gered by changing expectations without derailing the whole enterprise. Functions are listed on ReM Information Work sheets in the left hand col umn. Primary functions are listed first and the functions
RCMII INFORMATION WORKSHEET
© 1996 ALADON LTO 1
To channel all the hot turbine exhaust gas without restriction to a fixed point 10 metres above the roof of the turbine hall
2
To reduce exhaust noise levels to ISO Noise Rating 30 at 150 metres
3
To ensure that the surface temperature of the dueting inside the turbine hall does not exceed 60"C
(These functions apply to the exhaust system of a 5 megawatt gas turbine.) A complete Information W orksheel is shown at the end of Chapter 4. 4
'Figure 2.9:
Describing functions
SUB-SYSTEM
Chapter 2 dealt at length with the first question. After a brief introduction to the general concept of failure, this chapter considers the second question, which deals withfun ctionalfail u res
.
FUNCTION
are numbered numerically, as shown in Figure 2.9.
SYSTEM
5
To transmit a warning signal to the turbine control system if the exhaust gas temperature exceeds 475"C and a shutdown signal if it exceeds 500"C at a point 4 metres from the turbine To allow free movement of the ducting in response to temperature changes
3.1 Failure In the previous chapter, we saw that people or organisations acquire assets because they want the assets to do something. Not only that, but they also expect their assets to fulfil the intended functions to an acceptable stan dard of performance. Chapter 2 went on to explain that for any asset both to do what its users want and to allow for deterioration, the initial capability of the asset must exceed the desired standard of performance. Thereafter, as long as the capability of the asset continues to exceed the desired standard of per formance, the user will be satisfied. On the other hand, if for any reason the asset is unable to do what the user wants, the user will consider it to have failed.
46
Reliability-centred Maintenance
Functional Failures
Conversely, it is equally possible for the pump to deteriorate to the paint where
This leads to a basic definition of failure:
it cannot pump the required amount (failed in terms of its primary function), while it still contains the liquid
'Failure' is defined as the inability of any asset to do what its users want it to do.
t I
What the uO sderO
......w�a.nltis.it.t..IIl�. What the asset
can do
LU () Z
47
(not failed in
terms of the secondary function).
This shows why it is more accurate to define failure in terms of the loss
This is illustrated in Figure 3.l .
of specific functions rather than the failure of an asset as a whole. It also shows why the ReM process uses the term 'functional failure' to describe
For instance, if the pump shown in Figure
failed states, rather than 'failure' on its own. However, to complete the
2.1 on Page 22 is unable to pump 800
definition of failure, we also need to look more closely at the question of
litres per minute, it will not be able to keep the tank full and so its users will regard it as 'failed'.
performance standards. Performance Standards and Ii'ailure As discussed in the first part of this chapter, the boundary between satis
u.
a: LU ��------
factory performance and failure is specified by a perfonnance standard.
Figure 3. 1 :
Given that performance standards apply to individual functions, 'failure'
The general failed state
can be defined precisely by defining a functional failure as follows:
A functional failure is defined as the inability of any asset to fulfil a function to a standard of performance which is acceptable to the llser
3.2 Functional Failures The above definition treats the concept of failure as if it applies to an asset as a whole. In practice, this definition is vaguc because it does not distin guish clearly between the failed state (functional failure) and the events which cause the failed state (failure modes). It is also simplistic, because it does not take into account the fact that each asset has more than one function, and each function often has more than one desired standard of performance. The implications are explored in the following paragraphs. Functions and Failures
The following paragraphs discuss different aspects of functional fail ure under the following headings: •
partial and total failure
•
upper and lower limits
•
gauges and indicators
•
the operating context.
Partial and total failures
We have seen that an assct is failed if it doesn't do what its users want it
The above definition of functional
to do. We have also seen that what anything must do is defined as a func
failure covers complete loss of func
tion and that every asset has more than one and often several different
tion. It also covers situations where
functions. Since it is possible for each one of these functions to fail, it
the asset still functions, but perrom1s
follows that any asset can suffer from a variety of different failed states.
outside acceptable limits.
w () Z
a: LU � �---------Figure 3.2:
Functional failure
For instance, the pump in Figure 2.1 has at least two functions. One is to pump
For example, the primary function of the pump discussed earlier is 'to pump water
water at not less than 800 litres/minute, and the other is to contain the water. It is
from tank X to Tank Y at not less than 800 litres/minute'. This function could suffer
perfectly feasible for such a pump to be capable of pumping the required amount
from two functional failures, as follows:
(not failed in
•
fails to pump any water at all
•
pumps water at less than 800 litres per minute.
terms of its primary function) while leaking excessively
terms of the secondary function).
(failed in
48
Reliability-centred Maintenance
Functional Failures
49
Partial failure is nearly always caused by different failure modes from total
The function of a crankshaft grinding machine was listed as To finish grind main
failure, and the consequences are different. This is why all the functional fail
bearing journals in a cycle time of 3.00 ±0.03 minutes to a diameter of 75 ±0. 1 mm with a surface finish of Ra 0.2'.
ures which could affect each function should be recorded.
•
Record all the functional failures associated with each function. Note that partial failure should not be confused with the situation where
INITIAL CAPABILITY
the asset deteriorates slightly but its capability remains above the level of performance required by the user. For example, the initial capability of the pump in Figure 2.1 is 1 000 litres per min ute. Impeller wear is inevitable, so this capability will decline. As long as it does not decline to the point where the pump is unable to pump 800 litres per minute, it will still be able to fill the tank and so keep the users satisfied in the context described.
1
What it can d
Actual deterioration
Margin for deterioration
Completely unable to grind workpiece
•
Grinds workpiece in a cycle time longer than 3.03 minutes
•
Grinds workpiece in a cycle time less than 2.97 minutes
•
Diameter exceeds 75.1 mm
•
Diameter is below 74.9 mm
•
Surface finish too rough.
Of course, if only one limit applies to a particular parameter, then only one failed state is possible. For instance, the absence of a lower limit on the roughness specification in the above example suggests that it is not pos sible to make the item too smooth. In some circumstances, this may not
UJ () z « :2 a: o
actually be true, so care needs to be taken to verify this point when analys ing functions of this type. In practice, the failed states associated with upper and lower limits can manifest themselves in two ways. Firstly, the spread of capability could
u..
a: UJ 0...
breach the specification limits in one direction only. This is illustrated in Figure 3.4, which shows that this type of failed state can be likened to a number of shots hitting a target in a tight group but way off center.
However if the capability of the as
Figure 3.3:
set deteriorates so much that it falls
Asset still OK despite some deterioration
below the desired performance, its users will consider it to have failed.
Upper and lower limits The previous chapter explained that the performance standards associated with some functions incorporate upper and lower limits. Such limits mean that the asset has failed if it produces products which are over the upper limit or below the lower limit. In these cases, the breach of the upper limit usually needs to be identified
separately from
the breach of the lower
limit. This is because the failure modes andlor the consequences associa ted with going over the upper limit are usually different from those asso ciated with going below the lower limit. For example, the primary function of a sweet-packing machine was listed in Chap ter 2 as being 'To pack 250±1 gm of sweets into bags at a minimum rate of 75 bags per minute' This machine has failed: •
if it stops altogether
•
if it packs more than 251 gm of sweets into any bags
•
if it packs less than 249 gm into any bags
•
if it packs at a rate of less than 75 bags per minute.
Figure 3.4:
Capability breaches upper limit only The second failed state occurs when the spread of capability is so broad that it breaches both the upper and the lower specification limits. Figure
3 . 5 shows that this can be likened to shots scattered all over the target. Note that in both of the above cases, not all of the products produced by the processes in question will be failed. If the breach is minor, only a small percentage of out-of-spec products will be produced. However, the further off centre the grouping in the first case, or the broader the spread in the second case, the higher will be the percentage of failures.
50
Reliability-centred Maintenance
51
Functional Failures Who should set the standard?
An issue which needs careful consideration when defining functional failures is the 'user'. To this day, most maintenance programs in use around the world are compiled by maintenance people working on their own. These people usually decide for themselves what is meant by 'failed' . In practice, their view of failure often turns out to be quite different from that of the users, with sometimes disastrous consequences for the effectiveness of their programs. Figure 3.5:
Capability breaches upper and lower limits Figure 2.6 illustrated a process which is in control and in specification. Figures 3 .4 and 3 . 5 show that processes which are out of control and out of spec are in a failed state. The failure modes which can cause these failed states are discussed in the next chapter. (Chapter 7 deals with the implica tions of a process which is out-of-control but within specification.)
Gauges and indicators The above discussion has tended to focus on product quality. Chapter 2 men tioned that upper and lower limits also apply to the performance standards associated with gauges, indicators, protection and control systems. De pending on failure modes and consequences, it may also be necessary to treat the breach of these limits separately when listing functional failures.
For instance, the function of a temperature gauge could be listed as 'to display the temperature of process X to within (say) 2% of the actual process temperature'. This gauge can suffer from three functional failures, as follows: fails altogether to display process temperature displays a temperature more than 2% higher than the actual temperature displays a temperature more than 2% lower than the actual temperature. • •
•
Functional failures and the operating context The exact definition of failure for any asset depends very much on its operating context. This means that in the same way that we should not generalise about the functions of identical assets, so we should take care not to generalise about functional failures.
For example, we saw how the pump shown in Figure 2 . 1 fails if it is completely unable to pump water, and if it is able to pump less than 800 litres/minute. If the same pump is used to fill a tank from which water is drawn at 900 litres/minute, the second failed state occurs if the throughput drops below 900 litres/minute.
For example, one function of a hydraulic system is to contain oil. How well it should fulfil this function can be subject to widely differing points of view. There are pro duction managers who believe that a hydraulic leak only amounts to a functional failure if it is so bad that the equipment stops working altogether. On the other hand, a maintenance manager might suggest that a functional failure has occurred if the leak causes excessive consumption of hydrauliC oil over a long period of time. Then again, a safety officer might say that a functional failure has occurred if the leak creates a pool of oil on the floor in which people could slip and fall or which might create a fire hazard. This is illustrated in Figure 3.6. Figure 3.6:
Different views about failure
L
g
Leak deteriorates
o
TIME
-+
l::!� ! H C..QNSUMPI!Qt:! :--
...
� ."�__._�
"FAILED" says.ml;lintalner
�o
'
�
EQUIPM!=NI"�TQE.§ 'yy'ORKIf;:!G "FAILED" says production manager
_ � ��_
The maintenance manager (who controls the hydraulic oil budget) may ask the operators for access to the hydraulic system to repair leaks 'because oil consump tion is excessive'. However access may be denied because the operators think the machine 'is still working OK'. When this happens, the maintenance people (1 ) record that the machine 'was not released for preventive maintenance', and (2) form the opinion that their production colleagues 'don't believe in PM'. For similar reasons, the maintenance manager might not release a maintenance person to repair a small leak when requested to do so by the safety officer. In fact, all three parties almostcertainly do believe in prevention. The real problem is that they have not taken the trouble to agree exactly what is meant by 'failed' so they do not share a common understanding of what they are seeking to prevent. This example illustrates three key points:
52
•
Reliability-centred Maintenance
the performance standard used to define functional failure
in other
words, the point where we say 'so far and no further' - defines the level of proactive maintenance needed to avoid that failure (in other words, to sustain the required level of performance) •
Failure Modes and Effects Analysis (FM EA)
much time and energy can be saved if these performance standards are clearly established
•
4
before the failures occur
the perfonnance standards used to define failure
must be set by opera
tions and maintenance people working together with anyone else who
We have seen that by defining the functions and desired standards of per
has something legitimate to say about how the asset should behave.
formance of any asset, we are defining the objectives of maintenance with respect to that asset. We have also seen that defining functional failures
How Functional Failures should be Listed
enables us to spell out exactly what we mean by 'failed'. These two issues
Functional failures arc listed in the second eolumn of the RCM Informa tion Worksheet. They are coded alphabetically, as shown in Figure 3 . 7 .
RCMII INFORMATION WORKSHEET
© 1996 ALADON LTD
SYSTEM
._---
'El(fiaust System
To channel all the hot turbine exhaust gas without restriction to a fixed point 10 metres above the rool of the turbine hall
are
This chapter describes the main elements of an FMEA, starting with a definition of the term 'failure mode '
A
Unable to channel gas at all
B
Gas flow restricted
C
Fails to contain the gas
D
Fails to convey gas to a point 10 m above the roof
asset (or system or process) to fail. However, Chapter 3 explained that it whole. It is much more precise to distinguish between a 'functional fail
2
To reduce exhaust noise levels to ISO Noise Rating 30 at 150 metres
A
Noise level exceeds ISO Noise Rating 30 at 150 metres
3
To ensure that duct surface temperature inside turbine hall does not rise above 60"C
A
Duct surface temperature exceeds 60''C
4
To transmit a waming Signal to the control system if exhaust gas temperature exceeds 475"C and a shutdown signal if it exceeds 500''C at a point 4 metres from the turbine
A
Incapable 01 sending a warning signal
B
Incapable of sending a shutdown signal il exhaust temperature exceeds 500°C
To allow free movement of ducting in response to temperature changes
A
Does not allow free movement of ducting
5
modes which
FUNCTIONAL FAILURE
.... u,"'-' " Vf\I
1
The next two questions seek to identify the failure
reasonably likely to cause each functional failure, and to ascertain thefclil
ure efjects associated with each failure mode.This is done by performing a failure modes and effects analysis (FMEA) for each functional failure.
5 MW
rsUB-SYSTEM
were addressed by the first two questions of the RCM process.
if exhaust temperature exceeds 475°C
4.1 What is a Fail ure Mode ? A failure mode could be defined as apy event which is likely to cause an is both vague and simplistic to apply the term 'failure' to an asset as a ure' (a failed
state)
and a 'failure mode' (an
e vent which could cause a
failed state). This distinction leads to the following more precise defini tion of a failure mode:
A failure mode is any event which causes a functional failure The best way to show the connection and the distinction between failed states and the events which could cause them, is to list functional failures
Figure 3.7: Describing functional failures
first, then to record the failure modes which could cause each functional failure, as shown in Figure 4.1 overleaf.
54
Reliability-centred Maintenance
RCM II INFORMATION WORKSHEET
© 1996 ALADON LTD
SYSTEM
55
4.2 Why Analyse Failu re Modes?
Coofing 'Water Pumping S!fsten --
SUB-SYSTEM
A single machine can fail for dozens o f reasons. A group o f machines or
FUNCTION
1
Failure Modes and Effects Analysis (FMEA.)
FUNCTIONAL FAILURE
(Loss of Function)
To tran::.>fer water from tank X to tank Y at not than 800
A
I Unable to tra��fer any water at all
htresiminute
FAILURE MODE
5 1I 1 2 3 4 6
B
Tf ansfers !ess than 800
! litms
rninute
1 2
(Cause of Failure) reunng seizes
system such as a production line can fail for hundreds of reasons. For an entire plant, the number can rise into the thousands or even tens of thousands.
Impeller comes adrift Impeller jammed by foreign object
Most managers shudder at the thought of the time and effort likely to
Coupling hub shears due to fatigue Motor burns out
be involved in identifying all these failure modes. Many decide that this
Inlet valve jams dosed
type of analysis is just too much work, and abandon the whole idea entirely.
' " etc
I Impeller
I
r-
-----
worn
Pal1ially blocked suction line ... etc
Figure 4. 1 : Failure modes of a pump
In doing so, these managers overlook the fact that on a day-to-day basis.
maintenance is really managed at the failure mode level. •
work orders or job requests are raised to cover specific failure modes
•
day-to-day maintenance planning is all about making plans to deal with
Figure 4.1 also indicates that at the very least, a description of a failure
specific failure modes
mode should consist of a noun and a verb. The description should contain enough detail for it to be possible to select an appropriate failure manage
•
discussions about what has failed, what caused it (and who is to blame),
wasted on the analysis process itself.
what is being done to repair it and
In particular, the verbs used to describe failure modes should be chosen
sometimes
what can be done to
stop it happening again. In short, the entire meeting is spent discussing
with care, because they strongly influence the subsequent failure man
failure modes
agement policy selection process. For instance. verbs such as 'fails ' or little or no indication as to what might be an appropriate way of managing
in most industrial undertakings, maintenance and operations people hold meetings every day. The meetings usually consist almost entirely of
ment strategy, but not so much detail that excessive amounts of time are
' breaks' or 'malfunctions' should be used sparingly, because they give
For instance:
•
to a large extent, technical history recording systems record individual failure modes (or at least, what was done to rectify them).
the failure. The use of more specific verbs makes it possible to select from
Sadly, in too many cases, these failure modes are discussed, recorded or
the full range of failure management options.
otherwise dealt with after they have' occurred. Dealing with failures after
For example, a term like 'coupling fails' provides no clue as to what might be done
they have happened is of course the essence of
to antiCipate or prevent the failure. However, if we say 'coupling bolts come loose'
Proactive management on the other hand, means dealing with events before they occur - or at least, deciding how they should be dealt with if
or 'coupling hub fails due to fatigue', then it becomes much easier to identify a possible proactive task.
reactive maintenance.
they were to occur. In order to do this, we need to know beforehand what
In the case of valves or switches, one should also indicate whether the loss
events are likely to occur. The ' events ' in this context are failure modes.
of function is caused by the item failing in the open or closed position
So if we wish to apply truly proactive maintenance to any physical asset,
'valve jams closed' says more than 'valve fails'. In the interests of complete
we must try to identify all the failure modes which are reasonably likely
clarity, it may sometimes be necessary to take this one step further.
to atfect that asset. Ideally, they should be identified before they occur at
For instance, 'valve jams closed due to rust on lead screw' is clearer than 'valve
all, or if this is not possible, before they occur again.
jams closed'. Similarly, one might need to distinguish between 'bearing seizes due
Once each failure mode has been identified, it then becomes possible
to normal wear and tear' and 'bearing seizes due to lack of lubrication'.
to consider what happens when it occurs, to assess its consequences and
These issues are discussed at length later in this chapter, but first we look
to decide what (if anything) should be done to anticipate, prevent, detect
at why we need to analyse failure modes at all.
or correct it
or perhaps even to design it out
Copyrighted Material
So Ihl' m:tillll'n:tnel' lask seleclion process qucnl l11an;I);CrnCIll onhcsc la�k�
-
-
and mnch of lhe subsc
i� carried OUI at Ihe failure modc kl·cl.
Thi� is briefly iIlustratcd in the following enmple. and then di>cussed al lengl h in the rel11aining
ch ap t�n.
of Ihis book:
ConsKier again the ReM Information WOf\<shei!t shown ,n Figure 4.1. This appies to the p
Impeller worn out
. I
:
USEflJlllFE
-I ;I _
Manage rh,s faj/ure by: chang'f19 impellers before end of ·useful Me-? Impeller
iammed Manaf1& this failure by: installing screen in suction line?
Fi�ure 4.2: Failures of too impeller of a centrifugal pump
•
Impellel wear is likely to be an age·related phenomenon. As shown in Figure 4.1 Ihis meanslhal il is likely 10 conform to lhe second ofthe six lailure panerns introduced In Figure 1.5 on Page 12 (Failure Pattern B). So ,I we know roughly •
what the useful lile of the impeller is. and � the consequences ot the failure are serious enough. lhen � may dec� to prevent thiS fa,lure by changing the impeller juSI before Ihe end 01 the useful lile. •
ImpelJeriammedby foreign object. The likelihood of a fO'eign object appearing in the S\lClOOIl line will almost certainly ha�e noth'ng 10 do wilh how long the im· peller has boon in selVice. As a result il stands to reason lhat this lailure mode w,1I occur on a random bas's (Pattern E in Figure 1 .5)_ There W<.Iuid also be no .
warning 01 the fact that the failure is about to (fC(:Uf SO � Ihe consequences were Serious enough. and the failure happened often enough, we would be likely to consider modifyrng the system perhaps by installing some sort 01 lilte, ,
or screan in the sOOion linll_
Copyrighted Material
Failure Modes and Effects Analysis (FMEA) •
57
Impeller adrift: If the impeller fastening mechanism is adequately designed and it still keeps coming adrift, this would almost certainly be because it wasn't put on properly in the first place. (If we knew that this was so, then perhaps the failure mode should actually be described as 'Impeller fitted incorrectly'.) This in turn means that the failure mode is most likely to occur soon after start up, as shown in Figure 4.2 (Pattem F in Figure 1 .5), and we would probably deal with it by improving the relevant training or procedures.
This example reinforces the point that the level at which we manage the maintenance of any asset is not at the level of the asset as a whole (in this case, the pump), and not even at the level of any component (in this case, the impeller), but at the level of each failure mode. So before we can de velop a systematic, proactive maintenance management strategy for any asset,
we must identify what these /ailure modes are
(or could be).
The example also suggests that one of the failure modes could be elim inated by a design change and another by improving training or proce dures. So not every/ailure mode is dealt with by scheduled maintenance. Chapters 5 to 9 describe an orderly approach to deciding what is likely to be the most suitable way of dealing with each failure. Note also that the failure management solutions proposed in Figure 4.2 represent only one of several possibilities in each case.
For instance we could monitor impeller wear by monitoring the pump perform ance and only change the impeller when it needs it. We also need to bear in mind that adding a screen to the suction line adds three more failure possibilities, which need to be analysed in turn (it could block up, it could be holed and therefore cease to screen, and it could disintegrate and damage the impeller.) Chapters 6 to 9 examine these alternatives in more detail. These points all indicate that the identification of failure modes is one of the most important steps in the development of any program intended to ensure that any asset continues to fulfil its intended functions. In prac tice, depending on the complexity of the item, its operating context and the level at which it is being analysed, between one and thirty failure modes are usually listed per functional failure. The next two sections of this chapter consider two of the key issues in this area under the following headings: •
categories of failure modes
•
level of detail.
Thereafter, the last three parts of the chapter consider failure effects, sources of information for an FMEA, and how failure modes and effects should be listed.
58
Reliability-centred Maintenance
Failure Modes and Ejfects Analysis (FMEA )
asset to deteriorate by lowering its capability, or more accurately, its re
4.3 Categories o f Fail ure Modes
sistance to stress.
Some people regard maintenance as being all about - and only about dealing with deterioration. Some even go so far as to specify that FMEA's carried out on their assets should deal only with failure modes caused by deterioration, and should ignore other categories of failure modes (such as human errors and design flaws). This is unfortunate, because it often transpires that deterioration causes a surprisingly small proportion of fail ures. In these cases, restricting the analysis to deterioration only can lead to a woefully incomplete maintenance strategy. On the other hand, if one accepts that maintenance means ensuring that physical assets continue to do whatever their users want them to do, then a comprehensive maintenance program must address all the events that are reasonably likely to threaten that functionality. Failure modes can be classified into one of three groups, as follows: when capability falls below desired performance
•
when desired performance rises above initial capability
•
when the asset is not capable of doing what is wanted from the outset.
Each of these categories is discussed in the following paragraphs.
with, but then drops below desired service, as illustrated in Figure 4 . 3 . Th e five principal causes o f reduced capability are listed below: •
deterioration
•
lubrication failures
•
dirt
•
•
disassembly 'capability reducing' human en·ors.
abrasion, erosion, evaporation, degradation of insulation, etc). These failure modes should of course be included in a list of failure modes wherever they are thought to be reasonably likely. The level of detail with which they need to be recorded is discussed in the next part of this chapter.
Lubrication failure Lubrication is associated with two types of failure modes. The first con cerns lack of lubricant, and the second the failure of the lubricant itself. With regard to lack of lubrication, things have changed considerably in the last two decades. Twenty years ago, the majOlity of lubrication points were replenished manually. The cost of lubricating each of these points to the cost of analysing the lubrication requirements of each point in detail. This meant that it was just not worth carrying out an in-depth analytical exer cise to set up a lubrication program. Instead, these programs were usually set up on the basis of a quick survey by a lubrication specialist. cation systems have become the norm in most industries. This has led to
The first category of failure modes
performance after the asset is put into
in other words, it fails.
Deterioration covers all fonns of 'wear and tear' (fatigue, corrosion,
Nowadays however, 'sealed-for-life' components and centralised lubri
Falling Capability
above desired performance to begin
Eventually resistance drops so much that the asset can
no longer deliver the desired performance
was tiny compared to the cost of not doing so. It was also tiny compared
•
covers situations where capability is
5<)
a massive reduction in the number of points where a human has to apply oil or grease to a machine, and a massive increase in the consequences of failure (especially the failure of ce n tralised lubrication systems). From
1
the analytical viewpoint, this means that it is now cost-effective to:
w o z « 2 a: o
•
use ReM to analyse centralised lubrication systems in their own right
•
consider the loss of lubricant in the few remaining manually lubricated points as individual failure modes.
The second category of failures associated with lubrication concerns
lL.
a: w 0...
deterioration of the lubricant itself. It is caused by phenomena such as shearing of the oil molecules, oxidation of the base oil and additive deple Figure 4.3:
Failure Mode Category 1
tion. In some cases, deterioration of the oil may be aggravated by the build up of sludge or the presence of water or other contaminants. A lubricant may also fail to do its job simply because the wrong lubricant has been
Deterioration
used. If any or all of these failures modes are considered to be likely in the
Any physical asset that fulfils a function which brings it into contact with
context under consideration, then they should be recorded and subjected to
the real world is subject to a variety of stresses. These stresses cause the
further analysis. (This also applies to transformer oil and hydraulic oil.)
60
Reliability-centred Maintenance
Dirt Dirt or dust is a very common cause of failure. It interferes directly with machines by causing them to block, stick or jam. It is also a principal cause of the failure of functions which deal with the appearance of assets (things which should look clean look dirty). Dirt can also cause product quality problems, either by getting into the clamping mechanisms of ma chine tools and causing misalignment, or by getting directly into products such as food, pharmaceuticals or the oilways of engines. As a result,
Failure Modes and Effects Analysis (FMEA) Increase in Desired Performance
Disassembly
(or Increase in Applied Stress)
The second category of failure modes occurs when desired performance is within the envelope of the capability of the asset when it is first put into service, but then the desired performance increases until it falls outside the capability envelope. This causes the asset to fail in one of two ways: •
•
failures caused by dirt should be listed in the FMEA whenever they are likely to interfere with a significant function of the asset.
61
the desired performance rises until the asset can no longer deliver it, or the increase in stress causes deterioration to accelerate to the extent that the asset becomes so unreliable that it is effectively useless.
An example of the first case occurs if the users of the pump shown in Figure 2 . 1 were t o increase the offtake from the tank t o
1 050 litres per minute. Under these
circumstances, the pump is unable to keep the tank full. (Note that in this case,
If components fall off machines, assemblies fall apart or whole machines
the users are not forcing the pump to work any faster - they have simply opened
come adrift, the consequences are usually very serious, so the relevant
a valve a bit wider somewhere else in the system.)
failure modes should be listed. These are usually the failure of welds, sol dered joints or rivets due to fatigue or conosion, or the failure of threaded components such as bolts, electrical connections or pipefittings which can also fail due to fatigue or corrosion, or which simply come undone. Also take care to record the functions and associated failure modes of locking mechanisms such as split pins and lock nuts when considering the integrity of assemblies.
The second case occurs for instance if the owner of a motor car whose engine is 'red-lined' at 6 000 rpm persists in revving the motor to 7 000 rpm. This causes the engine to deteriorate more quickly than if the user keeps the revs within the prescribed limits, so it fails more often.
This phenomenon is illustrated in Fig ure 4.4. It occurs for four reasons, the first three of which embody some kind of human enor: •
sustained, deliberate overloading
Human errors which reduce capability
•
sustained, unintentional overloading
The final subset of the 'falling capability' category of failure modes are
•
those caused by human enor. As the name implies, these refer to enors
sudden, unintentional overloading
•
inconect process material.
which reduce the capability of the process to the extent that it is unable to function as required by the user.
Sustained, deliberate overloading
Examples include manually operated valves left shut causing a process to be
In many industries, users quickly give
unable to start, parts incorrectly fitted by maintenance craftsmen or sensors set
in to the temptation simply to speed
in such a way that a machine trips out when nothing is wrong.
If failure modes of this type are known to occur, they should be recorded
up equipment in response to increased demand for existing products. In other
w o z « ::2! a: o u.. a: w 0... ........ .. .. _ ........ .. .... .. _ .. ........ .. .... .. .. Figure 4.4:
Failure Mode Category 2
in the FMEA so that appropriate failure management decisions can be
cases, people use assets acquired for
made later in the process. However, when listing failure modes caused by
one product to process a product with different characteristics (such as
people, take care simply to record what went wrong and not who caused
larger, heavier unit sizes or higher quality standards). People do this in the
it. If too much emphasis is placed on 'who' at this stage, the analysis could
belief that they will get more out of their facilities without any increase
become unnecessarily adversarial, and people begin to lose sight of the
in capital investment. This may even be true in the short term. I lowever,
fact that it is an exercise in avoiding or solving problems, not attaching
this solution canies long-term penalties in terms of reduced reliability
blame. For instance, it is enough to say 'control valve set too high', not
and/or availability, especially when the increased stresses begin to approach
'control valve incorrectly set by instrument technician' .
or exceed the ability of the asset to withstand them.
62
Reliability-centred Maintenance
Failure Modes and Effects Analysis (FMt.,A)
(This phenomenon causes some of the most ferocious disputes between
63
A production line with 1 2 operations and supplied with four services
maintenance and operations people. When it occurs, operations people tend to claim that 'there must be something wrong with our maintenance', while maintenance accuses operations of 'flogging the machine to death'. These disputes occur because operations people usually focus on what they want
New
out of each asset, while maintenance people tend to think in terms of what it can do. Neither of them are 'wrong' - they are simply considering the problem from two different points of view.) In these cases, implementing 'better' maintenance procedures will do little or nothing to solve the problem. In fact, maintaining a machine which
Now
400
<Jl (,) c '" "O E Q) �
300
'� .g
<Jl <Jl Cl o.
cannot deliver the desired performance has been likened to rearranging the deck chairs on the
Titanic. In such
cases we need to look beyond mainte
nance for solutions. The two options are to modify the asset to improve its inherent capability, or to lower our expectations and operate the machine within its existing capabilities.
200 100 0
Figure 4.5: The Destabilising Impact of Debottlenecking'
(Some industrial organisations have found that despite the best efforts of their engineers, debottlenecking usually causes so much instability that it is forbidden in all but the most tightly controlled and heavily restricted
circumstances. In these cases, growth is handled by allowing for it in the design of the original plant and/or by building new plants. )
Sustained, unintentional overloading Many industries respond to increased demand by undertaking formal 'debottlenecking' programs. These programs entail increasing the capa bility of a production facility - such as a production line
to accommo
Sudden, unintentional overloading Many failures are caused by sudden and (usually) unintentional increases
date a new level of desired performance. However, much to the chagrin
in applied stress, usually caused in turn by one of the following:
of their sponsors, these programs often seem to end up causing more prob
•
lems than they solve. This usually happens because a few small subsys tems or components get left out of the overall upgrade program, with sometimes devastating results. How this occurs is illustrated in Figure 4.5.
•
•
Demand for the products produced by the facility illustrated in the example has increased to the extent that its users wish to increase output from 400 to 500 tons per week. The dotted lines represent the capability of each operation. They show that most of the operations are already capable of meeting the new requirement. However, operations 3, 8 and 1 0 are capable of less than 500 tons, so they are the 'bottlenecks'. To achieve the new target, the users 'debottleneck' these operations by installing new machines or components which are capable of producing well over
500
tons per week. They also upgrade the power supplies to match.
However, in this example the need to upgrade the instrument air supply was
incorrect operation (for instance, if a machine is put into reverse while moving forward) incorrect assembly (for instance, overtorquing a bolt) external damage (for instance, if a fork lift truck smashes into a pump or lightning strikes a poorly protected electrical installation).
These are not actually increases in desired performance, because no-one wants the operator to put the machine into reverse at the wrong moment
or the fork lift to smash the pump. However, they belong in this category because applied stress rises above the ability of the asset to withstand it. If any of these failure modes are thought to be reasonably likely in the context under consideration, they should be incorporated in the FMEA.
overlooked, so the plant begins to suffer intermittent instrument problems when demand for instrument air is at a maximum. (Note also that although the unchanged operations were already capable of more than 500 tons, their margin for deteriora tion is reduced by the upgrade program, so they also begin to fail more often.)
Clearly, if a plant is suffering from failure modes of this type, they should be recorded in the FMEA so that they can be dealt with appropriately.
Incorrect process or packaging materials Manufacturing processes often suffer functional failures caused by pro cess materials which are out of specification (in terms of such variables as consistency, hardness or pH). Similarly, packaging plants often suffer from inadequate or incompatible packaging materials.
64
Failure Modes and Effects Analysis ( FMbA)
Reliability-centred Maintenance
In both cases, the machines fail or run badly because they cannot handle the out-of-spec material. This can be seen as an increase in applied stress. In practice, these 'failure modes' are seldom the result of a failure of the
asset under review, but are nearly always the effect of a failure elsewhere in the system. This means that remedial action has to be applied to a differ
ent asset. However, acknowledging these failures in the analysis of the affected asset helps to ensure that they will receive attention when the system which is really causing the problem is analysed. As a result, these failure modes should be incorporated in the FMEA where they are known
to affect the asset under review, with a comment in the failure effects column which directs attention to the real source of the problem. Initial incapability Chapter 2 explained that for any asset to be maintainable, its desired performance must fall within the envelope of its initial capability. It went on to mention that the majority of assets are in fact built this way. How
ever, situations do arise where desired perfonnance is outside the envelope of the initial capability right from the outset, as shown in Figure 4.6. This incapability problem seldom affects entire assets. It usually affects just one or two functions of one or two components, but these weak links upset the operation of the whole chain. The first step towards rectifying design problems of this nature is to list them as failure modes in an FMEA. Figure 4.6: Failure Mode Category 3
In practice, it can be surprisingly difficult to find an appropriate level of detail. However, it is important to do so, because the level or detail pro foundly affects the validity of the FMEA and the amount of time needed to do it. Too little detail and/or too few failure modes lead to superficial and sometimes dangerous analyses. Too many failure modes and/or too much detail causes the entire RCM proeess to take much longer than it needs to. In extreme cases, excessive detail can cause the process to take two or even three times longer than necessary (a phenomenon known as
analysis paralysis). This means that it is essential to try to strike the right balance. Some of the key factors which need to be taken into account are discussed in the following paragraphs. Causation The causes of any functional failure can be defined to almost any level of detail, and different levels are appropriate in different situations. At one extreme, it is sometimes enough to summarise the causes of a functional failure in one statement, such as 'machine fails' . At the other, we may need to consider what goes wrong at the molecular level ,md/or explore the remoter corners of the psyche of the operators and maintainers in a bid to define so-called root causes of failure. The extent to which failure modes can be described at different levels
i
LU () z « � ex: o u..
4.4 How Mu ch Detail?
65
ex: LU a..
Earlier in this chapter, it was mentioned that failure modes should be de scribed in enough detail for it to be possible to select an appropriate failure management strategy, but not in so much detail that excessive amounts of time are wasted on the analysis process itself.
Failure modes should be defined in enough detailfor it to be possible to select a suitable failure management policy
of detail is illustrated in Figure 4.7 on the next three pages. Figure 4.7 is based on the pump set shown in Figure 4.2, some of whose failure modes were listed in Figure 4.1. Figure 4.7 lists ways in which the pump set might suffer from the functional failure 'unable to transfer any water at all'. These failure modes are considered at seven different levels of detail. The top level (Level
1 ) is failure of the pump set as a whole. Level 2 recognises
the failure of the five major components of the pump set - the pump, the drive shaft, the motor, the switchgear and inlet/outlet. Thereafter failures are considered in progressively more detail. When considering this example, please note that •
levels have been defined and failure modes allocated to each level for the purpose of this example only. They are not any kind of universal classification.
•
Figure 4.7 does not show all failure possibilities at each level so don't use this example as a definitive model
•
it is possible to analyse some of the failure modes at even lower levels than level
•
the failure modes listed only apply to the functional failure 'unable to transfer
7,
but it would very seldom be necessary to do so in practice
water at all'. Figure 4. 7 does not show failure modes which would cause other functional failures, such as loss of containment or loss of protection.
0\ 0\
LEVEL 4 Impeller comes adrift
LEVEL 1 Pump set fails
�
� \:l
[ ,�. �
"" ;::s
�
\:l..
�
raslngrlJpturea·Casing b()lts come100se
� � (i; :3 o
S· ;;, See Appendix 2
Assembly error
�
�
9:
iil' (i; "l. Cii
Pump seaT fails
(§�
0.'"
�:L.. til � :::::.: -.
Normal wear and tear -Pump runs dry Seal misaligned Seal faces dirty Wrong seal fitted Damaged seal installed
Operating error -Pump in vulnerable position -S"m"-a::Cs:Ch-:CedTL : byc-ocbC-:je�c"t fc c ro:C:m:C-sTkC Cy--"'Casing hit by meteorite Cilslng hit by part of aircraft Seal abraded
See "wilier supply fails" below
Assembly error Assembly error Wrong seal supplied
Wrong seal specified· Pump seal dropped in stores Pump seal damaged in transit
-SeeAppendix'1 See Appendix 2
;::s \:l ;::s
" '"
See Appendix 2
Design error
________
Procurement error Storekeeping error -Design error Storekeeping error Procurement error
See Appendix 2 See Appendix 2 See Appendix 2 --; pp"- e:C:n-:;dl�x "2See AC:: See Appendix 2
LEVEL 1 Pump set fails
� � .::
�
� �
� � til
"" \:l
� '� "'.
� � til
-� 9:�'
��
(i; '" :::0
""
::�
<1l " <1l
,.., c>
____
til::;:
.... . :::. """" -
o
0.;:: <1l ""
��
" ...., ""
::0;::s �
.::;� c.-:
�
�
� .... .. 0\ --J
Failure Modes and Effects Analysis (F'MEA)
Reliability-centred Maintenance
68
I I
I II I
I�
\-0
The first point to emerge from this example is the connection between the level of detail and the number of failure modes listed. The example shows the further one 'drills down' in an FMEA, the larger the number of failure modes that can be listed.
�jlli1J �T'" i
I galig � �-
°
'-£��
rei
::oE�g? ctS 010 E (I,)"'O<..) () U) if>
�
Q)
2:
«: Q) "" '" JD .J:: if> O c:
m -, 0. wi!! '"
if>
� i:75
�
Root causes
113�'o" .f; B .8
The term 'root cause' is often used in connection with the analysis of fail ures. It implies that if one drills down far enough, it is possible to anive at a final and absolute level of causation. In fact, this is seldom the case.
":Lsl'::: 1-0 I 11
For instance, in Figure 4.7 the failure mode 'impeller nut overtightened' is listed at level 6, which in turn i s caused by an 'assembly error' at level 7. If we were to go down one level further, the assembly error might have occurred because the 'fitter was distracted' (level 8). He might have been distracted because his 'child was ill' (level 9). This failure might have occurred because the 'child ate bad food in restaurant' (level 10).
�
I�
-'if> $ if> "§
'"
�
g l'l � 21
I
:>'" w
�i-0
J
Two more key issues which arise from Figure 4.7 concern 'root causes' and human enor. They are discussed below.
��� �� 0g 8-88
�
,.g -'.E'
4.7 but 64 at level 6.
x
�< ..§�I8.�al� I
::l
For instance, there are five failure modes listed at level 2 forthe pump set i n Figure
°1°10
<'(co co aJ
l ,-�m2:8(ij I
""
'"
XIX
�0 Q
$ �
�
0 "11:1:
�
0
°5
0if> ::J
°
";:: ::> a. (f)
(!) if>
.2
-0
'0 � ::l
'" 0> '"
$
E ro
0 O§ if> ::>
0. (f) .91 .Q '"
U
$
0:;,;
£L
Ij
Ii !� I� I� !B ]1
c
8
�
c: 0 (.)
69
21
:£§ $
03:0-
0> c:
"E
0
u E:
Figure 4.7 (continued):
Failure modes at different levels of detail
Clearly, this process of drilling down could go on almost forever way beyond the point at which the organisation doing the FMEA has any con trol over the failure modes. This is why this chapter stresses repeatedly that the level at which any failure mode should be identified is the level at which it is possible to identify an appropriate failure management policy. (This is equally true whether one is canyillg out an FMEA before failures occur or a 'root cause analysis' after a failure has oecurred.) The fact that the level which is appropriate varies for different failure modes means that we do not have to list all failure modes at the same level on the lnfonnation Worksheet. Some failure modes might be identified at level 2, others at level 7, and the rest somewhere in between. For instance, in one particular context, it may be appropriate to l ist only those failure modes shaded in grey in Figure 4.7. In another context, it may be appro priate for an entire FMEA for an identical pump set to consist of the single failure mode 'pump set fails'. Another context may call for yet another selection.
Obviously,in order to be able to stop at an appropriate level, the people doing such analyses need to be aware of of the full range of t�lilure manage ment policy options. These are diseussed at length in Chapters 6 to 9. Other factors which influence the level of detail are considered in the rest of this part of this chapter and again in part 7.
70
Failure Modes and Effects Analysis (fME'\)
Reliability-centred Maintenance
However, a review of existing schedules should only be carried out as a final check after the rest of the ReM analysis has been completed in order to reduce the possibility of perpetuating the status quo. (Some users of ReM are tempted to assume that all reasonably likely failure modes are covered by their existing PM systems, and hence that these are the only failure modes which need to be considered in the FMEA. This assumption leads these users to develop the entire FMEA by working backwards from their existing maintenance schedules, and then work ing forwards again through the last three steps of the ReM process. This approach is usually adopted in the belief that it will speed up or" stream line" the process. In fact this approach is notrecommended because among other shortcomings, it leads to dangerously incomplete ReM analyses.)
Human error
Part 3 of this chapter mentioned a number of general ways in which human error could cause machines to fail. It went on to suggest that if the associated failure modes are thought to be reasonably likely, they should be incorporated in the FMEA. This has been done in Figure 4.7, where all the failure modes ending with the word 'error' are some form of human error. Appendix 2 provides a brief summary of key issues involved in the classification and management of such errors. Probability
Different failure modes occur at different frequencies. Some may occur regularly, at average intervals measured in months, weeks or even days. Others may be extremely improbable, with mean times between occur rences measured in millions of years. When preparing an FMEA, deci sions must be made continuously as to what failure modes are so unlikely that they can safely be ignored. This means that we do not try to list every single failure possibility regardless of its likelihood. When listing failure modes, do not try to list every s ingle failure possibility regardless of its likelihood
Only failure modes which might reasonably be expected to occur in the context in question should be recorded. A list of 'reasonably likely' fail ure modes should include the following: •
failures which have occurred before on the same or similar assets.
These are the most obvious candidates for inclusion in an FMEA un less the asset has been modified so that the failure cannot occur again. As discussed later, sources of information about these failures include people who know the asset well (your own employees, vendors or other users of the same equipment), technical history records and data banks. In this context note the comments in Part 6 of this chapter about the shortcomings of most technical history records,and in Chapter 12 about the danger of too much reliance on historical data.
•
failure modes which are already the subject ofproactive maintenance routines, and so which would occur if no proactive maintenance was being done. One way to ensure that none of these failure modes have been overlooked is to study existing maintenance schedules and ask "what failure mode would occur if we did not do this task?"
71
•
any otherfailure modes which have not yet occurred but which are con sidered to be real possibilities. Identifying and deciding how to deal with failures which have not happened yet is an essential feature of pro active management in general and of risk management in particular. It is also one of the most challenging aspects of the ReM process, because it calls for a high degree of judgement. On the one hand, we need to list all reasonably likely failure modes, while on the other we don't want to waste time on failures which have never occurred before and which are extremely unlikely (incredible) in the context in question. For example, 'sealed-for-life' bearings are installed on the motor driving the pump shown in Figure 4.7. This means that the likelihoOd of lubrication failure is low so low that it would not be included in most FMEA's. On the other hand failure due to lack of lubricant probably would be inCluded in FMEA's prepare for manually lubricated components, centralised lube systems and gearboxes.
d
However, the decision not to list a failure mode should be tempered by careful consideration of the failure consequences. Consequences
If the consequences are likely to be very severe indeed, then less likely failure possibilities should be listed and subjected to further analysis. For instance, if the pump set in Figure 4.7 was installed in a food factory or a vehicle assembly plant, the failure mode 'casing smashed by an object falling from the sky' would be dismissed immediately as being laughably unlikely. However, if the pump were pumping something really nasty in a nuclear installation , this failure mode is more likely to be taken seriously even though it is still highly improbable. (Appropriate failure management policies might be either to ban aircraft from flying over the facility, or to design a roof Which can withstand a crashing aircraft.)
72
Reliability-centred Maintenance
Failure Modes and Effects Analysis (FMEA)
Another example from Figure 4.7 is 'motor not switched on'. This failure mode is likely to be dismissed on the grounds of improbability in most situations. Even if it does occur, the consequences may be so trivial that it is excluded from the FMEA. (On the other hand, if it could occur and it does matter- especially in cases where things must be switched on in a particular sequence and something could be damaged if they are not then this failure mode should be considered.)
Similarly, a vehicle operating in the Arctic would be subject to different failure modes from the same make of vehicle operating in the Sahara desert. Similarly, a gas turbine powering a jet aircraft would have different failure modes from the same type of turbine acting as a prime mover on an oil platform.
These differences mean that great care should be taken to ensure that the operating context is identical before applying an FMEA developed in one set of circumstances to an asset which is used in another. (Note also the comments regarding the use of generic FMEA's in part 6 of this chapter.) The operating context affects levels of analysis as well as the causes and consequences of failure. As discussed earlier, it might be appropriate to identify failure modes for two identical assets at one level in one opera ting context and at another level in another.
Cause vs Effect
Care should be taken not to confuse causes and effects when listing failure modes. This is a subtle mistake most often made by people who are new to the RCM process. For example, one plant had some 200 gearboxes, all of the same d�sign and all . performing more or less the same function on the same type of eqUipment. Initi ally, the following failure modes were recorded for one of these gearboxes: Gearbox bearings seize • Gear teeth stripped. These failure modes were listed to begin with because the people carrying out the review recalled that each failure had happened in the past to their knowledge (some of the gearboxes were twenty years old). The failures did not affect safety but they did affect production. So the implication was that it might be worth dom? preventive tasks like 'check gear teeth for wear' or 'check gearbox f?r backlash , . and 'check gearbox bearings for vibration'. However, further diSCUSSion revealed that both failures had occurred because the oil level had not been checked when it should have been, so the gearboxes had actually failed due to lack of oil. What is more, no-one could recall that any of the gearboxes had failed if they had been properly lubricated. As a result, the failure mode was eventually recorded as: Gearbox fails due to lack of oil. This underlined the importance of the obvious proactive task, which was to check the oil level periodically. (This is not to suggest that all gearboxes should be ana lysed in this way. Some are much more complex or much more heavily loa?ed, and so are subject to a wider variety of failure modes. In other cases, the failure consequences may be much more severe, which would call for a more defenSive view of failure possibilities.)
73
•
•
Failure Modes and the Operating Context
4.5 Failure Effects The fourth step in the RCM review process entails listing what happens when each failure mode occurs. These are known as fai/u re effects. Failure effects describe what happens when a fa ilure mode occurs
(Note that failure effects are not the same as failure consequences. A failure effect answers the question "what happens?", whereas a failure consequence answers the question '�(how) does it matter?".) A description of failure etIects should include all the infomlation needed to support the evaluation of the consequences of the failure. Specifically, when describing the effects of a failure, the following should be recorded: •
•
•
what evidence (if any) that the failure has OCCUlTed in what ways (if any) it poses a threat to safety or the environment in what ways (if any) it affects production or operations
We have seen how the functions and functional failures of any item are influenced by its operating context. This is also true of failure modes in terms of causation, probability and consequences.
•
For example, consider the three pumps shown in Figure 2.7. The failure modes which are likely to affect the stand-by pump (such as brinelling of the beanngs, stagnation of water in the pump casing and even the 'borrowing' of key compo nents to use elsewhere in an emergency) are different from those which might affect the duty pump, as set out in Figure 4.7.
These issues arc reviewed in the following paragraphs. Note that one of the objectives of this exercise is to establish whether proactive maintenance is necessary. If we are to do this correct! y, we cannot assume that some sort of proactive maintenance is being done already, so the effects of a failure should be described as if nothing was being done to prevent it.
•
what physical damage (if any) is caused by the failure what must be done to repair the failure.
_
z
74
Reliability-centred Maintenance
Failure Modes and Effects Analysis (FNIEA)
75
Evidence of Failure
Safety and Environmental Hazards
Failure effects should be described in a way which enables the team doing the RCM analysis to decide whether the failure will become evident to the operating crew under normal circumstances.
Modern industrial plant design has evolved to the point that only a small proportion of failure modes present a direct threat to safety or the environ ment. However, if there is a possibility that someone could get hurt or killed as a direct result of the failure, or an environmental standard or regulation could be breached, the failure etTect should describe how this could happen. Examples include: • increased risk of fire or explosions the escape of hazardous chemicals (gases, liquids or solids) electrocution • falling objects • pressure bursts (especially pressure vessels and hydraulic systems)
For instance, the description should state whether the failure causes warning lights to come on or alarms to sound (or both), and whether the warning is given on a local panel or in a central control room (or both).
Similarly, the description should state whether the failure is accompanied (or preceded) by obvious physical effects such as loud noises, fire, smoke, escaping steam, unusual smells, or pools of liquid on the floor. It should also state whether the machine shuts down as a result of the failure. For example, if we are considering the seizure of the bearings of the pump shown in Figure 3.5, the failure effects might be described as follows (the italics describe what would make it evident to the operators that a failure has occurred): Motor trips out and trip alarm sounds in the control room. Tank Y low level alarm sounds after 20 minutes, and tank runs dry after 30 minutes. Downtime required to replace the bearings 4 hours. In the case of a stationary gas turbine, a failure mode that occurred in practice was the gradual build up of combustion deposits on the compressor blades. These deposits could be partially removed by the periodic injection of special materials into the air stream, a process known as 'jet blasting' The failure effects were de scribed accordingly as follows: Compressor efficiency declines and governor compensates to sustain power output, causing exhaust temperature to rise. Exhaust temperature is displayed on the local control panel and in the central control room. If no action is taken, exhaust gas temperature rises above 475°C under full power. A high exhaust •
•
gas temperature alarm annunciates on the local control panel and a warning
Above 500°C, the control system (Running at temperatures above 475°C shortens the creep life of the turbine blades.) The blades can be partially cleaned by jet blasting, and jet blasting takes about 30 minutes. light comes on in the central control room. shuts down the turbine.
This is an unusually complex failure mode, so the description of the fail ure effects is somewhat longer than usual. The average description of a failure effect usually amounts to between twenty and sixty words. When describing failure effects, do not prejudge the evaluation of the failure consequences by using the words 'hidden' or 'evident' They are part of the consequence evaluation process, and using them prematurely could bias this evaluation incorrectly. Finally, when dealing with protective devices, failure effect descrip tions should state briefly what would happen if the protected device were to fail while the protective device was unserviceable.
•
•
•
•
•
•
•
•
•
exposure to very hot or molten materials the disintegration of large rotating components vehicle accidents or derailments
exposure to sharp edges or moving machinery increased noise levels the collapse of structures the growth of bacteria
ingress of dirt into food or pharmaceutical products t1ooding . When listing these effects, do not make qualitative statements like "this failure has safety consequences" or '�this failure atTects the environment". Simply state what happens, and leave the evaluation of the consequences to the next stage of the RCM process. Note also that we are not only concerned about possible threats to our own staff (operators and maintainers), but also about threats to the safety of customers and the community as a whole. This may call for some research by the team doing the analysis into the environmental and safety standards which govern the process under review. •
•
Secondary Damage and I)roduction Effects
Failure effect descriptions should also help with decisions about opera tional and non-operational failure consequences. To do so, they should indicate how production is affected (if at all), and for how long. This is usually given by the amount of downtime associated with each failure.
Reliability-centred Maintenance
76
Failure Modes and Effects Analysis (l'MEA)
In this context, downtime means the total amount of time the asset would normally be out of service owing to this failure, from the moment it fails until the moment it is fully operational again. As indicated in Figure 4.8, this is usually much longer than the repair time. OIl
Macb�QfJ
stops{� ..'
Figure 4.8.
Find the person who can repair it
DOWNTIME Diagnose the fault
Find the spare parts
Repair the fault
Revalidate or test the machine
,.,u�tl't�
m�chine bact< into service
--..- REPAIR ..-TIME
77
In addition to downtime, any other ways in which the failure could have a significant effect on the operational capability of the asset should be listed. Possibilities include: •
whether and how product quality or customer service is affected, and if so whether any financial penalties are involved
•
whether any other equipment or activity also has to stop (or slow down)
•
whether the failure leads to an increase in overall operating costs in addition to the direct cost of repair (such as higher energy costs)
•
what secondary damage (if any) is caused by the failure .
Corrective Action
Downtime vs repair time
Downtime as defined above can vary greatly for different occurrences of the same failure, and the most serious consequences are usually caused by the longer outages. Since it is consequences which are of most interest to LIS, the downtime recorded on the information worksheet should be based on the 'typical worst case'. For instance, if the downtime caused by a failure which occurs late on a weekend night shift is usually much longer than it is when the failure occurs on a normal day shift, and if such night shifts are a regular occurrence, we list the former.
It is of course possible to reduce the operational consequences of a failure by taking steps to shorten the downtime, most often by reducing the time it takes to gel hold of a spare part. However, as discussed in Chapter 2, we are still in the process of defining the problem at this stage so the analysis should be based at least initially on current spares holding policies. Note that if the failure affects operations, it is more important to record downtime than the mean time to repair the failure (MTIR), for two reasons: •
in many people's minds, the word 'repair time' has the meaning shown in Figure 4.8. If this is used instead of downtime, it could upset the sub sequent assessment of the operational consequences of failure
•
we should base the assessment of consequences on the 'typical worst case' and not the 'mean' , as discussed above.
If the failure does not cause a process stoppage, then the average amount of time it takes to repair the failure should be recorded, because this can be used help establish manpower requirements.
Failure effects should also state what must be done to repair the failure. This can be included in the statement about downtime, as shown in italics in the following examples: •
• •
Downtime to replace bearings about four hours Downtime to clear the blockage and reset the trip switch about 30 minutes Downtime to strip the turbine and replace the disc about 2 weeks.
4.6 Sources of Information about Modes and Effects When considering where to get information needed to draw up a reason ably comprehensive FMEA, remember the need to be proactive. This means that as much emphasis should be placed on what could happen as on what has happened. The most common sources of information arc dis cussed in the following paragraphs, together with a brief review of their main advantages and disadvantages.
The manufacturer or vendor of the equipment When carrying out an FMEA, the source of information which usually springs to mind fIrst is the manufacturer. This is especially so in the case of new equipment. In some industries, this has reached the point where manu facturers or vendors are routinely asked to provide a comprehensive FMEA as part of the equipment supply contract. Apart from anything else, this request implies that manufacturers know everything that needs to be known about how the equipment can fail and what happens when this occurs. This is seldom the case in reality.
78
Reliability-centred Maintenance
In practice, few manufacturers are involved in the day-to-day opera tion of the equipment. After the end of the warranty period, almost none get regular feedback from the equipment users about what fails and why. The best that many of them can do is try to draw conclusions about how their equipment is performing from a combination of anecdotal evidence and an analysis of spares sales (except when a really spectacular failure occurs, in which case lawyers tend to take over from engineers. At this point, rational technical discussion about root causes often ceases.) Manut�'lcturers also have little access to infoffi1ation about the operat ing context of the equipment, desired standards of perfonnance, failure consequences and the skills of the user's operators and maintainers. More often the manufacturers know nothing about these issues. As a result, FMEA's compiled by these manufacturers are usually generic and often highly speCUlative, which greatly limits their value. The small minority of equipment manufacturers who are able to produce a satisfactory FMEA on their own usually fall into one of two categories: •
•
they are involved in maintaining the equipment throughout its useful life, either directly or through closely assoeiated vendors. For instance most privately-owned motor vehicles are maintained by the dealers who sold the vehicles. This enables the dealers to provide the manufac turers with copious failure data. they are paid to carry out formal reliability studies on prototypes as part of the initial procurement process. This is a common feature of military procurement, but much less common in industry.
In most cases, the author has found that the best way to access whatever knowledge manufacturers possess about the behaviour of the equipment in the field is to ask them to supply experienced field technicians to work alongside the people who will eventually operate and maintain the asset, to develop FMEA's which are satisfactory to both parties. If this sugges tion is adopted, the field technicians should of course have unrestricted access to specialist support to help them answer difficult questions. When adopting this approach, issues such as warranties, copyrights, languages which the participants should be able to speak fluently, techni cal support, confidentiality, and so on should be handled at the contract ing stage, so that everyone knows clearly what to expect of each other. Note the suggestion to use field technicians rather than designers. De signers are often surprisingly reluctant to admit that their designs can fail, which reduces their ability to help develop a sensible FMEA.
Failure Modes and Effe cts Analysis ( FMEA)
79
Generic lists offai/Lire modes
�
'G neric' lists of failure modes are lists of failure modes or sometimes entlre FMEA's - prepared by third parties. They may cover entire systems , but more often cover individual assets or even single components. These generic lists are touted as another method of speeding up or 'streamlinina' this part of the maintenance program development process. In fact, th y should be approached with great caution, for the following reasons:
:
�
•
th level ofanalysis may be inappropriate: A generic list may identify . fmlure modes at a level equivalent to (say) level S in Figure 4.7, when all that may be needed is level I. This means that far from streamlining the process, the generic list would condemn the user to analysin g far more failure modes than necessary. Conversely, the generic list may focus on level 3 or 4 in a situation where some of the failure modes really ought to be analysed at level 5 or 6.
•
the operating contextmay be diff erent: The operating context of your asset may have features which make it susceptible to failure modes that do not �ppe r in the generic list. Conversely, some of the modes in the genede 11st J11lght be extremely improbable (if not impossible) in your context.
•
performance standards may dijfer: your asset may operate to standards of
�
performance which mean that your whole definition of failure may be completely difterent from that used to develop the generic FMEA.
These three points mean that if a generic list of failure modes is used at all, it should only ever be used to supplement a context-specific FMEA, and never used on its own as a definitive list.
Other users of the same equipment Other users are an obvious and very valuable source of information about what can go wrong with commonly used assets, provided of course that competitive pressures permit the exchange of data. This is often done through industry associations (as in the offshore oil industry), throuah regulatory bodies (as in civil aviation) or between different branches f the same organisation. However, note the above comments about the dan gers of generic data when considering these sourees of information.
�
Technical history records Technical history records can also be a valuable source of information. However, they should be treated with caution for the following reasons:
80
Reliability-centred Maintenance
•
they are often incomplete
•
more often than not, they describe what was done to repair the failure ('replaced main bearing') rather than what caused it
•
they do not describe failures which have not yet occurred
•
they often describe failure modes which are really the effect of some other failure.
These drawbacks mean that teehnical history reeords should only be used as a supplementary source of information when preparing an FMEA, and never as the sole source.
The people who operate and maintain the equipment Tn nearly all cases, by far the best sources of information for preparing an FMEA are the people who operate and maintain the equipment on a day to-day basis. They tend to know the most about how the equipment works, what goes wrong with it, how much each failure matters and what must be done to fix it and if they don't know, they are the ones who have the most reason to find out. The best way to capture and to build on their knowledge is to arrange for them to pmticipate formally in the preparation of the FMEA as part of the overall ReM process. The most efficient way to do this is under the guidance of a suitably trained facilitator at a series of meetings. (The most valuable source of additional information at these meetings is a compre hensive set of process and instrumentation drawings, coupled with ready access to process and/or technical specialists on an ad hoc basis.) This approach to ReM was introduced in Chapter 1 and is discussed at much greater length in Chapter 13.
Failure Modes and Effec ts Analysis (FMEA)
81
The detail used to describe failure modes on Information Worksheets is also influenced by the level at which the FMEA as a whole is carried out. This in tum is governed by the level at which the entire ReM analysis is performed. For this reason, we review the principal factors which influence the overall level of analysis (which is also known as 'level of indenture') before considering how this affects the detail with which failure modes should be described. Level of analysis
ReM is defined as a process used to determine what must be done to ensure that any physical asset continues to do whatever its users want it to do in its present operating context. In the light of this definition, we have seen that it is necessary to define the context in detail before we can apply the process. However, we also need to define exactly what the 'physical asset' is to which the process will be applied. For example, if we apply ReM to a truck, is the entire truck the 'asset'? Or do we subdivide the truck and analyse (say) the drive train separately from the braking system, the steering, the chassis and so on? Or should we further subdivide the drive train and analyse (say) the engine separately from the gearbox, propshaft, differentials, axles and wheels? Or should the engine not be divided into engine block, engine management system, cooling system, fuel system and so on before s!arting the analysis? What about subdividing the fuel system into tank, pump, pipes and filters?
This issue needs careful thought because an analysis carried out at too high a level becomes too superficial, while one done at too low a level can become unmanageable and unintelligible. The following paragraphs ex plore the implications of carrying out the analysis at different levels.
Starting at a low level
4.7 Levels of Analysis and the Information Worksheet Part 4 of this chapter showed how failure modes can be described at almost any level of detail. The level of detail which is ultimately selected should enable a suitable failure management policy to be identified. In general, higher levels (less detail) should be selected if the component or sub-system is likely to be allowed to run to failure or subject to failure finding, while lower levels (more detail) need to be selected if the failure mode is likely to be subjected to some sort of proactive maintenance.
One of the most common mistakes in the ReM process is carrying out the analysis at too Iow a level in the equipment hierarchy. For example, when thinking about the failure modes which could affect a motor vehicle, a possibility which comes to mind is a blocked fuel line. The fuel line is part of the fuel system, so it seems sensible to address this failure mode by raising a Worksheet for the fuel system. Figure 4.9 indicates that if the analysis is carried out at this level, the blocked fuel line might be the seventh failure mode to be identified out of a total of perhaps a dozen which could cause the functional failure 'unable to transfer any fuel at all'.
82
Reliability-centred Maintenance
RCMII INFORMATION WORKSHEET
© 1996 ALADON LTD 1
SYSTEM
�:<W
To transfer fuel from the fuel tank to the engine at a fate of up to 1 litre per minute
Failure Modes and Effects Analysis (FMEA)
A
Figure 4.9:
If special attention is not paid to this issue, the same function ends up being analysed three times in slightly different ways, and the same failure-finding task prescribed more than once for the same loop.
:Fue[ system No fuel in tank
Unable to transfer any fuel at all
Fuel filter blocked Fuel line blocked by foreign object
a new worksheet has to be raised for each new sub-system. This leads to the generation of vast quantities of paperwork for the analysis of the entire vehicle, or the consumption of equally large amounts of compu ter memory space. The associated manual or electronic filing systems have to be very carefully structured if the information is to remain manageable. In short,the whole exercise starts to become much more extensive and much more intimidating than it needs to be.
, .. etc
Failure modes of a fuel system
system, the following problems begin to arise:
the further down the hierarchy one progresses, the more diffieult it be comes to conceptualise and define performance standards. (One could also ask who actually cares about the precise amount of fuel passing through the fuel system as long as the fuel economy of the vehicle is within reasonable limits and the vehicle has enough power.)
•
at a low level,it becomes cqually difficult to visualise and hence to ana lyse failure consequences.
•
the lower the level of the analysis, the more difficult it becomes to decide which components belong to which system (for instance, is the accelerator part of the fuel system or the engine control system?)
•
some failure modes can cause many sub-systems to cease to function simultaneously (such as a failure in the supply of power to an industrial plant). If each sub-system is analysed on its own, failure modes of this type are repeated again and again.
•
•
12 Fuel line severed
When the decision worksheet has been completed for this sub-system,the RCM review group proceeds to the next system, and so on until the main tenance requirements of the entire vehicle have been reviewed. This seems to be straightforward enough until we consider that the vehicle can actually be sub-divided into literally dozens if not hundreds of sub systems at this level. If a separate analysis is carried out for each sub •
83
control and protective loops can become very difficult to deal with in a low-level analysis, especially when a sensor in one sub-system drives an actuator in another through a processor in a third. For instance, a rev limiter which reads a Signal off the flywheel in the 'engine block' SUb-system might send a signal through a processor in the 'engine control' SUb-system to a fuel shut-off valve in the 'fuel' sub-system.
FMEA's are often carried out at too Iow a level in the equipment hierarchy because of a belief that there is a correlation between the level at which we identify failure modes and the level at which the FMEA (or the RCM analysis as a whole) should be performed. In other words, it is often said that if we want to identify failure modes in detail, then we ought to carry out a separate FMEA for each replaceable component or sub-system. In fact, this is not so. The level at which failure modes can be identified is independent of the level at which the analysis is performed, as shown in the next section of this chapter.
Starting at the top Instead of starting the analysis towards the bottom of the equipment hier archy, we could start at the top. For example, the primary function of a truck was listed on page 28 as follows: To transport up to 40 tons of material at speeds of up to 75 mph (average 60 mph)
first functional failure asso ciated with this function is 'Unable to move at all'.The four failure modes shown in Figure 4.9 could all cause this functional failure, so instead of being listed on an Information Worksheet for the fuel system, they could have been listed on a Worksheet covering the entire truck, as shown in Figure 4.10. from Startsville to Endburg on one tank of fuel. 'The
RCMII INFORMATION WORKSHEET 1996 ALADON LTD
©
FUNCTION
1
SYSTEM
SUB·SYSTEM
To transfer up to 40 tons of
material from Startsville to I
!
Endburg speeds of up to 75
60 mph) on one
40 ton truc/(.
FUNCTIONAL FAILURE (1..0$$ of F�nctl0l!::") _-+-,_�. '-.=:.::.::..:. : ::.::�: .::: c.. " ���"-l-
A Unable to move at all
Figure
73 114
Fuel line blocked by foreign object
4.10: Failure modes of a truck
84
The main advantages of starting the analysis at this level are as follows: •
functions and performance expectations are much easier to define
•
it is easier to identify and analyse control loops and circuits as a whole
•
it is not necessary to raise a new information worksheet for each new sub-system, so analyses carried out at this level consume far less paper.
•
Failure Modes and Effe cts Analysis (FMEA)
Reliability-centred Maintenance
failure consequences are much easier to assess
40 ton truel( A
1
Unable to transfer any material at all
For instance, we have seen how the blocked fuel system might have been the seventh failure mode out of twelve to be identified in the analysis carried out at the 'fuel system' level. However, at the truck level, Figure 4.10 shows that it might have been 73rd out of several hundred failure modes.
FAILURE MODE
18 No fuel in tank 42 i I 73 114
1
there is less repetition of functions and failure modes
However, the main disadvantage of performing the analysis at this level is that there an� hundreds of failure modes which could render the truck effectively unable to move. These range from a flat front tyre to a sheared crankshaft. So if we were to try to list all the failure modes at this level, it is highly likely that several would be overlooked altogether.
8S
'13raf(Jng system
Steering system
2
object
Cabill FAILURE MOD E ---I ::: ..
9 I No fuel in tank
16 , 33 . 71
I
(jearbo?(
Propsliaft
'Diffe ren tia[s
'Engine Moel(
CooLIng system
'E;t;/wust system
3
Intermediate levels The problems associated with high- and low-level analyses suggest that it may be sensible to can'y out the analysis at an intermediate level. In fact, we are almost spoiled for choice, because most assets can be sub-divided into many levels and the RCM process applied at any one of these levels.
1
For example, Figure 4.11 shows how the 40 ton truck could be divided into at least five levels. It traces the hierarchy from the level of the truck as a whole down to the level of the fuel lines. It goes on to show how the primary function of the asset might be defined at each level on an ReM Information Worksheet, and how the blocked fuel line could be appear at each level.
Given the choice of five (sometimes more) possibilities, how do we select the level at which to pelform the analysis? We have seen that the top level usually embodies too many failure modes per function to permit sensible analysis. In spite of this however, we still need to identify the main functions of the asset or system at the highest levels in order to provide a framework for the rest of the analysis. For example an operator acquires a truck to carry goods from A to S, not to pump fuel along a fuel line. Although the latter function contributes to the former, the overall performance of the asset and hence of its maintenance - tends to be judged at the highest levels. For instance, the chief executive of a truck fleet is much more likely to ask 'how is truck X performing?' than 'how is the fuel system on truck X performing?' (unless the fuel system is known to be causing a problem).
FUNCTION
FUNCTIONAL FAILURE
To transfer fuel from the fuel tank to A Unable to transfer any the engine at a rate of up to 1 litre per fuel at all minute
'.!ue[ tanl( 5
FUNCTION
1 I To carry fuel from the fuel tank to the engine at a rate of up to 1 litre per minute
1 3
7 12
'.!uefpump
JileLjifter
Unable to carry any fuel at all
object
etc.... Figure 4.11:
Functions and failures at different levels
Chapter 2 explained that in practice. a statement of the operating context provides a record of the main functions and associated performance stan dards of any asset or system at levels above the level at which the RCM analysis is to be carried out.
86
Reliability-centred Maintenance
On the other hand, we have seen the initial inclination is nearly always to start too low in the asset hierarchy. For this reason, a good general rule (especially for people new to RCM) is to carry out the analysis one level or even two levels higher than at first seems sensible. This is because it is always easier to break complex sub-systems out of a high-level ana lysis than it is to go up a level when one has started too low. This is dis cussed in more detail in the next section of this chapter. With a bit of practice (especially concerning what is meant by 'a level at which it is possible to identify a suitable failure management policy' ), the most suitable level at which to carry out any analysis eventually be comes intuitively obvious. In this context, note that it is not necessary to analyse every system at the same level throughout the asset hierarchy. For instance, the entire brakingsystem could be analysed at level 2 as shown in Fig ure 4.11, but it may be necessary to analyse the engine at level 3 or even level 4. How Failure Modes and Effects Should be Recorded
Once the level of the entire RCM analysis has been established, we then have to decide what degree of detail is necessary to define each failure mode within the framework of that analysis. There is no technical reason why all the failure modes cannot be listed (together with their effects) at a level which enahles a suitable failure management policy to be selected. However, even intermediate level analyses sometimes generate too many failure modes per functional failure, especially for pIimary functions. This usually happens when the asset incorporates complex subassemblies which could themselves suffer from a large number of failure modes. Examples of such subassemblies include small electric motors, small hydraulic systems, small gearboxes, control loops, protectivecircuits and complex couplings.
Depending as usual on context and consequences, these sub-assemblies can he handled in one of four different ways, as discussed below.
Option 1 List all the reasonably likely failure modes of the subassembly individu ally as part of the main analysis in other words, at levels equivalent to level 3,4, 5 or 6 in Figure 4.7. For example, consider an asset which could stop completely as a result of the failure of a small gearbox. On the Information Worksheet for this asset, this gear box failure could be listed as shown below:
Failure Modes and Effects Analysis (FMEA) FAILURE MODE Gearbox bearings seize
87
FAILURE t::.t-I-t::.c I
Motor trips and alarm sounds in control room. 3 hours downtime to replace gearbox with spare. New bearings fitted in vvv f\�"VfJ
2 Gear teeth stripped
Motor does not trip but machine stops. 3 hours downtime to replace gearbox with spare. New gears fitted in workshop
3 Gearbox seizes due
Motor trips and alarm sounds in control room. 3 hours downtime to replace gearbox with spare. Seized gearbox would be scrapped
to lack of oil
......etc
In general, the the failure modes which could affect a subassembly should be incorporated in a higher level analysis if the subassembly is likely to suffer from no more than about 6 failure modes which are considered to be worth identifying and which will cause any one functional failure of the higher level system.
Option 2 List the failure of the sub-assembly as a single failure mode on the Informa tion Worksheet to begin with, then raise a new worksheet to analyse the functions, functional failures, failure modes and effects of the sub-assembly as a separate exercise. For example, the failure of the gearbox discussed above could have been listed as follows:
FAILURE MODE Gearbox fails
Gearbox analysed separately
..... . etc
A sub-assembly is usually worth treating in this way if more than 1 0 fail ure modes of the subassembly could cause the loss of any one function of the main assembly. (If there are between 7 and 9 failure modes per functional failure, use option one or option two, bearing in mind that separate analyses mean more analyses, but fewer failure modes per analysis.)
Option 3 List the failure of the sub-assembly on the Information Worksheet as a single failure mode - in other words, at a level equi valent to level one or two in Figure 4.7 record its effects, and leave it at that. -
For example, if it was considered appropriate to treat the failure of the gearbox discussed above in this fashion, it would be listed as shown overleaf:
Reliability-centred Maintenance
88
FAILURE MODE Gearbox fails
Failure Modes and Effects Analysis (FMEA)
89
FAILURE EFFECT
Motor trips and alarm sounds in control room. Downtime to
replace gearbox 3 hours
. . . . . . etc
This approach should only be adopted for a component or subassembly which has the following characteristics: •
it is not subject to detailed diagnostic and repair routines when it fails, but is simply replaced and either discarded or subjected to later repair
•
it is quite small but quite complex
•
it is not likely to be susceptible to any form of proactive maintenance.
•
it does not have any dominant failure modes
Option 4 In some cases,a complex subassembly might suffer from one or two dom inant failure modes which are readily preventable, and a number of less common failures which may not be worth preventing because the frequency and/or the consequences of the failures do not warrant it. For example, a small electric motor operating in a dusty environment might be certain to fail due to overheating if the grille covering its cooling fan gets blocked, but failures for other reasons might be few, far between and not very serious if they do occur. In this case, the failure modes for this motor might be listed as follows: motor fan blocked by dust • motor fails (for other reasons). •
This option is really a combination of options 1 and 3.
Services The failure of services (power, water, steam, air, gases, vacuum, etc) are treated as a single failure mode from the point of view of the asset which is supplied by that serviee, beeause detailed analysis of these failures is usually beyond the scope of the asset in question. Such failures are noted for infonnation purposes (,Power supply fails' ), their effects reeorded and they are then analysed in detail when the service is analysed as a whole.
A Completed lrzj()rmation Worksheet Failure effects are listed in the last column of the Information Worksheet alongside the relevant failure mode, as shown in Figure 4. 13.
w li>' a s
a ,=: ::H!
="
t: �
�
(J)
(J) <1l
�
W !ii
'f "
E
0
�
�
�"
J2
"'
�
��
" " 0
V) � �
<1l .c:
It)
>'"
\f{1
0'
(!j S: �
::;
We a: " => ..
�
"
w
0>
�
Ui
�
::;
w
... '" >-
"l
III => '"
z 0 ,... 0
- >.tJ f-
f-! � < ::J:: _ � CIl - rx: � � o rx:
U�O
� o � ;1 � �
rx: � � @
z
Q
... 0 z =>
u.
Figure 4. 13:
The ReM Informa tion Worksheet
Failure Consequences
5
Fai l u re Conseq uences
Previous chapters have explained how the RCM process asks the follow ing seven questions about each asset: •
what are the functions and associated performance standards of the asset in its present operating context?
•
in what ways does it fail to fulfil its functions ?
•
what happens when each failure occurs?
•
what causes each functional failure ?
•
in what way does each failure matter?
•
what if a suitable proactive task cannot be found?
•
what can be done to predict or prevent each failure?
The answers to the first four questions were discussed at length in Chap ters 2 to 4. These showed how RCM Information Worksheets are used to record the functions of the asset under review, and to list the associated functional failures, failure modes and failure effects. The last three questions are asked about each individual failure mode. This chapter considers the fifth question: •
in what way does each failure matter?
5.1 Technically Feasible and Worth Doing Every time a failure occurs, the organisation which uses the asset is affec ted in some way. Some failures affect output, product quality or customer service. Others threaten safety or the environment. Some increase oper ating costs, for instance by increasing energy consumption, while a few have an impact in four, five or even all six of these areas. Still others may appear to have no effect at all if they occur on their own, but may expose the organisation to the risk of much more serious failures. If any of these failures are not prevented, the time and effort which need to be spent correcting them also affects the organisation, because repair ing failures consumes resources which might be better used elsewhere.
91
The nature and severity of these effects govern the way in which the failure is viewed by the organisation. The precise impact in each case in other words, the extent to which each failure matters depends on the operating context of the asset, the performance standards which apply to each function, and the physical effects of eaeh failure mode. This combination of context, standards and effects means that every failure has a specific set of conscl/uences associated with it. If the conse quences are very serious, then considerable efforts will be made to prevent the failure, or at least to anticipate it in time to reduce or eliminate the consequences. This is especially true if the failure could hurt or ki II some one,or if it is likely to have a serious effect on the environment. It is also true of failures which interfere with production or operations, or which cause significant secondary damage. On the other hand, if the failure only has minor consequences, it is possible that no proactive action will be taken and the failure simply cor rected each time it occurs. This suggests that the consequences of failures are more important than their technical characteristics. It also suggests that the whole idea of proactive maintenance is not so much about preventing failures as it is about avoiding or reducing the consequences of failure. Proactive maintenance has much more to do w ith avo iding or reducing the consequences offa illlre than it has to do w ith preventing the failures themselves
If this is accepted, then it stands to reason that any proactive task is only worth doing if it deals successfully with the consequences of the failure which it is meant to prevent. A proactive task is worth doing if it deals successfully w ith the consequences of the failure which it is meant to prevent
This of course presupposes that it is possible to anticipate or prevent the failure in the first place. Whether or not a proactive task is technically feasible depends on the technical characteristics of the task and of the failure which it is meant to prevent. The criteria governing technical feasibility are discussed in more detail in Chapters 6 and 7. If it is not possible to find a suitable proactive task,the nature of the failure consequences also indicate what default action should be taken. Default tasks are reviewed in Chapters 8 and 9.
92
Failure Consequences
Reliability-centred Maintenance
The remainder of this chapter considers the criteria used to evaluate the consequences of failure, and hence to decide whether any form of proac tive task is worth doing. These consequences are divided in two stages into four categories. The first stage separates hidden functions from evident functions.
5.2 Hidden and Evident :Functions We have seen that every asset has more than one and sometimes dozens of functions. When most of these functions fail, it will inevitably become apparent to someone that the failure has occurred. For instance, some failures cause warning lights to flash or alarms to sound, or both. Others cause machines to shut down or some other part of the process to be interrupted. Others lead to product quality problems or increased use of energy, and yet others are accompanied by obvious physical effects such as loud noises, escaping steam, unusual smells or pools of liquid on the noor. For example, Figure 2.7 in Chapter 2 showed three pumps which are shown again in Figure 5 . 1 below. If a bearing on Pump A seizes, pumping capability is lost. This failure on its own will inevitably become apparent to the operators, either as soon as it happens or when some downstream part of the process is interrupted. (The operators might not know immediately that the problem was caused by the bearing, but they would eventually and inevitably become aware of the fact that something unusual had happened.)
Figure 5. 1:
Stand Alone
Three pumps
Failures of this kind are classed as evident because someone will even tually find out about it when they occur on their own. This leads to the following definition of an evident function: A 11 evident function is one whose failure w ill on
its own eventually and inevitably become evident to the operating crew u nder normal circumstances
However, some failures occur in such a way that nobody knows that the item is in a failed state unless or until some other failure also occurs.
93
For instance, if Pump C in Figure 5. 1 failed, no-one would be aware of the fact because under normal circumstances Pump B would still be working. In other words, the failure of Pump C on its own has no direct impact unless or until Pump B also fails (an abnormal c ircumstance).
Pump C exhibits one of the most important characteristics of a hidden function, which is that the failure of this pump on its own will not become evident to the operating crew under nonnal circumstances. In other words, it will not become evident unless pump B also fails. This leads to the following definition of a hidden function: A h idden function is one whose failure will not become e vident to the operating crew under normal c ircumstances if it occurs Oil its own.
The first step in the ReM process is to separate hidden functions from evident functions because hidden functions need special handling. They are discussed at length in Part 6 of this chapter. We will see later that these functions are associated with protective devices which are not fail safe. Since they can account for up to halfthe failure modes which could 4fect modern, complex equipment, hidden functions could well become the dominant issue in maintenance over the next ten years. However, to place hidden functions in perspective, we first consider evident failures. Categories of Evident Failures
Evident failures are classified into three categories in descending order of importance, as follows:
safety and environmental consequences.
•
A failure has safety conse quences if it could hurt or kill someone. It has environmental conse quences if it could lead to a breach of any corporate, regional or national environmental standard
•
operational consequences.
•
Iloll-operatiollal consequences. Evident failures in this category affect
A failure has operational consequences if it affects production or operations (output, product quality, customer service or operating costs in addition to the direct cost of repair) neither safety nor production, so they involve only the direct cost of repair.
By ranking evident failures in this order, RCM ensures that the safety and environmental implications of every evident failure mode are considered. This unequivocally puts people ahead of production.
94
Reliability-centred Maintenance
This approach also means that the safety, environmental and economic consequences of each failure are assessed in one exercise, which is much more cost-effective than eonsidering them separately. The next four sections of this chapter consider each of these categories in detail, starting with the evident categories and then moving on to the rather more complex issues surrounding hidden functions.
5.3
Safety and Environmental Consequences
Safety First
As we have seen, the first step in the consequence evaluation process is to identify hidden functions so that they can be dealt with appropriately. All remaining failure modes - in other words, failures which are not clas sified as hidden must by definition be evident. The above paragraphs explained that the RCM process considers the safety and environmental implications of each evident failure mode first. It does so for two reasons: •
a more and more firmly held belief among employers, employees, cus tomers and society in general that hurting or killing people in the course of business is simply not tolerable, and hence that everything possible should be done to minimise the possibility of any sort of safety-related incident or environmental excursion.
•
the more pragmatic realisation that the probabilities which are tolerated for safety-related incidents tend to be several orders of magnitude lower than those which are tolerated for failures which have operational con sequences. As a result, in most of the cases where a proactive task is worth doing from the safety viewpoint, it is also likely to be more than adequate from the operational viewpoint.
At one level, safety refers to the safety of individuals in the workplace. Specifically, RCM asks whether anyone could get hurt or killed either as a direct result of the failure mode itself or by other damage which may be caused by the failure. A failure mode has safety consequences if it causes a loss offUllction or other damage which could hurt or kill someone
Failure Consequences
95
At another level, 'safety' refers to the safety or well-being of society in general. Nowadays, failures which affect society tend to be classed as 'environmental' issues. In fact, in many parts of the world the point is fast approaching where organisations either conform to society' s environ mental expectations, or they will no longer be allowed to operate. So quite apart from any personal feelings which anyone may have on the issue. environmental probity is becoming a prerequisite for corporate survival. Chapter 2 explained how soeiety' s expectations take the form of muni cipal, regional and national environmental standards. Some organisations also have their own sometimes even more stringent corporate standards. A failure mode is said to have environmental consequences if it could lead to the breach of any of these standards. A fai/ure mode has en vironmental consequences if it causes a loss offllllction or other damage which could lead to the breach of any known environmental standard or regulation
Note that when considering whether a failure mode has safety or environ mental consequences, we are considering whether one failure mode on its own could have the consequences. This is different from part 6 of this chapter,in which we consider the failure of both elements of a protected system. The Question of Risk
Much as most people would like to live in an environment where there is no possibility at all of death or injury, it is generally accepted that there is an element of risk in everything we do. In other words, absolute zero is unattainable, even though it is a worthy target to keep striving for. This immediately leads us to ask what is attainable. To answer this question, we first need to consider the question of risk in more detail. Risk assessment consists of three elements. The first asks what could happen if the event under consideration did occur. The second asks how likely it is for the event to occur at all. The combination of these two ele ments provides a measure of the degree of risk. The third and often the most contentious element asks whether this risk is tolerable.
96
Reliability-centred Maintenance
Failure Consequences
For example, consider a failure mode which could result in death or injury to ten . people (what could happen). The probability that this failure mode could occur IS one in a thousand in any one year (hOW likely It IS to occur). On the baSIS of these figures, the risk associated with this failure is: 1 0 x (1 in 1 000) 1 casualty per 1 00 years . Now consider a second failure mode which could cause 1 000 casualties, but the probability that this failure could occur is one in 1 00 000 in any one year. The risk associated with this failure is: 1 000 x (1 in 1 00 000) 1 casualty per 1 00 years. =
?
In these examples, the risk is the same although the figures upon whic . it is based are quite different. Note also that these examples do not mdl cate whether the risk is tolerable - they merely quantify it. Whether or not the risk is tolerable is a separate and much more difficult question which is dealt with later.
Note that throughout this book, the terms 'probability ' (1 in 10 chance of afailure in any one period) and 'failure rate ' (once in ten periods on aver age, corresp;mding to a mean time between failures of 1 .o.periods) are . used as ifthey are interchangeable when applied to randomjazlures. Strzctly speaking, this is not true. However, if the MTBF is greater than about 4 periods, the dUJerence is so small that it can usually be ignored. . The following paragraphs consider each of the three elements of nsk
in more detail.
What could happen if the failure occurred? Two issues need to be considered when considering what could happen if a failure were to occur. These are what actually happens and whether anyone is likely to be hurt or killed as a result. What actually happens if any failure mode occurs should be re orded on the RCM Information Worksheet as its failure effects, as explamed at length in Chapter 4. Part 5 of Chapter 4 also listed a number of typical effects which pose a threat to safety or the environment. The fact that these effects could hurt or kill someone does not neces sarily mean that they will do so every time they occur. Some may even occur quite often without doing so. However, the issue is n t whether such consequences are inevitable, but whether they are possIble.
�
�
�
For example, if the hook were to fail on a travelling crane used to carry s eel coils, the falling load would only hurt or kill anyone who happened to be standing under it or very close to it at the time. If no-one was nearby, then no-one would get hurt. However the possibility that someone could be hurt means that thiS failure mode should b treated as a safety hazard and analysed accordingly.
�
97
This example demonstrates the fact that the RCM process assesses safety consequences at the most conservative level. lf it is reasonable to assume that any failure mode could affect safety or the environment,we assume that it can, in which case it must be subjected to further analysis. (We see later that the likelihood that someone will get hurt is taken into conside ra tion when evaluating the tolerability of the risk.) A more complex situation arises when dealing with safety hazards that are already covered by some form of built-in protection. We have seen that one of the main objectives of the RCM process is to establish the most effective way of managing each failure in the context of its consequences. This can only be done if these consequences are evaluated to begin with as if nothing was being done to manage the failure (in other words, to pre dict or prevent it or to mitigate its consequences). Protective devices which are designed to deal with the failed or the failing state (alarms, shutdowns and relief systems) are nothing more than built-in failure management systems. As a result, to ensure that the rest of the analysis is carried out from an appropriate zero-base, the conse quences of the failure of protected functions should ideally be assessed as if protective devices of this type are not present. For example, a failure which could cause a fire is always regarded as a safety hazard, because the presence of a fire-extinguishing system does not necessar ily guarantee that the fire will be controlled and extinguished.
The RCM process can then be used to validate (or revalidate) the suitabi lity of the protective device itself from three points of view: •
its ability to provide the required protection. This is done by defining the function of the protective device, as explained in Chapter 2
•
whether the protective device responds jclst enough to avoid the con sequences, as discussed in Chapter 7
•
what must be done to ensure that the protective device continues to function in its turn, as discussed in part 6 of this chapter and Chapter 8 .
How likely i s the failure t o occur? Part 4 of Chapter 4 mentions that only failure modes which are reason ably likely to occur in the context in question should be listed on the RCM Information Worksheet. As a result, if the Information Worksheet has been prepared on a realistic basis, the mere fact that the failure mode has been listed suggests that there is some likelihood that it could occur, and therefore that it should be subjected to further analysis.
98
Reliability-centred Maintenance
Failure Consequences
The figures given in this example are not meant to be prescriptive and they do not necessarily reflect the views of the author they merely illustrate what one individual might decide that he or she is prepared to tolerate. Note al o that they are based on the perspective of one individual going about hIs or her daily business. This view then has to be translated into a degree of risk for the whole population (all the workers on a site,all the citizens of a town or even the entire population of a country).
(Sometimes it may be prudent to list a wildly unlikely but nonetheless dangerous failure mode in an FMEA, purely to place on record the fact that it was considered and then rejected. In these cases, a comment like "This failure mode is considered too unlikely to justify further analysis" should be recorded in the failure effects column.)
�
Is the risk tolerable ? One of the most difficult aspects of the management of safety is the extent to which beliefs about what is tolerable vary from individual to indivi dual and from group to group. A wide variety of factors influence these beliefs, by far the most dominant of which is the degree of control which any individual thinks he or she has over the situation. People are nearly always prepared to tolerate a higher level of risk when they believe that they are personally in control of the situation than when they believe that the situation is out of their control. For example, people tolerate much higher levels of risk when driving their own cars than they do as aircraft passengers. (The extent to which this issue governs perceptions of risk is given by the startling statistic that 1 person in 1 1 000000 who travels by air between New York and Los Angeles in the USA is likely to be killed while doing so, while 1 person i n 1 4 000 who makes the trip by road is likely to be killed. And yet some people i nsist o n making this trip by road because they believe that they are 'safer'!)
This example illustrates the relationship between the probability of being killed which any one person is prepared to tolerate and the extent to which that person believes he or she is in control. In more general terms, this might vary for a particular individual as shown in Figure 5.2. o � "' .. - '"
1 0-4 l!! » .52 � 1 0.5 .s o - » ..c: c £1 .. ..c: .c 1 0--6 ;:: -
.:::--g
= =
:5 32 "' r:» c .CI 0 ·-
......
�
1 0-7
� '" CL .CI
Figure 5.2: Tolerability
of fatal risk
I believe I have
complete control (driving my car or in my home workshop)
I�
believe I have some control and some choice about exposing myself (on the site where I work) I
�
I believe I have
no control, but I don't have to expose myself (in a passenger aircraft)
I--
have no control, and no choice about exposing myself and/or my family (off-site exposure to industrial accidents) I
99
In other words, if I tolerate a probability of 1 in 100 000 (10.5) of being killed at work in any one year and I have 1 000 co-workers who all share the same view, then we all tolerate that on average 1 person per year on our site will be killed at work every 100 years - and that person may be me, and it may happen this year.
Bear in mind that any quantification of lisk in this fashion can only ever be a rough approximation. In other words, if I say I tolerate a probability of 1 0.5, it is never more than a ballpark figure. It indicates that I am prepared to tolerate a probability of being killed at work which is roughly 1 0 times lower than that which I tolerate when I use the roads (about 1 0 4). Always bearing in mind that we are dealing with approximations, the next step is to translate the probability which myself and my co-workers are prepared to tolerate that any one of us might be killed by any event at work into a tolerable probability for each single event (failure mode or multiple failure) which could kill someone. For example, continuing the logic of the previous example, the probability that any one of my 1 000 co-workers will be killed in any one year i s 1 in 100 (assuming that everyone on the site faces roughly the same hazards) . Furthermore, if the activities carried out on the site embody (say) 10 000 events which could kill someone, then the average probability that each event could kill one person must be reduced to 10.6 i n any one year. This means that the probability of an event which is likely to kill ten people must be reduced to 10.7, while the probability of an event which has a 1 in 10 chance of killing one person must be reduced to 10.5.
The techniques by which one moves up and down hierarchies of probabil ity in this fashion are known as probabilistic or quantitative risk assess ments. This approach is explored further in Appendix 3. The key points to bear in mind at this stage are that: •
the decision as to what is tolerable should start with the likely victim. How one might involve such ' likely victims' in this decision in the industrial context is discussed later in this chapter
•
it is possible to link what one person tolerates directly and quantita tively to a tolerable probability of individual fai lure modes.
100
Reliability-centred Maintenance
Failure Consequences
Although perceived degree of control usually dominates decisions about the tolerability of risk, it is by no means the only issue. Other factors which help us decide what is tolerable include the following: •
•
•
individual values: To explore this issue in any depth is well beyond the scope of this book. Suffice it to contrast the views on tolerable risk likely to be held by a mountaineer with those of someone who suffers from veltigo, or those of an underground miner with those of someone who suffers from claustrophobia.
industry values: While every industry nowadays recognises the need to operate as safely as possible,there is no escaping the fact that some are intrinsically more dangerous than others. Some even compensate for higher levels of risk with higher pay levels. The views of any individual who works in that industry ultimately boil down to his or her perception of whether the intrinsic risks are 'worth it' in other words, whether the benefit justifies the risk.
�
the efJect on 'future generations ': The safety of children - espec ally unborn children has an especially powerful effect on peoples' VIews about what is tolerable. Adults frequently display a surprising and even distressing disregard for their own safety. (Witness how much time has to be spent persuading some people to wear protective clothing.) How ever, threaten their offspIing and their attitude changes completely.
?
For example, the author worked with one group which had occasion to iscuss the properties of a certain chemical. Words like 'toxic' and 'carcinogenic' were treated with indifference, even though most of the members of this group were the people most at risk. However, as soon as it emerged that the chemical ,:,as also mutagenic and teratogenic, and the meaning of these words was explained to the group, the chemical was suddenly viewed with much greater respect. •
knowledge: perceptions of risk are greatly influenced by how much people know about the asset, the process of which it forms part and the failure mechanisms associated with each failure mode. The more they know, the better their judgement. (Ignorance is often a two-edged sword. In some situations people take the most appalling risks out of sheer ignorance, whilG in others they wildly exaggerate the risks also out of ignorance. On the other hand, we need to remind ourselves constantly of the extent to which familiarity can breed contempt.)
A great many other factors also influence perceptions of ris�, such as the . value placed on human life by different cultural groups, rehgIOus values and even factors such as the age and marital status of the individual.
101
All of these factors mean that it is impossible t o specify a standard of tolerability for any risk which is absolute and obj ective. This suggests that the tolerability of any risk can only be assessed on a basis which is both relative and subjecti ve - 'relative' in the sense that the risk is compare d with other risks about which there is a fairly clear consensus, and 'subjec tive' in the sense that the whole question is ultimately a matter of judgement. But whose judgement?
Who should evaluate risks ? The very diversity of the factors discussed above mean that it is simply not possible for any one person or even one organisation to assess risk in a way which will be universally acceptable. If the assessor is too con servative, people will ignore and may even ridicule the evaluation. If the assessor is too relaxed, he or she might end up being accused of playing with people's lives (if not actually killing them). This suggests further that a satisfactory evaluation of risk can only be done by a group. As far as possible, this group should represent people who are likely to have a c lear understanding of the failure mechanism, the failure effects (especially the nature of any hazards), the likelihood of the failure occurring and what possible measures can be taken to anticipate or prevent it. The group should also include people who have a l egitimate view on the tolerability or otherwise of the risks. This means represen tatives of the likely victims (most often operators or maintainers in the case of direct safety hazards) and management (who are usually he Id account able if someone is hurt or if an env·ironmental standard is breached). If it is applied in a properly focused and structured fashion, the collec tive wisdom of such a group will do much to ensure that the organisation does its best to identify and manage all the failure modes that could affect safety or the environment. (The use of such groups is in keeping with the worldwide trend towards laws which say that safety is the responsibility of all employees, not just the responsibility of management.) Groups of this nature can usually reach consensus quite quickly when dealing with direct safety hazards, because they include the people at risk. Environmental hazards are not quite so simple, because society at large is the 'likely victim' and many of the issues involved are unfamiliar. So any group which is expected to consider whether a failure could breach an environmental standard or regulation must find out beforehand which of these standards and regulations cover the process under review.
Reliability-centred Maintenance
102
Safety and Proactive Maintenance
If a failure could affect safety or the environment, the RCM process stipu lates that we must try to prevent it. The above discussion suggests that: For failure modes which have safety or environmental con sequences, a proactive task is only worth doing if it reduces the probability of the failure to a tolerably low level
If a proactive task cannot be found which achieves this objective to the satisfaction of the group performing the analysis, we are dealing with a safety or envirorimental hazard which cannot be adequately anticipated or prevented. This means that something must be changed in order to make the system safe. This 'something' could be the asset itself, a process or an operating procedure. Once-off changes of this sort are classified as 'redesigns' , and are usually undertaken with one of two objectives: • to reduce the probability of the failure occurring to a tolerable level • to change things so that the failure no longer has safety or environmen tal consequences. The question of redesign is diseussed in more detail in Chapter 9 . Note that when dealing with safety and environmental issues, RCM does not raise the question of economics. If it is not safe we have an obli gation either to prevent it from failing, or to make it safe. This suggests that the decision process for failure modes which have safety or environ mental consequences can be summarised as shown in Figure 5 . 3 below: Doe$ tM fl.tilure mode ca�se a loss <.>f function or <.>tne.r damage which r- No cOl,Jldhurt Or kill · . . someone? Figure 5.3: Identifying and developing a maintenance strategy for a failure which affects safety or the environment
/ If
I
-
Does the failure mode cause a loss of functio" .or other damage which eQuId breach any �own environmental standa.r d or regulation?:
I
Yes
Yes
I
I
Proac�ye maintenance. is worth dOing if. it reduces the risk of til, failure tQ a tolerably loW leVel
I
I
No
I
See Parts 4 and 5 of this chapter
a proactive task cannot be found which reduces the risk of the failure to a tolerably low level, redesign is compulsory
I
Copyrighted Material
The basis on which we determine Ihe I�chnical feasibililY and frequency of differem Iypes of pro;,clive task is discussed in Chapler.; 6 and 7.
RCM and SaMy Legislation A question ofleo afi,es concerning Ihe rcialionship between RCM 'Ind safely legisblion (environmental legisl3lioll is deall with dircclly), Nowadays. most lcgblation gOl'eming safety merely demands thm users are able to demon,trate that they afe doing whatever is pruden! to ensure that their assets are safe. This has led 10 "'pidiY increasing empha· sis on the concept of an ",,,/il Imil. whid basically rcquiIC.� U..en. ofassets .
to be able to produce documentary cvidence thai Ihere is a "'tionaL defen sible basis for their mainlcnance programs. In Ihe vast majorily of cases.
RCM wholly s;l\isfies this Iype of requirement However. some regulmiollS demand Ihal specific lasks should be done on specific types of equipmenl al specific imervals. If the RCM proce." suggests a differenl lask and/or a diffcrem imerval. il is wise to conlinu!' doing the task specified by I� legislation and to discu" Ihe .,ugge,led change with Ihe appropriate regul;llory aUlhority.
5.4 Operational Consequences How Failures Affect Operations The primary funelion of rnostl'quipmelll in industry is conneCied in .'10m!' way wilh lhe need 10 cam rewnue or 10 soppon fCVCllue earning aclivilies. For example, the primary lunct,on 01 most ot the asselS used in manulacturing is to add value to malerials, while customers pay directly loraccess 10 telecommuni_ calions and transport equipment (buses, Irucl<s, Irains or aircrah). Failures which affect Ihe primary funelion, of these aSSCIS affcci Ihe revenue-earning capabi lily of Ihe org;1nization. The magnitude of these effects depends on how heavi Iy Ihe equipllleni is loaded and Ihe availabil· ity of aliemali ves. However. in nemly all cases the effeels are greal(""f
_
often much greater - Ihan Ihe CO>I of repairing 11K: failures. This i, also Inre of equipmcnl in servi<:c ;nduslries stich as comenainmenL <'ommen:e and c"en banking. For example. � the lights lail at a ball game. fans lend to want the" money back. The same applies it projectOfs lail at the movies II the air-conditiomng fails in a shop Or restaurant. CUSlomers walk out. 8anl<s Jose busif\eSS � the" AnA's lail.
Copyrighted Material
104
Failure Consequences
Reliability-centred Maintenance
As we have seen, these consequences tend to be
In general, failures affect operations in four ways:
they qffect total output.
o
This occurs when equipment stops working
altogether or when it works too slowly. This results either in increased production costs if the plant has to work extra time to catch up, or lost sales if the plant is already fully loaded. o
they affect product quality.
If a machine can no longer hold manufac
so
they are usually evaluated in economic terms. However, in certain more extreme cases (such as losing a war), the 'cost' may have to be evaluated on a more qualitative basis. Avoiding Operational Consequences
turing tolerances or if a failure causes materials to deteriorate, the likely
The overall economic effect of any failure mode which has operational consequences depends on two factors: o
tems, the accuracy of targeting systems and so on. Failures affect customer service in many of ways, ranging from the late delivery of orders to the late departure heavy attract s passenger aircraft. Frequent or serious delays sometime of penalties, but in most cases they do not result in an immediate loss
they (!ffect customer service.
increased operating costs in addition to the direct cost of repair.
instance, the failure might lead to the increased use of energy or involve switching to a more expensive alternative process.
how much the failure costs each time it occurs, in terms of its effect on operational capability plus repair costs
s revenue. However chronic service problems eventually cause customer e. elsewher business their take and to lose confidence
o
economic in nature '
result is either scrap or expensive rework. In a more general sense, "quality" also covers concepts such as the precision of navigation sys
o
105
For
it might
failures In non-profit enterprises such as military undertakings, certain function primary can also affect the ability of the organisation to fulfil its sometimes with devastating results.
•
how often it happens.
In the previous section of this chapter, we did not pay much attention to how often failures are likely to occur. (Failure rates have little bearing on
�
�
sa 'ety-related ailures, because the objective in these cases is to avoid any fmlures on WhICh to base a rate.) However, if the failure consequences are economic, the total cost
failures, we need to assess how much they are likely to cost over if period
of time. Consider for example the pump shown in Figure 2.1 and again in Figure 5.4. The pump is controlled by
"For want of a nail, a shoe was lost. For want of a shoe, a horse was lost. For want
one float switch which ac
of a horse, a message was lost. For want of a message, a battle was lost. For want
tivates it when the level in
of a battle, a war was lost. All for want of a horseshoe nail."
Tank Y drops to 120 000
failures While it may be difficult to cost out the results of losing a war, If level. mundane more a at ons implicati c economi have of this sort still order in horses two (say) keep to necessary be may they occur too often, it tanks to ensure that one will be available to do the job - or sixty battle this on cy Redundan five. of instead of fifty -- or six aircraft carriers instead
litres, and another that tums
scale can be very expensive indeed. does The severity of these consequences mean that if an evident failure focuses not pose a threat to safety or the environment, the ReM process next on the operational consequences of failure.
A failure has operational consequences if it has a
direct adverse effect on operational capability
is affected by how often the consequences are
likely to occur. In other words, to assess the economic impact of these
Pump can deliver up to 1 000 litres of water per minute
y Offtake from tank: 800 litres/minute
it off when the level in Tank
Y reaches 240 000 litres. A
Figure 5.4: Stand-alone pump
low level alarm is located just below the 120 000 litre level. If the tank runs dry, the downstream process has to be shut down. This costs the organisation using the pump £5 000 per hour.
FAILURE MODE 1
Bearing seizes due to normal wear and tear
FAILURE EFFECT ------�----. ..---Motor trips but no alarm sounds in control room. Level in tank drops until low level alarm sounds at 120 000 litres. Downtime to replace the bearing 4 hours. (The mean time between occurrences of this failure mode is about 3 years.)
Figure 5.5: FMEA for bearing failure on the stand-alone pump
106
Assume that it has already been agreed that one failure mode which can affect this pump is 'Bearing seizes due to normal wear and tear'. For the sake of simpli
•
reducing the frequency (and hence the total cost) of the failure
•
reducing or eliminating the consequences of the failure
•
making a proactive task cost-effective.
city, assume that the motor on this pump is equipped with an overload switch, but there is no trip alarm wired to the control room. This failure mode and its effects might be described on an ReM Information Worksheet as shown in Figure5.5 above. Water is drawn out of the tank at a rate of800 litres per minute, so the tank runs dry2.5 hours after the low level alarm sounds. It takes 4 hours to replace the bear ing, so the downstream process stops for1.5 hours. So this failure costs: 1.5 x £5 000 :: £7500 in lost production every three years, plus the cost of replacing the bearing. Assume that it is technically feasible to check the bearing for audible noise once a week (the basis upon which we make this kind of judgement is discussed at length in the next chapter). If the bearing is found to be noisy, the operational
107
Failure Consequences
Reliability-centred Maintenance
Redesign is discussed in more detail in Chapter 9. Note that in the case of a failure mode with safety and environmental consequences, the objective is to reduce the probability of the failure to a very low level indeed. In the case of operational consequences, the objective
i
is to reduce the probability (or frequency) to an economically to erable level. As mentioned at the start of part 3 of this chapter, this frequency is likely to be several orders of magnitude greater than we would tolerate for most safety hazards, so the RCM process assumes that a proactive task
consequences of failure can be avoided by ensuring that the tank is full before
which reduces the probability of a safety-related failure to a tolerable
starting work on the bearing. This provides five hours of storage so the bearing can
level will also deal with the operational consequences of that failure.
now be replaced in four hours without interfering with the downstream process. Assume also that the pump is located in an unmanned pumping station. It has been agreed that the check should be carried out by a maintenance craftsman, and that the total time needed to do each check is twenty minutes. Assume further
To begin with, we again only consider the desirability of making changes after we have established whether it is possible to extract the desired performance from the asset as it is currently configured. However, in this
that the total cost of employing the craftsman is £24 per hour, in which case it costs
case modifications also need to be cost-justified, whereas they were the com
£8 to perform each check. If the MTBF of the bearing is 3 years, he will do about
pulsory default action for failure modes with safety or environmental
150 checks per failure. In other words, the cost of the checks is: 150 x £8 :: £1 200 every three years, again plus the cost of replacing the bearing.
to the In this example, the scheduled task is clearly cost-effective relative repair. of cost the plus failure the of nces cost of the operational conseque basis for This suggests that if a failure has operational consequences, the follows: as c, economi is doing deciding whether a proactive task is worth
consequences. In the light of these comments, the decision process for failures with operational consequences can be summarised as shown in figure 5.6: Does the failure mode have a direct adverse effect on operational capability? I Yes
For failure modes with operational consequences, a proactive task is worth doing if, over a period of time, it costs less than the cost of the operational consequences plus the cost of repairing the failure which it is meant to prevent
I
Proactive maintenance is worth doinQ if it costs less over a period of time than the cost of the operational consequences plus the cost of repai ring the failures which it is meant to prevent
then it is Conversely, if a cost-effective proactive task cannot be found, prevent or anticipate to not worth doing any scheduled maintenance to try
costs by: the asset (or to change the process) in order to reduce total
I
See Part 5 of this chapter
1
cost-effec the failure mode under consideration. In some cases, the most the failure. tive option at this point might simply be to decide to live with
conse However, if a proactive task cannot be found and the failure of design the change to e desirabl be may it le, quences are still intolerab
I
No
If a cost-effective Jjroactive task cannot be found, the default decision is no scheduled maintenance...
\
I but it might be worth redesigning the asset or changing the process to reduce total costs ...
.�
Figure 5.6: Identifying and developing a maintenance strategy for a failure which has operational consequences
108
Reliability-centred Maintenance
Failure Consequences
Note that this analysis is can-ied out for each individual failure mode, and not for the asset as a whole. This is because each proactive task is designed
��
to prevent a specific failure mode, so the economic feasibility o ach task . can only be compared to the costs of the failure mode WhICh It IS meant
�
to prevent. In each case, it is a simple go/no go decisio . . . In practice, when assessing individual failure modes 111 thIS way, It IS .
�
not always necessary to do a detailed cost-benefit study based on actu l downtime costs and MTBF's as shown in the example on Page 106. ThIS
is because the economic desirability of proactive tasks is often intuitively obviolls when assessing failure modes with operational consequences. However, whether or not the economic consequences are evaluated fOJ1nally or intuitively, this aspect of the ReM process must still be applied
thoroughly. (In fact, this step is surprisingly often overlooked by people new to the process. Maintenance people in particular have a tendency to
implement tasks on the basis of technical feasibility alone, which results in elegant but excessively costly maintenance programs.) . Finally, bear in mind that the operational consequences of any failure are heavily influenced by the context in which the asset is operating. This
is yet another reason why care should be taken to ensure that the context one is identical before applying a maintenance program developed for
2. asset to another. The key issues were discussed in Part 3 of Chapter
109
The duty pump is switched on by one float switch when the level in Tank Y drops to 120 000 litres, and switched off by another when the level reaches240 000 litres. A third switch is located just below the low level switch of the duty pump, and this switch is designed both to sound an alarm in the control room if the water level reaches it, and to switch on the stand-by pump. If the tank runs dry, the downstream process has to be shut down. This also costs the organisation which uses the pump £5 000 per hour. As before, assume that it has been agreed that one failure mode which can affect the duty pump is 'bearing seized', and that this seizure is caused by normal wear and tear. Assume that the motor on the duty pump is also equipped with an over-load switch, but again there is no trip alarm wired to the control room. This failure mode and its effects might be described on an RCM Information Work sheet as shown in Figure5.8:
PAil liRE EFFECT
FAILURE MODE Bearing seizes due to normal wear and tear
Motor trips but no alarm sounds in control room. Level in tank drops until low level alarm sounds at 120 000 litres, and stand-by pump is switched on automatically. Time required to replace the bearing 4 hours. (The mean time
between occurrences of this failure is about 3 years)
Figure 5.9: FMEA for failure of bearing on duty pump with stand-by In this example, the stand-by pump is switched on when the duty pump fails, so the tank does not run dry. So the only cost associated with this failure is: the cost of replacing the bearing. Assume however that it is still technically feasible to check the bearing for audible noise once a week. If the bearing were found to be noisy, the operators would switch over manually to the stand-by pump and the bearing would be replaced. Assume that these pumps are also located in an unmanned pumping station,
5.5 Non-operational Consequences
and that it has again been agreed that the check
The consequences of an evident failure which has no direct adverse effect on safety, the environment or operational capability are classified as non
:
which also takes twenty min
utes - should be done by a maintenance craftsman at a cost of £8 per check. So once again, he will do about 150 checks per failure. In other words, the cost of the proactive maintenance program per failure is:
the operational. The only consequences associated with these failu es are economic. also are direct costs of repair, so these consequences
In this example, the cost of doing the scheduled task is now much greater
Consider for example the pumps shown in Figure57 . . This set-up is similar to that
than the cost of not doing it. As a result, it is not worth doing the proactive
shown in Figure5.4, except that there are now two pumps (both identical to the pump in Figure5.4).
150 x £8
=
£1 200 plus the cost of replacing the bearing.
task even though the pump is technically identical 10 the pump described in Figure 5.3. This suggests that it is only worth trying to prevent a failure which has non-operational consequences if, over a period of time, the cost of the preventive task is less than the cost of correcting the failure. If it is
Figure 5.7: Pump with stand-by
x
Pumps can deliver up to1 000 litres of water per minute
y Offtake from tank: 800 litres/minute
not, then scheduled maintenance is not worth doing.
For failure modes with non-operational consequences, a pro active task is worth doing if over a period of time, it costs less than the cost of repairing the failures it is meant to prevent
5.6 Hidden Failure Consequences
If a proactive task is not worth doing, then in rare cases a modification might be justified for much the same reasons as those which apply to failures with operational consequences.
Hidden Failures and Protective Devices Chapter 2 mentioned that the growth in the number of ways in which equipment can fail has led to corresponding growth in the variety arid severity of failure consequences which fall into the evident categories. It also mentioned that protective devices are being used increasingly in an attempt to eliminate (or at least reduce) these consequences, and explained how these devices work in one of five ways: to alert operators to abnormal conditions to shut down the equipment in the evcnt of a failure to eliminate or relieve abnormal conditions which follow ;I failure and which might otherwise cause much more serious damage * to take over from a function which has failed * to prevent dangerous situations from arising. In essence, the function of these devices is to ensure that the consequences of the failure of the protected function are n ~ i ~ less c h serious than they would be if there were no protection. So any protective device is in fact part of a system with at least two components: the protective device the protected function.
Further Points Concerning Non-operational Consequences Two more points need to be considered when reviewing failures with non-operational consequences, as follows: s e c o n d a damage: ry Some failure rnodes cause considerable secondary damage if they are not anticipated or prevented, which adds to the cost of repairing them. A suitable proactive task could make it possible to prevent or anticipate the failure and avoid this damage. However, such a task is only justified if the cost of doing it is less than the cost of repairing the failure and the secondary damage. For example, in Figure 5.7 the description of the failure effects suggests that the seizure of the bearing causes no secondary damage. If this is so, then the analysis is valid.However, if the unanticipated failure of the bearing alsocauses (say) the shaft to shear, then a proactive task which detects imminent bearing failure would enable the operators to shut down the pump before the shaft is damaged. In this case the cost of the unanticipated faiture of the bearing is: the cost of replacing the bearing and the shaft. On the other hand, the cost of the proactive task (per bearing failure) is still: f 1 200 plus the cost of replacing the bearing. Clearly, the task is worth doing if it costs more than f 1 200 to replace the shaft. If it costs less than f 1 200, then this task is still not worth doing.
protected functions: it is only valid to say that a failure will have nonoperational consequences because a stand-by or redundant component is available if it is reasonable to assume that the protective device will be functional when the failure occLIrs. This of course means that a suitable maintenance program must be applied to the protective device (the stand-by pump in the example given above). This issue is discussed at length in the next part of this chapter. If the consequences of the multiplc failure of a protected system are particul;trly serious, it rnay be worth trying to prevent the failure of the protected function as well as the protective device in order to reduce the probability of the multiple failure to a tolerable level. (As explained on Page 97, if the rni~ltiplefailure has safety consequences, it may be wise to assess consequences as if the protection was not present at all, and then to revalidate the protection as part of the task selection process.)
For example, Pump C in Figure 5.7 can be regarded as a protective device, because it 'protects'the pumping function if Pump B should fail. Pump B is of course the protected function.
I
The existence of such a system creates two sets of failure possibilities, depending on whether the protective device is fail-safe or not. We consider the implications of each set in the following paragraphs, starting with devices which are fail-safe. Fuil-saje protective devices In this context, fail-safe means that the failure of the device o n its o w n will become evident to the operating crew ~lndernor~nalcirctrnixtances
Zrz the context of this hook, a 'fail-safe' device is one whose failure on its own will become evident to the operating crew under norrnal circurnstances
112
Reliability-centred Maintenance
This means that in a system which includes a fail-safe protective device, there are three failure possibilities in any period, as follows. The first possibility is that neither device fails. In this case everything proceeds normally . The second possibility i s that the protected function fails before the protective device. In this case the protective device carries out its inten cled function and, depending on the nature of the protection, the conse quences of failure of the protected function are reduced or eliminated. The third possibility is that the protective device fails before the pro tected function. This would be evident because if it were not , the device would not be fail-safe in the sense defined above . If normal good practice is followed , the chance of the protected device failing while the protective device is in a failed state can be almost eliminated, either by shutting down the protected function or by providing alternative protection while the failed protective device is being rectified.
Failure Consequences
113
The second possibility is that the protectedfunctionfails at a time when the protective device is still fullctional. In this case the protective device also carries out its intended function, so the consequences of the failure of the protected function are again reduced or eliminated altogether. For instance, consider a pressure relief valve (the protective device) mounted on a pressure vessel (the protected function). If the pressure rises above tolerable limits, the valve relieves and so reduces or eliminates the consequences of the over pressurisation. Similarly, if Pump B in Figure 5.7 fails, Pump C takes over.
The third possibility is that the protective devicefails while the protected function is still working. In this case, the failure has no direct consequen ces. In fact no-one even knows that the protective device is in a failed state. For example, if the pressure relief valve was jammed shut, no-one would be aware of the fact as long as the pressure in the vessel remained within normal operating limits. Similarly, if Pump C were to fail somehow while Pump B was still working, no-one would be aware of the fact unless or until Pump B also failed.
For instance, an operator could be asked to keep an eye on a pressure gauge and his finger by a stop button - while a pressure switch is being replaced.
The above discussion suggests that hidden functions can be identified by asking the following question:
This means that the consequences of the failure of a fail-safe protective device usually fall into the 'operational' or 'non-operational ' categories . This sequence o f events i s summarised i n Figure 5 .9.
Will the loss of function caused by this failure mode on its own become evident to the operating crew under normal circllmstances?
2: Protected function is shut down or other protection provided while protective device is under repair. This reduces probability of multiple fai/ure to near zero.
Protected function Protective device
Protected '" function made T 4: If protected function fails here, safe while I protective device acts to reduce or eliminate consequences protective I device.is 1: Failure of "fail under repai� 3: Protective device safe" device is reinstated: situation evident immediately back to normal
Figure 5.9: Failure of a "fail-safe" protective device
Protective devices which are not fail-safe In a system which contains a protective device which is not fail-safe, the fact that the device is unable to fulfil its intended function is not evident under nornutl circumstances. This creates four failure possibilities in any given period, two of which are the same as those which apply to a fail-safe device. The first is where neither device fails, in which case everything proceeds normally as before.
If the answer to this question is no, the failure mode is hidden. rf the answer is yes, it is evident. Note that in this context, 'on its own' means that nothing else has failed. Note also that we assume at this point in the analysis that no attempts are being made to check whether the hidden function is still working. This is because such checks are a form of scheduled maintenance, and the whole purpose of the analysis is to find out whether such mainte nance is necessary. These two issues are discussed in more detail later in this chapter. The fourth possibility during any one cycle is that the protective device fails, then the protected.fimction fails while the protective device is in a failed state. This situation is known as a multiple failure. (This is a real possibility simply because the failure of the protective device is not evident, and so no-one would be aware of the need to take corrective or alternative action to avoid the multiple failure.) A multiple failure only occurs if a protected fullction fails while the protective device is in a failed state
114
Reliability-centred Maintenance 2: No action taken to shut down the protected
Time
function or to provide other p�otection
Protected funct ion Protect ive device
Failure Consequences
1: Failure of non-fail-
safe device is not evident to operators
function operating without protection because no-one knows that the protective device has f ailed
1
I I I I
3: If protected function fails here, the result is a multiple failur e
Further examples of hidden failures and the multiple failures which could follow if they are not detected are: •
vibration switches: A vibration switeh designed to shut down a large fan might be configured in such a way that its failure is hidden. How ever, this only matters if the fan vibration rises above tolerable limits (a second failure), causing the fan bearings and possibly the fan itself to disintegrate (the consequences of the multiple failure).
•
ultimate level switches: Ultimate level switches are designed to activate an alarm or shut down equipment if a primary level switch fails to operate. In other words, if an ultimate low level switch jams, there are no consequences unless the primary switch also fails (the second fail ure), in which case the vessel or tank would run dry (the consequences of the multiple failure).
•
fire hoses: The failure of a fire hose has no direct consequences. It only matters if there is a fire (a secondfai/ure), when the failed hose may re sult in the place burning down and people being killed (the consequence of the multiple failure).
Figure 5.10:
Failure of a protective device whose function is hidden
The sequence of events which leads to a multiple failure is summarised in Figure 5. 10. In the case of the relief valve, if the pressure in the vessel rises excessively while the valve is jammed, the vessel will probably explode (unless someone acts very quickly or unless there is other protection in the system). If Pump B fails while Pump C is in a failed state, the result will be total loss of pumping.
Given that failure prevention is mainly about avoiding the consequences of failure, this example also suggests that when we develop maintenance programs for hidden functions, our objective is actually to prevent - or at least to reduce the probability of the associated multiple failure. The objective of a maintenance program for a hidden function is to prevent - or at least to redllce the probability of - the associated mllltiple failllre
How hard we try to prevent the hidden failure depends on the consequen ces of the multiple failure. For example, Pumps B and C might be pumping cooling water to a nuclear re actor. In this case, if the reactor could not be shut down fast enough, the ultimate consequences of the multiple failure could be a melt-down, with catastrophic safety, environmental and operational consequences. On the other hand, the two pumps might be pumping water into a tank which has enough capacity to supply a downstream process for two hours. In this case, the consequence of the multiple failure would be that production stops after two hours, and then only if neither of the pumps could be repaired before the tank ran dry. Further analysis might suggest that at worst, this multiple failure might cost the organisation (say) £2 000 in lost production.
In the first of these examples, the consequences of the multiple failure are very serious indeed , so we would go to great lengths to preserve the inte grity of the hidden function. In the second case, the consequences of the multiple failure are purely economic, and how much it costswould influence how hard we would try to prevent the hidden failure.
115
Other typical hidden functions include emergency medical equipment, most types of fire detection, fire warning and fire fighting equipment, emergency stop buttons and trip wires, secondary containment structures, pressure and temperature switches, overload or overspeed protection devices, stand-by plant , redundant structural components, over-current circuit breakers and fuses and emergency power supply systems. The Required Availability of Hidden Functions
So far, this part of this chapter has defined hidden failures and described the relationship between protective devices and hidden functions. The next question involves a closer look at the performance we require from hidden functions. One of the most important conclusions which has been drawn so far is that the only direct consequence of a hidden failure is increased exposure to the risk of a multiple failure. Since it is the latter which we most wish to avoid, a key element of the performance required from a hidden function must be connected with the associated multiple failure. We have seen that where a system is protected by a device which is not fail-safe, a multiple failure only oceurs if the protected device fails while the protective device is in a failed state, as illustrated in Figure 5.10.
116
Failure Consequences
Reliability-centred Maintenance
So the probability of a multiple failure in any period must be given by the probability that the protected function will fail while the protective device is in a failed state during the same period. Figure 5.11 shows that this can be calculated as follows: Probability of a multiple failure
Probability of failure of the protected function
x
Average unavailability of the protective device
The tolerable probability of the multiple failure is determined by the users of the system, as discussed in the next part of this chapter and again in Appendix 3. The probability of failure of the protected function is usually a given. So if these two variables are known, the allowed unavailability of the protective device can be expressed as follows: Allowed unavailability of the protective device
Probability of a multiple failure Probability of failure of the protected function
So a crucial element of the performance required from any hidden func tion is the availability required to reduce the probability of the associated multiple failure to a tolerable level. The above discussion suggests that this availability is determined in the following three stages: •
first establish what probability the organisation is prepared to tolerate for the multiple failure
•
then determine the probability that the protected function will fail in the period under consideration (this is also known as the demand rate)
•
finally, determine what availability the hidden function must achieve to reduce the probability of the multiple failure to the required level.
When calculating the risks associated with protected systems, there is sometimes a tendency to regard the probability of failure of the protected and protective devices as fixed. This leads to the belief that the only way to change the probability of a multiple failure is to change the hardware (in other words, to modify the system), perhaps by adding more protec tion or by replacing existing components with ones which are thought to be more reliable. In fact, this beliefis incorrect, because it is usually possible to vary both the probability of failure of the protected function and (especially) the unavailability of the protective device by adopting suitable maintenance and operating policies. As a result, it is also possible to reduce the proba bility of the multiple failure to almost any desired level within reason by adopting such policies. (Zero is of course an unattainable ideal.)
117
Figure 5.11: CALCULATING THE 'PROBABILITY OF A MULTIPLE FAILURE
The *probability that a protected function will fail in any period is the inverse of its mean time between failures, as illustrated in Figure 5.11a below: Figure 5.lla: Probability
If the mean time between protected function is 4 years and one year, then the 'probability that will fail in this period is 1 in 4 Protected '-i =':ii:=-=+-------------------....·. F il a s fu n ctjo n
and protected functions
ive Protect ·..:..: ::; : c::. 7 :..: +----- ---------4I Fails dev ice .------ Measuring
The probability that the protective device will be in a failed state at any time is given by the percentage of time which it is in a failed state. This is of course measured by its unavailability (also known as downtime or fractional dead time), as shown in Figure 511b below: I<E---------------- ---- Measuring period -.-------Protected �;:: :::-::;;:: :::+--------------------� Fa ils f u nctio n Protective device
Figure 5.11b: Probability and protective devices
The probability of the multiple failure is calculated by multiplying the probability of failure of the protected function by the average unavailability of the protective device. For the case described in Figure 5.11(a) and (b) above, the probability of a multiple failure would be as indicated in Figure 511(c) below: <E----------------..---
One year
-------.--------'3>.
--- ---
Probabilityof failur e in anyone year = 1 in 4 Protected �;:: :::-::; ;::�+------...:.------t--...:.---------l� · Fa ils fu nct io n Unavailability 33% Availability 67% Protective -----�� -�����-- ------� � � �� Faood device Figure 5.11c: Probability of a multiple failure
The 'probability of a multiple failure in any one year: 1 in 4 x 1 in 3 = 1 in 12 *
See note on page 96
118
Reliability-centred Maintenance
Failure Consequences
For example, the consequences of both pumps in Figure 5.7 being in a failed state may be such that the users are prepared to tolerate a probability of multiple failure of less than 1 in 1 000 in any one year (or 10.3). Assume that it has also been estimated that if the duty pump is suitably maintained, the mean time between unanticipated failures of the duty pump can be increased to ten years, which corresponds to a probability of failure in any one year of one in ten, or 10-'. So to reduce the probability of the multiple failure to less than 10-3, the unavail ability of the stand-by pump must not be allowed to exceed 10-2, or 1%. In other words it must be maintained in such a way that its availability exceeds 99%. This is illustrated in Figure 5.12 below.
119
As before, these levels of tolerability are not meant to be prescriptive and do not necessarily reflect the views of the author. They are meant to demon strate that in any protected system, someone must decide what is tolerable before it is possible to decide on the level of protection needed, and that this assessment will differ for different systems. Part 3 of this chapter suggested that if the multiple failure could affect safety, 'someone' should be a group which includes representatives of the likely victims together with their managers. This is also true of multiple failures which have economic consequences.
Figure 5.12:
For instance, in the case of the spelling error, the 'likely victim' is the author of the correspondence. In most organisations, the consequences are likely to be no more than mild embarrassment (if anyone even notices the error). In the case of the electric motor, the person most likely to be held accountable (in other words, the 'likely victim') will either be the manager responsible for the maintenance budget, or the maintenance manager in person. In the case of loss of pumping, the larger sums involved mean that higher levels of management Should become involved in setting tolerability criteria.
In practice, the probability which is considered to be tolerable for any multiple failure depends on its consequences. In the vast majority of cases the assessment has to be made by the users of the asset. These consequen ces vary hugely from system to system, so what is deemed to be tolerable varies equally widely. To illustrate this point, Figure 5.13 suggests four possible such assessments for four different systems:
Figure 5.13 also suggests that the probabilities which any organisation might be prepared to tolerate for failures which have economic conse quences tend to decrease as the magnitude of the consequences increases. This further suggests that it should be possible for any organisation to develop a schedule of tolerable 'standard' economic risks which could in tum be used to help develop maintenance programs designed to deliver those risks. This might take the form shown in Figure 5.14.
II 1I,,,,,,nlll:rv
99%
Desired availability of a protected device
Failure of Protected Function
Failed State of Protective Device
Multiple Failure
Spelling error in interoffice memo or e-mail
Spell-checker in a word-processing program unable to detect errors
Spelling mistake undetected
10kW motor on pump B overloaded
Trip switch jammed in closed pOSition
Motor burns out: £500 to rewind
Tolerable Rate of Multiple Failure
10 per month?
1 in 50 years?
1 ---10-1 0>---:; ai 10-2 �c: .c:::: 10-3 �ffi �a; �� 10'" :g � 10-5 .QC:� £ -s �10-6 . . Up to TriVial '"
�� '" 0
(no cost)
Duty pump B fails
Stand-by pump C failed
Total loss of pumping capability: £10 000 in lost production
1 in 1 000 years?
BOiler overpressurised
Relief valves jammed shut
Boiler blows up: 10 people die
1 in 10 000 000 years?
Figure 5.13: Multiple failure rates
£100
-----�
� -� l'---
£1 000
£10 000
£100 000
Cost of anyone event
!--r--
�-
r-
£1 million £10 million and over
Figure 5.14: Tolerability of economic risk
Yet again, please note that these levels o f tolerability are not meant to be prescriptive and are not meant to be any kind of proposed universal standard. The economic risks which any organisation is prepared to tolerate are quite literally that organisation's business.
120
Reliability-centred Maintenance
Figures 5 .2 and 5 . 14 suggest that it might be possible to produce a schedule of risk which combines safety risks and economic risks in one continuum. How this might be done is discussed in Appendix 3. In some cases, it may be unnecessary indeed it is sometimes impossible -to perfonn a rigorous quantitative analysis of the probability of multiple failure in the manner described above. In such cases, it may be enough to make a judgment about the required availability of the protective device based on a qualitative assessment of the reliability of the protected func tion and the possible consequences of the multiple failure. This approach is discussed further in Chapter 8. However, if the multiple failure is parti cularly serious, then a rigorous analysis should be performed. The following paragraphs consider in more detail how it is possible to influence: - the rate at which protected functions fail - the availability of protected devices. Routine Maintenance and Hidden Functions
In a system which incorporales a non-fail-safe protective device, the pro bability of a multiple failure can be reduced as follows: •
•
reduce the rate of failure of the protected function by: * doing some sort of proactive maintenance * changing the way in which the protected function is operated * changing the design of the protected function. increase the availability of the protective device by * doing some sort of proactive maintenance * checking periodically if the protective deviee has failed * modifying the protective device.
Prevent the failure of the protected function We have seen that the probability of a multiple failure is partly based on the rate of failure of the protected function. This could almost certainly be reduced by improving the maintenance or operation of the protected device, or even (as a last resort) by changing its design. Specifically, if the failures of a protected function can be anticipated or prevented, the mean time between (unanticipated) failures of this func tion would be increased. This in turn would reduce the probability of the multiple failure.
Failure Consequences
12 1
For example, one way to prevent the simultaneous failure of Pumps B and C is to try to prevent unanticipated failures of Pump B. By reducing the number of these failures, the mean time between failures of Pump B would be increased and so the probability of the multiple failure would be correspondingly reduced, as shown in Figure 5.12.
However, bear in mind that the reason for installing a protective device is that the protected function is vulnerable to unanticipated failures with serious consequences. Secondly, if no action is taken to prevent the failure of the protective device, it will inevitably fail at some stage and hence cease to provide any protection. After this point, the probability of the multiplefailure is equal to the probability of the protectedfil1lction failing 011 its own. This situation must be intolerable, or a protective device would not have been installed to begin with. This suggests that we must at least try to find a practical way of preventing the failure of protective devices which are not fail safe. Prevent the hidden failure In order to prevent a multiple failure, we must try to ensure that the hidden function is not in a failed state if and when the protected function fails. r f a proactive task could be found which was good enough to ensure 100% availability of the protective device, then a multiple failure is theoreti cally almost impossible. For example, if a proactive task could be found which could ensure 100% avail ability of Pump C while it is in the stand-by state, then we can be sure that C would always take over if B failed. (In this case a multiple failure is only possible if the users operate Pump C while B is being repaired or replaced. However, even then the risk of the multiple failure is low, because B should be repaired quickly and so the amount of time the organisation is at risk is fairly short. Whether or not the organisation is prepared to take the risk of running Pump C while Pump B is down depends on the conse quences of the multiple failure and on whether it is possible to arrange other forms of protection, as discussed earlier.)
In practice, it is most unlikely that any proactive task would cause any function, hidden or otherwise, to achieve an availability of I OO{Yc; indefi nitely. What it must do, however, is deliver the availability needed to reduce the probability of the multiple failure to a tolerable level. For example, assume that a proactive task is found which enables Pump C to achieve an availability of 99%. If the mean time between unanticipated failures of Pump B is 10 years, then the probability of the multiple failure would be 10 (1 in 1000) in any one year, as discussed earlier.
122
Failure Consequences
Reliability-centred Maintenance
If the availability of Pump C could be increased to 99.9% then the probability of the multiple failure would be reduced to 104 (1 in 10 000), and so on.
So for a hidden failure, a proactive task is only wOlth doing if it secures the availability needed to reduce the probability of the multiple failure to a tolerable level. For hidden failures, a proactive task is worth doing if it secures the availability needed to reduce the probability of a multiple failure to a tolerable level
The ways in which failures can be prevented are discussed in Chapters 6 and 7. However, these chapters also explain that it is often impossible to find a proactive task which secures the required availability. This applies especially to the type of equipment which suffers from hidden failures. So if we cannot find a way to prevent a hidden failure, we must find some other way of improving the availability of the hidden function.
12]
Hidden Functions: The Decision Process
All the points made so far about the development of a maintenance strate gy for hidden functions can be summarised as shown in Figure 5.15:
Figure 5.15:
Identifying and developing a maintenance strategy for a hidden failure
Will the loss of function caused by this failure mode on its own become evident to the operating crew under normal circumstances?
I
I Proactive maintenance is worth dOing if it secures the availability needed to reduce the probability of a multiple failure to a tolerable level
Yes
I
Thefailure is evidenl. See Parts3t05 of this chapter
I Detect the hidden failure If it is not possible to find a suitable way of preventing a hidden failure, it is still possible to reduce the risk of the multiple failure by checking the hidden function periodically to find out if it is still working. If this check (called a 'failure-finding' task) is carried out at suitable intervals and if the function is rectified as soon as it is found to be faulty, it is still possible to secure high levels of availability. Scheduled failure-finding is discus sed in detail in Chapter 8. Mod([v the equipment In a very small number of cases, it is either impossible to find any kind of routine task which secures the desired level of availability, or it is im practical to do it at the required frequency. However, something must still be done to reduce the risk of the mUltiple failure to a tolerable level, so in these cases, it is usually necessary to 'go back to the drawing board' and reconsider the design. If the multiple failure could affect safety or the environment, redesign is compulsory. If the multiple failure only has economic consequences, the need for redesign is assessed on economic grounds. Ways in which redesign can be used to reduce the risk or to change the consequences of a multiple failure are discussed in Chapter 9.
I a suitable proactive task cannot be found, f check periodically whether the hidden function is working (do a scheduled failure-finding task)
I
If a suitable failure-finding task c annot be found: • redesign is compulsory if the multiple failur e could affect safety or the environment • if the multiple failure does not affect safety or the environment, redesign must be jus tified on economic grounds
Further Points ahout Hidden Functions
Six issues need special care when asking the first question in Figure 5. 1 5. They are as follows: • the distinction between functional failures and failure modes • the question of time • the primary and secondary functions of protective devices • what exactly is meant by 'the operating crew' • what are 'normal circumstances' • fail-safe' devices. These are all discussed in more detail in the following paragraphs.
124
Reliability-centred Maintenance
Failure Consequences
Functional failure and failure mode At this stage in the ReM process, every failure mode which is reasonably likely to cause each functional failure will already have been identified on the ReM Information Worksheet. This has two key implications: •
•
firstly, we are not asking what failures could occur. All we are trying to establish is whether each failure mode which has already been identi fied as a possibility would be hidden or evident if it did occur. secondly, we are not asking whether the operating crew can diagnose the failure mode itself. We are asking if the loss of function caused by the failure m )de will be evident under normal circumstances. (In other words, we are asking if the failure mode has any effects or symptoms which under nonnal circumstances, would lead the observer to believe that the item is no longer capable of fulfilling its intended function or at least that something out of the ordinary had occurred.)
t
For example, consider a motor vehicle which suffers from a blocked fuel line. The average driver (in other words, the average "operator") would not be able to diag nose this failure mode without expert assistance, so there might be a temptation to call this a hidden failure. However, the loss of the function caused by this failure mode is evident, because the car stops working.
The question of time There is often a temptation to describe a failure as 'hidden' if a consider able period of time elapses between the moment the failure occurs and the moment it is discovered. In fact, this is not the case. If the loss of function eventually becomes apparent to the operators, and it does so as a direct and inevitable result of this failure on its own, then the failure is treated as evident, no matter how much time elapses between the failure in ques tion and its discovery. For example, a tank fed by Pump A in Figure 5.4 may take weeks to empty, so the failure of this pump might not be apparent as soon as it occurs. This might lead to the temptation to describe the failure as hidden. However, this is not so be cause the tank runs dry as a direct and inevitable result of the failure of Pump A on its own. Therefore the fact that Pump A is in a failed state will inevitably become evident to the operating crew. Conversely, the failure of Pump C in Figure 5.7 will only become evident if Pump B also fails (unless someone makes a point of checking Pump C from time to time.) If pump B were to be operated and maintained in such a way that it is never necessary to switch on Pump C, it is possible that the failure of Pump Con its own would never be discovered.
125
This example demonstrates that time is not an issue when considering hidden failures. We are simply asking whether anyone will eventually be come aware of the fact that the failure has occurred on its own, and not if they will be aware when it occurs. Primary and secondary functions Thus far we have focused on the primary function of protective devices, which is to be capable of fulfilling the function they are dcsigned to fulfil when called upon to do so. As we have seen, this is usually after the pro tected function has failed. However, an important secondary function of many of these devices is that they should not work when nothing is wrong. For instance, the primary function of a pressure switch might be listed as follows: to be capable of transmitting a signal when pressure falls below 250 psi The implied secondary function of this switch is: • to be incapable of transmitting a signal when pressure is above 250 psi. •
The failure of the first function is hidden, but the failure of the second is evident because if it occurs, the switch transmits a spurioLls shutdown signal and the machine stops . If this is likely to occur in practice, it should be listed as a failure mode of the function which is interrupted (usually the primary function of the machinc). As a result, thcre is usually no need to list the implied second function separatcly, but the failure mode would bc listed under the relevant function if it is reasonably likely to occur. The operating crew When asking whether a failure is evident, the term operatinR crew refers to anyone who has occasion to observe the equipment or what it is doing at any time in the course of their normal daily activities, and who can be relied upon to report that it has failed. Failures can be observed by people with many different points of view. They include operators, drivers, quality inspectors, craftsmen, supervi sors, and even the tenants of buildings. However, whether any of these people can be relied upon to detect and report a failurc depends on four critical elements: •
the observer must be in a position either to detect the failure mode itself or to detect the loss of function caused by thc failure mode. This may be a physical location or access to equipment or information (including management information) which will draw attention to thc fact that
something is wrong.
126
Reliability-centred Maintenance
Failure Consequences
1 27
•
the observer must be able to recognise the condition as a failure.
'Fail-safe'devices
•
the observer must understand and accept that it is part of his or her job to report failures.
•
the observer must have access to a procedure for reporting failures.
It often happens that a protective circuit is said to be fail-safe when it is not. This usually occurs when only part of a circuit is considered instead of the circuit as a whole.
Normal circumstances Careful analysis often reveals that many of the duties performed by operators are actually maintenance tasks. It is wise to start from a zero base when considcring these tasks, because it may transpire that either the tasks or their equencies need to be radically revised. In other words, when asking if a failure will become evident to the operating crew under 'normal' circumstances, the word normal has the following meanings:
fr
•
•
that nothing is being done to prevent the failure. If a proactive task is cun'ently successfully preventing the failure, it could be argued that the failure is 'hidden' because it does not occur. However in Chapter 4 it was pointed out that failure modes and effects should be listed and the rest of the RCM process applied as if no proactive tasks are being done, because one of the main purposes of the exercise is to review whether we should be doing any such tasks in the first place.
An example is again provided by a pressure switch, this time attached to a hydro static bearing. The switch was meant to shut down the machine if the oil pressure in the bearing fell below a certain level. It emerged during discussion that if the electrical signal from the switch to the control panel was interrupted, the machine would shut down, so the failure of the switch was initially judged to be evident. However, further discussion revealed that a diaphragm inside the switch could deteriorate with age, so the switch could become incapable of sensing changes in the pressure. This failure was hidden, and the maintenance program for the switch was developed accordingly.
To avoid this problem, take care to include the sensors and the actuators in the analysis of any control loop, as well as the electrical circuit itself.
5.7 Conclusion
that no specific task is being done to detect the failure. A surprising number of tasks which already fonn part of an operator's normal duties are in fact routines designed to check if hidden functions are working. For example, pressing a button on a control panel every day to check if all the alarm lights on the panel are working is in fact a failure-finding task.
We shall see later that failure-finding tasks are covered by the RCM task selection process, so once again it should be assumed at this stage in the analysis that this task is not being done (even though the task is currently genuinely part of the operator's normal duties). This is be cause the RCM process might reveal a more effective task , or the need to do the same task at a higher or lower frequency. (Quite apart from the question of maintenance tasks, there is often con siderable doubt about what the 'normal' duties of the operating crew actually are. This occurs most often where standard operating procedures are either poorly documented or do not exist. In these cases, the RCM review process does much to help clarify what these duties should be, and can do much to help lay the foundations of a full set of operating proce dures. This applies especially to high-technology plants.)
if not...
if not...
if not...
if not...
Figure 5.16: The evaluation of failure consequences
This chapter has demonstrated how the RCM process provides a compre hensive strategic framework for managing failures. As summarised in Figure 5 . 1 6, this framework:
128
Reliability-centred Maintenance
•
classifies all failures on the basis of their consequences. In so doing it separates hidden failures from evident failures, and then ranks the con sequences of the evident failures in descending order of importance
•
provides a basis for deciding whether proactive maintenance is worth doing in each case
•
suggests what action should be taken if a suitable proactive task cannot be found.
The different types of proactive tasks and default actions are discussed in the next four chapters, together with an integrated approach to consequence evaluation and task selection.
6
Proactive Maintenance 1 : Preventive Tasks
6.1 Technical Feasibility and Proactive Tasks As mentioned in Chapter 1 , the actions which can be taken to deal with failures can be divided into the following two categories: •
proactive tasks: these are tasks undertaken before a failure occurs, in order to prevent the item from getting into a failed state . They embrace what is traditionally known as 'predictive' :md 'preventive' maintenance, although RCM uses the terms scheduled restoration, scheduled discard and on-condition maintenance
•
default actions: these deal with the failed state , and are chosen when it is not possible to identify an effective proactive task. Default actions in clude failure-finding. redesign and run-to-failure.
These two categories correspond to the sixth and seventh of the seven questions which make up the basic RCM decision process, as follows: •
what call be dOlle to predict or prevent each failure?
•
what if a suitable predictive or preventive task canllot be fOllnd?
Chapters 6 and 7 focus on the sixth question. This deals with the criteria used to decide whether proactive tasks are technically feasible. They also look in more detail at how we decide whether specific categories of tasks are worth doing. (Chapters 8 and 9 review default actions.) When we ask whether a proactive task is technically feasible, we are simply asking whether it is possible for the task to prevent or anticipate the failure in question. This has nothing to do with economics econom ics are part of the consequence evaluation process which has already been considered at length. Instead, technical feasibility depends on the techni cal characteristics of the failure mode and of the task itself. Whether or Ilot a proactive task is technically feasible depends on the technical characteristics of the failure mode and of the task
130
Two issues dominate proactive task selection from the technical view point. These are: •
the relationship between the age of the item under consideration and how likely it is to fail
•
what happens once a failure has started to occur.
The rcst of this chapter considers tasks which could apply when there is a relationship between age (or exposure to stress) and failure. Chapter 7 considers the more difficult cases where there is no such relationship.
6.2 Age and Deterioration Any physical asset which is required to fulfil a function which brings it into contact with the real world will be subjected to a variety of stresses. These stresses cause the asset to deteriorate by lowering its resistance to stress. Eventually this resistance drops to the point at which the asset can no longer deliver the desired performance in other words, it fails . This process was first illustrated in Figure 4.3, and is shown again in a slightly different form in Figure 6. 1 . Exposure to stress is measured in a variety of ways including output, distance travelled, operating cycles, calendar time or mIming time. These units are all related to time, so it is common to refer to total exposure to w o stress as the age of the item. This z « connection between stress and time :2: a:: suggests that there should be a direct o u. relationship between the rate of de a:: w terioration and the age of the item. If 0.. this is so, then it follows that the point at which failure occurs should also Figure 6.1: depend on the age of the item, as Deterioration to failure shown in Figure 6.2. However, Figure 6.2 is based on two key assumptions, as follows: • deterioration is directly proportional to the applied stress, and
1
•
Preventive Tasks
Reliability-centred Maintenance
the stress is applied consistently.
it can do (resistance to stress)
1
••• .'
131
Figure 6.2: .
Absolute predictability
If this were tme of all assets, we would be able to predict equipment life withgreatpre cision. The classical view of preventive maintenance sug gests that this can be done all we need is enough infor mation about failures. Tn reality, however, the situation is much less clear cut. This chapter starts looking at the real world by considering a situation where there is a clear relationship between age and failure. Chapter 7 moves on to a more gen eral view of reality. w o z « :2: a:: o u. a:: w 0..
Age-related Failures
Even parts which seem to be identical vary slightly in their initial re sistance to failure. The rate at which this resistance declines with age also varies . Furthermore, no two parts are subject to exactly the same stresses throughout their lives. Even when these variations are quite small, they can have a disproportionate effect on the age at which the part fails. This is illustrated in Figure 6.3, which sho�s what happens to two components that are put into service with similar resistance to failure. Figure 6.3: A realistic view of age-related failures
1
2
3
Age (x 10 000) --Jio-
4
5
6
7
8
Part B is exposed to a generally higher level of stress throughout its life than part A, so it deteriorates more quickly. Deterioration also accelerates in response to the two stress peaks at 8 000 km and 30 000 km. On the other hand, for some reason part A seems to deteriorate at a steady pace despite two stress peaks at 23 000 km and 37 000 km. So one component fails at 53 000 km and the other at 80 000 km.
132
Reliability-centred Maintenance
Preventive Tasks
This example shows that the failure age of identical parts working under apparently identical conditions varies widely . In practice, although some parts last much longer than others, the failures of a large number of parts which deteriorate in this fashion would tend to congregate around some average life, as shown in Figure 6.4. >-
t
1
2
3
Age (x 10 000) -+
A
4
5
Frequency of fai/ure and "average life"
Figure 6.5:
til
c:
Conditional
g
probability of
o
"useful life "
'8
c: o
fai/ure and
2
3
Age (x 10 OOO)-+-
4
5
6
7
8
Figure 6.6:
til c:
The effect of
S
premature
'8
c: o
failures
1
2
3
4
6
7
8
Figure 6.7:
C
Failures which a re age-related
These patterns were introduced in Chapter 1 and are discussed at much greater length in Chapter 12. The characteristic shared by patterns A and B is that they both display a point at which there is a rapid increase in the conditional probability offailure. Pattern C shows a steady increase in the probability of failure, but no distinct wear-out zone. The next three parts of this chapter consider the implications of these failure patterns from the viewpoint of preventive maintenance.
6.3 Age-Related Failures and Preventive Maintenance
If large numbers of apparently identical age-related failure modes are analysed in this fashion, it is not unusual to find a number which occur prematurely. Why this occurs is also discussed in Chapter 12. The result of such premature failures is a conditional probability curve as shown in Figure 6.6. This is the same as failure pattern B in Figure 1.5.
Age (x 10 OOO)-+-
8
.---,..-..,.---r--,--
So even when resistance to failure does decline with age, the point at which failure occurs is often much less predictable than common sense suggests. Chapter 12 explores the quantitative implications ofthis situation in more depth. It also explains that the failure frequency curve shown in Figure 6.4 can be drawn as a conditional probability of failure curve, as shown in Figure 6.5 below. (The term useful life defines the age at which there is a rapid increase in the conditional probability of failure. It is used to dis tinguish this age from the average life shown in Figure 6.4.)
o
Even this is actually a somewhat simplistic view of age-related fail ures, because there are in fact three sets of ways in which the probability of failure can increase as an item gets older. These are shown in Figure 6.7.
Figure 6.4:
g� � .2 g:§ U:OL-.
133
For centuries certainly since machines have come into widespread use mankind has tended to believe that most equipment tends to behave as shown in Figures 6.4 to 6.6. In other words, most people still tend to assume that similar items performing a similar duty will perform reliably for a period, perhaps with a small number of random early failures, and then most of the items will 'wear out' at about the same time. In general, age-related failure patterns apply to items which are very simple, or to complex items which suffer from a dominant failure mode. In practice, they are c011lIl1only found under conditions ofdirect wear (most often where equipment comes into direct contact with the product). They are also associated with fatigue, corrosion, oxidation and evaporation.
134
Wear-out characteristics most often occur where equipment comes into direct contact with the product. Age-related failures also tend to be associated with fatigue, oxidation, corrosion and evaporation.
Examples of points where equipment comes into contact with the product include furnace refractories, pump impellers, valve seats, seals, machine tooling, screw conveyors, crusher and hopper liners, the inner surfaces of pipelines, dies and so on. Fatigue affects items especially metallic items -which are subjected to reasonably high-frequency cyclic loads. The rate and extent to which oxidation and corrosion affect any item depend of course on its chemical composition, the extent to which it is protected and the environment in which it is operating. Evaporation affects solvents and the lighter frac tions of petrochemical products. For items which conform to one of the failure patterns shown in Figure 6.7, classical theory suggests that it is possible to determine an age at which it is possible to take some sort of action to prevent the failures from hap pening again in the future, or at least to reduce the consequences of the failures. Two preventive options which are available under these circumstances are scheduled restoration tasks and scheduled discard tasks. They are considered in more detail in the next two sections of this chapter.
6.4
Preventive Tasks
Reliability-centred Maintenance
Scheduled Restoration Tasks
As the name implies, scheduled restoration entails taking periodic action to restore an existing item or component to its original condition (or more accurately, to restore its original resistance to failure). Specifically: Scheduled restoration entails remanufacturing a single component or overhauling an entire assembly at or before a specified age limit, regardless of its condition at the time
Scheduled restoration tasks are also known as scheduled rework tasks. As the above definition suggests, they include overhauls which are done at pre-set intervals.
1 35
The Frequency of Scheduled Restoration Tasks
If the failure mode under consideration conforms to Pattern A or E, it is possible to identify the age at which wear-out begins . The scheduled res toration task is done at intervals slightly less than this age. In other words: The frequency of a scheduled restoration task is governed by the age at which the item or component shows a rapid increase in the conditional probability of failure.
In the case of Pattern C, at least four different restoration intervals need to be analysed to determine the optimum interval (if one exists at all). In practice, the frequency of a scheduled restoration task can only be determined satisfactorily on the basis of reliable historical data. This is seldom available when assets first go into service, so it is usually impos sible to specify scheduled restoration tasks in prior-to-service mainte nance programs. (For example scheduled restoration tasks were only assigned to seven components in the initial program developed for the Douglas DC 10). However, items subject to very expensive failure modes should be put into age exploration programmes as soon as possible to find out if they would benefit from scheduled restoration tasks. The Technical Feasibility of Scheduled Restoration
The above comments indicate that for a scheduled restoration task to be technically feasible, the first criteria which must be satisfied are that •
•
there must be a point at which there is an increase in the conditional probability of failure (in other words, the item must have a 'life') we must be reasonably sure what the life is.
Secondly, most of the items must survive to this age. If too many items fail before reaching it, the nett result is an increase in unanticipated fail ures. Not only could this have unacceptable consequences, but it means that the associated restoration tasks are done out of sequence. This in turn disrupts the entire schedule planning process. (Note that if the failure has safety or environmental consequences, all the items must survive to the age at which the scheduled restoration task is to be done, because we cannot risk failures which might hurt people or damage the environment. In this context, the comments about safe-life limits which are made in the next part of this chapter apply equally to scheduled restoration tasks.)
136
Reliability-centred Maintenance
Preventive Tasks
Finally, scheduled restoration must restore the original resistance to failure of the asset , or at least something close enough to the original con dition to ensure that the item continues to be able to fulfil its intended function for a reasonable period of time.
>-t
g� g; .2 1ir:§ u: a
For example, no-one in their right mind would try to overhaul a domestic light bulb, simply because it is not possible to restore it to its original condition (regardless of the economics of the matter). On the other hand, it could be argued that retread ing a truck tyre restores the tread to something approaching its original condition.
Scheduled restoration tasks are technically feasible if: there is an identifiable age at which the item shows a rapid increase in the conditional probability of failure • most of the items survive to that age (all of the items if the failure has safety or environmental consequences) they restore the original resistance to failure of the item.
Even if it is technically feasible, scheduled restoration might still not be worth doing because other tasks may be even more effective. Examples showing how this might occur in practice are discussed Chapter 7 . I f a more effective task cannot b e found, there i s often a temptation to select scheduled restoration tasks purely on the grounds of technical feasibility. An age limit applied to an item which behaves as shown in Figure 6.6 means that some items will receive attention before they need it while others might fail early, but the nett effect may be an overall reduc tion in the number of unanticipated failures . However even then sched uled restoration might not be worth doing, for the following reasons: •
•
as mentioned earlier, a reduction in the number of failures is not suffi cient if the failure has safety or environmental consequences, because we want to eliminate these failures altogether. if the consequences are economic, we need to be sure that over a period of time, the cost of doing the scheduled restoration task is less than the cost of allowing the failure to occur . When comparing the two, bear in mind that an age limit lowers the service life of any item, so it increases the number of items sent to the workshop for restoration. Why this is so is shown in Figure 6.S.
FUL LlF E -+ 12 months '-----r--,--.,.---i--1'Il 3 6 9 12 15
"average life"
1
18
21
24
In this example, the useful life is 1 2 months, while the average life is 1 8 months. In a period of 3 years, the failure occurs twice if no preventive maintenance is done, while the preventive task would be done three times. In other words, the preventive task has to be done 50% more often than the corrective task which would have to be performed if the failure was allowed to occur on its own. If each failure costs (say) £2 000 in lost production and repair, failures would cost £4 000 over a three year period. If each scheduled restoration task costs (say) £1 1 00, these tasks would cost £3 300 over the same period. So in this case, the task is cost-effective. On the other hand, if the average life was 24 months and all other figures remained the same, failures only occur 1 .5 times every three years, and would cost £3 000 over this period. The scheduled restoration task still costs £3 300 over the same period, so it would not be cost-effective.
•
The Effectiveness of Scheduled Restoration Tasks
Figure 6.8: "Useful life" and
Age (months) �
These points lead to the following general conclusions about the tech nical feasibility of scheduled restoration:
•
137
When considering failures which have operational consequences, bear in mind that a scheduled restoration task may itself affect operations. In most cases, this effect is likely to be less than the consequences of the failure because : •
•
the scheduled restoration task would normally be done at a time when it is likely to have the least effect on production (usually during a so called production window) the scheduled restoration task is likely to take less time than it would to repair the failure because it is possible to plan more thoroughly for the scheduled task.
If there are no operational consequences, scheduled restoration is only justified if it costs substantially less than the cost of repair (which may be the case if the failure causes extensive secondary damage).
6.5
Scheduled Discard Tasks
Again as the name implies, scheduled discard means replacing an item or component with a new one at pre-set intervals. Specifically:
138
Preventive Tasks
Reliability-centred Maintenance Scheduled discard tasks entail discarding an item or component at or before a specified age limit, regardless of its condition at the time
These tasks are done on the understanding that replacing the old compo nent with a new one will restore the original resistance to failure.
Ideally, safe-life limits should be determined before the item is put into service. It should be tested in a simulated operating cnvironment to deter mine what life is actually achieved, and a conservative fraction of this life used as the safe-life limit. This is illustrated in Figure 6.9. AGE AT WHICH FAILURES START TO OCCUR
The Frequency of Scheduled Discard Tasks
Like scheduled restoration tasks, scheduled discard tasks are only tech nically feasible if there is a direct relationship between failure and opera ting age, as shown by the graphs in Figure 6.7. The frequency at which they are done is determined on the same basis, so: The frequency of a scheduled discard task is governed by the age at which the item or component shows a rapid increase in the conditional probability of failure
In general, there is a particularly widely held belief that all items 'have a life ' , and that installing a new part before this 'life' is reached will automati cally make it 'safe '. This is not always true, so ReM takes special care to focus on safety when considering scheduled discard tasks. For this reason, ReM recognises two different types of life-limits when dealing with scheduled discard tasks. The first apply to tasks meant to avoid failures which have safety consequences, and are called safe-life limits. Those which are intended to prevent failures which do not have safety consequences are called economic-life limits. Safe-life limits Safe-life limits only apply to failures which have safety or environmental consequences so the associated tasks must prevent all failures. In other words, no failures should occur before this limit is reached. This means that safe-life limits cannot apply to items which conform to pattern A , because infant mortality means that some items must fail prematurely. In fact , they cannot apply to any failure mode where the probability of fail ure is more than zero when the item enters service. In practice, safe-life limits can only apply to failure modes which occur in such a way that no failures can be expected to occur before the wear out zone is reached.
139
1
Age �
2
4
5
Figure 6. 9: 6
7
8
Safe-life limits
There is never a perfect correlation between a test environment and the operating environment. Testing a long-lived part to failure is also costly and obviously takes a long time, so there is usually not enough test data for survival curves to be drawn with confidence . In these cases safe-life limits can be established by dividing the average by an arbitrary factor as large as three or four. This implies that the conditional probability of failure at the life limit would essentially be zero. In other words, the safe life limit is based on a 100% probability of survival to that age . The function of a safe-life limit is to avoid the occurrence of a critical failure, so the resulting discard task is worth doing only if it ensures that no failures occur before the safe-life limit . Economic-life limits Operating experience sometimes suggests that the scheduled discard of an item is desirable on economic grounds. This is known as an economic life limit. It is based on the actual age-reliability relationship of the item, rather than a fraction of the average age at failure. The only justification for an economic life limit is cost-effectiveness. In the same way that scheduled restoration increases the number of jobs passing through the workshop , so scheduled discard inereases the con sumption of the parts which are subject to discard. As a result , the cost effectiveness of scheduled discard tasks is determined in the same way as it is for scheduled restoration tasks. In general, an economic life-limit is worth applying if it avoids or reduces the operational consequences of an unanticipated failure, or if the failure which it prevents causes significant secondary damage. Clearly, we must know the failure pattern before we can assess the cost effective ness of scheduled discard tasks.
140
Reliability-centred Maintenance
For new assets, this means that a failure mode which has major econo mic consequences should also be put into an age-exploration program to find out if a l ife limit is applicable. However, as with scheduled restora tion, there is seldom enough evidence to include this type of task in an initial scheduled maintenance program. The Technical Feasibility of Scheduled Discard Tasks
The above comments indicate that scheduled discard tasks are technically feasible under the following circumstances:
Scheduled discard tasks are technically feasible if: •
there is an 'identifiable age at which the item shows a rapid increase in the conditional probability of failure • most of the items survive to that age (all of the items if the failure has safety or environmental consequences). There is no need to ask if the task will restore the original condition be cause the item is replaced with a new one.
6.6 Failures which are Not Age-related One of the most challenging developments in modem maintenance management has been the discovery that very few failure modes actually conform to any of the failure patterns shown in Figure 6.7. As discussed in the following paragraphs, this is due primarily to a combination of vari ations in applied stress and increasing complexity. Variable stress Contrary to the assumptions listed in part 2 of this chapter, deterioration is not always proportional to the applied stress, and stress is not always applied consistently. For instance, part 3 of Chapter 4 mentioned that many failures are caused by increases in applied stress, which are caused in turn by incorrect operation, incorrect assembly or external damage. Examples of such increases in stress given in Chapter 4 included operating errors (starting up a machine too quickly, accidentally putting it into reverse while it is going forward, feeding material into a process too quickly) assembly errors (over torquing bolts, misfitting parts) and external damage (lightning, the 'thousand year flood', and so on).
Preventive Tasks
14 1
In all of these cases, there is little or no re lationship between how long the asset has been in service and the likelihood of the failure occuring, This is shown in Figure 6. 10, which is basically the same as Figure 4.4 with a time dimension added. ( Ideally, Figure 6.10 'preventing ' failures of this sort should be a matter of preventing whatever causes the increase in stress levels, rather than a matter of doing anything to the asset.) In Figure 6. 1 1, the stress peak penna nently reduces resistance to failure, but does not actually cause the item to fail (an earthquake cracks a structure but does not cause it to fall down). The reduced failure resistance makes the part vulnerable to the next peak, which may or may not occur be Figure 6.11 fore the part is replaced for another reason. In Figure 6. 12, the stress peak only tem porarily reduces failure resistance (as in the case of thermoplastic materials which sof ten when temperature rises and harden again when it drops). Finally in Figure 6. 13 a stress peak acce lerates the decline of failure resistance and Figure 6.12 eventually greatly shortens the life' of the component. When this happens, the cause and effect relationship can be very difficult to establish , because the failure could occur months or even years after the stress peak. This often happens if a part is damaged during installation (which might happen if a ball-bearing is misaligned), if it is damaged priorto instalFigure 6. 13 lation (the bearing is dropped on the floor in the parts store) or if it is mistreated in service (dirt gets into the bearing). In these cases, failure prevention is ideally a matter of ensuring that maintenance and in stallation work is done correctly and that parts are looked after properly in storage.
In all four of these examples , when the items enter service it is not possible to predict when the failures will occur. For this reason, such failures are de scribed as 'random '.
142
Preventive Tasks
Reliability-centred Maintenance
143
Complexity The failure processes depicted in Figure 6.7 apply to fairly simple mech anisms . In the case of complex items, the situation becomes even less predictable. Items are made more complex to improve their performance (by incorporating new or additional technology or by automation) or to make them safer (using protective devices) . For example, Nowlan and Heap'978 cite developments in the field of civil aviation. In the 1 930's, an air trip was a slow, somewhat risky affair, undertaken in reason ably favourable weather conditions in an aircraft with a range of a few hundred miles and space for about twenty passengers. The aircraft had one or two recipro cating engines, fixed landing gear, fixed pitch propellers and no wing flaps. Today an air trip is much faster and very much safer. It is undertaken in almost any weather conditions in an aircraft with a range of thousands of miles and space for hundreds of passengers. The aircraft has several jet engines, anti-icing equip ment, retractable landing gear, moveable high-lift devices, pressure and tempe rature control systems for the cabin, extensive navigation and communications equipment, complex instrumentation and complex ancillary support systems.
In other words, better performance and greater safety are achieved at the cost of greater complexity. This is true in most branches of industry. Greater complexity means balancing the lightness and compactness needed for high performance, with the size and mass needed for durabil ity . This combination of complexity and compromise: •
increases the number of components which can fail, and also increases the number of interfaces or connections between components. This in turn increases the number and variety of failures which can occur. For example, a great many mechanical failures involve welds or bolts, while a significant proportion of electrical and electronic failures involve the connec tions between components. The more such connections there are, the more such failures there will be.
•
reduces the margin between the initial capability of each component and the desired performance (in other words, the 'can' is closer to the 'want'), which reduces scope for deterioration before failure occurs.
These two developments in turn suggest that complex items are more likely to suffer from random failures than simple items. Patterns D, E and F The combination of variable stress and erratic response to stress coupled with the increasing complexity mean that in practice, a high and rising proportion of failure modes conform to the failure patterns shown in Figure 6. 14.
E
Figure 6.14: F
Failures which are not age related
The most important characteristic of patterns D , E and F is that after the initial period, there is little orno relationship between reliability and oper ating age. In these cases, unless there is a dominant age-related failure mode, age limits do little or nothing to reduce the probability of failure. (In fact, scheduled overhauls can actually increase overall failure rates by introducing infant mortality into otherwise stable systems. This is borne out by the high and rising number of nasty accidents around the world which have occurred either while maintenance is under way or immediately after a maintenance intervention. It is also borne out by the machine operator who says that "every time maintenance works on it over the weekend, it takes us until Wednesday to get it going again".) From the maintenance management viewpoint, the main conclusion to be drawn from these failure patterns is that the idea of a wear-out age simply does not apply to random failures, so the idea of fixed interval replacement or overhaul prior to such an age cannot apply . As mentioned in Chapter I , an intuitive awareness of these facts has led some people to abandon the idea of preventive maintenance altogether . Although this can be the right thing to do for failures with minor conse quences, when the failure consequences are serious, something must be done to prevent the failures or at least to avoid the consequences . The continuing need to prevent certain types o f failure, and the grow ing inability of classical techniques to do so, are behind the growth of new types of failure management. Foremost among these are the techniques known as predictive or on-condition maintenance. These techniques are discussed at length in the next chapter.
Predictive Tasks
7
Proactive Maintenance 2: Predictive Tasks
Examples of potential failures include hot spots showing deterioration of furnace refractories or electrical insulation, vibrations indicating imminent bearing failure , cracks showing metal fatigue, particles in gearbox oil showing imminent gear fail ure, excessive tread wear on tyres, etc .
7.1 Potential Failures and On-condition Maintenance The previous chapter explained that there is often little or no relationship between how long an asset has been in service and how likely it is to fail. However, although many failure modes are not age-related, most of them give some sort of warning that they are in the process of occurring or are about to occur. If evidence can be found that something is in the final stages of failure, it may be possible to take action to prevent it from failing completely and/or to avoid the consequences. Figure 7. 1 illustrates what happens in the final stages of failure. It is called the P-F curve, because it shows how a failure starts, deteriorates to the point at which it can be detected (point 'P') and then, if it is not detected and conected, continues to deteriorate - usually at an accelerating rate until it reaches the point of functional failure CF). Point where failure starts to occur (not necessarily related to age)
'\
t
Point where we can find out that it is failing ("potential failure")
----
145
If a potential failure is detected between point P and point F in Figure 7. 1, it may be possible to take action to prevent or to avoid the consequences of the functional failure. (Whether or not it is possible to take meaningful action depends on how quickly the failure occurs, as discussed in part 2 of this chapter.) Tasks designed to detect potential failures are known as
on-condition tasks.
On-condition tasks entail checking for potential failures, so that action can be taken to prevent the functional failure or to avoid the consequences of the junctional failure On-condition tasks are so called because the items which are inspected are left in service on the condition that they continue to meet specified performance standards. This is also known as predictive maintenance (because we are trying to predict whether and possibly when the item is going to fail on the basis of its present behaviour) or condition-based maintenance (because the need for conective or consequence-avoiding action is based on as assessment of the condition of the item.)
Point where
it has failed (functional
Figure 7.1: The P-F curve Time�
7·) F
7.2 The P-F Interval In addition to the potential failure itself, we need to consider the amount of time (or the number of stress cycles) which elapse between the point at which a potential failure occurs in other words, the point at which it becomes detectable and the point where it deteriorates into a functional failure. As shown in Figure 7.2, this interval is known as the P-F interval. -
The point in the failure process at which it is possible to detect whether the failure is occuning or is about to occur is known as a potential failure.
The P-F interval is the interval between the occurrence of a potential failure and its decay into a fUllctional failure
A potential failure is an identifiable condition
which indicates that a functional failure is either about to occur or in the process of occurring In practice, there are thousands of ways of finding out if failures are in the process of occurring.
Figure 7.2: The P-F interval
146
Predictive Tasks
Reliability-centred Maintenance
The P-F interval tells us how often on-condition tasks must be done. If we want to detect the potential failure before it becomes a functional failure, the interval between checks must be less than the P-F interval.
On-condition tasks must be carried out at intervals less than the P-F interval
147
Figure 7.3 shows that if the item is inspected monthly, the nett P-F interval is 8 months. On the other hand, if it is inspected at six monthly intervals as shown in Figure 7.4, the nett P-F interval is 3 months. So in the first ease the minimum amount of time available to do something about the failure is five months longer than in Inspection interval the second, but the inspection 6 months '" task has to be done six times more often. =
The P-F interval is also known as the warning period, the lead time to failure or the failure development period. It can be measured in any units which provide an indication of exposure to stress (running time, units of output, stop-st<;lrt cycles etc), but for practical reasons, it is most often measured in terms of elapsed time. For different failure modes, it varies from fractions of a seeond to several decades. Note that if an on-condition task is done at intervals which are longer than the P-F interval, there is a chance that we will miss the failure alto gether. On the other hand, if we do the task at too small a percentage of the P-F interval, we will waste resources on the checking process. For instance, if the P-F interval for a given failure mode is two weeks, the failure will be detected if the item is checked once a week . Conversely, if it is checked once a month, it is possible to miss the whole failure process. On the other hand,
Figure 7.4: Nett P-F interval
(2)
F Time�
The nett P-F interval governs the amount of time available to take what ever action is needed to reduce or eliminate the consequences of the fail ure. Depending on the operating context of the asset, warning of incipient failure enables the users of an asset to reduce or avoid consequences in a number of ways, as follows •
if the P-F interval is three months it is a waste of effort to check the item every day.
downtime: corrective action can be planned at a time which does not disrupt operations. The opportunity to plan the com::ctive action properly also means that it is likely to be done more quickly.
In practice it is usually sufficient to select a task frequency equal to half the P-F interval. This ensures that the inspeetion will detect the potential failure before the functional failure occurs, while (in most cases) provid ing a reasonable amount of time to do something about it. This leads to
burns out, it may be possible to replace it when the machine is normally idle.
the concept of the nett P-F interval.
ure are avoided.
For example , if an electrical component is found to be overheating before it Note that in this case, the failure of the component is not 'prevented' be doomed whatever happens
•
The Nett p-�' Interval
The nett P-F interval is the minimum interval likely to elapse between the
it might
but the operational consequences of the fail
repair costs: users may be able to take action to eliminate the secondary damage which would be caused by unanticipated failures. This would reduce the downtime and the repair costs associated with the failure.
discovery of a potential failure and the occurrence of the functional failure.
For instance, a timely warning might enable users to switch a machine off
This is illustrated in Figures 7.3 and 7.4, which both show a failure with a P-F interval of nine months.
before (say ) a collapsing bearing allows a rotor to touch a stator .
Figure 7.3:
Nett P-F interval (1)
Inspection interval: 1 month
�
I I �,,,I I I I I
� P-F interval: � 9 months I
� Nett P-F
interval: 8 months
--)0-
•
safety: warning of failure provides time either to shut down a plant before the situation becomes dangerous, or to move people who might otherwise be injured out of harm's way. For instance, if a crack in a wall is discovered in good time , it may be possible to shore up the foundations and so prevent the wall from deteriorating so much that it falls down. It is highly likely that we would have to vacate the premises while this work is done, but at least we avoid the safety consequences which would arise if the wall fell down.
148
Reliability-centred Maintenance
Predictive Tasks
For an on-condition task to be technically feasible, the nett P-F interval must be longer than the time required to take action to avoid or reduce the consequences of the failure. If the nett P-F interval is too short for any sensible action to be taken, then the on-condition task is clearly not tech nically feasible. In practice , the time required varies widely. In some cases it may be a matter of hours (say until the end of an operating cycle or the end of a shift ) or even minutes (to shut down a machine or evacuate a building ). In other cases it can be weeks or even months (say until a major shutdown ).
In general, longer P-F intervals are desirable for two reasons: •
it is possible to do whatever is necessary to avoid the consequences of the failure (including planning the corrective action) in a more consid ered and hence more controlled fashion.
•
fewer on-condition inspections are required.
This explains why so much energy is being devoted to finding potential failure conditions and associated on-condition techniques which give the longest possible P-F intervals. However, note that it is possible to make lise of very short P-F intervals in certain cases.
Clearly, in these cases a task interval s hould be selected which is sub stantially less than the shortest of the likely P-F intervals. In this way, we can always be reasonably certain of detecting the potential failure before it becomes a functional failure. If the nett P-F interval associated with this minimum interval is long enough for suitable action to be taken to deal with the consequences of the failure, then the on-condition task is technically feasible. On the other hand, if the P-F interval is wildly inconsistent as some of them can be - then it is not possible to establish a meaningful task inter val, and the task in question should again be abandoned in favour of some other way of dealing with the failure.
7.3 Technical Feasibility of On-condition Tasks In the light of the above discussion, the criteria which any on-condition task must satisfy to be technically feasible can be summarised as follows:
Scheduled on-condition tasks are technically feasible if'
For example, failures which affect the balance of large fans calise serious prob lems very quickly , so on-line vibration sensors are used to shut the fans down when such failures occur . In this case , the P-F interval is very short , so monitoring
•
it is possible to define a clear potential fai/ure condition
•
the P-F interval is reasonably consistent
•
is continuous. Note also that once again, the monitoring device is being used to
The P-F curves illustrated so far in this chapter indicate that the P-F interval for any given failure is constant. In fact, this is not the case some actually vary over a quite considerable range of values, as shown in Figure 7.5.
it ispractical to monitor the item at intervals less than the
P-F interval
avoid the consequences of the failure.
P-F Interval Consistency
149
Figure 7.5: Inconsistent P-F intervals
•
Longest P-F � interval
For example , when discussing the P-F interval associated with a change in noise levels, someone might say: " This thing rattles away for anything from two weeks
7.4 Categories of On-condition Techniques The four major categories of on-condition techniques are as follows: •
condition monitoring techniques, which involve the use of specialised equipment to monitor the condition of other equipment
•
techniques based on variations in product quality
•
primary eflects monitoring techniques, which entail the intelligent use of existing gauges and process monitoring equipment
•
inspection techniques based on the human senses.
to three months before it collapses." In another case, tests might show that any thing from six months to five years elapses from the moment a crack becomes de tectable at a particular point in a structure until the moment the structure fails .
the nett P-F interval is long enough to be of some llse (in other words, long enough for action to be taken to reduce or eliminate the consequences of the flinctional failllre).
These are each reviewed in the following paragraphs.
150
Reliability-centred Maintenance
Predictive Tasks
151
Condition monitoring
Product qnality variation
The most sensitive on-condition maintenance techniques usually involve the use of some type of equipment to detect potential failures. In other words, equipment is used to monitor the condition of other equipment. These techniques are known as condition monitoring to distinguish them from other types of on-condition maintenance. Condition monitoring embraces several hundred different techniques, so a detailed study of the subject is well beyond the scope of this chapter. However, Appendix 4 provides a brief summary of about 100 of the better known techniques. All of these techniques are designed to detect failure effects (or more precisely, potential failure effects, such as changes in vibra tion characteristics, changes in temperature, particles in lubricating oil, leaks, and so on). They are classified accordingly in Appendix 4 under the following headings: • dynamic effects • particle effects • chemical effects • physical effects • temperature effects • electrical effects. These techniques can be seen as highly sensitive versions of the human senses. Many of them are now very sensitive indeed, and a few give several months (if not several years) warning of failure. However, a major limita tion of nearly every condition monitoring device is that it monitors only one condition. For instance, a vibration analyser only monitors vibration and cannot detect chemicals or temperature changes. So greater sensitiv ity is bought at the price of the versatility inherent in the human senses. The P-F intervals associated with different monitoring techniques vary from a few minutes to several months. Different techniques also pinpoint failures with different degrees of precision. Both of these factors must be considered when assessing the feasibility of any technique. In general, condition monitoring techniques can be spectacularly ef fective when they are appropriate, but when they are inappropriate they can be a very expensive and sometimes bitterly disappointing waste of time. As a result, the criteria for assessing whether on-condition tasks are techni cally feasible and worth doing should be applied especially rigorously to condition monitoring techniques.
In some industries, an important source of data about potential failures is the quality management function. Often the emergence of a defect in an article produced by a machine is directly related to a failure mode in the machine itself. Many of these defects emerge gradually, and so provide timely evidence of potential failures. If the data gathering and evaluation procedures exist already, it costs very little to use them to provide warn ing of equipment failure One popular technique which can often be used in this way is Statistical Process Control (SPC). SPC entails measuring some attributc of a product such as a dimension, filling level or packing weight, and using the meas urements to draw conclusions about the stability of the process. Figure 2.6 in Chapter 2 showed how such measurements might appear for a process which is in control and in specification. Figures 3.4 and 3.5 in Chapter 3 showed two ways in which a process could be out of control and out of specification (in other words, failed). In a great many cases, the transition from being in control to failed takes place gradually, SPC charts frequently track this transition. For instance, Figure 7.6 overleaf shows a typical SPC chart on which the readings are in control to start with. A failure mode occurs which causes the measurements to start drifting in one direction. For example , as a grinding wheel wears , the diameter of successive workpieces increases until the wheel is ad justed or replaced .
In zone 2 in Figure 7.6 the process is ,out of control but still within speci fication. (Oaklandl991 describes how itis possible to identify very gradual shifts of this sort using a 'cusum chart'.) This shift in the mean is a clearly identifiable condition which indicates that a functional failure is about to occur. In other words, it is a potential failure. If nothing is done to rectify the situation the process eventually begins to produce out-of-spec product';, as shown in zone 3 in Figure 7.6. This example describes only one of many ways in which SPC can be used to measure and manage process variability. A full description of all the techniques is well beyond the scope of this book, However, the key point to note at this stage is that if deviations on charts like these can be related directly to specific failure modes, then the charts arc sources of on-condition data which can make a valuable contribution to overall pro active maintenance efforts.
152
Predictive Tasks
Reliability-centred Maintenance
153
Normal Potential failure Figure 7. 7: Using gauges for on-condition maintenance
In control and in specification OK
Out-of-control and in specification potential failure
Out-ol-control and out-ol-spec lunctional failure
Figure 7.6: On-condition maintenance and SPC
Functional failure
The process of taking readings can be greatly simplified if gauges are marked up (or even coloured) as shown in Figure 7.7. In this case, all the operator - or anyone else needs to do is look at the gauge and report if the pointer is in the potential failure (yellow?) zone, or take more drastic action if it is in the functional failure (red?) zone. However the gauge must still be monitored at intervals which are less than the P-F interval. (For obvious reasons, this suggestion only applies to gauges which arc measuring a steady state. Also take care to ensure that gauges marked up in this way are not taken off and remounted in the wrong place.)
Primary effects monitoring
The human senses
Primary effects (speed, flow rate, pressure, temperature, power, current, etc) are yet another source of information about equipment condition. The effects can be monitored by a person reading a gauge and perhaps recording the reading manually, by a computer as part of a process control system, or even by a traditional chart recorder. The records of these effects or their derivatives are compared with refer ence information, and so provide evidence of a potential failure. How ever, in the case of the first option in particular, take care to ensure that:
Perhaps the best known on-condition inspection techniques are those based on the human senses (look, listen, feel and smell). The two main disadvan tages of using these senses to detec� potential failures are that:
•
the person taking the reading knows what the reading should be when all is well, what reading corresponds to a potential failure and what cor responds to functional failure
•
the readings are taken at a frequency which is less than the P-F interval (in other words, the frequency should be less than the time it takes the pointer on the dial to move from the potential failure level to the func tional failure level when the failure mode in question is occurring)
•
that the gauge itself is maintained in such a way that it is sufficiently accurate for this purpose_
•
by the time it is possible to detect most failures using the human senses, the process of deterioration is already quite far advanced. This means that the P-F intervals are usually short, so the checks must be done more frequently than most and response has to be rapid
•
the process is subjective, so it is difficult to develop precise inspection criteria, and the observations depend very much on the experience and even the state of mind of the observer.
However, the advantages of using these senses are as follows: •
the average human being is highly versatile and can deted a wide vari ety of failure conditions, whereas any one condition monitoring tech nique can only be used to monitor one type of potential failure
•
it can be very cost-effective if the monitoring is done by people who are at or near the assets anyway in the course of their normal duties
154 •
Reliability-centred Maintenance
Predictive Tasks
a human is able to exercise judgement about the severity of the potential t�lilure and hence about the most appropriate action to be taken, whereas a condition monitoring device can only take readings and send a signal.
Selecting the Right Category
Many failure modes are preceded by more than one - often several dif ferent potential failures, so more than one category of on-condition task might be appropriate. Each of these will have a different P-F interval, and each will require different types and levels of skill. For example , conSider a ball bearing whose failure is described as 'bearing seizes due to normal wear and tear'. Figure 7.8 shows how this failure could be preceded by a variety of potential failures , each of which could be-detected by a different oncondition tas k.
Figure 7.8: Different potential failures which can precede one failure mode
Point where failure starts to occur
-
t
/
mean that all ball bearings will exhibit these potential failures , nor
F Functional failure (Bearing seizes)
The extent to which any technique is technically feaSible and worth doing depends very much on the operating context of the bearing. For instance: the bearing may be buried so deep in the machine that it is impossible to monitor
7.5 On-condition Tasks: Some of the Pitfalls When considering the technical feasibility of on-condition maintenance, two issues need special care. They concern the distinction between poten tial and functional failures, and the distinction between potential failure and age. These issues are discussed in more detail below.
Potential and functional failures
it is only possible to detect particles in the oil if the bearing is operating in a totally
In practice, confusion often arises over the distinction between potential and functional failures. This happens because certain conditions can
bac kground noise levels may be so high that it is impossible to detect the noise made by a failing bearing
•
apply the RCM task selection criteria rigorously to determine which (if any) of the tasks is likely to be the most cost-effective way of anticipa ting the failure mode under consideration.
its vibration characteristics enclosed oil-lubricated system •
•
consider all the warnings which are reasonably likely to precede each failure mode, together with the full range of on-condition tasks which could be used to detect those warnings.
As with so much else in maintenance, the 'right' choice ultimately depends on the operating context of the asset.
will they necessarily have the same P-F intervals.
•
•
9 months
Particles which can be detected by oil analysis: P-F interval 1 - 6 months
This does not
•
In fact, if RCM is correctly applied to typical modern, complex indus trial systems, it is not unusual to find that condition monitoring as defined in this part of this chapter is technically feasible for no more than 20% of failure modes, and worth doing in less than half these cases. (All four categories of on-condition maintenance together are usually suitable for about 25 - 35% of failure modes.) This is not meant to imply that condition monitoring should not be used - where it is good it is very, very good but that we must also remember to develop suitable strategies for managing the other 90% of our failure modes. In other words, condition monitoring is only part of the answer and a fairly small part at that. So to avoid unnecessary bias in task selection, we need to:
Changes in vibration characteristics which can be detected by vibration analysis:
P-F interval 1
155
it may not be possible to reach the bearing housing to feel how hot it is .
This means that no one single category of tasks will always be more cost effective than any other. It is important to bear this in mind, because there is a tendency in some quarters to present condition monitoring in particu lar as 'the answer' to all our maintenance problems.
correctly be regarded as potential failures in one context and as functional failures in another. This is especially common in the case of leaks. For example , a minor lea k in a flanged jOint on a pipeline might be regarded as a potential failure if the pipeline is carrying water. In this case , the on-condition tas k would be 'Chec k pipe joints for lea ks'. The tas k frequency is based on the amount of time it ta kes for an 'acceptable' minor lea kto become an 'unacceptable' major lea k, and suitable corrective action would be initiated whenever a minor lea k was discovered.
156
Reliability-centred Maintenance
Predictive Tasks
However, if the same pipeline was carrying a toxic substance like cyanide, any leak at all would be regarded as a functional failure. In this case it is not feasible to ask anyone to check for leaks, so some other method would need to be found to manage the failure. This would almost certainly entail some sort of modification.
This example re-emphasises how important it is to agree what is meant by a functional failure before considering what should be done to prevent it.
The P-F interval and operating age When applying these principles for the first time, people often have diffi culty in distinguishing between the ' life' of a component and the P-F interval. This leads them to base on-condition task frequencies on the real or imagined ' life' of the item. If it exists at all, this life is usually many times greater than the P-F interval, so the task achieves little or nothing. In reality, we measure the life of a component forwards from the moment it enters service. The P-F interval is measured back from the functional failure, so the two concepts are often completely unrelated. The distinc tion is impOltant because failures which are not related to age (in other words, random failures) are as likely to be preceded by a warning as those which are not. Failures occur on a random basis
o Age (years) --.
Potential failure detected at least 2 months before functional failure
7.6 Linear and Non-linear P-F Curves Part 1 of this chapter explained that the final s tages of deterioration can be described by the P -F eurve. In this part of this chapter, we consider this curve in more detail, starting with a look at non-linear P-F curves and then going on to consider linear P-F curves.
The final stages of deterioration Figure 7. 1 on Page 144 suggests that deterioration usually accelerates in the final stages. To see why this is so, let us consider in more detail what happens when a ball bearing fails due to 'normal wear and tear' . Figure 7. 10 overleaf illustrates a typical vertically-loaded ball bearing which is rotating clockwise. The most heavily and frequently loaded part of the bearing will be the bottom of the outer race. As the bearing rotates, the inner surface of the outer race moves up and down as each ball passes over it. These cyclic movements are tiny, but they are sufficient to cause subsurface fatigue cracks which develop as shown in Figure 7.10. Figure 7. 10 also explains how these cracks eventually rise to detectable symptoms of deterioration. These are of course potential failures, and the associated P-F intervals are shown in Figure 7.8 on page 154. This example raises several further points about potential failures, as follows: •
3
4
5
Figure 7.9: Random failures and the P-F interval For example, Figure 7.9 depicts a component which conforms to a random fail ure pattern (pattern E). One of the components failed after five years, a second after six months and a third after two years . In each case, the functional failure was preceded by a potential failure with a P-F interval of four months. Figure 7.9 shows that in order to detect the potential failure, we need to do an
inspection task every 2 months. Because the failures occur on a random basis, we don't know when the next one is going to happen, so the cycle of inspections must begin as soon as the item is put into service. In other words, the timing of the inspections has nothing to do with the age or life of the component.
However, this does not mean that on-condition tasks apply only to items which fail on a random basis. They can also be applied to items which suffer age· related failures, as discussed later in this chapter.
157
in the example, the deterioration process accelerates. This suggests that if a quantitative technique such as vibration analysis is used to detect potential failures, we cannot predict when failure will occur by drawing a straight line based on just two observations. This in turn leads to the notion that after an initial deviation is ob served, additional vibration readings should be taken at progressively shorter intervals until some further point is reached at which action should be taken. In practice, this can only be done if the P-F interval is long enough to allow time for the additional readings. It also does not escape the fact that the initial readings need to be taken at a frequency which is known to be less than the P-F interval. (In fact, if the shape of the P-F curve is fairly well known and the P F interval is reasonably consistent, it should not be necessary to take additional readings after the first sign of deviation is discovered. This suggests that the process of deterioration should only be tracked by taking additional readings if the P-F curve is poorly understood or if the P-F interval is highly inconsistent.)
158
Predictive Tasks
Reliability-centred Maintenance •
159
different failure modes can often exhibit similar symptoms. For example, the symptoms described in Figure 7.10 are based on failure due to normal wear and tear. Very similar symptoms would be exhibited in the final stages of the failure of a bearing where the failure process has been initiated by dirt, lack of lubrication or brinelling .
In practice, the precise root cause of many failures can only be identi fied using sophisticated instruments. For instance, it might be possible to determine the root cause of the failure of a bearing by using a ferro graph to separate particles from the lubricating oil and examining the particles under an electron microscope. However, if two different failures have the same symptoms and if the P-F interval is broadly similar for each set of symptoms as it probably would be in the case of the bearing examples the distinction between root causes is irrelevant from the failure detection viewpoint. (The dis tinction does of course become relevant if we are seeking to eliminate the root cause of the failure.)
Strains on the outer race eventually cause subsurface fatigue cracks Cracks migrate to the surface of the outer race
•
Ball forces lubricant into the crack, causing a sliver of metal to stand proud of the surface. This is sheared off, forming a particle which can be detected by oil analysis in enclosed systems. The crater left behind changes the vibration characteristics of the bearing, and can be detected initially by vibration analysis. As the balls pass over the crater , they make it bigger. Soon the balls themselves get damaged because they are no longer rolling on a smooth surface. At some point, the bearing becomes audibly noisy , and then starts getting hotter . Deterioration continues at an accelerating pace until the balls eventually disintegrate and the bearing seizes .
Figure 7.10:
How a rolling element bearing fails due to 'normal wear and tear'
failure only becomes detectable when the fatigue cracks migrate to the surface and the surface starts breaking up. The point at which this hap pens in the life of any one bearing depends on the speed of rotation of the bearing, the magnitude of the load, the extent to which the outer race itself rotates, whether the bearing surface is damaged prior to or during installation, how hot the bearing gets in service, the alignment of the shaft relative to the housing, the materials used to manufacture the bear ing, how well it was made, etc. Effedively this combination of variables makes it impossible to predict how many operating cycles must elapse before the cracks reach the surfacc, and hence when the bem'ing will stml exhibiting the symptoms mentioned in Figure 7.10. (For those inter ested in pursuing this subject further, chaos theory in particular the 'buttertly effect' shows how tiny differences between the initial condi tions which apply to any dynamic system lead to dramatic differences after the passage of time. This may explain why minute vmiatiol)s between the initial conditions of two rolling element bearings can lead to huge differences between the ages at which they fail. See Gleickl9R7)
Deterioration accelerates in the final stages of most failures. For instance, deterioration is likely to accelerate when bolts start to loosen, when filter elements get blinded, when V -belts slacken and start slipping, when elec trical contactors overheat, when seals start to fail, when rotors become unbalanced and so on. But it does not accelerate in every case.
Predictive Tasks
Reliability-centred Maintenance
160
Linear P-F curves
n over its entire life, If an item deteriorates in a more or less linear fashio n will also be more oratio deteri of it stands to reason that the final stages sts that this is likely sugge 6.3 and or less linear. A close look at Figures 6.2 to be true of age-related failures.
�
:",�
ar in a mo e e of a tyre is likely to For example, consider tyre wear . The surfac legal mll1lmum. If thIS the s reache depth tread the until or less linear fashion a depth of tread greater than 2 minimum is (say) 2 mm, it is possible to specify nal failure is imminent. This functio that g warnin te adequa mm which provides is of course the potential failure level. then the P-F interval is the distance If the potential failure is set at (say) 3 mm, tread depth wears down from 3 mm the tyre could be expected to travel while its . 1 1 7. Figure to 2 mm, as illustrated in
16 1
Note that the P-F interval and the associated task frequency can only be deduced in this way if deterioration is linear. As we have seen, the P-F interval cannot be determined in this way if deterioration accelerates between 'P' and 'F' . A further point about linear failures concerns the point at which one should start to look for potential failures. For example, Figure 7.11 suggests that it would be a waste of time to measure the overall depth of the tyre tread at ten or twenty thousand km, because we know that it only approaches the potential failure point at 50 000 km. So perhaps we should only start measuring the tread depth of each tyre after it has passed the point where we know tread depth will be approaching 3 mm -in other words, when the tyre has been in service for more than (say) 40 000 km. However, if we want to ensure that this checking regime is adopted in practice, consider how the checks for a 4-wheeled truck would have to be planned if the actual history of a set of tyres is as follows:
Tread
Distance travelled by truck and by each tyre
Item
140 000
142 500
145 000
Left front tyre
4 7500
50 000
52 500
1 000'
Right front tyre
22 000
24 500
2 7000
29 500
Left rear tyre
12 500
2
Right rear tyre
38 000
40 500
Truck
*
t
OOOt
14 7500
4 500
7 000
43 000
45 500
Tread depth of UF tyre dropped below 3 mm and tyre replaced at depot 6 inch nail caused tyre to blow out at 13 000 km - replaced with new spare
If we are seriously going to try to ensure that the driver only checks each tyre after
Figure 7.11:
A linear P-F curve
?
service with a tread depth f (say) Figure 7.11 also suggests that if the tyre enters l based on th total distance interva P-F the t predic to le possib 12 mm, it should be instance, If the tyres last For ed. retread usually covered before the tyre has to be reasonable to conclude is it ed, retread be to have they at least50 000 km before 1 mm for every 5 000 km travelled. This that the tread wears at a maximum rate of associated on-condition task would The km. 000 amounts to a P-F interval of 5
�
call for the driver to:
report tyres whose tread 'Check tread depth every 2 500 km and depth is less than 3 mm.' detected before it exceeds he legal limit, Not only will this task ensure that wear is in this case -for the vehIcle operators km 500 2 time of but it also allows plenty s the limit. to plan to remove the tyre before it reache
�
' 'F' is only likely to be In general, linear deterioration between 'P and sically age-related intrin are nisms encountered where the failure mecha complex case. more hat somew (except in the case of fatigue, which is a later.) in This failure process is discussed in more detail
it passes 40 000 km in service, we have to devise a system which tells him to: • start
checking the UF tyre only when the truck reached 132 500 km
• check • and • but
the UF and RIR tyres when the truck reached 142 500 km
again at 145 000 km
check the RlR tyre only at 14 7500 km.
Clearly this is nonsense, because the cost of administering such a planning system would be far greater than the cost of simply asking the driver to check the tread depth of every tyre on the vehicle every 2 500 km. In other words, in this example the cost of fine-tuning the planning system would be far greater than the cost of doing the tasks. So we would simply ask the driver to check the tread depth of every tyre at 2 500 km intervals, rather than direct his attention to specific tyres.
However, if the process of deterioration is linear and the task itself is very expensive, then it might be worth ensuring that we only start checking for potential failures when it is really necessary. For instance, if an on-condition task entails shutting down and opening up a large turbine to check the turbine discs for cracks, and we are certain that deterioration
only becomes detectable after the turbine has been in service for a certain length of time (in other words, the failure is age-related), then we should only start taking the turbine out of service to check for the cracks after it has passed the age at
162
Reliability-centred Maintenance
which there is a reasonable likelihood that detectable cracks will start to emerge. Thereafter, the frequency of checking is based on the rate at which a detectable crack is likely to deteriorate into a failure. For the record, the age at which cracks are likely to start becoming detectable is known as the crack initiation life, whereas the time (or number of stress cycles) which elapse from the moment a crack becomes detectable until it grows so large that the item fails is known as the crack propagation life.
In cases like these, the cost of doing the task would be much greater than the cost of the associated planning systems, so it is worth ensuring that we only start doing the tasks when it is really necessary. However, if it is felt that this fine-tuning is worthwhile, bear in mind that the planning process has to employ two completely different timeframes, as follows: •
•
the first time-frame is used to decide when we should start doing the on condition tasks. This is the operating age at which potential failures are likely to start becoming detectable. the second time-frame governs how ofien we should do the tasks after this age has been reached. This time-frame is of course the P-F interval.
For example, it might be felt that the turbine disc is unlikely to develop any detec table cracks until it has been in service for at least 50 000 hours, but that it takes a minimum of ten thousand hours for a detectable crack to deteriorate into disc failure. This suggests that we don't need to start checking for cracks until the item has been in service for 50000hours, but thereafter it must be checked at intervals
of less than ten thousand hours.
Planning with this degree of sophistication requires a very detailed un derstanding of the failure mode under consideration, together with highly sophisticated planning systems. In practice, few failure modes are this well understood. When they are, even fewer organisations possess planning systems which can switch from one time frame to another as described above, so this issue needs to be approached with care. In closing this discussion, it must be stressed that all the curves - P-F and age-related which have been drawn in this part of this chapter have been drawn for one failure mode at a time.
Predictive Tasks
163
7.7 How to Determine the P-F Interval It is usually a fairly simple matter to determine the P-F interval for age related failure modes whose final stages of deterioration are linear. It is done by applying logic similar to that used in the tyre example above. On the other hand, the P-F interval can be surprisingly difficult to deter mine in the case of random failures where deterioration accelerates. The main problem with random failures is that we don't know when the next one is going to occur, so we don't know when the next failure mode is going to start on its way down the P-F curve. So if we don't even know where the P-F curve is going to start, how can we go about finding out how long it is? The following paragraphs review five possibilities, only the fourth and fifth of which have any merit.
Continuous observation In theory, it is possible to determine the P-F interval by continuously ob serving an item which is in service until a potential failure occurs, noting when that happens, and then continuing to observe the item until it fails completely. (Note that we cannot chart a full P-F curve by observing the item intermittently, because when we eventually discovered that it was failing we still wouldn't know precisely when the failure process started. What is more, if the P-F interval is shorter than the intermittent observa tion period we might miss the P-F curve altogether, in which case we would have to start all over again with a new item.) Clearly this approach is impractica.l, firstly because continuous obser vation is very expensive especially if we were to try to establish every P-F interval in this way. Secondly, waiting until the functional failure occurs means that the item actually has to fail. This might end up with us saying to the boss after (say) the compressor blew up: "Oh, we knew it was failing, but we just wanted to see how long it would take before it finally went so that we could determine the P-F interval!"
For instance, in the example concerning tyres, the failure process was 'normal' wear. Different failure modes (such as flat spots worn on the tyres due to emer
Start with a short interval and gradually extend it
gency braking or damage to the carcass caused by hitting kerbs) would lead to
The impracticality of the above approach leads some people to suggest that P-F intervals can be established by statting the checks at some quite short but arbitrary interval (say 10 days), and then waiting until "we find out what the interval should be", perhaps by gradually extending the intervaL Unfortunately, this is again the point at which the functional fail ure occurs, so we would still end up blowing up the compressor.
different conclusions because both the technical characteristics and the conse quences of these failure modes are different.
It is one matter to speculate on the nature of P-F curves in general, but it is quite another to determine the magnitude of the P-F interval in practice. This issue is considered in the next section of this chapter.
164
Reliability-centred Maintenance
This approach is of course potentially very dangerous, because there is also no guarantee that the initial arbitrary interval, no matter how short, will be shorter than the P-F interval to begin with (unless serious considera tion is given to the failure process itself).
Arbitrary intervals
The difficulties associated with the two approaches described above lead some people to suggest quite seriously that some arbitrary 'reason ably short' interval should be s elected for all on-condition tasks. This arbitrary approach is the least satisfactory (and the most dangerous) way to set on-conditi n task frequencies, because there is again no guarantee that the 'reasonably short' arbitrary interval will be shorter than the P-F intervaL On the other hand, the true P-F interval may be much longer than the arbitrary interval, in which case the task ends up being done much
O:
more often than necessary. For instance, if a daily task really only needs to be done once a month, that task is costing thirty times as much as it should,
Research
The best way to establish a precise P-F interval is to simulate the failure in such a way that there are no serious consequences when it eventually does occur. For example this is done when aircraft components are tested to failure on the ground rather than in the air. This not only provides data about the life of the components, as discussed in Chapter 6, but it also enables the observers to study at leisure how failures develop and how quickly this happens. However, laboratory testing is expensive and it takes time to yield results, even when it is accelerated. So it is only worth doing in cases where a fairly large number of components are at risk such as an aircraft fleet - and the failures have very serious consequences.
A rational approach
The above paragraphs indicate that in most cases, it is either impossible, impractical or too expensive to try to determine P-F intervals on an empi rical basis. On the other hand, it is even more unwise simply to take a shot in the dark. Despite these problems, P-F intervals can still be estimated with surprising accuracy on the basis of judgement and experience. Thefirst trick is to ask the right question, It is essential that anyone who is trying to determine a P-F interval understands that we are asking how quickly the item fails. In other words, we are asking how much time (or
Predictive Tasks
165
how many stress cycles) elapse from the moment the potential failure be comes detectable until the moment it reaches the functionally failed state. We are not asking how often it fails or how long it lasts. The second trick is to ask the right people people who have an inti mate knowledge of the asset, the ways in which it fails and the symptoms of each failure. For most equipment, this usually means the people who operate it, the craftsmen who maintain it and their first-line supervisors. If the detection process requires specialised instruments such as condi tion monitoring equipment, then appropriate speeialists should also take part in the analysis. -
In practice, the author has found that an effective way to crystallise thinking about P-F intervals is to provide a number of mental 'coat-hooks' on which people can hang their thoughts. For instance, one could ask: "do you think that the P-F interval is likely to be of the order of days, weeks or months?" If the answer is (say) weeks, the next step is to ask: "One, two, four or eight weeks?"
If everyone in the group achieves consensus, then the P-F interval has been established and the analysts go on to consider other task selection criteria such as the consistency of the P-F interval and whether the nett interval is long enough to avoid the failure consequences. If the group cannot achieve consensus, then it is not possible to provide a positive answer to the question "what is the P-F interval?". When this happens, the associated on-condition task must be abandoned as a way of detecting the failure mode under consideration, and the failure must be dealt with in some other way. The third trick is to concentrate on,one failure mode at a time. In other words, if the failure mode is wear, then the analysts should concentrate on the characteristics of wear, and should not discuss (say) con'osion or fatigue (unless the symptoms of the other failure modes are almost identi cal and the rate of deterioration is also very similar). Finally, it must be clearly understood by everyone taking part in such an analysis that the objective is to arrive at an on-condition task interval which is less than the P-F interval, but not so much less that resources wiII be squandered on the checking process. The effectiveness of such a group is redoubled if management ex presses an appreciation of the fact that it is made up of human beings, and that humans are not infallible, However, the analysts must also be aware that if the failure has safety consequences, thc price of getting it badly wrong could (literally) be fatal for themselves or their colleagues, so they need to take special care in this area,
166
Predictive Tasks
Reliability-centred Maintenance
167
7.S When On-condition Tasks are Worth Doing
7.9 Selecting Proactive Tasks
On-condition tasks must satisfy the following criteria to be worth doing:
It is seldom difficult to decide whether a proactive task is technically feasible. The characteristics of the failure govem this decision, and they
•
if a failure is hidden, it has no direct consequences. So an on-condition task intended to prevent a hidden failure should reduce the risk of the multiple failure to an acceptably low level. In practice, because the function is hidden, many of the potential failures which normally affect evident functions would also be hidden. What is more, much of this type of equipment suffers from random failures with very short or non existent P-F intervals, so it is fairly unusual to find an on-condition task which is technically feasible and worth doing for a hidden function. But this does not mean that one should not be sought.
•
if the failure has safety or environmental consequences, an on-condition task is only worth doing if it can be relied on to give enough warning of the failure to ensure that action can be taken in time to avoid the safety or environmental consequences.
•
if the failure does not involve safety, the task must be cost-effective, so over a period of time, the cost of doing the on-condition task must be less than the cost of not doing it. The question of cost-effectiveness applies to failures with operational and non-operational consequences,
are usually clear enough to make the decision a simple yes/no affair. Deciding whether they are worth doing usually needs more judgement. For instance, Figure 7.8 indicates that it may be teehnically feasible for two or more tasks of the same eategory to prevent the same failure mode. They may even be so closely matched in terms of cost-effectiveness that which one is chosen becomes a matter of personal preference. The situation is complicated further when tasks from two different categories are both technically feasible for the same failure mode. For example, most countr ies nowadays spec if y a m in imum legal tread dept h for t yres (usuall y about 2 mm). T yres w hic h are worn below t his dept h must e it her be replaced or retreaded . In pract ice , truck t yres -espec iall y t yres on s im ilar ve h icles in a s ingle fleet work ing t he same routes -s how a fa irl yclose relat ions h ip between age and fa ilure . Retread ing restores nearl y all t he or ig inal fa ilure res istance, so t he t yres could be sc heduled for restorat ion after t he y have covered a set d is tance . T his means t hat all t he t yres in t he truck fleet would be retreaded after t he y had covered t he spec if ied m ileage, w het her or not t he y needed it . F igure 6.4 in c hapter 6, repeated below as F igure 7. 12, could have been drawn for just suc h a fleet . T his s hows t hat in terms of normal wear, all t he t yres last between 50 000 and 80 000 km . If a sc heduled restorat ion pol ic y were to be
as follows:
adopted on t he bas is of t his informat ion, t here is a rap id increase in t he cond it ional
- Operational consequences are usually expensive, so an on-condition
before t h is age, so all of t he t yres would be retreaded at 50 000 k rn . However, if
task which reduces the rate at which the operational consequences occur is likely to be cost-effective. This is because the cost of inspection is usually low. This was illustrated in the example on pages 104 and 1 05. - The only cost of a functional failure which has non-operational con sequences is the cost of repair. Sometimes this is almost the same as the cost of correcting the potential failure which precedes it. In such cases, even though an on-condition task may be technically feasible, it would not be cost-effective, because over a period of time, the cost of the inspections plus the cost of correcting the potential failures would be greater than the cost of repairing the functional failure (see pages 107 and 108.) However, an on-condition task may be justified if the functional failure costs a lot more to repair than the potential failure, especially if the former causes secondary damage.
probab il it y of t h is fa ilure mode at 50 000 km and none of t hese fa ilures occur t h is pol ic y were adopted man y t yres wo �ld be retreaded long before it was reall y necessar y. In some cases, t yres w hic h could . have lasted as muc h as 80 000 km would be retreaded at 50 000 km, so t he y could lose up to 30 000 km of useful l ife . On t he ot her hand, as d iscussed in part 6 of t his c hapter, it is poss ible to def ine a potent ial fa ilure cond it ion for t yres related to tread dept h. C heck ing tread dept h is qu ick and eas y, so it is a s imple matter to c heck t he t yres ever y 2500 km and to arrange for t hem to be retreaded only w hen t he y need it. T his would enable t he fleet operator to get an average of 65 000 km out of his t yres (in terms of normal wear) w it hout endanger ing his dr ivers, instead of t he 50 000 km w hic h he gets if he does t he sc heduled restorat ion task descr ibed above
an increase in useful
t yre l ife of 30%. So in t h is case on-cond it ion tasks are muc h more cost -effect ive t han sc heduled restorat ion . >-
Figure 7.12:
Failure of tyres due to normal wear in a hypothetical truck fleet
t
"AVERAGE LIF E"
g� � .2 cy'ca Q) -
u: o '----._--,.._--,-- __-,...... 1
2
3
Age (x 10 000) -+
4
5
6
7
8
168
Reliability-centred Maintenance
This example suggests the following basic order of preference for selec ting proactive tasks :
On-condition tasks On-condition tasks are considered first in the task selection process, for the following reasons: •
they can nearly always be performed without moving the asset from its installed position and usually while it is in operation, so they seldom interfere with the production process. They are also easy to organise.
•
they identify specific potential failure conditions so corrective action can be clearly defined before work starts. This reduces the amount of repair work to be done, and enables it to be done more quickly.
•
by identifying equipment on the point of potential failure, they enable it to realise almost all of its useful life (as illustrated by the tyre example).
Scheduled restoration tasks If a suitable on-condition task cannot be found for a particular failure, the next choice is a scheduled restoration task. It too must be technically feasible, so the failures must be concentrated about an average age. If they are, scheduled restoration prior to this age can reduce the incidence of functional failures. This may be cost-effective for failures with major economic consequences, or if the cost of doing the scheduled restoration task is significantly lower than the cost of repairing the functional failure. The disadvantages of scheduled restoration are that: •
it can only be done when items are stopped and (usually) sent to the workshop, so the tasks nearly always affect production in some way
•
the age limit applies to all items, so many items or components which might have survived to higher ages will be removed
•
restoration tasks involve shop work, so they generate a much higher workload than on-condition tasks.
However, scheduled restoration is more conservative than scheduled dis card because it involves restoring things instead of throwing them away.
Predictive Tasks
Safe-life limit'! may be able to prevent certain critical failures, while an economic-lifc limit can reduce the frequency of functional failures that have major economic consequences. However, these tasks suffer from all the same disadvantages as scheduled restoration tasks.
Combinations of tasks For a very small number of failure modes which have safety or environ mental consequences, a task cannot be found which on its own reduces the risk of failure to an acceptably low level, and a suitable modification does not readily suggest itself. In these cases, it is sometimes pos Is an · on�ndition task sible to find a combination of tasks technically feasible (usually from two different task cate and worth doing? gories, such as an on-condition task No and a scheduled discard task), which reduces the risk of the failure to an ac Do the on-condition ceptable level. Each task is carried out task at intervals less than the P-F interval at the frequency appropriate for that task. However, it must be stressed that Is a scheduled restoration situations in which this is necessary task teChnically feasible are very rare, and care should be taken and worth doing? not to employ such tasks on a 'belt and No braces' basis.
The task selection process The task selection process is summa rised in Figure 7. 13. This basic order of preference is valid for the vast major ity of failure modes, but it does notapply in every single case. If a lower order task is clearly going to be a more eost effective method of managing a failure than a higher order task, then the lower order task should be selected.
Scheduled discard tasks Scheduled discard is usually the least cost-effective of the three proactive tasks, but where it is technically feasible, it does have a few desirable features.
169
Figure 7.13:
The task selection process
Do the scheduled restoration task at intervals less than the age limit
Do the scheduled discard task at intervals less than the age limit DefauH action depends on the failure consequences (See chapters 4, 8 and 9)
Failure-finding Tasks
8
Default Actions 1 : Failure-finding Tasks
Will the loss of function caused by this failure mode on its own become evident to the operating crew under normal circumstances?
IN 8.1
Default Actions
Previous chapters have mentioned that if a proactive task cannot be found which is both technically feasible and worth doing for any failure mode, then the default a<;tion which must be taken is governed by the consequen ces of the failure, as follows: •
•
•
if a proactive task cannot be found which reduces the risk of a failure which could atIect safety or the environment to a tolerably low level, the if a proactive task cannot be found which costs less over a period of time than a failure which has operational consequences, the initial default decision is 110 scheduled mailltenallce. (If this occurs and the opera tional consequences are still unacceptable, then the secondary default decision is again redesign). if a proactive task cannot be found which costs less over a period of time than a failure which has non-operational consequences, the initial default decision is 110 scheduled mailltenallce, and if the repair costs are too high, the secondary default decision is once again redesign.
The location of the default actions in the RCM decision framework is shown in Figure 8 . 1 opposite. At this point, we are answering the seventh of the seven questions which make up the RCM decision process : •
I
what should b e dOlle if a suitable proactive task call1lot b e foulld?
This chapter considers failure-finding. Chapter 9 deals with redesign and run-to failure, and also considers routine tasks whichfall outside the ReM decision framework such as walk-around checks_
}
"iT
Proactive maintenance is worth doing if it reduces the probability of the failure to a tolerable level
Does this failure mode have a direct adverse effect on operational capability?
IY
Proactive maintenance is worth doing if over a period of time it costs less than the cost of the operational WII�"yU"'1 ces plus the of repairing the failure
N Proactive maintenance is worth doing if over a of time it costs than the cost of repairing the failure
I
I .,----- if noL ----- ifnot. ., ----- if noL
if a proactive task cannot be found which reduces the risk of the multi ple failure associated with a hidden function to a tolerably low level, then a periodic failure-filldillg task must be performed. If a suitable failure-finding task cannot be found, then the secondary default decision is that the item may have to be redesigned.
item must be redesiglled or the process must be challged. •
Proactive maintenance is worth doing if it reduces the probability of a multiple failure to a tolerable level
Y
Could this failure mode cause a loss of function or secondary damage which could hurt or kill someone or lead to the breach of any known environmental standard?
17 1
DEFAULT
ACtiONS Figure 8. 1 : Default actions
8.2 Failure-finding Why bother?
Much of what has been written to date on the subject of maintenance stra tegy refers to three and only three types of maintenance: predictive, preventive and corrective. Predictive tasks entail checking if something is failing. Preventive maintenance means overhauling items or replacing components at fixed intervals. Corrective maintenance means fixing things either when they are found to be failing or when they have failed. However, there is a whole family of maintenance tasks which falls into none of the above categories. For example, when we periodically activate a fire alarm, we are not checking if it is failing. We are not overhauling or replacing it, nor are we repairing it. We are simply checking if it still works. Tasks designed to check whether something still works are known as failure-finding tasks orfunctional checks. (In order to rhyme with the other three families of tasks, the author and his colleagues also call them detective tasks because they are used to detect whether something has failed. )
172
Reliability-centred Maintenance
Failure-finding Tasks
Failure-finding applies only to hidden or unrevealed failures. Hidden failures in turn only affect protective devices. If RCM is correctly applied to almost any modem, complex industrial system, it is not unusual to find that up to 40% of failure modes fall into the hidden category. Furthermore, up to 80% of these failure modes require failure-finding, so up to one third of the tasks generated by comprehensive,
correctly applied maintenance strategy development programs arefailure finding tasks.
A more troubling finding is that at the time of writing, many existing maintenance programs provide for fewer than one third of protective devices to receive any attention at all (and then usually at inappropriate intervals). The eople who operate and maintain the plant covered by these programs are aware that another third of these devices exist but pay them no attention, while it is not unusual to find that no-one even knows
p
that the final third exist. This lack of awareness and attention means that most of the protective devices in industry our last line of protection when things go wrong are maintained poorly or not at all. This situation is completely untenable. If industry is serious about safety and environmental integrity, then the whole question of failure-finding needs to be given top priority as a matter of urgency. As more and more maintenance professionals become aware of the importance of this neglected area of maintenance, it is likely to become a bigger maintenance strategy issue in the next decade than pre dictive maintenance has been in the last ten years. The rest of this chapter explores this issue in some detail. Multiple failures and failure-finding
A multiple failure occurs if a protected function fails while a protective device is in a failed state. This phenomenon was illustrated in Figure 5. 10 on page 1 14. Figure 5. 1 1 on page 1 17 showed that the probability of a multiple failure can be calculated as follows: Probability of a multiple failure
Probability of failure of the protected function
x
Average unavailability of the protective device
.... 1
This led to the conclusion that the probability of a multiple failure can be reduced by reducing the unavailability of the protective device in other words, by increasing its availability. Chapter 5 went on to explain that the best way to do this is to prevent the protective device from getting into a failed state by applying some sort of proactive maintenance.
173
Chapters 6 and 7 described how to decide whether any sort of proactive maintenance is technically feasible and worth doing. However, when the criteria described in these two chapters are applied to hidden functions, it transpires that fewer than 10% of these functions are susceptible to any form of predictive or preventive maintenance. Nonetheless, although proactive maintenance is often inappropriate, it is still essential to do something to reduce the probability of the multiple failure to the required level. This can be done by checking periodically whether the hidden function is still working. For example, we cannot prevent the failure of a brake li ght bul b . So if there i s no warnin g circuit to show that a bul b ha s failed, the only way to reduce the po ssi bility that a burnt-out bul b will fail to warn other driver s of our intention s i s to check if it i s still workin g and replace it if it ha s failed.
Such checks are known as failure-finding tasks.
Scheduled failure-finding entails checking a hidden function at regular intervals to find out whether it has failed This chapter looks at key technical aspects of failure-finding, describes how to determine failure-finding intervals, defines the formal technical feasibility criteria for failure-finding and considers what should be done if a suitable failure-finding task cannot be found. Technical aspects of failure finding
The objective of failure-finding is to satisfy ourselves that a protective device will provide the required protection if it is called upon to do so. In other words, we are not checking whether the device looks OK we are checking whether it still works as it should. (This is why failure-finding tasks are also known as functional checks.) The following paragraphs consider some of the key issues in this area.
Check the entire protective system A failure-finding task must be sure of detecting all the failure modes which are reasonably likely to cause the protective device to fail. This is especially true of complex devices such as electrical circuits. In these cases, the function of the entire system should be checkedfrom sensor to actuator. Ideally, this should be done by simulating the conditions the circuit should respond to, and checking if the actuator gi ves the right response.
174
Failure-finding Tasks
Reliability-centred Maintenance
For exa mple, a pressure switch may be designed to shut down a machine i f the lubricating oil pressure drops below a certain level. Wherever possible switches li ke this should be chec ked by dropping the oil pressure to the required level and chec king whether the machine shuts down . Si milarly, a fire detection circuit should be chec ked fro m s mo ke detector to fire alar m by blowing s mo ke at the detector and chec king i f the alar m sounds.
Do not disturb Dismantling anything always creates the possibility that it will be put back together incorrectly. If this happens to a hidden function, the fact that it is hidden means that no-one will know it has been left in a failed state until the neRt check (or until it is needed). For this reason, we should always look for ways of checking the functions of protective devices with out disconnecting or otherwise disturbing them. This having been said, some devices simply have to be dismantled or removed altogether to check if they are working properly. In these cases , great care must be taken to do the task in such a way that the devices will s till work when they are returned to service. (The mathematical implica tions of the fact that a failure-tinding task might induce a failure are con sidered later in this chapter.)
It must be physically possible to check the function In a very small but still significant number of cases, it is impossible to carry out a failure-finding tasks of any sort. These are: •
where it is impossible to gain access to the protective device in order to check it (this is almost always a result of thoughtless design).
•
when the function of the device cannot be checked without destroying it (as in the case of fusible devices and rupture discs). In most such cases, other technologies are available (such as circuit breakers instead of fuses). However, in one or two cases our only options are to find some other way of managing the risks associated with untestable protection until something better comes along, or to abandon the processes concerned.
175
If a protective device has to be disabled in order to carry out a failure finding task, or if such a device is checked and found to be in a failed state, then alternative protection should be provided or the protected function should be shut down until the original protection is restored. This issue is discussed in more detail later. Failure-finding should not be carried out on systems where it is called for but would simply be too dangerous, (If society is serious about safety, it is debatable whether such systems should be allowed to exist at all.)
The frequency must be practical It must be practical to do the failure-finding task at the required intervals. However, before we can decide whether a required interval is practical, we need to determine what interval is actually ' required ' . This issue is considered next.
8.3 Failure-finding Task Intervals This section of this chapter describes how to determine the frequency of failure-finding tasks. It will start by explaining that this frequency depends on two variables - the desired availability and the frequency of failure of the protective device. It goes on to look at how we establish the 'desired' availability, and then examines different methods which can be used to establish failure-finding intervals under different circumstances. Failure-finding intervals, availability and reliability
We have seen that predictive and preventive maintenance task intervals are each based on just one variable (P-F interval and useful life respec tively). The following paragraphs will show that not one but two variables - availability and reliability - are used to set failure-finding intervals. Figure 8.2 shows a situation in which ten motorbi kes have been in service for four years. This means that the total service life of the fleet of bi kes in this period is :
Minimise risk while the task is being done It should be possible to carry out a failure-finding task without signifi cantly increasing the risk of the multiple failure.
10 bi kes x 4 years
40 years .
The bra ke light on each motorbi ke has been chec ked once a year for four years . ( This exa mple assu mes that no atte mpt is made to chec k the lights between the annual chec ks .) Over the four year period, the lights have been found to be in a
An exa mple of a borderline tas k is overspeeding so mething in order to chec k
failed state on four occasions, as shown in Figure 8.2 . So the mean ti me between
whether the overspeed protection mechanis m wor ks .
failures ( MT BF) of the bra ke lights is: 40 years in service
4 failures
=
10 years .
Reliability-centred Maintenance
176 Bike
1 998
1
2
2001
1 999
";,,
3
1
4 5
;::';;Wl
6
''''
::
EB
=
Checked/OK
®
=
Checked/failed
I l =
. A,.,
Failed during this year
7
Figure 8.2: Brake light failures
9 10
In this c ase, t he failu re- finding inte rv al o f one ye a r is equ al to 10% o f the MT BF o f ten ye a rs. Howeve r, we don 't know ex actly when e ach failed light ce ased to function. One might h ave failed the d ay afte r the l ast check, anothe r the d ay befo re the cu rrent check, and the rest at some time in between. All we know fo r su re is
�
th a e ach o f the fou r lights failed some time du ring the ye ar p receding the check . So In the absence o f any bette r in fo rm ation, we assume th at on average e ach
�
failed Iight failed h al f w ay th rough the ye ar. In othe r wo rds, on ave rage, e ch o f . the failed lights w as out o f se rvice fo r h al f a ye ar. This me ans th at ove r the fou r ye a r pe riod, ou r failed lights we re in a failed st ate fo r a tot al o f: 4 failed lights x 0.5 ye a rs e ach in a failed st ate
=
the unavailability required to carry out the failure-finding task and to effect any repairs is likely to be very small indeed relative to the unre vealed unavailability between tasks, to the extent that it will usually he
•
LEGEND
8
�
177
Failure�tinding Tasks
2 ye ars.
negligible on purely mathematical grounds •
both the failure-finding task and any repairs which might be needed should be carried out under tightly controlled conditions. Thesc condi tions should greatly reduce - if not completely eliminate - the chance of a multiple failure while the intervention is under way. This entails either shutting down the protected system or arranging alternative pro tection until the system has been fully restored. If this is done properly, the unavailability resulting from the (controlled) intervention can be ignored in any assessments of the probability of a multiple failure.
In the RCM decision process, the latter point is covered by the criteria for assessing whether a failure-finding task is worth doing. If there is a signi ficant increase in the likelihood of a multiple failure while the task is under way, the answer to the question "Does the task reduce the probabili ty of a multiple failure to a tolerable level" will be 'no' , and the RCM deci sion process defaults to the secondary default actions discussed later.
So on he b asis of the above in fo rm ation, it seems th at we c an ex pect an ave rage un av ail ability from ou r b rake lights o f: 2 ye a rs in a failed st ate
+
40 ye ars in se rvice
=
5 %.
This co rres ponds to an av ail ability o f 95 %.
The above example suggests that there is a linear correlation between the unavailability (5%), the failure-finding interval ( 1 year) and the reliability of the protective device as given by its MTB F 0 0 years), as follows: Un av ail ability
=
0.5 x failu re -finding inte rv al
+
MT BF o f the p rotective device
It can be shown that this linear relationship is valid for all unavailabilities of less than 5%, provided that the protective device conforms to an expo nential s urvival distribution (failure pattern E or random failure). (See Cox & Tait!99!, Pp 283 - 284 or Andrews & MOSS1993, Pp 1 10 - 1 12)
Calculating FFI using availability and reliability only
If we use the abbreviation 'FFJ' to describe the failure-finding interval and ' MT1VE' to describe the MTBF of the protective device, the above un availability equation can be rearranged to give the following formula: FF I
=
2 x un av ail ability x MT1VE
•••••
,.
2
This tells us that in order to determine the failure-finding interval for a single protective device, we need to know its mean time betweenfailures and the desired availability of the device (from which we can determine the unavailability to be used in the formula). Fo r inst ance, assume th at the ride rs of ou r moto rbikes decide they are not s atis fied with an av ail ability o f 95 %, and would p re fe r to see it inc re ased to 99%. The associ ated un av ail ability is 1 %. I f the MT BF o f the b rake lights st ays unch anged
Excluding task time and repair time N ote that the 'unavailability' of the protective device does not include any unavailability incurred while the failure-finding task is being carried out, nor does it include any unavailability caused by the need to repair the device if it is found to be failed. This is so for two reasons:
at fou r ye ars, checking inte rv al needs to be ch anged from once a ye ar to : FF I
=
2 x 1 % x 4 ye ars
=
2% o f 48 months
1 month .
In othe r wo rds, b ased on thei r av ail ability ex pec t ations and the existing failu re d at a, the bike rs need to check whethe r thei rb rake lights are wo rking once a month . I f they w ant an av ail ability o f 99.9%, they need to check about twice a week .
178
Reliability-centred Maintenance
Failure�finding Tasks
(Strictly s peaking, the above calculations are only valid if the brake lights on all the bikes are used about the same number of times each week. If there is a wide variation, both the MTBF and the failure-finding interval should be calculated in terms of distance travelled, or even more precise ly, in terms of the number of times the brakes - and hence the brake lights are used. However, the key point to note at this stage is the connection between the checking interval, the desired availability and the MTBF). For people who are uncomfortable with mathematical formulae, form ula (2) above can be used to develop a simple table, as follows
-<E--------
MTBF = 10
0. 1 %
0.2%
1%
2%
4%
tective device
Figure 8.4: DeSired availability of a protected device
this me�nt that the navail ability of the standby pump must not excee d 1 %, so � the availab ility of thiS pump has to be 99% or better (step 3). Figure 8.3 sugge sts that to achiev e an avai 'ability of 99% for the stand- by pump, . someo ne would need to carry out a failure -findin g task (in other words , check that It I� fully functio nal) at an Interva l of 2% of its mean time betwe en failure s. Recor ds might show that the stand- by pump has a mean time betwe en fai l u res of 8 years (or about 400 weeks ), so the failure -findin g task freque ncy should be .. 2% of 400 weeks 8 weeks 2 month s.
Figure 8.3: Failure-finding intervals, availability and reliability
•
Required availability Having established the relationship between availability, reliability and failure-finding intervals, the next issue to consider is how we decide what availability we require. Part 6 of Chapter 5 explained that this can be done in three stages, as follows :
=:
1: first ask what probability the organisation is prepared to tolerate for the multiple failure which could occur i f the hidden function was not work 2: then determine the probability that the protected function will fail in the period under consideration
3: finally determine what availability the hidden function must achieve to reduce the probability of the multiple failure to the desired level In addition to carrying out these three steps, we need to find out the mean time between failures of the hidden function. Once this has been done, we are in a position to look at Figure 8.3 and select the task frequency which cones ponds to the level of availability established in step 3. This process is illustrated in the following example: Figure 8.4 summarises the duty/stand-by pump example in Chapter 5 , where: •
•
in step 1 above, the users decided that they wanted the probability of the multi ple failure to be less than 1 in 1 000 in any one year in step 2 they established that the rate of unanticipated failures of the duty pump could be reduced to an average of 1 i n 10 years
I����i.L!4\�<
likely to need the pro
1 0%
ing when called upon to do so
"""
------
STEP TWO: Determinel estimate how often the protected function Is
99 .99% 99.95 % 99.9% 99 .5 % 99% 98% 95 %
0.02%
One year
179
Rigorous Methods for Calculating FFI
The above example suggests that it is possible to develop a sino-Ie formula for determ ning failure-finding intervals which incorp orates the vari ables conSIdered so far. In fact, this can be done by combining equations ( I ) an above, as explained in the following paragraphs. Let us begin . by defmmg a few key terms:
�11
�
� (�)
•
�
a probab lity of a multiple failure of l in I 000 000 in any one year implies a r:zean tune between multiple failures of I 000 000 years. Let us call his If this i so, then the probability of a multiple failure OCCUlTing m any one year IS l IM . (See again note on page 96)
� �r·
•
�
�
�
MF
we h ve seen that if t e demand rate of the protec ted function is (say) once m 200 years, thIS conesponds to a probability of failure for the protected function of 1 in 200 in any one year, or a
mean time between failures of the protectedjimctiofl of 200 years. Let LIS call this M ... 1' s� the probability of failure of the protected functi on in any one y , . wIll be I lMy . This is also known as the deman d rate. ED
� �:
180
FailureJinding Tasks
Reliability-centred Maintenance
ive as, before, MTIVE is the mean time between failures of the protect device and FFI is the failure-finding task interval.
•
UTIVE is the allowed unavailability of the protective device. es: If we substitute the above expressions, equation ( 1) becom
•
1/
MMF
=
(1 /M TED)
X
UT1VE
This can be rearranged as follows:
UT1VE
=
MTEO
+
Equation (2) above states that: FFI
2 x
g
UT1VE
X
2 x
MMF MT1VE
MT1V E X MTED MMF
.... . ... 2 . ...
4
So substitutin U TIVE from equation 4 into equation 2 gives: FFI
.... 3
If we apply this formula to the figures used in the duty /stand-by pump system
MMF is 1000 years, MT1VE is FFI
�y_�100
8 years and
MTEO is 10 years, so :
2 months
1000
Multipl e failure modes in a single protective device
cause Throughout this chapter, all the failure possibilities which could single each protective device to fail have been grouped together as one ve de failure mode ('stand-by pump fails ' ) . The vast majority of protecti could which modes failure the all because way, this in vices can be treated func the when checked are function to cease to cause a protective device d. checke is tion of the device as a whole of However, it is sometimes appropriate to carry out a detailed FMEA on might which modes the device in order to identify individual failure on. their own cause the device to be unable to provide the required protecti This is usually done in two sets of circumstances : •
to pro when some of the failure modes arc known to be susceptible able. active maintenance, but others are neither predictable nor prevent ion! restorat ed schedul or ition on-cond iate appropr In these cases, the and , qualify which modes failure the to discard task should be applied modes failure the of er remaind the failure-finding tasks applied to
when the protective device is new and the only failure data which are available (from data banks, component suppliers or wherever) apply to parts of the device but not to the device as a whole.
In these cases, equation (5) above can be modified to accommodate the MTBF of each component of the device.
When the failure-finding task can cause the failllre A major practical problem which affects the whole question of failure finding is that the task itself can cause the very failure which it is supposed to detect. This usually happens in one of two ways: •
5
This formula allows a failure-finding interval to be determined in a single step, as follows: mentioned above,
•
18 1
•
the task stresses the system in such a way that it eventually causes it to fail (as might be the case when a switch is tested, where the mere act of switching imposes stresses on the mechanism of the switch) if the system needs to be disturbed to do the task, there is always a chance that the person doing it will leave the system in a failed state.
In both cases, the device will be in a failed state from the moment the test is completed. If p is the probability that it will be left in such a state after a test, then p (as a decimal) will be its unavailability caused by the testing process. If MOTHER is the mean time between failu res caused by phenom. ena other than the test, it can be shown (for a single system) that: FFI
�._MOTHER
x
( MTEO ) p
MMF
(1 - p )
......
6
In this formula, the expression ( 1 - p) .can be ignored if p is less than 0.05. If the act of switching is the only cause of failure (in other words, there is no MOTHER) and if the failure conforms to an exponential survival distri bution, the probability of a multiple failure is the demand rate (in years) multiplied by the number of cycles between failures of the protective device. For example, if the demand rate is 40 years and the switch lasts an average of 600 000 cycles, then the probability of a multiple failure i s: 1 in (40 x 600 OOO)
=
1 in 24 000 000 years .
This is so because if the failure is caused only by switching, then the act of operating the switch to check if it has failed will simultaneously: •
•
enable you to find out if the last operation of the switch caused it to fail stress the switch and so create the possibility that it will fail as a result of the check.
182
solely So under this unique set of circumstances (random failure caused the ng operati s involve by operating the item), a failure- finding task which pro the on all at item to check whether it has failed will have no effect is done. In bability of a multiple failure, regardless of how often the task task at the other words, the answer to the question "Is it practical to do the L So in this required intervals?" is 'no' , because there is no suitable interva of the failure e multipl a of lity probabi the wants case, if the organisation only the years, 000 000 100 in 1 than less switch described above to be switch, the on rate demand the g way they can achieve this is by reducin switch. and/or by installing either more switches or a more reliable number of a as given are All of this il dicates that failure rates which reasons: ng followi operations should be treated with great caution, for the failure under consideration is hidden • they seldom indicate whethe r the Of evident underlying failure-pattern is age • they do not indicate whether the scheduled restoration or scheduled of form some related, in which case r it is random whethe or discard might be appropriate,
l
•
Failure-finding Tasks
Reliability-centred Maintenance
by the oper despite the previous comment, a failure mode caused solely is equally ation of a switch is likely to be age-related. If this is so, then it s the pro reduce which ed identifi be could task likely that a preventive level. d require the to bability of a multiple failure
big circuit This sugges ts that as a rule, important switches - especially should they , Rather modes. failure breakers should not be treated as single ance mainten riate approp be subjected to a detailed FMEA, and the most policy developed for each failure mode. Sources of Data for FFI Calculations
d protected Most modern industrial undertakings possess several thousan multiple fail systems, most of which incorporate hidden functions. The enough to ures associated with many of these systems will be serious . finding failureto ches approa s rigorou the of necessitate using one n functio ed protect the of failure of ility If accurate data about the probab le, availab are n functio hidden the of and the mean time betwcen failures ation is not the calculations can be performed quite quickly. If this inform e what estimat to ry available and very often it is not - then it is necessa In rare . eration these variables are likely to be in the context under consid following: cases, it might be possible to obtain data from one of the
• • •
183
the manufacturers of the equipment commercial data banks other users of similar equipment.
More often, however, the estimates have to be based on the knowledae b and experience of the people who know the most about the equipment. In many cases these are operators and maintenance craftsmen. (When using data from external sources, take special note of the operating context of the items for which the data was gathered compared to the context in which your equipment is operating. ) Once a failure-finding task frequency has been established and the tasks are being done on a regular basis, it becomes possible to verify the assumptions used to determine the frequency quite rapidly. However, this does require the keeping of absolutely meticulous records. not only about when each failure-finding task is done, but also about: •
whether or not the hidden function is found to be functional each time the task is done
•
how often the protected function fails (this can often be infened from the number of times the protected function mnkes use of the protective device for instance, from the number of times a pressure relief val ve actually has to relieve the pressure in the system).
On the basis of this information the actual mean time between failures can be calculated and, if necessary, the task frequency revised accordingly. Failure modes where the MTBF and/or the associated failure patterns are completely unknown - and a satisfactory guess cannot be made should be put into an age-exploration program right away to establish the true picture. If the situation is such that the uncertainty cannot be tolerated while the data are being gathered -- in other words, if the consequences of guessing wrong are simply too serious for the organisation (or in some cases, society as a whole) to accept - then every effort should be made to change the consequences. This in turn will nearly always necessitate some form of redesign. An Informal Approach to Setting Failure-finding Intervals
Not every hidden function is important enough to warrant the time and effort needed to do a full rigorous analysis. This applies mainly to multi ple failures which do not affect safety or the environment. It could also apply to multiple failures which could affect safety but where the protec ted function is inherently very reliable and the threat to safety is marginal.
1 84
Reliability-centred Maintenance
In these cases, it may be sufficient to take a general view of the entire protected system in its operating context, and go straight to a decision on a desired level of availability for the hidden function. This decision is then used in conjunction with the MTBF of the hidden failure to set a task inter val, using the table in Figure 8 . 3 . (Some organisations even go so far as to use an availability of95% for all hidden functions where the associated multiple failure cannot affect safety or the environment. However, gene ral policies of this nature can be dangerous so they should only be used by people who have extensive experience with this type of analysis.) Once again, if adequate records about hidden failures are not available and they seldom will be it will be necessary to guess at the MTBF's to begin with.e But again these records should be compiled as quickly as possible to validate the initial estimates. Other Methods of Calculating Failure-finding Intervals
The range of techniques for setting failure-finding intervals described so far in this chapter is by no means exhaustive. Many additional variants have been developed by the Aladon network of RCM special ists. These include formulae for: • voting systems • multiple, independent, fully redundant systems • deriving cost-optimised intervals for systems where the multiple failures do not affect safety or the environment. As this book is only intended to provide an introduction to this subject, these formulae are not included in this chapter. The Practicality of Task Intervals
The methods described so far for calculating failure-finding intervals some times produce very short or very long intervals. In some cases, these inter vals might be too long or too short, as follows: •
a very short failure-finding task interval has two main implications: - sometimes the interval is simply far too short to be practical. Exam ples would be failure-finding tasks which call for major items of plant to be shut down every few days - the task could cause habituation (which might happen if a fire alarm is tested too often). In these cases, the proposed task is rejected and we move on to the next stage of the RCM decision-making process, as discussed later.
Failure-jlnding Tasks •
1 85
we also encounter very long intervals sometimes as long as a hundred years or more. Here the process is clearly suggesting that we really need not worry about doing the task at all. In these cases the proposed 'task' should be stated as follows: 'the risk/reliability profile is such thatfailure
finding is felt to be unnecessary '. •
in rare cases, task intervals emerge which are significantly longer than the demand rate (M D). It makes little sense to carry out a failure-finding TE task at intervals (FFI) which are longer than the system is effectiv ely testing itself ( M ) , so in these cases, the answer to the questio n TED "Is it practical to do the task at the required interval?" will be 'no'. Howeve r bear in mind that if a failure-finding task is not done on a protecte d system, (and i f M T1vE is more than 4 or 5 times greater than � E. D' which is usually the case), it can be shown that: M MF MTED + MT1VE If this value of M is too low to be acceptable, then the protecti on is MF inadequate and the system will almost certainly have to be redesign ed, as discussed in the next chapter. =
8.4 The Techn ical Feasibility of Failure-finding The issues discusse d in Parts 2 and 3 of this chapter mean that for a failure finding task to be technically feasible, it should be possible to do the task at all, it should be possible to do it without increasing the risk of the multi ple failure, and it should be practica l to do the task at the required interval.
Failure-finding is technically feasible if •
it is possible to do the task
•
the task does not increase the risk of a multiple failure
•
it is practical to do the task at the required interval.
The objective of a failure-finding task is to reduce the probability of the multiple failure associate d with the hidden function to a tolerable level. It is only worth doing if it achieves this objective.
Failure-finding is worth doing if it reduces the probability of the associated multiple failure to a tolerable level
1 86
Reliability-centred Maintenance
9
Failure-finding is a Default Action!
Bear in mind that successful proactive maintenance prevents things from failing, whereas failure-finding accepts that they will spend some time albeit not very much in a failed state. This means that proactive mainte nance is inherently more conservative (in other words, safer) than failure finding, so the latter should only be specified if a more effective proactive task cannot be found. For this reason, it is wise to avoid RCM decision diagrams which put failure-finding ahead of proactive maintenance in the task selection process.
Three default actions are shown at the foot of Figure 8.1 . The first o f these - failure-finding was covered in Chapter 8 . This chapter focuses on no scheduled maintenance and redesign. It also briet1y reviews the role of
walk-around checks.
What if Failure-finding is Not Suitable?
If it transpires that a failure-finding task is not technically feasible or worth doing, we have exhausted all the possibilities which might enable us to extract the requ ired perfonnance from the existing asset. Where this leaves us is once again governed by the consequences of the multiple failure, as follows: •
if a suitable failure-finding task cannot be found and the multiple failure could affect safety or the environment, something must be changed in order to make the situation safe. In other words, redesign is compulsory
•
if a failure-finding task can not be found and the multi ple failure does not affect safety or the environment, then it is acceptable to take no action, but redesign may be justified if the multiple failure has very expensive consequenees.
This decision process is sum marised in Figure 8 . 5 . (This diagram is a fuller descrip tion of this aspect of the proc ess than the two boxes at the foot of the left hand column in Figure 8. 1 ):
Is a scheduled failure-finding task to detect the functional failure te c hnically feasible a\1d\Nort� doing?
CO!,itdllle multiple Jallyre�.ffect$afety or lh. enVlronme.nt?
Figure 8.5:
Failure-finding: the decision process
Other Default Actions
9.1
No Scheduled Maintenance
We have seen that failure-finding is the initial default action if a suitable proactive task cannot be found for a hiddenfailure. Howev er, if a suitable failure-finding task cannot be found in tum, then redesig n is the compul sory secondary default action if the multip le failure has safety or environ mental conseq uences . We have also seen that if an eviden t failure has safety or environmental consequences and a suitable preventive task can not be found, something must also be changed to make the situation safe . However, if the failure is eviden t and it does not affect safety or the environment, or if it is hidden and the multip le failure does not affect safety or the enviro nment, then the initial defaul t decisio n is to do no scheduled maintenance. In these cases, the items are left in service until a functio nal failure occurs , at which point they are repaire d or replaced. In other words, 'no scheduled mainte nance' is only valid if: •
a suitable scheduled task cannot be found for a hidden function, and the associated multiple failure does not have safety or environmental con sequences
•
a cost-effective preventive task cannot be found for failures which have operational or non-operational consequences .
Note that i f a suitable preventive task cannot be found for a failure under either of these circumstances, it simply means that we do not carry out scheduled mainte nance on that compo nent in its present form. It does not mean that we simply forget about it. As we see in the next section of this chapter, there may be circum stances under which it is worth changi ng the design of the compo nent to reduce overall costs.
188
9.2
Reliability-centred Maintenance Redesign
The question of equipment design has arisen again and again as we have traced the steps which must be followed to develop a successful mainte nance program. In this part of this chapter, we consider two general issues which affect the relationship between design and maintenance, and then
consider the part played by redesign in the task selection process. The tenn 'redesign' is used in its broadest sense in this chapter. Firstly, it refers to any change to the specification of any item of equipment . This means any action which should result in a change to a drawing or a parts list. It includes changing the specification ofa component, adding a new item, replacing an entire machine with one of a different make or type, or relocating a machine. It also means any other once-off change to a process or procedure which affects the operation of the plant. It even covers training as a method of dealing with a specific failure mode (which can be seen as 'redesigning' the capability of the person being trained.) Design and Maintenance
Changing anything is expensive. It involves the cost of developing the new idea (designing a new machine, drawing up a new operating procedure), the cost of turning the idea into reality (making a new part, buying a new machine, compiling a new training program) and the cost of implementing the change (installing the part, conducting the training program). Further indirect costs are incurred if equipment or people have to be taken out of service while the change is being implemented. There is also the risk that a change will fail to eliminate or even alleviate the problem it is meant to solve. In some cases, it may even create more problems. As a result, the whole question of modificat ions should be approached with great caution. Two issues need particular attention: nce? • what do we consider first - design or maintena desired performance. and reliability • the relationship between inherent
Which comes jlrst - redesign or maintenance? Reliability , design and maintenance are inextricably linked. This can lead to a temptation to start reviewing the design of existing equipment before considerin g its maintenance requirements. In fact, the RCM process con siders maintenance first for two reasons.
Other Default Actions
1 89
Most modifications take from six months to three years from concep tion to commissioning, depending on the cost and complexity of the new design. On the other hand, the maintenance person who is on duty today has to maintain the equipment as it exists today, not what should be there or what might be there some time in the future. So today ' s realities must be dealt with before tomorro w ' s design changes. Secondly, most organisations are faced with many more apparently desirable design improvement opportunities than are physically or eco nomically feasible. By focusing on failure consequences, RCM does much to help us to develop a rational set of priorities for these projects, especially because it separates those which are essential from those that are merely desirable. Clearly, such priorities can only be established after the review has been carried out.
Inherent reliability vs desired performance Among other things, Part 2 of Chapter 2 stressed that the inherent reI iabi lity of any asset is established by its design and by how it is made, and that maintenance cannot yield reliability beyond that inherent in the design. This led to two important conclusions. Firstly, if the inherent reliability or built-in capability of an asset is greater than the desired performance, maintenance can help achieve the desired performance. Most equipment is adequatcly specified, designed and built, so it is usually possible to develop a satisfactory maintenance program, as described in previous chapters. In other words, in most cases, RCM helps us to extract the desired performance from the asset as it is currently configured. On the other hand, if desired performance exceeds inherent reliability, then no amount of maintenance can deliver the desired performance. In these cases 'better' maintenance cannot solve the problem, so we need to look beyond maintenance for the solutions. Options include: • modifying the equipment • changing operating procedures • lowering our expectations and deciding to live with the problem. This reminds us that maintenance is not always the answer to chronic reliability problems. It also reminds us that we mllst establish as soon and as precisely as possible what we want each piece ofequipment to do in its operating context before we can starting talking sensibly about the appro priateness of its design or its maintenance requirements.
1 90
Other Default Actions
Reliability-centred Maintenance
Redesign as the Default Action
Figure 8. 1 shows that redesign appears at the bottom of all four columns of the decision diagram. In the case of failures which have safety or environmental consequences, it is the compulsory default action, and in the other three cases, it 'may be desirable' . In this part of this chapter, we consider each case in more detail, starting with the safety case.
Hidden failures In the case of hidden failures, the risk of a multiple failure can be reduced by modifying the equipment in one of four ways: •
For example, a battery used to power a s m o ke detector is a classical hidden function if n o additional p rotection is provided. However, a warning light i s fitted
Safety or environmental consequences
to most such detectors in such a way that the light goes out if the battery fails. In this way the additional p rotection makes the function of the battery evident. (Note that the light only tells us about the condition of the battery, not about the ability of the detector to detect smoke.)
Special care is needed in this area, because extra functions installed for this purpose also tend to be hidden. If too many layers of protection are added, it becomes increasingly difficult if not impossible to define sensible failure-finding tasks. A much more effective approach is to substitute an evident function for the hidden function, as explained in the next paragraph.
quately prevented. In these cases, redesign is usually undertaken with one of two objectives:
•
to reduce the probability of the failure mode occurring to a level which is acceptable. This is usually done by replacing the affected component with one which is stronger or more reliable. to change the item or the process in such a way that the failure no longer has safety or environmental conseque nces. This is most often done by installing one or more of the five types of protective devices which were categorised as follows in Chapter 2: to alert operators to abnormal conditions - to shut down the equipment in the event of a failure - to eliminate or relieve abnormal conditio ns which follow a failure and which might otherwise cause more serious damage - to take over from a function which has failed to prevent dangerous situation s from arising. Remember that if such a device is added, its maintenance requirements must also be analysed. Safety or environmental consequences can also be reduced by eliminat ing hazardous materials from a process, or even by abandoning a dangerous process altogether.
As mentioned in Chapter 5, when dealing with safety or the environment, RCM does not raise the question of econom ics. If the level of risk asso ciated with any failure is regarded as unacceptable, we are obliged either is to to prevent the failure, or to make the process safe. The alternative . unsound entally environm or unsafe be to accept conditions that are known es. industri This is no longer acceptable in most
make the hiddenfunction evident by adding another device: Certain
hidden functions can be made evident by adding another device which draws the attention of the operator to the failure of the hidden function.
If a failure could affect safety or the environment and no preventive task or combinatioq of tasks can be found which reduces the risk of the failure to an acceptable level, something must be changed, simply because we arc dealing with a safety or environmental hazard which cannot be ade
•
191
•
substitute an evident/unction/or the hidden/unction: In most cases
this means substituting a genuinely fail-safe protective device for one which is not fail-safe. This is surprisingly difficult to do in practice, but if it is done, the need for a failure-finding task falls away at once. For example, one commonly used way to warn the driver of a vehicle that his brake lights have failed is to install a warning light which is switched on if the brake lights fail. (In many cases, this 11ght i s also switched on for a short while when the ignition i s switched on. However, so are all the other lights on the dashboard. Under these circumstances one missing warning light is likely to be overlooked, so its function i s effectively hidden.) The system might also be configured i n such a way that its full function can only be tested by disabling a brake light and seeing if the warning light comes on. This i s a clumsy and invasive task which i s likely to cause more problems than it solves, so it i s likely to be dismissed o n the grounds of impracticality. The . m u ltiple failures associated with this system could have serious safety conse quences, so it i s necessary to reconsider the design. One way to eliminate this problem i s to make the function of the brake lights
and of the warning system evident. This can be done by substituting fibre-optic cables for the warning light, and mounting the cables so that the driver looks through them at the brake lights every time he uses the brakes. ( I n fact, he sees a pinprick of light at the end of each cable.) In this situation, it i s apparent to the driver if either a brake light o r a cable fails. In other words, the function of this protective device i s now evident, so fai l u re-finding is no longer necessary.
1 92 •
Reliability-centred Maintenance
Other Default Actions
substitute a more reliable (but still hidden) device/or the existing hid
---
den/unction: Figure 8 . 3 suggests that a more reliable hidden fu�ction
Duty pump -+
_
to reduce the probability of the multiple failure and increase the task intervals , giving increased protection with less effort. •
duplicate the hidden/unction: If it is not possible to find a single pro
tective device which has a high enough MTBF to give the desired level of protection, it is still possible to achieve any of the ab�ve t�ree objec tives by duplicating (or even triplicating) the hidden functIon. Let us return to the example of a duty pump with a stand-by. It was explained on page 179 that if the users want the probability of a multiple failur� to be less than 1 in 1000, and the unanticipated failure rate of the duty pump IS reduced to 1 in 10 years, then the availability of the stand-by pump has to be 99% or better. This led to the conclusion that a failure-finding task should be done on the stand-by pump every 2 months in order to achieve an availability of 99% (based on an MTBF for this pump of 8 years). .. However, now let us assume that someone has decided that the probability of a multiple failure in this system should not exceed 1 in 100 000 (or 10.5), rather than 1 in 1 000. If the mean time between unanticipated failures of the duty pump (MTED ) is unchanged at 10 years, applying formula 4 in Chapter shows that the unavailability (UT1VE) of the stand-by pump should not exceed.
�
UTIVE
=
MTED/M MF
=
10/100 000
=
10.4
So the unavailability of the stand-by pump must now not exceed 10-4 (0.010//0) . . formula 2 If the MTBF of the stand-by pump is unchanged at 8 years, applYing from Chapter 8 yields the following: FFI 2 x 104 X 8 years 14 hours Activating a stand-by pump this often is plainly impractical, so more thought has to be given to the design of this system. In fact, Figure 9.1 opposite shows that if we were to add a s�cond stand-by pump, and ensure that the availability of each stand-by pump on ItS ow� .exceeds 99% (corresponding to an unavailability of 1%, or 10-2), the probability of the multiple failure would be: =
=
10.1
x
10-2
X
10-2
=
10.5
or 1 in 100 000. Figure 8.3 suggests that this can be achieved by doing a failure finding task on each stand-by pump at the original fre�ue�cy of once every 8 . weeks. In other words, a much higher level of protection IS achieved Without changing the task interval.
---
St and-by pump 1 -+
A�aHa b H fty 99 % __ _ _ __�_ __o
.------.--
St and-by pump 2 -+
A�a Ha bi Hty 99 %0 _ _ _ _ _ _ �_ _ _ _
U_na % ty__ 1� � __ a _ ff _ a bjf_l_�
__
to reduce the probability of the multiple failure without changing the failure-finding task intervals. This increases the level of protection to increase the interval between tasks without changing the probabil ity of the multiple failure. This reduces resource requirements
One year ------
·� Pro - b-a- b- m� f f.a iW �e in any �a - - - �-���- � � �� 1 year unchanged at 1 in 10
____ __
(in other words, one which has a higher mean time between faIlures) will enable the organisation to achieve one of three objectives: _
-----
1 93
Figure 9.1:
The effect af duplicating a hidden function
__
______
Una�a ffa b m � 1 % _ _ _ _ _ _� �
________
__
/
Probability of a multiple failure in any 1 year: 1 in 10 x 1 in 100 x 1 in 100 = 1 in 100 000
Operational and non-operational consequences If a technically feasible preventive task cannot be found which is worth doing for failures with operational or non-operational consequences, the immediate default decision is no scheduled maintenance. However, it may still be desirable to modify the equipment to reduce total costs. To achieve this, the plant could be modified to: •
reduce the number of times the failure occurs, or possibly eliminate it altogether, again by making the component stronger or more reliable
•
reduce or eliminate the consequ ences of the failure (for example, by providing a stand-by capability)
•
make a preventive task cost-effective (for instance , by making a component more accessible) .
Note that in this case the failure consequ ences are purely econom ic so modifications must be cost-justified, whereas they were the compuls ory default action if the failure had safety or environmental consequ ences. There is no one way to determine whether a modification will be cost effective . Each case is governed by a different set of variable s, which in clude a before-and-after assessme nt of maintenance and operating costs, the remaining technologically useful life of the asset, the likelihood that the modification will work, the number of other projects competing for the capital resources of the company and so on. A detailed cost-benefit study which takes all these factors into account can be very time-con suming, so it is helpful to know beforehand whether this effort is likely to be worthwh ile. To help make a quick preliminary assessment, Nowlan & Heapl978 developed the decision diagram shown in Figure 9.2.
1 94
Other Default Actions
Reliability-centred Maintenance
Figure 9.2:
Decision diagram for a preliminary assessment of a proposed modification
·1�ttl� ...ell:ll1lhin�t,c\lnOI()gi�aIlY us�ful life 6,ftoeequiPhleht high"?
.[)oesthe failure have major .op�r:ational consequences? . . . Yes No .
. '
-
. Is t��cost ()fsQh�(tf.lled andior . corrective maintenance high"? Yes Redesign is not justified
··Met
thl:}t".
. ...... elim\n�ted
by:thedesi{Jnctu'lIlge?
No matter how reliable, all Yes assets are eventually super seded by new technology. I,)h�rea. hi�h pr:pbabilitYt Redesign is wiJh e�flltln� tecllnology, not justified So the first question to ask thIilMtt9.de.si91l·chang& is whether the asset under wil.l.De�u�ce�sful? consideration is going to be rendered obsolete in the near future. If it is, then it Redesign is .. �·· Doesa fdrm@{ C�$t. not justified ��"(leflt'st��y'f$fJi;)Wtin is clearly not worth modi ov:etall .Costt �dl!C�iol'l fying it. On the other hand, if it is going to be around for a while longer, the modification might have Redesignis Redesign is desirable not justified a chance to pay for itself. This is why the first question in Figure 9.2 asks:
Is the remaining technologically useful life of the equipment high? Some organisations demand that modifications should pay for them selves within a specified period say, two years. This effectively sets the operational horizon of the equipment at two years. This type of policy reduces the number of projects initiated on the basis of projected cost benefits and ensures that only projects which will pay for themselves quickly are submitted for approval. So if the answer to the first question in Figure 9.2 is 'no ' , rcdesign is probably not justified.
For example Figure 9.3 shows a stainless steel hopper which is periodically blocked by lumps . So far, the Re M process has revealed that this failure mode costs £400 in lost production every time it occurs, and that it cannot be pre vented by maintenance . It has been sugges ted that one way to eliminate the failure mode might be to install a stainless steel grid above the hopper outlet at a cost of £6 000 . If the hopper were due to be superseded within two years, it is highly unlikely that this modification would be worth doing, especially in view of the fact that several months would elapse before it could be commissioned. On the other hand, if the hopper were to remain in service for several more years, the modifica tion would be worth further consideration.
1 95
Proposed modification: a stainless steel grid
• • • • • • • • • • • • •
A
Figure 9.3:
stainless steel hopper
If the answer to the first question is 'yes', the next 4uestion to consider is whether the failure is happening often enough to be a real problem:
Is the functional failure rate high? This question eliminates items which fail so seldom that the cost of redesign would probably be greater than the benefits to be derived from it (unless of course a preventive task is the reason for a low failure rate. This is why a 'no' answer to this question does not immediately abort the modification the maintenance task itself might be so expensive that the modification is still justified.) For example if the blockage in the hopper occurred once every two or three years, no-one would pay much attention to it. If it occurred once a month, it would be worth investigating further.
If the failure rate is high, we start considering the economic implications of the failure:
Does the failure involve major operational consequences? If the answer is yes, then the question of redesign should be taken further. A 'no' to this question means that the failure only has a minor effect on operating costs, but we must still consider the maintenance costs associ ated with the failure by asking:
Is the cost of scheduled and/or corrective maintenance high? Note that this question is approached from two directions. As we have seen, we may get a 'no' answer to the failure rate 4uestion only because a very costly preventive task is preventing functional failures.
1 96
Reliability-centred Maintenance
A 'no' answer to the question of operational consequences means that failures might not be affecting operating capability, but they may result in excessive repair costs. So a 'yes' answer to either of these two ques tions brings us to the design change itself:
Are there specific costs which might be eliminated by the design change? This question refers to the operational consequences and the direct costs of proactive and/or corrective maintenance. However, if these costs are not related to a specific design feature, it i s unlikely that the problem will be solved by a design change. So a 'no' answer to this question means that it may be necessary to live with the economic consequences of the failure. On the other lwnd, if the problem can be pinned down to a specific cost element, then the economic potential of redesign is high. In the case of the hopper, it is hoped that the grid would prevent the lumps from reaching the hopper outlet, and so eliminate the cost of £400 per blockage.
But will the new design work? In other words:
Is there a high probability, with existing technology, that the modification will be succes.ljitl? Although a particular design change might be very desirable economi cally, there is a chance that it will not have the desired effect. A change directed at one failure mode may reveal other failure modes, requiring several attempts to solve the problem. Any design change which entails adding hardware also adds more failure possibilities maybe too many. So if a cold-blooded assessment of the proposed change indicates a low probability of success, the change is unlikely to be economically viable. For instance, in the case of the hopper we would need to be sure that lumps would not simply accumulate o n the grid and coagulate into a possibly much more costly p roblem in the long term.
Any proposed design change which makes it this far deserves a detailed cost-benefit study:
Does an economic trade-ojJstudy show an expected cost saving? Such a study compares the expected reduction in costs over the remaining useful life of the equipment with the costs of carrying out the modifica tion. To be on the safe side, the expected benefit should be regarded as the projected saving if the first attempt at improvement is successful, multi plied by the probability of success at the first try. Alternatively it might be considered that the design change will always be successful, but only some of the savings will be achieved.
Other Default Actions
1 97
If we are certain that the modification to the hopper will work, a discounted cash flow analysis o n the figures provided for the hopper (at a discount rate of 1 0%) shows that the modification will pay for itself * in five years if the blockage occurs four times per year, *
in seven years if it occurs three times per year and
*
in more than ten years if it occurs twice per year.
This type of justification is not necessary, of course, if the reliability characteristics of an item are the subject of contractual warranties or if the changes are needed for reasons other than cost (such as safety) .
9.3 Walk-around Checks Walk-around checks serve two purposes. The first is to spot accidental damage. These checks may include a few specific on-condition tasks for the sake of convenience, but damage in general can occur at any time and is not related to any definable level of failure resistance. As a result, there is no basis for defining an explicit potential failure condition or a predictable P-F interval. Similarly, the checks arc not based on the failure characteristics of any particular item, but are intended to spot unforeseen exceptions in failure behaviour. Walk-around checks are also meant to spot problems due to ignorance or negligence, such as hazardous materials or foreign objects left lying around, spillage, and other items of a housekeeping nature. They also give managers an opportunity to ensure that general standards of maintenance are satisfactory, and can be used to check whether maintenance routines are being done correctly. Again, there are rarely any explicit potential failure conditions and no predictable P-F interval. Some organisations distinguish between formal scheduled tasks and walk-around checks on the pretext that one is mainly technical and the other predominantly managerial, so they are sometimes done by different people. In fact it does not matter who does them, as long as both are done frequently and thoroughly enough to ensure a reasonable degree of pro tection from the consequences of the failures concerned.
The ReM Decision Diagram
10
The ReM Oecision Diagram
'0
�
0
10.1
•
in what way does each failure matter?
•
what can be done to prevent each failure? what should be done if a suitable preventive task cannot be found?
This chapter summarises the most important of these criteria. It also describes the RCM Decision Diagram, which integrates all the decision processes into a single strategic framework. This framework is shown in Figure LO. l overleaf, and is applied to each of the failure modes listed on the RCM Information Worksheet. Finally, this chapter describes the RCM Decision Worksheet, which is the second of the two key working documents used in the application of RCM (the Information Worksheet shown in Figure 4. 1 3 being the first).
10.2
!�
CCIl IUC
o.g
-m �m i!
�4)
..5£
Integrating Consequences and Tasks
Chapters 5 to 9 have provided a detailed explanation of the criteria used to answer the last three of the seven questions which make up the RCM process. These 'questions are:
•
�
0
1 99
!
'c '" u.
Z E
�,..
'"
.�."
" «
Z E
�,., .b" '"
.II(
(I)
.19
"
�
0 a..
2
Q.
The ReM Decision Process
The RCM Decision Worksheet is illustrated in Figure 1 0.2 opposite. The rest of this chapter demonstrates how this worksheet is used to record the answers to the questions in the Decision Diagram, and in the light of these answers, to record: •
what routine maintenance (if any) is to be done, how often it is to be done and by whom
•
which failures are serious enough to warrant redesign
•
cases where a deliberate decision has been made to let failures happen. Figure 10.2: The ReM Decision Worksheet
200
Reliability-centred Maintenance
The ReM Decision Diagram
m Does the failure mode have
a direct adverse effect on operational capability (output,
20 1
No
quality, customer service or operating costs in addition to the direct cost of repair)?
No Is a task to detect whether the failure is occurring or about to occur technically feasible and,worth doing?
Scheduled on-condition task Is a scheduled restoration task to reduce the failure rate technically feasible and worth dOing?
Yes Is a task to detect whether the failure is occurring or about to occur technically feasible and worth doing?
Scheduled on-condition task
Scheduled on-condition task
Is a scheduled restoration task to avoid failures technically feasible and worth dOing?
Is a scheduled discard task to reduce the failure rate technically feasible and worth
Is a scheduled discard task to avoid failures technically feasible and worth dOing?
Is a failure-finding task to detect the failure technically feasible and worth doing?
Scheduled failure-finding task
Could the multiple failure affect safety or the environment?
Is a scheduled restoration task to reduce the failure rate technically feasible and worth dOing?
Is a scheduled restoration task to reduce the failure rate technically feasible and worth doing?
Scheduled restoration task
Is a scheduled discard task to reduce the failure rate technically feasible and worth doing?
Scheduled discard task
Scheduled discard task
Scheduled discard task
Is a task to detect whether the failure is occurring or about to occur technically feasible and worth dOing?
Scheduled on-condition task
Scheduled restoration task
Scheduled restoration task
Scheduled restoration task
Is a task to detect whether the failure is occurring or about to occur technically feasible and worth dOing?
No scheduled maintenance • • •
Is a scheduled discard task to reduce the failure rate technically feasible and worth doing?
Scheduled discard task
Redesign may be desirable
Combination of tasks
No scheduled maintenance • • •
Redesign may be desirable
Redesign is compulsory
Redesign is _iii........._ No scheduled..... Redesign may be desirable compulsory maintenance
Figure 10.1:
U ll-iJ � 00 © ��J UU 1.0 � � J ® J:{2) ij�
iDJ l�\t�) i.r:3 J.\ 1\]J ©
1991 Aladon Ltd
202
Reliability-centred Maintenance
The R eM Decision Diagram
The decision worksheet is divided into sixteen columns. The columns headed F, FF and FM identify the failure mode under consideration. They are used to cross-refer the information and decision worksheets, as shown in figure 1 0. 3 below: RCMII INFORMATION WORKSHEET
© 1990ALADONLTD
SYSTEM
Failure Consequences
The precise meanings of questions H, S, E and 0 in Figure 1 0. 1 are dis cussed at length in Chapter 5. These questions are asked for each failure mode, and the answers recorded on the decision worksheet on the basis shown in Figure 1 0.4 below.
Cooung Water Pumping Systen
--=---!::�"---=SUB-SYSTEM -----------
FUNCllQNAL FAILURE' (L�ss of FunctlimJ
A Unable to transfer any water at all
Will the. loss of function caused by this failure mode ()n its own become evident to the operating crew under normal circumstances?
F.A1LURE MODE (Cause of FlllrUfe)
Yes
------
No
RCMII �S�YS�T�EM�--Cooung Water DECISION ..::... WORKSHEET SUB-S;vY<>'ST;:;:: M---E.... -c-
.-
Write the letter N -- in column Hand go to question H1
- Write the letter Y in column Hand go to question S
--
1-----IYes
© 1990ALADONLTD
203
-
Write the letter Y - in column Sand go to question S1
No ---------- Write the letter N in column S and go to question E
Figure 10.3:
Cross-referring the information and decision worksheets
Could this. failure mode cause a loss of func" tlol'1 or other damage which couldbreach any known environmental standard or regulation?
Write the letter Y -- in column E and go to question S1
---------- Write the letter N in column E and go to question 0 failure mode have a direct hal capability
The headings on the next ten columns refer to the questions on the RCM Decision Diagram in Figure 1 0. 1 , as follows: •
the columns headed H, S, E , 0 and N are used to record the answers to the questions concerning the consequences of each failure mode
•
the next three columns (headed HI, H2, H3 etc) record whether a pro active task has been selected, and if so, what type of task
•
if it becomes necessary to answer any of the default questions, the columns headed H4 and H5, or S4 are used to record the answers.
The last three columns record the task which has been selected (if any), the frequency with which it is to be done and who has been selected to do it. The 'proposed task' column is also used to record the cases where re design is required or it has been decided that the failure mode does not need scheduled maintenance. In the following paragraphs, each of these four sections of the decision worksheet is reviewed in the context of the associated questions on the Decision Diagram.
---------- Write the letter N in column 0 and go to question Nt
-
Write the letter Y - in column 0and go to question 01
Figure 10.4: Using the decision worksheet to record failure consequences
Figure 1 0.5 shows how the answers to these questions are recorded on the decision worksheet. Note that: •
each failure mode is dealt with in terms of one category of consequen ces only. So if it is classified as having environmental consequences, we do not also evaluate its operational consequences (at least when per forming the first analysis of any asset). This means that if for instance a 'Y' is recorded in column E, nothing is recorded in column O.
•
once the consequences of the failure mode have been categorised, the next step is to seek a suitable preventive task. Figure 7 . 5 also summa rises the criteria used to decide whether such tasks are worth doing.
204
Reliability-centred Maintenance
Inform.all.on Consequence. evaluation reference � � ,,-..-.. -,.._--:..
F
..
FF FM
H�rS�rE
0
N I___________ �---------- A hidden failure: To be worth doing, any preventive task must reduce the risk of a multiple failure to an acceptable level
3
A
1
5
B
2
Y
Y
2
C
4
Y
N
1
A
5
Y
N
1
B
3
Y
N
�------�-------
Y
N
N
- Safety consequences: To be worth doing, any preventive task must reduce the risk of this failure o n its own to an acceptable level
- --
---- ---
Y
N
Environmental consequences: To be worth doing, any preventive task must reduce the risk of this failure o n its own to an acceptable level
- Operational consequences: To be worth doing, over a period of time any preventive task must cost less than the cost of the operational consequences plus the cost of repair of the failure which it i s meant to prevent
---
I
----
Non-operational consequences: To be worth doing, over a period of time any preventive task must cost less than the cost of repairing the failure which it i s meant to prevent
The ReM Decision Diagram
205
In each case, a task is only suitable if it is worth doing and technically feasible. Chapters 6 and 7 explained in detail how to establish whether a task is technically feasible,. These criteria are summarised in Figure 1 0.6. In essence, for a task to be technically feasible and worth doing it must be possible to provide a positive answer to all of the questions shown in Figure 1 0.6 which apply to that category of tasks, and the task must fulfil the 'worth doing' criteria in Figure 1 0 . 5 . If the answer to any of these questions is 'no' or unknown, then that task as a whole is rejected. If all of the questions can be answered positively, then a Y is recorded in the appropriate column. Figure ZO.6:
H1 H2 H3
S1 82 S3
Technical feasibility criteria
01 02 03 N1 N2 N3 Y
N
......--
Y
-- - --- -
-
-----
---
Figure ZO.5:
Failure consequences a summary
Is a task to detect whether a failure is occurring or about to occur technically feasible?: Is there a clear potential failure condition? What is it? What is the P-F i nterval? Is this interval long enough to be of any use? I s the P-F interval reasonably con sistent? Is it p ractical to monitor the item at intervals less than the P-F i nterval? Is a scheduled restoration task to reduce the failure rate (avoid all failures in the case of safety) technically feasible? Is there an age at which there is a rapid increase i n t h e conditional probability o f failure? What is this age? Do most of the items survive to this age (aI/ in the
case of safety or environmental consequences)? Is it possible to restore the original resistance to failure of the item?
Proactive Tasks
The eighth to tenth columns on the decision worksheet are used to record whether a proactive task has been selected, as follows: •
the column headed H liS 1 10 lIN I i s used to record whether a suitable on-condition task could be found to anticipate the failure mode in time to avoid the consequences
•
the column headed H2/S2/02IN2 is used to record whether a suitable scheduled restoration task could be found to prevent the failures
•
the column headed H3/S3/03IN3 is used to record whether a suitable scheduled discard task could be found to prevent the failures.
N
N
Y --- Is a scheduled discard task to reduce the failure rate (avoid all failures in the case of safety) techni cally feasible? Is there an age at which there is a rapid i ncrease i n t h e conditional probability o f failure? What is this age? Do most of the items survive to this age (all in the
I
case of safety or environmental consequences)?
If a task is selected, a description of the task and the frequency with which it must be done are recorded as explained later in this chapter, and the analysts move on to the next failure mode. However, as mentioned in Chapter 7, bear in mind that if it seems that a lower order task may be more cost-effective than a higher order task, then the lower order task should also be considered and the more effective of the two chosen.
206
Reliability-centred Maintenance
The ReM Decision Diagram
The Default Questions
207
For example, consider a situation where an on-condition task has been specified
The columns headed H4, H5 and S4 on the decision worksheet are used to record the answers to the three default questions. The basis on which these questions are answered is summarised in Figure 1 0.7. (Note that the default questions are only asked if the answers to the previous three questions are all 'no ' .) Figure 10.7:
The default questions Y ---------- Is a failure-finding task technically
for a rOiling element bearing. Chapter 7 explained how such bearings can suffer from a variety of potential failure conditions, including noise, vibration, heat, wear and so on. Many machines have more than one and often several such bearings. Consequently, at the very least, the 'proposed task' should specify which bearing . is to be checked for what condition. In other words, if a particular bearing is to be checked for noise, the proposed task should read 'check bearing X for audible noise', and not just 'check bearing'.
This issue is discussed in more detail in the next chapter. If the decision process calls for a design change, then the proposed task should provide a brief description of the design change. The actual form of the new design should be left to the designers.
I I I feasible and worth doing? Record yes if it is possible to do the task and it is practical to do it at the required
a guard has to be redesigned for safety reasons, the 'proposed task' should state
frequency andit reduces the risk of the multiple failure to an acceptable level.
something like 'more secure fastening mechanism required for guard'. Do not
4 B 4 N 4 C 2 N
J
N N N N Y I__ Could the multiple failure affect N N N N N I----� safety or the environment? I I (This question is only asked if the
answer to question H4 is no.) If the answer to this question is yes, r edesign is compulsory. If the answer is no, the default action is no
nance, but redesign may be desirable. 5 B 2 Y Y 2 A 5 Y Y
N N N N N N
Y N
I 1-
s cheduled maint e
Is a combination of tasks technically feasible and worth doing?
or more preventive tasks will reduce t h e risk o f the failure t o an acceptable level (this is very rare). If the answer is no, r edesign is compulsory.
1 1
I
Yes if a combination of any two
I I I I
A 5 Y N N Y N N N -------------- In these two cases, the consequences B 3 Y N N N N N N ,____'___L__!_ of the failure are purely economic and I I I I no suitable preventive task has been found. As a result, the initial default decision is no s ch eduled maint enance, but redesign may be desirable.
Proposed Task
If a proactive task or a failure-finding task has been selected during the decision-making process, a description of the task should be recorded in the column headed 'proposed tas k ' . Ideally, the task should be described as precisely on the decision worksheet as it will be on the document which reaches the person doing the task. If this is not possible, then the task should at least be described in enough detail to make the intent absolutely clear to whoever writes up the detailed task description.
For example if the RCM process reveals (say) that the fastening mechanism of
simply write 'redesign required'. On the other hand, it should be left to the design ers to decide exactly what sort of fastening mechanism will be used.
This issue is also discussed further in the next chapter. Finally, if a decision has been taken to allow the failure to occur, in most cases the words 'no scheduled maintenance' should be recorded in the 'proposed task' column. The only exception is hidden failure where 'the risk/reliability profile is such that failure-finding is not required ', as explained on page 1 85 . Initial Interval
Task intervals are recorded on the' decision worksheet in the 'Initial Interval' column. We have seen that they are based on the following: •
on-condition task intervals are governed by the P-F interval
•
scheduled restoration and scheduled discard task intervals depend on the useful life of the item under consideration.
•
failure-finding task intervals are governed by the consequences of the multiple failure, which dictate the availability needed, and the mean time between occurrences of the hidden failure.
When completing the decision worksheet, record each task interval on its own merits - in other words, without reference to any other tasks. This is because the reason for doing a task at a particular frequency can change over time indeed the reason for doing the task at all could disappear. So if the frequency of task X is based on the frequency of task Y and task Y is later eliminated, the frequency of task X becomes meaningless.
208
The ReM Decision Diagram
Reliability-centred Maintenance •
As explained in the next chapter, if we are confronted with a number of tasks which need to be done at a wide range of different frequencies, the time to consider consolidating them into a smaller number of work packages is when compiling maintenance schedules. However, the initial task frequencies should always remain on the decision worksheet to remind us how the schedule frequencies were derived (in other words, to preserve the 'audit trail' .). Note also that task intervals can be based on any appropriate measure of exposure to stress. This includes calendar time, running time, distance travelled, stop-start cycles, output or throughput, or any other readily measurable variable which bears a direct relationship to the failure mech anism. However, calendar time tends to be used where possible because it is the simplest and cheapest to administer.
•
•
if the task interval is short, the inspection frequency will be very high sometimes more than once per shift. This can lead to so many high frequency tasks that maintainers do little more than travel from one task to the next. This travelling time plus the cost of planning and controlling the tasks makes the use of maintainers for this purpose expensi ve, often to the point where it is simply not worth using them in this capacity. many skilled people find high-frequency tasks boring and are often reluctant to do them at all.
skilled craftspeople are very scarce in many parts of the world, so it is often difficult to spare them for this kind of work in the first place.
A second option is to use operators to do high frequency tasks. This option can be attractive because it is usually more economical and organisation ally easier to use people who are near the equipment most of the time to do high frequency tasks. Operators are also often more highly motivated to look after 'their' machines . However, three conditions must be satis fied before operators can be used with confidence to do these tasks:
Can Be Done By
The last column on the decision worksheet is used to list who should do each task. Note that the RCM process considers this issue one failure mode at a time. In other words, it does not approach the subject with any preconceived ideas about who should (or should not) do maintenance work. It s imply asks who has the competence and confidence to do this task correctly . The answer could b c anyone a t all. Tasks might b e allocated t o main tainers, operators, insurance inspectors, the quality function, specialist technicians, vendors, structural inspectors or laboratory technicians. A sometimes controversial issue which arises at this stage concerns simple high-frequency on-condition and failure-finding tasks. It some times makes sense to allocate these tasks to maintainers, but in many cases, using maintainers to do these tasks has the following drawbacks (especially if they are skilled tradespeople):
209
•
they must be properly trained in how to recognise the appropriate pot ential failure condition s in the case of on-condition tasks, and must be properly trained to do high-frequency failure-finding tasks safely
•
they must have access to simple and reliable procedures for reporting any defects which they do find. (The design of these procedures is dis cussed in more detail in Chapter I I )
•
they must be sure that action will be taken on the basis of their reports, or that they will receive constructive feedback in cases of misdiagno sis.
Using operators for this purpose can also have profound implications in terms of industrial relations and reporting relationships, so it is an issue which needs to be handled with care. Ingeneral, as with most ofthe other decisions in the RCM process, who exactly is in the best position to do each task is best decided by the people who know the equipment best. This issue is discussed at greater length in Chapter l 3.
10.3
Completing the Decision Worksheet
To illustrate how the decision worksheet is completed, we consider three failure modes which have been discussed at length in previous chapters. These are: •
the bearing which seizes on a pump with no stand-by, as discussed on pages 1 05 and 1 06
•
the bearing which seizes on an identical pump which does have a stand by, as discussed on pages 1 08 and 1 09
•
the failure of the stand-by pump set as a whole, as discussed on pages 1 1 8 and 1 79.
z ;; 1 '"
The ReM Decision Diagram
Reliability-centred Maintenance
210
.8 � 1: 01
15
The associated decisions are recorded on the decision worksheet shown in Figure 1 0. 8 . Please note three important points about this example:
�
al e:
u .s
o
- 'li .!l! 1:
�
0
'2 !
- .5
�13
�::I
Z E
Z E $ ., >-
>'"
'"
.. U.
f£
«
i
! e
Q.
,£
(.)
0> c:
c: Cd c
.�
.$ c 'ca E
.D 0.
E
::J 0. c:
""0
:;
"ca E -" (.)
al ..c:
�
•
a number of other preventive tasks could have been chosen to anticipate the failure of the bearing - the decisions in the example are for the pur pose of illustration only.
•
the stand-by pump is treated as a 'black box ' . In practice, if such a pump were known to suffer from one or more dominant failure modes, these failures would be analysed individually.
10.4 Computers and RCM .
o
z
>z
z
z
z
z
z
z
!;j g
-.8 z (j)
z - .8 z �
� N
� N
>-
Z Z "' o ::I: 8
:?1 u � � u ", o � O � @ ��1�
the first two pumps could suffer from many more failure modes than the failure under consideration. Each of these other failures would also be listed and analysed on its own merits.
was selected. This information is invaluable if the need to do any main tenance task is challenged at any time. The ability to trace each task right back to the functions and desired performance of the asset also make it a simple matter to keep the main tenance program up to date. This is because users can readily identify and reassess tasks which are affected by a change in the operating context of the asset (such as a change in shift arrangements or a change in safety regulations), and avoid wasting time reassessing tasks which are unlikely to be affected by the change.
:tl ....
= V) � �
•
In essence, the RCM worksheets not only show what course of action has been selected to deal with each failure mode, but they also show why it
�
.b"
211
«
�
__ __
>«
Q.
�
Q. -- :::.. til �
�
�
(J)
z < N
�
__ __ __ __ __ __ __ __ __ __ __ __ __ __
entries Figure 10.8: An ReM deCision workSheet with sample
The information contained in the RCM and Decision Worksheets lends itself readily to being stored in a computerised database. In fact, if a large number of assets are to be analysed, it is almost essential to used a com puter for this purpose. A computer can also be used to sort the proposed tasks by interval and skillset, and to generate a variety of other reports (failure modes by consequence category, tasks by task category, and so on). Finally, storing the analyses in a database makes it infinitely easier to revise and refine the analyses as more is learned and as the operating context changes (as it;;urely will see part 7 of Chapter 1 3) . However, note that a computer should only ever be used to store and sort RCM infornlation, and perhaps to assist with the more complex failure finding interval calculations. For reasons discussed in Chapter 1 4, com puters should never be used to drive the RCM process.
Implementing ReM Recommendations
11
Implementi ng ReM Recom mendations
Can be Proposed Initial noto:= sk eo:: rVo= al:_f -
__
No sCheduled maintenance Check coupling bolts
Monthly
Mechanic
1
No sCheduled maintenance Redesign guard Check agitator gearbox oil level
11.1 Implementation - The Key Steps
213
Weekly
--+-
Operator
Check tension of main drive chain
Monthly
Calibrate gauge
Annually
E&I technician
4 yearly
Operator
Mechanic
AIIDIT THE DECISJON WORKSHEET
No SCheduled maintenance
We have seen how the formal application of the RCM process ends with completed deeision worksheets. These specify a number of routine tasks which need to"be done at regular intervals to ensure that the asset contin ues to do whatever its users want it to do, together with the default actions which must be taken if an appropriate routine task cannot be found. The people who participate in this process learn a great deal about how the asset works and about how it fails. This on its own frequently causes the participants to change their behaviour in ways which often lead direct1y to remarkable improvements in asset performance. However, in order to derive the maximum long-term benefit from RCM, steps must be taken to implement the recommendations on a formal basis. These steps should ensure that: •
all the recommendations are approved formally by the managers with overall responsibility for the assets
•
all routine tasks are described clearly and concisely
•
all actions which call for once-off changes (to designs, to the way the assct is operated or to the capability of operators and maintainers) are identified and implemented correctly
•
routine tasks and operating procedure changes are incorporated into appropriate work packages
•
the work packages and once-off changes are implemented. Specifically, this in tum entails: incorporating the work packages into systems which ensure that they will to be performed by the right people at the right time and that they will be done correctly
Drain main tank and check if low level alarm sounds at 50 litres
2.
3 .
UPGAAbi: AOUTINE TASK DESCR I PTIONS (Wfite leocl task instruc o on8)
IDENTII=Y ONCE�OFF CHANG ES (to capabUity or. tp o perating pfocedures)
4
INCOAPORATE oo ooo ROOT' N E r K D ESC �I 0 .0 o .S' INTOWOFtK · ..
PACf5.AGf!S
r Operatin91 LP!ocedur:�
·
�;
·· ·· 5lMPl.E"'E l' SVStElVIs WHI.CHJ;NSUR� rHAT THE WORK IS DONE
- ensuring that any faults found are dealt with speedily. These steps are summarised in Figure 1 1 . 1 opposite. The most important of them are discussed in more detail in the rest of this chapter.
Figure 11.1: After ReM
214
Implementing ReM Recommendatiolls
Reliability-centred Maintenance
11.2 The ReM Audit If it is correctly applied, the RCM process provides the most robust frame s. work currently available for formulating asset management strategie integrity These strategies profoundly affect the safety, environmental r, and economic well-being of the organisation using the assets. Howeve people the of efforts best the of spite in wrong badly go if something does and applying the process, every decision will be subjected to a thorough regula from ranging tions organisa by often intensely adversarial review of tory authorities through insurers and shareholders to representatives RCM uses which tion organisa victims (or their survivors). As a result, any what should take great care to ensure that the people who apply it know sensible are s they are doing, and also to satisfy itself that their decision and defensible. The latter step is known as the RCM audit. RCM audits entail a formal review of the contents of the RCM In at formation and Decision Worksheets. This section of this chapter looks entails. it what and done be should it who should do the audit, when Who should do the audit
Senior management bears the overall respons ibility for the asset if some their thing goes badly wrong, so it is in the interest s of themsel ves and to taken being are steps ble reasona employ ers to satisfy themsel ves that rs manage senior , 1 Chapter in prevent such occurrences. As mentioned delegate may but ves, do not necessa rily have to do the audits themsel How them to anyone in whose judgement they have enough confidence. are ever, if this is done, it should always be understo od that the auditors acting on behalf of senior management, so the latter still bear the ultimate should respons ibility for the decision s. (Whoever carries out the audits RCM.) in also be thoroughly trained If the auditors disagree with any findings or conclusi ons, they should doing, discuss the matter with the people who performed the analysis . In so be may ves themsel they that the auditors should be prepared to accept .) queried are s decision wrong. (In most cases, no more than 5% of the When the audit should be done
has Audits should be carried out as soon as possible after each review : reasons three for been completed (preferably within two weeks),
•
215
the people who did the analysis are keen to see the results of their efforts put into practice. (If this happens too slowly, they start to lose interest, and more seriously, they begin to question whether manage ment was serious about involving them in the first place. )
•
people can still recall easily why they made specific decisions
•
the sooner the decisions are implemented, the sooner the organisation derives the full benefits of the exercise.
When overall agreement is reached about each analysis, the decisions are implemented as described in the rest of this chapter. What the audit entails
An RCM analysis needs to be audited from the point of view of method and content. When reviewing the method, the auditor seeks to ensure that the RCM process has been correetly applied. When reviewing the content, the auditor seeks to ensure that the correct information has been b crathered and conclusions drawn both about the asset itself and the process of which it forms part. Issues which most often need attention are as follows:
Levels of analysis The analysis should be carried out at the right level. The most common fault is to analyse assets at too Iow a level, and the usual symptom is large numbers of items with only one or two functions defined per item.
Functions All the functions of the asset should be clearly and con·ectly described. Key points to look for include the following: •
by and large, each function statement should define only one function, although it may incorporate more than one performance standard. As a rule, each function statement should contain only one verb (unless it is a protective device)
•
performance standards should be quantified, and should indicate what the asset must be able to do in its present operating context rather than its rated capacity (what it can do)
•
all protective deviees should be listed and their functions correctly de scribed ( 'to do X ifY occurs ' )
•
the functions o f all gauges and indicators should be listed, together with desired levels of accuracy,
216
Reliability-centred Maintenance
Implementing RCM Recommendations
2 17
Functional failures
Consequence evaluation
All the functional failures associated with each function should be listed (usually complete failure plus the negative of each performance standard in the function statement).
Special care should be taken to ensure that the hidden function question (question H on page 200) has been answere d correctly . In particula r, the correct meanings should have been attached to the terms on its own and
Failure modes Ensure that failure modes which have happened or which are reasonably likely have not been omitted. Failure mode descriptions should also be specific. In particular, •
they should include a verb, not just specify a component
•
the verb should be a word other than fails or malfunctions unless it is appropriate to treat the failure of a sub-assembly as single failure mode (option 3 on page 87)
•
switch and valve failures should indicate whether the item fails in the open or closed position
Failure modes should relate directly to the functional failure under consi deration, and failure modes and effects should not be transposed, as in: Failure Mode
Failure Effect
Motor trips out
Pump impeller jammed by rock
Another common mistake is to combine two substantially different fail ure modes in one description, as follows: Wrong
Right
1 Screens damaged or worn
1
Screens damaged
2 Screens worn.
Failure effects Failure effect descriptions should make it possible to decide: •
whether (and how) the failure will be evident to the operating crew
•
whether (and how) the failure poses a threat to safety
•
what effect (if any) the failure has on production or operations (output, product quality, customer service).
Failure effects should not incorporate actual 'consequence words' like 'This failure affects safety' or ' This failure is evident' . However, they should list likely total downtime as opposed to repair time, and should indicate what must be done to rectify the failure (replace, repair, reset, etc) Finally, auditors should satisfy themselves that anything which is said to be 'analysed separately' actually is analysed separately.
under normal circumstances in this question , as explaine d on pages 1 24 and 1 26. Special attention should also be paid to the evaluation of the safety and environmental consequ ences of evident failures, and to the effective ness of any tasks which might have been selected to manage failures in these two categorie s. Task selection Any tasks which have been selected should not only satisfy the criteria for technical feasibility as explained in chapters 6, 7 and 8, but they should also address the conseque nces of the failure. Key points to look out for: •
if the answer to question H is 'No' and the answer to question H4 is 'No ' , then question H5 must b e answered. I f the answer t o H5 is ' Yes ' , the proposed task should not be 'no scheduled maintenan ce'
•
if the answer to question S or E is ' Yes' , the proposed task should not be 'no scheduled maintenance'
•
if the failure has operational or non-operational conseque nces, the task must be cost-effective.
Proposed tasks or default actions should be described in enough detail to leave the auditor in no doubt as to what is intended. In particular, routine task descriptions should not simply list the type of task ( , scheduled on condition tasks' or ' scheduled failure-finding', etc). The task description should also relate directly and solely to the failure mode in question. It should not incorporate a combination of tasks be cause this usually signifies two different failure modes (unless the answer to question S4 is yes). For example: Wrong
Right
Inspect chain for wear and adjust tension
Adjust tension of chain
or Inspect chain for wear
Initial interval Task intervals should clearly have been set according to the criteria set out in Chapters 6, 7 and 8. In particular, look out for a tendency to confuse P-F intervals with useful life in on-conditio n task intervals .
218
Reliability-centred Maintenance
Implementing ReM Recommendations
11.3 Task Descriptions Before any task reaches the person who has to do it, it must be described in enough detail to leave no doubt at all as to what is to be done. Clearly, the degree of detail required will be influenced by the overall level of skill and experience of the workers involved. However, bear in mind that the more that is left out of a task description, the greater the chance that some one will miss out a key step or choose to do the wrong task altogether. In this context, special care needs to be taken with the description of any failure-finding task which calls for a hazardous situation to be simulated in order to test 'the function of a protective device. Task descriptions should also explain what action must be taken if a defect is encountered. (For instance, should the defect be reported to a supervisor or to the maintenance department - or should it be rectified immediately?) Instructions like 'check component A for condition B and replace if necessary' should be used with caution, because the 'check' part of the task might only take a few seconds, while the 'replace' part could take several hours. This can play havoc with the duration of planned downtime. Instructions of this sort should in fact be written as 'check component A for condition B and report defects to supervisor' . Only use 'if necessary' for quick servicing routines, such as 'check gearbox oil level using dipstick and top up with Wonderoil Type 900 if necessary' . Examples of the right and the wrong way to specify tasks are shown in Figure 1 1 . 2 below: Wrong
Check coupling
Check feedscrew coupling for loose bolts and replace if necessary Visually check agitator coupling flange for cracks and report defects to the maintenance supervisor
..
. etc
Fit 0 - 20 bar test gauge to test point and check if reading o n pressure gauge P I 1 204 i s within 0.5 bar of the reading o n the test gauge when the test gauge
reads 8 bar. Arrange to replace out-of-spec gauges when plant i s shut down for cleaning
or Remove pressure gauge P I 1 204 to workshop and calibrate following procedure in manual 27A
Figure lL2: Task descriptions
Basic information In addition to a clear descripti on of the task itself, the documen t on which the task is listed should also clearly state the following : •
a description ofthe asset to which it applies together with an equipment
•
who should do the task (operator, electrician, fitter, technician, etc) the frequency with which the task is to be done whether (and ifnecessary, how) the equipment should be stopped and/
number where relevant • •
or isolated while the task is done, together with any other safety pre cautions which must be taken
special tools and prescribed spares. Listing these items can save much unproductive walking to and fro after the job has started.
ISO 9000 and R(.'M
or
Calibrate gauge
Pages 206 and 207 explaine d that each task should be defined as clearly as possible on the decision worksheet. This saves the duplication of effort which occurs if detailed procedures have to be written up later by some one else. It also reduces the possibili ty of transcrip tion errors. Howeve r, if time does not permit the procedures to be specifie d during the RCM analysis , then they must be specifie d later. As mention cd below, this can often be done as part of an ISO 9000-typ e initiative . Note that if detailed task descript ions are to be prepared later, this should ideally be done by someone who participated in the oIiginal RCM analysis . If this is not possible , the third party should understand clearly that he or she is being asked to define the tasks on the decision worksheet in more detail, and not to re-audit the analysis.
•
Right
219
A maj or objective of RCM is to identify what work people should be doing. (In other words, to ensure that 'they do the right job' .) On the other hand, a major thrust of quality systems like ISO 9000 i s to define what people should be doing as clearly as possible in order to minimise the chance of errors. (In other words, to ensure that 'they do the job right ' . ) This suggests that the process of transferring tasks from RCM decision worksheets to end-user documen ts can be seen as the point where the output of an RCM analysis becomes the input to an ISO 9000 procedure writing exercise. It also suggests that if both initiati yes are to be under taken, it makes sense to apply ReM first.
220
Reliability-centred Maintenance
1 1.4 hnplementing Once-off Changes At the end of a typical RCM analysis, it is not unusual to find that between 2% and 1 0% of the failure modes default to redesign. Part 2 of Chapter 9 mentioned that in the context ofRCM, redesign means a once-off change in any of the following three areas: • a change to the physical configuration of an asset or system • a change to a process or operating procedure • a change to the capability of a person, usually by training. Once they have been accepted by the auditors, these changes need to be implemented as �horoughly and as quickly as possible. Key issues in each of these three areas are discussed below
Implementing RCM Recommendations
22 1
Changes to the capability ofpeople As explain ed in Chapter 4, the RCM process frequen tly reveals failure modes caused by slips or lapses on the part of operators or maintainers (skill-based human errors) . These immediately become appare nt to any operators or mainta iners who participate directly in the process , and they usually modify their behaviour appropriately as soon as they learn what they are doing wrong. Howev er, we also need to ensure that people who have not particip ated directly in the process acquire the relevant skills. In most cases, the most efficien t way to do this is to revise or extend existing training programs, or to develop new program s. In most organisations, this will be done in consultation with the training department.
Changes to the physical asset All modifications should be: •
•
•
Work Packages
j ustified in terms of their consequences. Modifications intended to deal with single or multiple failures which have safety or environmental consequences should reduce the risk (frequency and/or severity) of the consequences to a level which is acceptable. Figure 9.2 showed an algo rithm which can be used to justify modifications intended to deal with failures that only have economic consequences
Once the maintenance procedures have been fully specified, they need to be packaged in a form which can be planned and organised without too much difficult y, and which can be presented in a neat and compac t form to the people who will be doing the tasks. This can be done in two ways:
correctly designed by suitably qualified engineers. As a rule, attempts
•
should not be made to redesign assets during the RCM process, but the designer should consult afterwards with the people who did the review in order to develop a correctly focused specification
Standard Operating Procedures
properly implemented. Steps must be taken to ensure that modifications are carried out as intended in terms of time, cost and quality, and that all drawings, manuals and parts lists are updated correctly
•
1 1 .5
properly justified. Chapter 9 explained that modifications should be
properly managed. Modifications should not interfere with essential routine maintenance activities in other parts of the plant, and the main tenance requirements of every modified item of equipment should be correctly assessed and implemented.
Changes to the way in which the plant is operated Once-off changes to the way in which plant must be operated are handled in the s ame way as routine tasks which are incorporated into operating procedures, as explained in the next part of this chapter.
•
high-frequency maintenance procedures to be done by operators can he incorporated into the operating procedures of the equipment
the balance of the maintenance routines are packaged into separate schedules and checklists .
The previou s part of this chapter mentioned that any changes which must be made to the way in which an asset is operated should he docume nted in standard operatin g procedures, or SOP's. (In situation s where they do not exist already , it will almost certainl y be necessary to develop them in order to ensure that the changes are implemented.) In many cases, SOP's are also the simplest and cheapes t way to manage high frequenc y tasks which need to be done by operators, as illustrated in Figure 1 1 . 3 overleaf . As a rule, tasks should only be incorporated into operating procedures if they need to be done at intervals of one week or less. Tasks which need to be done by operators at longer intervals should be packaged into sepa rate schedules and planned, organised and controlled in the same way as maintenance schedules, as described in Part 6 of this chapter.
222
Reliability-centred Maintenance lriiliar
Interval
C
No scheduled maintenance
No scheduled maintenance Redesign guard
Operator
Check tension of main drive chain
Monthly
Mechanic
Calibrate gauge
Annually
No scheduled maintenance Drain main tank and check it low level alarm sounds at 50 Htres
WIDGET-�.--.. WASHING MACHINE -"--.- .--Fill feed hopper
•
Open air valve and wait until pressure reaches 50 psi
•
Weekly
leliel
.Stllndard �rating proceq\lfe At the start of the shift
Mechanic
Monthly
ChecK coupling bolts
Implementing ReM Recommendations
4 yearly
E&I technician
Operator
•
(Monct,iy m0f!1iIJ9$ only} Check agitatorgtl'�rbox oil leVel using dlpstklk and report if it is bE)foW. level 2 Press start button
• •
Open detergent valve
•
Start widget feed
etc
•
Figure 11.3:
Transferring a task from a decision worksheet to an SOP Maintenance Schedules A
maintenance schedule is a document listing a number of maintenance tasks to be done by a person with a specified level of skill on a specified asset at a specified frequency. Figure 1 1 .4 shows the relationship be tween these schedules and the RCM decision worksheets: Can done
. Maiote0!!l�e S(;hed.u�_
'-.-�.-.� '�'-'-7"�;� c·. I;!Olle Ill: .
WIDGET WASHING MACHINE
.C. �.--.-'�.' ..
.... .Iotervaf
MONTHLY
.. . .... .
�.�- ----
MECHANIC
---�-
Stop machine and follow lock-out procedure X, then 1
Calibrate gauge
Annually
E&I technician
4 yearly
Operator
No scheduled maintenance Drain main tank and check if low level alarm sounds at 50 litres
Figure 11.4:
Transferring tasks from a decision worksheet to a maintenance schedule
Compiling schedules from RCM decision worksheets is a fairly straight forward process. However, a few additional factors need to be taken into account as explained in the following paragraphs.
223
Consolidating frequencies In Chapter 7 it was mentioned that if a wide range of different task inter vals appear on a decision worksheet, they should be consolidated into a smaller number of work packages when compiling the schcdules based on the worksheets. Figure 1 1 .5 gives an extreme example of the variety of task intervals which could appear on a decision worksheet, and how they might be consolidated into a smaller number of schedule frequencies . The most expensive tasks, in terms Intervals Intervals o f of the direct cost of doing them and o f tas ks o n maintenance the amount of downtime needed to do decisio n s chedules wor ks heets them, tend to dictate basic schedule Daily intervals. However, planning is sim Daily Weekly Weekly plified if schedule intervals are multi 2- weekly ples of one another, as shown in the Monthly Monthly example. 6-weekly Note also that if a task frequency is 2-monthly changed in this fashion, it should al 3-monthly 3- monthly ways be incorporated into a schedule 4-monthly of a higher frequency. Task intervals 6-monthly 6- monthly should never be arbitrarily increased, 9-monthly because doing so could move an on 1 2-monthly 1 2-monthly condition task frequency outside the P-F interval for that failure, or it could Figure 11.5: m(�ve a scheduled discard task past Consolidating task frequencies the end of the 'life' of the component
Contradictions When a low frequency schedule incorporates a higher frequency sched ule, should the latter be incorporated as a global instruction, or should it be rewritten in full? In other words, should (say) an annual schedule in clude an instruction like 'do the three monthly schedule' , or should all the tasks in the three monthly schedule be written out in the annual schedule? In fact it is wise to rewrite the schedules in order to avoid the problem of contradictions. For i n stance, consider what could happen i n a situation where a three monthly schedule includes the i nstruction 'check gearbox oil and top u p if necessary', and the annual schedule for the same machine starts with the instruction 'do the three monthly schedule', and later says 'drain, flush and refil l gearbox'.
Mainlenance and contradictions of this nature system ill the cyes of the little extra time to ensure that
schedulcs on the basis described above, there is often a schedule. This is
start adding tasks to the on the basis that 'when we do A and . This should be avoided
we might as well
for thc following reason,,:
tasks increase tbe routine workload.
too many tasks
workload is increased to the point where there ic. either insufficient labour to do all the amount
or the equipment cannot be released for the to do
or both
the schedules soon rClllise that
and
are not
for reasons ule
a whole. When they find
tasks
and the whole maintenance program This
common in shutdowns.
�hutdown tasks are
needed, but because the it is
to
at' the equipment. This adds
is
ami
to the cost and work also leads
sometimes to the duration of the shutdown.
to an incrense in infant mortality when the plant starts up does nol mean that people who do routine tasks should concentrate only 011 the
tasks and ignore any other potential and func-
tiollal failures which they may encounter. Of course they should their
open. The point is that the schedule itself should only needs to be done at that frequency,)
the next
into sensible work which cnsure that
influences
Formal the
valuable source of
taken to ensure thal arc noticed. This detail later.
failures
the schedule plans and so there is no need for any not used for allY tasks whieh
done at intervals
than a checkli s t involves one or two documents per week amounts to no more than
1000 items
documents (0 these checks.
readings to be or
(vibration
while the checklist described above is
start Possible alternati ves are as follows: document for (Ill the
in each
and
attach this one document to the checklist for that section each week person to take all such the
in the elltire those
which
limits in the remarks column of the cheek list in the
of certain
discussed ill
MailltellwlCI:'
schedules can be
can be
to
we consider the but cOllsiderllle
I ill/£'
The principles of elapsed time planning are weI! many purpo,>es in addition to maintenance time planning is usually based on a similar to that shown ill
! J .7 (or its
7: A
board
Most of these systems use an overall planning horizon or one
. divi-
ded into 52 weeks. However, bear in mind when setting up such that some failure-finding tasks in particular can have
times of up to
ten years, and the planning horizon of any associated planning must accommodate such tasks. When setting up these schedules
always involve
and these ean have operational consequences in which
the
care lIlllSt
are supposed to prcvent, so
taken to minimise these consequences. Points to watch for include: and
in the
schedules should be
two machines which a
same
at
record is fed into
/VIaintC'IWJl('C' automated it can
to agree on which which
schedules subject to the same any (Jther
of maintenance work. This
do the schedules. standards of workmanship. and so on. Two additional factors necd to be considered.
the planning
should indieate when any schedule is overdue. As mentioned automatically, but
earlier, such sehedulcs should not be should be
on an
basis.
maintenance schedules should be reviewed continuollsly i n which and new information. ill this context, associated with the to
is
1
lV/aintellance
ivl (I iIII C!la/l( 'C
IS whieh has survived to This is known is a 14% chance that an
12 will fail in that period. Slrn''''rlll of period 1 failure of 87°;;), to Ihe
pattern
!'vIa iIII CfUlilce
the have two
'life' can
mean time between failures (which j::, the same as the average has run to
whole a
The second is the
at which
increase in the conditional
allY other term, this has
named the
to overhaul or
if
time
half would fail before
reached it. In olher words,
halfof the failures, which is
J 2.2 shows that the useful life is shorter than the mean time between failures
it can be very nHlch shorter.
if the bell curve is
As a result, it can only be concluded that the mea/1 time IIres is
the
little or 110 lise in
of failure. if
the
do
{'n:rn'-If)!:,,,n
the average service I if
let it run to failure. As
the cost of rnaintenanee ( associated with the failures). For instance, if we were to
12.2
all the surviving impellers in
tile end of period I 0, the average service life of tile impellers would be about instead of 12.3 periods if tiley were allowed to run to failure.
9.5 �
The fact that there are two 'lives' associated with pattern B -type failures means that we must take care to ever we use the term we
For
which one we Illean when-
'life'.
pi lOne the manufacturer of a certain component to ask
what its 'life' is. We may have in mind the useful life, but if we don't spell out what
mean, he
in all
failures. If this is then used to establish problems arise, often
in
faith give us the mean time between all kinds of unnecessary unpleasantness.
problem associated with behave in this fashion.
MailllcfWIlCC
'elilred lvtuillienance if
of IS
on
random basis
many failure modes which
IS trical ami electronic items, This does not detract in any way from the means that no form of scheduled restoration or these
and
default
WeibuJ/ distributions
A At this distribution is enable it to fit many kinds of data. The Weibull
MaintellwlCC
c
/vJ Uill/CIW}/('e
to almost flat.
failures
conform to fail-
12.6 were skewed towards the the
limit, the failure This would
towards
limit and curve would
a Weibull shape parameter
normal distribution and so gives a condi-
curve which resembles pattern S.
tional On the other hand, if the
"'<::',!fTlf1U,tl(' then
limit is below the point at which
becomes
a long 'tail' on the right.
distribution will
Weibull distribution where f3 is between 1 and 2, which
in of any distribution very wide, which would in turn lead to an curve. Add to this the failore
of additional
listed on page 159 which have the same symptoms as
and the overall probability
effectively becomes fully exponen
tial, which leads to failure pattern E as we have seen. So
could manifest itself
failure pattern S, C, 0 or even E.
As mentioned above, failure pattern D is the conditional probability curve associated with a Weibull distribution whose shape parameter f3 is than I and less than
Pattern declines
13.
Mainteno/lce
and work
up ewryone involved in to lhal it does so.
routine maintenance tasks latter
it will fall
quickly. This
lhat anyone who is ealled upon w
or corrective maintenance task is trained alld motivated the first time.
and
The ahove discussion solved the
that infant
problems are usually
once�otr actions rather than by scheduled maintenance (with where it may be fcasible to use on-condition minimal role before
Mainiemllle!'
sLlmrnarised helow under the and evolution failure the ultimate contradiction.
Most industrial
consist of
if not thousands of
different assets. These are made up of dozens of different components, which between them exhibit every extreme and intermediate reliabilit y behaviour. This combination of that it is simply not
of
aml diversity means
t o develop a complete analytical description
of the reliability characteristics of an entire major asset within the Even at the level of individual functional
or even any
Maintenallce
it. the result is that manufacturer talks about what u/lil/WIC ('olllradicliOIl
which be(kvils the whole
statement: thuught to be most ahout critica/failures failllre These may even reveal some failures don't matter lIlueh, it is uled restoration or scheduled discard the actuarial to be
it
waste of time.
vehicle
into this category are street
many 01 the components used
the
distribution industries. Items numbers for precise actuarial analyses to be (in many cases, if only the
or failures whieh merit actuarial still
to he
and
task and the cost of the failure are
for a limited number of or the relationship between
actuarial
and failure is of very little use from the maintenance man the most serious shortcoming of historical it is rooted in the
whereas
or
focused on
with the definition of the functions and the standards of each asset, ReM enables us to what we mean by 'failed'. By distinguish between built-in capability and desired performance, and state) and functional failures (the
failures
. ReM breaks each asset clown into its functions and each
fUllction into functional failures, ancl only then identifies the failure modes which calise each fUllctional failure. This provides an framework within which to con"idcr each failure mode. This in turn makes them much failure mode levld
tu manage than if we
to start out at the
of most classical
the
and •
record of all the
el'ollitiotl:
functional failures and failure modes associated wilh out
how any
to make
105 106
108 -109
105 106 108 109
98 - 101
144 -145 Mean tlf118 be tween failures of
interval
MailltenallCI:'
Table 12.1 only describes data which are used CIes
to deal with
to formulate poli-
failure modes.
not include data
used to track the overall performance of the maintenance function and c1assified as
of such infor-
of the maintenance fUllelion is of course an essential
This
cussed in more detail in note
Oil
is dis-
14.
the
In recent times. the concept oflhe 'mean time between failures' seems to have
to its real value in
maintenance
it has
to do with the
of on·conditioll tasks. and nothing to do with
of
scheduled restoration and scheduled discard tasks.
it does have
certain very specifie uses. Table 12.1 men! ions three of these: to establish
In additioll to
of failure-finding tasks.
to help decide whether scheduled maintcnancc is worth doinR in the
about maintenance
case of failure modes which have operational or nOll-operational con
the ()('currellce
sequences only. (In other words it helps us to decide whetlwrsuch tasks
sequen(,es. This information can
need to be done. but not how often they need to be done.)
tween the failures in order
to help establish the desired availahility of a
device.
I n the first case. the MTBF is always needed to make the
cision, but in the second two it is
de-
used if the nature and consequences
of the failures are such that a
must be carried out
The MTBF also has a number of uses outside the field of maintenance formulation. as follows: in lhefield modificatiOll.
. to carry out a detailed mentioned
of
260
Maintenance
of a boiler
as
This can be done
task (or who discovers the failure in the
is used [0 store it
enter the data "w,,,,,,,,
the records themselves
m:
or or manual maintellance
'The used to
data for
reaSOllS, rather than to record every-
in the hope that it will eventually tell us something, they become useful and powerful contributors to the practice of maintenance management rather thall the
white elephants that so many of them tend to be,
jll/aintenance
deliver.
maintenance
and a valuable source of information about Ihe vendor and who
technician vvho is has worked on the same or similar equipment.
recorded as failure effects. These include: the evidence that the failure has occurred, which is most often obtained from the
of the equipment
the amount of time the machine is failure occurs,
out of action each time
obtained from operators
obtained
erators alld their
!VJainfenallce
increase in costs other than the direct costs of who arc in the best COSIUCColilltan/s.
managers cOlI s c quenccs : the
direct
costs
llIainte-
of different
The information needed to assess the technical tasks was discussed in Chapters 6 to
types uf
9, and
205. If clear actuarial data are not
summariscd on
be answered
answers, then the questions must
available tu
as follows:
and ofl-condilion tusks:
consider
many different
on-condition tasks. The
ill thc
similar group w()uld need to consider the duration and consistency of the associated
illter vals, as explained 011 pages
164 and 165.
The amount of time needed to avoid the consequences of the failure other words. the nett
interval) is
maintl:'
should not consult
nance and
jllst that it should not
scheduled restorll/ioll and scheduled discard: in the absence of suit-
able historical
who are usually most
the
whether any failure mode is
to know
and if so \vhether am! whell
at which there is a rapid increase in the e()llditional prothe
altd
to failure of III
review
Maintenance
In the light of the issues raised in Part 1 of this who should
we now
review group. what each out of this process.
Who should
mentioned most line supervisors,
ill Part I of this and craftsmen. This suggests that a
review group should include
What e(fch grouJI does
the people shown in
Facilitator
13.1.
of
In the
this
un every
lS
consenslis.
I:ach group mcmber
group do not have to filled hy exactly the same people as showIl in Figure
13. 1. The objective
is to as:,emble a gruup which can provide most if not all of tile information described in Part I of this chapter. These are the people who have the most extensive knowledge and
of the asset and of the process of
which it forms part. To ensure that all the different viewpoints are taken into account. this group should include a cross-section of users and main
1�
u)
contrihute what· he or she can at each
in the
process, as shown
1:1.2.
taillcrs, and a cross-sect ion of the people who do the tasks and the people who manage them. In
it should consist of not less than
not more than seven
the ideal
The group should consist of the same individuals of any one asset. If the faces much lime is lost
over
at each
1 of this
the
whole group.
too
discussed in Part the
which has
the benefit of the newcomers. As
four and
five or
call
3 of
A1(/intellllflce
This work is done at a each, ,md each
which las! for about
of
meets
hours
rate
an
week. If the group includes shift with
need to
care.
The asset should be subdivided and allocated to that any one group can
the entire process in no!
not more thall fifteen
no morc than
gel out
What
procc.Is
at these
'rlle flow of information which lakes
is no!
into the database. W hen allY one member of the group makes a contribu-
the others more about
learn three asset, more about the process of which it forms
more about what lUust be done to
it
As a
who eaeh know a bit
of
five or six
little bit
about the asset under review, the
and instead
often a surprisingly five
on the
SIX
more ahout the
and
maintenance
of their
learn more about what their
to achieve, while
learn much more about
how maintenance
them to
more about the indi vidual ber On
it.
and weaknesses of each team mem-
much more tends to be learned about which has a salutary effect on mutual
than as well as
mutual understanding. In short, participants in this process gain a much better understanding of what each group member (themselves included) should be doing •
what the group is trying to achieve by doing it and how well each group member is equipped to make the attempt. the group from a collection of highly disparate individuals
This
adversarial into a team. The fact thai
have each played a pari in
solutions also leads 10 a much of the
For while maintenance
offer constructive criticism of 'their' schedules.
and
mentioned that the facilitator of ReM. The facilitator is to facilitate the of a group of people chosen that the group reaches
270
MainlCfUlIlCC
administration,
and
points about each ski!lscl are discllssed in the
ReM process is
The facilitator must ensure that the the review group. This entails the
ReM process are asked
have been
correctly
that all
embodied in
in the correct sequence, that
understood
all the group members and that the
group reaches consensus about the answers.
the facilitator and/or
the following decisions are made
and
lhe facilitator alone does the work. facilitator should This includes now
collect basic information about the
and circuit drawings . •
Select levels o/analvsi.\/de./ine boundaries: The equipment to be analysed by each review group will be identified durillg the planning phase. How ever, it may become necessary to group the equipment differently in
review groups. Before this can
order to carry out a sensible analysis. This means that the final decision
ReM Worksheets in
about equipment grouping/levels of analysis is Inade by the facilitator,
them into a formal document called the audit
who then has to define the boundaries of the allaly�is accordingly. Handle cmnplexfailure modes appropriately: Deeide when to choose which of the four options listed in Chapter 4
7
86 88) when -
failure modes modes:
Know when 10 stop the failure modes that
cause each functional failure is one of the
elements of successful facilitating, and
careful judgment.
on too soon to the next functional failure means that critical failure modes may described.
overlooked or that failure etlects are too lllallY failure modes icads to
do
coherent fashion.
Illc.
Maintenance
deal with the way in which the facilitator on a
human level.
the scene: At the first
of each group, the facilitator must
norms with the group
such as use of names,
etc.) and ensure that every group member understands the scope and
of the exercise and
been asked to take
he or she
At the start of all
should
the faci-
reeap what has been doue to elate and The faeilitator should also ensure that or
. How the facilitator conducts him
the herself in
to
has a profound effect on
way the other group
members behave. In particular, the facilitator should set a attitude to the process, take care to preserve of group members and
feedback ill
contributions.
sponse to Ask the ReM
role or the facilitator is 10 ask the process. It is essential to avoid any tendency to
questions or to take
answers for granted. (In particular, take care not to ignore or overlook to establish whether any task is worth doing.)
questions
has heen correct(v understood: In
Ensure thaI each
of the
fact that they should all have attended a basic RCM training course, group members are not as familiar with the RCM process as the facili-· they often misunderstand the questions, especially in and the facilitalor must be alert to sllch misunderstandCommon mistakes were discussed in eVl'rvone 10
11
Coach the group or individual the facilitator to is
formal and
substitute for formal
IV!Lli III elUll we
(ludit: As
arrange have been COlll
Since recommendations cannot be implemented until they have also beell audited, this An
should also be caITicd out as quickly facilitator shoulu be able to have an
more than two weeks after the last
of the
not
in
been trained
decide when, - decide when the
276
Mainlencmcc
CummuniCitle
informed about progress which you cannot solve such
facilitator should usu-
EnslIre Ihat ReI!;! worksheets ilre alldited:
note corrections
in person, to answer and
to the auditors on the
10
(although the auditor� must to audit an
forma!
in
process before
ReM analysis). The facilitator must also ensure
that consensus is achieved between tlle auditors and the review group the alldit process. This entails the group, and
audit findings back to
that differences are resolved. Finally, the facili
tator must update the worksheets to incorporate the results of lhe audit. ",'yHlfnnn'
A short, high .. quality summary of at least
should
'M
to the senior managers of should show how the or will be
and what
of
could threaten
sinal! illl'l:stment.
none of of lend
to
also have somc asscts which arc more which are di!Ticult to pour poor customer
or excessive maintcllullce costs. Othcr
confronted with ullaeeeptable
areas
hazards which need to
tackled on a
Given hundreds if Hot thousands of items to choose from in a it
with the power
it does not contain a sation to the risk of
IvlII ill {{'n(lm'e
280
called 'critical ity assessments', known as the a numerical value [0 or failure rate
the
olan asset (the
and another value t o
ofthe consequences o r the failure
the more serious thc failure, the higher the value). The two numbers a third, which is the
P RN. Assets with the highcst
then those with lowcr scores and so on until assets are encountt'red where the likely return does notjustify detailed analysis. variations of this process build lip composite PRN' s different numerical weightings to different categories of railure consequences (typically, high for safety or environmental conse quences. intermediate
for operational, and lowcr for direct repair costs).
historical failure rates further relined
Pareto
of this sort call be useCul in
available.
grams as eli rrerent groups select different same failure. One way in which this
7.8 on page
I
lVluillft'IlUllce
resource intensive so it attention. be considered if a number of other initiatives are to with
ReM. is still to review all the
to do so in
should not undertaken
on tbe site, but
Perhaps four or five groups arc activated at a
under the direction of one or two facilitators. Oil this basis, it could take fi ve to ten years to to four OIl a smaller olle). 'I'he
all the
site
process, commitments are needed to achieve it. However, before
Mailltenall('e
of RC
of which have been
scientific view of what must be done to cause them to continue to fulfil will not be
their intended functions. However, the
and
for two reasons:
will
orrailure modes which haven't In an environment like this, it is inevitable tbat some failure while some failure
modes and effects will be overlooked consequences and task
will be assessed
and the processes of which they form part will be continuously. This means that even
of the analysis whieh are com··
valid today may become invalid tomorrow. involved in the process will also change. This is and
which
has anyone become aware of a to one of those selected
with time, and partly because people others
are. All these factors mean that
which most
"' .. "",."
because
of those who take part in the original analysis leave and their
there any reason to believe Ih,l! not in fact technically feasible
failure.
IS
lv/a in I ('Il(l/i ce
of data such
evaluation of of data sllch as P-F intervals.
a
individual
Muilltl!fWIlCe
barrier
starts to
often for the first time ever, as a
maintenance programs. on
equipmellt manufacturers
little ofthe information neededto draw up truly contextmailltenanee programs. such programs
also have other agendas when
least of which is to sell
What is more.
either commilling the users' resources to doing the maintenance (in which case
don't have to pay for it, so they have little interest in
it) or tbey may even be bidding to do the maintenance thell1which case
may be keen to do as much as
combination of extraneous commercial about lhe
and
context means that maintenance programs oftell
level of over-maintenance of spares.
computers to dril'(' the process.
]() I1lcntioncd
databases should be used to store
aud sort the informal ion
the
with so mueh in the world of information what
Climb (0 the
to have' uses. tt�",""mn
to computerise ReM
200
201. This up the system so that a 'no' answer
H1 while a
answer leads to one mistaken belief that
Lip or 'streamline' the process. In ff't,orr',nn
to
succession of twelve to twenty
sheet of paper, so
a computer in
process down.
to drive of Clse
can a lso have much
III
asset under review.
Ihis rl,ason the aurhor agrees with Smithl"'" when
he says that there is no "software code to do the
thinking for
us", alld that the computer "doesn't replace the need for solid . ill short.
kllOW�how
ReM is thollghtware, not software.
Conclusion
that the surest way to achieve most if not all of is to do so on a formal basi" tht'
the process at the right
gruups 01"
and maintenance of the
under review.
5
293
What ReM Achieves
14
What ReM Achieves
14.2 Maintenance Effectiveness Chapters I and 2 emphasised that the objective of maintenance is to en sure that any physical asset continues to fulfil its intended functions to the s t ndards of pe formance desired by the user. As a result, any assessment of how well mall1tenance is achieving its objectives must entail an assess
�
14.1 Measuring Maintenance Performance
�
ment of how well the assets are continuing to fulfil their functions to the desired standard. This i s intluenced i n turn by three issues: 'continuity' can be measured i n several different ways • users have different expectations of different functions individual assets can have more than one and often several functions ' as explained i n Chapter 2 . These issues are considered in more detail in the following paragraphs. •
results i n As discuss ed at length i n Chapter 1 1 , the applica tion of RCM : three tangible outcom es, a s follows mainten ance schedul es to be done by the maintenance department •
•
•
revised operating procedures for the operators of the asset of the a list of areas where once-off changes must be made to the design ns asset or the way in which it is operated, in order to deal with situatio current where the asset cannot deliver the desired performance in its configuration.
J3 Two other less tangible outcom es which were mention ed i n Chapter asset the how about deal great a learn process are that particip ants in the works, and also tend to functio n better as teams. effort, Achiev ing all these outcom es requires a great deal of time and r, if Howeve . 3 1 especia lly if RCM is applied as describe d in Chapter costs the h RCM is applied correctl y, i t yields returns which far outweig involve d. Most applicat ions pay for themsel ves i n a matter of months, although some have paid for themselv es i n two weeks or less. The wide in variety of ways i n which RCM pays for itself are discusse d at length we tive, perspec n i on discussi this place to order n I . part 4 o f this chapter the first need to conside r different ways in which i t i s possible to measure . perform ance of the mainten ance function Mainten ance perform ance can be considered from two quite distinct assets viewpoints. The first focuses on how well maintenance ensures that referred continue to do what their users want them to do. This is usually to to as mainten ance effectiveness, and it is likely to be of most interest nt the users or 'customers' of the maintenance service. The second viewpoi is This used. being are es resourc ance mainten concentrates on how well to interest more of usually is It y. referred to as mainten ance efficienc issues manage rs who are directly respons ible for mainten ance. These two chapter. this of are considered separately in the next two sections
•
Different Ways of Measuring Maintenance Effectiveness The primary function of any highly mechanised and fully loaded manu fa turing facility is to produce at least as many units of saleable product as It was expected to produce when it was built. (,Fully loaded' means that
�
it is operating seven days per weekl24 hours per day and that there is a ready market for every unit which the facility can produce. ) In this con text, any failure which reduces output results in lost sales. In cases like these, the simplest overall measure of the operational per formance of the facility as a whole is total output per period. ·
If the facility is not producing what the users usually the owners feel it should be producing on a regular basis, they will not be satisfied until the si uat on is p t right. At least until then, the users will be incl ined to judge e fectIveness 1l1 tenns of total output against targets. This should be recog . l1lsed when settll1g up any system for tracking maintenance effectiveness
: � �
·
�
All is not necessarily well if overall output is o n target. A facility which
�
IS produ lI1g the right number of units could still be experiencing prob lems WhICh affect safety, product quality, operating costs, environmental integrity, customer service, ,md so on, so these also need to be measured and dealt with appropriately. Ther are a great any ways in which we can measure how effecti ve1y . . an asset is fulfillIng Its functions. Five of the most common are as follows:
�
•
�
how often it fails. This is the most widely understood meanin
f
294 •
how long it lasts. This is usually thought of as its 'life' or its 'lifespan', at the end of which the item under consideration fails and is eitherrebuilt or discarded and replaced with a new one. Strictly speaking, this phenom enon should be described as 'durability'.
•
how long it is out (�fservice when it does/ail. This is usually referred
to as 'downtime' or 'unavailability ', and measures how much of the time the item is incapable of fulfilling a stated function to the satisfac tion of the user, in relation to the amount of time the user would like i t to b e capable of doing so. Unavailability ( o r the converse, avail-ability) is usually expressed as a percentage.
•
how likely it is to/ail in the nextperiod assuming that it has survived to
the beginning of that period. We have seen that this is the conditional probability of failure. This could perhaps be described as a measure of 'dependability', if only to distinguish it from the other three variables. One common variation of this measure is the ' B 10 life'. Chapter 1 2 explained that this is usually measured from the moment the item i s put into service, and is the period before which not more than 1 0% of the items can be expected to fail. (In other words, the conditional probabil ity of failure in the stated period is 1 0 % . )
•
What ReM A chieves
Reliability-centred Maintenance
efficiency. I n common business usage, the term efficiency actually has two quite distinct meanings. The first measures output relative to input, while the second measures how well something is performing against how well it should be performing.
For instance, in a power station, energy efficiency measures the amount ofenergy exported in relation to the amount of energy released by the fuel. Depending on the technology used (coal, gas, combined cycle, etc), this usually ranges from about 35% to about 58%. However, if a station which should average 40% energy efficiency is only averaging 38%, it will be exporting 95% of the energy which it should be exporting. In the context of this book, the first (40%) measure is a functional performance standard. As explained in Chapter 3, this is used to judge whether the item has failed. The second (95%) measure is used to judge the effectiveness with which the organisation is achieving the desired performance on an on-going basis. ' Efficiency' also refers to pace of working, and also does so in two ways -how fast an asset shOUld work relative to the pace at which it could work (desired performance versus initial capability), and how fast it actually works relative to the pace at which it should work (actual performance versus desired perform ance). We have seen that desired performance must be less than initial capa bility because allowance must be made for deterioration. So in the context of this chapter, efficiency compares the pace at which an asset actually works with the pace at which it should work, not with the pace at which it could work.
295
'Efficiency'-type measurements can also apply in a slightly different fashion to the consumption of maintenance consumables (such as lubricating oil and hydraulic oil) and process consumables (such as solvents and reagents used in chemical plants and in the extraction of minerals).
All five of these measures are valid. It is simply a matter of deciding which is the most appropriate in the context under consideration. For example, if a turbo-generator set has the lowest energy costs per unit of output among all those used by an electric utility, it is likely that its users would want it to generate (base load) power for as much of the time as possible. In terms of this function, the most appropriate measure of maintenance effectiveness is availability. (The operators may occasionally choose to run the set at less than full load. They may even choose to shut it down completely from time to time for purely operational reasons. Slowdowns or shutdowns of this nature affect the utilisation of the asset as opposed to its availability. In essence, availability measures what percentage of time the machine is available to fulfil its primary performance re quirement, while utilisation measures how much it actually fulfils it.) On the other hand, the generator set might only be used periodically to satisfy peak demands for power (peak loads). In this case, the primary concem of the users will be that the generator comes on stream as soon as it is required, so a primary measure of effectiveness will be how often it does so (or conversely, how often it fails to do so, expressed by a failure rate). When measuring safety , performance is usually measured in terms of number of days or number of manhours worked between lost time incidents (or fatalities). This is a form of 'mean time between failures'. Similar measures are used for en vironmental incidents. On the product quality front, a scrap rate of (say) 4% can be seen as a measure of unavailability, in the sense that while a machine is producing scrap, it is not 'available' to produce first grade product. (A. scrap rate of 4% corresponds to a yield of 96%). Scrap rates can also be expressed as (say) 20 parts per million, which is another way ofexpressing a failure rate. Both are valid measures of maintenance effectiveness, especially in highly mechanised or automated processes.
Different Expectations Every function has associated with it a unique set of continuity (reliability andlor durability and/or availability and/or dependability) expectations. For instance, two of the functions associated with the bodywork of a car are 'to . Most isolate the occupants of the car from the elements' and 'to look car owners expect the bodywork to be able to fulfil the first function throughout the expected life of the car (unless the car is a convertible or unless they open a door or a window). On the other hand, everyone knows that cars get dirty -- and hence start 'to look unacceptable' in the space of a few days or weeks. So in the first case we have a continuity expectation which might be measured in hundreds of thousands of kilometres or decades, while in the second case, the continuity ex pectation is measured in hundreds of kilometres or days.
296
Reliability-centred Maintenance
This issue is complicated by the fact that the loss of nearly every function can be caused by more than one -- sometimes dozens of failure modes. Each failure mode has associated with it a specific failure rate (or MTBF), and each will take the function out of service for an amount of time which i s specific to that failure mode. A s a result, the eontinuity characteristics of any function will actually be a composite of the continuity character istics of all the failure modes which could cause the loss of that function.
For instance, take the function 'to look acceptable' which was mentioned above. In addition to the accumulation of dirt, this function could be lost due to rust or corrosion, fading of the paintwork, external damage (sideswiped in a parking lot) and vandalism, among others. It should also be apparent that some of these failure modes have little or nothing to do with maintenance. For instance, external damage is mainly a function of how this car - or the other vehicle involved - is operated, although design may play a small part by adding rubbing strips to re duce damage and/or by making it easier and cheaper to replace damaged panels. The probability of vandalism is also a function of where the car is used (the opera ting context), so it is almost completely beyond the control of the designer and the maintainer. The rate of accumulation of dirt is a function of where and when a car is used (road conditions and climatic conditions), and it is managed by a suitable maintenance program (washing the car). Corrosion and fading of paintwork can be influenced substantially at the design stage (although yet again the operating context climatic conditions and the provision of shelter - and to some extent maintenance activities - polishing and chassis washing - can play a part in mod erating the severity and frequency of these failures).
This example leads to two important conclusions, as follows: •
•
we need a thorough understanding of all the failure modes which are likely to cause each loss of function in order to be able to design, operate and maintain an asset in such a way that the effectiveness expectations which we have of each function will be achieved it is unreasonable to hold the maintainer of an asset alone accountable for the achievement of any continuity (reliability/availability/dur abil ity/dependability) targets for any asset or any function of any asset. The
achievement of these targets is also a function of how it is designed, built and operated. Accountability for achieving the associated targets should be divided jointly between the people responsible for all ofthese functions. (In other words, 'maintenance' effectiveness as i t is being defined in this chapter is not only a measure of the effectiveness of the
maintenance department. It measures how effectively everyone associ ated with the asset is playing their part i n doing whatever i s necessary to ensure that it continues to do what its users want it to do. )
What ReM Achieves
297
Different Functions Perhaps the most important point about measuring the effectiveness of maintenance activities is the fact that every asset has more than one and sometimes dozens of functions. A s explained above, a unique set of con tinuity expectations is associated with each function. This means that if an sse has ten functions, then the etTecti veness with which the asset is being
� �
mamtamed can be measured in (at least) ten different ways. For instance, let us consider how maintenance effectiveness might be measured by t e ow ��r of a typical suburban gas station. For the purpose of this example, the asset IS a storage and pumping system used for gasoline. In this system, unleaded gasoline IS stored in an underground tank with a capacity of 50 000 litres. It is periodically filled by a road tanker to a level of 48 000 litres. An upper level sWitch In the tank sWitches on a local warning light if the tank has been filled to a lev � 1 of 48 500 litres, and another switches on another warning light in the main office If the level drops to 5 000 litres. A low level alarm sounds in the office if the tank level drops to 2 000 litres, and a local ultimate high level alarm sounds if the tank level reaches 49 000 litres. The tank is double skinned to ensure that gasoline is contained in the event of a leak in the inner skin. A level indicator indicates the fuel level in the tank. The tank supplies gasoline to five pumps. Each pump is switched on and off by pressing and releasing a handle in the nozzle. The nozzle also incorporates a pressure sWitch which tnps the pump when the vehicle fuel tank is filled to the tip of the nozzle. A flow meter measures the amount of fuel delivered each time the pump is activated and displays the volume and value of the fuel delivered to the customer. This meter is zeroed each time the nozzle is returned to its cradle. (This system embodies additional secondary functions which deal with access onto and into the tank, drains, venting, valving, ease of use by a customer, other protection, appearance and so on. These would also be listed in a real-life situa tion. However, for the purpose of this example, we only conSider the functions described above.) On this basis, a list of functions might read as follows: to pump between 25 and 40 litres/minute of gaSOline to the vehicle • to indicate volume and value of fuel delivered to customer to within 0.03% of actual volumelvalue to shut off pump when required by customer or when customer's fuel tank is full. to contain the gasoline • to store between 2 000 and 48 000 litres of gasoline to switch on a warning light in main office if the tank level drops to 5 000 litres to switch on a local warning light if the tank level reaches 48 500 litres to sound an alarm in main office if the tank level drops below 2 000 litres to sound an alarm if the tank level reaches 49 000 litres • to contain the contents of the tank in the event of a leak to indicate the level of fuel in the tank to within 0.05% of the actual level When assessing the maintenance effectiveness of this system, the owner of the gas station will have different criteria for each of the above functions. For instance:
�
•
•
•
•
•
•
•
•
298
•
Reliability-centred Maintenance
What ReM Achieves
Function 1: to pump between 25 and 40 litres/minute of gasoline to the vehicle.
- Functional failure B: indicates that more than 0.03% more fuel has been delivered
This function can fail in three ways with three quite different sets of consequences, so each functional failure needs to be considered on its own merits, as follows:
than actual: If this happens and it comes to the attention of either the customers or the trading standards authorities (probably both), the station owner would be in serious trouble. Many of his customers would regard him as a crook and take their business elsewhere. The authorities would probably fine him and, depending on the seventy of the discrepancy, might even revoke his license to trade (thus putting him out of business). Whatever else happens, his standing in the com mUnity would take a beating. The severity of these consequences would lead him to seek a v�ry low failure rate- say once in 50 000 years on any one pump. (Whether thiS IS achievable or not IS another issue entirely.)
- Functional failure A: fails to pump at all:
Obviously, if a pump isn't working, it cannot be used to pump gasoline. However, there are five pumps in the station so the level of availability required depends on the pattern of demand. For example, the station owner may tell us that he "hardly ever" has all five pumps in use at once so seldom that we can ignore the possibility. He might also tell us that four pumps tend to be in use simultaneously for a total of not more than one hour a day, and then never for more than about ten minutes at a time. If each pump has an average availability of 95%, two pumps will be out of service simultaneously for no more than 2% of the time. In other words, four pumps would be available 98% of the time, while there is a demand for four pumps 4% of the time. Under these circumstances, only a tiny fraction of customers would need to wait for gasoline, and then not for very long. This might tempt the owner to accept an availabilityof 95%. (If he regularly had five or more customers want ing to buy gas at the same time, he would expect a much higher availability. But it may cost him somewhat more to achieve it, especially jf he has to pay a premium for rapid response when calling out technicians to deal with failures.)
•
carnes on pumping after the customer releases the handle, the back pressure sensor should shut it off when the tank is full. As a result, the customer will end up with much more gas in the tank than he or she wanted. This would almost certainly lead to a row about how much should be paid for and possible loss of a customer. As a result, the station owner would probably require a fairly low failure rate - say once in 1 000 years on any one pump. - Functional failure B: fails to shut off when tank is full: Many customers rely on
the sensor to tell them when the tank is full. If it fails to do so, the pump should shut off when the customer releases the handle. However, it is likely that the tank will overflow onto the shoes of the customer before he or she is able to react, leading to a lot of unpleasantness and perhaps a demand for compensa tion. This too would lead the owner to expect a low failure rate-· say again once in a 1 000 years for any one pump
- Functional failure C: pumps more than 40 litres/minute: If the pump pumps too
-
- Functional failure C: both local switches unable to switch off pump:
If the sensor and the handle both fail to shut 9ft the pump, it will carry on pumping gasoline all over the forecourt until the electrical supply is shut off at the main circuit breaker. This would create a nasty fire hazard, so the owner would expect a very low failure rate- say once in 1 000 000 years. (This is attainable if each switch independently achieves 1 in 1 000.)
•
Function 2: to indicate volume and value of fuel delivered to customer to within 0.03% of actual volume/value.
This function can fail in two ways, as follows:
- Functional failure A: indicates that more than 0.03% less fuel has been delivered than actual: If this happens, the station owner appears to be selling less fuel than he is actually selling, so he loses money. The failure becomes apparent after a while, because the ratio of fuel sold to fuel received will start to decline. Nevertheless, the owner would probably still seek a low failure rate - say not more than one in 1 000 years on any one pump. (If the indicator fails completely, it shows that nothing has been delivered. If this happens, one lucky customer might get a free tank of fuel, then the station manager would shut down the affected pump until the problem is rectified.)
This function can also fail in three ways, as follows:
- Functional failure A: fails to shut off when required by the customer: If the pump
might find slow pumps sufficiently irritating to take their business elsewhere, especially if there are faster alternatives nearby. Consequently, the owner is likely to want any of his pumps which wasn't failed completely to pump at the required rate "all the time or at least, as close to all the time as you can make it". This might turn out to mean (say) 99.8% of the time that the pump is not otherwise out of action another form of 'availability'. fast, it is likely to generate sufficient back pressure to keep tripping the 'tank full' pressure senSing mechanism in the nozzle. Customers would have to learn to throttle back the filling rate by not depressing the handle so much, which many regulars might also find irritating enough to cause them to take their busi ness elsewhere. As a result, the owner is likely to say that he wouldn't want this failed state to occur "too often". He might then quantify this expectation as a failure rate say not more than once in fifty years on any one pump.
Function 3: to shut off pump when required by customer or when customer's fuel tank is full.
- Functional failure B: pumps less than 25Iitres/minute: Some regular customers
•
299
•
Function 4: containment: When asked about this function, the station owner might say something like "we have had one leak in the gasoline system in the last ten years - and that was one too many." Here the user is measuring effect iveness in terms of a failure rate. When pressed, he might accept a rate of (say) one In 500 years for a 'small' leak, which he might choose to define as less than 5 litres per hour. (It is highly unlikely that anyone would measure containment in terms of availability, because (say) 99% availability means that the system would be leaking 1 % of the time - about 800 hours out of ten years. Even 99.9% still means that it would leak for 80 hours. Clearly this is nonsense.) Function 5: to store between 2 000 and 48 000 litres of gasoline. This function can also fail in three ways, each of which must again be considered separately, as follows:
300
Reliability-centred Maintenance
What ReM Achieves
- Functional failure A: level drops below 2 000 litres: Based on normal patterns
of demand, fresh supplies of gasoline are ordered when the tank level ap proaches 5 000 litres, and we are told that they are nearly always delivered before the level reaches 2 000 litres. If the level in the tank drops much below 2 000 litres, there is a greatly increased chance that the tank will empty, caus ing the station to lose business. As a result, the station manager expedites the delivery if the level drops to 2 000 litres (as indicated by the low level alarm). He says he needs to expedite deliveries about once a year, which he says is "just about acceptable". Here he is again judging effectiveness in terms of a failure rate. (Note that this failed state is caused by increased demand and/or slow delivery. It has nothing to do with the maintenance department in the clas sical sense. Nonetheless, dealing with this failure can be seen as 'maintenance' because we are seeking to 'cause the business to continue'.) - Functional failure B: level rises above 48 000 litres: The level in the tank is only
likely to rise above 48 000 litres if the delivery driver is not paying attention to the tank level indicator when filling the tank or if the level indicator itself has failed. In both cases the warning light comes on at 48 500 litres. We are told that this happens "about once every six months" - another failure rate which the people involved might say they accept.
•
Function 8: to sound an alarm if the level in the tank drops below 2 000 litres.
If the level in the tank drops to 2 000 litres and the low level warning does not sound, the delivery is not expedited. We are told that under these circumstances, there is a 50% chance that the tank will run dry before the tanker arrives, and the station would be out of gaSOline for about one hour on average under such circumstances. This leads the station owner to conclude that he will not accept this multiple failure (level drops below 2 000 litres while low level alarm is failed) more than "once in a hundred years" (MMF 1 00 years). As discussed above, MTED is one year, so the station can tolerate a maximum unavailability for the low level alarm of MTED/MMF 1 / 1 00 1 %. In the light of this objective, the low level alarm is being maintained effectively if its availability remains above 99%,. =
S imilar logic would be followed to determine availabilities for functions 9. 1 0 and II in the above example. It would also be used to estahlish effec tiveness measures for the functions of this system that were not included i n the ahove list. However, for the functions discussed. the effectiveness
expectations of the gas station owner can be summarised as follows: Function
FUnctional Failure
1
A B C A B A B C A A B C A A A
- Functional failure C: tank contains something other than gasoline: The
tank can only contain something other than gasoline if it is filled with something else - (say) diesel. If this happens, customers could fill their tanks with the wrong fuel and cause serious damage to their engines. The station owner figures that the resulting bad publicity and claims for damages could put him out of bus iness, so he would rather this didn't happen at all. When reminded that 'never' is an unattainable ideal, he might decide to accept a failure rate of (say) once in 1 00 000 years.
•
Function 6: to switch on a local warning light if level drops to 5 000 litres. The station manager usually logs the level in all the fuel tanks every day in order to track consumption, and orders more fuel when levels approach 5000 litres. The low level warning light serves as a reminder if the level indicator fails or if there is a sudden surge in demand between readings. This light is needed about once every two years (MTED 2 years). If it does no� work when needed, t e low level . alarm sounds when the level drops to 2 000 htres. If an Initial order IS placed at this late stage, the tank will almost certainly run dry and the station will be out of gasoline for several hours. The owner says he will accept a mean time between occurrences of this multiple failure (MMF) of 400 years. In the light of this expecta tion, the formula on page 1 16 tells us that the maximum unavailability the station can tolerate for the low level warning light is MTEJMMF 2 1 400 0.5%. This means that the low level alarm is being maintained effectively if its availability remains above 99.5%. =
•
2 3
4 5
�
=
6 7 8
Measure of Effectiveness Availability MTBF ;> 95% ?: 99.8% ?: 50 years ?: 1 000 years
50 000 years ?: 1 000 years ;> 1 000 years 1 000 000 years
500 years ?: 1 year :� 6 months 1 00 000 years
?: 99.5% ?: 97.5% ?: 99%
Comments Each pump Each pump Each pump Each pump Each pump Each pump Each pump Each pump Whole system Tank Tank Tank UL warning light H/L warning light UL alarm
The example illustrates several important points about the measurement of maintenance effectiveness, as follows:
=
Function 7: to switch on a local warning light if level rises to 48500 litres. The high level warning light is backed up by an audible alarm, so following similar logic to the above example, the owner might come to the conclusion that he will accept an availability of 97.5% for this warning light.
30]
•
when measuring maintenance performance, we m-e not measuring equip ment effectiveness - we are measuring functional effectiveness. The distinction is important. because sh ifting emphasis from the equipment to its functions helps people maintainers i n particular what the equipment does rather than what i t is.
to focus on
302 •
•
What ReM A chieves
Reliability-centred Maintenance
even quite simple assets have a surprisingly large number of functions . Each of these functions has a unique s e t of performance expectations. Before it is possible to develop a comprehcnsive maintenance eflec ti veness reporting system, we need to know what all these functions are, and we must be prepared to establish what the user thinks is acceptable
303
For instance, the primary function of a milling machine might be: To mill 1 01 ± 1 workpieces per hour to a depth of 1 1 ±. 1 0101. If this machine is out of action completely for (say) 5% of the time, its availability is 95%. If it is only able to produce 96 pieces per hour when it is running, its effici ency is 96%. If 2% of its output are rejects, its yield is 98°/". Applying the above formula gives an overall effectiveness of 0.95 x 0.98 x 0.96 0.894, or 89.4%. •
or otherwise in each case.
This particular composite measure is sometimes refelTed to as 'overall
This means that it is not possible to list a single continuity statement for an entire asset, such as "to fail not more than once every two years" or "to last at least eleven years". We need to be specific about which function must not be lost more than once every two years or must not fail for at least eleven years (or more precisely, which functional failure must not occur more than once every two years, or which functional failure must not occur before eleven years).
equipment effectiveness', or OEE. Composite measures of this sort are popular because they allow users to assess maintenance effectiveness at a glance. They also seem to offer a basis for comparing the performance of similar assets (so-called 'benchmarking'). However, these measures
there is often a tendency to focus too heavily on primary functions when assessing maintenance effectiveness. This is a mistake, because in practice apparently trivial secondary functions often embody bigger threats to the organisation if they fail than primary functions. As a result, every
actually suffer from numerous drawbacks, as follows: •
For instance, in the milling machine example above, the workpiece may have a work-in-process value of£200 at that point in the process. The organisation might be making a gross profit of £ 1 00 on a finished product sale price of (say)£500. This means that 1 % downtime or 1 % loss of efficiency costs the company one sale per hour-a lost profit of£1 00 per hour. On the other hand, 1 scrap means that the organisation has to write off 1 workpiece per hour, representing £200worth of work-in-process in addition to£ 1 00 lost profit - a total loss of £300 per hOUL Consequently, the machine in the above example is losing: (5 x 1 00) + (4 x 1 00) + (2 x 300) £ 1 500 per hour due to downtime, slow running and rejects. However, an identical machine pro ducing the same product might suffer from 4% downtime, run at 98% of its rated speed and produce 4% scrap. In this ca$e the 'overall effectiveness' would be 0.96 x 0.98 x 0.96 0.903, or 90.3%. This is apparently a better performance than the first machine. However, this machine is losing: (4 x 1 00) + (2 x 1 00) + (4 x 300) £ 1 800 per hour which is actually a significantly worse performance than the first machine!
function must be considered when setting up maintenance effectiveness measures and targets. For instance, the primary functions listed for the gasoline system are to pump and to store fuel ( Functions 1 and 5 respectively). However, two of the highest expectations of the owner centred around two secondary functional failures 2-8 (a failure which could put him out of business) and 3-C (a failure with serious safety implications).
Multiple Performance Standards and the OEE If a function embodies multiple performance standards, it is tempting to try to develop a single composite measure of effectiveness for the entire function. For instance, the primary function of a machine performing a conversion operation in a manui�lcturing facility usually incorporates three performance standards, as follows: i t must work at all • i t must work at the right pace it must produce the required quality. The effectiveness with which it continues to meet each of these expecta tions is measured by availability, efficiency and yield. This suggests that a com-osite measure of the effectiveness with which this machine i s •
•
fulfilling i t s primary function o n an on-going basis could be determined by multiplying these three variables, as follows: overall effectiveness availability x efficiency x yield.
the use of three variables in the same equation implies that all three have equal weighting. This may not be the case in practice.
=
•
It is possible for many assets to operate too fast as well as too slowly. Overspeeding an asset would increase the OEE as defined above, whieh means that i t possible to obtain an apparent improvement in 'overall' performance by forcing the asset to operate in a failed state. For instance, a primary performance standard of the milling machine was that it should produce 1 0 1 ±1 workpieces per hour. The '+ l' means that if the machine produces more than 1 02 units per hour, it is in a failed state (perhaps because it starts going faster than a bottleneck assembly process, leading to a build-up of work-in-process, or because going too fast causes the milling cutter to over heat and damage the workpieces or because it leads to excessive tool wear.) However, if it operates at 1 03 workpieces per hour, the apparent 'efficiency' is 1 02%. This increases the 'overall equipment effectiveness' as defined above, at a time when the machine is actually in a failed state. This is clearly nonsense.
304
What RCM A chie ves
Reliability-centred Maintenance
functio n of any the GEE as defined above only relates to the primary gasolin e storage asset. This is mislea ding, becaus e as in the case of the more func system, every asset machin e tools included - have many have their own tions than the primary functio n, and each of these will i s not a meas GEE the uently, Conseq ations. unique performance expect the effecti ve of e measur a only but ure of 'overall' effecti veness at all, fulfille d. being is asset the of ness with which the primary functio n
•
maintenance finally, for the reasons discuss ed earlier, truly user-oriented ent' effec 'equipm from away n attentio their turn to enterprises need this sort of es measur if o S . veness effecti nal tivene ss towards functio es of measur as them to refer to te accura must be used, it is much more ent equipm 'overall than rather (PFE) ' prim,uy functional effectiveness'
•
effect ivenes s' .
Conclusion
The two most important conclu sions to emerge from part are that:
•
Maintenance costs The costs referred to in this part of this chapter are the direct costs of main tenance labour, materials and contractors, as opposed to the indirect costs associated with poor asset performance. The latter issues were discussed in Part 3 of this chapter. In many industries, the direct cost of maintenance is now the third highest element of operating costs, behind raw materials and either direct production labour or energy. In some cases, it has risen to second or even first place. As result, controlling these costs has become a top priority. Some industries offer scope for substantial reductions in direct main tenance costs, especially those whose processes embody mature or stable technologies and/or which have a large legacy of second generation think ing embodied in their maintenance practices. However, in other industries, especially those which are newly mechanising or automating their pro
2 of this chapte r
is making to the when evalua ting the contrib ution which mainte nance each function perform ance of any asset, the effecti veness with which This in turn is being fulfilled must be measured on an on-goi ng basis. ns of the asset, require s a crystal clear unders tanding of all the functio meant when it is what of tanding unders clear y together with an equall is said to be 'failed ' .
•
305
expectations must the ultimate m'biter of effectiveness i s the user (whose legitim ately in turn be realist ic) . What users expect will vary quite depen ding on the from functio n to functio n and from asset to asset, operating conte xt.
14.3 Maintenance Efficiency maint enanc e efficie ncy A s menti oned at the begin ning of this chapter, the resour ces at its measu res how well the mainte nance functi on i s using be done are gener can this which in dispos al. The J arge numbe r of ways in this part of this briefly sed ally well understood, so they are only discus
chapte r for the sake of compl etenes s. ries, These are Efficie ncy measu res can be grouped into four catego ng and control. maintenance costs, labour, spares and materials and planni
cesses at a significant rate, the sheer volume ofmaintenance work to be done is often growing at such a pace that maintenance costs are likely to rise i n absolute terms over the next ten years or s o . As a result, take care to evalu ate the pace and direction of technological change before committing to substantial long-term reductions i n total maintenance costs. The most common ways i n which maintenance costs are measured and analysed are as follows: • Total cost of maintenance (actual and budgeted) - for the entire facility - for each business unit - for each asset or system • Maintenance cost per unit of output •
R atio of parts to labour expenditure.
Labour The cost of maintenance labour typically amounts to between one third and two thirds of total maintenance costs, depending on the industry and overall wage levels in the country concerned. In this context, mainte nance labour costs should include expenditure on contract labour (which i s often incorrectly - grouped under ' spares and materials' because it i s bought out). When considering maintenance labour, it is also wise not to make the common mistake of treating maintenance work done by op erators as a zero cost "because the operators are there anyway". In u sing operators for this work, the organisation i s still committing resources to maintenance, and the cost should be acknowledged accordingly .
306
Reliability-centred Maintenance
Common ways of measuring and analysing maintenance labour effi ciency include the following: Maintenance labour cost (total and per unit of output) Time recovery (time performing specific tasks as a percentage of total •
•
•
•
time paid for) Overtime (absolute hours and as a percentage of normal hours) Relative and absolute amounts of time spent on different categories of work (proactive tasks, default actions and modifications, and subsets of
•
•
these categories) B acklog (by number of work orders and by estimated hours) Ratio of expenditure on maintenance contractors to expenditure on fuUtime maintenance employees.
Spares and materials
Spares and materials usually account for the portion of maintenanc e expenditure which does not come under the heading of 'labour' . How well they are managed is usually measured and analysed in the following ways: Total expenditure on spares and materials (total and per unit of output) • Total value of spares i n stock Stock turns (total value of spares and materials in stock divided by the •
What ReM Achieves
307
Some of these efficien cy measur es are useful for making immed iate de cisions or initatin g short-term manage ment action (expen diture against budget s, time recove ry, schedu le comple tion rates, backlo gs). Others are more u seful for trackin g trends and compar ing perform ance with similar facilitie s in order to plan longer term remedial action (mainte nance costs per unit of output, service levels and ratios in general ). Togeth er, they are a great help in focusin g attentio n on what must be done to ensur that mainten ance resource s are used as efficien tly as possible . Mainten ance efficien cy is also quite easy to measure . The issues whic h i t addresse s are usually under the direct control o f mainten ance manage rs. For these two reasons, there is often a tendency for these manager s to focu s t o o much attention on efficienc y a n d not enough on mainten ance effectiv eness. This is u nfortuna te, because the issues discllsse d under the heading of mainten ance effectiv eness usually have a much greater impact on the overall physical and financia l well-bei ng of the organisa tion than those discusse d under maintenance efficiency. Truly 'customer'-oriented mainten ance managers direct their attention accordin gly. As Part 4 of this chapter explains , the greatest strength of RCM is the extent to which it helps them to do so.
�
•
•
•
total annual expenditure on these items) Service levels (percentage of requested stock items which are i n stock at the time the request is made) Relative and absolute values of different types of stocks (consumabl es, active spares, 'insurance' spares, dead stock).
Planning and control How well maintenance activities are planned and controlled affects all other aspects of maintenanc e effectivene ss and efficiency, from the over all utilisation of maintenance labour to the duration of individual stop
pages. Typical measures include: Total hours of predictive/preventive/failure-findi ng maintenanc e tasks issued per period The above hours as a percentage of total hours Percentage of the above tasks completed as planned Planned hours worked vs unplanned hours Percentage of jobs for which the time was estimated
14.4 What ReM Achieves Figure 1 . 1 in Chapter 1 , reproduced as Figure 1 4 . 1 , shows how expecta tions of the maintena nee function have evolved over the past fifty years. Third Generation:
Figure 14.1: Growing expectations of maintenance
•
HigherplantavailaQility and reliability
•
•
•
•
•
•
Accuracy of estimates (estimated hours vs actual hours for jobs which were estimated).
1940
1 950
1 960
1970
1980
1990
2000
The use of RCM helps to fulfil all of the Third Generation expectations. The extent to which it does so is summarised in the following paragraphs, starting with safety and environmental integrity.
308
Reliability-centred Maintenance
Greater Safety and Environmental Integrity RCM contributes to improved safety and environmental protection in the following ways: •
the systernatic review of the safety and environmental implications of every evidentfailure before considering operational issues means that
t
safe y and environmental integrity become - and are seen to become top maintenance priorities. •
from the technical viewpoint, the decision process dictates that fail ures which could affect safety or the environment must be dealt with in some fashion it simply does not tolerate inaction, As a result, tasks are selected which arc designed to reduce all equipment-related safety or environmental hazards to an acceptable level, if not eliminate them completely. The fact that these two i ssues are dealt with by groups which include both technical experts and representatives of the 'likely victims ' means that they are also dealt with realistically.
•
the structured approach to protected systems, especially the concept of the hidden function and the orderly approach to failure-finding, leads to substantial improvements in the maintenance of protective devices. This greatly reduces the probability of multiple failures which have serious consequences. (This is perhaps the most powerful single feature of RCM. Using it correctly significantly lowers the lisk of doing business.)
•
involving groups of operators and maintainers directly in the analysis makes them much more sensitive to the real hazards associated with their assets. This makes them less likely to make dangerous mistakes, and more likely to make the right decisions when things do go wrong.
•
the overall reduction in the number and frequency of routine tasks (especially invasive tasks which upset basically stable systems) reduce the risk of critical failures occurring either while maintenance is under way or shortly after start-up. This issue is particularly i mportant if we consider that preventive maintenance played a part in two of the three worst accidents in industrial histo ry (Bh �pal, . Chernobyl and Piper Alpha). One was caused directly by a proactive mainte nance intervention which was currently under way (cleaning a tank full of methyl isocyanate at Bhopal). On Piper Alpha, an unfortun �te seri �s of i�cidents and oversights might not have turned into a catastrophe If a crUCial relief valve had not been removed for preventive maintenance at the time.
What ReM Achieves
309
A s men tione d i n Part 2 of this chap ter, the mos t com mon way to track performance in the areas of safety and envi ronmental integ rity is to record the num ber of incid ents whic h occur, typic ally by reco rding the num ber of lost- time accid ents per milli on man -hou rs in the case of safety, and the num ber of excursions (inci dent s where a standard or regulation is breached) per year in the case ofthe environment. Whil e the ultimate target in both cases is usua lly zero, the short-term target is alwa ys to better the previOlls record.
To provi de an indica tion of what RCM has achie ved in the field of safety, Figur e 1 4.2 s hows the n u b r f accid ents per millio n take- ofts recor ded each year in � � � . the comm ercia l CIVil aViat ion Il1dus try over the perio d of deve lopm ent of the RCM philosophy (exclu ding accidents cause d by sabotage, military action o r turbU lence ). The percentage of these crash es which were cause d by equip ment failure also declin �d. Much of the i mpro ved reliab ility is of cours e due to the use of supe rior mate nals and great er redun danc y, but most of these impro veme nts were drive n i n turn by the realisation that main tenan ce on its own could not extract the required level of perfo rman ce from the assets as they were then confi gured . As expla ined In Chapter 1 2, this shifted attent ion from a heavy reliance on fixed time overhauls in the 1 960's to doing what ever is nece ssary to avoid or elimin ate the conse quences of failur es, be it maint enan ce or redeS ign (the corne rston e of the RCM philos ophy ). It also reduc ed the numb er of crash es which might other wise have been cause d by inapp ropria te maint enanc e interv ention s. 30
t
III :t: 0 Q)
25
20
.:.:: «I r::
�
15
'E ...
Q) 0.. III r:: Q)
"0
10
5
'(3 (.)
«
Year � Figure 14.2:
Safety in the civil aviation industry
Source: C A Shifrin: "A viation Safety Takes Center Stage Worldwide" A viation Week & Space Technology: Vol 145 No 19: Pp 46 48
310
Reliability-centred Maintenance
What ReM Achieves
Higher Plant A vailahility and Reliability
•
The scope for performance improvemen t clearly depends on the perform ance at the outset. For example, an undertaking which is achieving 9 5 % availability has l e s s improvemen t potential than o n e which i s currently only achieving 85%. Nonetheles s, if it is con'ectly applied, RCM achieves significant improvem ents regardless of the starting point.
•
For instance, the application of ReM has contributed to the following : a 16% increase in the total output of the existing assets of a 24-hour 7-day milk-processing plant. This improvement was achieved in 6 months, and most of it was attributed to an exhaustive ReM review done during this period. a 300-ton walking dragline in an open-cast coal mine whose availability rose from 86% to 92% in six months. a large holding furnace in a steel mill which achieved 98% availability in its first eighteen months of operation against an expectation of 95%.
For instance, a ReM enabled a major integrated steelworks to eliminate a/l fixed-interval ove.rhauls from its steel-making division. In another case, the inter �als between major overhauls of a stationary gas turbine on an oil platform were Increased from 25 000 to 40 000 hours without sacrificing reliability. •
the systematic review ofthe operational consequences ofeveryfailure which has not already been dealt with as a safety hazard, together with the stringent criteria used to assess task effectiveness, ensure that only the most effective tasks are selected to deal with each failure mode. the emphasis placed on on-conditi on tasks helps to ensure thatpotential failures are detected before they becomefunctionalfailures. This helps reduce operationa l consequen ces in three ways:
problems can be rectified at a time when stopping the machine will have the least effect on operations
•
it is possible to ensure that all the resources needed to repair the fail ure are available before it occurs, which shortens the repair time - rectification is only carried out when the assets really need it, which extends the intervals between corrective interventions. This in turn means that the asset has to be taken out of service less often. For instance, the example concerning tyres on page 1 30 shows that the tyres need to be taken out of service 20% less often for retreading if on-condition maintenance is used instead of scheduled restoration. In this case, the effect on the availability of the vehicle would be marginal, because removing a lyre and replacing it with a new one can be done very quickly. However, in cases where the corrective action requires extensive downtime, the improvement in availability could be substantial.
n:
time without a correspon ding increase in unschedu led downtime .
•
•
the previous example suggests that greater emphasis on on-condi tion aintenance reduces thefrequency ofmajor overhauls, with a correspon
anyfrequency. This leads to a reduction in previousl y scheduled down
•
•
by relating each failure mode to the relevant function al failure, the in formatio n workshe et provides a tool for quickfailure diagnosis, which leads i n turn to shorter repair times.
dmg long-term increase in availabili ty. In addition, a comprehe nsive list of all the failure modes which are reasonabl y likely together with a dispassion ate assessmen t of the relationsh ip between age and failure, reveals that there is often no reason at all to pel/Orin routine overhauls at
•
Plant performance is of course improved by reducing the number and the severity of unanticipat ed failures which have operational consequenc es. The ReM process helps to achieve this in the following ways:
311
in spite of the above comments , it is often necessary to plan a shutdown or an overhaul for any of the following reasons: - to prevent a failure which is genuinely age-related - to rectify a potential failure - to rectify a hidden functional failure - to carry out a modification. In these cases, the disciplined review of the need for preventi ve or cor rective action that is part of the ReM process leads to shorter shutdown worklists, which leads in turn to shorter shutdowns . Shorter shutdowns are easier to manage and hence more likely to be completed as planned. short shutdown worklists also lead tofewer i4ant mortality problems when the plant is started up again after the shutdown, because it is not
disrupted as much. This too leads to an overall increase in reliability. •
as explained on page 268, RCM provides an opportunity for those who participate i n the process to learn quickly and systematically how to
operate and maintain new plant. This enables them to avoid many of the
errors which would otherwise be made as a result of the learning pro cess, and to ensure that the plant is maintained cOiTectly from the outset. At least four organisations with whom the author has worked in the UK and the USA achieved what each described as 'the fastest and smoothest start-up in the company's history' after applying ReM to new installations. In each case, ReM was applied in the final stages of commissioning. The companies con cerned are in the automotive, steel, paper and confectionery sectors.
3 12
What R eM A chieves
Reliability-centred Maintenance
the elimination ofsuperfluous plant and hence of superfluous failures. As mentioned in Chapter 2, it is not unusual to find that between 5% and 20% of the components of a complex plant are utterly superfluous, but can still disrupt the plant when they fail. Eliminating such components leads to a corresponding increase in reliability.
•
•
by using a group of people who know the equipment best to c arry out a systematic analysis of failure modes, i t becomes possible to ide�tify and eliminate chronicfailures which otherwise seem to defy detectIOn, and to take appropriate action.
Improved Product Quality on pages 48 and By focusin g dircctl y on produc t quality issues as shown proces ses. ated 49, RCM does much to improve the yield of autom to reduce scrap rates For instance, an electronics assembly operation used RCM ppm. 50 to ) million per from 4% (4 000 parts
) Greater Maint enanc e Efficiency (Cost-effectiveness of growth of mainte ReM helps to reduce, or at least to control the rate nance costs in the follow ing ways:
Less routine maintenance:
g fully-developed Wherever ReM has been correctly applied to an existin to 70% of tion redu a to led has � preven tive mainte nance system , it partly IS IOn reduct ThIS oad. i n the percei ved routin e mainte nance workl l overal an to due mainly due to a reduct ion in the numb er of tasks, but s i RCM if that sts increa se in the interv als betwe en tasks. It also sugge equip ment o r for equip used t o develo p maint enanc e programs for new ntive maint enanc e ment which is currently not subjec t to a formal preve 70% lower than if the program, the routin e workload would be 40% means . maint enanc e program were develo ped b y any other enanc e means maint uled' 'sched or e' Note that in this contex t, 'routin g of the loggin daily the it be basis, any work undertaken on a cyclic annual an g, readin ion vibrat readin g on a pressure gauge, a month ly al interv fixedearly functi onal check of a temperature switch or a five-y sched tasks, n nditio overh aul. In other words , it cover s sched uled on-co g. -findin led failure uled restoration, schedu led discard tasks and schedu in routine maintenance For example, RCM has led to the following reductions s: workloads when applied to existing system
�Oo/�
• •
•
•
313
a 50% reduction in the routine maintenance workload of a confectionery plant a 50% reduction in the routine maintenance requirements of the 1 1 kV trans formers in an electrical distribution system an 85% reduction in the routine maintenance requirements of a large hydraulic system on an oil platform a 62% reduction in the number of low frequency tasks which needed to be done on a machining line in an automotive engine plant.
Note that the reductions mentioned above are only reductions in perceived routine maintenance requirements. In many PM systems, fewer than half of the schedules i ssued by the planning o ffice are actuall y completed. This figure is often as low as 30%, and sometimes even lower. In these cases, a 70% reduction in the routine workload will only bri ng what i s i ssued into line with what is actually being done, which means that there will be no reduction in actual workloads. Ironically, the reason why so many traditionally-dcrived PM systcms suffer from such low schedule completion rates is that much of the routine work i s perceived --correctly - to be unnecessary. Nonetheless, if only a third of the prescribed work is being done in any system, that system is wholly out of control. A zero-based ReM review does much to bring situations like these back under controL
Better buying of maintenance services Applying RCM to maintenance contracts leads to savings in two areas. Firstly, a clear understanding of failure consequences allows buyers to specify response times more precisely even to spec ify different res ponse times for different types offailures or different types of equipment. Since rapid response is often the most costly aspect of contract mainte nance, j udicious fine-tuning i n this area can lead to substantial savings. Secondly, the detailed analysis of preventive tasks enables buyers to reduce both the content and the frequency of the routine portions of main tenance contracts, usually b y the same amount 70%) as any other schedules which have been prepared on a tradi tional basis. This leads to corresponding savings i n contract costs. -
Less need to use expensive experts If field technicians employed by equipment suppliers attend ReM meetings as suggested on page 269, the exchange of knowledge which takes place leads to a quantum j ump in the ability of the maintainers employed by the users to solve difficult problems on their own. This leads to an equally dra matic drop in the need to call for (expensive) help thereafter.
314
Reliahility-centred Maintenance
Wha t ReM Ach ieve s
,A� �
Clearer guidelinesf()r acquiring new maintenance technology
�
quicker failure diagnosis means that less time i s spent on each repair detecting potentialfailures before they hecome functionalfailures not only means that repairs can be planned properly and hence carried out more efficiently, but it also reduces the possibility of the expensive secondary damage which could be caused by the functional failure
� :
the reduction or elimination of overhauls together with shorter work listsfor the shutdowns which are necessary can lead to very substantial
•
the elimination ofsuperj7uous plant also means the elimination of the need either to prevent it [rom fail ing in a way which i nterferes with production, or to repair it when i t does fail
The most spectacular case of this phenomenon encountered by the author con cerned a single failure mode caused by incorrect machine adjustment (operator error) in a large process plant. It was identified during an ReM review and was thought to have cost the organisation using the asset just under US$1 million in repair costs alone over a period of eight years. It was eliminated by asking the operators to adjust the machine in a slightly different way.
Longer Useful Life of Expensive Items
�
the ReM
process does much to help ensure that just about any asset can be made to last as long as its bas i c supporting structure remains intact and spares remain available.
� �
�
Better Teamwork In a curio us way, teamwork seem s to have become both a means to an cnd and an end in itsel f in man y organisat ions. The ways in which the highly struc tured ReM approach to maintena nce prob lem anal ysis and deci sion mak ing con ribut es to team build ing were summ arise d on page 268. Not only does thIS appr oach foster teamwork with in the revie w grou ps them selve s, but it also impr oves com mun icati on and co-o pera tion betw een: • prod uctio n or oper ation s departments and the main tenan ce func tion mana geme nt, supe rviso rs, techn ician s and operators equip ment desig ners, vendors, users and main taine rs.
�
•
By ensuring that each asset receives the b are minimum of essential main tenance i n other words, the amount of maintenance needed to ensure that what it can do stays ahead of what the users want i t to do
� �
?
�
learning how the plant should be operated together with the identifica tion ofchronicfililures leads to a reduction in the number and severity of failures, which leads to a reduction in the amount of money which must be spent on repairing them.
�
n: �
savi ngs in expenditure on parts and labour (usually contract labour) •
•
ReM helps to impr ove the moti vatio n of the peop le who are invo lved i n the revie w roce ss i n a num ber o f way s. First ly, a clear er unde rstan ding of the func tIOn s of th asse t and of wha t they mus t do to keep it work ing grea tly enha nces theIr com pete nce and henc e their conf idenc e . Seco ndly , a clea r unde rstan ding of the issue s whic h are beyo nd the cont rol of each indiv idua l in other words, of the limit s of wha t they can reas onab ly be expe cted to achi eve enab les them to work more com fort a ly with in th se limit s. (For insta nce, no long er are rnaintenance supe r nsors a tomat cally held responsib le for every failure, as so often happens m prac tIce. ThIS enab les them - and thos e abou t them to deal with fail ures ore calm ly and ratio nally than migh t othe rwis e be the case . ) Th rdly , the know ledg e that each grou p member played a part in for mula tmg goal s, i n deci ding wha t shou ld be done to achie ve them and i n deci d ng wh sh uld d o i t lead s a stron g sens e o f own ersh ip. ThIS com bma tlOfl of com pete nce, conf iden ce, com fort and own ersh i p aIl m an t at the peop le conc erne d are muc h more likel y t o wan t t o d o the . fight Job fight the first time .
Most of the items listed in the previous section of this chapter also i mprove maintenance cost-effec tiveness. How they do so is summarised below:
•
.
Greater motivation of individuals
Most of tile items listed under 'improved operating pelformance '.
•
�� � �
men tione d o n se ra o casi ons, ReM also help s users to enjoy (he . ma llnu m useful lIfe of mdl vldu al com pone nts by sele ctinn b on-c ondi tion mam tena nce m preference to othe r tech niqu es wherever poss ible.
The criteria used to decide whether a proactive task is technically feasible and worth doing apply directly to the acquisition of condition monitoring equipment. If these criteria are applied dispassionately to such acquisi tions, a number of expensive mistakes can be avoided.
•
3 15
•
A Maintenance Database The ReM Infoflnatio n and Deci sion Worksheets prov ide a numb er of addit ional bene fits, as follo ws:
316
•
Reliability-centred Maintenance
adapting to changing circumstances: the RCM database makes it pos
sible to track the reason for every maintenance task right back to the functions and the operating context of the asset A s a result, if any aspect of the operating context changes, it is easy to identify the tasks which are affected and to revise them accordingly , (Typical examples of such changes are new environmental regulations, changes in the operating cost structure which affect the evaluation of operational consequences, or the introduction of new process technology. ) Conversely it is equally
easy to identify the tasks which are not affected b y such changes, which means that time is not wasted reviewing these tasks. In the case of traditionally-derived maintenance systems, such changes often mean that the whole maintenance program has to be reviewed in its entirety. A s often as not, this i s seen as too big an undertaking, so the system as a whole gradually fall s into disuse. •
an audit trail: Part 3 of Chapter 5 mentioned that rather than prescrib ing specific tasks at specific frequencies, more and more modern safety
legislation is demanding that the users of physical assets must be able to produce documentary evidence that their maintenance programs are built on rational, defensible foundations. The RCM worksheets provide this evidence the audit trail in a coherent, logical and easily under stood form . •
more accurate drawings and manuals: the RCM process usually means that manuals and drawings are read in a completely new light. People start asking "what does it do?" instead of "what is it?". This leads
them to spot a surprising number of errors which may have gone un noticed in as-built drawings (especially process and instrumentation drawings). This happens most often if the operators and craftsmen who work with the machines are included in the review teams. •
reducing the effects of staff turnover: all organisations suffer when experienced people leave or retire and take their knowledge and experi ence with them. By recording this information in the RCM database, the organisation becomes much less vu lnerable to these changes.
For example, a major automotive manufacturer was faced with a situation w here a plant was to be relocated and most of t he workforce had chosen not to move with t he eqUipment to the new s ite. However, by using ReM to analyse the equipment before it was moved, the company was able to transfer much of t he knowledge and experience of the departing workers to the people who were recruited to operate and maintain t he equipment in its new location.
What RC'M A chieves •
31
the introdu�tion orexpert systems: the infor
mati on on the Information Worksheet In partI cular provides an exce llent foun datio n for an expert syste m. I fact, any users regard this work shee t as a simp le expe 11 . syste m In ts own nght , espe cially if the infor mati on is stored in a si mple comp uten sed database and sorted appr opria tely.
� �
n:
An Integrative Framework
�
As me ltion ed in Chap ter 1 , all of the issue s discu ssed abov e are part of the maIn"tream of main tenan ce man agem ent, and man y are already the : target mpro vem ent program s. A key feature ofRC M is that it provides . n ffect I e step- by-st ep fram ework for tackling all of them at once , and for Invo lvmg everyone who has anyth ing to do with the equip ment in the proces s .
� �
0:�
�
A Brief Histor)' (?lR CM
1 5 A Brief History of ReM
15.1 The Experience of The Airlines In 1 974, The United States Department of Defense commissioned United Airlines to prepare a report on the processes used by the civil aviation in dustry to prepare maintenance programs for aircraft. The resulting report was entitled Reliability-centered Maintenance. Before reviewing the application of RCM in other sectors, the follow ing paragraphs summarise the history of RCM up to the time of publica tion of the report by N owlan and Heap1978. The italicised paragraphs quote extracts from their repOli.
319
resp onsi�lefor regulating airline maintena nce practices, was frustrated b� expe �lences :howing that it was notposs ible to control thc /ililure rate oj certall1 unrelzable type s ofengines by anyfeasible changes in either the content orfrequency ofscheduled o verh auls. As a result, in f 960 a task force w�lsjormed, consisting of represen talivesfrom hoth the FAA and . l1es, to the azril investigate the capabilities ofpreventiv e maintenance. ��e ��ork of this group �ed t� the establishm ent of the FAA/Ind ustry . RelIabIhty Program, descnbed tn the introduction to the authorising doc . ume nt as jollo ws:
"T�l� development ofthis program is towa rds the control {
The Traditional Approach to Preventive Maintenance The traditional approach to scheduledmaintenanceprograms was based on the concept that every item on a piece of complex equipment has a 'right llge ' at which complete overhaul is necessary to ensure safety and operating reliability. Through the years, however, it was discovered that many types offailures could not be prevented or effectively reduced by such maintenance activities, no matter how intensively they were per formed. In response to this problem, airplane designers began to develop designfeatures that mitigatedfailure consequences - that is, they learned how to design airplanes that were :lailure tolerant '. Practices such as the replication ofsystemfunctions, the use ofmultiple engines and the design ofdamage tolerant structures greatly weakened the relationship between sqfety and reliability, although this relationship has not been eliminated altogether. Nevertheless, there was still a question concerning the relationship of preventive maintenance to reliahility. By the late 1 950's, the size of the commercial airlinefleet had grown to the point czt which there was ample datafor study, and the cost C?lmaintenance activities had become suffici ently high to warrant a searching look at the actual results of existing practices. At the same time the Federal Aviation Agency, which was
The History of ReM Ana lysis The next st�p ��s an attempt to organise what had been learnedFom the . vano us relzabllzty programs and develop a logic al and generallv appli cab�e approach to the design of preventive maintenance programs. A rudlmentary decision-diagram techn ique was devised in J<J65, and in June 1 967 a pape r on its use was presented at the AIAA CommercialA ir craft Design and Operations Meeting. Subs equent refinements this
3 20
A Brief History of RCM
Reliability-centred Maintenance
e evaluation and techn ique were embodied in a handbook on maintenanc steering group formed program developme nt, drafted by a maintenance the new Boemg 747 for am progr l to oversee development of the initia by special teams of used was -1, airplane. This document, known as MSG uled-mamtenance sched first industl}' and FAA personnel to develop the maintenance. The � entre( program based on the principles ofreliability-c ful. ss Boeing 747 maintenance program has been succe er improveme�ltS, furth to led ique techn m iagra ion-d Use of the decis d document, MSG -2: which were incorporated tyvo years later in a secon Plann ing Docu ment. Airlin e Manu facturer Main tenan ce Program
enance programs for MSG-2 was used to develop the scheduled maint These programs nes. airpla DCJO las the Lockheed 1011 and the Doug to tactical mili ed appli been also have also been successful. MSG-2 has the Lockheed as such ft aircra tary aircraft; thefirst applications werefor prepa red in ent docum r S-3 and P-3 and the McDonnell F41. A simila t as the Air f aircra Europe was the basisfor the initial programsfor such bus lndustrie A -300 and the Concorde. and MSG-2 �as to The objective ofthe techniques outlined in MSG- l assured the maxunum develop a scheduled-maintenance program that capable and also pro was ment equip safety and reliability of which the economic benefits the of le vi�led them at the lowest cost. As an examp e policies the enanc maint achieved with this approach, under traditional scheduled red requi ne initial programme for the Douglas DC-8 airpla C- J 0 pro D the items in overhaulfor 339 items, in contrast to seven such in aul limits the later gram. One af the items no longer subject to overh nation of schedu led . programs was the turbine prop ulsion engine. Elimi in labour and matenals overhauls for engines led to major reductions tory required to cover costs, and also reduced the spare-engine inven esfor larger airpl�nes engin Since shop main tenan ce by more than 50%. a respectable savmg. was this then cost more than US$l million each, for the Boeing 747, am As another example, under the MSG -1 progr major structural on United Airlines expended only 66 000 manh ours for the first hours 000 inspections before reaching a basic interval of 20 e ?oli enanc maint heavy inspections ofthis airplane. Under traditional n manh ours to arnv e at des it took an expenditure ofmore than 4 millio er and less complex the same structural inspection intervalfor the small of obvious impor are itude magn this Doug las DC-8 . Cost reductions of large fleets of aining maint tance to any organisation responsible for complex equipment. More important:
•
321
Such cost reductions are achieved with no decrease in reliability. On the contrary a better understanding ofthe failure process in complex equipment has actually improved reliability by making it possible to direct preventive tasks at specific evidence ofpotentialfailures.
A lthough the MSG- 1 and MSG-2 documents revolutionised the proce dures followed in deVeloping maintenance programs for transport air craft, their application to other types (�f equipment was limited by their brevity and :.;pecialised focus. In addition, the formulation of certain concepts was incomplete . For example, the decision logic began with an evaluation ofproposed tasks, rather than an evaluation ofthej(lilure con sequences that determine whether they are needed, and ijso, their actual purpose. The problem of establishing task intervals was not addressed, the role of hidden-function failures was unclear, and the treatment of structural maintenance was inadequate. There was also no guidance on the use of operating info rmation to refine or modify the initial program after the equipment entered service or the i!�formation systems neededfor effective management of the on-going program. All these shortcomings, as well as the need to clarilJ many of the underlying principles led to analytic procedures of broader scope and their crystallisation into the logical discipline now known as Reliability centred Maintenance. The Nowlan and Heap report provided the basis for MSG3 , which was promulgated in 1 980 and revised in 1 988 and 1 99 3 . MSG3 remains to this day the process used to develop and refirie maintenance program s for all maj or types of civil aircraft.
15.2
ReM in Other Sectors
Since 1 978, ReM has been applied extensively by the US Navy. In 1 984, three nuclear power undertakings in the U nited States commenced a series of pilot applications under the auspices ofthe Electric Power Research In stitute in San Diego. These projects led to the widespread adoption of ReM by nuclear facilities in North America, by EdF in France and now to an increasing extent by nuclear facilities throughout the world. The author and his associates began working with the application of ReM in the mining and manufacturing sectors in the early 1 980s. Sinee then, we have worked on more than 500 sites in 32 countries spanning all six continents.
322
Reliability-centred Maintenance
The projects concerned range from in-company awareness training for senior operations and maintenance managers to the full-scale application of ReM to all the equipment on a site. The sectors in which projects have been carried out in the fifteen years prior to the date of publication of this edition of this book include the following (sectors where we have worked on more than 5 sites are shown in italics); • aerospace
•
agrochemicals (fertiliser, pesticides) aluminium processing (mining, smelting, refining, forming) automobile manufacture (engines, assembly, components, tyres) base metal refining (nickel, aluminium, platinum)
•
banking
•
breweries and soft drinks
•
buildings and building services
• • •
• •
chemicals (industrial and household) cigarette manufacture
•
coalmining cosmetic manufacture
•
computer manufacture
•
•
electric utilities (coal, steam, gas and nuclear generation, distribution, transmission)
•
food numufacture (bakeries , canned fruit and vegetables, cereals, coffee, confectionery, frozen food, ice cream, margarine, meat products, milk)
•
foundries
A Brief History ofRCM •
•
postal services
pulp & paper (tissue paper, fine paper, photographic paper, newspr int) railways (passenger, freight, underground, i nfrastl1lcture, signalling) steelmaking (integrated mills, rolling m i ll s, stainles s steel)
•
•
323
water distribution woodworking. The fact that ReM has been enthusi asticall y receive d by people at all levels, and has enabled users to achieve some remarkable success es in aU these countrie s, suggest s that it is much less affected by cultural differ ences than many other particip ative techniq ues of this nature. •
•
The organis ations listed opposite all became aware of and started apply ing RCM at differen t times, so implementation is more advance d on some sites than others. The overall situatio n can be summarised as follows : • i n about 25% of the cases, senior managers have undergone prelimi n ary training •
•
about 1 0 % of the organis ations have applied ReM to all of the equip ment on at least one site the remaini ng 65% have reviewe d some of their equipm ent, and most plan to continu e using RCM to analyse most if not all of their assets.
Space does not permit a detailed consideration of the work done in each case. Howeve r, Chapter 1 4 provide s a general summary of the results achieve d to date, together with a brief review of some of the highligh ts.
gas distribution glass making
15.3 Why ReM 2?
•
h arbours
•
iron mining
•
mass housing
There are now a l arge number of interpret ations and variation s of the RCM deeision logic in existenc e. Some are more rigorous than others, but fou r account for the majority of the RCM work c lllTently being done on the planet. The first i s shown on pages 91 and 92 of the report by Now Ian and HeapI978. This is the version first used by the author, and it is also the
• •
•
metalworking microelectronics military undertakings (armies, navies and air forces)
•
nuclear facilities
• •
• •
• •
offshore oil and gas oil refining petrochemicals pharmaceuticals
•
photographic equipment
•
plastics (polymers, films, cellulose, fibreglass)
version originally used by most other ReM practition ers. The second version of the decision diagram is the official M S G3 version currently used by the civil aviation industry. It is shown as the 'SystemJPowerpla nt Logic Diagram ' on pages 4 and 5 of the Mainte
nance Program Developm ent Documen t published by the Air Transport Associati on of Americal 993• The third, which is similar to MSG3, is US Mil-Std- 2 1 7 3 used by the US Naval Air S ystems Comrnan d l 986. The fourth is RCM2, which is the subject of this book.
Copyrighted Material
324
ReliabiIiIY'{'{'lIlaed Mllinle/lWIi'f'
I n sel"""lf' bill Il
cal po,,'a Mililies, ellrri,'d 011/ 111'0 pi/ol (/pplic{l/;()/IS I)fRCM ill Ihe US
""c'lcar I"m'er ;"dllsll)'. Their inlere,'1 omse /mm " be/i<'f IhOl Ihis indlwry ...m aehie"i"g lulell"m{' lew)., "".III/elY wul relil/bilitv, bill "'(/.1 IIIl1ssi"e/), OI'<,mll/illlllillillg ils <'lIlfip"'<'III. As" res"II, IIU'ir IIIlIi/l IIIru5 1 "'lIS simply10
redltCf' m(lilllellallC<' COSIS, ralher Ihall l() imp"Ne rdiabil·
ily, aml lll<')' mmlii f l'lilhe RCM "wI.'e"" lIc,.'ordillgly, (So mllcII
,'0. illftwl,
111(1/ it be(lrS /illie rn,'mbltmee 10 lite wigimtl RCM process d<,scribed by Now/tm Imd Heap: il sholiid be IIwre Cl)fr<'Clll' dl'ScriiJed lI.' plwmed mllint<'1I<mce oplimiZalioll, or PMO. ralher l/tllll RCM.) Tlli,,- ,.,(xlified pmee,u "'IIJ"dol,ted "" lIn illdllslry'...ide basi., 11.1' the Amerinm III1e/ellr
power indllstry ill 1987, ",ul "llfilll1<m,,- o/II'is"f'prO/wlt \I'e,,' 1I/tI'lWlIrds ado!"nl h)' "ariOllS 01111'1' III1e/f'lIr Illili/ies, uther brwwhes o/Ihe tlnlric· il}' geliemliOlI Imd dislrilm/;oll illdllslf)'. lim/ PllriS o/Ihe uil indll;uy. AII"e slime timf', ,'uldill SIH'cilllisls ill Ihe/orlllllllllilm o/",,,imenwlI"f' .Ilfl1legie.� became imeralnl ill Ihe apI,/i"lIIion 0/ RCM ill imlllsiries ml,er thall lll'ia'it",. ForemO,,1 among Ihesf' "'af' Jolm MOllhra)' IIIII/Itis 1l"-S(J(·illlf'S. TI,is gmlll' illilillliy "'orkf'll ",illi RCM in lIIilling 111111III/mil·
/1l/"/IIri" g
illdllslries ill Somhi'm Afrim lIIuler Ihe menlorship of Sllln
No",lall. 1III
its originlll foell;- 01/ eqllil'ml'lll
slIfel)' IIIld
relillbilil),. For ('XIIIIIl'le, 1I1f')' 1I1ll'f' incorfXmllf'll "th'ironm"III,,/ issues il1lo /111' df'cisioll·mllki/lg prof.'ess. clarified Ihe WI'),S in ",II'-ch eqail'menl flllteliOlIS slIollld be d. /illt'l/, de,'dop�d 1111)'" fireei,... ",Ies /01' selec'lill8 ' mailllf'nllnCf' tasks 1IIIIIIask illlen'lIls, alld in,'of{XJfIIled IIU(mlilalil'e r"'k criluia direclly il1lo Ihf' selling 0/ /lI imff'-Jimlillg Illsk inlen·lIls. Tltf'ir ellhllliud rersioll 0 / RCM i.lllo", kllowll (U RCM2. The Nud
/or II SllIIu/(ml:
Ihe f990s
Sillu Ihe el/rly J 990 's, a grt'lll IIIml)' moff' orgalliZlllifms b"..e de,'doped varia/ioll,w/the RCM process. SOllie, slIcilllS Ihe US NIH'III A ir Commlllld ",ilh ils "Gllidf'iillf'S for lite NIH'1l1 A"illlioll ReliabililY Cmlf're,1 Mllilllf" 1I(llIff' Process (NA VA IR 00·25·403) " and tlte British Ro)'a/ Nav)" ",ilh its RCM·orif'lIIed Nal'lll Ellgillurillg Slalldard (NES45), hlll'e ff'ma;lIed Iruf'1O Ihe process origillally f'x{>OIIIU/f'd by Now/all lllld !lellp, lIow('\'('(, (IS
II,e ReM bam/wagoll h,1S .Ilarll'd mllill8_ a .,-I",le lIew e(}IiI','liOIl
rif
P"X'f'SSt's hilS emergedIh,l! ore calfed "RCM" by Iheir I'mpollelll.<, hili
Copyrighted Material
Copyrighted Material
A Brit'f Hi,'IOI)' lif RCM
325
111111 oflell bear lilllt' or I/O rfSt'mbl{lllct' t" tht' origillal "'t'lica/ml,'}), rnearciled, hilf"')' slru1'llI,,'d wul llwfOlIgl,I." 1''''''t'1l P"X''',,'X dt'"t'loped by NOII'/all Wil/ Heap. A,' a re,m/l, if {lit {"ga"i�mioll ,mid Ihm il 1I'(IIHed I,elp ill 1I_.illlf or Il'llrtlillg ho'" to liS.. RCM, il rOllld 1I0t I>t' surt' \\'11111 pfOass 1I',mid bt' uift"t'(I,
II/dud, ",hell II,e US N{wy ren'lIIly asked fur t'qlliPlllt'1I1 ,',,"durs to uSt' RCM ",I'ell buiiding a lit'", ,.hi{, daH, mil' us 1'01111""'.1' offt"t''' a pm cns closely relaled to Ihe 1970 MSG·l {,mee",._ ft dt'ft'lIdt'd ils 11ft'rillg b}' 1I0lillg thm it) , prtKe.n u,"ed (I dt'cisirm-Iogir diagram, Sill'" H.eM also
U.•es (I
decisioll-/ogie dillgram, the cmlll"lII)' IIrg"ed, ils {'fOaSS I\'IIS (III
RCM pfOass, Tltt' US Nal'y had 110 all5",'" 10 titiJ arg
",
711(' US Air fora ,'tlll"{ard lI'aS ClIllalled ill 1995,
Tht' US NUl'}, has bUll ",,,,ble 10 im'akt' ils xumd"rds ""d ,.pel'ificati""" I\'itll eqllipmenl "t'mlo", (Iltollgh il 1'""lillll<'.< 10 liSt' 11tt'1II for ils in/ema/ work) -
ill tile illlilistrilli I\'orid. Durillg Ihe 1990s, magllzill..s alit! 1'11llferell(,(,_1 dt'l'oll'd to rqlliplllelll mailllt'lIl111Ct' "1"'" IIIlIltiplied, allll 1IIag{/�ille arti cles alld cmifert'IICt' p"f't'rs aOOIll RCM berallle 1II0re IIlII/lIIore 1II1 1IIt" "liS, Thes(' hili'" slw"'lI lltat "t'I)' dijJeft'1II processes lire beillg gil'ell the Slllllt' lIallli', "RCM ", So l}(Jth II,e US milila')' m,d ('(,mmarill/ illl/llst,)' saw " IIet'd /(J defilll' wh(ll {III RCM pruress is,
{II his (994 lIIt'moflllldlllll, Pal)' said, '" eIICOllfllg(' llit' Ulllia Sffretaf)' of Dt'feliSt' (IIrQ11isiIiOIl (/lid Tt'rII/w/og\') 10 fom, pam't'fships lI'i III ilidustf)' associatiOlIS to del'elop/wlI-gol't'fllmt'lll sl""dords for repit,u",em of lIIi/itol)' s/(llldards wltere pmetieable. " {lldud, lite Tedlllii'lli Srlll,dards 8m"d of the SAE lias I",d" 100'g {md clost' rt'latim'sitip wilh lit" ,<"mdards eo"'m,mit}' ill the US mililll')', allli lIos beel! h'orkillgfor $f"eral years 10 heJp del'eJdp c"III",ercial sUllldards 10 replact' lIIilitill)' S/{I11(/urds alld specijieatiolls, II'ltell lluded alii/ "'ltt'll IIOIlt' "Irt'ad)' aislei/,
Copyrighted Material
Reliability-centred Maintenance
326
Appendix 1: Asset Hierarchies and Functional Block Diagrams
Plant registers and asset hierarchies Most organisations own, or at least use, hundreds if not thousands of physical assets. These assets range in size from small pumps to steel rolling mills, aircraft carriers or office blocks. They may be concentrated on one small site or spread over thousands of square kilometres. Some of these assets will be mobile, others will be fixed. Before any organisation can apply RCM
a process used to determine
what must be done to ensure that any physical asset continues to do what ever its users want it to do - it must know what these assets are and where they are. In all but the smallest and simplest facilities, this means that a list of all the plant, equipment and buildings owned or used by the organisa tion, and which require maintenance of any sort, must be prepared. This list is known as the plant ...x: (f)
(]) ::l (f)
c:--.£9 o
.<:Q
E
ctl Cl.. (])
:S
o o
(f) ctl (f) ctl ""0 (])
(]) ""0
"11 ::l 0
(ij -C (]) 0
:;:
:.=
CC1
c:
:;:
.2 .,g .�
Q
is
4T � .2 '(0
0> c: "0
..c: � Q (])
E 1ii :;:
Cl.. (f) (])$ 0> ctl
�
(f)
Table IS.I:
u...
";::
(f) 0> c:
. ;::
::l 0 � (f) Q Cl.. c: c: Cl.. 0 UJ ctl Q (f) (]) Q c: (]) ::l
c:
(])
yet to be analysed and those that are not going to be analysed. (The plant register is also needed for other aspects of maintenance management, such as the planning and scheduling of routine and non-routine maintenance
(])
0> c:
1B
keep track of the assets that have been analysed using RCM, those that have
.
0 .2: ::l "E 0"" (]) � (]) .c (f) (]) > (f) � ::l ::l
..2:"-
� S �.� g>� (]) . 1:: -1":: -g 0 @ u...""oQ Q<;::::;:
.2 ctl "a1 05
(f)
$
(5��
o.. (J)
� -5 g"� 0 -0 c: ""0 (]) -g:g:J Io- ......
"0 Z
::l
'0 (]) :: .c (]) .sQ
'-
c: ctl ..c:
�
�
(f) (])
;:: 0 :S ':;1: 0> -"" * Q (f) Q (]) a:; .£9 "E (f) (])
ctl O (])g'(ij c: 0 .- Q (f) .""0 :2 'c
'0
:S ':;1: -
� -
.
"5
0 (]) :S
ctl ..a-'ia
vr£ .� � 55 >
�
(])
0
(f)
�
�-o ::l c:
(]) 0""
» (f)
(]) c: ctl 0 WQ
_
registe r.
The register should be designed in a way which makes it possible to
tasks, history recording and maintenance cost allocation. As a result, it ""0
�
(]) ""0
'(i)
c: 0 Q
(5 Z
"E
(])
E c:
2 '5
c: (])
(]) ..c: f-
should be set up and the associated numbering systems designed in such a way that it can be used for all these purposes.) Chapter 4 explained that ReM can be applied at almost any level in a hierarchy. It also suggested that the most appropriate level is the level which leads to a reasonably manageable number of failure modes per function. 'Appropriate' levels become much easier to identify if the plant register is set up as a hierarchy which makes it possible to identify any system or any asset at any level of detail, down to and including individual compo nents ( 'line replaceable items' ) or even spare parts. The truck on page 85 provides one example of such a hierarchy. Figure A 1.1 overleaf shows another example covering a boiler house in a food
Comparing Nowlan & Heap, MSG3 and RCM 2
factory.
328
Appendix 1: Asset and FUllctional Hierarchies
Reliability-centered Maintenance
329
:Functional hierarchies and functional block diagrams It is possible to develop a hierarchy showing the primary functions of each of the assets in the asset hierarchy. Figure A 1 .3 overleaf shows how this might be done for the asset hierarchy shown in Figure A 1. 1 . Variations of the functional hierarchy in Figure A 1 .3 are used t o show the relationships between functions at the same level. These arc usually known as 'functional block diagrams' , and they can be used to depict the relationships in a number of different ways. For instance, SmithlYG3 defines a functional block diagram as 'a top level representation of the major functions that the system performs'. On the other hand, Blanchard & Fabryckyl990, who prefer the term 'functional t10w diagram', that these diagrams can be prepared at many different levels. Smith tends to use the diagrams to show the movement of materials. energy and control signals through and between different elements of a system, whereas Blan chard & Fabrycky use them to depict the movement of single asset through different mission phases (such as an aircraft moving from start-up to taxi,
Figure AI.] Asset Hierarchy
A list drawn up for the assets in this hierarchy, together with a hierarchical numbering system for each asset, might appear as shown in Figure A l .2.
Figure A 1.2 Plant Register and Hierarchical Numbering System
Number 01 02 03 0301 0302 0303 030301 030302 030303 030304 03030401 03030402 0303040201 0303040202 0303040203 030304020301 030304020302 030304020303 0303040204 03030403 03030404 0304 04 05
Asset FoodColnc Factory 1 Factory 2 Factory 3 Preparation Dept Packing Dept Site Services Power Supply Compressed Air System Water System Boiler House Coal Handling System Boiler No 1 Shell & Tubes Grate Feed Pump Pump Motor Switchgear FD Fan Boiler No 2 Ash Handling System
Maintenance Dept Distribution Head Office
take-off, climb, cruise, descent, landing and so on.) A functional block diagram for the boiler house in Figure A1 .1 shows that the coal flows from the coal handling plant to the two boilers, and the residue to the ash handling plant. It also shows what materials and services flow across the system boundaries. This is illustrated in Figure A1.4 on page 332, which goes on to show a more detailed functional block diagram for one of the boilers. A more complex version of these diagrams could also be used to show what control and indication signals pass across the system boundaries.
Functional hierarchies and functional block diagrams are an essential part of the equipment design process, because design starts with a list of desired functions and designers have to specify an entity (asset or system) which is capable of fulfilling each functional requirement. As mentioned in Chapter 2, functional block diagrams can also be of some help when RCM is applied to facilities where the processes or the relationships between them are not intuitively obviolls. These tend to be large, poorly accessible, very complex, monolithic stmctures such as naval vessels, combat aircraft and the less accessible parts of nuclear facilities. However, in most other industrial applications (such as thermal power stations, food and automotive manufacturing plants, offshore oil plat forms, petrochemical and pharmaceutical plants and vehicle t1eets), there is usually no need to draw up functional block diagrams before embarking on an RCM project, for the following reasons:
330
Reliability-centered Maintenance
Appendix 1: Asset and Functional Hierarchies
WHATIT IS ...
WHAT IT DOES ...
Figure AI.3
Figure AI.3 (continued)
Asset Hierarchy .....
..... with Corresponding Functional Hierarchy
331
Reliabili(v-centered Maintenance
332
Appendix 1.' Asset and Functional H ierarchies
f[EVEL.. . . 4. :Boiier-House]
In cases of uncertainty, equipment is usually accessible enough for it to be
I�,�.�,." .� . -� ..�-��.-"-�.AC Power
AC Power Steam
333
easy to go and see what goes where, and if not, the required information
Flue Gas
can be extracted from a set of process and instnlll1entation drawings. In
{-
fact, a good set ofP&ID's nearly always climinates the need for functional
Coal-
block diagrams entirely as a precursor to the application of RCM. In such cases, block diagrams add significantly to the time, effort and cost of the RCM process while adding nothing to its value.
Water •
functional block diagrams only identify primary functions at each level, so they only tell part of the story. (For instance, nearly all the assets at the fourth level and below in Figure A l . I have a secondary containment function. This cannot be shown in a functional block diagram without
,'Treatment Chemicals water
! ..
AC Power:
;}
To supply f water to drum
making it unmanageably cumbersome.)
Dirty Water· , •
as explained in Part 3 of Chapter 2, the principal functions of the assets in the hierarchy above the level chosen for the analysis should be sum
_
marised in suitably worded operating context statements. These state
_
ments are written
-.--�-
Air
only for those assets which are relevant to the analysis
in question. As a result, time is not wasted defining the functions of
..;-. Steam
assets which are not germane to the asset under consideration. (If large numbers of assets are analysed, these high-level context statements evolve into a de facto functional hierarchy for the entire organisation
one that
is far more detailed than a cmde, single-statement-pcr-asset diagram.) AC Power
Steam
•
assets at or below the level chosen for the analysis are dealt with as part of the normal RCM process. Part 7 .of Chapter 4 showed that the func tions of lower level assets are either listed as secondary functions in the
Figure A 1.4:
Functional Block Diagrams
main analysis, or dealt with as failure modes, or in the case of exception ally complex subsystems, broken out for separate analysis.
•
in most industries, the relationships between different processes is usu ally well enough understood by the participants in RCM review groups to make these diagrams unnecessary. For example, the boiler house operators and maintainers would be fully aware of the fact that coal, water and air go into the boiler at one end and that steam, flue gas and ash (and occasionally dirty water) come out of the other. Most of
For instance, the example of the truck shown in Figure 4.11 in Chapter 4 showed how a blockage in the fuel line could simply be treated as a failure mode of either the engine or the drive system, without needing a separate function statement for the fuel system or the fuel line.
(In the author's experience, functional block diagrams tend to he of most value to outsiders seeking to apply ReM on behalf of equipment users. Be
them would probably regard the notion that these simple facts should be drawn
cause they are outsiders, they need these diagrams
in a diagram at best as a waste of time. As discussed at length in Chapter 2, the
expense of the owners of the assets
usually prepared at the
to improve their own understanding
real challenge is not to identify the simple and usually obvious relationships be
of the processes which they are about to analyse. The best way to avoid
tween processes, but to define the desired performance relative to the initial
this expense is not to employ outsiders as analysts in the first place, but
capability for all the key elements of each system, and then to define what must be done to ensure that the system continues to deliver the desired performance.
rather to train people who have a reasonable first-hand working knowledge of the plant as RCM facilitators.)
334
Reliability-centered Maintenance
System boundaries When applying RCM to any asset or system, it is of course important to define clearly where the 'system' to be analysed begins and where it ends.
Appendix 2: Human Error
If a comprehensive asset hierarchy has been drawn up and a decision taken to analyse a particular asset at a particular level, then the 'system' usually automatically encompasses all the assets below that system in the asset hierarchy. The only exceptions are subsystems which are judged to
Chapter 4 mentioned that a great many equipment failures are caused by
be so insignificant that they will not be analysed at all, or very complex
'human error' . It went on to mention that if a specific human error is con
subsystems which are set aside for separate analysis.
sidered to be a credible reason why a functional failure could occur, then
Care is needed with control loops which consist of a sensor in one sys
that error should be included in the FMEA. However, human error is an
tem which sends a signal to a processor in a second system, which in tum
enormous subject in its own right. The purposc of this appendix is to pro
activates an actuator in a third. Chapter 4 explained that this issue can often
vide a brief summary of the mqjor categories of human cow, and to suggest
be dealt with either by conducting the analysis at a high enough level to en
how they might be dealt within the framework of RCM.
sure that the 'system' encompasses the entire loop, or by analysing control systems separately (after the controlled systems have been analysed). How ever, sometimes this is not practical, in which case a decision must be made as to which system will encompass the control loop in its entirety. Care is also needed to ensure that assets or components right on the boundaries do not 'fall between the cracks'. This applies especially to items like valves and f1anges. It is wise not to be too rigid about boundary definitions, because as understanding grows during the RCM process, perceptions about what should or should not be incorporated in the analysis frequently change. This means that boundaries may need to be extended to incorporate some subsystems, others may be dropped and yet others which are included initially may be set aside for later analysis. (Again, the strongest exponents of rigid boundary definitions tend to be external contractors seeking to apply RCM on behalf of end-users, because system boundaries must be defined precisely in order to define the com mercial scope of the contracts. The fact that the analysis is the subject of a formal contract means that boundaries have to be defined much more pre cisely -- and much more rigidly - than is necessary from a purely technical point of view. Contracts of this type then either have to be renegotiated every time a boundary needs to be changed, or the boundary is not moved when it should be, resulting in a suboptimal analysis. The best way to avoid the time and cost associated with these commercial manoeuvrings is to avoid contracting out this aspect of maintenance policy formulation altogether.)
Principal Categories of Human Error When considering the interaction between people and machincs, Blan chard et aP995 group the main factors under four headings: •
anthropometric factors
•
human sensory factors
•
physiological factors
•
psychological factors.
Nearly every 'human error' can be traced to a failure or a problem which has occurred in at least one of these four areas. As a result, we review them brief1y in the first part of this appendix, before looking in more detail at the fourth category.
Anthropometric factors Anthropometric factors are those which relate to the size and/or strength of the operator or maintainer. Errors occur because a person (or part of a person, such as a hand or arm): - simply cannot fit into the space available to do something - cannot reach something - is not strong enough to lift or move something If a failure is occurring or is reasonably likely to occur for any of these reasons, it is highly unlikely that a proactive maintenance task will be found to deal with it. Note also that if a human error occurs for onc of these reasons, the human error is not the root cause. The failure mode is actually poor design and the resulting human error is a failure effect.
336
Appendix 2: Human Error
Reliability-centered Maintenance
lf the consequences are such that something must be done about a fail ure which is occurring for anthropometric reasons, the only viable course of action is likely to be redesign. This will nearly always involve recon figuring the asset in such a way that it becomes more accessible or easier to move. In this context, Figure A2. 1 shows some dimensions which are considered by the US Navy to be adequate for reasonable human access in confined spaces.
337
Human sensory factors Human sensory factors concern the ease with which people can see, hear, feel and even smell what is going on around them. In the case of operators, this tends to apply to the visibility and legibility of instruments and con trol consoles. For maintainers, it relatcs to the visibility of components in the nooks and crannies of complex systems. The volume and variability of background noise levels also affect the ability of both operators and maintainers to discern what is happening to their equipment.
6f
Note again that if errors are occurring or thought to be likely for these reasons, the human error is not the root cause, but is the effect of some other failure. The remedies also usually entail redesigning the asset (making things easier to see, reducing noise levels).
1---�-188cm
�-73cm-1
-I
Physiological factors The term 'physiological factors' refers to environmental stresses which
1T85
190 em
� '--/33cml-
J�I
affect human periormance. The stresses include high or low temperatures, loud or irritating noises, excessive humidity, high vibration, exposure to toxic chemicals or radiation, or simply working for too long at a physically or mentally demanding task
especially
without an adequate break,
Sustained exposure to these stresses leads to reduced sensory capacity, slower motor responses and reduced mental alertness. These are all mani festations of (human) fatigue, and all greatly increase the chances that the people concerned will make a slip, lapse or mistake. (These three tetl1)s are defined in the next section of this appendix.) If errors occur or are thought to be likely for any of these reasons, the human is once again not the root cause, but the error is the effect of some
T
130
other failure. Again, if the consequences wan-ant it, the remedy is likely to be some form of once-off change. Either the design of the physical en 1'-'---1"
vironment can be changed in such a way that the error-inducing stresses are reduced (for instance, by reducing temperatures or by providing hearing protection), or operating procedures could be changed in a way that overstressed people a chance to recover (longer, more frequent
1- 90cm�.j
or
more
carefully timed rest breaks). Another environmental stress factor is a relentlessly hostile or adver sarial organisational climate. While this does not necessarily have a physio
Figure A2.1: Where people fit (From NAVSHIPS 94234, Maintainability Design Criteria Handbook for Designers of Shipboard Electronic Equipment. US Navy, Washington DC)
logical effect, it can lead to an increased predisposition towards psycho logical errors. In many cases, it boils down to excessive and inappropriate use of high task/low relationship leadership styles. Unfortunately, there is not much that ReM can do about this problem.
338
Appendix 2: Human Error
Reliability-centered Maintenance
However, what RCM can do is alleviate -- if not eliminate - the hostile relationship which so often exists between maintenance and operations
•
•
339
unintended errors are subdivided into slips and lapses intended errors are subdivided into mistakes and violations.
people, as explained on page 268. This makes people less inclined to blame
These categories are illustrated in Figure A2.2 and are discllssed brietly
each other for errors, and more inclined to find solutions.
in the following paragraphs.
Psychological factors
Slips and lapses
The three sets of factors discussed so far all relate to external phenomena which cause the human to make an error. As a result, they are relatively
somebody who is fully qualified to do a job
easy to identify and to deal with (although doing so may sometimes be expensive). A far more complex and challenging category of errors are
Slips and lapses are also known as skill-based errors. They OCCllr when it eorrectly many times in the past
,md who may even have done
does the job incorrectly.
Slips occur
when somebody does something incorrectly (for instance, if an electrician
those which find their roots in the psyches of the humans themselves. As a result, these psychological factors are discussed in more detail in the
when someone misses out a key step in a sequence of activities (for in
next section of this appendix.
stance, if a mechanic leaves a tool behind after working in a machine or
wires a motor incOlTectly, causing it to run backwards). Lapses occur
forgets to fit a key component while reassembling it.)
Psychological errors
These errors usually happen because the person concerned was distracted,
Reason 1991 divides the psychological categories of human error into those which are unintended and those which are intended. An unintended error is one which occurs when someone does a task which he or she should be doing, but does it incorrectly ("does the job wrong"). An intended error occurs when someone deliberately sets out to do something, but what they do is inappropriate ("does the wrongjob"). Reason divides these two catego ries further as follows:
SLIP
!.APSE
AttentiQnalf�i'�re13 Cartyol.lt a planned task incorrectlyor in the wrang sequence Memoryfall�r�� Miss outEtstepll1a
planned seCiYf1I)ce
alevents .
preoccupied or simply 'absent minded'. As a result, they are lIsually un predictable, although their likelihood increases if the person is working in a physically hostile environment, or if the task is exceedingly complcx. However, if the environment is reasonably benign and the task is fairly simple, then this category of human errors is perhaps the only one where it is fair to describe the error as the root cause of failure. The possibility of a great many slips and lapses can be reduced if opera tors and maintainers are involved directly in the RCM process (especially the FMEA) , This leaves them with a much broader and deeper understanding of the effects and consequences of their actions, which in turn results in greater motivation to do the job 'right first time'. This applies especially to tasks where the consequences of failure are likely to be most sellous. Another approach to slips which occur during assembly is based on the assumption that if something can be installed incorrectly, it will be. The remedy is to go back to the drawing board and: •
redesign systems in such a way that they can only be assembled in the correct sequence
•
redesign individual components in such a way that they can only be installed the right way round and in the right place.
This is the essence of the Japanese concept of poka yoke ('mistake proof
Figure A2.2: Categories of psychological errors
"-------I
VIOLATION
Routine violations Exceptional violations Acts of sabotage
ing'). Ideally this philosophy should be applied to original designs rather than retrofitted to existing assets, because it is usually cheaper to build in good practice initially than to modify out bad practice later.
340
Reliability-centered Maintenance
Appendix 2: Human Error
The chances of bad habits developing are also reduced if care is taken during the FMEA to identify failure modes which give rise to spurious alarms, and to take steps subsequently to reduce them to a minimu m. (In cases where the frequency and/or the possible consequences of a false
Mistakes I: Rule-based mistakes Rule-based mistakes occur when people believes that they are following the correct course of action when doing a task (in other words, applying a 'rule'), but in fact the course of action is inappropriate. Rule-based mistakes are further subdivided into
misapplication of a good rule
34 1
alarm warrant it, the most appropriate remedy usually entails redesign.) ReM helps to reduce the possibility ofapp/ying bad rules because the whole ReM process is all about defining the most appropriate' rules' for maintaining any asset.
and
application of a bad rule. In the first case, under a given set of conditions, a person selects a course of action which seems appropriate, usually because it has been successful
Of course, care must be take to ensure that the rules of ReM itself are not applied badly. This is best done by ensuring that everyone involved in the application of ReM is adequately trained in the underlying principle s.
in dealing with similar conditions in the past- hence the term 'good rule' . However, some subtle variations on this occasion mean that the course of action, undertaken deliberately, is wrong. For instance, a protected system might be set up in such a way that excessive pressure should cause an alarm to sound and a warning light to illuminate. How
Mistakes 2: Knowledge-based mistakes
ever, a situation might arise where the alarm is failed, the pressure increases and
Knowledge-based mistakes occur when someone is confronted with a situation which has not occurred before and which has not been anticipa
the light comes on. The absence of the alarm may lead the operator to believe that the warning light on its own is only a false alarm, especially if it has a history of spurious failures. In this case, the operator may choose to take no action until the light is repaired
a course of action which has been appropriate in the past. On
this occasion however, it is not the right thing to do.
The application of a bad rule means just what it says. The normally chosen A classic example of a bad rule is a maintenance program which schedules items for fixed interval overhauls in order to deal with failure modes which conform to failure pattern E or F (see Figure 1.5 or 12.1). In the case of Pattern F especially, an action deSigned to improve reliability will in fact make it worse, by upsetting a stable system and inducing infant mortality.
In these cases, the 'root cause' of the failure is the rule itself or the process by which it is selected. If the rule is promulgated or selected by someone other than the person who performs the task -in other words, if the person then the mistake is really the effect
of another failure. The ReM process helps to reduce the possibility of misapplying
good
rules in two ways: •
the thorough analysis of failure effects, especially what could happen if a hidden function is in a failed state when it is needed, means that people are less likely to jump to inappropriate conclusions when the situation does arise (especially if they have been involved in the ReM process)
•
by focusing attention on the functions and maintenance of protective devices, the ReM process greatly reduces the probability that these devices will be in a failed state in the first place.
and a mistake occurs if this decision is wrong. In practice, the author has found that a common problem which occurs
in this context is a belief on the part of senior managers and engineers
or prescribed course of action is just plain wrong.
doing the task is only following orders
ted (in other words, one for which there are no 'rules'). In situations like this, the person has to make a decision about an appropriate course of action,
that "I know, therefore my company knows". In fact, if a crisis occurs late at night when all the senior people are off-site, the requisite knowledge is lIse
less if it is not in the mind of the person who has to take the first steps to deal with the crisis. This suggests that first and most obvious way to avoid knowledge-based mistakes is to improve the knowledge of the people who have to make the
decisions. In most cases, these are the operators and mainlainers. Opera tors and maintainers are likely to make appropriate decisions more often
if they clearly understand how the system works (its functions), what can go wrong (functional failures and failure modes), and the symptoms of each
failure (failure effects). As mentioned several times in chapters 2, 13 and 14, this understanding is hugely enhanced if operators and maintainers are involved directly in the ReM process. The most important findings can be
disseminated subsequently to people who do not pmticipate in the analysis by incorporating the findings into training programs. If necessary, the possibility of knowledge-based mistakes can also be
reduced by designing (or redesigning) systems in ways which: •
minimise complexity, so that there is less to know
342
Reliability-centered Maintenance
minimise novelty, because new and alien technologies put people at the bottom of the learning curve, where mistakes are most likely to happen
•
•
avoid tight coupling. This means designing systems in such a way that if failures do occur, consequences develop slowly enough to give people time to think and hence more opportunity to make the right decisions.
Chapter 5 suggested that it might be possible to produce a schedule of
Vio lations
A violation occurs when someone knowingly and deliberately commits an error. Violations fall into three categories: •
routine v iolations. For instance, when people make a habit of not wear
ing items of protective clothing (such as hard hats) despite rules which clearly state that they should •
•
For instance, if someone who usually wears a hard hat knowingly rushes outside without the hat on "because they couldn't find it and didn't have time to look for it"
exceptiona l v iolations.
sabotage.
Appendix 3: A Continuum of Risk
This occurs when someone maliciously causes a failure.
The remedy for routine and exceptional violations usually consists of ap propriate enforcement of the rules by management. However, once again,
involvement in the RCM process gives people a clearer understanding of the need for safety procedures and the risks they are running if they violate
them. The management of sabotage is beyond the scope of this book.
acceptable risks which combines safety risks and economic risks in one continuum. It suggested that this might be made possible by combining Figures 5.2 and 5. 14 in some way. Figure 5.1 4, repeated as Figure A3. 1, showed what an organi
prepared to tolerate in a
The most important conclusions to emerge from this appendix are that: •
•
not all human errors are necessarily the fault of the person who made
Complete control, full choice Some control, some choice
event which could
prove fatal in that situation, as summarised in Figure A3.1. Acceptability of fatal risk
(at work)
No control, some choice
(in
No control,no choice
(off-site)
1'\
"\ 1'\
I'"
10 4 10-5 10-0 10-7 10s
the elTor. In many cases, the elTOl' is either forced by external circum
In fact, these two charts cannot be combined as they stand, because Figure
stances or by inappropriate lUles. So if blame is to be allocated for any
A3.1 is based on the probability of a single event while Figure A3.2 de
error, care must be takcn to idcntify the real source
picts what one individual might consider to be tolerable for any event.
human elTor is at least as common a reason why equipment fails to do what its users want it to do as deterioration, if not more so. As a result,
However, with respect to thc latter, part 3 of Chapter 5 went on to show
it should be dealt with as part of the RCM process, either as afailure mode when it is a root cause, or as afailure efrect when it consists of inappropriate responses to other failures •
"""
10' 102 10' 104 10-\ 10'
(in my car or home worksliop)
specific situation from
Figure A3.2:
""
+
1
Acceptability of economic risk
one individual might be
l� -- �
£1 000 000 £10 000 000
Figure 5.2 depicted what
""
£100 000
economic consequences only.
----.--
�--
£10 000
accept for one event that has
Figure A3.1:
""
£1000
sation might decide that it can
any Conclusion
Trivial upto £100
in the industrial context, it is only possible to come to grips with human errors if the people involved in committing the elTors are involved . directly in identifying them, and developing appropriate solutions
that it is possible to use what one individual tolerates from any event in a given situation as a basis for deciding what probabilities apply to each event which could place him or her at risk in that situation, as follows: The first step is to convert what one person tolerates to an overall figure for an entire site. In other words, if I tolerate a probability of 1 in 100 000 (10-5) of being killed at work in any one year and I have 1 000 co-workers who all share the same view, then we all accept that on average 1 person per year on our site will be killed at work every 100 years - and that person may be me, and it may happen this year.
344
Reliability-centred Maintenance
Appendix 3: A Continuum of Risk
The next step was to translate the probability which myself and my co workers are prepared to tolerate that any one of us might be killed by any event at work into a tolerable probability for
each single event
(failure
mode or multiple failure) which could kill someone. For example, continuing the logic of the previous example, the probability that any
345
A detailed examination of fault trees is beyond the scope of this book. The purpose of this appendix is only to suggest how it might be possible
to convert risks whieh individual members of society might be prepared to tolerate (another manifestation of 'desired performance') into mean ingful information which can be used to establish a maintenance proaram b
one of my 1 000 co-workers will be killed in any one year is 1 in 100 (assuming
designed to deliver that performance.
that everyone on the site faces roughly the same hazards). Furthermore, if the
The process described above can be used to produce a graph showinn the probabilities of a single fatal event at work which would flow from th
activities carried out on the site embody
�
(say) 10 000 events which could kill some one, then the average probability that each event could kill one person must be reduced to 106. This means that the probability of an event which is likely to kill ten people must be reduced to 10.7, while the probability of an
Area 2x103
10'
10.3
1-+-1
103
5 5x10"
I
I
Line 1 Line 2 Line 3 Line 4 Line 5 2x10' 4x10"
event that has a 1 in 10 chance of killing one
100 ...... (Could kill 10) (Could kill 1) 1(JB 101
person must be reduced to 10.5 On a site that is divided into several areas and where each area is further divided into several sec tions, this process of subdividing acceptable
Figure A3.3:
risk could be carried out in stages, as shown
From whole site to one event
in Figure A3.3.
a
single failure mode
(as defined in the FMEA) which on its own has
lethal consequences. The probability allocated to this type of event de fines the 'tolerable level' which is referred to when the RCM process asks the question "Does this task reduce the probability of the failure to an acceptable level?". See page 102. •
a
multiple failure
more accurately, the annual failure rate.)
1in 10chance ofkilling 1employee
Figure A3A: Acceptability of one lethal event where I have some control
In the example shown, an 'event' is either: •
risks which one individual is prepared to accept, on the assumption that his. or her judgement is accepted by everyone else on the site. This is illustrated in Figure A3.4. Note that in the next four graphs, the X-axis represents the probability of any one event occUlTing in any one year, (or
where a protected system fails and the protective
device which should have rendered the system non-lethal is itself in a failed state. The probability allocated to this type of event defines the 'acceptable level' which is referred to when the RCM process asks "Does this task reduce the probability of the multiple failure to a toler able level?" See page 122. It is also the probability used to establish
MMF when setting failure finding intervals. See page 179.
and some choice
� �
Likely to kill 100employee s Likely to kill 1 000employee s
"'-
�
or people visiting large buildings (shops, offices, sports stadiums, theatres,
and so on). In general, these people could be called 'customers'. In this case, if they all tolerate the same risk as the individual in Figure
A3.2 (and there are the same number of potentially life-threatening events inherent in the system), the process of apportioning risk used in Figure
A3.3 could lead to the single-event probabilities shown in Figure A3.5.
10G 101 108 10.8 10,010"
!- -
1in 10chance o f killing 1cu stomer
Figure A3.5: Acceptability of one
jectives for each safety-oriented proactive task and to determine failure
lethal event where
finding task intervals, rather than upwards to determine a top-event pro
I have no control
bability based on an existing maintenance program.
"'1"'-
Likely to kill 10employee s
nsk. Th example in Figure A3.2 suggests that an airline passenger might be a tYPIcal example of someone in this situation. From the maintenance viewpoint, such people are likely to be users of mass transport systems.
lysis would be used to allocate probabilities (see Andrew and MOSS1993). (the probability of a fatal accident anywhere on the site) to establish ob
-
I",
'I he same process could be applied to the situation in which the likely . lCtImS have no control but some choice about exposing themselves to the
In complex systems, it is likely that an approach similar to a fault tree ana However, in this case, we work downwards from a top-event probability
Likely to kill 1employee
10.5 10� 10' 10B 10" 10,0
and some choice
Likely 10kill 1cu stomer
'f'"
r""
Likely to kill 10cu stomer s Likely 10kill 1 00cu stomer s Likely to kill 1 000cu stomer s
-
""
I""
-'-
346
Reliability-centred Maintenance
Appendix 3: A Continuum olRisk
Similar reasoning applied to the no control/no choice scenario might yield the single event probabilities shown in Figure A3.6. (In practice, most individuals are likely to accept an even lower probability of being killed for this reason than is shown in Figure A3.2
the so-called 'dread'
factor. However, in most facilities, fewer events would be likely to have off-site consequences, so the probability for each event might end up about the same.) 10-7 108 10-9 10-10 10-11 10-12 1 in 10 chance of killing 1 person o ff site
Figure A3.6:
Likely to kill 1 person o f ·f site
Acceptability of one
Likely to kill 10 people off·site
lethal event where
I"::
'"
Likely to kill 1 00people off·site
I have no control and no choice
Likely to kill 1 000people of f-site
'" '"
i'
Once acceptable probabilities have been determined for single events as shown in Figures A3.1, A3.4, A3.5 and A3.6, it is of course possible to combine them into a single 'continuum of risk' , as shown in Figure A3.7. 1 10' 10.2 103 104 105 106 107 108 10-9
Figure A3.7:
Trivial -Pl.-+----+ --j- ---+---+---t
A "continuum of risk" Upto £100
-I-l----"!.----+-+--+--t----1--t--1-+
£ 1 0 000 -1-+-+ - -+-�'k-+-t-+---I--t--£100 000 4-+---+-+---'\.: £ 1m or 1 in 1 0chance 4-+---+-+----t--rof killing 1 employee >£ 10m orlikely to kill 1 employee 4_+---+_+ or 1 in 10 chance of killing 1 customer Likely to kill 1 0employees or 1 customer or 1 in 10 chance of killing 1 person off-site Likely to kill 1 00employees or 10 customers or 1 person off·site Likely to kill 1 000employees or 100 customers or 10 people off-site
_+-__-1------'
-I-+-+--+-+---+-+---!-+--!--t ++
+--+
__
+-_-+_+---1_+---"1..-+
__
..I....j-+--l-+--r-+-r--r1� 1�1�1�1�1�1�1�1�
Please note once again that these figures are not meant to be prescriptive and do not necessarily reflect the views of the author or any other organi sation or individual as to what should or should not be acceptable.
347
Figure A3.7 is also not intended to imply that I employ ee is worth £10 million. That figure represents a point at which two different value systems happen to coincide. The financial risks which YOllror ganisation is willing to accept and the personal risks which your employ ees and customers (and society as a whole in the case of no control/no choice hazards) are prepared to accept may lead to a completely different set of figures in your opera ting context. The key point is that the criterion upon which the whole RCM philo sophy is based is what is tolerable, not what is practica ble or what is a current industry norm (although these may coincid e). Part 3 of Chapter 5 suggested that the people who are both morally and practically in the best position to decide what is tolerable are the likely victims. These are the shareholders and their management representatives in the case of financial risks, and employees, customers and the managers who have to clear up afterwards (and bear the responsibility) in the case of personal risks. As mentioned above, this appendix shows one way in which it may be possible to turn informed consensus about tolerabl e risk into a frame work for setting targets for maintenance programs designe d to deliver it. Finally, please bear in mind that the approach outlined in this appendix
is not intended to be prescriptive. If you have access to a different frame work which satisfies all the parties involved, then by all means use it.
Appendix 4: Condition Monitoring Techniques Smaller deviation from "normal" needs more sensitive monitoring equipment but gives longer P-F interval
Appendix 4: Condition Monitoring Techniques
Larger deviation from "normal" but shorter P-F interval /
t
Figure A4.1: P-F intervals and deviations from "normal" conditions
1 Introduction some warning of the Chapter 7 explained at length that most failures give is called a potentialfailure, fact that they are about to occur. This warning
identifiable physical condition which indicates that s of occurring. afunctionalfailure is either about to occur or is in the proces ty of an item inabili the as d define is lure lfai iona On the other hand, ajimct potential detect to iques Techn rd. standa to meet 11 specifIedperformance are items e becaus tasks, nance failures are known as on-condition mainte ed specifi meet they that inspected and left in service on the condition
349
Time�
and is defined as an
inspections is determined performance standards. The frequency of these n the emergence of the by the P-F interva l, which is the interval betwee
failure. potential failure and its decay into a functional have existed as long as ues techniq nance mainte Basic on-condition , touch and smell). sound (sight, senses human mankind, in the form of the of using people age advant al technic As explained in Chapter 7, the main ial failure potent of range wide in this capacity is that they can detect a very s are that antage disadv the conditions using these four senses. However, P-F in ated associ and the inspection by humans is relatively imprecise, tervals are usually very short. ed, the longer the P-F But the sooner a potential failure can be detect tions need to be done less interval. Longer P-F intervals mean that inspec ver action is needed to often and/or that there is more time to take whate much effort is being so why is This . avoid the consequences of the failure and develop tech ions condit spent on trying to define potential failure le P-F intervals. possib t niques for detecting them which give the longes interval means P-F long However, Figure A4.1 opposite shows that a is point which higher up the that the potential failure must be detected at a the smaller the deviation P-F curve, But the higher we move up this curve, final stages of deterioration from the "normal" condition, especially if the more sensitive must be the the ion, deviat are not linear. The smaller the potential failure. the detect monitoring technique designed to
2 Categories of Condition Monitoring Techniques Most of the smaller deviations tend to be beyond the range of the human
�
�
sen es and an only be detected by special instruments. In other words, eqll1pment IS used to monitor the condition of other equipment, which is . w y the techmques are called condition monitoring, This name distin
�
� �
gUIshes th m rom the other types of Oil-condition maintenance (perform ance momtonng, quality variation and the human senses).
:
�
�
As entioned in Chapte 7, con ition monitoring techniques are really . . no mOle than hIghly senSItIve verSIOns of the human senses. In the same
:
way as the human sen es react to the symptoms of a potential failure (noise,
�
sme ls, etc), so condItion monitoring techniques are designed to detect
�
speCIfic SYI ptoms (vibration, temperature, etc). For the sake of simplicity,
�
:
th se techmques a e classified according to the symptoms (or potential
failure effects) whIch they monitor, as follows: •
d�namic effects.
?ynam�c monitoring detects potential failures (espe
CIally those aSSOCIated WIth rotating equipment) which calise abnormal
amounts of energy to be emitted in the form of waves such as vibration, pulses and acoustic effects. •
particle effects.
Particle monitoring detects potential failures which
cause discrete particles of different sizes and shapes to he released into the environment in which the item or component is operating. •
chemical effects. Chemical monitoring detects potential failures which cause traceable quantities of chemical elements to be released into the environment.
350
Appendix 4: Condition Monitoring Techniqlles
Reliability-centred Maintenance
changes in the physico! ejji'cts. Physical failure effects encompass
•
ent which can be detec physic al appearance or structure of the equipm ques detect potential techni ring ted d irectly, and the associated monito effects of wear and visible failures in the form of cracks, fractures, the dimensional changes. ques look for po tempe rature efleets. Temperature monitoring techni
•
rature of the equipment tential failure s which cause a rise in the tempe of the material being itself (as opposed to a rise in the temperature processed by the equipment).
electrical c.ffccts: Electrical monitoring techniques look
•
for changes in
potential. resistance, conductivity, dielectric strength and
developed and more are An enormous variety of techniques has been produce an exhaustive list appearing all the time, so it is not possible to appendix provides a very This time. of all the techniques available at any available. Some of these tly brief summary of96 of the techniques curren others are in their infancy or are welJ-known and well-established, while even still under development. is technically feasible However, whether or not any of these techniques ed with the same rigour and worth doing in any context should be assess this process, this appendix as any other on-condition task. To help with
lists the following for each technique: que is meant to detect the potential failure conditions which the techni •
(conditions monito red)
•
•
•
•
s) the equipment it is designed for (application the technique (P-F interval) the P--F intervals typically associated with rough 'ballpark' guide for obvious reasons, this can only be a very how it works (operation) the technique (skill) the training and/or level of skill needed to apply
the advantages of the technique (advantages) es). the disadvantages of the technique (disadvantag worth noting that a is it ques, Finally, before considering specific techni monitoring now ion condit great deal of attention is being focused on ed as being regard often is it adays. Because of its novelty and complexity, . How nance mainte led completely separate from other aspects of schedu is only oring condition monit ever, we should not lose sight of the fact that is used, it should be designed one form of proactive maintenance. When it systems wherever possible, into normal schedules and schedule planning s. system l paralle sive amI not made subject to expen •
•
351
3 Dynamic Monitoring A Preliminary Note on Vibration Analysis Equipment which contains moving parts vibrates at a of These frequencies are governed by the nature of the vibration sources. ami can vary across a very wide range or ,lj)ectrull1_ For instance, the vibration cies associated with a gearbox include the prilllary frequencies of rotation (lf thc shafts (and their harmonics), the tooth contact frequencies of different �ear sets, the ball passing frequencies o f the bearings and so OIl. I f any o f these COI;lpollents starts to tall. Its Vibration characteristics change. and vibratiun is all about detecting and analysing these changes. This is done by measuring how much the item as a whole and then using spectrum analysis techniques to hOllle in on the frequency of vibration of each individual component in order to see whether anything is Howeve r the situation is complicated by the fact that it ,
three different characteristics of vibration. These are amplitude. acceleration. So step one is to decide which of the characteristics is
and
to be measured - and what measuring device is going to he lIsed and then step two I S to deCide which technique will be used to analyse the signal generated the measuring device (or sensor). In general . amplitude (or sensors are more sensitive at lower frequencies, velocity sensors acros\ the middle ran[!es � and accelerometers at higher frequencies. The strength of the at any j re quency is also influenced by how closely the sensors are mounted to the sourcc of the signal at that frequency. Another important characteristic o f vibration is 'phase·. Phase refers to ·the position of a vibrating part at a given instint with reference to a fixed point or another vibrating part'. As a rule, phase measurements are not taken during normal routine vibration measurements. but can provide valuable informatio when a problem has been detected (such as imbalance, bent shafts. misalignment. mechal1lcal looseness. reciprocating forces and eccentric pulleys and gears). Fourier analysis also plays an important part in vibration
�
The role of expert systems in vibration analysis is rapidly coming of age. SOllie systems can now find and diagnose problems as as vibration analysts. They are great time savers, and also enable tIsers to compare readl!1gs WIth a l llhe data from previous measurements. The rest or this part o f this chapter looks at vibration in more detail.
352
3.1
Appendix 4: Conditioll Monitoring
Reliabilitv-celltred Maintenance
B road Band Vibrati on Analysis
in vibration al characte ristics caused by fatigue, mechan ical loosenes s, turbulen ce, etc. ment, misalign ce, imbalan wear,
35 3
Limited i n formation for also l imited by logarithmic frequency scale:
Conditiolls mon itored:
Applications: Shafts, gearboxes, belt drives, compressors, engines, roller bear ings, journal bearings, electric motors. pumps, turbines, etc.
3.3
Constant Bandwidth Analysis
Conditions lIlonitored: Changes in vibrational characteristics caused
wear, imbalance, rnisalignment. mechanical looseness and turbulence, and to iden
P-F interval : L imited warning of fai l ure
tify m u l tIple harmonics and sidebands
primaril y of two parts: a Operati on: A broad band vibratio n system consists
Applications: Shafts, gearboxes. belt drives, compressors, journal bearings, electric motors, pumps, turbines. and and experimental work (especially on gearboxes)
convert mech transduc er which i s mounted on the point of measure ment to indicatin g device anical vibration s i nto an e l ectrical signal, and a measurin g and Monitor s the called a vibration meter, which is calibrate d i n vibration al units. mean square root the simply is which overall reading o f the v ibration signature primaril y suited is and value mple i s a is t I signal. ( R M S ) value of the broad band have a meters Such wave. complex a than rather wave al t o a s i ngle sinusoid z H 000 1 to Hz 20 range the over response y frequenc constant a semi-ski lled worker Skill: T o u s e the equipme nt a n d record t h e v i bration: imbalanc e of rotating Advanta ges." Can be very effective in detecting a major
compact . Can equipme nt. Can be used by inexperie nced personne l . Cheap and ation and Interpret logging. data inimal M . installed ntly permane or be portable criteria such as lity i acceptab n conditio d publishe on based be can ls appraisa VOl 2056 from German y
information about the nature Disadv({ntages: The broad band signal provide s l i ttle
contribute very of the fault. In initial spectra, spectral peaks are much lower and do grow the peaks spectral l ittle to the overall broad band signatur e. When these l t to set fficu i D ation. deterior f o state d equipme n t i s normall y i n an advance ity. alarm level s . Lacks sensitiv
3.2 Octave Band Analysis broad band vibratio n Conditi ons mon itored and applica tions: As for ion P-F interval : Days to weeks depend ing o n applicat octave filters divide the fre Operation: Fixed contiguo us octave and fraction al
a constan t width quency spectrum i n to a series of bands of interest, which have i s measure d filter each from output average The . mically when plotted logarith a recorde r on plotted or meter a by isplayed d are values the and success ively, a s u i tabl y trained technic ian Sk ill: To operate equipm ent and interpre t results: paramet ers have been previ A dl'(lIltages: S imple to use when the measure ment
sive: Good detectio n ously determin ed by an enginee r: Portable : Relative ly inexpen perman ent record. a s provide r Recorde filters: octave al fraction abil ities
P-F illlervai: Usuall y several weeks to months Operation: An accelerometcr detects the vibration and converts it i n to an e1ec .
trical signal which is amp l i fied, subjected to a constant bandwidth filter and then fed 1l1to an analyser. The constant bandwidths are between 3 . 16 Hz and 1 000 Hz �n d the frequency range from 2 H z to 200 kHz. Both l inear and logarithmic frequency sweeps may be selected, but l inear is chosen when identifying har mOl1lCs. In order to analyse the peaks in Illore detaiL bandwidths and ranges can be changed to suit requirements Skill: To operate the equipment:
11
suitably traincd ski lled worker To
the resu lts: an experienced technician A dvantages. S i mple to use when measurement parameters have been set. Good
tor l arge frequency ranges and for detailed i n ve s ti l.w tion at h10h Identifies mUltiple harmonics and side bands which )ccur at con ' tant intervals. Equipment portable
�
�
Disadvantages: Relatively long anal y s i s time. I n-depth
of the
machine harmonics and side bands required to inte rpret results .
3 . 4 Constant Percentage Bandwidth Analysis Conditions monitored: Shock and vibration Applications: Shafts, gearboxes, belt drives, compressors,
rolier bear-
ings, j o urnal bearings. e l ectric motors, pumps, turb i nes, etc. P-F interval: Usually several weeks to months
i s '1<',nnn r,,,,. Operation: High resolution narrow bandwidth frequency by sweeping through the desired frequency range (2 HI to 20 kHz) a con stant percentage fi l ter handwidth ( I 'ft , 6%, I which separatcs c losely spaced frequenCIes or harmonics. A constant percentage fi ltcr bandwidth �arrow as I 't,) allows very fine resolution analyses. Continuous sweeps through the frequency range can be made ill rcal lime.
354
Appendix 4: Condition Monitorillfi Techniques
Reliability-centred Mainte lzance
e d worker, to interpre t the Skill: To operate the equipm ent a suitably trained skill results an experien ced technici an is therefore faster than Advanta ges: Analysi s can be done in ' real time' and
FFf
the batch nature of analysis and does no! suffer from certain pitfalls caused by good for rapid vcry are spectra CPB FFT, such as loss o f data by window ing. faul t detection . Equipme nt portablc Disadvantages: High ski l l required to interpret results.
3.5 Real Time Analysis Conditions monitored: Acoustic and vibrational signals; measurement and analysis o f shock and transient signals
xes etc Applica tions: Rotatin g machin es, shafts, gearbo P-F interl'al: Several weeks to months
played back Operation: A signal is recorded on magneti c tape and
through a real
the frequen cy domain . timer analyse r. The signal is sampled and transformed into 400 equally spaced at d measure d, produce A constant bandwid th spectrum is 0- J 0 Hz to 0-20 from le selectab range cy frequen a frequen cy interval s across also be adj u s ted can scan the and , seleeted be can mode on kHz. A high resoluti baseban d spectrum the in ehanges any allowing , analysis motion' 'slow a to to be observe d as the time window is stepped along an experie nced enginee r Skill: To operate equipm ent and interpre t results: the entire analysis range simul Advantages: Analyse s all frequen cy bands over d spectra is continu ously analyse f o display al graphic neous Instanta taneous ly: to analysis o f short updated : No need to wait for level recorde r readout : Suited rs provide a recorde X-Y shoek: duratio n signals such as transien t vibratio n and perman ent record expensi ve: Needs high level Disadvantages: Equipm ent not portable and very of skil l: Off-line analysi s.
3.6 Time Waveform Analysis gear teeth; pump eavitati on, Conditions monitored: Chipped , craeked , broken misalign ment, mechani cal loosenes s, eccentricity, etc.
Applicatiolls: Gearboxes, pumps, rol l cr bearings, etc
I'-F interval: Usually several weeks to months
ed to a standar d vibration analyse r or a Operation: An oscil loscope is connect
cope vertical in real time analyse r. A vibratio n signal is applied to the oscillos horizon tal axis the and de amplitu in scaled is put. The vertical axis o n the CRT
:i55
is scaled in time such as seconds or milliscC()nd�. The vertical is adjusted until the peak or peak-to-peak value of the waveform displayed o n t h e C R T corresponds to t h e amplitude reading on t h c vibration rueter. W h e n a machine is generating a single frequency, the time waveform is simply a sine wave with the repetition rate of the running speed of the equipment. When the equipment generates more than one frequency, a complex composite wave form is generated. Additional frequencies can be generated in the form o f pulse, transient, beat, modulation, etc. which add to the complexitv o f t he waveform. To reduce the complexity of the wavef<)fIll. it is useful and e � en essential to use variabl e high pas s , low pass and band pass filters Skill: Needs cOllsiderable practice and experieIlce to interpret
wave limns
Advantages: Good for looking at transients, s low heats,
non-linearities, sine waves, amplitude modulation, frequency modulation, instabilities, ete. Often provides more information than frequency analysis. The wave forIll can be used to distinguish between spectra resulting from impacts or random noise Disadvantages: Machines that generate multiple frequencies often generate
noise which makes time wave forms so complex and confusing that they are difficu l t to break down into component parts. To examine a wave form which might have s low beats, a long time record is required
3.7 Time Synchronolls Averaging Analysis Conditions monitored: Wear, fatigue, stress waves emitted as a res u l t of metal
to-metal impacting, microwelding, etc Applications: Gearbox gear teeth. rol ler bearings, shafts, bank of fans, rol l s on
a paper machine, rotating machines
.
P-F interval: Usuall y several Weeks to months a s lightly varied with each rotation. ( S tatistically this is termed ' stochastic ' , in eomparison with identically repeated signals which are 'deterministic' . ) The c loser the tolerallees o f sliding and rolling parts, the less the variation, but nevertheless there is a variation. In many systems, this difference can be so great that it masks any changes due to a developing fault. The presence o f random noise can also COIl fuse the signal. These problems can be overcome by performing a level check at precisely the same part of the cycle rotation using a tachomder to or time records arc initiate data capture in a data collector. A number of averaged together. Signals not related to the RPM of the shaft are (lut, leaving a very clear ' real-time' wave representing the eomponcnts related to a single turning speed. The averaged wave for m can be examined directlv or a spectrum can be generated from it. I t is devoid of random and wi I I' show whether one part o f the cycle is changing more than another
Operation: Most rotating mechanical systems produce
356
Reliability-centred Maintenance
considera ble Skill: A suitably trained and experienc ed skilled worker. Requires practice and experiene e to i nterpret the resuits
speci fically the individua l gears can be analysed i n detai l . Very useful for analysing equipmen t that has many componen ts rotating at nearly the same speed
Advantag es: Gearboxe s
bearings Disadvan tages: Care must be taken w i th machines w i th roller element as the bearing tones are not synchron ous w i th the RPM and w i l l be averaged out.
3.8 Frequency Analysis by fatigue, COllditiollS monitored: Changes in vibration c haracteris tics caused etc e, turbulenc , looseness al mechanic ent, misalignm wear, imbalanc e, engines, roller bearApplicatio ns: Shafts, gearboxes , belt drives , compresso rs, ings, journal
electric motors, pumps, turbines, etc
P-F illlaval: Several weeks to months
time domain and Operation: Data i s collected from measurem ent points in the
transformed i n to the frequency domain u s i ng a Fast Fourier Transform (FFT) algorithm , by eit her the data collector itself or a host computer . The required frequency range of the measurem ents is dependen t on the speed of,the machine. es. Each machine which has moving parts w i l l produce a spectrum of trequenCi an to A baseline spectrum o f the machine in excellent condition i s compared Any load. and speed actual spectrum of the same machine running at the same forcing increases over the baseline of more than one standard deviation at any is analysis ueney frel of feature One l frequency can indicate a potential problem. same the at taken s signature are ls Waterfal s. signature the " waterfall " o f FFf point over a time i n terval allowing the signature s to be trended Requires consider able Skill: A s u itably trained and experienc ed s k i l led worker. practice and
to interpret the results
and easy to use. Expert A dvantage s: Data collectin g equipme nt i s portable plots small software systems makes data interpreta tion easy. Using waterfall stage early an at detected be can condition machine in changes
Appendix 4: CondiliOIl lVlonitoring
357
A s a machine w ears, i t develops l1on-lineanties that cansI.' harmoll ics of the primary force frequencies, and sum and diffcrencc appear in the vibration spectra. Cepstrum (pronounced "kepstrtlm ") separates the harmonics and s idebands present in the spectrum so be individuall y trended over time. In simple terms cepstrum could be defined as the FFT of the logarithmic spcetrum obtained from the FFT 'a spectrum u f a spec trum· . All technical words in cepstrum analysed are reversed due to the lip of the transform, i.e. spectrum becomes cepstrum, frequency becomes ljue frency and harmonies become rah rnonics Skill: I n-depth understanding of machine behaviour (harmonics and
and the expert software Advantages: Can analyse harmonics and sidebands that
![]
complex machines. S idebands are easy to find in the spectra of roller element bearings. Can be performed with some expert system software. Disadvwlt(lges: Skill and expelience needed to interpret harmonics and sidebands.
3. 1 0 Amplitude Demodulation Conditions monitored: Bearing tones masked by noise, cracks in eccentric or damaged gears, mechanical looseness Applications: Steam turbine;" bearings anel
low ponents of paper machines. reciprocating machines, etc
races, COI11-
P-F interval: Several weeks to months Operation: The acceleration analog signal (time dOl11uin) i s to h i gh pass filtering and then amplitude demodulation . Thi s is where a discrete frequ ency often called the "carrier" in the spectrum may be modulated by another frequency called the modulator. The res u lting signal i s then subjected to a low frequency range spectrum analysis. The amplitude demodulation i s performed in the data collector before the signal is digitised. Skill: A s uitably trained and experienced technician A dvantages: Early detection for bearing and gearbox problems
noise can look similar. Disadvantages: Spectra resulting from impacts and random
bearings completely masked by can easi l y be identified. Works w e l l in low-speed applications such as paper machines
3.9 Cepstrum
Disadvantages: High skill and experience needed to understand and
s in vibration spectra Conditions monitored: Wear causing harmonic s and sideband
results. D i fficu l t to i mplement on s low speed bearings because the stress are short-term transient events (less than a few the narrow pulse output from the demodulation c i rcuit i\ stage of signal conditioning (the low pass/anti -alias i ng fraction of the stress wave i s filtered out, making faul t detectioll less l ikely.
gear meshing , bel t rota Applicat ions: R o l l ing c lement bearings , shafts, gears, tion, vane and blade pass frequenc ies of pumps and fans
P-F imaml: Several weeks to months
35�
Reliability·centred Maintenance
3.1 1 Peak Value (PeakVue) Analysis Conditions monitored: Strcss waves caused by metal-to-metal impacts or metal tearing, stress cracking or scuffing, spalling an d abrasive wear Applicatiolls: Anti-friction bearings and gearbox shafts and gearing systems
P-F interval: Several weeks to months depending on the application Operation: Separates low energy faults such as those that occur i n anti-friction
bearings and gears, and enhances their signal causing the faults to stand above the spectral noise floor. This makes them easier to recognise. PeakYue first separates the stress waves from the v i bration waveform using a h i gh pass filter. It is then conditioned to enhance its ampli tude and pulse width, making it FFT friendly. The conditioned waveform is then processed using an FFT to deter mine the frequency at which the stress wave occurs Skill: Experienced vibration technician Advantages: R e veals some faults that may have gone undetected in their earlier
stages or which are buried in the noise noor (bottom) of the vibration spectrum. More consistent than demodulation. Outputs are independent of machine speeds and i nstrument Fmax settings. Applicable to a broad range of frequeneies, from very slow speed bearings to gear meshing i n excess of I kHz Disadvantages: High skill alld experience required to interpret res u l ts.
3. 1 2 Spike EnergyTM Conditions monitored: Dry running pumps, cavitation, flow change, bearing loose fit, bearing wear causing metal to metal contact, surface flaws of gear teeth, high
Appendix 4: Condition Monitoring Techniques
359
SAil!. A suitably trained and experienc ed skilled worker. Requi res practice and experience to interpret the res u l ts Adv(tllIages: Sensitive h igh frequenc y measurem ent paramete rs suited to the
detection of seal less pump problems w hich are often difficult to detect using ' conventio nal v ibration sensors such as velocity meters and accelerom eters Disadmll tages: High s k i l l and experienc e needed to interpret results .
3.13 Proximity A nalysis Conditions lIlonitored: M i salignme nt, o i l whirl, rubs. i mbalance /bent shafts. resonance. reeiprocati ng forces. eccentric pulleys and gears, ell" Applications: Shafts, motor assemblie s, gearboxes , fans, couplings , etc
P-F interval: Days to weeks
Op eration: In the basic mode, a signal from a transduce r operates as the ordinate
against a time base. W i th a single impulse. sinusoida l curves indicate imbalanc e, bent shafts, oil whirl, misalignme nt, adhesive hearing ntbs. Two a polar diagram which provides more character istic informati on than an X · Y dia gram. More informati on can be obtained by introducin g a phase indicating mark on the wave forms of the osci ll oscope display. These marks are generated at the rate of one revol ution by a pick up incorporat ed ill the tachomcter Skill: A suitably trained and experienced technician A dvantages : Pinpoints specific problems . Can be lIsed for balanci m, : Portable: �
Very simple to use
pressure steam or air now, control valves noise, poor bearing lubrication
IJisadv(l/1 t({ges: P-F interval short: Long analysis time: Diagnosti c ability l imited.
Applications: Seal-less pumps used in the chemical and petrochemieal industry,
3.14 Shock Pulse Monitoring
gearboxes, rol ler e l ement bearings, ete P-F interval: Several weeks to months Operation: Some faults excite the natural frequencies o f components and struc
tures. The intense energy generated by repetitive transient mechanical impacts causes a signal to appear as periodie spikes of high frequency energy i n a spectrum which can be measured by an accelerometer. A h i gh frequency hand pass filter is used to fil ter out low frequency vibration signals. The high frequency signals pass through a peak-to-peak detector that deteets and holds the peak-to-peak amplitudes o f the signal. This is c a lled enveloping, and measurement res u l ts are units. Pulses w i th large ampli tudes and high repetition rates expressed i n produce high overall readings. The enveloped signal can be subjected to a FFT analysis displaying a Spike Energy Spectrum. In the gSE spectrum, the fau l t frequency shows u p as certain defeet frequency and its harmonics.
Conditions monitored: S urfaee deteriorati on and lack o f lubrication meehanic al shock waves. W i th data trending can identify incorrect i nstall ation or replaceme nt, using the wrong type of lubricant, poor lubricant handling or dispensing practices, or i ncorrect i nstallation or maintenan ce of oil seals and pac kings, etc. Applications: Roll i ng element bearings, anti-fi
tools, valves of internal combustion engines P-F interval: Weeks to several months Operation: The type and size of the bearing is entered into the A piezoelectr ic accelerome ter placed on a bearing housing detects the m" chauical impact o f shock impUlses, caused by the impact o f two masses (such as the rota tional contaet between the surfaces of the ball or rol ler and the raceway ) . The
360
Reliability-centred Maintenance of the shock
depend on the surfacc condition and the peripheral
velocity of the bearing (rpm and size). The pulses set up a dampened oscil l ation in the transducer at its resonant frequency. The transducer is tuned mechanically and electricall y to a resonant frequency of 3 2 kHz. The peak ampl itude o f this oscillation i s directly proportional to the impact velocity. A s the bearing condi tion deteriorates from good to i m minent fail ure, shock pulse measurements can i ncrease lip to 1 000 times. 5'kill: A trained and suitably experienced technician Advantages: Relatively easy to operate. Portable. Can be used on virtually any
rol ler e lement bearing. Bearing condition and lubrication status analysed within seconds. Shock impulses are not significantly intluenced by background vibra tion and noise. Identifies subtle ehanges in bearing condition or l ubrication which might not be differentiated by eonventional vibration analysis Disadvllnt(lges: Needs ace urate bearing s ize and speed i n formation prior to
taking measurement Limited to rol ler e lement bearings.
3.15 Ultrasonic Analysis COllditions monitored: Changes in sound patterns (sonic signatures) caused by
leaks, wear, fatigue or deterioration Applications: Leaks in pressure and vacuum systems (ie. boilers, heat exehangers,
condensers, chillers, dist i l l ation columns, vacuum furnaees, specialised gas systems): bearing wear or fatigue: steam traps: valve and valve seat wear: pump cavitation: corona in switehgear: static discharge : the integrity of seals and gaskets in tanks, pipe systems and large walk-in boxes: underground pipe or tank leaks P-F interval: Highly variable depending on the nature of the faul t Operatioll: U l trasound technology i s concerned w i th high frequency sound
waves above human pereeption (20 H z to 20 kHz) ranging between 20 kHz to 1 00 kHz. High frequency sound waves are extremely short and tend to be fairly directional, so it i s easy to isolate these signals from background noises and de teet their exact location. A l l operating equipment and most leakage problems produce a broad range of sound. A s subtle changes begin to occur w i th deteri oration, the nature o f the airborne u l trasound allows these warning signals to be detected early. U ltrasonic translators convert the ultrasound sensed by the instru ment into the audible range where users can hear and recognise them through headphones. The u ltrasonie monitoring equipment filters ont surrounding noise and other unwanted frequencies. U l trasonic readings may be displayed visually on a V D U or a moving coil meter, as an audible signal on headphones or as traces on an electronic monitor or eOl11puter Skill: A s u itably trained skil le d worker
Appelldix 4: Condition Monitoring r
36 1
. Quick and easy. Can be llsed in very noisy areas screen ambient noise). Microphones highly direetional and enable the operator to detect a souree of noise at long range. Equipment portable Disadvantages: Does not indicate the size of leaks.
tanks can
be tested under vacuum.
3.16 Kurtosis Conditio liS monitored: Shock pulses Applicatiolls: Rolling e lement bearings, anti-friction
P-F illterml: Several weeks to months Operatioll. Restrieted almost exclusively to where a few frequency ranges are examined (3-5 kHz, 5 1 0 kHz. 10 1 5 kHz). K urtosis is a statistical analysis of the time-based (ti me domain) signal and looks at the fourth moment or the spectral amplitude difference from the l1lean leveL A llormal dis tribution has a kurtosis ( K ) value of 3 Skill: A suitably trained sem i -skil led worker Advantages: i\pplicable to any materials with a hard surfacc. Equipment
Very simple to use Disadvantages: Limited application and significantly affected by impact nuise
from other sources. Considered by SCHne u sers to be too sensitive.
3.17 Acoustic Emission '
Conditiolls monitored: Plastie deformation and crack formation caused
stress and wear Applications: Metal materials used in struetures, pressure vessels.
and
underground m i ning excavations P·F interval: Several weeks depending on application Operation: Audible stress waves, due to crystallographic are emitted from materials subjeeted to loads. Thesc stress waves arc picked up a trans ducer and fed via an amplifier to a pulse analyser, then either to an X-Y re�'order or to an oscil loscope. The displayed signal is then evaluated Skill: A suitabl y trained and experienced technician. Advantages: Remote detection of flaws. Covers entre structures.
system set up very quickly. High sensitivity. Requires limited access to test Detects active flaws. Only relative loads are required. Can sometimes llsed to forecast fail ure loads.
362
Reliability-centred Maintenance
The structure has to be loaded. A -E activity dependent on mate rials. Irrelevant electrical and mechanical noise can interfere with measurements. Gives l imited information on the type of flaw. Interpretation may be difficult.
4 Particle Monitoring 4.1 Ferrography COllliitioliS lIlonitored.- Wear, fatigue and corrosion particles Applications: Greases: Oil s used in diesel and gaso l i ne engines, gas turbines, transmissions. gearboxes, compressors and hydraul i c systems
P-F interval: Usually several months Operation: A representative sample i s dil uted with a fixer solvent ( tetrachloro e thylene) and then passed over an inclined glass slide under the inf1uence of a graduated magnetic field. The particles are distributed along the length of the
Appendix 4. Condition Monitorillg
363
P F intam!: Usually several months
Opowioll: An analytical ferrograph is used to prepare a ferrogram as described under ferrography. After the particles have been deposited on the a wash is used to flush away any remaining oil or water-based lubricant. Once the wash evaporates, the partieles remain penmmently attached to the substrate on the ferrogram. A Fe!Togram Scanner scans the ferrogram in less than 2 0 seconds and generates standard output values that correspond to the wear mechanism. Various partieles are graded by their types and shapes which revcal problems. For example, laminar metals ( having a 'peeled l oo k ' . and thin) often indicate a problem with roller bearings. Red oxicks typically are rust ( l i kely water contamination). The software then report� tbe wear levels and changes in condition of the component
Ski/!: To draw sample and operate ferrograph: a suitably trained semi-skilled worker. Tll analyse and interpret the n;s u lts: an
technician
51 ide according to their size. Larger particles are deposited near the entry. while fine
Advantages: Available in a ,vide range of on-l inc syste m s . I n- depth evaluation.
particles are deposited near the exit of the s lide. The s lide, known as a ferrogram, is treated so that the particles adhere to the surface when the oil is removed. Ferrous particles are separated magneticall y and are distinguished by their a lignment to
Disadvantages: H i gh level o f operator cxperience required. Time
photographic recording and data-base management. Less affected tluiLi opa city and water contamination than many other techniques_ Equipment
the magnetic fields l i nes. while non-magnetic and non-metallic particles are distri buted in a random fashion over the entire s lide. The total density of the particles and the ratio o f large to small particles indicate the type and extent of wear. Analysis
sample preparation and analysis. The need to dilute s
is done by a techniquc known as hichromatic microscopic examination. Thi s uses both ret1ected and transmitted I ight sources (which may be used simultaneousl y ) . Green, red and polarised filters are a l s o used t o distinguish t h e size, composition,
4.3 Direct Reading Ferrograph (DRF)
shape and texture of both meta l l ic and non-metallic particles. An electron micro scope can also be used to determine particle shapes and provide an indication o f t h e cause o f fai lure
Applications: Oils used in dicsel and gasoline engines, gas turbines, transmis
Skill: To draw sample and operate ferrograph: a suitably trained semi-ski l l e d worker. T o analyse a n d interpret the results: a n experienced technician
Advantafies: More sensitive than emission spectrometry at early stages of wear: Measures particle shapes and s i zes: Provides a permanent pictorial record
Disadvantages: Not an on-line technique: Time consuming, and needs some very expensive analytical support equipment: Measures generall y only the ferromag netic particles: Requires an electron microscope for an in-depth analysis.
that the sample w i l l actually be representative of actual wear.
Conditions monitored: M achine wear. fatigue and corrosion particles sions. gearboxes, compressors and hydraul i c systems,
P-F interval: Usually several months Operation: A D R F quantitatively measures the concentration of ferrous part icles in a fluid sample by precipitating these particles onto the bottom o f a tube subjected to a strong magnetic fielel. Fibre optic bundles direct light through and the glas s tube at two positions correspondi ng to the location where small particles are deposited by the magnet. The light is reduced in relation to the number o f particles deposited in the glass tube, and this reduction i s monitored and displayed electronicall y . Two sets of readings are obtained fur small particles (above and below 5 microns) which are plnlted on a
4.2 Analytical Ferrography
Skill: A suitably trained semi-skil l ed worker
Conditions Inonitored: Wear, fatigue and corrosion particles
Advantages: Compact, portable. on-l i ne technique,
Applications: Greases .. O i l s used in diesel and gaso l i ne engines, gas turbines, transmissions, gearboxes, compressors and h y draulic systems
and
to operatc. Less sensitive to fluid opacity and water contamination than some techniques.
.364
Reliability-centred Maintel1ance
Appendix 4: Condition MOllitoring
365
Disadvantages: Measures only ferromagnetic paJ1icles: Requires further analytical
a tlow-decay time curve. The hand held computer uses a mathematical program
ferrographic analysis when readings are high.
to convert the now-decay time curve i n to a particle size distributioll. This is used to compute an ISO cleanliness code.
4.4 Mesh Obscuration Particle Counter (pressure Differential) Conditiulls monitored: Particles i n lubricating and hydraulic oil systems caused
by wear, fatigue, corrosion and contaminants Applicatiolls: Enclosed l ubricating and hydraul i c oil systems such as cngines,
transmissions, compressors, etc. P-F interval: Usually several weeks to months. Operation: This instrument measures the differential pressure across three h i gh
precision 5, 1 5 , 2S micron screens, each w i th a known number of pores. As the oil passes through each screen, particles larger than the pores are trapped o n the mesh surface, which reduces the open area of the sereen and increases the pressure drop across the scrcen. Sensors measure the pressure change which is converted to reflect the number of particles larger than the screen size. Thi s i s converted in turn i n to ISO 4406 clean l i ness codes. Skill: To operate thc portable u nit: a s uitably trained semi-ski l l ed worker. To
interpret thc rt:sults: a s u i tabl y trained and experienced technician Advantages: No pre -sample preparation. Equipment is portable and can be used
in the field or the laboratory. A n i n - l i ne version o f the equipment can be used for real time continuous monitoring. Particle counts are ealibrated to an ISO 4406 cleanl iness standard. M os t o i l s can be analysed in a matter of m inutes. Not atTec ted by bubbles, emulsions or dark o i l s that limit laser-based analyser:; Disadvantages: Provides no indication of the chemical composition of particles.
Only applicable to circ u lating oil systems. Equipment moderately expensive.
4.5 Pore-blockage (Flow Decay) Technique Conditions monitored: Particles in lubricating and hydraul i c o i l caused by wear,
Skill: To operate the portable unit: a suitably trained skil led worker. To
the results: a suitabl y trained and experienced tech nician Adl'alliages: No pre-sample preparation. Equipment i s
and can b e llsed in the field or the laboratory. An i n - line version of the equipment can he llsed for continuous mon itori ng. Particle counts arc calibrated to an I S O 4406 cleanl iness standard. Most oils can be analysed in a matter of minutes,
Disadvantages: : Provides no indication of the chemical composition
Only applicable to circulating o i l systems. Equipment 4.6
Light Extinction Particle Counler
Conditiolls lIlonitored: Particles i n lubricating and hydrau lic oil caused fatigue, corrosion and contaminants Applicatiolls: O i l s used i n diesel and gasoline cngines.
weaL
turbines. transmis-
sions, gearboxes. compressors and hydraulic systems. P-F interval: Usually several weeks to months Operation: The l ight extinction particle counter consists of an i ncandescent
light source, an object cell and a photo detector. The sample fluid moves through the object cell under controlled flow and volume conditioll s . When opaque part icles in the Ouid pass through the beam it blocks an amount of l ight to the particle s i zes. The n u mber and size uf the particles i n the oil sample deter mine how much light is blocked or reflected, and how much l ight passes through to the photo diode. The res ultant change in the electrical signal at the photo diode i s analysed against a calibration standard to caleulate the number uf in predetermined s i ze ranges a n d displays t h e count. From t h i s i nformation a direct reading o f the ISO c l ean l iness value is determined automatically.
fatigue, corrosion and contaminants
Skill: To operate the portable unit: a s u itably trained ski l led worker
Applications: O i l s used in diesel and gasoline engines, gas turbines, transmis
Advantages: Considerably faster than v isual graded fi ltration. Tcst resu l t s avai l
siems, gearboxes, compressors and hydraul i c systems.
able within minutes. Generally the test is quite accurate and reproducible.
P-F Interval: Usually several weeks to months
Disadvantages: Lacks the i n tensity and consistency of laser and fai l s to m·er
(}peration: A fluid sample is pressurised between 30 and I SO psi (can go as h i gh
as :lOO() psi) and allowed to flow through a selected precision calibration screen 1 0, 1 5 micron) depending on oil viscosity, in a sensor assembly. Particles than the screen start to accumulate, restricting the flow. S maller particles particles restricting the flow even further. The result is around the
come reaction of the many different wavelengths of on flu i d opacity. the n umber of translucent particles, air bubbles and water con tamination. The count and size may also v ary depl�nding on the orientation of long, thin or unusually shaped particles in the light beam. Resol utions l i mited to S microns particle range. Provides no i n formation on the chemical of the contaminants.
366
Appendix 4: Conditiull Monitoring Tt'c/miqucs
Reliability-centred Maintenance
4.7 Light Scattering Particle Counter COllditions monitored: Particles in lubricating and hydraulic o i l caused by wear, fatigue, corrosion and contaminants Applications: Enclosed l ubricating and hydraul i c oil systems such as engines,
gearboxes, transmissions, compressors, etc. P-F interval: Usually several weeks to months Operation: The light scattering particle counter consists of three primary
components; a laser light source, an object cell and a photo diode. The sample t1uid moves through the object cell under controlled now and volume conditions. When opaque particles in the fluid pass through the beam, the scattering of l i gh t i s measured and translated into a particle count. From t h i s information a direct reading of the ISO cleanliness value is determined automaticall y . Skill: A s u i tably trained skill e d worker Advantages: Good performance in settings where conditions are contro l l ed .
H i gh accuracy. Measurcs particles a s small as 2 microns . Faster than t h e visual graded filtration test res u lts available within m inutes. Generally the test is quite accurate and reproducible. Continuous monitoring i s possible. Disadvantages: Accuracy dcpendent on fluid opacity, the n u mber of translucent
particles, air bubbl es and water contamination. The count and size may also vary depending on thc orientation of long, thin or unusuall y shaped particles in the ligh t beam. Provides no information on the chemical composition of contami nants. Dilution i s often required for high particle concentrations to avoid coinci
Disadvwlwges: L i m i ted to collecting ferromagnetic
367 Indicates
lllass of ferromagnetic particles only
4.9 All-Metal Debris Sensors Conditiolls monitored: Ferrous and non-!"elTouS particles due to wear and Applications: Designed speci fically for the protection of gas turbille
P-F interval: Weeks to lllonths Operation: The sensor head consists uf three coils wound around an section of pipe . The outer stimulus coil s an' with an frequency signal. The sense coil (middle) is placed exactly at the null point between the stimulus coils. When a ferrous particle passes through the sensor. it di�turbs the first fie l d anel then the second. generating a readi l y detectable signature i n t h e sellse coil. A non-ferrous particle generates a unique and opposite s ignature. The sensor w i l l detect and measure most o f the severe wear particle range. These
signatures are captured anel stored as time domain plots and arC' used real time to ,den/advise operators, or to signal automatic responses from control systems Skill: E,xperienced skil l ed worker/technician to trend results Advalltages: Detects and quantifies both ferrous and non-ferrous wear metal
particles. Low probabil ity o f a false indication. On· board sensors can capture anel store the time domain plots o f vario lIS damage modes which can be used for identification of wear sources i n near feul time. Disadvantages: Cannot determine chemical composition and size of particles.
dence error where several particles bunch together and appear as one l arge particle. 4.10
4.8 Real Time Ferromagnetic Sensor
Graded Filtration
Conditions monitored: Particles in lubricating and hydrau lic oil caused
wear.
Conditions monitored: Ferromagnetic particles caused by wear and fatigue
fatigue, corrosion and contaminants
Applications: Oils u sed in diesel and gasoline engines, gas turbines, transmis
Applications: O i l s used in diesel and gasol i ne engines, gas turbines. transmis
sions, gearboxes, compressors and hydraul i c systems.
sions, gearboxes, compressors and hydraulic systems.
P-F interval: Weeks to months
P-F interval: Usually several weeks to months
O peration: A n analog felTomagnetic sensor uses an inductive or magnetic prin ciple to measure the quantity o f ferrous particles passing the sensor. The sensor attracts the ferrous particles with an e lectromagnet. The partic les collect around a sense coil causing a change in an o sc i llator frequency. The frequency is cali brated to indicate the mass o f ferrou s particles collected. After a measurement has been taken, the particles are released. Measurements can be trended over t i me.
Operatioll. A small amount o f oil ( I O(J ml) is diluted and o f standard filter disks. Each disk i s then examined under a particles are counted manually. The results are expressed as the n umber in a particular size range. Their statistical d istribution i s showll in the form of a graph. Analysis of the particles distribution profiles indicatcs whether wear i s normal or not.
Skill: Experienced skil led worker/technician.
Skill: Sampli ng : a l aboratory assistant. Examination of part ide d istribution pro
On-line technique
files: an experienced laboratory technician or
368
Reliability-centred Maintenance . Contaminants such as mctal chips , pieces of seal materiaL or dirt
can be identified visually. Relatively eheap
Appendix 4: Condition MOllitoring hours is needed for the o i l to "blot" fully, after which the re�ults may be photometricall y . The test provides an indication when
Disad\'(/ntilges: S ubjeetive because the operator has to determine v i sually the
oil are reaching their end of their useful l i fe. Some portable te,l kits have reference
si.1e of the particles, even though there are grid markings for reference. Setting
standards which can provide an indication of t he level of
u p and examining each fi lter disk sample takes several hours. Specialist s k i l l s
present
Skill: Oil blotting: a s u itably trained semi-skilled worker.
required to i nterpret t h e t e s t results. Ident i fication o f particle elements difficult.
trained experienced technician
4. 1 1 Magnetic Chip Detection
Adl'illltages: Cheap, and easy to use and set up: Provides a record:
COllditions monitored: Wear and fatigue . Oils used i n diesel and gasoline engines, gas turbines, tran s m i s sions. gearboxes, compressors and hydraulic s ystems
accurate indicator of oil oxidation 24 hours needed for the oil to blot: COll,iderable s k i l l
level. Does not indicate
to interpret the results: Only a rough indication chemical composition of particles.
"" F interval: Days to weeks Operation: A magnetic plug i s mounted i n the l u bricating system so that the
4.13
Patch Test
magnetic probe is exposed to the circulating l ubricant. Fine metal particles
Conditions monitored: Wear metals , fatigue and corrosion particles,
suspended i n the uil and metal flakes Crom fatigue break up are captured by the
Applications: Oils used i n diesel and gasuline engines. gas turbines. transmis
probc. The probe is removed regularly for microscopic examination of the cap tured particles An increase particle size indicates i m m i nent fai l ure. The debris has different characteristics ( shape, colour, and texture) depending on its source
Skill: To collect the sample: a su itably trained semi-skilled worker To analyse the debris : a suitably trained and experienced technician
Advantages: Cheap. Low powered microscope only required for the anal y s i s of
de.
sions, gearboxes, compressors and hydraulic systems
P-F interval: Days to weeks a
Opemtioll: A vac u u m is used to draw a standard volumc of test fluid 5 micron M i l lipore 47 mm disc filter. The
of discoloration Ull the lilter is
compared with a standard membrane filter colour rating scale and
the debri s : Some probes can be removed without loss o f lubricant
men! scale to determine contamination levels. H igh particle Ievcls pmdllce a darker gray or a more h ighly coloured spot. free water appears ei ther as
DislIdvallIages: Short P-F interval: High skill required to i nterpret the debris
during the test procedure, or as a stain on the test filter The patch i s examined
4 . 1 2 Blot Testing
c1es and to give a quick i m pression of the type and size of particles. An
Conditions mOil ito red: Wear metals, fatigue and sometimes corrosion particles. s ludge, etc
Applications: O i l s used i n diesel and gasoline engines, gas turbines, transmis sions, gearboxes, compressors and hydraulic systems
P-F Interval: A few days to a few weeks Operation: One or two drops of oil are placed o n a flat piece of blotting paper or filter paper. The o i l drops spread out and dry, the large particles remain within a centre circular corona o f a small radius. This removes many organometal l i c a n d detergent-dispersant additives. Further dispersion leads t o o i l penetration and filtration through the paper, so e i rcular zones corresponding to the size o f particles transported by t h e filtering o i l are c learly defined. A sharply defined around the oil welled area indicates the present of s ludge. A period of 24
using a microscope to determine whether the system is heavily loaded with mate (qualitative) c l ean liness rating can be determined
the patch
to a picture chart
Skill: A suitably trained and experienced skilled worker Advantages: Test results are dependable, repeatable and sensitive
to detcct
any significant change in cleanliness. Good qualitative measure of c()ntami na� tion. Portable
DisadvCllltaf{es: U�ing a microscope to count wear or contaminant
i:.; tedious, eannot be calibrated, and is subject to high levels of user-to-llscr variance. 4.14
Sediment (ASTM D - 1 698)
Conditions monitored: Inorganic sediment from contaminatiun,
�ediment
from oil deterioratio n or contamination: soluble sludge from oil deterioration
370
PetroleUll1 based insulating oils ill transformers, breakers, and cables
37 1
Appendix 4: COllditioll Moniloring T
Reliahilitv-centred Maintenance
Wear metals : the following wear mctals are measured in lubricating o i ls
Aluminium from pistons, journal bearings, shims, thrust washers, accessory
P-F intermi: Sevcral weeks
casings, bearing cages of planetary, pumps, gears,
lube pumps etc
The upper, sediment-free portion is decanted and used to measure soluble sludge
Antimony from somc bearing alloys and grease Chromium from the wear plated components slich as shafts, sl'als,
by dilution with pcntane to precipitate pentane insol ubles and filtration through
rings, cylinder l iners, bearing cages and somc bemings
Operation: An oil sample is centrifuged to separate the sediment from the oiL
a filtering crucible, The sediment is dislodged and filtered through a filtering
Coppcr from journal bearings, thrust bcarings, cam and rOl'ker arlll
crucible, After drying and weighing to obtain total sedi ment the crucible i s
piston pin bushings, gears, valves, clutches, and
Prcsent in brass or bronze alloys and often detected in conjullctioll with line in the
ignited a t 500°C and reweighed, Loss i n weight is organic and t h e remainder i s
former and tin in the latter
inorganic content of the sediment
Skill: Taking t he sample an e l ectrician, To conduct the test: a suitably trained
- Iron from cast cyl inder l i ners, piston rings, pistons, camshafts, crankshafts, valve guides, anti-friction bearing rollers and races, gears, shafts, lube pumps
laboratory technician
Advantages: Test quick and easy. Transformer does not have to be taken of11 i ne to monitor the insu lating fluid
Disadvalltages: Test suitable for low viscosity oils only, for example 5 , 7 to 1 3 ,0 cSt at 40°C ( I 04°F), Test has to be conducted i n a laboratory, Pentane i s mildly toxic and flammable.
4.15 LIDAR (LIght Dctection And Ranging)
and machinery structures, etc
Lcad from journal beari ngs and seals - J\:lagncsium from turbine accessory casings, shafts and \'alvc') Manganesc from valves and blowers - Molybdcnum from wear to plated uppcr piston rings i n some diesel Nickcl from valves, turbine blades, turhochargl'r cam plate:s and - Silver from locomotive engines, solder and needle - Tin from bearin g alloys, brass, oil seals and solder - Titanium found in bearing hubs, turbine blaclcs and compressor discs of gas •
•
•
COllditiolls monitored: Presence of particles in the atmosphere Applicatiolls,' Quality and dispers i on of plumes of smoke from smokestacks P-F illferril/.' Highly variable depending on the application
Operatioll: S ingle wavelength light is directed to the area under investigation,
The quantity o f particulate matter is assessed by measuring backscatter. Loca tions are determined by triangulation based on readings taken from two points.
Skill: A n experienced e n gineer A dvantai\es, A remote sensing technique which can cover large areas Disadvanfai\i?s: V ery expensive: Requires a h igh level of skilL
5 Chemical Monitoring A Preliminary Not.c on the Chcmical Detcction of Contaminants in J;'luids The teehniqlles described i n this section of part 5 are used to deteet elements i n fluids usually lubricating o i l which indicate that a potential failure has oecurred elsewhere in the system, as opposed to incipient failure of the fluid itself The elements most commonly detected by these techniques are l isted below, and they can appear as a result of wear, leaks or corrosion.
turbine aircraft engines
- Zinc from brass components, neoprene seals , Leaks: t h e fol lowing elements are associated with leaks
Aluminium from atmosphere contamination Boron from coolant leaks i n o i l Calcium w h e n found i n fuel, generally indicates contam inat ion seawater Copper from oil cooler cores - cooling water il] oil - Magnesium from seawater contamination - Phosphorus from a coolant leak in o i l - Potassium from contamination b y seawater i n o i l Silicon ii'om contamination by silica from induction systems or cleaning fluids Sodium from anti-corrosiol] agents in engine cooling solutions usually as a -
result of a coolant leak. Corrosion: the fol lowing elements are associated with corrosion
- Aluminium from engine block corrosion Iron fro m corrosion in storage tanks and piping - Manganese sometimes found along with iron as a result of corrosion of steel
Copyrighted Material
Cm"/ilio,,, ",,,,,il()m/; Wc:" melals such as "on, ,llImi""I1I. chr""II"I1I. copl'cr. lead. tin. nic�cl. and ,ih'er oil additi,'es <'{)Iuaining boron. zinc. pho'poom",. c'al cium. magnesium. or barium; cxtmrl<'OuS l"OOl:nTllnanl, such as ,ilicon: c"OlT!JSion Al'l'liclIIio"s: OIls u",d In dic",1 mw gasoline "nglnc'. gas IlIrbm",. tTml,mi,· sio",. gearl","",. c'''''pre",,,, and hydraulic 'y,lem'
Op<"",li"", AI: ncite, tn. ",,"' n"'tal clemen!, in the ,ampic hy .....i'ing Iheir alol1lic energy Slalcs in a IlIgh ,"ullage ( 15kV) lelll]lC"" ure ."'u....·c. The clements arc '''lOmized' and emil 'heill'h"...""" cri,'ic radiation Thc rc'"Itan! light enefgy
1>-1'_'C' thmugh a 'lit to a diffraction grallng which "'parJIC, lhe indi' idual emi.' siOll lil1<" foreach ekl11<'nl. n..... emi-sion inlensity at" chamctcri\lJc wavelenglh of an clen",m is proponional lo Ihe ,'oncen""';On of Ihe clemen! in the 'ample, A photomultiplier dck'Clor mca,ure, the intcn,ily of ca,:h emi"ion and transfe...
In. values 10 a re:odoul
' Ihe "mple: � ,uitabl)' !r�iflCd "'mi·,�illcd "·or�cr. To "p"rai<' Ihe spt.'Ctmmcter: a suitably tr�il1<>d laboratory Ic-chnjcian. To analFe the Ie," re.'"It.<: an �.\]lCricnced chemical IlImly.,t Ad"""w/:e.: Can perfoml ""quenlial or
10
"ap
Can"'� determine the t)P" of wear proce" thai may he "ccurring.
5.2 AE· Kota!i,,!: nisk Ek..., tmde Cm"li'm".- """"IOud: Tr.u:c ie,'ct, of wcar metal-. c,t""",ou, con,"minan!.,. and :l/iralio"s: En<:io",d lubri'''ting ,y" ell" in die",1 and ga",IIn" engi"," . g:" turbil"" . Iron,n""ion•. gcarbous. comlll"C"o"S, , and hydraulic 'y,'enlS I'·F illlen·"I.· Usually ""'e.....1 wc..,ks It, months
O'�'r(Jrio,,: A mlal ing graphile di,\; i, immersed into a """pic "c,,,,1 piding up " '111,11 ,"''''pie of oil. grc",," or fud as il Ium,. The' 'ampic IS inlrodlK'eJ inlo:1 high tcmp<:m!ure d""lric an: created in Ihe gap hetwecn the disc' cll'Ctrodc and a rOO cOlmler ck"lruJc. The 'alllple " complcldy \'ulalilc>d. Creal IIlg a pla"na which emit, lighl charaCleristic of the elements In the sample.l"h< emis,ion lines ofeach clemo," are tnea,ured by an optical 'ystem, alld troe re,ult, are displayed on a
CRT and a prinler in Ihe I",n, per ",illion (ppm) range
Copyrighted Material
Copyrighted Material
Ap{'('IIdi... -I: COII/Ii,ioll /HollilOrin!: T�c1l11i'l"es
373
Skill: To dr�'" the
nician. To analyze tllC Icst TV'ultS: a suitably Irai""d laboralory lechnieian Admntage.<: Simple 10 oper.IIC. No pn',smnplc r><"paralioo. Analy,i' lake, alloul
30 seconds, Equipment p<Jnablc, Up!O 32 elemem, can be analy>.<:,1 at the S
panicles alx",e 5 microns. 5.3 AE · [nducli>'CIy Coupled I'[a�ma (ICI') C"",/iliom mOllilOft'd: Wear metal, from 1110' ing pam ("",h as iron, alu111mum,
.hrornium. copper, lead, lin, nickel, and sihcr): oil additives rontaining hol'OO. >,inc. phospiloroos, calt-ium, magne,ium or bariunr e:
sion.<. gearlxJxe.<, rompre<ts and hydr.wlic 'y'tems {'·F ;nun'a!: U<"ally ",,'cral ,,'eels to months O"",alior!" Argon gas i, pa<,;,.-.I through a rdllio frequency induction coil and
heated (0 a Ic"'per�turc of 8.())() K to 10,000 K pn)(!ucing a pia",,;', The oil s�mple i. diluted by a low viscosity ,olvenl su"h as ,yicnc or kcr operale lite
SpectTOl'T\t:tcr: a suitably trained technician, To analyze "" ull" an experienced lechnician Ad,''''u"S�s; More accurate. reliable and repeatable than Ihe rmary electrode
mclhod, A large dynamic range permits singic ellli<sion lines to 0. u."",d for the TIlCasurement of a fnl\ge of ,:one'en!"'tio" Ic"cls, rr""ide, pam per billion (ppb) ",n.itivily for compounds ,ueh as metal-organics and wear metal particle, Ie" tilan 3 micmns in size, FaS! and easy to oper.J1e. No ne<:d for oper�t(Jf to dilule samples ma"ually prior to ""aly.,is. Automaled opemtion Dis"d",mtuge.: ICI' spectn)me!er is more complc.\ 'md c'pensiw. "nd has a
higher opemting co
Copyrighted Material
374
Appendix 4: Conditioll MOllituring
Reliabilitv-centred Maintenallce
5,4 Atomic
Spectroscopy
Conditiol1S mOllitored: Wear metals (such as iron, aluminium, chromium, lead, tin, copper. nickel and silver): oil additives containing boron, phosphorous, zinc, calcium, magnesi u m , or bariu m : contaminants such as s i licon: corrosion Applicatiolls: Oils used in diesel and gasoline engines, gas turb ines, t ransmis s ions, gearboxes, compressors and hydraulic systems, P-F illfi:'J'\'aI: Usually several weeks to months
Operatioll: Works on the principle that every atom absorbs l ight o f a speeific wavelength . The oil sample is d i l uted and burned i n an acetylene flame or other atomizer hot enough to dissociate the sample i n to i ts constituen t atoms, The name is i rradiated by a hollow eathode lamp at the characteri stie wavelength of the desired metal, The higher the concentration o f the metal, the higher the absorption of the l ight. The of absorption is measured and cOllverted i nto ppm val ues for that metal by a readout computer. Graphite furnace spectrometer uses an electricall y heated hol l ow cyl i nder to contain the samplc and can b e used for ultra low trace wear metal levels, This can i nerease measurement sens i ti v i ty from 1 0 0 to 1 000 times over the acetylene flame method, Skill: '1'0 draw the sample: a suitably trained semi-sk i l l ed worker. To operate the spectrometer: a suitably trained laboratory technician, To analyse the results: an experienced chcmical analyst Popular with s malier oil analysis faei l i tics for determin i n g wear metal concentrations i n used oil analysis, High accuraey, precision and repeata b i l i ty at low cost. AA does not suffer from spectral i n terference,
Disadvantages: Samples require preparation, Analysis t i me is longer. Requires a flammable gas, May fai l to vaporise particles above 5 microns, 5.5 X -Ray Fluorescence Spectroscopy
Cont/itiol/S ITWllito/,i:'d: Wear metals such as iron, aluminium, chromium, lead, tin, copper, llickel and s i lver: oil additives containing boron, phosphorous, zinc, caJ-cium, magnesiuIll , or barium: contaminants such as s i licon: corrosion Applications: Oils used in diesel and gasoline engines, gas turbi nes, transmiscompressors and hydraulic systems, sions, P-F interval: Usually several months
Operation: An oil sample is exposed to a high energy X -Ray source which raises the energy level o f the atoms in the sample, This causes the eontaminants to emit eharacteristic secondary X-Ray energy, except that the radiation measured i s the characteristic florescence of the chemical e lements in the sample which i s converted into their respcctive elemental data b y a multi -channel signal analyser.
Skill: To draw sample: a suitably trained semi-skil lt'd workcr To llperate thl' equipment: a suitably trained technician, ru inlcrpret the re:,lt l t s : an engineer Advantages: Good accuracy, precision and repeatability, Current softwarc has simpli fied its operatio n and data i nterpretation, Covers a widcr rangc uf dlt:!ll ical elements than AA or AE. Can see any particlt' size Disaril'(ll1tages: Requires a cryogenic cooled detector detection l im its, Longer analysis time, The analy,i� of h igher X-Ray encrgies and hence i ncreased precau t ionary
A B OI' AA
5.6 Energy Dispersive X-Ray Spectrometry
Conditions monitored: Wear metals (such as iron, alulIlinium, chromiulll, lead, till, copper, nickel and silver): oil additives containing boroIl, line, cal cium, magnesium, or barium: contaminants such as silicon: COlTosiun Applications: Oils used in diesel and gasoline gas turbincs, transmiss iems, gearboxes, compressors and hydraulic systelW>, P-F
interval: Usually several lllunths
Operatiol1: An energy dispersive speetrometer (EDS ) attachment to a e l ectron microscope (SEM) permits thc detection uf t he produced impact of the electron beam on a sample, thereby allowing qual itat i ve and quan titative analysis, The electron beam of the SEM is lIscd to excite the atoms in [he surface of a solid, These excited atoms produee characteristic \",hich are readily detected. By uti l i si n g the seanning feature of the SEM, a spatial distribu tion of the elements can be obtained, Skill: To draw sample: a suitably trained semi-ski l led worker. To do the test: a suitably trained technician, To i nterpret the results: an experienced Advantages: Rapid identification of particles: V cry fast e 1 e lllental line scans
and
Disadvantages: Not an on-line technique: Requires equipment: High degree of skill to interpret the results. 5.7 Dielectdc StI'ength (ASTM D-877 lind D· 1 8 1 6 )
Conditiolls monitored: The abili ty of i nsulating oil [0 \\ i t hst and eledril' stress caused by conductive contaminants such as metallic fibre:" or free wa[er Applications: J nsulating oils in transformers, breakers and cable" P-F
inti:'lwt/: S everal months
376
Appendix 4: Condition MOllitoring
Reliability-celltred Maintenal1ce . The
container is invcrted and swirlcd several limes before
filling the test cup. The test cup is fil led to the top of brass electrodes aud an increasing voltage applied at a rate of 3 k Vis (D-877) or 5 k V/s ( D- 1 8 1 6) to two electrodes spaced 2,54 mIll ( D-877), 2 mm ( D- 1 8 1 6) apart, until breakdown occurs. This value is recorded and trended. Five breakdowns are made with one CllP filling at one m inute intervals . The average of the five breakdowns is con
sidered the dielectric breakdown voltage of the sample. H i gh and medium volt
5.9 DIAL (Differential A bsorption LIDAR)
COllditions lI1onitored: The chemical eomposi tion and
of gases ill the
atmosphere
Applications: Gases emi tted by smokestaeks and leaks
tanks or
P-F interVLlI: M inutes to months, depending un t he
age transformers should observe the fol lowing l i mit, > 25k V for i n service oiL
Operation. Si m i l ar to Ll D A R (see 4. I 5 above), exeept that two differential
>30k V for new oi I . D-877 test used for rated voltages below 230 k V, D- I 8 1 6 test
wavelengths are used. One wavelength is set to correspond to a
llsed for voltages rated above 2 3 0 kV.
wavelengtlt is absorbed and the other reflected. The quantity of gas present is
Skill: Taking lhe sample: all electrician. To conduct the test: a suitably trained
determined by measuring the amount of l ight refledcd. The lueation uf the ean be determined by triangulation based on read i ngs taken fWIll t wo
l aboratory technician.
Advilntages: Test quick and simple. Transformer does not have to be taken off
so oIle
Skill: An e x perienced engineer
line to draw sample. A good overall i ndicator of transformer condi tion
Adl'{lIltages: Can eover large areas
Disadvantages: Test results dependent on sampling technique. Test sensitive to
Disadl'Cllltages: Must be calibrated for indi vidual gases:
ambient temperature and h umidity. Some risk involved i n handling PCBs. Uses
alld unlikely to be eeonomie for a s i ngle site: Operating the equipment rt�quires a
hazardous materials and equipment. Not an on-line technique.
level of ski l l .
5.S Interfacial Tension (ASTM D-971 )
A Preliminary Note o n t h e Chemical Measurement o f Fluid
Conditions lIlonitored: Presenee o f hydrophilic compounds ( a compound soluble
The techniques described in this section of part 5 arc u,ed lu detect incipient failure of the fluids themselves. They apply to fuels.
i n water or which attracts water to its surfaee )
Applications: Petroleum based insulating oils i n transformers, breakers and cables.
Operatioll. I nterfacial tension i s determined by measuring the foree needed to of platinum wire from the i nterface between a sample of o i l
a n d distilled water. After zeroing t h e device (known a s a tensiometer), the plati n u m ring is i m mersed i n the water to a depth of 5 mm. A filtered oil sample is poured on the water to a depth of 1 0 mm. The o i l-water interface is aged for about 3 0 seconds, then the eontainer i s lowered until the film ruptures. The i nter faeial tension is then ea1eulated. H igh and medium voltages transformers should not exceed >27 dyneslcm for in-service oil and > 40 dynesJcm for new oil
Skill: Taking t he oil sample an eleetrician. To conduct the test, a suitably trained laboratory technician Reliable indieation of compounds soluble in water. Test takes about I mi nute. Transformer does not have to taken otnine to monitor insu l ating oil
Di.lodl'{li1f([gl's: Test dependent on sampling technique. Hazardous and tlam mabIe materials are lIsed to conduct the test. Not an on-line technique laboratory equipment.
eondition of additives (although some also detect contaminants). The elements most commonly deteeted by these techniques are listed below.
P-F interval: Months detach a planar
oils and/or
They are used mainly to analyse the properties of the base fluid and/or the
requires
- Antimony from grease compounds. - A I'senic from anti-corrosion or biocide agents - Barium from detergent, dispersant and anti-oxidant additives for fuels and oils. Boron from anti-corrosion additive for engine coolants and as an anti-knoek o
agent i n fuels.
- Calcium from detergent <md/or dispersant additives Chromium from an anti-oxidant i n jet fuels Cobalt from natural trace levels in crude oils Copper from natural traee levels i n crude oils and lubrieant additi\"c�s - Iron from natural traee levels in crude oils - Lead from anti-wear additive i n some lubricants. some t i mes added to fuel as o
anti-knock agent
- Magnesium frum detergent andlor dispersant add i t i ves - Molybdenum from natural t raee levels i n crucie oj Is and
an anli - fricti oll
additive in some l ubrieants
- N ickel from natural traee levels i n crude oils vanadium
in
wllil
37R
Reliability- centred Maintenance
Phosphorus from natural trace levels in crude oils and as all anti-wear additive in sOllle l uhricants Potassium from natural trace levels in crude oils Selenium from natural trace levels in some crude oils and coal. Silicon from anti-foaming agent in some oils Sodium from natural trace levels in crude oils and seawater - Sulphur from natural trace levels in crude o i l and some fuels. Used as an anti corrosion agent in gear l ubricants and as anti-oxidants i n l ubricating oils. Vanadium from natural trace levels in some crude oils Zinc found naturally in some crude oils. Found as an anti-wear additive in automotive lubricants and as an anti-oxidant in marine lubricants. 5. 1 0 Fourier Transform Infrared (FT-IR) Spectroscopy
COllditions monitored: Deterioration, oxidation, water content and depletion of anti-wear additives in mineral oils and synthetic l ubricants
Applicatiolls: Lubricating oils from combustion engines, hydraul i c systems, etc P-F
intervul: Usually several weeks to months
Operation: Like atomic absorption spectroscopy, FT-IR measures absorbent l i ght enagy at spec i fic wavelengths to determine the level of the e lements in a sample, Uses a l o w power broadband i n frared beam converted into a uniform pattern of constructive and destructive interference hy a Michelson interfero meter. The interference pattern is passed through a sample where it is altered by the charac teristic absorbance levels of the elements o f the oil and eontaminants. The altered i nterference pattern enters a detector where it is converted into an audible frequency electronic signal, then converted into i ndi vidual wavelength/ amplitude data by a Fourier transform. The absorbance of the oil, additives and contaminants at thei r respective wavelengths is measured, generating a scalar spectrum, often called a ' fi n gerprint' . The sample fingerprint is compared w i th an unllsed o i l sample fingerprint using inte l l i gent software
Skill: To draw the sample : a suitably trained semi-ski l le d worker. To operate the spectrometer: a suitably trained laboratory technician. To analyse the test results: an experienced chemical analyst Does not usc dangerous chemicals, Lower energy levels do not alter the molecular structure of the compounds in the sample, u n like A A . Data can be converted intu ASTM equivalent parameters. Good repeatability. Total acid num ber (TAN) or total base number ( TB N ) data can be synthesized from FT-IR data
Disadvalltages: Uses l1annnable solvent for cleaning. D ifferent manufacturers of FT-I R equipment lise different data extraction algorithms for oil condition parametcrs and contaminants. Only sensitive to I 000 ppm water contaminatiol1.
Appendix 4: Conditio/l Monitoring
379
5 . 1 1 Infrared Spectroscopy
COllditiollS lIIonitorcd: The presence of gases such as fluoride, n i trogen, methane, carbon monoxide and
Applications: As for gas chromatography P-F interval: Highly dependent un the appl ication
Operation: The atoms of a molecule vibrate about their equilibrium w i th different but precisely determinable frequencies, A beam of i n frared l ight ahsorbs these characteristic bands, plotted again s t wavelength. the infrared spect ru m . The positiou of the absorption points on the wavelength scale is a ljlla l i tat ive cfwrac teristic and concl usions can be drawn fmm the i n tensity of the absorptiun bands Skill: To operate a preset i n fran'd spectrometer: a traillcd
assistan t . To i n terpret and evaluate the results : an experienced laboratory technician
A dvantages: Rapid analysis: H i gh sensitivity: Can be
laboratory assistant when equi pment is preset: Graphs provide a permanent renml
Disadvalltagcs: Considerable experience and ski l l neeckd t o
res u l t s : Laboratory based equipment: W i d e range of appl ications rc'quired t o j nstify [he cost of the equipment. 5.12 Gas Chromatography
Conditiorls monitored: Gases emitted as a resul t of faults. I'here are over 200 gascs present in e lectrical insulating oils of which nine are of interest. In order of criticality, these are nitrogen, oxygen, carbon dioxide xide (CO), methane, ethane, ethylene, hydrogen and acetylene. amounts otTO and CO, indicate overheating in the windings: CO. CO, and methane i ml i cate hot spots in t h e i nsulation: hydrogen, ethane a n d rnethLlne indicate corona discharge; methane is a sign of i nternal arcing Applicatiolls: Nuelear power systems, turbine generators, sulphur llexatlouride or n i trogen sealed systems, transformer oils, breakers etc P-F interval: Highly variable depending on the nat ure of the fault
Operatiol1: A gas sample is i njected through a silicolll' rubber septum port maintained at a temperature higher than the boiling point of the least volatile element in the sample. A carrier gas ( usually an inert gas such as helium, argon or n i trogen) sweeps the vaporised sampl e out of the port and into a column located in a thermostatically controlled oven. Elements witl! a wide range of boiling points are separated starting at a loyv oven telllperature and the temperature over time to clute the high temperature elemcnts. The
Reliability-centred Maintenance column contains absorbent materials such as diatomaceou, earth to separate the gases. Gases emerging from the column flow over a detector which can be direc ted into a mass spcctrometer or Fourier tranSf0I111 infrared spectrometer to record the spectrum as eluted fiT)m the column. D i fferent detectors are used for diflerent separation applications
Skill: Taking the sample: an electrician. To conduct the test: a suitably trained lahoratory techmcian. To trend and analyse the results: an electrical engineer H igh sensitivity dctection (one part i n 1000 m i l lion, b y volume): Once the equipment has been set up, i t can be operated by a laboratory assistant Adequate samples for sensitive analyses are diffieu l t to obtain: In
systems any fault gases Illay be rapidly d i l uted: Considerable skill
needed to i nterpret the results: Equipment i s not portable: Wide range of appli cations required to j usti fy purchase: Not widely used i n maintenance.
Appendix 4: Cundition Monitoring
38 1
P-F interval: Months
Operatioll : A thin layer o f atoms in the surface of thc material to be monitored is made radioactive by bombarding it with a beam
particles. Monitor
ing systems are calibrated to take radioactive decay i n to aL�C(llmt. Material losses
of u p to I �lm can be measured up to four years after activation
Skill: To take readings: a suitably trained semi-ski lled worker Advalltages: Wear can be measured duri ng normal plant
even with
substantial intervening material Components have to be removed to ac:t i v atcd unless cllupons can be used: Reactivation i s required every fuur years. 5 . 1 5 Scanning Electron Microscopy (SEM)
Conditions lI1onitored: Fractured surfaces for the prescnce of llnusLw! e l e ments 5. 1 3 U ltra-violet and Visible Absorption Spectroscopy
Conditions monitored: Changes in oil properties (alkalinity, acidity, insolublcs) . Applic(I{iolls: Oils u s e d i n diesel a n d gasolinc engines, gas turbines, transmissions, P-F
compressors and hydraulic systems
illterva/: Scveral months
Operation: An oil sample i s s u bjected to intense ultraviolet l ight, usually from a hydrogen or deuteriulIl lamp, or to v i s ible light from a tungsten lamp. Ultra violet and visible l ight are energetic enough to promote the outer electrons of the samplc e lements to h i gher energy levels, causing l ight at speci fic wavelengths to bc absorbed. The absorption can be monitored using a wavelength separator snch as a prism or a grating monochromator. The amount of l ight absorbed i s related to t h c concentration of each element. Quantitative measurements c a n b e made by scanning t h e spectrum or a t a s i ngle wavclength
Skill: A trained and experienced laboratory technician Advantages: Useful [or quantitative measurements Disadmntages: The u ltraviolet and visible spectra have broad features that are of limitcd LIse [or sample identification. Considerable skill and experience needed to analyse the res u l ts . Equipment is laboratory-based and is expe n s i vc . 5.14 Thin-layer Activation
Applicatiolls: A ny s urface types, thin films and interfaces found in raw ,ellli conductors, fin i shed semiconductors, metal and steel surface\. medical de viccs, ceramics, polymers, etc P-F intervul: Application dependcnt
Op(,rCltion: A focuseci beam o f electrons i s rastercd across a
s urface.
This causes a secondary e l ectron current to be emitted from the sample which varies accordi n g to the angle o f incidence of the bt�alll onto the sample . The secondary electron intensity i s uscd to vary [he brightness o f a cathode ray tube � which i s synchronous w ith the raster scan, y ielding a topographical of the sample surface. D ifferent detectors can be used to provide other i n formatio n . For instance, a backscattered electron detcctor provides average atomic n u mber i n formation , while an auxiliary energy d ispersive X - ray dctector can identify elements such as boron and uran i u m .
Skill: Skilled laboratory technician Advantages: High resolution with little sample prcparatioll. allows lise with rough samples. Rapid qualitative analysis areas coupled with an energy dispers i ve X -ray detector
Disadvantages: More of a diagnostic tcchniquc to determine root C:l uses o f fai l ures. Samples must be coated with a conductive fil m . 5 . 1 6 Scanning Auger Electron Spectroscopy
C'OllditiUI1S lIlonitored: Wear
Conditiolls monitored: Fractured surfaces for thlc presence (if unusual elements.
Applications: Turbine blades , engine c y l i nders, shafts, bcarings, electrical con
elemental mapping o f fine particles, corrosion and oxidatit)f! sc�tles
tacts, rail s and cool ing systems
382
Reliabilitv "centred Maintenance
Appendix 4: Condition Monitoring
surf:lce types, thin films and i nterfaces found ill raw ano finished semiconductors, metal and steel surfaces, medical devices, ceramics, polymers.
P-F interval: Weeks to m()nths
P-F intl!rl'ill: Application dependent
catalytic convertor. D i rt and oil are removed by a and moist ure a water separator. Gas sensors pick up the gas concentrations and are displayed as percentages (He in parts per million). CO mean, that t ile engine i s running rich. H igh 0, indicates a lean rni�fire or an exhaust leak. CO, is at its highest at the optimum air-fuel ratio (AFR), and when til" AFR
Opc/'(/{ioll: A finely focused electron beam irradiates the sample and creates a core hole by ejecting a core e l ectron from a sample atom . The resulting ion then de-excites when an electron from an upper level fil l s the core hole and a third electron i s e m itted to conserve energy. This electron has electron thc a kinetic e nergy characteristic o f the e m itting atom, which allows elements to be ident i fied to a depth o f between 2 and 20 atomic layers
Operation: A sampling probe is i n serted i nto the exhaust
upstream o f (he
Skill: Skil led laboratory techn i cian
is too rich or too lean. High HC indicates misfires or combustion. Lambda' readings are also calculated on most Hnalysers. Lambda is the name given to the ratio o f the actual AFR over the ideal ralio o f 1 :1 . 'rhe i deal lamhda reading i s one, and leaner ratios are greater than one
AdFClI1tages: S E M capab i lities are usually i ncorporated into A uger instruments.
Skill: Trained and experienced automotive l1lechanic
S urface sensitive. Elemental mapping. Rapid analysis
Advantages: Pinpoints cmissio n fai l ures. Portable
Distllil'wltages: More of a d i agnostic technique to determine root causes o f fail ures. Laboratory technique.
Disadvantages: Equipment needs to be taken off- line to C(ll1lll'ct to tht' 5.] 9 Colour indicator Titratioll
5.1 7 Electro-chemical Corrosion Monitoring
Conditions monitored: Corrosion of material embedded in concrete Applicatiolls: Structural steel pylons, gantries, etc
0974)
Conditiolls mOll itored: Lubricant deterioration by acidity and alkalinity in an oj I sample Applicatiolls: Oils used i n diesel and gasolinc
P-F illlervill: Months
sions, gearboxes. compressors and hydraulic systems
Operatioll: S ll1all currents arc passed between the structure and a probe i n serted
P-F il/terml: Weeks to months
in the ground nearby. Thcse eurrents affect the potential of the structure at any point wherc corrosion i s taking place. The changes in the potential are measured by a halfcel l in contact with the ground and close to thc structure. The degree ofeorrosion is directly related to the eurrent required to displace the leg potential. High currents indicate the need for a physical i nspection
Skill: A s u i tably trained technician Advantages: Structures do not have to bc excavated for inspection u n less this
Skill: Laboratory teehnician
Disadvantages: Docs not measure t h e e x tent or precise location of corrosion:
A dvantages: Test accurate to within 1 5 Cfr Disadvantages: Can only be used for petroleum based oils. Poisonolls, n alll � mable, corrosive chemicals lIscd in the test. Cannot be used for dark o i b .
5.18 Exhaust Emission A nalysers (Four-gas Analysis)
Conditions monitored: Combustion efficiency by measuring the concentrations of oxygen (0,), carbon monoxide ( C O ) , ear'bon diox ide (CO,) and hydrocarbons ( HC ) in exha il s t emissions. Exhaust leaks �
Applications: I n ternal combustion engines
turbi ncs. traw,mis "
Opcration: The sample is dissol ved i nto a mixture of toluenc, alcohol and water and titrated w i th an a lcoho l ic 'base or acid solut i o ll. to the end point indicated by a colour dlange of the added naphtholhenlcin solution. The acidity or alkalinity is expressed as milligrams of potassium hydroxide needed to neutra lise a gram of o i l . The higher the acid or base number the greater the o i l deteri oration. High and medium voltages transfo rmers should be O.SmgKOII/grn for new o i l and < 0 . 1 mgKOH/gm for in service oil
technique reveals a real need t o do s o Ground must be moist.
the l e v e l , of
S.20 Potentiometrk Titration TANrrUN
Conditions II/ollitored: Lubricant deterioration by ity of an oil sample
3 84
Reliability-centred Maintenance
Appendix 4: Conditioll Monitoring
385
Appiir.'aliofls: Oils used in diesel and gasoline engines, gas turbines, transmis
Operation: A thoroughly mixed sample is poured into a c lean beaker and heated
ers. sic)f]s, gearbox es, compres sors, hydraul i c systems and transform
to 2° below desired test temperature. The cell i s removed from the test chamber and fil led w ith the heated sample. The inner electrode i s inserted into the cel l to
P-F illterl'(ll: Weeks to months
gether with a mercury thermometer The electrical connections are then made to the cell. The sample i s then electrically stressed by a throUQJl the cell and the power factor i s caleulated. For high and medi lUll transfo�mers the power factor limit should be less than I at 25"C
l alcohol Operation: The sample is dissolved i nto a m ixture of toluene, isopropy ed by determin is acidity The e. hydroxid potassium and waler titrated with alcoholic e is hydroxid m potassiu the as ivity conduct ectrical l e in change measurin g the the number, acid added. The val ue is expresse d as mgKOH/g. The h i gher the greater the breakdo wn of the o i l .
Skill: Laboratory technician Advilntages: Can be lIsed for o i l s that are too dark to use a colour change indi cator. Test accur<Jte to w ithin 4(1
5,21 Potentiometric Titration TBN (ASTM D2896) inity. Conditions monitored: Lubrica nt deterior ation by measuri ng alkal , transmi s · ApplicatiollS: Oils used ill diesel and gasoline engines, gas turbines mers. sions. gearbox t's, compre ssors, hydraul i c systems and transfor
P-F llllerval: Weeks to months which is Operation: The sample is dissolve d i n to a mixture of titration solvent
ivity) readings titrated with perchlor ic acid. The potentio metric (electric al eonduct The alkalinit y (base . solution trating i t of volumes ve respecti against plotted art' to titrate the solution nllmber ) i s calculat ed Ii'om the quantity o f acid needed nt (mgKO H/g). equivale gram per de express ed i n milligrams of potassiu m hydroxi formed during acids e corrosiv e neutralis to The test is it measure of an o i l ' s ability use. ed continu for lity i suitab operatio n, indieati ng its
Skill: Laboratory technic ian Advantages: Can be llsed regardless of colour of oil. Accurate to within 1 5°/r). Disadvantages: Can only be used for petroleum-based oils. Dangerous chem iutls used in the test.
Skill: Taking the sample: an electrician. To conduet the test: a suitably trained l aboratory technician simple. Transformer does not have to be Test q uick and taken offline to monitor the insulating fluid
Disadvantages: Uses hazardous materials and equipment. Ttcst must be conduc· ted in a laboratory and depends on sampling A Preliminary Note on IVIoisture Monitoring Water in oil rapidly reduces machinery and component l i fe . For insl:lllce, it call reduce roller element bearing l i fe by as llllIch as 1 00 limes. It also interferes seriously wth the l ubricating properties o f oil for instance, a drop 01 water in 5 htres of o i l at 85°C totally destroys l.inc 'lIl ti ·wear additives. Water affects the o i l itsel f in the fol lowing ways: it increases oxidation, and i n so doing forms s l i mes and resins it increases conducti v ity, which i s especially undesirable in trans/(mner o i l it reacts with anti-oxidants to form acids a n d precipi tate salts i t reacts with zinc eli-alkyl di-thio phosphate (ZDDP) anti-wear additi ves to form hydrogen sulphide and sulphuric acid i t promotes the growth o f microbes - it changes the v i scosity of the o i l i t degrades viscosity improvers Water also affects other aspects of the system as foll o w s : : - i t rusts a n d corrodes metal surfaces i t j ams valves b y forming icc crystals - it increases wear it gums up valves and orifices it shortens the l i fe of filters it t'ntrains more air, which affects bul k modu lus.
5.22 Power Factor (ASTM D- 924) caused by Conditiolls monitored: Dielectr ic losses i n electrica l insulatin g oils
5.23 Karl Fischer Titration Test
contami nation and o i l detenor ation
Conditions /Ilonitored: Water in oil
, and cables Applications: Petroleu m based insulatin g oils in transformers, breakers
Applications: Enclosed oil systems such as
P·F interval: Several weeks
D- 1 744)
compressors, hydraul ic systems, turbines. transformers. elc.
transmissions.
386
Appelldi.\ 4: COllditioll Monitoring Techniques
Reliability-centred Maintenance
P-F interval: Days to weeks
5.25 Crackle Test (Human senses)
Operatioll: A measured sample is reacted with a Karl Fischer reagent which
Conditiolls monitored: Water in o i l
contains iodine. When iodine is present, current will pass between two platinum electrodes. M o isture entrained i n the sample reacts with the iodine, perpetuatin g t h e test as l o n g as water which h a s n o t reacted w i t h the iodine rem a i n s . Once depleted, the e lectrodes are depolmlsed by the iodine and the test is complete. The cOiTesponding potentiometric change is used to determine the titration end point and calculate the water concentration. The duration of the test indicates the water content. High and medium voltage transfonners should not exceed 25ppm at 200e
Skill: Laboratory technician Advantages: Accurate for small quantities of water (parts per million ) . Accuracy within 1 0%. Test is relatively fast.
Disadvantages: Adequate samples for sensitive analyses are difficult to obta i n :
App!ic(ltir7l1s: Oils L1sed i n diesel and
t urbines, transmissions. gearboxes, eompressors. hydraulie systems and transformers.
P-F
illterv({/: Days to weeks
Operation: A few drops of o i l are placed o n a hot plate ( about 250°F) If water i s present i t quickly vaporises and makes a crack ling or popping sound. Ski!!: A trainee! semi-skilled worker Advantages: Cheap, quick and easy to use. Effective and cconomical Disadvantages: Moisture under 300 400 ppm cannot be
thc amuun t of water present. Require s a quiet area to hear the cr:tckk's. Danger o f hand I ing oil around a hot surface.
I n l arge systems any fault gases may b e rapidly diluted: Considerable s k i l l needed to interpret results: Equipment not portable: Wide range of applications required to j ustify purchase: Not widely used in the maintenance environment.
5.26 Crackle Tt'st ( A udio detector)
5.24 Moisture Monitor ( Vapour Induced Scintillation)
Conditions lIlonitored: W ater in oil
Conditiolls monitored: Water in o i l
Applications: Oils used i n diesel and gasoline
Applications: O i l s u s e d i n diesel a n d gasoline engines, g a s turbines, trans missiems,
compressors, hydraulic systems and transformers.
P-F interval: Several weeks
Operation: A probe with a m i niature heating element i s submerged i nto a n o i l sample. D u r i n g t h e t e s t t h e heating element g l o w s at a constant temperature, causing slIspended moisture in the sample to vaporise and emit a distinctive acoustic signal known as crackling. A microphone mOllnted near the heating element picks u p this s ignal and electronically passes i t to the data collector for analysis. The algorithm in the data collector is calibrated to convel1 signal per unit time into moisture levels in ppm or percentage. The unit i s able to detect suspended moisture to as low as 25 ppm and as h i gh as 1 0,oon ppm. A typical test takes 3 0 seconds
threshold
Skill: A trained semi-skilled worker Advantar;es. No sample preparation needed. Quick and easy. Detects a wiele range of concentrations. Requires only 70 m i l l i l i tres of fluid to conduct the test. Contains n o moving parts. Not affected by fluid' s viscosity, colour, density, contamination, conductivity , or flow. Portable.
Disadvantages: Equipment expensive.
heard crackling.
Test subjective, from test-to ·lest and llser-to ,user. Docs not
gas t urbines. transmis-
siems, gearboxes, compressors, hydraulic systems and transformers. P-F
interval: Weeks
Operatioll. A m icrophone is mounted adjacent to it hot plate (heating clement). A few drops o f o i l are placed on a hot p l ate (about
If water i s present it
quickly vaporises and makes a crackling or popping sound. The microphone picks up the sound. converts it into a n e l ectronic signal and passes it to a data collector for analysis. The algorithm i n the data collector i s calibrated to convert signal threshold crossings per unit time into moisture le\'els in ppm or percentage.
Skill: A trained and experienced technician Advantages: Can detect moisture levels as low as 25 ppm and high as I () ()()() ppm. Test takes 30 seconds. Easy to u se .
Disadvantages: Danger of handling oil around a h o t surface.
test.
5.27 Clear and Bright Test
Conditiolls monitored: Water in oil Applicatiolls: O i l s uscd i n diesel and gasoline engincs. gas turhincs. (mll s m i s sions, gearboxes, compressors, hydraillic systcms a n d transformers.
388
389
Appendix 4: COllditioll Monitoring
Reliability-centred Maintenance
P-F interval: Several days
6.2 Electrostatic Fluurescent Pt'lIctnUit
Operation: As moisture becomes entrained in oil, the o i l becomes visibly hazy in other words, it is no longer clear and bright. However, note that some oils can dissolve significant amounts of water (depending on v iscosity and the additive package) and still remain clear and bright. lt is only when the oil reaches a more advanced stage of emulsification, where the oil and water eombine (not mix), that i t is no longer elear and bright.
Conditions monitored and applications: As for l iquid
penetrants
P-F interval: S l ightly longer than l iquid dye penetrants
Operatioll: As for liquid penetrant dyes, except that e l ectrostatic pola. rity must be induced between the workpiece and testing materials Skill: As for l iquid dye penetrants
Skill: Experienced technician . No test equipment required. Cheap, quick, simple and economical .
Admlltages: The polarity ensures more complete and even trant and developer than with ordi nary penetrants, which
Disadvantages: Oil colour can bring error i n to the test. Subjective.
Disac/l'alltilges: As for ordinary fluorescent penetrants.
6 Physical Effects Monitoring
6.3 Magnetic Particle Inspection
6.1 Liqnid Dye Penetrants
COllditions lIIonitored: S urface and near-surface crack s and discontinuities caused by fatigue, wear, laminations, inclusions, surface shrinkage. grinding. heat treat and corrosion strt'ss. ment, hydrogen embrittlemenl. laps, seams. corrosioll
Conditions monitored: S urface discontinuities or cracks due to fatigue, wear, surface shrinkage, grinding, heat-treatment, corrosion fatigue, corrosion stress and hydrogen embritllemen t. Applications: Ferrous and non-ferrous materials such as welds, machined sur faces, steel structures. shafts. boilers, plastic structures, compressor receivers P-F illterval: Several days to several months, depending on the application
Operatio1l: The l iquid penetrant i s applied to the test surface and sufficient time is allowed for penetration into surface discontinuities. Excess surface penetrant is removed. A devcloper is appl ied which draws the penetrant from the discon t inuity to the test surface. where it is i nterpreted and evaluated. Liquid penetrants are categorised according to the type of dye (visible dye, fluorescent or dual sen sitivity penetrants) and the processing required to remove them from the test surface ( water washable. post emulsified or solvent removed).
Applicatiol1s: Ferromagnetic metab such as compressor receivl'rs, welds, ma chined surfaces, shafts. steel structures, boilers. etc, P-F il1terval: Days to months depending Oll the application,
Operation: A test piece is magnetised and thell sprayed with a solution cuntain ing very fine iron particles over t h e area t o be inspected. I f a tT
technician
Advantages: Reliable and sensitive: Very widely used.
Skill: To apply penetrant: suitably trained semi-skilled worker. I nterpretatiun: suitably experienced technician
DisadwlIltagcs: Detects only surface and near-surface cracks: Time l'onsurning: ' ('ontaminates c lean surfaces: Not an on-line monitoring technique.
Advantages: Visible dye penetrant kits are very cheap (but the more expensive fluorescent kits are far more sensitive): Detects surface discontinuities on non ferrous materials.
6.4 Strippable Magnetic Film
D isadvLlnfages: Fluorescent penetrants require a darkened area for inspection: Highly qualified personnel required to evaluate resul ts : Not an o n- l i ne monitor technique: Monitors surface-breaking defects only: Cannot lest materials w i th very porous surfaces.
Conditions monitored: S urface discontinuities and crack, caused by wear, surface shrinkage. grinding, heat treatment. hydrogen embriltlement. lami .. nations. corrosion fatigue, corrosion stress, laps and seams, Applications: Ferromagnetic metals sudl as compressor receivers. welds, ma chined surfaces. shafts, gears. steel structures. boilers. etc. 1'.1' inrerml:
Several weeks to months .
390
Reliability"centred Maintenance . A self curing s i l icon rubber sol ution containing fine i ro n oxide par
tides is poured into or onto the area under i n spection and a magnetic field induced by a magnet. The magnetic partides in the solution migrate to cracks under the i nfluence of the magnetic field. After curing, the rubber is removed as a pJu g from holes or as a coating from surfaces. Cracks appear on the cured rubber as i n tense black l ines. I nvestigation o f small cracks may need a microscope.
Appendix 4"" Condition Monitoring 6.7 Ultrasonics - Resonance Technique
Conditions monitored, applicatiull ulld P "F illle n'a!: As for pulse echo tech nique. (Also used for testing the bond strength between thin P-F imerm/: A s for pulse echo technique.
Operation: A transmitter is moved over the test surface and the
experienced technician.
Resonance in the absence of discontinuities the trallsmitted Discontinuities CaLISe the transmitted signal to fade or
A dvantages: Can be used o n areas with li mited visual access: Provides a record.
Skill, advantages, and disadl'Lllltages: As for pnlse echo techniqne.
Skill: Application o f solution: s u i tably trained semi-ski l l ed worker. Evaluation:
Disadvantages: Detects only s u rface cracks: Not an on-line technique.
39 1
6.8 Ultrasonics
Frequency IV[odul
6.5 Ultrasonics - Pulse Echo Technique
Conditions monitored, applications IIlId VF in terl'u! : As for pulse echo"
Cunditions monitored: S u rface and subsurface discontinuities caused by fa
Operation: A transducer is used to send ultrasonic waves continuouslv at chan ging radio frequencies. Echoes return at the i nitial frequency and i n t rrupt the new changed frequency. By measuring the phase betwecn li"equcncies the location
tigue. heat treatment, inclusions, lack of penetration and gas porosity i n welds, lamination; The thickness of materials subject to wear and corrosion.
Applications: Ferrous and non-fe!T()Us materials related to welds, steel struc tures, boilers, boiler tubes , plastic structures, shafts, compressor recei vers, etc" P· F
interval: Several weeks to several months.
Operation: A transmitter sends an u ltrasonie pulse to the test surface. A recei ver amplifier feeds the return pulse to an oscilloscope. The echo i s a combination of return pulses from the opposite sidc o f the workpiece and from any i ntervening discontinuity. The time elapsed between the i nitial and return signals and the relative height indicate the l ocation and severity of the discontinuity. A rough idea o f the size and shape o f the defect can be gained by triangu l ation .
Skill: A s u i tably trained and experienced technician. Advantages: Applicable to the majority of materials . Disadvantages: D ifficult to d i fferentiate types o f defects . 6.6 Ultrasonics - Transmission Technique
Conditions monitured, applications and P-F interval: As for pulse echo technique. Operation: A transmitter emits continuous waves from one transducer which are passed right through the test piece. Discontinuities reduce the amount of energy reaching thc receiver and so the i r presence can be detected.
Skill and advantages: As for pulse eeho technique. Disadvantages: As for pu l sc echo technique: Problems o f modulation associated with standing waves cause false rcadings to be obtained.
�
of the defect can be determined.
Skill, wh'anrages alld disadwlIltagl's: As for pulse echo techniqut'" 6.9 Coupon Testing
Conditions monitored: General and localised erosion and corrosioll. A.pplicatiolls: As for electrical resistance method, exc,�pt paper m i l ls" P - F interval: S everal months.
Operation: Coupons are usually produced from m i ld, low carbon steel or from a grade o f material that duplicates the wall o f a vessel or pipe. The coupons are carefu l l y prepared, weighed and measured before exposure . After the coupons have been i mmersed i n the process stream for a period o f time (several ,'leeks to several months) they are removed and checked for weight Joss and pitting. From these measurements, relative metal loss from the pipe wall can be calculated and pitting ean be estimated.
Skill: A suitably trained technician. Advantages: Very satisfactory when corrosion is steady: Usefu l where electrical devices are prohibited: Fairly cheap: Indicates corrosion t ype:
used.
Disadvantaf?es: Long duration of exposure required: Response to corrosive conditions is slow: Use of coupons i s labuur intensive: Corrosion ratc determination usual l y takes several weeks: Provides no allowance for unusual or temporary conditio n s : Coupons inadequate for pulp and paper industry .
392 6. 1 0
Appendix 4: Condition MOl1itorinR Techniljucs
Reliability-centred Mainlerwnce Current
393
6.12 X-ray RHdiogra l}hk Fluoroscopy
Conditiolls monitored: Surface and subsurface discontinuities caused by wear, fatigue and stress; detection of dimensional changes through wear. strain and
Conditiolls monitored, applic(l{ions alld P-F interl'!ll: as for
corrosion : determination of material hardness.
Operation, The transmitted radiation produces a fl uorescence of sity on the coated screen i nstead of darker patches, The
Applicatiolls: Ferrous materials used for boiler tubes, heat exchanger tubes.
is proportional to the intensity of the transmitted radiation.
hydrau lic tubing, hoist ropes, railway l ines, overhead conductors, etc.
P1 i!llcn'al: Several weeks depending on the application. OpcmlioJl: A test coil carrying alternating current at 1 00 kHz to 4 MHz induces cddy currents i n the part being inspected. Eddy currents detour around discoll tinuities, becollling compressed, delayed and weakened. The clectrical reaction on the test coil is ampl i fied and recorded on a CRT or a di rect reading meter.
Skill: A suitably trained and expericneed technieian .
intell-
Skill: As for X -ray radiography . A dl'Ulltliges: Quick results : Scanning capab i l i t y : Detects defects i n parts or structures not visual l y accessible: Most widely appl icable technique: Low cost.
Disadvantages: No record produced: (jenerally i n ferior sensitive than X - ray radiography .
qual i t y : Less
6 . 1 3 Rigid Borescopes
Advantages: Applicable to a wide range of conducting materials: Can work w i t h
Conditiuns monitored: S urface cracks and their orientation, oxide films. weld
out surface preparation. High defect detection sensitivity: strip chart recorder provides a permanent record.
defects. corrosion. wear. fatigue.
Disadvantages: Poor response from non-ferrous materials.
Applications: I n ternal v isual inspection of narrow tubes, bores and chamber, of engines, pumps, turbines, compressors, boilers. etc i n automoti ve, aircraft, power generation, chemical and related industries.
6. U
Radiography
/'-F illterml: Several weeks depending on application.
Cunditions monitored: Surface and subsurface discontinuities caused by stre,s, fat igue, incl usions, lack of penetration in welds, gas porosity, intergranular cor� rosion and stress corrosion. Semiconductor discontin u i ties such as l oose wires.
Applications: Welds, steel structures, plastic structures, metallic wear compo"" llellt, of engines, compressors, gearboxes, pumps, shafts, etc.
to bc taken.
Skill: A suitably trained and experienced teehnician. Advantages: I n spection done with clear i llumination: Parts not visible to the naked eye can be photographed and magnified.
P-F interval: Several I11n nths. Operation: A radiograph is produced b y passing X-rays or gamma rays through materials which are optical ly opaque. The absorption o f the initial X-ray depends on t h ickness, nature of the material and intensity o f the i nitial radiation. Film exposed to these rays becomes dark when it is developed
Operation: Light is channelled from an extemal l ight source cable to the boreseope. Very intense light (300 W ) enables
how dark depends o n
DisadwlIItages: Provides surface inspection only: Rcsolution limited: Lens sys tems relatively inflexible: Operators can suffer 'optic
during
inspections_
6. 1 4 Cold Light Rigid Probes
t h e amount o f radiation reach i ng it. T h e fi l m i s darkest where t h e object is thin
ConditiollS lIlonitored, applicatiol1s (flld P-F intervul: As for
nest. A crnck, inclusion or a void i s observed as a dark patch.
(also used in combustible and heat sensitive areas) .
Skill: Use of equipment: a suitably trained and skilled technician. To interpret the
Operatioll: H i g h intensity white l i g h t is channelled from a cold light uilit ( 1 50 W ) via a tlexible fibre cable into a rigid borcscope. The cOlllains a lens relay system sheathed by fibres which passes the l ight to the No
results: a highly ski lled technician or Provides a permanent record: Detects defects in parts or structures not visually accessible: Most widely applied X-ray technique.
l ight is wasted and no heat is emitted. Forward, fore- oblique,
and retro-
viewing versions of these probes are availahle. Prohe diameter� range from I
j)isadnlfltages: Sensitivity often low for crack-like defects: Two-sided aecess
mm to 1 0 111m ancl lengths frolll g cm to 1 33 Cill . Parts not visible to the naked
sOll1ctimes needed .
eye can he photographed and magnifiedorrecorded by it m i n iat url: v idcu camera .
394
Appendix 4: Condilion Monitoring
Reliability-centred Maintenance
Skill: As for rigid horescopes.
6. 1 7 Electron Fractography
Advan{{/fies: A s for rigid borescopes. No heat i s generated when cold l ight
COl1ditiollS monitored: The growth of fa tigue cracks
supply used: Detailed inspection of s u rface fin i s h in inaccessible areas can be obtained without dismantl ing: Photographs provide permanent records: Equip mcnt portabl e : With the use of the video camera/endoscope technique, inspec tion time i s reduccd to a quarter of the t i me required for direct viewing
395
Applications: Metallic components subj ected to P-F
interval: Depends on the application
Operation: Every fracture h a s its o w n 'fingerpri n ' ,. i n that the
Disadmll {{/fies: As for rigid borescope: Probe int1exible: Not an on-line technique.
fracture process i s i mprinted on the fracture surface. By study i n g a
6.1 5 Deep-Probe Endoscope
circumstanees of fai l ure
Conditiolls monitored, applicatiolls and f'-F interval: As for rigid borescopes.
Skill: Replica of the fracture s urface: suitably trained technician.
(Also llsed for the i n spection of pipework in boilers and heat exchangers)
reading: experienced engineer
of the fracture with an electron microscope, it i s possible to establish the causes and
Operation: These are speeial modular endoscopes available i n lengths of up to
Advantages: Failures can be analysed with a high
2 I rn. They are made of stainless steel and screw together to provide a viewing system which can penetrate bores with severely restricted entry. Illumination i s
caused to fracture surface when replica is made
provided by a h i gh i n tensity quartz halogen light
Skill: As for rigid borescopes Advantages alld disadvantages: As for rigid borescopes.
of thc
and
of certainty: No
Disadvantages: Eleetron m i croscope i s expensive: H igh
of specialisation required to read the results: Not an on-line monitoring technique: I naccessible
components must be dismantled.
6 . 1 8 Colour (ASTM D- 1 524)
6 . 1 6 Pan-view Fihrescopes
Conditiol1s monitored: O i l colour and condition
Conditiolls monitored, applications and P-F intf'rval: As for rigid borescopes
Applications: Petroleum based insulating oib i n transformers, breakers and cables
Operation: White light of h igh i ntensity from a cold l ight supply unit is trans
P-F interval: Weeks to months
mitted by total i n ternal retlection through a flexible fibre cable into a fibrescope. The fibrescope contains optical fibres bundled together to form t1exible l ight pipes. The fibrescope has a remotely controllable prism built i n to its tip which can be made to view forwards or sideways as required. The instrument can be i nserted using forward viewing and can be stopped to take a detailed sideways look at any pass i ng defect by s i mply rotating a control knob built i n to the s i de of the eyepiece. Adaptors can be used to take photographs or mount TV v iewers or cine cameras. An ultraviolet l ight of h igh i n tensity can also be used with t1uorescent penetration to detect m inute flaws i n inaccessible areas
Skill: As for rigid borescopes
Operation: A test tube is filled w i th the oil sample and placed next to the colour comparator. Colour is compared by revolving the colour standard disk until a 111
colour match i s made with the sample. The figure seen in the upper the front cover gives the direct reading. For high and medium
transformers
the colour limit should not exceed 3 .0 on the ASTM 0- 1 524 colour scale
Skill: A n experienced electrician Adval1tages: Provides rapid field screening of test samples for further Transformer does not have to be taken offline to monitor the insulating o i l
Disadvantilfies: Depends on sampling technique. C a n be affected
s u n l i ght..
Advantages: As for cold l ight rigid probe s : Flexibility makes more detailed in spections possible
6. 1 9 Oil Appearance
Disadvantages: Not an on-line monitoring technique: Provides surface inspec
Conditiol1S monitored: Oil oxidation, water contami nation, wear metal
tion only: Resolution l i m ited: Operators can suffer from 'optic eye' during
and particulate contamination
prolonged inspections: Ultra violet fibrescopes are expensive.
A.pplications: Lubricating o i l s
396
Reliability-centred Maintenance
P-F interval: Days to weeks
Appendix 4: Conditioll Monitoring
3')7
P·F illfl'rvai: \Veeks to months
Operation: Perhaps the simplest of all tests, appearance can provide distinct indi
OpPl'atioll: Res istance wire, foi l and semiconductor strain gauges work 011 the
cations of oil condition and contamination. Most industrial oils are gold coloured
prineiple that when an electrical conductor i s s t retched. its electrical resistance
l iquids that are bright and free of suspended solids when new. Greases, coolants
increases. By bonding the conductor to provide intimate Illechanical contact with
and fuel s also have a distinct appearance prior to use. A hazy or clouded appearanee
the sllrface u nder test, any strain on that surface will be reflected i n a
often indicates water eontamination, while gradual darkening often occurs as in
the resistance i n the strain gauge. Sensitive indicating or
o i l is oxidised in serviee, Particles as smal l as 40 microns can be seen by the
needed to monitor the strain s in most structures.
unaided eye, providing an i ndication of large particulate contam i nation.
Skill: A trained semi-skil led worker
of IS
Skill: Operation of equipment: a suitably trained technician:
of
results: a structural engineer.
Test s i mple, quick and cheap. No test equipment required.
Read i ly attached to almost any s urface.
Disadvalltages: Subjective. Particles less than 4() microns cannot be seen by the
Disadvwzlages: The strain gauge mllst be compatible with both the material
unaided eye. Particle concentration levels and souree of contamination cannot
uncler test and the e n vironment in which it is
be determined. Test dependent on sampling technique.
A Preliminary N ote on Viscosity Monitoring 6.20 Oil Odour
V i scosity is an i mportant physical property of lubricating fluids. It is essential in
Conditions /Ilonitored: Oil ox idation Applications: Lubricating oils P-F in/k'rvul: Days to weeks
Op eration: Most o i l s have a bland or llondescript odour when new and develop a more pungent or ' burned' odour as they oxid i se i n service. An unusual odolll' may indicate contamination sllch as fuel dilution. Often , the stronger the odour, the greater the level of oxidation or contamination. Thi s technique is also l i m i ted by the subjective nature of the observation. Some people have a more sensitive sense o f s mell and react differently to strong odours. I n addition, vapours can
providing critical clearances between moving anel
surfaces. I mproper
v iscosity is an early i ndicator of general lubrication fai l ure. Illay also be a warning sign o f many potential failure conditions. An increase in v iscosity can cause s l uggish valve control, pump cavitation, reduced mechani caL volumetric. and energy efficiencies and increased temperature. A decrease in viscosity lOan cause i ncreased internal and external
increased tempera
ture. excessive wear due to poor lubrication ami reduced control and precision.
6.22 Viscosity Monitor
Conditiolls monitored: V iscosity changes caused by overheating, additive fai l
collect i n closed reservoirs or storage tanks, and give o ff a strong odour when the
ure, mixed lubricants, fuel and glycol dilution, oxidation. moisture and particulate
tank is first opened. These concentrated vapours may therefore suggest higher
contamination.
level of contamination or oxidation than normally exists.
Applications: Oils used in diesel and gasoline engines. gas turbines, transmis
Skill: Experienecd semi-skil led worker
5ions, gearboxes, compressors, hydraulic systems and transformers.
Adl'(mtages: Quick, easy and cheap. No test equipment required.
P·F
Suhjective.
intavai: Several wceks to months.
Operation. A sensor i s attached directly to a portable condition monitor (peM) which controls the test sequence and displays the results. The sensor tests direc
(,. I S Strain Ganges
('onditioIlS monitored: Strain Large c i v i l structures slich as bridges, tunnels, the load bearing buildings
tly from the sample bottle and gives on- t lle ,spot
results that can be
saved into the PCM for later viewing. Results can be computer for trending and graphing. The fluid temperature is measured digital t e mperature probe.
Skill: A trained skilled technician.
398
Appendix 4: Condition Monitoring
Reliability-centred Maintenance
Advantages: Fast rel iable, on-site testing. Can be calibrated to ASTM v iscosity standards. Measures absolute viscosity directly. Kinematic viscosity can be deter m i ned by entering speci fi c gravity. Results can be d isplayed in SSU, centipoise, centistoke, or ISO viscosity grades and can be recorded at 40°C or 1 00°C.
7
399
Temperature Monitoring
A Preliminary Note o n Thermography
D isadvantages: Eqllipment very expensive.
Thermography is the measurement of the radiation emitted from the surface of an object in real time. producing a visible i mage of the i nvisible i n frared radi
6.23 Falling Hall Comparator
atio n . It is based on the pri nciple that all objects above absolute zero (- 273')F) emit i n fra-red radiation. Thermal imaging systems are electronic cameras that
Conditions monirored: Oil v i scosity. Applications: Oils used in diesel and gasoline engines, gas turbines, transmis sions, gearboxes, compressors and hydraulic systems. P-F
interval: Weeks to months.
Operation: An oil sample is compared to a reference oil. Identical ball s are allowed to fall freely through the sample and the reference oil. The time required to fal l a specific distance provides a comparison of the viscosities. One kit provides
make the radiation optically visible i n an
ofdiffcrcnt colours
These images can be recorded on conventional
scales).
or electronic media.
7.1 Infra-red Scanners
C'Ol1ditions monitored: Electrical: c u rrcnt/resistance relationships from loose_ oxidised or corroded connections or malfunction o f t he component itself.
Mechanical: heat generated from friction caused by faulty
inadequate
l u brication, misalignment, misuse, and normal wear
a direct reading of the sample o i l v iscosity, while others require calculations.
Applicariolls: Electrical. power distribution and high tension lines. transformers,
Skill: A trained laboratory tcc h n ician.
transformer bushings. capacitor bank connections. thyristor banks. disconnects.
Advantages: S i mple and easy to use. Accurate to within I 'Yo i n most cases.
relays and circuit breakers, meter and control connections_ circuit bus and fuse connections, fuse clips and stab conllections_ moulded cast' and air
Disadvantages: Oil sample needs to be tran s l ucent enough to see the ball as it
breakers, motor windings, thermal overloads. conductor
fal l s , dark or oxidised o i l s may be unsuitable. Not a field portable technique.
windings. generator brush riggings, generator feeders to primary. excitcrs,
generator
reglliators, motor control centres. Mechanical: boilers and refractories, steam
6.24 Kinematic Viscosity ( A ST M D445)
Conditions monitored: Oil viscosity.
piping, heat exchangers, radiators, cooling towers, diesel
l�xhaust mani
folds. hydralllic systems, gas mains, bearings, bearing lubrication. conveyor belts, drive gears, drive belts, couplings, plastii.: s. metals, gears, shafts,
extru
Applications: Oils used i n diesel and gasol i n e engines, gas turbines, transmis
sions, turbine blades, welds, buried steam lines. steam traps, brick refractories, wall
sions, gearboxe s , compressors and h ydraul i c systems.
and roof i n s u lation, ducting, rotating kilns, tyre defects . Continuous processes:
P - F interval: Weeks to months.
Operation: Thi s test (resistance to flow) measures the time it takes for a given voillme o f oil to pass through a calibrated glass capil l ary v i scometer under a speci fied head (gravity) at a given temperature (usually l OO°F or 3 8°C). The test can be used to monitor oil deterioration over time or to indicate the presence of contamination by fuel or other oils. The k i nematic viscosity i s the product of the
glass, paper, metal. plastic and rubber manufacture. P-F
interval: A few days to several months depending on the appl ication.
Operation: Infra-red scanners employ sets o f mirrors and/or prisms rotating at h i gh speed and a collimating lens to collect the radiat ion and deli ver it to a few a current, the detectors. The detectors respond to the radiation by amount of current being proportional to the amount of radiation. Th is output is
t i me o f flow and the calibration factor o f the i n strument. The dynamic v i scosity is the product of the kinematic v iscosity value and the density of the l iquid.
the n processed by an on-board processor into a visible colour
Skill: A trained l aboratory technician.
Skill: A suitably trained and experienced technician
A dvantages: Can be used for both transparent and opaque oils. Good repeata
Advantages: N on-contact safe to view e nergised electrical system_
bility . Can be used for most l ubricating o i l s .
or moving processes without i nfluencing the temperature o f the
Disadvantages: Flammable solvents are used i n t h e test. Not a f i e l d technique.
and presen
ted on a v i ew finder or monitor as a thermogram.
sensitive, can see temperature di fferences as small as O. I 'F or less
400
Appendix 4: Condition Monitoring
Reliabilitv-centred Muinterwllce Equipment expensive and can be cumbersome to move around. to interpret the results . Camera has several moving parts.
Applications: Pipelines, engi nes, transformer P-F
40 1 power cables.
interval: Hours to months.
Operation: Li ght i s passed down a fibre-optic cabl e . A certain arnollllt uf bac'\"'
7.2 Focal Plan Arrays (FPA's)
Conditions mOllitored: Electrical: current/resistance relationships form loose, oxidiscd or corroded connections, or malfunction of the component i tself.
Mechanic(/l: heat generated by friction caused by fau l ty bearings, inadequate
scatter is renected back towards the light source and diminishes the of the olltgoing signaL There i s a direct mathematical relationship between t he tiIlle i t takes the l ight t o travel down the cable for a given distance. the amount o f back
lubricatio n , misalignment, misuse and normal wear
scatter and the temperature of the cable. This relationship can be used to determine the temperature at given points along the cable.
Applications: J;Jectrical: power distribution and high tension lines, transfonners.
Ski!!: A suitably trained and experienced technician.
transformer bushings.
bank conncctions, thyristor banks, disconnects,
relays and circuit breakers. meter and control connections, circuit breaker contacts, bus and fuse connections. fuse clips and stab connections, moulded case and air breakers. motor w i ndings. thermal overloads, condnctor fatigue, generator windgenerator brush
generator feeders to primary, exciters, voltage regu-
lators, motor control centres. Mechanical: boilers and refractOlies, steam piping, heat
radiators, cooling towers, diesel engines, exhaust manifolds, gas
mains, hydraulic systems. bearings, bearing lubrication. conveyor belts, couplings, drive gears, drive belts, plastics, metals, gears, shafts, castings, extrusions, turbine
Advantages: U naffected by the presence o f electromagnetic interfercnce' : able i n hazardous environments: Can reach inaccessible locations: Combines temperature sensing and data transmission i n a s i ngle component ( the cable ): Continues function i n g e v en i f a cable break occurs: Accurate up to 4 kll1.
D isadvantages: Uneconomic in small i nstal lations. 7.4 Temperature Indicating Paint
blades, welds, hUlied stearn lines, steam traps, brick refractories, ducting, wall and
ConditiollS ilionitored: Surface temperature.
roof insulation, rotating kilns, tyre defects. COllfinllOU.I' processes: glass, paper.
Applications: l-l.ot spots, insulation fai l u re .
metal, plastics and rubber manufacture P-F
l'-F
interTal: A few days to several months depending on the application.
interval: Weeks to months, depending on the
Operation: A silicon-based pain t which changes colour as temperatures rise. rhe
Operation: A lens i n the FPA focuses the radiation onto a matrix of detectors
colour starts out green, changes to b l ue at 204 uC and turns white at 3 1 6"e. The
which deliver spatial and thermal resolutions that were previously u n known.
colours do not change back again as the .temperature drops.
Each detector is composed of many small elements. The detectors convert the radiation into electrical energy. which i s amp l i fied and processed i nto a visible image and presented on a view finder or monitor as a thermogram. FPA ' s have only one moving part
a cooler.
Advantages: S imple: Permanent record o f the highest temperature reached. Disadvantages: Colours do not change back again : Only useful at two fi xed
Skill: A suitably trained and experienced technician Advantages: H ighly versatile. Non·contact
Skill: No train ing required for observers.
safe to v i ew energised electrical
systems, stationary or moving processes without int1uencing the temperature of the o bject. Can see temperature differences as small as 0. 1 OF or less. Small and compact. Radiometric. Equipment more expensive than IR scanners. Needs speciali st to i nterpret the results.
temperature s : Service l i fe o f each coat only one to two years (provided it does not change colour i n the interim ) .
8 Electrical Effects M onitoring 8. 1 Linear Polarisation Resistance (Corrator)
Conditions mOllitored: Rate of corrosion in systems
to
con
ductive COITosive fluids.
7.3 Fibre Loop Thermometry
Applications: Cooling water systems, Illunicipal water systems, n uclear power
Conditions munitored: Temperature variations caused by i nsulation deteriora�
heat exchange waters, geothermal power
lion. leaks, blocked cooling systems, etc.
and pul p and paper m i l l s .
systellls. de�alillation
402
Appendix 4: Condition lHollitoring
Rcliabili(v-C('fltrcd Maintenance
P - F iIltClval.·
�cvcral months in most applications.
Operatioll.- The e l ectro-chemical polarisation is based on the fact that a s ma l l voltagc applied between a metal specimen and a corrosive solution produces a CUtTent. The ratio of applied voltage to current is i nversely proportional to the corrosion rate. so this ratio provides a measure o f the corrosion rate increase.
Skill: A suitably trained technician. Advantages. Provides a quick and direct indieation of corrosion rate and pitting tendency: Measures corrosion as i t occurs: Some i n struments record the corro sion condition: Automatic and portable systems available: Sensitive to corrosion rates as low as a fraction o f a mil per year: Easy to i nterpret results.
Disadvantages: Portable equipment does not provide a permanent reeord: Readmllst be adjusted when taken i n h igh sensitivity corrosive media: Gives no information o n total corrosion.
8.2 Electrical Resistance ( Corrometer)
403
Applications: Electrolyte environments sllch as chemical process plants. paper m i l l s , electrical generating plant. pollution control planls. desalination plants. etc; best suited to stainless stcel , nickel-based alloys and titanium P-F
interval: Depends o n thc material and the rate of ellrrosion.
Operatioll: Thi s technique takes advantage of the fact that. from the point o fvicw o f corrosion, a metal which i s in a passi ve statc ( low corrosion rate) has a noblc cOlTosion potential, while the same metal in an active statc ( h ighcr com)sion rate) has a much less noble potential. The potential changes when breaks down. and measurements ean be made using a voltmeter of about 1 0 megohm input i m pedance and full-scale detlection of 0.5 (0 2 volts.
Skill: Usually a trained technieian. but sometimes needs an Advantages: Monitors localised attack: Fast response to Disadvantages: Small potential changes can be influenced
changcs in tcmpera- ture and acidity: Does not give a direct measure of corrosion rate or total corrosion:
Expert assistance may be requ i red for interpretation.
COllditiollS monitored: Integrated metal loss (i.e. total corrosion) . Applicatiolls: Petroleum refineries, process plants, gas transmission plants, under
8.4 Power Factm' Testing
ground or undersea structures, cathodic protection monitoring, abrasive slurry
Conditions monitored: Power loss through the insulation sy:-,tem Gllbed by
transport, water distribution systems, atmospheric corrosion, electrical generat
to ground, moisture in cables_
ing plants, paper m i l l s , etc.
Applicatiol1s: Electrical circuits, transi()ffller windings. high
P-F interval.- As for linear polarisation resistance.
bushings, high and medium voltage cables.
OpemtiolL' The system is composed of a probe and an instrument to read the probe.
P-F interval: Several months.
The probe consists of a wire, strip or tube of the same metal as the plant being monitored. The e l ectrical resistance of the probe. measured by a bridge circ u i t, increases as the probe cross-section decreases with corrosion. The increased resis tance corresponds to total metal l os s , which is easily converted to cOlTosion rate.
transformer
Operation: Power factor is circuit resistance divided by c i rcuit i m pedance. A known voltage i s applied to the winding insulation and the res u l ting current is measured. The cosine o f the angle between the voltage and curren! i s called the power factor. The measured current squared times the insulation resistance i s
S'kill: As for l inear polarisation method.
called the watts loss. These values are measured and recorded when t h e insulation
Advantages: When plotted against a time scale, yields both corrosion rate and
system is first i nstalled to establish a baseline. S ubsequent tests results are compared to the initial readings. As the circuit impedance due (0 mois-
total metal loss: Can be used i n any environment: Portable equipment availab l e : On-line monitoring possible: In-plan! equipment provides permanent record s : I nterpretation normally easy.
Disadvantages: Does not indicate whether the corrosion rate at a particular time i s high or low: Portable equipment provides no permanent record.
ture, contami n ation, i n s u l ation shorts or physical i n - service o i l fil led transformer under 2%.
Skill: Conducting the tests: field technician. Analysing tht� data: an Advantages: One of the best predictive tests_
8.3 Potent.ial Monitoring
Conditions monitored: Corrosive states (active or passive) such as stress-corro sion
pitting corrosion, selective phase cOlTOsion. impingement attack etc.
the power factor rises.
A newly fi lled oil trans former should have a power factor of under n.Y/( and an
Disadvantages: Not an on-line technique.
404 8.5
Reliabilitv-ccntred Maintenallce and Other Voltage Generators
Conditiolls mOllitored: I n s ulation resistance. AppficatiollS: Electrical circ uits, P-F intervlll: Months to years. Operation: A known DC (250 volts to I OkY) voltage is applied to the equipment under test. resulting i n a (hopefu l l y ) s mall current flow. If there i s n o current return to the test set from the equipment under test. this current must be flowing to ground. The c u rrent flowing to ground is cal led ' leakage cUITent' The insu lation res istance can then be calcu l ated using Ohms Law.
Skill: Technicians or engineers.
Appendix 4: Condition Monitoring Techniques
405
Ohm s Law. R e s i stances of about 200 micfU-ohms are normaL although manu facturers routinely publish their own design limits. This valut� i s trended OVt�r time to assess deterioration. Maximum l i mi ts can be obtained fwm manufacturers,
Skill: Conducting the tests: field technician. A nalysing the data: an Advantages: Resistance values can be trended over t i me to detect
fai l -
ures before the breaker contacts deteriorate significantly.
Disadvantages: N o t a n on- line technique. Normal resistance meter cannot b c used due to the resistance being in the order of micro ohms. Not
as a t rue
predictive techniqlle,
8.8 Motor Circui t Analysis ( MCA)
Ai/I'lliltages: A s imple and very well understood technique.
Conditions monitored: Changes i n conductor path resistance caused
DislidFalltages: Test cannot be carried out on-line.
corroded connections, loss of copper (turns) i n the stator: Phase to phase i nduc
8.6 B reaker Timing Testing
affected by rotor position, rotor porosi ty and eecen t ri c i ty. stator turn. coil and
COllditiollS l11onitored: B reaker contact travel, speed, wipe and bounce, Applications: H igh and medi um voltage c i rcuit breakers. p-p interval: Weeks to months.
Operation: A transducer is mechanically attached to the breaker mechanism, then electrically connected to a timing set. The breaker circuit i s then operated through its entire cycle of ope n i ng and closing. The test set measures contact trave l , speed, wipe and bounce. These results are compared to the last test and to the manufacturers' recommendations. Trending this i n formation indicates whether adj u s tments to the breaker are necessary.
Skill: Conducting the tests: field technicians. Analysing the data: an engineer. A dvantages: H igh and mediu m voltage breakers can benefit from this test. Disadvantages: Not an on-line technique. Not applicable to moulded case breakers and/or low voltage breakers.
8.7 B reaker Contact Resistance Test
Conditions monitured: B reaker contact wear and deterioration. Applications: C ircuit breakers. P-P interval: Several weeks. Operation: A DC current, usually 1 0 or 1 00 amps is applied to the contacts. The voltage across the contacts is measured and the resistance can be calculated using
loose or
tance caused by magnetic i nteraction between stator and rotor: Stator inductance phase shorting; W i n d i ng cleanliness and res istance
to
ground.
ApplicatiolJs: Electric motors (DC, AC induction, synchronous and wound rotor), P , F intervlII: Several weeks to mOIlths depending on thc application. Operatiol1: A n umber o f tests arc taken together to
a
complete picture of t he
motor circuit condition. Algorithms and rules arc llSt'd t() measure the of any defects whieh may be prcsent. The test applies low DC and At' ( resistance to ground test uses 500 or 1 000 volts DC) at the motor control centrc (MCC) power bus to measure the following: resistance to ground, c i rc u i t rc sistance (phase-to-phase), capac itance to ground, inductance rotor i nfluence. DC bar to bar and polarisation i ndex/die lectric absorption, I n the conductor path o f the motor c i rcuit, the res istance o f each and compared to the other phases. Readings are usually lower for
is measured motor
cireuits and h i gher for smaller motors. U nequal resistance i n any part of the circuit unbalances the voltages in the phases, which in turn causes heating of the motor windings. Inductive i mbalance i s also measured. This indicates i mbalanced magnetic fields and unequal current flows i n the and i s most often associated with stator w i ndi ngs (but can be i ntluenced stator iron and rotors ). Increased capacitance values are normally assoeiated with the moto!'. When the void between the stator and thc lllotor becomes dirty and/or damp. the capacitive effect between the conductor path i ns ide t he i nsulation and the outer ' s k i n ' of the insulation is increa�ed, AC c urrent can pass across this ' natural capacitance' and then to ground via the
, wet COflnec-
tions to the motor casing.
Skill: Conducting the tests: field technician.
the data: an
406
Appendix 4: Condition Monitoring Techniques
Reliability-centred Maintenance
Advantages: Tests are done at low voltages and minimum current test signals which are non-destructive. Lightweight and portable equipment can be used in the field. Tests can be done at the MCC requiring no break in motor connections. Disadvantaf?es: Not an on-line tech nique. Motor circuit must be non-energised.
407
Applications: A C or D C motors. P-F interval: S everal weeks to months. Operation : This technique is based on the principle that an electric motor driving . a mechamcal load acts as an efficient, continuously available transducer. Th�
:� � �
mo o senses
�echanical load variations and converts them into electric current
8,9 Electrical Surge Comparison
var at ons whIch are transmitted along the motor power cables. These current
to-phase insulat ion deterioration, Conditions monitored: Turn-to-turn mld phaseone or more coils or coil group s. of ction conne the and reversal or open circuit in
electrIc motor,
motor s, DC armatures, synchr onous Applications: Induct ion or synchr onous
vanatlOns, though very small in relation to the average Cllrrent drawn by the
�an be �nonitored and recorded at a cOllvenient location away ? eqmp:nent. Analysis of the variations provides an indication
from the operat1l1
�
of m chme conditIOn, whIch may be trended over time to provide a warning of
field poles.
detenoratlOn or process alteration. The test is done by placing a single split-jaw
on motor cycle freque ncy and starts P·F interval: Week s to month s, depen dent
current transf rmer probe on one of the power leads at the motor control centre
under loaded condi tions. at high freque ncy to two separate but Opera tion: A transie nt surge i s applie d e waveforms reHect ed from each voltag ng resulti equal parts of a windin g. The gs are identic al, each wave windin both If scope. oscillo part are display ed on an so a single trace appears on the screen. form is exactly superimposed on the other, a short -circui t, or a reversed or open If one of the two windin g segme nts contai ns Jfthis problem is found, it is necess ary coil, the waveforms are visibly differe nt. be done by comparing each segment can to establi sh which segment is at fault. This nation produces the wavefonn de combi to a third segme nt, and noting which eause fairly small differences in turns g missin or d fleetio ns. Gener ally, shorte as coil reversal or interphase shorts waveform amplit ude. M i s -conne ctions such larities in waveform shape. With this tend to cause large differe nces or irregu ine the voltag e at which turn-to -turn or method i t i s also often possib le to detenn shortin g i s near operat ing voltag e, phase- to-pha se condu ction begins . I f this and should be replac ed as soon as fault tion then the motor has a seriou s insula operat ing voltag e plus 1 000 Y, twice to up ed possib le. I f shortin g is not detect motor can be return ed to servic e. the windin g is consid ered good and the operator. Skill: A traine d and an exper ienced test to-phase shorting often occur before Advantaf?es: Portab le. Turn-to-turn and phaselonger P-F interva ls. Most equip giving tion insula deterioration of ground wall 8. 1 4 below ). (see test tial poten high m ment can also perfor sive. Cannot evaluate one coil b y itself. Disadvantages: Quite comp lex and expen the locatio n and severi ty of a fault Requi res careful repeti tion to determ ine
8.1 0 Moto r Curre nt Signa ture Analy sis bares) or shorti ng rings, high resista nce Conditions monitored: Broke n rotor stator air gaps, rotor mispo sition, deteri betwe en bars and rings, uneve n rotortion. orated or shorted rotor or stator core lamina
?
or starter cab1l1et. The raw waveform signal is amplified, fi ltered and fUlther . proce sed t obta1l1 a measurement of the instantaneollS load variations within
�
�
the dnve trm n and the ultimate load. In general, the current in the three phases . hould not dIffer by mor than 3%. If the variation exceeds 3% for any phase,
�
�
stator problems could eXIst. The amplitudes at line frequency can also be COll1-
��
?
pole pa s frequency immediately to the left of line frequency. A p are d �lth t . . slgmf'Jcant dIfference III amphtude between these two frequenCies indicates a cracked or broken rotor bar, end ring, or slip ring, or resistance joint problems. Skill: To clamp current transformer around one of the
�
�
power line leads
an xper enced electrician. To conduct the test and interpret the results: a tcch lllClan WIth an understanding of electric motors.
�
�dvantages:' On-line r easurements can be taken without breaking any electrical
�
conn ctlOns. No electncal connectIOns are required which reduces the h azard of
�
electncal shock . Readings can be taken remotely and safely on large, hiuh b speed or otherWise hazardous machines.
:
Disadvantage : Complex due to the relatively subjective nature of in terpreting
�h e spectra (thiS has re�entl� be.en improved from a data collection and analysi� IllterpretatlOn standpolllt). Eqlllpment expensive.
8.1 1 Power Signature Analysis
��
Con( it ons monitored: Rotors, broken bars, cracked or broken end
bad
!
cage omts, bowed or bent rotors; Stators, shorted lamination, eccentricity; Single phasmg, phase current and voltage balance, resistive and inductive irnb'.! lance'
�
�
Torque variations, wear or deterioration of machine clearances, now or l achill output restrictions, machinery alignment; Machinery efficiencies .
Applications: AC induction motors, synchronous motors, com pressors, pumps and motor operated valves.
408
Reliability-centred Maintenance
P-F interval: Several weeks to months.
Appendix 4: Condition Monitoring Techniques
409
and voltage signals while the motor is running. A signal conditioning unit condi
the power of the pulses (inten� ity) the rate of change over timc of thc powcr (trend analysis ) When trendin g data three levels o f PD thresho lds are set. Green means good for contmued use, yellow means increased activity and red means failure is immine nt.
tions and filters analog signals sensed from the feed lines. Data files are compiled
Skill: Experienced electrical technician.
Operation: Probes are attached to motor feed lines either at the Motor Control Centre (MCC), at a breaker box or locally at the motor, to gather electric c urrent
and analysed using application based software tool s (Fast Fourier Transform) to plot variables slIch as lotal real power, tolal reactive power and total power factor. Analysis of thc plots enables the motor and overall system performance to be evaluated in dctail. The plots can also be compared to baseline fingerprints to detect deviations. Skill: To attach probes to livc motor feed lines: an electrician. To conduct the test and interpret thc results: an experienced technician. Advantages: Tests can be done without shutting down the eqnipment. One of the few techniques that enables broken rotor bars to be detected under load. Allows
•
•
:
Advan ages: Allows
�
uick and more informed decision s. Can bc applied to any type of electnca l eqIllpme nt.
Disadvantages: Current availabl e on-line technolo gy cannot locate the exact source of the PD wh le the equipment is energised. One single data point provides . lIttle or no mforma tlOn. Several data points are needed to trend information. No current standards availabl e on maximn m acceptable levels of PD activity, except for cables. Expert knowledge and statistical analysis required to set PD threshol ds. Need special sensors off-line to determi ne the exact location of PD activity.
!
equipment efficiencies to be determined.
8.l3 High Potential (Hi·}>ot) Testing
Disadvantages: Skill and care required when connecting probes to live motor
Conditions mon itored: Motor winding gronnd wall insulatio n deterioration.
feed lines. Interpreting and analysing the data takes some practice and an under standing ofelectric motors and the driven equipment i s necessary. Limited number
Applications: AC and DC motors.
of industry-widc applications for comparison. Equipment expensive.
P-F interval: S e veral weeks.
8.12 Partial Discharge Conditions monitored: Insulation breakdown. Applications: All typcs of medium voltage electrical equipment including switch gear, bus ducts, transformers, arresters, bushings, switches, motor starters, pot heads, motors, generators, cable terminations, cable splices, and the cables them selves. D istribution systems and equipment > 2,000 volts AC.
-
P F interval: Several weeks to months (voltage levels, the shape of the void, am
��
Operation: High voltage i s applied to the stator winding s in graduated steps or ramps up to a hm t, usually twice the line vo ltage. Tcst voltages arc usually denved from th IEEE Standard 95. At the first sign of non-Iincarity in the test current .or drop 111 I I1SUlatlo n resistanc e with further voltage increase , the test voltage IS recorded and the voltage removcd in ordcr to avoid complet e insulatio n breakdown. If the insulation withstan ds the voltage, it is considered to be safc and the motor can be return d to service. Any trend in voltage at which non-linc arity . m current drop or lIlsulatlOn resistance occurs can be used to predict remaining life.
�
�
�
bient temperature, system losses all influence how quickly the insulation fails).
Skill: An experienced electrical technician.
Operation: A partial discharge (PO) occurs when a small void, crack, or irregu
Advantages: Tests normally correlate with surge compari son tests.
larity in an insulation system causes an electric field to build up. Sensofs are used to pick-up the PD. On switchgear, the sensor i s connected between the grounded side of the metering CT circuit. On cables, the sensor is connected around the ground wire that connects the cable shield or placed around the insulated con ductor. On motors, sensors are placed on the motor frame Of around the ground connection or around the insulated motor lead. In analysing the data, three issues are considered: •
the number of pulses per cycle and the magnitude of the pulse (field strength in pi co coulombs)
I?isadvantages: Motors have to be taken out of service to conduct the test. Test mg potentially destructive .
8.14 Magnetic Flux Analysis Conditions monitored: B roke'l rotor bars, unbalanced phases and anomalies in stator windings such as turn-to-turn, phase-to-phase and phase-tn-ground shorts. Applications: AC induction motors. P-F interval: Several weeks to months
4 10
Reliability-centred Maintenance
Operation: A flux coil sensor i s placed at the centre of the axial outboard end of the motor. (Consistent positioning of this sensor i s essential for reliable and trend
Appendix 4: Condition Monitoring Techniques
411
Skill: Field technician. Advantages: The test can be perfonued without removin g the
domain using an FFr analyser. A trend of certain magnetic flux frequencies will
hattery from service, as the AC signal is low level and 'rides' on top of the DC of the battery .
indicate electrical asymmetries associated with the rotor and stator windings.
Disadvantages: Test could take a long time on large battery banks.
able data.) The signal received from the sensor is transformed into the frequency
Most of the peaks in a flux coil spectrum occur at frequencies which have some relationship to running speed. Broken rotor bars increase the sideband activity around running speed. Unhalanced supply voltage (which causes motor heating and eventually leads to premature deterioration of the stator windings) shows no change except around the peak occurring at line frequency + 1 x RPM . One of the first faults a winding will encounter is tum-to-tum shorts, which then migrate into phase-to-phase or phase-to-ground shorts. A winding faul t can be indicated around the 3 x running speed sideband of line frequency. A variation of this technique i s used to detect turn-to-tum shorts by looking at the family of 'slot pass' frequencies from measurements taken with a flux coil . Flux measure ments are taken as mention above, and the resulting signature i s analysed at the 'slot pass' frequencies. The principle slot pass frequency occurs at the product of the number of rotor bars and running speed. The technique involves compar ing spectra over t i me to determine when changes occur. Skill: To record the spectrum: lin electrician/technician with an understanding of motors. To interpret the results: an engineer. Advantages: One of the few techniques that can detect faults associated with elec trical insulation of electric motors while the motor is on-line. Disadvantages: High degree of skill and knowledge of electric motors required to interpret results.
8.15 Battery Impedance Test Conditions monitored: Cell deterioration . Applications: Emergency power and DC control power batteries. P-F interval: Several weeks. Operation: A s a battery ages and begins to lose capacity, its internal impedance rises. A battery impedance set i njects an AC signal hetween the terminals of the battery. The resulting voltage is measured and the impedance calculated. Two comparisons can then he made: first, the impedance i s compared with the last reading for that battery ; and second, the reading i s compared with other batteries in the same bank. Each battery should be within \0% of the others and 5 % of its last reading. A reading outside these values indicates a cell problem or capaeity loss. There are no set guidelines and limits for this test. Each type, style and COIl figuration of hattery has its own impedance, so it is important to take a baseline reading early in the battery ' s life .
9 A Note on Leaks With the exceptio n of ultrason ic leak detectio n, a topie which has not heen covered i n much detail i n this Append ix i s leaks., especia lly in undergr ound storage tanks. This i s because a publica tion which provide s a compre hensive descript ion of 3 6 differen t leak detectio n method s i s already availahl e. It is ealled " Underg round Leak Detectio n Method s A State of the Art Review ", and is i n the form of a report prepare d in 1 9R6 by Shahzad Niaki and John Broscio us of the IT Corpora tion in Pittsbur gh and commis sioned by the Hazardo us Waste Enginee ring Researc h Laboratory, Edison, New Jersey. Copies of the report are availabl e from the Nationa l Technic al Informa tion Service , a division of the United States Departm ent of Comme rce based in Springf ield, Virginia , USA.
Bibliography
Bibliography
4I3
Mercier J-P. Nuclear Power Plant Maintenance. Maisons- Alfort France. Editions Kirk. 1 987 Mohr G . "Technology Overview - Ultrasonic Detection". P/PM
8
(5) 1 995 Moubray JM . " Maintenance Management
A New Paradigm " .
7/1ird Anl/ual
Conference ofthe Society ofMaintenance & Reliahilitv Illinois. 2 ofASTM Standards. Phil America n Society of Testing Materia ls. Annual Book 995 1 ASTM. vania. Pennsyl adelphia . Harlow, Essex. LongA ndrew s JD & Moss TR . Reliability andRiskAssessment man. 1 993
New York. Wiley. 1 995 Blanchard B S , Verma D & Peterson EL. Maintainability. Analysis. Englewo od and ring Enginee Systems WJ. y B lanchard B S & Fabryck 990 1 Hall. Prentice Jersey. New Cliffs, Englewood Cli ffs, New B lanchard BS. Logistics Engineering and Management. Jersey. Prentice Hall. 1 986 on Inductio n Motors using B erry JE. " Detection of Multiple Cracked Rotor Bars Technology. 9 (3) 1 996 P/PM s" both Vibratio n and Motor Current Analysi ance of AC Inductio n Mainten ve Predicti for Bowers S V . " Integrated Strategy Conference. Indiana l Nationa gy Technolo ance Mainten e Motors" Predictiv polis, Indiana. 4 .' 6 December 1 995 ment. Oxford. B utterworth Cox SI & Tait NRS . Reliability, Safety and Risk Manage Heinemann. 1 99 1
Technology National Dalley R. "Wear Patticle Analysis " . Predictive Maintenance 995 1 er Decemb 6 4 . Indiana COI!ierence. Indiana polis,
of the American StatisDavis D J: " An Analysi s of Some Failure Data " . Journal tical Association,
47 (258) 1 95 2
Analysis t o Detect the DelZingaro M & Matthew s C . " Using Power S ignature s" . P/PM Technology. Machine Behavio ur of Electric Motors and Motor-driven
8 (5) 1 995
Maintenance. Palo Alto, Gaertne r J P. Demonstration of Reliability-Centered 989 1 . Institute h Researc Power Californ ia: Electric
Making a New Science. New York. Penguin . 1 987 ance Technology National James R. " B asic Oil Analysi s " . Predictive Mainten 1 995 Conference. Indiana polis, I ndiana. 4 - 6 Decemb er 1 995 Gulf. Texas. , Jones R B . Risk-based Management. Houston Prioritis e Mainte Help Can logies Techno ance Kane CF: " Predicti ve Mainten Show. Merrill Trade & ble Roundta Business Indiana st Northwe . nance Dollars" Gleick J. Chaos
ville, Indiana. 3 October 1 996 Program Development Mainten ance Steering Group - 3 Task Force . Maintenance tion (ATA) of Associa rt Document MSG-J. Washin gton DC: Air Transpo America 1 993
Chicago
4 October 1 995
Moubray J M . "Maintenance and Product Quality" . International Total Quality, Hong Kong: 1 6
-
on
17 November 1 989
Moubray J M . " Maintenance and S afety
a Proactive Approach " . Annual Con-
ference ofthe Accident Prevention andAdvisOlY [lnil
UK National fIealth
and Safety Executive, Li verpool, U K ; 1 9 May 1 989 Moubray 1 M . "Developments in Reliability-centred Maintenance " . The Factory Efficiency & Maintenallce Show and COI\ferellce. NEe, B i rlllingham, U K; 27
- 30 September 1 988 Moubray J M . " Maintenance Management - T h e Third Generation " . The 9th European Maintenance Congress, Helsinki, Finland; 24
27 May 1 988
Moubray J M . "Reliability-centred Maintenance " . A
on Condition
Monitoring, Gol, Norway; 2 - 4 November 1 987 Nelson W. Applied Life Data Analysis. N e w York: Wiley, 1 982 Niaki S and Broscious 1 A . Underground Tank Leak
Detection Methods
A
State-of-the-A rt Review. Springfield, Virginia: National Technical I n fonna tion Service, US Department of Commerce, 1 985 Nicholas J R l r ( 1 995): " A C & DC Motor Circuit Testing and Predictive Analysis" . Predictive Maintenance Technology National Conference. Indianapolis, I n diana. 4
6 December 1 995
Nowlan F S & Heap H. Reliability-centered Maintenance. Springfield, V irginia: National Technical Information Service, US Department of Commerce. 1 978 Oakland J S . Total Quality Management. Oxford. B utterworth Heinemann. 1 1)89 Perrow C. Normal A ccidents. New York. H arper-Coll ins. 1 984 Reason IT. Human Error. Cambridge, England. Cambridge Uni versity Press . 1 990 Resnikoff H L. Mathematical Aspects ofReliability-centered Maintenance. Los Altos, California: Dolby Access Press. 1 978 Robinson DrlC & Piety Dr KR. "Peak Value (PeakVue) Analysis: Advantages over Demodulation for Gearing Systems and Slow-speed Bearings". North west Indi, (lna Business Roundtable & Trade Show. Merrillville, I nd iana. 3 October 1 996 Rose A. "How to Set Up an Electrical Predictive Maintenance Program " . Predi c · tive Maintenance Technology National Conference. Indianapolis, Indiana.
9 November 1 994 Sandtorv H & Rausand M. " RCM - Closing the Loop between Design Reliability and Operational Reliab il i ty " . Maintenance, 6( I), 1 3
2 1 . 1 99 1
Shiftiin CA "Aviation Safety Takes Center Stagc Worldwide " . A viation Space Technology.
145 (19). 1 996
Week &
414
Reliability-centred Maintenance
York. McGraw-HilI. 1 993 Smith AM. Reliability-centered Maintenance. New Butterworth Heinem ann. Oxford. Risk. Smith OJ. Reliability, Maintainability and
Index
1 993
Butterworth Heine Snow DA (editor). Plant Engineer's Reference Book. Oxford. mann . 1 99 1 ntation. Website Tissue B M . SCIMEDIA - Analytical Chemistry and Instrume . 1 996 Published by author. 1 995 Toms LA. Machinery Oil Analysis. Pensaco la, Horida. ent). Mil-Std-2 1 73: Departm tandards S & ations U S Navy (Eng ineering Specific ircraft, Weapon NavalA or ment.�f Require ance Mainten ered Reliability-Cent U S Departm ent of Systems and Support Equipment. Lakehur st, New Jersey. Defense . 1 986
ve Mainten ance A van der Horn G & Woyshn er W. " Electric Motor Predicti ogy National Technol ance Mainten ve Compre hensive Approac h" . Predicti - 6 Decemb er 1 995 4 . Indiana polis. Conference. Indiana ogy. 8 (5) 1 995 Weaver C. "Time Waveform Analysi s" PlPM Technol ve Maintenance Predicti . " rs Analyse and rs Collecto Data n Vibratio White G . " 6 4 - Decemb er 1 995 Technology National Conference. Indiana polis, Indiana. ve Maintenance Predicti . " White G . " Designi ng and Installin g a Monitoring System er 1 995 Decemb 6 4 . Indiana polis, Technology National Confere nce. Indiana Spike using Pumps Sealless of ing Monitor Xu Miug & Le B leu 1. "Condit ion EnergyTM" . P/PM Technology.
8 (6) 1 995
AE: see Atomic emission
Bathtub curve: 1 2. 249
Abrasion : 5 9
B attery impedance test: 4 1 0-4 1 1
Accelerometer: 35 1
Benchmarking: 303
Acceptable:
Benefits o f ReM: 307 ·3 1
looks: 25, 296 risk: 30. 98- 1 0 1 , 256, 343-347 Acoustic emission: 3 6 1 -362 Actuarial analysis: 250-255, 257
Bhopal: 3()8 Blot testing: 368-36<) B o iler: 3 3 2 house: 227. 328-332
Administering RCM: 275-277
Boundaries: see System boundaries
Age-related failures: 1 1 , 1 3 1 - 1 33 , 1 60-
Brake lights; 1 73 , I 7 S- 1 76, 1 9 1
1 6 1 , 235-238, 243-246 Air Transport Association o f America (ATA); 3 2 3
Breaker contact resistance lest: 404-405 Breaker
404
Broad band vibration analysis: 352
Aircraft accidents: 3 0 9
Broken bel! detector: 4 1
Airlines a n d ReM: 3 1 8 - 3 2 1
B uttertly effect: 1 5 9
All-metal debris sensor: 366-367 AmplitUde demodulation: 357
"Can be done by": 208-20<)
Amplitude sensors; 35 1
Capability; sec Initial capability
Analysis paralysis: 65 Analytical ferrography: 362-363
of people: 22 1 Car:
Anthropometric factors: 3 3 5 - 3 3 6
desired performance: 3 7 , 3 9 . 43
Anxiety: 4 0
bodywork: 295-296
Appearance: 40
Causation: 65-70
Applying RCM: 1 6- 1 8, 26 1 -2 9 1
Cause and effec t: 72
Asset hierarchy: 327-330
Centralised lubrication systems: 59
Atomic absorption spectroscopy: 374
Cepstrum: 356-357
Atomic emission spectroscopy: 372
Challenges
Attentional failures: 3 3 8
Chaos: 22
maintenance: 5
Attractiveness: 25
Chaos theory: 1 59
Audit trail: 20, 3 1 6
Cbecklist: 227-229
Auditing RCM: 1 8, 214-217, 27 1 , 274,
Chemical effects: I SO, 349. 370-381\
276 Availability: 3 , 294-304, 3 1 0-3 1 2 of hidden functions: 1 1 5 - 1 1 8, 257
and failure-finding: 1 75- 1 79 Average l i fe : 1 32, 1 37, 2 3 8
Chemical measurement o r fluid properties: 377.· 378 Civil aviation:
anri RCM: 3 1 1\ -3 2 1 safety record ; '109 Clear and bright lest: .1 87-381\
" B I O" life: 24 1 -242. 294
Coal mine: 3 1 0
B a l l bearing: 1 4 1 , 1 5 7 1 59, 207
Cold ligbt rigid probes: 393-394
Batch processes: 29
Colour indicator titration: .183
Reliability-centred Maintenance
416
Index
Combination of tasks: 1 69 , 206
Crankshaft grinding: 26, 27, 34, 49
Comfort: 40
Criticality assessment: 280
Effects: see Failure effects
Commissioning: 248
Customer service: 3, 1 9, 29, 1 04, 20 \ , 280
Efficiency: 42, 294-295
Complex equipment: 1 42
Cusum chart: 1 5 1
see Maintenance effectiveness
energy: 294
lind actuarial analysis: 2 5 0
maintenance; 304-307
see
Computers a n d R C M : 2 1 1 , 27 1 , 2 9 0
Damage:
Condition-based maintenance: 1 45
Database: 1 9, 267-268, 3 1 5-3 1 7
Electrical effects: I SO, 350, 40 1 -4 1 1
Condition monitoring: 5, 1 49- 1 50, 348-41 1
Data: 2 55-260
Electrical resistance meter: 402
Conditional probability of failure: 237
Secondary damage
failure: see Technical history data
Consequence evaluation: I I , 2 1 7
Debottlenecking: 62-63
Consequences of failure: 1 0- 1 1 , 1 5 , 7 1
Deep probe endoscope: 394
Elapsed time: 230-2 3 1
Failed state: 46, 5 3 Fail-safe: 1 1 1 - 1 1 2, 1 27 Failure: 4546 age-related: see Age-related failures conditional probability of: 1 1 - 1 2, 1 32 consequen ces: see Consequen ces of failure data: see Technical history data
Electrical surge comparison: 406
definition: 46
Electro-chemical corrosion monitoring:
development period: 1 46
382
effects: see Failure effects
72, 90- 127, 203-204, 2 8 5
Decision Diagram: 200-201, 267
Electron fractography: 395
Decision support: 5
evidence of: : 74
avoidance: 9 1
Electrostatic t1uorescent penetrants: 389
categories: 1 0
Decision Worksheet: 198-2 1 1 , 267 , 27 1
evident: 92-93
Emergency medical equipment:
economic: 1 5, 1 05 , 1 08
Default actions: 1 4, 9 1 , 170-197
Emergency power generators:
-finding task: see Failure-find ing
environmental: see Environmental hidden failure: see Hidden operational:
see
Operational
non-operational: see Non-operational recording decisions: 203-204
recording decisions: 206 Defect reporting: 233-234 Design changes:
see
Redesign
Desired performance: 23-24, UO, 1 89 , 256
Empowerment: Energy dispersive X-ray spectrometry: 37 5
multiple: see Multiple failure non-age-related: 1 40- 1 43
Environmental consequences: 3, 1 0, 1 5 ,
pattern: see Patterns of failure
Detective tasks: 1 7 1 Deterioration: 58-50, nO- 1 3 3
93, 94-103, 204, 263
Device:
definition of: 95
protective: see Protective devices
manageme nt policy: 64, 69
Environment: 324
safety: see Safety consequences
Consistency of P-F interval: 148- 1 49, 205
functional: see Functional failure hidden: see Hidden failure modes: see Failure modes
Entropy: 2 2
strategic framework: 1 27 Consensus: 1 7 , 267, 273
and redesign: 1 02 , 1 90
potential: see Potential failure random: see Random failure reporting policies: 2 3 3 -234, 2 5 1 -252 traditional view of: I I
Consolidating task intervals: 223
DIAL: see Differential absorption LTDAR
Constant bandwidth analysis: 353
Dielectric strength: 375-376
Environmental hazards: 7 5
Constant percentage bandwidth analysis:
Differential absorption lidar: 377
Environmental integrity: 1 9, 3 8 , 308-309
Direct reading ferrograph: 363-364
Equipment effectiveness: 3 0 I
1 7 1 - 186, 206, 259, 265
Disassembly: 60
Equipment vendors: 77-78, 288-289
decision process: 1 86
Ergonomics: 40
definition: 1 73
Erosion: 59
intervals: 175·)85
353-354 Contaminants in tluids: 370-3 7 1
see
Continuum o f risk: 1 20, 343-347
Discard:
Containment: 26, 40, 299
D istribution:
Contradiction:
in maintenance schcdules: 223 -224 the ultimate: 252 Control: :�9
Scheduled discard
standards: 30, 9 5 , 279
Failure effects: 9, 73-79, 9 1 , 2 1 6, 262, 285 Failure-find ing: I I , 1 4, 1 5, 1 22 , 1 23 , 1 70,
exponential: 239-241
'ESCAPES' functions: 38-44
normal: 1 32 , 236-237
Evaporation: 59, 1 34
survival: 236, 240, 244
a rigorous approach: 1 79- 1 8 3
Evidence of failure: 74
Weibull: 242-243
other approaches: 1 84
Evident failures: 92-93
technical feasibility: 1 85
Control limits: 2 7
Dirt: 60
Corrator: 401 ,,402
Disassembly: 60
categories: 93 Evident function: 92
a less formal approach: 1 83 - 1 84
worth doing: 1 85 Failure modes: 9, 53-73, 2 1 6, 1 24, 262, 270, 2 8 5
Corrective maintenance: 17 j
Downtime: 3 , 75-77, 1 47 , 256, 2 7 8 , 294
Exceptional violations: 3 3 8 , 342
COlTometer: 402
Drawings: 20, 3 1 6
Exhaust emission analysers: 382-383
categories: 5 8 ,64
Corrosion: 1 34
Dynamic effects: 1 50, 349, 3 5 1 -362
Existing assets: 1 9
definition: 5 3
Existing maintenance schedules: 7 1 -7 2
detail required: 64-73, 80-88
Economy: 42
Expectations of maintenance: 3
probability: 70-7 1
Economic:
Expert systems: 3 1 7
reason for analysng: 5 5 - 5 7
Exponential distribution: 239-24 1
sources of information: 77-80
Coupon testing: 3 9 1 Cost: 4 effectiveness: 3 , 1 9 , 3 1 2-3 1 4 maintenance: 278, 305
consequences: 1 5 , 1 05 , 1 08
operating: 1 04, 20\
-life limits: 1 39- 1 40
repair: 1 04, 1 08 - 1 10, 1 47 , 256
risk: 1 1 9 , 343, 346
417
Failure rate: 293 , 295 FAA: 3 1 8-323 s e e Fast
Crackle test: 387
Eddy cUlTent testing: 392
FFT:
Craftsman: see Maiutenance craftsman
Effectiveness:
Facilitators: 1 7 - 1 8, 269-277
Fourier transform
see also: Mean time between failures and note on P96 Falling ball comparator: 398
Reliability-centred Maintenance
418
Index
Fast Fourier transform: 3 5 1
Graded filtration: 367-368
Fatigue: 5 9 , 1 34
Groups:
see
Review Groups
Knowledge-based mistakes: 338, 34 1 -342 Kurtosis: 3 6 1
Hidden failures: 1 0, 1 5 , 92-93, 1 1 1 -128,
Manufacturer of equipment: 77-78, 288289
and bearing failure: 1 57 - 1 59 Fault-tree analysis: 344-345
419
Margin for deterioration: 23 "L 10" life: s e e B 1 0 life
1 72 , 204, 263
Market demand: 3 2
Fedcral Aviation Agency: see FAA
Lapses: 338-339
normal circumstances: 1 26
Mean t i m e between fai l ures: 1 06. 1 75 -
Lead time to failure: 1 46
Perrography: 362
operating crew: 1 25 - 1 26
Leak detection: 4 1 1
Fibre loop thermometry: 400-401
primary & secondary functions: 1 25
Mean time to
Field technicians: 78, 3 1 3
and redesign: 1 22, 1 9 1 - 1 92
Legislation:
Measming maintenance performance: 292
First Generation: 2
the question of time: 1 24
Feasible:
see
Technically fcasible
Flow processes: 2 9
Hidden functions: 1 1 1 - 1 28, 1 70
Fluorescence spectroscopy (X-ray): 3 7 4
decision process: 1 23
FMEA: 53·89
required availability: 1 1 5 - 1 1 8
see also Failure modes alui Failure effects
Focal plan arrays: 400 Fourier analysis: 35 1 Fourier transform infrared spectroscopy:
Hierarchy: asset: 85, 327-330 functional: 329-333 High-frequency:
see Maintenance schedules
High potential (Hi-pot) testing: 409
378
see
Fractional dead time: 1 1 7
H istory:
Frequencies:
Hopper: 1 95- 1 97
Technical history
1 8 2 , 23 8 , 24 1 , 257-260, 263, 293-301
safety: !O3
76
Meetings: 266-269, 272-276
Level of analysis: 8 1-88, 2 1 5
Mcgger testing: 404
LIDAR: 370
Memory failures: 3 3 8
Life: 294
Mesh obscuration particle counter: 364
average: 1 32, 1 37, 23g
Mil-Std- 2 1 73 : 323
safe-: 1 38 - 1 3 9
Military undertakings: 1 04
usefu l : 1 9 , 1 32 , 1 37 , 238, 3 1 4
Milk processing: 3 1 0
Light detection and ranging:
see
LIDAR
Milling machine: 29 , 303
Light extinction particle counter: 365
Mistakes: 338-34 1
Light scattering particle counter: 366
Modes:
Likely victims: 99- 100, 263, 308, 347
Modify:
Linear P-F curves: 1 60- 1 62
Moisture monitoring: 385
see
Failure modes
see
Redesign
consolidation: 223
Human error: 60, 6 1 , 70, 335·342
Linear polarisation resistance: 40 1 -402
Human senses: 1 49, 1 5 3 - 1 54
Motivation: 20, 3 1 5
failure-finding: 1 75- 1 85
Liquid dye penetrants: 3 8 8
Human sensory factors: 3 3 7
Motor circuit analysis: 405-406
Living R C M program: 2 7 7
Hygiene: 39
Motor current signature analysis: 406-407
Lubrication failure: 5 9 , 7 2
high: low:
see
see
Maintenance schedules
Maintenance schedules
on-condition: 1 45 - 1 49 , 1 63 - 1 65 scheduled discard: 1 38 - 1 40 scheduled restoration: 1 35
Implementing RCM recommendations: 1 8, 2 12-234, 276
MSG- l : 320- 3 2 1 MSG-2: 320- 3 2 1
MTBF:
see
Mean time between failures
see
MSG- 3 : 322-326
MTTR :
Mean time to repair
Multiple failure: 1 1 3 - 1 j 5 , 206
Frequency analysis: 356
Inductively coupled plasma: 373
Magnetic chip detection: 368
Initial capability: 23-24, 47-48, 64, 130
pmbabi lity o f: 1 1 5 - 1 20, 1 72- 1 73 , 1 79-
Fuel line: 8 1 -85
Magnetic flux analysis: 409-4 1 0
Initial incapability: 64
1 80, 257, 263
Functions: 7-8, 2] -44, 2 1 5 , 26 1 -262
Magnetic particle inspection: 3 8 9
definition of: 22
Initial interval: 207-208, 2 1 7
Maintain:
different types : 35-44
Infant mortality: 1 3 , 1 43 , 247-249, 3 1 1
evident: hidden:
see see
Evident functions Hidden functions
Information Worksheet: 3 8 , 5 2 , 54, 89, 202, 267, 2 7 1
Nett P-F interval: 1 46- 1 48
definitions: 6
New expectations: 3
Maintainable: 24
New assets: 1 9 , 269
Maintenance: 1 88- 1 89
New research: 4
primary: 8 , 35-37
Infra-red scanners: 399-400
challenges: 5
Infra-red spectroscopy: 379
New techniques: 5 6
secondary: 8, 34, 37-44
costs: 3 1 2-3 1 4
No scheduled maintenance: 1 4 , 1 5 , 1 07 ,
superfluous: 43 Functional block diagrams : 36, 329-333
Initial capability: 23-24
craftsmen: 1 7, 222, 226-232, 262-268,
Initial interval: 207-208
291
1 09 , 187, 206 Non-age-related failures: 1 40 - 1 43
Functional checks: 1 7 1 , 1 73
Installation and infant mortality: 247-248
definition: 6
Integrative framework: 3 1 7
Non-maintainable: 24
Functional hierarchies: 329-333
effectiveness: 293-304, 3 1 2- 3 1 4
Fnnctional effectiveness: 297-304
Interfacial tension: 376
Non-operational consequences: IO, 1 5 , 93 ,
efficiency: 304-307
ISO 9000: 2 1 9
108- 1 1 0, 1 70, 204, 264
Functional failures: 8-9, 45-52, 1 24 , 1 55 ,
labour: 305
and redesign: 1 10 , 1 93
2 1 6, 262 definition: 47 Gas chromatography: 379-380
planning & control systems: see Planning Job card: 233 -234
Normal circumstances: 1 26, 263
and redesign: 1 88- 1 89
Just-in-time: 3, 3 1
Normal distribution: 1
schedules: 1 8 , 222- 224
Gas station: 297-302
supervisors: 1 7, 228-234, 262-268, 29 1
Octave band
236-237 3 52 - 3 5 3
Gearbox failure: 72, 86-88
Karl Fischer titration test 385-386
Management information: 258, 292-304
OEE: 302 304
Generic failure data: 79
Kinematic viscosity test: 398
Manuals: 20, 3 1 6
O i l appearance: 395-396
Index
Reliability-centred Maintenance
420
Oil colour: 395
Patch test: 369
Oil odour: 396
Patterns of failure: 1 2, 235-249
Once-off changes: 1 8, 220-221
A: 1 2, 1 33, 249
On-condition tasks: 1 1 , 14, 145-169, 205,
age-related: 1 1 , 1 3 1 - 1 33
264
B : 1 2, 1 32, 1 33, 235-238
definition: 1 45
C: 1 2, 1 33, 243-246
intervals: 1 45 - 1 49
D: 1 2, 1 43 , 246
pitfalls: 1 55 - 1 56 Operating context: 7 , 28 - 35, 50, 72-73, 9 1 , 1 08 , 2 8 1 , 285 hierarchy: 34-35
E: 1 2, 1 43, 239-242 F: 1 2, 1 43, 246-249 non age-related: 1 40- 1 43 Peak value analysis: 3 5 8
Operating costs: 1 04, 20 1
PeakVne analysis: see Peak value analysis
Operating crew: 1 25 - 1 26, 263
Performance standards: 7, 22-27, 47, 9 1 ,
Operating procedure: 1 8 Operational consequences: 1 0, 1 5 , 93, 103-108, 1 70, 204, 259, 264, 3 1 0
256 Phase: 3 5 1 Physical:
definition of: 1 04
asset: 6, 2 1 -24, 8 1 ,302, 304, 327-330
and redesign: 1 07, 1 93 - 1 97
effects: 1 50, 350, 388-398
Operations managers: 26 1 -268, 2 9 1
Physiological factors: 337-338
Operations supervisors: 1 7, 262-268, 29 1
Piper Alpha: 30S
Operator: 17, 209, 225-226, 262-26S, 2 9 1
Plant register: 1 6, 327-330
Outcomes o f RCM: I S
Planning:
Output: 1 9, 1 04, 20 I , 293
an RCM project: 1 6, 275, 2 8 3
Overall equipment e ffectiveness: 302-304
board: 2 3 0
Overhaul: 1 3, 1 43, 3 1 1
& control systems: 224-232, 306
see also Scheduled restoration tasks
high-frequency schedules: 224-225,
Overloading: 6 1 -64
226-229
Oxidation: 1 34
low-frequency schedules: 224-225, 230-232
P&ID: 3 2 9 P - F curve: 144- 1 45, 1 5 7 - 1 62, 348-349
elapscd time: 230 running time: 2 3 1
linear: 1 60- 1 6 1
Poka yoke: 3 3 9
and random failures: 1 56
Pore-blockage technique: 364-365
P F interval: 145-149, 1 56, 205, 256, 348-
Potential failure: 1 4, 144-1 45, 1 54, I SS,
4 10
205, 256, 3 1 0, 348-349
consistency of: 1 48 - 1 49, 205
definition: 1 44
definition: 1 45
Potential monitoring: 402-403
determining the: 1 63 - 1 65
Potentiometric titration:
nett: 1 46- 1 48 vs
operating age: 1 56, 1 60- 1 62
Packaging tasks: 2 2 3 Pan-view fibrcscopes: 3 9 4 Pareto analysis: 280 Partial discharge: 408A09
TANffBN: 383-384 TBN: 384 Power factor: 384-385 testing: 403 Power signature analysis: 407-408 Predictive tasks: 144-169, 1 7 1
Partial failure: 47
Pressure relief valve: see Relief valve
Particle effects: 1 50, 349, 362-370
Pressure switch: 1 25
Particle monitoring: 349, 362-370
Preventive tasks: 133-140, 1 7 1
Participation: 5, 266-269, 282-285, 3 1 5
Primary effects monitoring: 149, 1 52- 1 5 3
Primary functions: 8 , 35-37, 1 25
Quality checks: 226
PRN: 280
Quality standards: 3, 2 9
Proactive maintenance: 129-169
o f products:
42 1
Product quality
Proactive tasks: 1 1 - 14, 9 1 , 1 02 - 1 03 , 1 061 07, 1 70, 2 8 5 combination: 1 69 hidden failures: 1 2 1
Random failure: ]3, 1 40- 1 43 , 239-242
and P-F intcrval s : 1 56 Raw material supply: 3 3
order o f preference: 1 69
ReM:
recording decisions: 204-206
ReM:
selection: 1 67 - 1 69 Probability see
Mean Time between Failures and
note on P96 Probabiltyllisk number: 280
Reliability-centred Maintenance
facilitators:
see
Facilitators
in perpetuity: 284-286 meetings:
see lV'":"""'"
review gwups:
see
Process:
200-201
& instrumentation drawings: 329 batch: 29 flow: 29 Product quality: 3 , 1 9, 60, 1 04, 1 49, 20 l , 277, 3 1 2 monitoring: 1 5 1 - 1 52 Production effects: 75-77
Review groups
RCM 2: 323-326 Information Work sheet and Decision Worksheet Real time analysis: 354 Real time ferromagnetic sensor: 366 Redesign: I I , 1 4, 1 5 , 1 8, 188. 1 97, 206, 207, 220-22 1 , 26 5 , 2 7 1 economic justification: 1 9:\ - 1 97
Production managers: 26 1
hidden fai lures: 1 22 , 1 23 , 1 86, 1 92- 1 93
Project managers: 1 7, 275-277
cnvironmental consequences: 1 02 , 1 90
Proposed task: 206-207
and maintenance: 1 1\11 - I S9
Protected function: 1 1 0
non-operational consequences: 1 09- 1 1 0,
mean time between failures: 1 79- 1 8 1 ,
1 93 - 1 97
1 85
operational consequences: lO7, 193- 1 97
Protective devices: 4 1 -42, I I I - l i S , 1 721 85
safety consequences: 1 02 Redundancy: 29, 1 04 , 280 see
Proven technology: 247
Register:
Proximity analysis: 3 5 9
Regulation: 3 0
Psychological factors: 338-342
Rcliability: 3, 43, 2 9 3 , 3 10 - 3 1 2
Pulse-echo technique: 390 Pump:
Plant registcr
and failure-finding: 1 75- 1 85 Reliability-centrcd Maintenance (ReM):
desired performance: 22-23, 24
Decision Diagram: 200-20 1
different operating contexts: 28, 29
definition:
hidden failures: 92, 93
implementation strategies: 277-284
failure modes: 54, 56-57, 6 1 , 65-69
seven basic questions: 7
fail ure-finding: 1 78- 1 79, 1 92
Relief valve: 1 1 3 , 1 1 4
functions: 22-23
Repair time: 3 2 , 77
functional failures: 46 impeller failure: 54, 56-57
costs: 1 04, 1 08- 1 1 0 , 1 4 7 Reporting defects: 233-234
multiple failure: 1 1 8 , 1 2 1 , 1 22
Reporting failures: 2 5 1 -252
non-operational consequcnces: l OS - I I 0
Research into equipment failure: 4
operational consequences: 1 05 - 1 06
Resistance to slress: 59, 1 3 1
protective devices: I I I
Resnikoff cOllundrum: 2 5 2
redesign: 1 92
Response t i m e : 3 2
422
Reliability-centred Maintenance
Restoration: scheduled: see Scheduled restoration
Scheduled restoration: 1 1 , 1 3, 134-137, 1 68, 205, 238, 265
Index TQM: 2 1, 2 8 8
Ultimate cuntradiction: 252
Task:
Ultimate level switch: 1 1 5
Review groups: 1 7 , 266-269
frequency: 1 35
descriptions: 2 1 8- 2 1 9
Ultrasonic analysis: 360- 3 6 1
Rigid borescopes: 393
technical feasibility: 1 35 - 1 36
frequencies: see Frequencies
Ultrasonics:
Risk: 95- 1 0 1
worth doing: 1 36- 1 3 7
force: 278
pulse ccho technique: 390
proactive: see Proactive tasks
transmission technique: 390-39 1 resonance technique: 3 9 1
acceptabi Iity: 98- 1 0 1
Scrap rate: 295
continuum of: 1 1 8- 1 20
Sealed-[or-life bearings: 59, 7 1
packaging: 223
economic: 1 1 9, 343, 346
Second Generation: 2, I I
proposed: 206, 207
evaluation: 1 0 1
Secondary damage: 75-77, 1 1 0
selection process: 1 3 - 1 4, 1 69, 2 1 7
Secondary functions: 8,35, 37-44, 1 25
Teamwork: 5 , 20, 268, 3 1 5
Root causes of failure: 69, 1 59, 3 3 5 - 3 4 1
Sediment: 369-370
Technical:
Rotating disc electrode: 3 7 2
Selective approach to RCM : 278-282
Routine maintenance:
fatal: 98, 99, 343-347
and hidden functions: 1 20- 1 22 reduction: 3 1 2- 3 1 3
frequency modulation: 39 1 Ultra-violet spectroscopy: 3 8 0 Unavailability: 1 1 6- 1 1 8 , 294 United Airli nL�s: 320
characteristics: 1 5 , 1 29
Useful li!'e: 1 9, 1 32, 1 37 . 2 3 8 , 3 1 4
Senses: see Human senses
feasibility: see Tcchnically feasible
Utilisation: 295
Scrvice:
h istory records: 79-80, 259-260
customcr: see Customer service
Technically feasible: 1 4, 90-9 1 , 129-130,
423
Vapour induced scintillation: 3 8 6
Routine violations: 3 3 8 , 342
Seven questions: 7
205, 324
Velocity sensor: 3 5 1
Rule,based mistakes: 338, 340-341
Shape parameter: 243
failure-finding tasks: 1 85
Vibration analysis: 1 5 7. 3 5 1 ,3 6 1
Run-to-failure: 1 1 , 1 4
Shift arrangements: 3 0
on-condition tasks: 1 49
V i bration switch: 1 1 5
Running time: 2 3 1 -2 3 2
Shifted Weib u l l distribution: 245
scheduled discard tasks: 1 40
Violations: 3 ,' 8 , 342
Shock pulse monitoring: 359-360
scheduled restoration tasks: 1 35- 1 36
V i scosity monitor: 397-398
S-N curve: 243-245
Shutdowns: 3 1 1
Techniques of maintenance: 5
Viscosity
SPC: see Statistical process control
Significant assets: 279-280
Temperature effects: 1 50, 350, 399AOI
Visible absorption spectroscopy: 3i\O
3 97
Sabotage: 338, 342
Simultaneous learning: 269
Temperature indicating paint: 40 1
Safe-life limits: 1 38- 1 39
Skill-based errors: 339
Templating: 2 8 1
S afety: 3, 1 9, 38, 1 47 , 1 70, 279, 308-309
Skills in RCM : 29 1
Thermography: 399
Wear-anel-lear' 59, 1 59, 1 60- 1 6 1
Safety consequences: 3 , 1 0, 1 5 , 93, 94-
Slips: 3 3 8--339
Thin-layer activation: 380
Wei bull distribution: 242-243 Work -in-process: 29, 3 1 -3 2
Walk-around checks: 1 7 1 , 197
103, 204, 263
Smoke detector: 1 9 1
Third Generation: 2-5, 307
definition of: 94
Spares: 20, 32 - 33, 306
Time and hidden failures: 1 24- 1 25
Work packages: 2 1 3 , 221-224
and redesign: 1 02 , 1 90
Spectrometric oil analysis:
Time synchronous averaging analysis:
Worksheet
3 5 5 - 366
Safety first: 93
Spectrum: 35 1
Safety hazards: 30, 75
Spelling error: 1 1 9
Time waveform analysis: 354-355
S afety legislation: J 03
Specification limits: 27
Total failure: 47
Safety valve: see Relief valve
Spike cnergyTM: 358-359
Total quality management: see TQM
failure-finding tasks: 1 85
Sample size and actuarial analysis: 25 1
S taff turnover: 20, 3 1 6
Traditional view of failure: 1 1
on-condition tasks: 1 66
Scale parameter: 243
Standard operating procedure: 2 2 1 -222
Traditional approach to maintenance: 1 6
scheduled discard tasks: 1 3 9
Scanning auger electron microscopy:
Stand-by pump: see Pump
Training 2 2 1
scheduled restoration tasks: 1 36
Decision: 198-2 1 1 Information: 3 8 , 5 2 , 54, 89, 202 Worth doing: 1 5 , 90-92, 203-204, 324
Statistical process control: 1 5 1 - 1 52
needs analysis: 269
Scanning electron microscopy: 3 8 1
Steel mill: 3 1 0, 3 1 1
in RCM: 2 9 1
Schedule: see Maintenance schedule
Strain gauges: 396-397
Trip wire: 42, 1 1 5
X-ray radiographic fluoroscopy: 393
Scheduled discard: 1 1 , 1 3 , 137-140, 168,
Stress: see Applied stress
Truck: 26, 28, 8 1 -85
X-ray radiography: 392
205 , 238, 265
Strippable magnetic film: 389-390
Truncated Weibull distribution: 245
frequency: 1 38- 1 40
Structural integrity: 39
Turbine:
technical feasibility: 1 40
S upert1uous functions: 43
worth doing: 1 40
Survival distribution: 237-245
3 8 1 -382
Scheduled on-condition tasks:
s{'(' On-condition tasks Scheduled overhauls: 1 3, 143
see also Scheduled restoration
disc failure: 1 6 1 - 1 62 exhaust system: 44, 5 2 , 8 9
Sweet packing: 2 7 , 48
Turnover: see S taff turnover
System boundaries: 1 7 , 270, 334
Tyres: 1 60- 1 6 1 , 1 62 , 25 1 , 3 1 0
X-ray fluorescence spectroscopy: 374
Yield: 295, 302-303