Every alarm has to earn its place on the console.
Alarm rationalization is the stage of the ANSI/ISA-18.2 (IEC 62682) alarm-management lifecycle where each proposed or existing alarm is tested against the plant's alarm philosophy and either justified, re-specified or removed. An alarm is a signal that tells an operator about an abnormal condition requiring a timely response — so the test is concrete: what is the consequence of doing nothing, what response is expected, and is there time to make it? An alarm with no distinct operator response is not a rationalized alarm; it is an event that has been given a horn.
What rationalization actually decides, alarm by alarm.
| Metric | What it means | Commonly used target |
|---|---|---|
| Average alarm rate | Annunciated alarms per operator during steady operation. | About 150 per day, roughly 1 per 10 minutes; over 2 per 10 minutes is generally considered unmanageable. |
| Alarm flood | A burst that arrives faster than anyone can read, let alone act on. | More than 10 alarms in 10 minutes per operator is the usual flood threshold; time spent in flood should be well under 1%. |
| Chattering alarm | Repeatedly annunciates and clears in quick succession — typically three or more times in a minute. | Target zero. A handful of chattering tags routinely generate the majority of an entire console's alarm count. |
| Fleeting alarm | Annunciates and clears so quickly the operator cannot respond, without repeating. | Target zero; usually a deadband, delay or signal-quality problem. |
| Stale alarm | Stays annunciated continuously for more than 24 hours. | A small handful at most on any day, each with an owner and a plan — not a permanent decoration on the display. |
| Top-10 contributors | Share of total alarm load coming from the ten most frequent tags. | Should be a small fraction of the load. When it is half, you have ten problems, not an alarm-management problem. |
What “rationalized” means
Not “we reduced the count”. It means every alarm on the console has a documented cause, consequence, response, response time and priority, agreed by operations, process and safety, and recorded in the master alarm database. That record, not the configuration, is what the system is measured against.
Where the reduction comes from
Mostly from three moves: deleting alarms with no operator response, fixing the small number of chattering and duplicate tags that generate most of the load, and suppressing by design — state-based alarming so that alarms which are meaningless while a unit is down do not arrive while it is down.
Shelving is not suppression
Shelving is an operator-initiated, time-limited, logged removal of a nuisance alarm that comes back automatically and appears on a review list. Suppression by design is an engineered, documented rule. A permanent shelf with no expiry is neither; it is a silently disabled alarm with a paper trail nobody reads.
Why alarm systems drift back.
- Rationalizing once as a project, then never auditing. Monitoring and assessment are lifecycle stages; without them the count climbs back within a couple of years.
- Priority inflation. Every owner argues their alarm is high priority, and a console where nothing is low priority is a console with no priorities.
- Shelving used as a mute button, with no expiry and no review list.
- Alarms added straight into the running system for a one-off situation, bypassing management of change and never removed.
- Using alarms as reminders, shift notes or event logging. If it is not an abnormal condition that needs a timely response, the log is the right place for it.
- Ignoring the flood problem because average rate looks acceptable. Averages hide the ten minutes after a trip, which is exactly when the operator most needs a readable console.
- Discarding alarm history after 30 days, which makes it impossible to prove either the problem or the improvement.
What the alarm system depends on.
Alarm performance is measured from event history, so it inherits the quality of the data layer: consistent timestamps, one time base, and the asset context that lets an alarm from the DCS join a state change from the line, which is what contextualization and a Unified Namespace exist to provide, with ISA-95 supplying the equipment hierarchy the counts roll up through. Downstream, alarm floods are usually the first evidence available to a root cause analysis, repeat nuisance alarms on the same asset are often a maintenance signal that nobody has read as one, alarm-driven stops show up as availability loss in the OEE breakdown, and which alarms are currently shelved or suppressed is one of the items that must survive a shift handover.
Most consoles are one honest week of analysis away from being usable.
Alarm Triage reads an alarm and event export you supply, groups the floods, ranks the chattering, fleeting and stale contributors, and drafts candidate rationalization actions your team can accept or reject. It reasons over the log you give it — it does not connect to your DCS or change any configuration. Five free runs, then it is part of the agents plan.
See Alarm Triage