Resources/A Buyer's Rubric for Evaluating Condition Monitoring Options
Maintenance Strategies

A Buyer's Rubric for Evaluating Condition Monitoring Options

Score condition monitoring options on workflow fit, data requirements, operating ownership, implementation dependencies, and financial assumptions using your own plant inputs.

13 min read
By Daniel Ortega

A vendor demo hits its high point. A bearing fault curve climbs on the dashboard, the presenter notes the anomaly was flagged weeks before the machine came apart, and the room nods. The reliability engineer asks about sensor sample rates. The IT lead asks about the gateway. Nobody asks who in the plant would have written the work order.

To evaluate condition monitoring options properly, score each one on five dimensions using your own plant data: workflow fit, data requirements, operating ownership, implementation dependencies, and financial assumptions. Agree on the weights before the first demo. Supply your own evidence for every score. Treat any option that cannot produce a planner-ready work order as incomplete, no matter how good the detection story looks.

That is the whole argument of this article. What follows is the rubric, the evidence you have to gather, and the questions that separate a program that survives its second year from a feed that everyone mutes.

The Demo That Answers the Wrong Question

Demos are built to answer "can this technology detect a developing fault." That question is largely settled. Commercial condition monitoring stacks combine vibration and temperature sensing with threshold rules and machine learning analysis, and they route abnormalities to a person for investigation [1][7]. Detection is table stakes.

The question your evaluation has to answer is different: can this option produce a maintenance decision our planners will trust and schedule, on our assets, with our data, using the people we actually have on shift?

This is why vendor rankings and pre-filled payback decks are poor inputs. Those numbers come from someone else's asset mix, duty cycles, staffing ratios, and shutdown calendar. A payback figure derived from a continuous-process refinery tells a discrete assembly plant with a strong planning function almost nothing. Worse, an outside number short-circuits the internal argument you need to have about which assets matter and who owns the response.

So build the rubric from what you can verify. Five dimensions, each scored 1 to 5 against evidence the plant supplies, each weighted by your evaluation team before anyone sits through a slide deck. If you cannot produce the evidence for a dimension, that is not a vendor problem. That is a gap you need to fund.

Five Dimensions That Decide Whether Condition Monitoring Sticks

Score each dimension 1 to 5. Multiply by the agreed weight. Record the evidence source next to every score so the number survives scrutiny in a capital review three months later.

DimensionWhat You ScoreEvidence You SupplyWeight GuidanceDisqualifier
Workflow fitDistance from alert to a reviewable work order with failure mode, asset ID, and recommended taskTrace one real alert end to end during trialHeaviest for single-site plants with mature planningNo CMMS write-back path
Data requirementsAsset registry quality, signal availability, usable work order historyHierarchy export, tag list, failure code fill rate20 to 25 percentOption depends on failure coding you do not have
Operating ownershipNamed reviewer, planner, and deferral approver per criticality tierSigned one-page alert policy20 percentNo named alert owner
Implementation dependenciesNetwork, OT access, mounting windows, credentials, approvalsOT architecture review and shutdown calendarHeaviest for multi-site rolloutsRequires flat network access to control layer
Financial assumptionsDowntime cost, event history, repair cost split, lead timeFinance and CMMS records, not vendor decks15 to 20 percentCase only survives optimistic scenario

Weighting is a leadership decision, not a spreadsheet default. A single plant with three experienced planners and a clean asset hierarchy should weight workflow fit heaviest, because the constraint is decision throughput. A district manager rolling out to nine sites should weight implementation dependencies heaviest, because the constraint is repeatability across sites with different networks and different taxonomy discipline.

Two scoring discipline rules. First, never score a dimension on features the vendor lists. Score it on evidence you verified in a trial or a reference call with a plant that runs similar assets. Second, a low score is information, not automatic rejection. A 2 on data requirements tells you the first budget line is asset registry cleanup, not sensors.

Workflow Fit: Follow the Alert to the Work Order

Workflow fit is the distance between a condition signal and a work order a planner will schedule without re-doing the diagnosis. Measure that distance by tracing a single alert.

Here is the trace to demand in a trial, on a real asset. A slurry pump develops a rising vibration signature. The option flags an abnormality and proposes a suspected failure mode, not just "anomaly detected." A reviewer opens the case and sees operating context alongside the signal: flow rate, suction pressure, and whether the line was running the abrasive product grade that week. The reviewer confirms, and the system creates a work order request carrying the asset ID, the failure mode, the evidence link, and a recommended task. The planner reserves the mechanical seal kit and schedules the job into the next production gap. The technician inspects, records what was actually found, and that finding goes back into the case record. Published condition monitoring workflows follow this same shape: readings produce alerts, a person investigates, corrective action is taken, and the technician records the resolution so the system learns from the outcome [3][8].

Ask every vendor what fields the option can populate in your CMMS. For SAP PM, IBM Maximo, Fiix, or UpKeep, the minimum set is:

  • Functional location or asset ID that matches your hierarchy exactly, not a vendor-side label
  • Failure mode code from your own catalog so the record is usable in reliability analysis later
  • Priority derived from criticality tier, not from raw alert severity
  • Evidence link back to the trend and the reviewer's note
  • Recommended task with craft and estimated duration so the planner can slot it
  • Parts referenced against the storeroom record

Then ask the question that reveals whether the option is a monitoring tool or a maintenance tool: what happens when a technician inspects and finds nothing wrong? A serious option gives the technician a way to record that outcome and uses it to adjust future sensitivity on that asset [8]. A weak option keeps firing the same alert until the crew stops looking. That feedback path is also what turns a monitoring feed into a shared reference for the reliability team, similar in spirit to how technician knowledge capture preserves inspection findings.

Score a 5 when the option writes a reviewable work order request into your CMMS with evidence attached and an outcome field. Score a 1 when it sends email to a shared inbox.

Data Requirements: What the Plant Has to Supply Before Anything Works

Every condition monitoring option depends on three plant-supplied inputs, and the model quality behind it cannot compensate for weakness in any of them.

The three inputs are asset registry quality, signal availability from sensors or the historian, and work order history with usable failure coding. Complete this readiness checklist before you schedule another demo:

  • Asset hierarchy completeness: what percentage of candidate assets have a parent, a functional location, and a criticality rating
  • Tag naming consistency: can you map historian tags to asset IDs without a person interpreting abbreviations
  • Historian sample rates: are vibration and process tags stored at a rate that shows a developing trend, or only shift averages
  • Failure code fill rate: on last year's corrective work orders, how many carry a usable failure mode versus a free-text note
  • Run hours and operating context: can you retrieve run hours, product grade, and line speed for the period around a past failure

Use standards as a guide to what the records should support. ISO 55000:2024 sets out asset management terminology and principles [9]. ISO 17359:2018 covers general procedures for setting up a machine condition monitoring program [10]. ISO 14224:2016 describes equipment, failure, and maintenance data categories for petroleum, natural gas, and petrochemical operations; teams in other sectors can use its taxonomy as a reference when defining their own failure codes [11]. Federal operations and maintenance guidance also emphasizes documented program practices [5].

The opinionated point: any option that requires clean failure coding you do not have will underperform, regardless of the analytics behind it. Google's production machine learning guidance makes the same argument for engineering teams, that measurement infrastructure and data plumbing precede model sophistication [6]. Score the gap honestly and put a number against closing it.

Operating Ownership: Name the People Before You Sign

The rule is simple. Every recommendation needs a named reviewer, a named planner, and a named approver who can authorize deferral. Write it on one page before purchase.

A workable policy snippet reads like this. Criticality 1 assets: alerts reviewed within one shift, reviewer is the area reliability engineer, deferral requires maintenance manager approval with a documented reason and a re-review date. Criticality 2: reviewed within three working days, reviewer is the planner, deferral documented in the work order text. Criticality 3: reviewed at the weekly planning meeting, batched, no individual documentation required. Every dismissed alert on a Criticality 1 or 2 asset gets a one-line reason recorded against the asset.

Ownership scales with consequence, so tier it explicitly. Constrained production assets and safety-related equipment get short review windows and a named engineer. Balance-of-plant fans and non-critical pumps get batch review. If you cannot say who reviews a Criticality 1 alert at 2am on a Saturday, you do not have an ownership model yet.

The Failure Mode That Kills Most Programs

The most common post-purchase failure has nothing to do with detection accuracy. Alerts route to a distribution list, no single person owns triage, the first few investigations find nothing actionable, and within two months the plant treats the entire feed as noise. Once a crew decides the alerts are noise, no amount of model tuning brings the trust back. Assign named owners and review windows before the first sensor is mounted.

Score ownership by asking whether the option can enforce assignment, track time from alert to first documented decision, and log the stated reason a recommendation was deferred. Governance frameworks for AI-enabled systems make the same point about assigned accountability and measurable oversight rather than informal review [2].

Implementation Dependencies: The Line Items Nobody Puts in the Proposal

The proposal covers hardware, software, and services. It rarely covers what the plant must supply. Write those items down as your own project scope:

  • Network segmentation and OT firewall rules consistent with your control system architecture and security zones [4]
  • Historian or PLC read access, including who approves the data path and what is explicitly read-only
  • Sensor mounting and cabling, including surface prep, access platforms, and permit requirements
  • Spare parts data so recommended tasks can reference real storeroom records
  • CMMS API credentials and a test environment for work order write-back
  • Change management approval for any touch on control-layer equipment

Then the sequencing trap. Sensor installation on rotating equipment usually requires a production window or a shutdown slot, so your realistic start date is a scheduling question, not a procurement question. Check the shutdown calendar before you commit to a go-live date.

A workable phase sequence, with the gate that must close before advancing: scoping (plant owns criticality ranking and asset list, gate is a signed asset scope), connectivity (plant owns network and data access, vendor owns ingestion, gate is verified signals on a test asset), pilot assets (vendor owns detection, plant owns review discipline, gate is documented outcomes on every alert), planner workflow integration (plant owns CMMS fields and approval, gate is planner acceptance of generated work orders), and multi-site extension (gate is a repeatable install and policy package). The typical blocker at each gate is the same: an approval nobody scheduled. The failure patterns that show up when programs cross from a handful of assets to hundreds are covered in why condition monitoring programs fail at scale.

For district managers and VP Operations, decide up front what is standardized centrally and what stays local. Asset taxonomy, the criticality rubric, failure code catalog, and alert policy belong at the center. Sensor placement, review meeting cadence, and craft assignment stay local. Define pilot exit criteria before the pilot starts: alert precision after review, planner acceptance rate on generated work orders, and the share of alerts with a verified finding recorded.

Financial Assumptions: Build the Case From Your Numbers

Refuse vendor payback claims as an input. Replace them with a small set of drivers you can defend line by line in a capital review.

You need five: cost per hour of downtime on the constrained asset, historical unplanned events per year on the candidate asset list, average repair cost split between planned and reactive execution, expected detection lead time for the failure modes in scope, and internal labor hours for program ownership including review and administration. Every one of those comes from finance, the CMMS, or the historian. None comes from a slide.

*Illustrative modeled estimate:* on a candidate list of twelve rotating assets, converting a portion of reactive repairs to planned work produces avoided downtime hours and a lower average repair cost, and the case turns on how many events per year you actually convert. *Model inputs:* plant-supplied downtime cost per hour on the constrained asset, unplanned event count per asset from the last 24 months of work orders, planned versus reactive repair cost from CMMS cost history, assumed conversion rate of detected events into scheduled work, detection lead time long enough to reach the next production window, and annual internal review labor hours. Change the conversion rate alone and you will see whether the case is real or decorative.

Validate each assumption against a specific record. Downtime cost per hour should trace to a production accounting figure, not a hallway estimate. Event counts should trace to closed work orders with a failure classification. Repair cost split should trace to actual cost postings. Then run two scenarios, conservative and expected, and reject any business case that only survives the optimistic one.

Frequently Asked Questions

Should we pilot on the worst asset or a representative one?

Representative. The worst asset teaches you about one failure mode and often has a known root cause. A representative asset tests your review workflow, which is what you are actually buying.

How many assets belong in a first phase?

Enough to generate regular alerts so review discipline forms as a habit, small enough that one reviewer can handle triage without falling behind. Use your criticality ranking, not sensor cost, to choose.

Do we need a historian to start?

Not always, but you need operating context from somewhere. Wireless vibration and temperature sensing can stand alone for detection [1], while process context is what makes a recommendation credible to a planner.

What if our failure coding is poor?

Score that dimension low, budget the cleanup, and start capturing structured outcomes on every new alert. The feedback loop from technician findings back into the case record is how the dataset improves [8].

What types of condition monitoring should we score first?

Start with the techniques that cover the assets on your criticality list. Commercial stacks commonly pair vibration and temperature sensing with threshold rules and machine learning analysis [1][7], so score those against the same five dimensions rather than treating any one sensing method as the decision. Our condition monitoring guide walks through the common techniques and where each one fits.

Run the Rubric This Week

Take 30 minutes today. Pull your top ten criticality-ranked assets and fill the data readiness checklist for each one: hierarchy complete, tags mapped, historian rate, failure code fill, run hours available. Do that before you schedule another demo, because those five columns will change which questions you ask.

Start tracking one metric now, before any purchase: time from condition alert to first documented maintenance decision, plus the share of alerts with a recorded outcome. Both are measurable with your current process and both predict whether a new option will stick. They also give you a baseline to compare against once alerts start routing into work orders automatically.

Go back to the demo. A bearing fault flagged 40 days early is worth exactly nothing until a named planner turned it into scheduled work with the seal kit staged and the outage window booked. Detection was never the hard part.

The deliverable from your evaluation is one page per finalist: five weighted scores, the evidence source behind each score, and a signed list of the internal gaps you agreed to fund. Monitory should sit on that page under the same rubric as every other option, scored on how its alert-to-work-order routing, CMMS integration, and recorded technician feedback perform against your assets and your planners. If any vendor, including us, asks you to score them on anything other than your own evidence, that is your answer.

References

[1] AWS, "What is Amazon Monitron? - Amazon Monitron", AWS documentation. https://docs.aws.amazon.com/Monitron/latest/user-guide/what-is-monitron.html

[2] NIST, "AI Risk Management Framework | NIST", 2023. https://www.nist.gov/itl/ai-risk-management-framework

[3] AWS, "The Amazon Monitron workflow - Amazon Monitron", AWS documentation. https://docs.aws.amazon.com/Monitron/latest/user-guide/deployed-workflow.html

[4] NIST, "Guide to Operational Technology (OT) Security | CSRC", 2023. https://csrc.nist.gov/pubs/sp/800/82/r3/final

[5] U.S. Department of Energy, "Operations and Maintenance in Federal Facilities | Department of Energy", DOE guidance. https://www.energy.gov/cmei/femp/operations-and-maintenance-federal-facilities

[6] Google, "Rules of Machine Learning: | Google for Developers", Google developer guidance. https://developers.google.com/machine-learning/guides/rules-of-ml?hl=en

[7] AWS, "How Amazon Monitron works - Amazon Monitron", AWS documentation. https://docs.aws.amazon.com/Monitron/latest/user-guide/how-monitron-works.html

[8] AWS, "Understanding sensor measurements and monitoring machine abnormalities - Amazon Monitron", AWS documentation. https://docs.aws.amazon.com/Monitron/latest/user-guide/anom-monitoring-chapter.html

[9] ISO, "ISO 55000:2024 - Asset management - Vocabulary, overview and principles." https://www.iso.org/standard/83053.html

[10] ISO, "ISO 17359:2018 - Condition monitoring and diagnostics of machines - General guidelines." https://www.iso.org/standard/71194.html

[11] ISO, "ISO 14224:2016 - Petroleum, petrochemical and natural gas industries - Collection and exchange of reliability and maintenance data for equipment." https://www.iso.org/standard/64076.html

Ready to put this into practice?

See how Monitory helps manufacturing teams implement these strategies.