Sensor Fusion That Works: Fusing Vibration, Thermal, Acoustic, and Oil Data
A practitioner's method for combining vibration, thermal, ultrasound, oil analysis, and operator rounds around named failure modes, plus planner rules for conflicting evidence.
It is Tuesday morning and a critical gearbox on the primary conveyor drive has three verdicts. The vibration analyst flags rising sideband energy around the stage-two mesh and wants the unit out at the next window. Thermography shot the same drive train the week before and saw nothing unusual. The oil report comes back with climbing iron and no water. The operator who has run that line for nine years says it sounds the same as always.
Sensor fusion is the practice of combining evidence from independent measurement physics (vibration, infrared thermography, ultrasound, oil analysis, motor current, and operator observation) around a single named failure mode, with explicit weighting rules and a written tie-break for when the signals disagree. It is not an averaging exercise across dashboards, and it is not a single asset health score. Done properly, it produces one decision owner, one evidence tier, and one work-order outcome.
What follows is the method I use: how to define the unit of fusion, which modality actually sees which failure mode, how to weight evidence without inventing a fake probability, what to do when signals conflict, and what the planner handoff must contain to survive a shutdown meeting.
The Gearbox That Three Technologies Disagreed About
In that hypothetical gearbox case, every specialist was correct inside their own modality. Sideband growth around gear mesh frequency is a legitimate mechanical indicator. An infrared survey taken at partial load on a well-ventilated housing can absolutely read normal while a tooth flank is degrading. Rising iron with no water points to dry wear rather than contamination ingress. The operator's ears are a real instrument, and they are tuned to the sound of the whole drive train, not one gear stage.
The failure was not analytical. It was organizational. Nobody owned the combined picture, so the scheduling decision defaulted to whoever spoke last in the planning meeting. When three technologies disagree and no one arbitrates, the plant does not choose the best evidence. It chooses the loudest voice or the cheapest option.
That is why I treat fusion as a definition and ownership problem first, and a math problem second. Before you compare any two signals, you need a named failure mode, a named operating state, a named arbiter, and a rule for what happens when the evidence splits. Skip those and you have built a more expensive way to argue.
Fusion Starts With a Named Failure Mode, Not a Dashboard
The unit of fusion is asset plus component plus failure mode plus operating state. Written out, that looks like: *motor drive-end bearing, spall propagation, loaded run above 70 percent of rated load*. Every signal you fuse must be qualified against that string. A vibration reading taken during unloaded coast-down is not evidence about spall propagation under load.
Pull those failure modes from work you already have. Most plants have FMEA worksheets, RCM analyses, or criticality studies sitting in a shared drive. Use them. If your equipment register already carries ISO 14224 style failure descriptions and failure codes, map to those instead of inventing a parallel taxonomy that only the reliability group understands. Federal facility operations and maintenance program guidance points teams toward documented program practices and planning resources rather than ad hoc method invention [5], and the same logic applies here: reuse the structure your planners and auditors already recognize.
Asset-level health scoring fails for a specific reason. A single index of 62 out of 100 tells the planner nothing about craft, task, duration, or parts. Decompose instead. A medium-voltage centrifugal pump might carry four monitored failure modes with four different evidence sources:
- Drive-end bearing spall, seen first by high-frequency vibration enveloping and airborne ultrasound, confirmed later by bearing housing temperature rise.
- Mechanical seal face wear, seen by seal pot level trend and leakage inspection on operator rounds, rarely visible in casing vibration until late.
- Impeller wear and recirculation, seen by head and flow deviation against the pump curve plus motor current trend, not by a single vibration band.
- Coupling misalignment, seen by running speed vibration and its harmonics with phase, confirmed by a laser alignment check during the next available window.
List the unmonitored modes explicitly. If nothing watches the seal faces on twelve of your eighteen pumps, write that down and show it to leadership. Silence on a dashboard reads as health, and that assumption is how plants get surprised by failures they never instrumented for.
Which Modality Actually Sees Your Failure Mode
Detection physics decides coverage, not sensor availability. Vibration resolves periodic mechanical defects because rolling contact and gear mesh generate repeating impulses. Ultrasound catches early friction, turbulence, and leakage because those events radiate high-frequency energy before bulk mechanical response changes. Infrared reads heat generation, which usually means a fault has already progressed to meaningful energy loss, though it is excellent for electrical connections and insulation. Oil analysis sees wear debris and additive depletion that no external sensor can infer. Motor current signature analysis sees electrical faults and load-path problems that casing sensors miss. Our condition monitoring guide walks through how each of these technologies gets set up on a route, which is the work that has to be right before fusing their output means anything.
| Failure mode | Primary evidence | Confirming evidence | Relative lead time | Best for |
|---|---|---|---|---|
| Rolling-element bearing spall | High-frequency vibration enveloping | Airborne ultrasound, housing temperature | Early to mid-stage | Medium and high-speed rotating assets |
| Gear tooth and mesh wear | Vibration sidebands and phase | Ferrography, gear oil iron trend | Mid-stage | Enclosed gearboxes with sampling ports |
| Lubricant degradation and contamination | Oil analysis (viscosity, water, particle count) | Ultrasound friction trend, bearing temperature | Earliest detectable | Assets with reliable sampling access |
| Electrical connection and insulation heating | Infrared thermography under load | Motor current signature, resistance testing | Mid to late stage | MCCs, switchgear, motor terminations |
| Steam trap and valve passing | Airborne and contact ultrasound | Thermal delta across the body | Early to mid-stage | Utility loops and compressed air systems |
| Misalignment and looseness | Running speed vibration and harmonics with phase | Thermography at coupling, operator rounds | Mid-stage | Coupled drive trains after rebuilds |
The lead-time ordering matters more than most fusion designs admit. Oil and ultrasound frequently move first. Thermal often moves last. If your rule weights all three equally, a late-stage thermal hit and an early-stage oil trend look like the same urgency, and the planner loses the ability to distinguish "schedule in the next window" from "get it out now."
Respect the blind spots. On low-speed assets, standard velocity spectra are weak and you will need acceleration enveloping or ultrasound. On sealed and grease-for-life units, oil sampling is impractical and you should stop pretending it is part of the plan. Our vibration analysis guide for rotating equipment covers the measurement setup choices that make these limits either manageable or fatal.
Weighting Evidence Without Inventing a False Probability
Do not publish a single percentage failure probability unless you have verified failure history and a validated model for that failure mode on that asset class. A confident numeric probability of bearing failure inside a fixed window is the fastest way to lose planner trust, because the early misses get remembered long after the model is corrected. Production machine learning guidance is blunt about this discipline: start simple, instrument the system, and measure before trusting model output [6].
Use evidence tiers instead. Plain language, defensible in a meeting:
- Tier 1, corroborated: two or more independent physics paths agree, operating context qualified, rate of change consistent across modalities.
- Tier 2, credible trend: single modality showing a consistent trend with plausible rate of change, context qualified.
- Tier 3, single reading: one reading at or above threshold, no trend confirmation yet.
- Tier 4, context only: operator concern, unreliable sensor, or data collected outside the defined operating state.
Independence is the rule people break most. Two accelerometers on the same bearing housing measuring the same energy path are not corroboration, they are redundancy. Vibration plus oil debris is independent. Vibration plus housing temperature is partially independent. Vibration plus a second vibration channel is one vote.
Gate every reading on operating context before it counts. Load, speed, ambient temperature, product changeover, and recent maintenance work all change what "normal" means. Commercial condition-monitoring services that pair vibration and temperature sensing in one device combine threshold logic with machine-learning analysis and still route abnormalities to a technician for investigation rather than automated action [7]. That is the right instinct. The human step exists because context rarely arrives with the reading.
Rules for When the Signals Disagree
Conflict is a trigger, not a cancellation
When two modalities disagree about a failure mode, the alert does not get closed. It gets an owner, a verification step, and a due date. Cancelling a Tier 2 vibration trend because thermography came back clean is the single most common way plants talk themselves out of a real finding, because thermal usually lags mechanical defect growth.
Write the conflict patterns down before they happen. These are the five I keep in the rule set:
1. Vibration up, oil clean. Likely a non-wear mechanical issue: looseness, resonance, alignment, or a mounting problem. Verification step is phase analysis and a physical inspection of mounts and coupling. Default action is an inspection work order, not a component replacement. 2. Oil debris up, vibration flat. Likely early distributed wear before a discrete defect forms. In slow wear progression, ferrography and particle count move before measurable spectral change. Verification is a repeat sample plus ferrography with wear particle typing. Default action is a shortened sampling interval and a scheduled inspection. 3. Thermal hot, everything else normal. Check load and ambient first, then lubrication quantity, then the electrical path. Verification is a repeat survey under comparable load with an emissivity check. Default action is a lubrication and electrical connection inspection. 4. Operator concern, all instruments normal. Treat this as Tier 4 evidence that earns an investigation, not dismissal. Verification is a targeted walkdown with the operator present and a route reading at the operating state they described. Sensor placement and route coverage are often the real gap. 5. All instruments alarming, no operator symptom. Suspect asset hierarchy or tag mapping errors before you suspect the machine. This is exactly the failure pattern described in why predictive alerts land on the wrong asset, and it burns credibility faster than a missed failure.
Name the arbiter. One reliability engineer, with a stated time limit (I use two working days for critical assets), records the conflict, the verification step chosen, and the reasoning. That reasoning goes in the CMMS so the next analyst inherits the logic instead of relitigating it.
The Planner Handoff: What Fused Evidence Must Contain
A fused alert that arrives as "high vibration, please check" is not planner-ready. Here is the minimum content set, written the way it should appear on the work order:
Asset: CV-102 Drive Gearbox (criticality A)
Failure mode: Stage-2 gear tooth wear, loaded run >70%
Evidence tier: Tier 1 (corroborated)
Contributing modalities:
- Vibration route 14 Mar, 28 Mar: mesh sidebands rising
- Oil sample 25 Mar: Fe 84 ppm trending up, water negative
Operating context: full load, ambient 22-26 C, no recent rebuild
Recommended task: Borescope stage-2 mesh, backlash check
Craft / duration: Millwright + vib tech, 4 h, line stopped
Parts required: Gasket kit GK-4471, 20 L ISO 320 gear oil
Verification: Photograph tooth flanks, record wear pattern,
resample oil post-inspection, close with finding codeEvery fused alert needs that verification line. Without it you never learn whether the fusion logic was right, and the rule set stops improving. Condition-monitoring workflows that close the loop follow the same path from reading to alert to investigation to corrective action to technician feedback [3], and the feedback step is what makes the next alert better. The measurement resolution record matters just as much as the alert itself [8].
Capture it in fields your system already has. In SAP PM, that means the notification damage code, cause code, and activity code plus the object part. In IBM Maximo, it is failure class, problem, cause, and remedy on the work order. In Fiix or UpKeep style systems, it is the failure code list plus a structured completion note template. Which fields you use matters less than whether the same three codes get filled every time, because coded outcomes are what let you test fusion rules later. CMMS integration earns its place here by eliminating manual evidence assembly, not by replacing the analyst who wrote the interpretation.
Planner-ready detail changes behavior in concrete ways. Tier 1 evidence with a named component justifies window selection and parts staging. Tier 3 evidence justifies a route recheck, not a shutdown scope change. That distinction is what stops predictive programs from becoming a queue of unschedulable suspicions.
A 90-Day Build Sequence for One Asset Class
Pick one asset class and 8 to 15 units. Fewer than that and your verification sample teaches you nothing. More than that and the rule writing stalls.
Weeks 1 to 3, failure mode inventory and modality gap review. List every credible failure mode per unit from existing FMEA or RCM material. Mark which modality currently sees each one, and mark the gaps in a separate column. Expect to find that half of your named modes have no coverage.
Weeks 4 to 6, baseline collection under known operating states. Collect readings with load, speed, and ambient recorded on every sample. A baseline without operating context is not a baseline. If you are adding new sensors and network paths in this phase, run the change through your OT architecture and security review rather than around it [4].
Weeks 7 to 10, evidence and conflict rules in writing. Draft the tier definitions, the independence rules, and the five conflict patterns for this asset class. Get the planner and the operations manager to read them out loud in a meeting. If they cannot repeat the rules back, the rules are too complicated.
Weeks 11 to 13, planner pilot with verification feedback. Route fused alerts through the real planning process. Require a verification outcome on every closure.
Set acceptance criteria before the pilot starts, not after:
- Verification coverage: share of fused alerts closed with a documented finding or no-finding outcome.
- Decision latency: median hours from fused alert to first documented maintenance decision.
- No-finding rate: share of alerts closed with nothing observed, broken out by evidence tier.
Do not add modalities until the first modality's data quality and asset hierarchy are clean. Adding oil analysis to a program with mismapped equipment tags multiplies the confusion rather than resolving it. Governance-minded frameworks for measuring and managing risk in AI-enabled systems make the same point about establishing measurement discipline before scaling reliance on automated output [2]. Set a quarterly review where the reliability engineer, planner, and operations lead examine no-finding rates by tier and decide whether any rule needs changing. Verified missed failures and repeated no-finding clusters are the only two triggers that should change a rule mid-cycle. The pilot-to-scale sequence for condition monitoring covers how this expands across sites once the first class is stable.
Frequently asked questions
Does sensor fusion require machine learning?
No. Written tier rules and conflict rules on trended data deliver most of the value. Machine learning helps when you have enough labeled failure history to validate it, and useful implementations combine threshold logic with learned models rather than replacing one with the other [7].
How many modalities are enough?
Two independent physics paths per critical failure mode. A third adds confidence but also adds cost and cognitive load. Single-sensor packages that combine vibration and temperature in one unit give you a starting point on assets that had nothing before [1].
What about low-speed and variable-load assets?
Gate on operating state and use acceleration enveloping, ultrasound, or oil analysis. Standard velocity spectra will mislead you on slow shafts.
Should fusion output a single number?
No. Output a tier, a named failure mode, and a recommended task. A number invites arguments about the number instead of decisions about the machine.
Which metric should I start tracking this week?
Median hours from fused alert to first documented maintenance decision, with the decision owner recorded. Alert counts measure activity. Decision latency measures whether fusion is actually changing maintenance behavior.
How do I fuse evidence when one modality has no coverage?
Write the gap down instead of inferring around it. An unmonitored failure mode is an open risk, not a healthy one, and the pilot-to-scale sequence for condition monitoring is where coverage gets added deliberately rather than opportunistically.
Failure Signals and Your Next 30 Minutes
Your next 30 minutes: pick one critical asset, write its top three failure modes on one page, and mark which modality currently sees each one. The gaps will be obvious and uncomfortable.
Back to the gearbox. Under tier rules, rising mesh sidebands across two routes plus rising iron with no water is Tier 1 corroborated evidence, because vibration and oil debris are independent physics. Clean thermography is expected, not contradictory, because thermal lags mechanical defect growth. The operator's "sounds fine" is Tier 4, logged and investigated during the walkdown. The decision was never close. It was just never owned. Write the rules, name the arbiter, and keep the quarterly review on the calendar.
References
[1] AWS, "What is Amazon Monitron? - Amazon Monitron", AWS documentation. https://docs.aws.amazon.com/Monitron/latest/user-guide/what-is-monitron.html
[2] NIST, "AI Risk Management Framework | NIST", 2023. https://www.nist.gov/itl/ai-risk-management-framework
[3] AWS, "The Amazon Monitron workflow - Amazon Monitron", AWS documentation. https://docs.aws.amazon.com/Monitron/latest/user-guide/deployed-workflow.html
[4] NIST, "Guide to Operational Technology (OT) Security | CSRC", 2023. https://csrc.nist.gov/pubs/sp/800/82/r3/final
[5] U.S. Department of Energy, "Operations and Maintenance in Federal Facilities | Department of Energy", DOE guidance. https://www.energy.gov/cmei/femp/operations-and-maintenance-federal-facilities
[6] Google, "Rules of Machine Learning: | Google for Developers", Google developer guidance. https://developers.google.com/machine-learning/guides/rules-of-ml?hl=en
[7] AWS, "How Amazon Monitron works - Amazon Monitron", AWS documentation. https://docs.aws.amazon.com/Monitron/latest/user-guide/how-monitron-works.html
[8] AWS, "Understanding sensor measurements and monitoring machine abnormalities - Amazon Monitron", AWS documentation. https://docs.aws.amazon.com/Monitron/latest/user-guide/anom-monitoring-chapter.html
Ready to put this into practice?
See how Monitory helps manufacturing teams implement these strategies.