A platinum concentrator I worked with spent close to four million rand on a predictive maintenance platform. Vibration sensors on the critical drives, oil analysis fed in, a dashboard that lit up beautifully in the boardroom. Eleven months later they had one confirmed catch — a mill pinion — and a maintenance team that had quietly gone back to running on the weekly route sheet.
The model was fine. The implementation was not, and the failure points were entirely predictable.
Your asset register is probably lying to you
Nothing kills a condition-monitoring programme faster than a register that does not match the plant. Duplicated equipment numbers, pumps that were swapped in 2019 and never updated, motors listed at the wrong frame size, functional locations that describe an area rather than an asset.
The model reads history against an asset ID. If three physically different pumps have shared that ID over six years, the failure history is fiction. I now insist on a register walkdown before any sensor is specified. It is unglamorous work — two technicians, a tablet and three weeks — and it is the highest-return activity in the whole project.
Failure codes nobody fills in properly
Ask your CMMS what caused your last fifty pump failures and count how many come back as "breakdown" or "other". In most South African plants it is over half.
Without a coded failure mode, you cannot train a model, and more importantly you cannot do the basic Pareto analysis that would have told you for free that 60% of your seal failures trace to one installation practice. Artificial intelligence is not a substitute for a fitter who writes down what he actually found.
The fix is cultural and small: a short, closed list of failure modes per asset class, no free-text-only closures, and a supervisor who bounces the ones that say "fixed".
The decision rights problem
This is the one that sinks otherwise good programmes. The system flags a rising bearing signature on a critical fan on a Saturday night. Who is allowed to take that fan out of service?
If the answer is "the engineering manager, on Monday", your predictive capability is decorative. The alert has to arrive with a pre-agreed action, a named authority level, and a defined production consequence. That is a governance exercise, not a technical one, and it has to be settled before commissioning — ideally written into the maintenance strategy and signed by production.
Where I have seen this done well, the plant agreed a simple three-tier response: advisory (log and inspect at next opportunity), planned (schedule within the next shutdown window), and urgent (shift supervisor may stop the machine without further approval). Three tiers, one page, no ambiguity at 03:00.
Start where failure is expensive and frequent
Sensor everything and you will drown. The economics only work where an unplanned failure carries real consequence and the failure mode develops slowly enough to be seen.
Good candidates in the plants I work in: large rotating equipment with long lead-time spares, transformers where dissolved gas trends give months of warning, critical process pumps in duty-standby arrangements where the standby is quietly also degraded. Poor candidates: cheap redundant equipment, anything with a fast-onset electrical failure mode, and assets you were going to replace in eighteen months anyway.
Run a pilot on ten to fifteen assets. Prove a catch. Put a rand value on the avoided event and publish it internally. Expansion becomes an easy conversation.
What this means for the people
The skills shift is real but it is not what the vendors imply. You do not need data scientists on site. You need condition-monitoring technicians who understand the physics of the failure mode and can interrogate a trend sceptically — someone who looks at a rising temperature and asks whether the ambient changed, whether the load profile shifted, whether the transmitter is drifting.
That is a trainable skill and it sits closer to a good millwright than to a statistician. The organisations getting value are the ones investing in that middle layer rather than hiring a modelling team and hoping the plant catches up.
The technology is genuinely capable. It just cannot compensate for a register nobody trusts, failure codes nobody completes, and a plant where nobody is allowed to stop the machine.

