Research · Faults in sensor history ·

How the historian benchmark is built

A pre-registered benchmark that scores fault detectors on data kept the way a process historian keeps it, from public fault datasets through a simulated historian, four ways of reading the archive and a shared false-alert budget to decision rules frozen before scoring.

A process historian does not keep every scan. It keeps a point only when the value moves outside a deadband (exception), and then only when a straight line through the kept points would miss it by more than a tolerance (swinging door). It stores that point some time after the scan. Most analytics tools read the archive by filling it back in onto a regular grid first.

The benchmark asks two things. Does reading the archive that way cost detection? And does a detector that reads the shape of a signal do better than the alarms and statistics already in use? This describes how it is built. Every dataset in it is public or simulated.

The parts

The pipeline runs from data in, through a simulated historian and its views, to detectors, calibration, scoring, and a record of what was found.

Stage Component What it does
Data Public fault datasets Tennessee Eastman (the Rieth simulation runs) and NASA C-MAPSS turbofan degradation
Data Dataset loaders Stream the files into arrays and split each dataset into fit, calibration, test-normal and test-fault runs
Historian Simulated historian Exception then swinging-door compression at a set tolerance, recording when each kept point is committed
Historian Archive The kept points, each with its commit time
Historian View builder Four readings of the archive (below)
Detectors Level detectors Limit alarm at 3 sigma, a limit alarm held to the shared budget, Shewhart and EWMA charts per tag
Detectors Variability and multivariate baselines Rolling variability, EWMA on the manipulated variables, PCA over all 52 process variables
Detectors Learned detectors Isolation Forest and the matrix profile, on gridded views
Detectors Shape watch Describes each window’s shape and scores its distance to the nearest normal window; a slow (2 hour) and a fast (1 hour) version
Scoring Alert shaping Takes the maximum over tags, applies a two-evaluation debounce, and stamps each score with when its inputs were available
Scoring Threshold calibration Finds the lowest threshold that keeps alerts on calibration normal time within budget
Scoring Scorer Merges alerts into episodes, then measures detection, lead over the limit alarm, delay, false alerts, AUROC and bootstrap intervals
Run Frozen protocol Every constant in one place, committed with a protocol page before scoring
Run Checkpointed runner Runs each compression setting across a process pool and saves it on finishing
Record Report script, log and results pages Applies the decision rules mechanically and writes tables, verdicts and an append-only log

The four views

View Reading
R Raw scans, before compression (the reference)
G1 The archive filled onto a 1-minute grid
G15 The archive interpolated onto a 15-minute grid
N1 The native archive, simplified at 2 percent of each tag’s span

The simulated historian

The compressor is the reason the benchmark exists, so it is checked against a reference implementation. Two details matter. First, its error bound is twice the swinging-door tolerance on that stage, and twice the tolerance plus twice the exception deadband end to end, not the single tolerance often quoted. Second, it returns each kept point’s commit time. A point is stored only once a later scan shows the door has closed, so a detector reading the archive cannot see a point before that moment. The alert-shaping stage carries that time forward, which keeps every lead honest about when a value could actually have been read.

Making the comparisons fair

One false-alert budget. Every detector, on every view, is calibrated to the same budget: one alert episode per asset per 720 hours of normal time. Thresholds are set on calibration runs that no test uses. A detector cannot look good by firing more often. The one exception is the 3-sigma limit alarm, kept as the familiar operator reference, with a budget-matched limit alarm beside it.

Same shaping for everyone. All detectors share the same debounce and the same rule for merging alerts within an hour into one episode.

Pre-registration. The questions, the rules that would refute each claim, the data splits and every constant are committed before any test data is scored. Any later change is logged with its reason on the protocol page.

Mechanical verdicts. A report script applies the frozen decision rules to the raw results and writes the tables. No verdict or table is edited by hand, and every table regenerates from its commit.

Resumable runs. The runner saves each compression setting as it finishes, so a restart resumes without losing results. Runs use a fork-based process pool with one numerical thread per worker.

Phases

Phase Data Question Status
1 Tennessee Eastman, C-MAPSS Does gridding cost detection? Does shape watch beat a limit alarm and a statistical baseline? Does one threshold serve every tag type? How much do answers depend on compression? Run
1b Tennessee Eastman, all 52 variables Does shape watch catch behaviour faults that level, variability and multivariate detectors miss? Is it earlier than a budget-matched limit alarm? Run
2 BattLeDIM (water network leaks), CARE to Compare (wind turbine SCADA) The same questions on heavier compression and real, mixed tags Planned
3 Several assets of one kind Can an alert be classified by its nearest labelled precedent on other assets of the same kind and a different size? Planned

What the public benchmarks showed

All of the following is on Tennessee Eastman (simulated) and C-MAPSS (public), and much of it is a null result, reported as such.

Question Finding
Does gridding the archive cost detection? Not on Tennessee Eastman; yes on C-MAPSS
Does shape watch beat a limit alarm and a statistical baseline at the same budget? No, on both datasets
Does one threshold serve every tag type? Yes on these simulated tags, which are cleaner than real ones
Does shape watch catch behaviour faults the other detectors miss? Only one Tennessee Eastman fault (fault 19); another was matched by a simple valve rule
Is shape watch earlier than a budget-matched limit alarm, and does a faster version close the gap on abrupt faults? No, and no
Does native reading keep false alerts steady under heavier compression? Contested (below)

Known weaknesses

Light compression. In Phase 1 the simulated historian kept 58 to 98 percent of scans, far more than a real archive keeps. Phase 1b added a heavier setting, to about 3 to 1. Phase 2 is planned on real SCADA data.

Look-ahead in N1. The native view simplifies a tag’s whole archive at once, so later data can reshape earlier windows, which a live system could not do. N1 has to be made causal before its leads are quoted. G15 has no look-ahead and serves as the cross-check.

Repeated normal time. The Tennessee Eastman faulty test runs share their first 8 hours across all 20 faults. In Phase 1b that meant 30 distinct normal segments rather than 600, and a single false alert was counted 20 times. That is why the compression question above is contested. The next protocol counts distinct segments only.

No labelled onset in C-MAPSS. Degradation is taken to start 125 cycles before failure, and the C-MAPSS lead figures are read against that convention.

One normal library per dataset. A site whose operating modes change may raise more false alerts than these datasets suggest.

What it is not

It is not a monitoring system: nothing here runs continuously. It is not a tuning harness, since no setting is chosen by looking at test results. It does not forecast remaining life, and it models no vendor’s historian beyond the published exception and swinging-door behaviour. It is also not a supervised classifier: every detector learns only from normal behaviour.