Research · Faults in sensor history ·
How the historian benchmark is built
A pre-registered benchmark that scores fault detectors on data kept the way a process historian keeps it, from public fault datasets through a simulated historian, four ways of reading the archive and a shared false-alert budget to decision rules frozen before scoring.
A process historian does not keep every scan. It keeps a point only when the value moves outside a deadband (exception), and then only when a straight line through the kept points would miss it by more than a tolerance (swinging door). It stores that point some time after the scan. Most analytics tools read the archive by filling it back in onto a regular grid first.
The benchmark asks two things. Does reading the archive that way cost detection? And does a detector that reads the shape of a signal do better than the alarms and statistics already in use? This describes how it is built. Every dataset in it is public or simulated.
The parts
The pipeline runs from data in, through a simulated historian and its views, to detectors, calibration, scoring, and a record of what was found.
| Stage | Component | What it does |
|---|---|---|
| Data | Public fault datasets | Tennessee Eastman (the Rieth simulation runs) and NASA C-MAPSS turbofan degradation |
| Data | Dataset loaders | Stream the files into arrays and split each dataset into fit, calibration, test-normal and test-fault runs |
| Historian | Simulated historian | Exception then swinging-door compression at a set tolerance, recording when each kept point is committed |
| Historian | Archive | The kept points, each with its commit time |
| Historian | View builder | Four readings of the archive (below) |
| Detectors | Level detectors | Limit alarm at 3 sigma, a limit alarm held to the shared budget, Shewhart and EWMA charts per tag |
| Detectors | Variability and multivariate baselines | Rolling variability, EWMA on the manipulated variables, PCA over all 52 process variables |
| Detectors | Learned detectors | Isolation Forest and the matrix profile, on gridded views |
| Detectors | Shape watch | Describes each window’s shape and scores its distance to the nearest normal window; a slow (2 hour) and a fast (1 hour) version |
| Scoring | Alert shaping | Takes the maximum over tags, applies a two-evaluation debounce, and stamps each score with when its inputs were available |
| Scoring | Threshold calibration | Finds the lowest threshold that keeps alerts on calibration normal time within budget |
| Scoring | Scorer | Merges alerts into episodes, then measures detection, lead over the limit alarm, delay, false alerts, AUROC and bootstrap intervals |
| Run | Frozen protocol | Every constant in one place, committed with a protocol page before scoring |
| Run | Checkpointed runner | Runs each compression setting across a process pool and saves it on finishing |
| Record | Report script, log and results pages | Applies the decision rules mechanically and writes tables, verdicts and an append-only log |
The four views
| View | Reading |
|---|---|
| R | Raw scans, before compression (the reference) |
| G1 | The archive filled onto a 1-minute grid |
| G15 | The archive interpolated onto a 15-minute grid |
| N1 | The native archive, simplified at 2 percent of each tag’s span |
The simulated historian
The compressor is the reason the benchmark exists, so it is checked against a reference implementation. Two details matter. First, its error bound is twice the swinging-door tolerance on that stage, and twice the tolerance plus twice the exception deadband end to end, not the single tolerance often quoted. Second, it returns each kept point’s commit time. A point is stored only once a later scan shows the door has closed, so a detector reading the archive cannot see a point before that moment. The alert-shaping stage carries that time forward, which keeps every lead honest about when a value could actually have been read.
Making the comparisons fair
One false-alert budget. Every detector, on every view, is calibrated to the same budget: one alert episode per asset per 720 hours of normal time. Thresholds are set on calibration runs that no test uses. A detector cannot look good by firing more often. The one exception is the 3-sigma limit alarm, kept as the familiar operator reference, with a budget-matched limit alarm beside it.
Same shaping for everyone. All detectors share the same debounce and the same rule for merging alerts within an hour into one episode.
Pre-registration. The questions, the rules that would refute each claim, the data splits and every constant are committed before any test data is scored. Any later change is logged with its reason on the protocol page.
Mechanical verdicts. A report script applies the frozen decision rules to the raw results and writes the tables. No verdict or table is edited by hand, and every table regenerates from its commit.
Resumable runs. The runner saves each compression setting as it finishes, so a restart resumes without losing results. Runs use a fork-based process pool with one numerical thread per worker.
Phases
| Phase | Data | Question | Status |
|---|---|---|---|
| 1 | Tennessee Eastman, C-MAPSS | Does gridding cost detection? Does shape watch beat a limit alarm and a statistical baseline? Does one threshold serve every tag type? How much do answers depend on compression? | Run |
| 1b | Tennessee Eastman, all 52 variables | Does shape watch catch behaviour faults that level, variability and multivariate detectors miss? Is it earlier than a budget-matched limit alarm? | Run |
| 2 | BattLeDIM (water network leaks), CARE to Compare (wind turbine SCADA) | The same questions on heavier compression and real, mixed tags | Planned |
| 3 | Several assets of one kind | Can an alert be classified by its nearest labelled precedent on other assets of the same kind and a different size? | Planned |
What the public benchmarks showed
All of the following is on Tennessee Eastman (simulated) and C-MAPSS (public), and much of it is a null result, reported as such.
| Question | Finding |
|---|---|
| Does gridding the archive cost detection? | Not on Tennessee Eastman; yes on C-MAPSS |
| Does shape watch beat a limit alarm and a statistical baseline at the same budget? | No, on both datasets |
| Does one threshold serve every tag type? | Yes on these simulated tags, which are cleaner than real ones |
| Does shape watch catch behaviour faults the other detectors miss? | Only one Tennessee Eastman fault (fault 19); another was matched by a simple valve rule |
| Is shape watch earlier than a budget-matched limit alarm, and does a faster version close the gap on abrupt faults? | No, and no |
| Does native reading keep false alerts steady under heavier compression? | Contested (below) |
Known weaknesses
Light compression. In Phase 1 the simulated historian kept 58 to 98 percent of scans, far more than a real archive keeps. Phase 1b added a heavier setting, to about 3 to 1. Phase 2 is planned on real SCADA data.
Look-ahead in N1. The native view simplifies a tag’s whole archive at once, so later data can reshape earlier windows, which a live system could not do. N1 has to be made causal before its leads are quoted. G15 has no look-ahead and serves as the cross-check.
Repeated normal time. The Tennessee Eastman faulty test runs share their first 8 hours across all 20 faults. In Phase 1b that meant 30 distinct normal segments rather than 600, and a single false alert was counted 20 times. That is why the compression question above is contested. The next protocol counts distinct segments only.
No labelled onset in C-MAPSS. Degradation is taken to start 125 cycles before failure, and the C-MAPSS lead figures are read against that convention.
One normal library per dataset. A site whose operating modes change may raise more false alerts than these datasets suggest.
What it is not
It is not a monitoring system: nothing here runs continuously. It is not a tuning harness, since no setting is chosen by looking at test results. It does not forecast remaining life, and it models no vendor’s historian beyond the published exception and swinging-door behaviour. It is also not a supervised classifier: every detector learns only from normal behaviour.