Datasets:
timestamp stringdate 2005-07-24 00:00:00 2005-08-13 23:30:00 | in_count int64 0 61 | out_count int64 0 40 | is_anomaly int64 0 0 |
|---|---|---|---|
2005-07-24 00:00:00 | 0 | 0 | 0 |
2005-07-24 00:30:00 | 0 | 1 | 0 |
2005-07-24 01:00:00 | 0 | 0 | 0 |
2005-07-24 01:30:00 | 0 | 0 | 0 |
2005-07-24 02:00:00 | 0 | 0 | 0 |
2005-07-24 02:30:00 | 0 | 2 | 0 |
2005-07-24 03:00:00 | 0 | 0 | 0 |
2005-07-24 03:30:00 | 0 | 0 | 0 |
2005-07-24 04:00:00 | 0 | 0 | 0 |
2005-07-24 04:30:00 | 0 | 0 | 0 |
2005-07-24 05:00:00 | 0 | 0 | 0 |
2005-07-24 05:30:00 | 0 | 0 | 0 |
2005-07-24 06:00:00 | 0 | 0 | 0 |
2005-07-24 06:30:00 | 0 | 0 | 0 |
2005-07-24 07:00:00 | 0 | 0 | 0 |
2005-07-24 07:30:00 | 3 | 2 | 0 |
2005-07-24 08:00:00 | 0 | 0 | 0 |
2005-07-24 08:30:00 | 0 | 0 | 0 |
2005-07-24 09:00:00 | 1 | 0 | 0 |
2005-07-24 09:30:00 | 1 | 1 | 0 |
2005-07-24 10:00:00 | 0 | 0 | 0 |
2005-07-24 10:30:00 | 0 | 0 | 0 |
2005-07-24 11:00:00 | 0 | 1 | 0 |
2005-07-24 11:30:00 | 0 | 0 | 0 |
2005-07-24 12:00:00 | 0 | 0 | 0 |
2005-07-24 12:30:00 | 0 | 0 | 0 |
2005-07-24 13:00:00 | 0 | 0 | 0 |
2005-07-24 13:30:00 | 0 | 0 | 0 |
2005-07-24 14:00:00 | 0 | 0 | 0 |
2005-07-24 14:30:00 | 0 | 0 | 0 |
2005-07-24 15:00:00 | 0 | 0 | 0 |
2005-07-24 15:30:00 | 2 | 0 | 0 |
2005-07-24 16:00:00 | 0 | 0 | 0 |
2005-07-24 16:30:00 | 1 | 0 | 0 |
2005-07-24 17:00:00 | 1 | 1 | 0 |
2005-07-24 17:30:00 | 1 | 2 | 0 |
2005-07-24 18:00:00 | 4 | 4 | 0 |
2005-07-24 18:30:00 | 0 | 1 | 0 |
2005-07-24 19:00:00 | 0 | 2 | 0 |
2005-07-24 19:30:00 | 2 | 1 | 0 |
2005-07-24 20:00:00 | 0 | 0 | 0 |
2005-07-24 20:30:00 | 0 | 0 | 0 |
2005-07-24 21:00:00 | 0 | 3 | 0 |
2005-07-24 21:30:00 | 0 | 2 | 0 |
2005-07-24 22:00:00 | 0 | 0 | 0 |
2005-07-24 22:30:00 | 0 | 0 | 0 |
2005-07-24 23:00:00 | 0 | 2 | 0 |
2005-07-24 23:30:00 | 0 | 0 | 0 |
2005-07-25 00:00:00 | 0 | 1 | 0 |
2005-07-25 00:30:00 | 0 | 0 | 0 |
2005-07-25 01:00:00 | 0 | 0 | 0 |
2005-07-25 01:30:00 | 0 | 0 | 0 |
2005-07-25 02:00:00 | 0 | 0 | 0 |
2005-07-25 02:30:00 | 0 | 0 | 0 |
2005-07-25 03:00:00 | 0 | 0 | 0 |
2005-07-25 03:30:00 | 0 | 0 | 0 |
2005-07-25 04:00:00 | 0 | 0 | 0 |
2005-07-25 04:30:00 | 0 | 0 | 0 |
2005-07-25 05:00:00 | 0 | 2 | 0 |
2005-07-25 05:30:00 | 2 | 2 | 0 |
2005-07-25 06:00:00 | 2 | 1 | 0 |
2005-07-25 06:30:00 | 1 | 1 | 0 |
2005-07-25 07:00:00 | 2 | 0 | 0 |
2005-07-25 07:30:00 | 0 | 0 | 0 |
2005-07-25 08:00:00 | 0 | 0 | 0 |
2005-07-25 08:30:00 | 7 | 2 | 0 |
2005-07-25 09:00:00 | 6 | 1 | 0 |
2005-07-25 09:30:00 | 13 | 5 | 0 |
2005-07-25 10:00:00 | 16 | 4 | 0 |
2005-07-25 10:30:00 | 5 | 9 | 0 |
2005-07-25 11:00:00 | 3 | 1 | 0 |
2005-07-25 11:30:00 | 5 | 7 | 0 |
2005-07-25 12:00:00 | 9 | 11 | 0 |
2005-07-25 12:30:00 | 2 | 4 | 0 |
2005-07-25 13:00:00 | 12 | 7 | 0 |
2005-07-25 13:30:00 | 18 | 3 | 0 |
2005-07-25 14:00:00 | 10 | 14 | 0 |
2005-07-25 14:30:00 | 7 | 7 | 0 |
2005-07-25 15:00:00 | 17 | 10 | 0 |
2005-07-25 15:30:00 | 10 | 6 | 0 |
2005-07-25 16:00:00 | 8 | 13 | 0 |
2005-07-25 16:30:00 | 5 | 8 | 0 |
2005-07-25 17:00:00 | 0 | 9 | 0 |
2005-07-25 17:30:00 | 3 | 3 | 0 |
2005-07-25 18:00:00 | 1 | 6 | 0 |
2005-07-25 18:30:00 | 0 | 5 | 0 |
2005-07-25 19:00:00 | 4 | 2 | 0 |
2005-07-25 19:30:00 | 3 | 5 | 0 |
2005-07-25 20:00:00 | 0 | 1 | 0 |
2005-07-25 20:30:00 | 0 | 3 | 0 |
2005-07-25 21:00:00 | 3 | 0 | 0 |
2005-07-25 21:30:00 | 2 | 0 | 0 |
2005-07-25 22:00:00 | 0 | 1 | 0 |
2005-07-25 22:30:00 | 0 | 2 | 0 |
2005-07-25 23:00:00 | 0 | 3 | 0 |
2005-07-25 23:30:00 | 0 | 3 | 0 |
2005-07-26 00:00:00 | 0 | 0 | 0 |
2005-07-26 00:30:00 | 0 | 4 | 0 |
2005-07-26 01:00:00 | 0 | 0 | 0 |
2005-07-26 01:30:00 | 0 | 2 | 0 |
mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale
Xiaona Zhou, Constantin Brif, Ismini Lourentzou
🎉 Accepted to TMLR 2026
📄 Paper · 📝 arXiv · 🌐 Project Page · 💻 Code
Overview
mTSBench is the largest and most diverse benchmark for multivariate time series anomaly detection (MTS-AD) and model selection. It contains 344 labeled multivariate time series from 19 datasets across 12 application domains, and evaluates 24 anomaly detectors and 3 unsupervised model selection methods under a unified evaluation suite.
No single anomaly detector performs consistently well across datasets, and current unsupervised model selection methods remain far from optimal.
Key Idea
Multivariate time series anomaly detection is challenging because of high-dimensional dependencies, cross-correlations between time-dependent variables, and the scarcity of labeled anomalies. Because no single detector dominates across datasets, choosing the right detector for a new time series is a problem in its own right.
mTSBench is the first benchmark to integrate unsupervised model selection methods and evaluate them alongside anomaly detectors under consistent settings, with a unified evaluation suite of point-based and ranking-based metrics.
Highlights
Largest MTS-AD benchmark: 344 labeled multivariate time series from 19 datasets across 12 application domains, including healthcare, cybersecurity, industrial monitoring, spacecraft telemetry, and IT infrastructure.
24 anomaly detectors: Spanning reconstruction-, prediction-, statistics-, and LLM-based approaches, including the only two publicly available LLM-based methods for multivariate time series.
3 unsupervised model selection methods: MetaOD, FMMS, and Orthus.
Unified evaluation suite: Point-based and ranking-based metrics for anomaly detection and model selection, including VUS-PR, VUS-ROC, AUC-PR, AUC-ROC, AUC-PTRT, and Precision@k, Recall@k, and NDCG@k.
Key finding: Even the strongest selectors remain far below the near-optimal baseline, and a single detector, PCA, surpasses all three selectors on 9 of 13 metrics.
Dataset Overview
| Dataset | Config | Domain | #TS | #Dim | Length | #AnomPts | #AnomSeqs | License |
|---|---|---|---|---|---|---|---|---|
| CIC-IDS-2017 | cicids |
Cybersecurity | 5 | 73 | > 100K | 0–8656 | 0–2546 | Citation Required |
| CalIt2 | CalIt2 |
Smart Building | 1 | 3 | > 5K | 0 | 21 | CC BY 4.0 |
| CreditCard | creditcard |
Finance / Fraud Detection | 1 | 30 | > 100K | 219 | 10 | Citation Required |
| Daphnet | Daphnet |
Healthcare | 26 | 10 | > 50K | 0 | 1–16 | CC BY 4.0 |
| Exathlon | Exathlon |
IT Infrastructure | 30 | 21 | > 50K | 0–4 | 0–6 | Apache 2.0 |
| GECCO | GECCO |
Industrial Process | 1 | 10 | > 50K | 0 | 37 | Citation Required |
| GHL | GHL |
Industrial Process | 14 | 17 | > 100K | 0 | 1–4 | Contact Authors |
| Genesis | Genesis |
Industrial Process | 1 | 19 | > 5K | 0 | 2 | CC BY-NC-SA 4.0 |
| GutenTAG | GutenTAG |
Synthetic Benchmark | 30 | 21 | > 10K | 0 | 1–3 | MIT |
| MITDB | MITDB |
Healthcare | 47 | 3 | > 500K | 0 | 1–720 | ODC-By v1.0 |
| MSL | MSL |
Spacecraft Telemetry | 26 | 56 | > 5K | 0 | 1–3 | BSD 3-Clause |
| Metro | metro |
Transportation | 1 | 6 | > 10K | 20 | 5 | CC BY 4.0 |
| OPPORTUNITY | OPPORTUNITY |
Human Activity Recognition | 13 | 33 | > 25K | 0 | 1 | CC BY 4.0 |
| Occupancy | room-occupancy |
Smart Building | 2 | 6 | > 5K | 1–3 | 9–13 | CC BY 4.0 |
| PSM | PSM |
IT Infrastructure | 1 | 27 | > 50K | 0 | 39 | CC BY 4.0 |
| SMAP | SMAP |
Spacecraft Telemetry | 48 | 26 | > 5K | 0 | 1–3 | BSD 3-Clause |
| SMD | SMD |
IT Infrastructure | 18 | 39 | > 10K | 0 | 4–24 | MIT |
| SVDB | SVDB |
Healthcare | 78 | 3 | > 100K | 0 | 2–678 | ODC-By v1.0 |
| SWAN-SF | swan |
Astrophysics | 1 | 39 | > 50K | 5233 | 1382 | MIT |
#AnomPts and #AnomSeqs are the numbers of anomalous points and anomalous sequences per time series. Licenses are those of the original datasets.
Each folder corresponds to one dataset and contains *_train.csv, *_val.csv, and *_test.csv files. Each CSV contains a timestamp column, dataset-specific feature columns, and a binary is_anomaly label. The *_val.csv files are held-out validation time series used to train the model selection methods. See data_summary.csv for per-file statistics.
Usage
This repository uses Git LFS for the CSV files:
git lfs install
git clone https://huggingface.co/datasets/PLAN-Lab/mTSBench
Or load one of the 19 configurations with 🤗 Datasets (one configuration per dataset, so files with different schemas are not combined):
from datasets import load_dataset
calit2 = load_dataset("PLAN-Lab/mTSBench", "CalIt2")
df_train = calit2["train"].to_pandas()
df_test = calit2["test"].to_pandas()
Resources
- Paper: mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale (TMLR 2026)
- arXiv: arxiv.org/abs/2506.21550
- Project Page: plan-lab.github.io/projects/mtsbench
- Code: github.com/PLAN-Lab/mTSBench
Citation
If you find mTSBench useful in your research, please consider citing our work:
@article{zhou2026mtsbench,
title={mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale},
author={Zhou, Xiaona and Brif, Constantin and Lourentzou, Ismini},
journal={Transactions on Machine Learning Research},
year={2026}
}
- Downloads last month
- 1,044