Dataset Viewer
Auto-converted to Parquet Duplicate
timestamp
stringdate
2005-07-24 00:00:00
2005-08-13 23:30:00
in_count
int64
0
61
out_count
int64
0
40
is_anomaly
int64
0
0
2005-07-24 00:00:00
0
0
0
2005-07-24 00:30:00
0
1
0
2005-07-24 01:00:00
0
0
0
2005-07-24 01:30:00
0
0
0
2005-07-24 02:00:00
0
0
0
2005-07-24 02:30:00
0
2
0
2005-07-24 03:00:00
0
0
0
2005-07-24 03:30:00
0
0
0
2005-07-24 04:00:00
0
0
0
2005-07-24 04:30:00
0
0
0
2005-07-24 05:00:00
0
0
0
2005-07-24 05:30:00
0
0
0
2005-07-24 06:00:00
0
0
0
2005-07-24 06:30:00
0
0
0
2005-07-24 07:00:00
0
0
0
2005-07-24 07:30:00
3
2
0
2005-07-24 08:00:00
0
0
0
2005-07-24 08:30:00
0
0
0
2005-07-24 09:00:00
1
0
0
2005-07-24 09:30:00
1
1
0
2005-07-24 10:00:00
0
0
0
2005-07-24 10:30:00
0
0
0
2005-07-24 11:00:00
0
1
0
2005-07-24 11:30:00
0
0
0
2005-07-24 12:00:00
0
0
0
2005-07-24 12:30:00
0
0
0
2005-07-24 13:00:00
0
0
0
2005-07-24 13:30:00
0
0
0
2005-07-24 14:00:00
0
0
0
2005-07-24 14:30:00
0
0
0
2005-07-24 15:00:00
0
0
0
2005-07-24 15:30:00
2
0
0
2005-07-24 16:00:00
0
0
0
2005-07-24 16:30:00
1
0
0
2005-07-24 17:00:00
1
1
0
2005-07-24 17:30:00
1
2
0
2005-07-24 18:00:00
4
4
0
2005-07-24 18:30:00
0
1
0
2005-07-24 19:00:00
0
2
0
2005-07-24 19:30:00
2
1
0
2005-07-24 20:00:00
0
0
0
2005-07-24 20:30:00
0
0
0
2005-07-24 21:00:00
0
3
0
2005-07-24 21:30:00
0
2
0
2005-07-24 22:00:00
0
0
0
2005-07-24 22:30:00
0
0
0
2005-07-24 23:00:00
0
2
0
2005-07-24 23:30:00
0
0
0
2005-07-25 00:00:00
0
1
0
2005-07-25 00:30:00
0
0
0
2005-07-25 01:00:00
0
0
0
2005-07-25 01:30:00
0
0
0
2005-07-25 02:00:00
0
0
0
2005-07-25 02:30:00
0
0
0
2005-07-25 03:00:00
0
0
0
2005-07-25 03:30:00
0
0
0
2005-07-25 04:00:00
0
0
0
2005-07-25 04:30:00
0
0
0
2005-07-25 05:00:00
0
2
0
2005-07-25 05:30:00
2
2
0
2005-07-25 06:00:00
2
1
0
2005-07-25 06:30:00
1
1
0
2005-07-25 07:00:00
2
0
0
2005-07-25 07:30:00
0
0
0
2005-07-25 08:00:00
0
0
0
2005-07-25 08:30:00
7
2
0
2005-07-25 09:00:00
6
1
0
2005-07-25 09:30:00
13
5
0
2005-07-25 10:00:00
16
4
0
2005-07-25 10:30:00
5
9
0
2005-07-25 11:00:00
3
1
0
2005-07-25 11:30:00
5
7
0
2005-07-25 12:00:00
9
11
0
2005-07-25 12:30:00
2
4
0
2005-07-25 13:00:00
12
7
0
2005-07-25 13:30:00
18
3
0
2005-07-25 14:00:00
10
14
0
2005-07-25 14:30:00
7
7
0
2005-07-25 15:00:00
17
10
0
2005-07-25 15:30:00
10
6
0
2005-07-25 16:00:00
8
13
0
2005-07-25 16:30:00
5
8
0
2005-07-25 17:00:00
0
9
0
2005-07-25 17:30:00
3
3
0
2005-07-25 18:00:00
1
6
0
2005-07-25 18:30:00
0
5
0
2005-07-25 19:00:00
4
2
0
2005-07-25 19:30:00
3
5
0
2005-07-25 20:00:00
0
1
0
2005-07-25 20:30:00
0
3
0
2005-07-25 21:00:00
3
0
0
2005-07-25 21:30:00
2
0
0
2005-07-25 22:00:00
0
1
0
2005-07-25 22:30:00
0
2
0
2005-07-25 23:00:00
0
3
0
2005-07-25 23:30:00
0
3
0
2005-07-26 00:00:00
0
0
0
2005-07-26 00:30:00
0
4
0
2005-07-26 01:00:00
0
0
0
2005-07-26 01:30:00
0
2
0
End of preview. Expand in Data Studio

mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale

Xiaona Zhou, Constantin Brif, Ismini Lourentzou

🎉 Accepted to TMLR 2026

📄 Paper · 📝 arXiv · 🌐 Project Page · 💻 Code


Overview

mTSBench is the largest and most diverse benchmark for multivariate time series anomaly detection (MTS-AD) and model selection. It contains 344 labeled multivariate time series from 19 datasets across 12 application domains, and evaluates 24 anomaly detectors and 3 unsupervised model selection methods under a unified evaluation suite.

No single anomaly detector performs consistently well across datasets, and current unsupervised model selection methods remain far from optimal.


Key Idea

Multivariate time series anomaly detection is challenging because of high-dimensional dependencies, cross-correlations between time-dependent variables, and the scarcity of labeled anomalies. Because no single detector dominates across datasets, choosing the right detector for a new time series is a problem in its own right.

mTSBench is the first benchmark to integrate unsupervised model selection methods and evaluate them alongside anomaly detectors under consistent settings, with a unified evaluation suite of point-based and ranking-based metrics.


Highlights

  • Largest MTS-AD benchmark: 344 labeled multivariate time series from 19 datasets across 12 application domains, including healthcare, cybersecurity, industrial monitoring, spacecraft telemetry, and IT infrastructure.

  • 24 anomaly detectors: Spanning reconstruction-, prediction-, statistics-, and LLM-based approaches, including the only two publicly available LLM-based methods for multivariate time series.

  • 3 unsupervised model selection methods: MetaOD, FMMS, and Orthus.

  • Unified evaluation suite: Point-based and ranking-based metrics for anomaly detection and model selection, including VUS-PR, VUS-ROC, AUC-PR, AUC-ROC, AUC-PTRT, and Precision@k, Recall@k, and NDCG@k.

  • Key finding: Even the strongest selectors remain far below the near-optimal baseline, and a single detector, PCA, surpasses all three selectors on 9 of 13 metrics.


Dataset Overview

Dataset Config Domain #TS #Dim Length #AnomPts #AnomSeqs License
CIC-IDS-2017 cicids Cybersecurity 5 73 > 100K 0–8656 0–2546 Citation Required
CalIt2 CalIt2 Smart Building 1 3 > 5K 0 21 CC BY 4.0
CreditCard creditcard Finance / Fraud Detection 1 30 > 100K 219 10 Citation Required
Daphnet Daphnet Healthcare 26 10 > 50K 0 1–16 CC BY 4.0
Exathlon Exathlon IT Infrastructure 30 21 > 50K 0–4 0–6 Apache 2.0
GECCO GECCO Industrial Process 1 10 > 50K 0 37 Citation Required
GHL GHL Industrial Process 14 17 > 100K 0 1–4 Contact Authors
Genesis Genesis Industrial Process 1 19 > 5K 0 2 CC BY-NC-SA 4.0
GutenTAG GutenTAG Synthetic Benchmark 30 21 > 10K 0 1–3 MIT
MITDB MITDB Healthcare 47 3 > 500K 0 1–720 ODC-By v1.0
MSL MSL Spacecraft Telemetry 26 56 > 5K 0 1–3 BSD 3-Clause
Metro metro Transportation 1 6 > 10K 20 5 CC BY 4.0
OPPORTUNITY OPPORTUNITY Human Activity Recognition 13 33 > 25K 0 1 CC BY 4.0
Occupancy room-occupancy Smart Building 2 6 > 5K 1–3 9–13 CC BY 4.0
PSM PSM IT Infrastructure 1 27 > 50K 0 39 CC BY 4.0
SMAP SMAP Spacecraft Telemetry 48 26 > 5K 0 1–3 BSD 3-Clause
SMD SMD IT Infrastructure 18 39 > 10K 0 4–24 MIT
SVDB SVDB Healthcare 78 3 > 100K 0 2–678 ODC-By v1.0
SWAN-SF swan Astrophysics 1 39 > 50K 5233 1382 MIT

#AnomPts and #AnomSeqs are the numbers of anomalous points and anomalous sequences per time series. Licenses are those of the original datasets.

Each folder corresponds to one dataset and contains *_train.csv, *_val.csv, and *_test.csv files. Each CSV contains a timestamp column, dataset-specific feature columns, and a binary is_anomaly label. The *_val.csv files are held-out validation time series used to train the model selection methods. See data_summary.csv for per-file statistics.


Usage

This repository uses Git LFS for the CSV files:

git lfs install
git clone https://huggingface.co/datasets/PLAN-Lab/mTSBench

Or load one of the 19 configurations with 🤗 Datasets (one configuration per dataset, so files with different schemas are not combined):

from datasets import load_dataset

calit2 = load_dataset("PLAN-Lab/mTSBench", "CalIt2")

df_train = calit2["train"].to_pandas()
df_test = calit2["test"].to_pandas()

Resources


Citation

If you find mTSBench useful in your research, please consider citing our work:

@article{zhou2026mtsbench,
  title={mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale},
  author={Zhou, Xiaona and Brif, Constantin and Lourentzou, Ismini},
  journal={Transactions on Machine Learning Research},
  year={2026}
}
Downloads last month
1,044

Paper for PLAN-Lab/mTSBench