Verbatim Duplication in a Public Urban Exposome Corpus: A Data Audit and Corrected Evaluation of Wellbeing Classification

Authors

  • Rabinarayan Panda
  • Thalaimalaichamy M
  • I Hameem Shanavas
  • Jeyakumar S
  • A. Sivakumar
  • Madhawa Surendar

DOI:

https://doi.org/10.51483/IJAIML.6.8s.2026.489-499

Keywords:

Data Leakage, Verbatim Duplication, Urban Exposome, Wellbeing Classification, Benchmark Integrity, Negative Results, Wearable Sensing, Evaluation Protocol.

Abstract

Public sensor corpora underpin an expanding body of work on environmental exposure and mental wellbeing, yet the integrity of those corpora is rarely audited before models are trained on them. This paper reports a duplication audit of the Digital Exposome dataset, a multimodal corpus of seven environmental and four physiological channels recorded from forty participants walking predefined urban routes, distributed through Mendeley Data and redistributed on IEEE Data-Port. The audit establishes that 14,352 of the 42,436 released rows, or 33.8 percent, are verbatim copies of rows appearing earlier in the same file, arranged in seven contiguous blocks, and that the label is identical within 99.9 percent of copy pairs. The consequence for evaluation is severe. Under the contiguous eighty-twenty partition natural to a time-ordered corpus, every one of the 8,488 held-out rows is present verbatim in the training partition; under the random partition adopted in the corpus’s own benchmark, 51.9 percent are. A gated recurrent classifier trained on the released file attains a macro-averaged F1-score of 0.941, whereas the identical architecture trained on the de-duplicated corpus under a purged split with a training-fitted scaler attains 0.299, below the 0.535 accuracy of constant majority-class prediction. Further, an oracle persistence predictor that simply repeats the preceding label attains a macro-F1 of 0.980, and only 49 of 5,558 held-out windows, or 0.88 percent, exhibit a label change at the prediction step. We therefore report a negative result: momentary wellbeing is not recoverable from the released exposome channels above trivial temporal baselines, and previously published performance on this corpus is not interpretable as predictive capability. A reusable audit procedure and a reporting protocol for time-ordered sensor corpora are supplied.

Downloads

Published

2026-08-01

How to Cite

Panda, R., M, T., Shanavas, I. H., S, J., Sivakumar, A., & Surendar, M. (2026). Verbatim Duplication in a Public Urban Exposome Corpus: A Data Audit and Corrected Evaluation of Wellbeing Classification. International Journal of Artificial Intelligence and Machine Learning, 6(8s), 489–499. https://doi.org/10.51483/IJAIML.6.8s.2026.489-499