Self-tuned healthy homogeneous core: Addressing heterogeneities in biomedical datasets
Filtering out borderline healthy samples improves disease diagnosis algorithms across heart and neurological conditions.
This methodological ML paper demonstrates that curating a stable, self-tuned healthy reference subset (H2C) using kernel density estimation reduces the impact of borderline healthy samples on disease classifier performance across three biomedical domains (cardiac, arrhythmia, migraine). The classifier-agnostic approach offers a practical upstream curation strategy relevant to any ML diagnostic system trained on healthy vs diseased cohorts.
What the study was
- Study design
- Technical/methodological study (validated on 3 biomedical datasets)
- Population
- PTB-XL+ cardiac ECG dataset; arrhythmia classification dataset; institutional migraine dataset (ASU-Mayo)
- Category
- Diagnostics
- Maturity
- Exploratory
- Journal
- Artificial Intelligence in Medicine
Why it surfaced
Solid methodological contribution to ML biomedical classification — addresses an often-ignored problem (healthy cohort heterogeneity) with a self-tuned approach showing consistent F1 improvements. Score of 4 (STANDARD borderline) reflects that clinical impact is indirect; applications in cardiac and neurological diagnostics are notable.
A plain-language summary of published research — not medical advice. Talk to a clinician about your own care.