Heartbeat Anomaly Detector β research release
This five-class PyTorch model classifies recorded heart-sound patterns. It does not diagnose diseases or estimate heart-attack or sudden cardiac-event risk.
Research and development only. Exploratory test accuracy is 76.52%, with macro F1 0.613. Always predicting the majority class achieved 60.87% accuracy (macro F1 0.151); this release exceeds that baseline on both metrics, unlike the previous release. The selected model has zero recall on extrasystole: all 10 extrasystole-labeled test recordings were predicted as normal. These results do not establish suitability for diagnosis, screening, triage, monitoring or treatment decisions.
Changes from the previous release
Previous (research_20260928_212714) |
Current (research_20260929_104950) |
|
|---|---|---|
| Selected model | wave_10sec_seed44 |
wave_unweighted_seed48 |
| Training device | CPU | CUDA (NVIDIA RTX 4050) |
| Configurations Γ seeds | 8 Γ 3 | 10 Γ 3, then re-run at 10 Γ 10 for this release |
| Test accuracy | 49.57% (below majority baseline) | 76.52% (above majority baseline) |
| Test macro F1 | 0.540 | 0.613 |
| Extrasystole test recall | 4 / 10 correct | 0 / 10 correct |
| Uncertainty rejection | Disabled (no threshold met the calibration target) | Enabled (threshold 0.366, ~97.4% coverage) |
| Original-only vs. combined-data comparison | Combined data appeared to hurt accuracy (68.18% vs. 63.64%, 3 seeds) | Combined data consistently helped (65.00% vs. 74.09% mean accuracy, 10 seeds, won 10/10 seed pairs) |
GPU training used the identical audited protocol (same fixed split, same audit checks, same epoch/patience budget) β it changed the compute device and seed count, not the methodology. The 10-seed re-run exists because the 3-seed comparisons above disagreed with each other on the data-combination question; more seeds were run specifically to check whether either conclusion was noise. See Training and calibration and Evaluation.
Two additional configurations (wave_balanced, wave_10sec_balanced β class-balanced batch sampling plus stronger augmentation, aimed at the minority-class weakness below) were tried and did not help: mean validation macro F1 0.45β0.47, clearly below every other configuration (0.57β0.66). Not selected.
Release
| Property | Value |
|---|---|
| Author / contact | Tathagata Mitra β tathagata.mitra@gmail.com |
| Release date | 2026-09-29 |
| Experiment | research_20260929_104950 |
| Selected model | wave_unweighted_seed48 |
| Best checkpoint | Epoch 54; cap 60 epochs, early-stopping patience 15 |
| Input | Native-rate mono/stereo recording, resampled to 2,000 Hz |
| Windows | 5 seconds / 10,000 samples |
| Output order | artifact, extrahs, murmur, normal, extrasystole |
| Training runtime | Custom PyTorch, CUDA (NVIDIA RTX 4050) |
| Inference runtime | Custom PyTorch, CPU (unchanged β single-recording latency does not need a GPU) |
| Checkpoint SHA-256 | f351df03efd74fccb2d06f86c7023949c76f35f6a55417842f69d899f344c75c |
extrahs and extrasystole remain separate source labels. A normal prediction does not establish cardiac health. Probabilities are model scores, not probabilities of a disease.
Usage
Python 3.12 was used for verification. Download matching code and weights. This is not a Transformers AutoModel or pipeline() package; config.json describes the custom implementation.
git clone https://huggingface.co/ai-mitra/heartbeat-anomaly-detector
cd heartbeat-anomaly-detector
pip install -r requirements.txt
Run from that directory:
from audio_processing import load_research_audio
from model_inference import HeartSoundPredictor
predictor = HeartSoundPredictor() # uses bundled active_model.json
audio, sample_rate = load_research_audio("your_recording.wav")
result = predictor.predict_with_details(audio, sample_rate)
print(result["predicted_class"])
print(result["probabilities"])
print(result["decision"])
Alternatively, install huggingface-hub and download a snapshot:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="ai-mitra/heartbeat-anomaly-detector",
local_dir="heartbeat-model",
allow_patterns=["*.py", "*.json", "*.pt", "requirements.txt", "README.md", "LICENSE"],
)
Install requirements and run from the downloaded directory. The two MP3s retained from the previous release are legacy examples with unverified provenance and diagnostic labels; their filenames are not clinical ground truth. They were not used to establish this release's performance.
Local API
uvicorn inference:app --host 127.0.0.1 --port 8000
curl "http://127.0.0.1:8000/hb?filename=your_recording.wav"
GET /health reports model availability. GET /hb accepts a server-local file path and returns ai_prediction, classification, descriptive acoustic measurements and plot information. GET /plot/{filename} retrieves a generated plot. This path-based API is for local use, not an authenticated public upload service.
Clients must accept five probability keys and the fields decision, uncertain and uncertainty_threshold. Previous disease-severity and cardiac-risk interpretations have been removed. Acoustic timing measurements are separate from classification and are not validated clinical measurements. Descriptive analysis uses up to 30 seconds; model inference processes the full recording in windows.
The project demo is heartbeat.kalkut.in; its deployed model version has not been verified against this release.
Model and preprocessing
Four Conv1d blocks have 32, 64, 128 and 256 channels with batch normalization, ReLU, pooling and dropout. Adaptive average pooling feeds a 256 β 128 β 64 β 5 classifier.
Inference averages stereo channels, resamples to 2 kHz, removes the mean, applies a fourth-order 20β500 Hz Butterworth bandpass with zero-phase filtering, divides by three times RMS and clips to [-1, 1]. Short recordings are zero-padded; longer recordings use consecutive 5-second windows and an end-aligned final window as needed. Softmax probabilities are averaged across windows before temperature calibration.
The package shares preprocessing with the training implementation. CPU inference was checked against all 115 saved test predictions (maximum absolute probability difference: 0.0). Training for this release used CUDA; the bundled version-2 predictor still runs on CPU regardless of local CUDA availability, since single-recording inference latency is negligible either way.
Data and provenance
| Source | Artifact | Extra HS | Murmur | Normal | Extrasystole | Total |
|---|---|---|---|---|---|---|
| Author-supplied original collection | 40 | 19 | 34 | 31 | 0 | 124 |
| PASCAL Dataset B | 0 | 0 | 95 | 320 | 46 | 461 |
The author reports obtaining the original collection from three clinics in Bengaluru, Karnataka, India, with labels assigned by clinic doctors. Review or adjudication details were unavailable. Collection duration was reported as between 8 and 21 months; exact dates, recording devices and anatomical sites are unavailable. Provenance is author-reported, not independently verified. The original class counts also match PASCAL Dataset A's published training counts; count agreement alone does not establish whether recordings are identical or independent.
The additional recordings came from Dataset B training archives of the PASCAL Classifying Heart Sounds Challenge. This evaluation uses a local split, not the official challenge test or leaderboard. Public availability does not establish independence from the original collection.
Two near-identical original recordings with conflicting labels were excluded without relabeling. 583 files remained: 300 training, 94 validation, 74 calibration, 115 test. This split is identical to the previous release β retraining did not change the data or the partition, only the compute device, configuration set and seed count. dataset_summary.json contains aggregate class counts. Filename/audio groups were kept together; they are not verified patient identities. Near-duplicate screening cannot exclude every partial or transformed copy.
For the original collection, the author reports research-use permission. Formal ethics approval or exemption documentation was not available to the author; patient consent status is unknown. Patient identifiers and demographics are unavailable. Training audio and record-level manifests are not distributed here, and this release does not claim that the unresolved matters have been cleared.
Training and calibration
Ten configurations were compared using seeds 42β51 (10 seeds): enhanced/simple preprocessing, weighted/unweighted loss, augmentation on/off, 5/10-second windows, MFCC CNN, log-mel CNN, engineered-feature XGBoost, and two variants adding class-balanced batch sampling plus stronger augmentation (random time-shift and time-masking) at 5 and 10 seconds. The two balanced-sampling variants scored below every other configuration and were not selected β see Changes from the previous release. Twenty further runs (10 seeds Γ 2 data conditions) compared original-only and combined-common-class training. Seeds share one split; this is not patient-level cross-validation.
| Configuration | Mean validation macro F1 (10 seeds) | SD |
|---|---|---|
wave_unweighted |
0.664 | 0.061 |
wave_10sec |
0.626 | 0.027 |
xgboost |
0.617 | 0.014 |
wave_simple |
0.586 | 0.031 |
wave_noaugment |
0.584 | 0.028 |
mfcc |
0.584 | 0.042 |
wave_enhanced |
0.575 | 0.030 |
logmel |
0.569 | 0.046 |
wave_balanced |
0.466 | 0.049 |
wave_10sec_balanced |
0.452 | 0.042 |
The selected model used unweighted cross-entropy (the previous release used class-weighted loss), Adam (learning rate 0.001, weight decay 0.0001), batch size 16, random 5-second crops, amplitude/noise augmentation and learning-rate reduction on validation plateaus. Selection maximized mean validation macro F1 across seeds, then selected the best validation seed within that configuration. The test set did not choose the checkpoint, but was inspected in earlier work, so these results remain exploratory. Unweighted loss tends to favor majority classes over minority ones; that tradeoff is visible directly in the extrasystole result below.
Temperature 1.670007 was fitted on the separate calibration partition (74 recordings; calibration accuracy 78.38%, macro F1 0.568). Unlike the previous release, a rejection threshold did meet the calibration target this time: uncertainty threshold 0.366381 (uncertain: true below this confidence), accepting 97.4% of test recordings at 76.8% accuracy among accepted records. uncertain: false still does not establish a reliable prediction.
Evaluation
| Metric | Value |
|---|---|
| Test recordings | 115 |
| Accuracy | 76.52% |
| Macro F1, five labels | 0.613 |
| Accuracy bootstrap 95% interval | 68.55β84.47% |
| Macro F1 bootstrap 95% interval | 0.479β0.703 |
| Majority-class accuracy / macro F1 | 60.87% / 0.151 |
Intervals resample filename/audio groups, conditional on the fitted model. They are not patient-level intervals and do not include total model-selection uncertainty. Small class counts, source differences and class imbalance limit interpretation.
Per-class test performance:
| Class | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| artifact | 0.89 | 1.00 | 0.94 | 8 |
| extrahs | 0.40 | 1.00 | 0.57 | 2 |
| murmur | 0.75 | 0.72 | 0.73 | 25 |
| normal | 0.78 | 0.86 | 0.82 | 70 |
| extrasystole | 0.00 | 0.00 | 0.00 | 10 |
The model never predicts extrasystole. All 10 extrasystole-labeled test recordings were classified as normal. This did not happen in the previous release (4/10 correct there), and it is a direct consequence of selecting the unweighted-loss configuration: it produces better accuracy and macro F1 overall, but at the cost of a whole class going unrecognized. Anyone relying on this release to flag extrasystole will not get a signal for it.
By data source, test performance was uneven: recordings from the author-supplied original collection scored 86.36% accuracy (macro F1 0.648, n=22), while PASCAL Dataset B recordings scored 74.19% accuracy but macro F1 only 0.295 (n=93) β largely because PASCAL B contributes all 10 test extrasystole recordings and none are recognized.
On the same original-source subset of 22 test recordings with four common labels, now run at 10 seeds (previously 3):
| Training data | Mean accuracy | Mean macro F1 |
|---|---|---|
| Original only | 65.00% | 0.545 |
| Combined (original + PASCAL B, common classes) | 74.09% | 0.669 |
Combined training data won on 10 of 10 seed pairs (mean advantage +9.1 accuracy points; paired comparison across seeds, t=4.74, p=0.0011). This reverses the previous release's finding, which reported that additional data did not help β that finding used only 3 seeds and appears to have been small-sample noise rather than a real effect. As before, seeds share one split rather than being independent patient cohorts, so this is a within-split repeated-optimization comparison, not a formal significance test over independent samples; the consistency across all 10 pairs is what makes it a meaningfully stronger signal than the previous 3-seed result, not the p-value in isolation.
External validation, verified patient-level separation and clinical expert error review remain unavailable. No PhysioNet 2016 challenge score is reported: validated binary references, complete clean/noisy annotations and required weights are unavailable. Five-class labels cannot substitute for those annotations.
Files and compatibility
heartbeat-anomaly-detector-model.pt: calibrated five-class checkpoint.active_model.json,config.json: location, hash, label order and preprocessing metadata.model.py,research_pipeline.py,model_inference.py,audio_processing.py: architecture and inference.inference.py: local FastAPI entry point.evaluation.json,dataset_summary.json,training_protocol.json,calibration.json: aggregate evidence.release_validation.json: export verification results.
The checkpoint filename is retained, but its window length and hash changed from the previous release (5-second windows now vs. 10-second previously); code and weights from this same revision must be used together β the previous checkpoint used class-weighted loss and 10-second windows and cannot be swapped in without also swapping config.json. The prior release remains accessible through repository history. Exact retraining requires the author's local training workspace and authorized data access; see MODEL_TRAINING_GUIDE.md.
Citation and license
@misc{mitra2026heartbeat,
author = {Mitra, Tathagata},
title = {Heartbeat Anomaly Detector: Research Heart-Sound Classification},
year = {2026},
howpublished = {https://huggingface.co/ai-mitra/heartbeat-anomaly-detector},
note = {September 2026 research release; cite the downloaded revision}
}
Dataset reference: Bentley, P., Nordehn, G., Coimbra, M. and Mannor, S. (2011), The PASCAL Classifying Heart Sounds Challenge 2011 (CHSC2011) Results. Challenge and citation.
The existing MIT license is retained. It does not grant rights to third-party datasets or resolve permissions for source recordings. Training-data redistribution is not included in this release.
- Downloads last month
- 40