Heartbeat Anomaly Detector β€” research release

This five-class PyTorch model classifies recorded heart-sound patterns. It does not diagnose diseases or estimate heart-attack or sudden cardiac-event risk.

Research and development only. Exploratory test accuracy is 76.52%, with macro F1 0.613. Always predicting the majority class achieved 60.87% accuracy (macro F1 0.151); this release exceeds that baseline on both metrics, unlike the previous release. The selected model has zero recall on extrasystole: all 10 extrasystole-labeled test recordings were predicted as normal. These results do not establish suitability for diagnosis, screening, triage, monitoring or treatment decisions.

Changes from the previous release

Previous (research_20260928_212714) Current (research_20260929_104950)
Selected model wave_10sec_seed44 wave_unweighted_seed48
Training device CPU CUDA (NVIDIA RTX 4050)
Configurations Γ— seeds 8 Γ— 3 10 Γ— 3, then re-run at 10 Γ— 10 for this release
Test accuracy 49.57% (below majority baseline) 76.52% (above majority baseline)
Test macro F1 0.540 0.613
Extrasystole test recall 4 / 10 correct 0 / 10 correct
Uncertainty rejection Disabled (no threshold met the calibration target) Enabled (threshold 0.366, ~97.4% coverage)
Original-only vs. combined-data comparison Combined data appeared to hurt accuracy (68.18% vs. 63.64%, 3 seeds) Combined data consistently helped (65.00% vs. 74.09% mean accuracy, 10 seeds, won 10/10 seed pairs)

GPU training used the identical audited protocol (same fixed split, same audit checks, same epoch/patience budget) β€” it changed the compute device and seed count, not the methodology. The 10-seed re-run exists because the 3-seed comparisons above disagreed with each other on the data-combination question; more seeds were run specifically to check whether either conclusion was noise. See Training and calibration and Evaluation.

Two additional configurations (wave_balanced, wave_10sec_balanced β€” class-balanced batch sampling plus stronger augmentation, aimed at the minority-class weakness below) were tried and did not help: mean validation macro F1 0.45–0.47, clearly below every other configuration (0.57–0.66). Not selected.

Release

Property Value
Author / contact Tathagata Mitra β€” tathagata.mitra@gmail.com
Release date 2026-09-29
Experiment research_20260929_104950
Selected model wave_unweighted_seed48
Best checkpoint Epoch 54; cap 60 epochs, early-stopping patience 15
Input Native-rate mono/stereo recording, resampled to 2,000 Hz
Windows 5 seconds / 10,000 samples
Output order artifact, extrahs, murmur, normal, extrasystole
Training runtime Custom PyTorch, CUDA (NVIDIA RTX 4050)
Inference runtime Custom PyTorch, CPU (unchanged β€” single-recording latency does not need a GPU)
Checkpoint SHA-256 f351df03efd74fccb2d06f86c7023949c76f35f6a55417842f69d899f344c75c

extrahs and extrasystole remain separate source labels. A normal prediction does not establish cardiac health. Probabilities are model scores, not probabilities of a disease.

Usage

Python 3.12 was used for verification. Download matching code and weights. This is not a Transformers AutoModel or pipeline() package; config.json describes the custom implementation.

git clone https://huggingface.co/ai-mitra/heartbeat-anomaly-detector
cd heartbeat-anomaly-detector
pip install -r requirements.txt

Run from that directory:

from audio_processing import load_research_audio
from model_inference import HeartSoundPredictor

predictor = HeartSoundPredictor()  # uses bundled active_model.json
audio, sample_rate = load_research_audio("your_recording.wav")
result = predictor.predict_with_details(audio, sample_rate)
print(result["predicted_class"])
print(result["probabilities"])
print(result["decision"])

Alternatively, install huggingface-hub and download a snapshot:

from huggingface_hub import snapshot_download
snapshot_download(
    repo_id="ai-mitra/heartbeat-anomaly-detector",
    local_dir="heartbeat-model",
    allow_patterns=["*.py", "*.json", "*.pt", "requirements.txt", "README.md", "LICENSE"],
)

Install requirements and run from the downloaded directory. The two MP3s retained from the previous release are legacy examples with unverified provenance and diagnostic labels; their filenames are not clinical ground truth. They were not used to establish this release's performance.

Local API

uvicorn inference:app --host 127.0.0.1 --port 8000
curl "http://127.0.0.1:8000/hb?filename=your_recording.wav"

GET /health reports model availability. GET /hb accepts a server-local file path and returns ai_prediction, classification, descriptive acoustic measurements and plot information. GET /plot/{filename} retrieves a generated plot. This path-based API is for local use, not an authenticated public upload service.

Clients must accept five probability keys and the fields decision, uncertain and uncertainty_threshold. Previous disease-severity and cardiac-risk interpretations have been removed. Acoustic timing measurements are separate from classification and are not validated clinical measurements. Descriptive analysis uses up to 30 seconds; model inference processes the full recording in windows.

The project demo is heartbeat.kalkut.in; its deployed model version has not been verified against this release.

Model and preprocessing

Four Conv1d blocks have 32, 64, 128 and 256 channels with batch normalization, ReLU, pooling and dropout. Adaptive average pooling feeds a 256 β†’ 128 β†’ 64 β†’ 5 classifier.

Inference averages stereo channels, resamples to 2 kHz, removes the mean, applies a fourth-order 20–500 Hz Butterworth bandpass with zero-phase filtering, divides by three times RMS and clips to [-1, 1]. Short recordings are zero-padded; longer recordings use consecutive 5-second windows and an end-aligned final window as needed. Softmax probabilities are averaged across windows before temperature calibration.

The package shares preprocessing with the training implementation. CPU inference was checked against all 115 saved test predictions (maximum absolute probability difference: 0.0). Training for this release used CUDA; the bundled version-2 predictor still runs on CPU regardless of local CUDA availability, since single-recording inference latency is negligible either way.

Data and provenance

Source Artifact Extra HS Murmur Normal Extrasystole Total
Author-supplied original collection 40 19 34 31 0 124
PASCAL Dataset B 0 0 95 320 46 461

The author reports obtaining the original collection from three clinics in Bengaluru, Karnataka, India, with labels assigned by clinic doctors. Review or adjudication details were unavailable. Collection duration was reported as between 8 and 21 months; exact dates, recording devices and anatomical sites are unavailable. Provenance is author-reported, not independently verified. The original class counts also match PASCAL Dataset A's published training counts; count agreement alone does not establish whether recordings are identical or independent.

The additional recordings came from Dataset B training archives of the PASCAL Classifying Heart Sounds Challenge. This evaluation uses a local split, not the official challenge test or leaderboard. Public availability does not establish independence from the original collection.

Two near-identical original recordings with conflicting labels were excluded without relabeling. 583 files remained: 300 training, 94 validation, 74 calibration, 115 test. This split is identical to the previous release β€” retraining did not change the data or the partition, only the compute device, configuration set and seed count. dataset_summary.json contains aggregate class counts. Filename/audio groups were kept together; they are not verified patient identities. Near-duplicate screening cannot exclude every partial or transformed copy.

For the original collection, the author reports research-use permission. Formal ethics approval or exemption documentation was not available to the author; patient consent status is unknown. Patient identifiers and demographics are unavailable. Training audio and record-level manifests are not distributed here, and this release does not claim that the unresolved matters have been cleared.

Training and calibration

Ten configurations were compared using seeds 42–51 (10 seeds): enhanced/simple preprocessing, weighted/unweighted loss, augmentation on/off, 5/10-second windows, MFCC CNN, log-mel CNN, engineered-feature XGBoost, and two variants adding class-balanced batch sampling plus stronger augmentation (random time-shift and time-masking) at 5 and 10 seconds. The two balanced-sampling variants scored below every other configuration and were not selected β€” see Changes from the previous release. Twenty further runs (10 seeds Γ— 2 data conditions) compared original-only and combined-common-class training. Seeds share one split; this is not patient-level cross-validation.

Configuration Mean validation macro F1 (10 seeds) SD
wave_unweighted 0.664 0.061
wave_10sec 0.626 0.027
xgboost 0.617 0.014
wave_simple 0.586 0.031
wave_noaugment 0.584 0.028
mfcc 0.584 0.042
wave_enhanced 0.575 0.030
logmel 0.569 0.046
wave_balanced 0.466 0.049
wave_10sec_balanced 0.452 0.042

The selected model used unweighted cross-entropy (the previous release used class-weighted loss), Adam (learning rate 0.001, weight decay 0.0001), batch size 16, random 5-second crops, amplitude/noise augmentation and learning-rate reduction on validation plateaus. Selection maximized mean validation macro F1 across seeds, then selected the best validation seed within that configuration. The test set did not choose the checkpoint, but was inspected in earlier work, so these results remain exploratory. Unweighted loss tends to favor majority classes over minority ones; that tradeoff is visible directly in the extrasystole result below.

Temperature 1.670007 was fitted on the separate calibration partition (74 recordings; calibration accuracy 78.38%, macro F1 0.568). Unlike the previous release, a rejection threshold did meet the calibration target this time: uncertainty threshold 0.366381 (uncertain: true below this confidence), accepting 97.4% of test recordings at 76.8% accuracy among accepted records. uncertain: false still does not establish a reliable prediction.

Evaluation

Metric Value
Test recordings 115
Accuracy 76.52%
Macro F1, five labels 0.613
Accuracy bootstrap 95% interval 68.55–84.47%
Macro F1 bootstrap 95% interval 0.479–0.703
Majority-class accuracy / macro F1 60.87% / 0.151

Intervals resample filename/audio groups, conditional on the fitted model. They are not patient-level intervals and do not include total model-selection uncertainty. Small class counts, source differences and class imbalance limit interpretation.

Per-class test performance:

Class Precision Recall F1 Support
artifact 0.89 1.00 0.94 8
extrahs 0.40 1.00 0.57 2
murmur 0.75 0.72 0.73 25
normal 0.78 0.86 0.82 70
extrasystole 0.00 0.00 0.00 10

The model never predicts extrasystole. All 10 extrasystole-labeled test recordings were classified as normal. This did not happen in the previous release (4/10 correct there), and it is a direct consequence of selecting the unweighted-loss configuration: it produces better accuracy and macro F1 overall, but at the cost of a whole class going unrecognized. Anyone relying on this release to flag extrasystole will not get a signal for it.

By data source, test performance was uneven: recordings from the author-supplied original collection scored 86.36% accuracy (macro F1 0.648, n=22), while PASCAL Dataset B recordings scored 74.19% accuracy but macro F1 only 0.295 (n=93) β€” largely because PASCAL B contributes all 10 test extrasystole recordings and none are recognized.

On the same original-source subset of 22 test recordings with four common labels, now run at 10 seeds (previously 3):

Training data Mean accuracy Mean macro F1
Original only 65.00% 0.545
Combined (original + PASCAL B, common classes) 74.09% 0.669

Combined training data won on 10 of 10 seed pairs (mean advantage +9.1 accuracy points; paired comparison across seeds, t=4.74, p=0.0011). This reverses the previous release's finding, which reported that additional data did not help β€” that finding used only 3 seeds and appears to have been small-sample noise rather than a real effect. As before, seeds share one split rather than being independent patient cohorts, so this is a within-split repeated-optimization comparison, not a formal significance test over independent samples; the consistency across all 10 pairs is what makes it a meaningfully stronger signal than the previous 3-seed result, not the p-value in isolation.

External validation, verified patient-level separation and clinical expert error review remain unavailable. No PhysioNet 2016 challenge score is reported: validated binary references, complete clean/noisy annotations and required weights are unavailable. Five-class labels cannot substitute for those annotations.

Files and compatibility

  • heartbeat-anomaly-detector-model.pt: calibrated five-class checkpoint.
  • active_model.json, config.json: location, hash, label order and preprocessing metadata.
  • model.py, research_pipeline.py, model_inference.py, audio_processing.py: architecture and inference.
  • inference.py: local FastAPI entry point.
  • evaluation.json, dataset_summary.json, training_protocol.json, calibration.json: aggregate evidence.
  • release_validation.json: export verification results.

The checkpoint filename is retained, but its window length and hash changed from the previous release (5-second windows now vs. 10-second previously); code and weights from this same revision must be used together β€” the previous checkpoint used class-weighted loss and 10-second windows and cannot be swapped in without also swapping config.json. The prior release remains accessible through repository history. Exact retraining requires the author's local training workspace and authorized data access; see MODEL_TRAINING_GUIDE.md.

Citation and license

@misc{mitra2026heartbeat,
  author = {Mitra, Tathagata},
  title = {Heartbeat Anomaly Detector: Research Heart-Sound Classification},
  year = {2026},
  howpublished = {https://huggingface.co/ai-mitra/heartbeat-anomaly-detector},
  note = {September 2026 research release; cite the downloaded revision}
}

Dataset reference: Bentley, P., Nordehn, G., Coimbra, M. and Mannor, S. (2011), The PASCAL Classifying Heart Sounds Challenge 2011 (CHSC2011) Results. Challenge and citation.

The existing MIT license is retained. It does not grant rights to third-party datasets or resolve permissions for source recordings. Training-data redistribution is not included in this release.

Downloads last month
40
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support