|
Download README.md from HitendraKawale/whispr-replication: direct link, hf CLI and curl.
- Browser
- Download file 8.25 kB
-
https://huggingface.co/HitendraKawale/whispr-replication/resolve/main/README.md
- Command line
-
hf download hf://HitendraKawale/whispr-replication/README.md
-
curl -L -o README.md https://huggingface.co/HitendraKawale/whispr-replication/resolve/main/README.md
8.25 kB
| license: mit | |
| language: | |
| - en | |
| tags: | |
| - automatic-speech-recognition | |
| - whisper | |
| - replication | |
| - educational | |
| - research-artifact | |
| - not-for-production | |
| datasets: | |
| - librispeech_asr | |
| metrics: | |
| - wer | |
| library_name: pytorch | |
| pipeline_tag: automatic-speech-recognition | |
| model-index: | |
| - name: whispr-100h | |
| results: | |
| - task: | |
| type: automatic-speech-recognition | |
| name: Automatic Speech Recognition | |
| dataset: | |
| name: LibriSpeech dev-clean | |
| type: librispeech_asr | |
| split: dev-clean | |
| metrics: | |
| - type: wer | |
| value: 103.4 | |
| name: WER | |
| verified: false | |
| # whispr — a from-scratch Whisper replication (these models do not work) | |
| > ### ⚠️ Read this before downloading | |
| > | |
| > **These models do not transcribe speech.** Word error rate is **103%** and | |
| > **150%** — worse than outputting nothing. They emit fluent English that loops | |
| > and largely ignores the audio. | |
| > | |
| > That is not a bug report, it is the result. They are published so the numbers | |
| > in the accompanying write-up are checkable and so nobody has to repeat five | |
| > hours of GPU time to reproduce them. | |
| > | |
| > **If you want to transcribe audio, use [`openai/whisper-tiny`](https://huggingface.co/openai/whisper-tiny)** | |
| > — same architecture, 6,800× more training data, and it actually works. | |
| > | |
| > These are also **not `transformers`-compatible.** `AutoModel.from_pretrained` | |
| > will not load them. See [Usage](#usage). | |
| ## What this is | |
| A from-scratch replication of | |
| [*Robust Speech Recognition via Large-Scale Weak Supervision*](https://arxiv.org/abs/2212.04356) | |
| (Whisper), built and trained on an M1 MacBook Air as a way of learning audio ML | |
| from the ground up. | |
| Code, a written note explaining every step, and the full method: | |
| **https://github.com/HitendraKawale/whispr_replication** | |
| The architecture is verified faithful: OpenAI's real `whisper-tiny` checkpoint | |
| loads into this implementation with `strict=True` and produces **bit-identical** | |
| outputs (max difference `0.0`). The log-mel frontend matches | |
| `whisper.log_mel_spectrogram` to `1.2e-7`. What differs is the training data — | |
| 100 hours instead of 680,000. | |
| ## The checkpoints | |
| | File | Train audio | Steps | Best val loss | WER (unseen speakers) | | |
| |---|---|---|---|---| | |
| | `whispr-100h-step20000.pt` | 100.3 h `train-clean-100` | 20,000 (~6 epochs) | 4.028 | **103.4%** | | |
| | `whispr-3.7h-step1250.pt` | 3.7 h `dev-clean` | 1,250 (~4.7 epochs) | 5.574 | **149.9%** | | |
| Each `.pt` requires its matching `tokenizer-*.json`. The vocabulary is fitted | |
| per-run on that run's training transcripts, so the token ids differ between | |
| them — **pairing a checkpoint with the wrong tokenizer decodes to nonsense | |
| silently**, without raising. | |
| ~18M parameters each (Whisper Tiny's shape: 4 layers, width 384, 6 heads, with a | |
| 2,048-token vocabulary instead of GPT-2's 50,257). | |
| ## What they actually output | |
| On speakers they have never heard: | |
| ``` | |
| REF SHE GOT UP ON HER KNEES AND WRUNG HER HANDS | |
| HYP I'LL NOT THINK IT'T YOU'T DO YOU'T (3.7 h) | |
| REF I SAW THAT INTERVIEW IN THE PAPER YESTERDAY TELLING WHERE YOU WERE | |
| HYP AND THE OTHER TWO OF THE MEN WHO WERE IN THE SAME WAY THE OLD MAN (100 h) | |
| ``` | |
| WER above 100% is possible because insertions are unbounded — the model emits | |
| more words than the reference. | |
| ## Why they fail, which is the interesting part | |
| The decoder learns two things at once: the **distribution of English text**, and | |
| the **mapping from audio to text**. The first is cheap — a few thousand | |
| transcripts teach you that "THE" is common and follows almost anything, and it | |
| requires no encoder at all. The second is expensive. | |
| So gradient descent does the cheap thing first. Cross-entropy falls from 7.6 to | |
| ~5.5 almost entirely by learning unigrams and bigrams, then plateaus, because | |
| further progress needs the encoder to become useful. | |
| At 3.7 hours the model never gets there and overfits — validation traced a | |
| textbook U, bottoming at step 1,250 then climbing while training loss kept | |
| falling. At 100 hours the U disappears: validation was still falling when the | |
| schedule ended, with training loss tracking it closely. The problem changed from | |
| **overfitting to underfitting**, which is a better problem — it says the next win | |
| is more epochs and more data, not more regularisation. | |
| Whisper used 680,000 hours. That gap is the paper's central claim, measured from | |
| the wrong end. | |
| ## Usage | |
| These use a custom model class, not `transformers`. | |
| ```bash | |
| git clone https://github.com/HitendraKawale/whispr_replication | |
| cd whispr_replication && uv sync | |
| ``` | |
| ```python | |
| import torch | |
| from huggingface_hub import hf_hub_download | |
| from whispr.audio import load_audio | |
| from whispr.config import AudioConfig, Config, ModelConfig, TrainConfig | |
| from whispr.decode import Decoder | |
| from whispr.mel import log_mel_spectrogram | |
| from whispr.model import build_model | |
| from whispr.tokenizer import WhisprTokenizer | |
| REPO = "HitendraKawale/whispr-replication" | |
| ckpt_path = hf_hub_download(REPO, "whispr-100h-step20000.pt") | |
| tok_path = hf_hub_download(REPO, "tokenizer-100h.json") # must match! | |
| ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=False) | |
| saved = ckpt["config"] | |
| cfg = Config( | |
| audio=AudioConfig(**saved["audio"]), | |
| model=ModelConfig(**saved["model"]), | |
| train=TrainConfig(**saved["train"]), | |
| ) | |
| model = build_model(cfg.model) | |
| model.load_state_dict(ckpt["model"]) | |
| model.eval() | |
| tokenizer = WhisprTokenizer.load(tok_path) | |
| mel = log_mel_spectrogram(load_audio("clip.flac"), pad_to=cfg.audio.n_samples) | |
| print(Decoder(model, tokenizer).greedy(mel)[0].text) | |
| ``` | |
| Each checkpoint carries its own config, so you don't have to remember that the | |
| 100 h model uses a 17 s window (850 encoder positions) and the 3.7 h model uses | |
| 15 s (750). Loading one with the other's config fails loudly. | |
| ### Resuming training | |
| Both include optimizer state, so training continues with the LR schedule intact: | |
| ```bash | |
| cp whispr-100h-step20000.pt checkpoints/run_100h/best.pt | |
| uv run python scripts/07_train.py --corpus train-clean-100 --window 17 \ | |
| --steps 20000 --mel-cache --resume best.pt | |
| ``` | |
| ## Training details | |
| Optimiser settings follow the paper's Table 17 exactly: AdamW, β=(0.9, 0.98), | |
| ε=1e-6, weight decay 0.1, gradient-norm clip 1.0, linear warmup then linear decay | |
| to zero, Gaussian fan-in initialisation. | |
| Deviations, all forced by a single laptop and all documented with reasoning in | |
| `whispr/config.py`: | |
| | | Paper | Here | | |
| |---|---|---| | |
| | Training audio | 680,000 h, multilingual | 100.3 h English read speech | | |
| | Updates | 1,048,576 @ batch 256 | 20,000 @ batch 8 | | |
| | Vocabulary | 50,257 (GPT-2 BPE) | 2,048, fitted on our transcripts | | |
| | Window | 30 s (1500 enc positions) | 17 s (850) | | |
| | Augmentation | none | ±6 dB gain jitter | | |
| | Tasks | transcribe + translate + langID + VAD | English transcription only | | |
| Trained on an Apple M1 (16 GB) via MPS. Roughly 6 hours for the 100 h run. | |
| ## Evaluation | |
| Speaker-disjoint throughout. The 100 h model trains on `train-clean-100` and | |
| validates on `dev-clean` — LibriSpeech's own partition, so the number is | |
| comparable to published results rather than only to itself. Greedy decoding, 400 | |
| held-out utterances, WER computed corpus-level (not averaged per utterance) after | |
| Whisper-style text normalisation. | |
| ## Limitations | |
| Everything. Specifically: English only, read audiobook speech only, no | |
| punctuation or casing (LibriSpeech labels have none), no timestamps, no | |
| translation, no robustness to noise or accents or spontaneous speech, and it does | |
| not produce correct transcripts of anything. | |
| ## Reproducibility | |
| MPS kernels are nondeterministic, so identical seeds do not give identical runs. | |
| An earlier 100 h run of this same code reached val 3.876 / 110.2% WER using a | |
| 25,000-step schedule. Expect the qualitative findings to reproduce and the third | |
| decimal not to. | |
| ## Citation | |
| The paper being replicated: | |
| ```bibtex | |
| @article{radford2022whisper, | |
| title = {Robust Speech Recognition via Large-Scale Weak Supervision}, | |
| author = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg | |
| and McLeavey, Christine and Sutskever, Ilya}, | |
| journal = {arXiv preprint arXiv:2212.04356}, | |
| year = {2022} | |
| } | |
| ``` | |
| ## Licence | |
| MIT, matching the repository. LibriSpeech is CC BY 4.0. | |