--- license: cc-by-4.0 base_model: nvidia/stt_ru_fastconformer_hybrid_large_pc language: - ru tags: - automatic-speech-recognition - forced-alignment - gguf - crispasr - fastconformer - ctc - nemo pipeline_tag: automatic-speech-recognition --- # stt-ru-fastconformer-hybrid-ctc-large-GGUF GGUF conversions of the **CTC branch** of [nvidia/stt_ru_fastconformer_hybrid_large_pc](https://huggingface.co/nvidia/stt_ru_fastconformer_hybrid_large_pc) for [CrispASR](https://github.com/CrispStrobe/CrispASR). The upstream model is a hybrid transducer+CTC Russian ASR release; the shared FastConformer encoder plus the auxiliary CTC head are extracted here as a standalone CTC model (the RNNT prediction network and joint are dropped), giving a compact Russian ASR **and forced-alignment** model with punctuation + capitalisation. | Quant | Size | Description | |---|---|---| | F16 | 219 MB | Full precision | | Q8_0 | 130 MB | 8-bit | | Q4_K | 82 MB | 4-bit K-quant (recommended) | ## Architecture 17-layer NeMo FastConformer encoder + Conv1d CTC head. d_model=512, 8 heads, SentencePiece vocab, 80 log-mel features, ~115M params. ## Usage ```bash # Russian ASR: crispasr --backend fastconformer-ctc -m stt-ru-fastconformer-hybrid-ctc-large-q4_k.gguf -f audio.wav # Forced alignment (word timestamps for known text, or re-timing an .srt): crispasr --align-only -am stt-ru-fastconformer-hybrid-ctc-large-q4_k.gguf \ -f audio.wav --text-file subtitles.srt --align-output retimed.srt ``` ## Attribution All credit for the model goes to NVIDIA's NeMo team; this repository only repackages the CTC branch in GGUF form under the same CC-BY-4.0 license. Conversion: `models/convert-stt-fastconformer-ctc-to-gguf.py` in CrispASR. ## Provenance and EU AI Act Art. 53 note - **Upstream model:** [nvidia/stt_ru_fastconformer_hybrid_large_pc](https://huggingface.co/nvidia/stt_ru_fastconformer_hybrid_large_pc) — published by `nvidia`. - **Upstream licence:** `cc-by-4.0`. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - **What was done here:** format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs. - **Training data:** documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository. - **Provider status:** under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.