balaji1312's picture
Upload curated model release
3adf1c0 verified
|
Raw
History Blame Contribute Delete
3.76 kB
metadata
library_name: transformers
pipeline_tag: automatic-speech-recognition
language:
  - en
base_model: openai/whisper-small
tags:
  - automatic-speech-recognition
  - whisper
  - child-speech
  - model-merging
  - compositional-domain-adaptation
  - 06_intersectional_dialect_small
  - supervised-fine-tuning

Whisper Small SFT: noisy LibriSpeech

This repository contains a minimal Hugging Face Transformers checkpoint for the manuscript Compositional Domain Adaptation for Automatic Speech Recognition with Headwise Selective Attention Merging.

Model Details

  • Model type: Whisper sequence-to-sequence ASR model
  • Base model: openai/whisper-small
  • Release group: Intersectional dialect generalization
  • Checkpoint kind: Single-source supervised fine-tuned checkpoint
  • Manuscript role: Noise robustness source model without AAE adaptation
  • Source artifact: 06_intersectional_dialect_small/libri_noise_small

Method Context

This is a single-source fine-tuned checkpoint. In the manuscript it serves as a source model, baseline, or task-vector endpoint for studying how distribution-shift factors can be recombined.

Training/adaptation context: Noise-robustness source adaptation based on noisy LibriSpeech.

The broader manuscript studies whether speech foundation model adaptations for different distribution shifts, such as acoustic condition, speaking style, speaker population, and dialect, can be recombined for low-resource and intersectional ASR without direct joint-supervision data.

Intended Use

Use this checkpoint to reproduce or extend the paper's ASR model-merging experiments. It is intended for research on child ASR, compositional domain adaptation, robustness, cross-corpus transfer, dialectal variation, and scaling behavior across Whisper model sizes.

How To Load

from transformers import WhisperForConditionalGeneration, WhisperProcessor

model_id = "balaji1312/whisper_small_sft_noisy_librispeech"
processor = WhisperProcessor.from_pretrained(model_id)
model = WhisperForConditionalGeneration.from_pretrained(model_id)

For local use before upload:

from pathlib import Path
from transformers import WhisperForConditionalGeneration, WhisperProcessor

model_dir = Path("final_release_models") / "06_intersectional_dialect_small" / "whisper_small_sft_noisy_librispeech"
processor = WhisperProcessor.from_pretrained(model_dir)
model = WhisperForConditionalGeneration.from_pretrained(model_dir)

Release Files

This model card was generated for the curated release tree. The model-loading payload consists of:

config.json, generation_config.json, preprocessor_config.json, tokenizer_config.json, tokenizer.json, vocab.json, merges.txt, normalizer.json, special_tokens_map.json, added_tokens.json, model.safetensors

Training state, optimizer state, decode logs, hypotheses, references, and intermediate experiment outputs were intentionally omitted.

Limitations

The checkpoint is released for research reproducibility. Results outside the paper's child ASR, robustness, cross-corpus, dialectal, and scaling-law settings are not characterized here. Reproducing WER numbers requires the manuscript evaluation pipeline and authorized access to the relevant speech corpora; no evaluation audio or transcripts are redistributed in this model folder.

Citation

If you use this checkpoint, please cite the manuscript:

@article{shankara2026compositional,
  title = {Compositional Domain Adaptation for Automatic Speech Recognition with Headwise Selective Attention Merging},
  author = {Shankara, Natarajan Balaji and Wang, Zilai and Eren, Eray and Alwan, Abeer},
  year = {2026},
  note = {Manuscript submitted to Computer Speech & Language}
}