| configs: | |
| - config_name: default | |
| data_files: | |
| - split: train | |
| path: train/**/*.parquet | |
| - split: validation | |
| path: validation/**/*.parquet | |
| - split: test | |
| path: test/**/*.parquet | |
| # Emilia S2S Mimi Q8 Named Speakers | |
| TinyAya codec-tokenized TTS dataset built from `gs://gp-s2s-data/s2s-data/emilia-s2s/chunk_00036_200k/parquet`. | |
| Rows are Mimi codec-only examples. The training `text` field already includes the | |
| user-facing speaker prefix: | |
| ```text | |
| Ira: transcript... | |
| ``` | |
| Speaker mapping: | |
| - `SP_SP010` -> `Ira:` | |
| - `SP_SP061` -> `Aisha:` | |
| - `SP_SP073` -> `Siya:` | |
| - `SP_SP079` -> `Zoya:` | |
| Reference conditioning is disabled for this artifact. No `ref_*` fields are emitted. | |
Xet Storage Details
- Size:
- 694 Bytes
- Xet hash:
- 06c5b5d3a46826dae6104d11a31c52f6bf80e64be6a0bba21ae488d8ed03b06a
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.