Commercial use of preset voices and training-data provenance

#3
by banbencochuyen - opened

Hello VieNeu-TTS author,

We are evaluating VieNeu-TTS v3 Turbo for a Vietnamese content project that may publish monetized videos.

Target versions:

The current model card states that the model package is distributed under Apache License 2.0. Before using generated audio in publicly distributed or monetized content, could you please clarify the following points?

  1. Does Apache-2.0 cover all artifacts in this revision, including model.safetensors, ONNX exports, configs/tokenizers, and the bundled preset-voice speaker embeddings and codes?

  2. May the bundled preset voices be used to generate synthetic voice-over audio for commercial and monetized videos? Were the preset-voice assets derived from speakers who granted appropriate rights or consent for commercial AI training and synthetic voice generation?

  3. The model is described as being trained on approximately 10,000 hours of English–Vietnamese speech, and the referenced VieNeu-TTS-10k-ENVI dataset is currently gated. Could you clarify:

    • the dataset license;
    • the sources and provenance of the audio;
    • whether the relevant speakers granted appropriate rights or consent for AI training and synthetic speech generation;
    • whether models trained on this dataset may be used commercially?
  4. We found a preset-voice mismatch for the pinned versions:

    • the source registry and GitHub README list 14 preset voices;
    • list_preset_voices() returns 14 voices;
    • the Hugging Face model card currently lists 10 default voices, with several different names.
      Could you confirm the official preset list for this revision and whether the model card should be updated?
  5. Apart from the attribution already stated in the model card, are there any additional license, attribution, or usage obligations relating to the preset voices, MOSS-Audio-Tokenizer-Nano, sea-g2p, or other bundled components?

If possible, we would appreciate a public clarification in the model card or repository so that the answer can be associated with the relevant revision.

Thank you for your work and for helping us comply properly.

Hi, just a gentle follow-up on this thread (opened about 9 days ago) — no rush, and thank you again for your work.

For our internal compliance review, the two points we most need confirmed for this revision (model 75ff82a72f54d55ed389e1eeb12041d3c4bac7d4, source f56ce97ffb3731aeafed623391587a1589ecb501, vieneu 3.2.3) are:

Whether the bundled preset voices and the generated audio outputs may be used in commercial / monetized content; and
The training-data provenance and speaker rights/consent for the VieNeu-TTS-10k-ENVI dataset.

A brief public note — here or in the model card, tied to this revision — would be enough for us. If it helps, we're happy to narrow the request down to just these two points. Thank you!

Hi, thanks for your patience and for the detailed follow-up.

To address your two main points for this revision (model 75ff82a72f54d55ed389e1eeb12041d3c4bac7d4, source f56ce97ffb3731aeafed623391587a1589ecb501, vieneu 3.2.3):

  1. Preset voices & commercial use
    The bundled preset voices are distributed under the same Apache-2.0 license as the rest of the repository, and may be used to generate synthetic voice-over audio for commercial and monetized content.

  2. Training data provenance
    We don't publicly disclose the detailed internal data collection and processing pipeline for the training dataset. What we can confirm is that the preset voices currently shipped in this revision do not, to our knowledge, infringe on third-party rights as of the release date.

We understand this may not cover every detail needed for your internal compliance review, and we appreciate you raising these questions — they're helpful for us to think through as we continue to document the project. If you need anything further clarified for your specific use case, feel free to let us know.

Thank you — that's really helpful, and the commercial-use confirmation for the preset voices and outputs is exactly what we needed.

On the training data, we completely understand not disclosing the internal pipeline. Just one narrower point for our records: can you confirm that the speakers whose voices are represented in the shipped preset voices provided consent for their voice to be used in AI training and synthetic speech generation? A simple confirmation is enough — we don't need any dataset details.

Also, if you're open to it, adding this commercial-use confirmation to the model card (tied to this revision) would help other users as well. Thanks again for your time.

Sign up or log in to comment