ChristophSchuhmann commited on
Commit
4e0170b
·
verified ·
1 Parent(s): c291e8b

Remove the blanket 'unusable for commercial purposes' claim; state the NC-corpus question neutrally

Browse files
Files changed (1) hide show
  1. README.md +10 -7
README.md CHANGED
@@ -30,11 +30,14 @@ benchmarks.
30
 
31
  ## Why this model exists
32
 
33
- The standard VoiceCLAP training mix contains **CC BY-NC** (non-commercial)
34
- corpora: Expresso and EARS (both Meta, CC BY-NC 4.0), and the bulk of
35
- [Emilia](https://huggingface.co/datasets/amphion/Emilia-Dataset) (the original
36
- 101k-hour split is CC BY-NC 4.0). That makes models trained on the full mix
37
- unusable for commercial purposes.
 
 
 
38
 
39
  This model removes every non-commercial source and keeps only commercially
40
  usable data. A controlled ablation (below) shows the small model loses
@@ -124,5 +127,5 @@ original open_clip implementation (cosine ≥ 0.99999 on both towers).
124
  **CC-BY-4.0** — all training data is commercially usable (Emilia-YODAS CC BY 4.0,
125
  LAION's Got Talent, in-house Majestrino), and the architecture/weights carry no
126
  non-commercial restriction. This is the distinguishing feature of this model
127
- versus the standard VoiceCLAP releases, which inherit a non-commercial
128
- restriction from their CC BY-NC training data.
 
30
 
31
  ## Why this model exists
32
 
33
+ The standard VoiceCLAP training mix contains corpora released under **CC BY-NC**
34
+ (non-commercial) terms: Expresso and EARS (both Meta, CC BY-NC 4.0), and the
35
+ bulk of [Emilia](https://huggingface.co/datasets/amphion/Emilia-Dataset) (the
36
+ original 101k-hour split is CC BY-NC 4.0). Whether such terms bind the weights
37
+ of a model trained on that data is a question each user must answer for their
38
+ own jurisdiction and use; LAION licenses its releases on its own assessment.
39
+ This model exists for users who prefer not to have to answer it: every corpus
40
+ in its training lineage permits commercial use.
41
 
42
  This model removes every non-commercial source and keeps only commercially
43
  usable data. A controlled ablation (below) shows the small model loses
 
127
  **CC-BY-4.0** — all training data is commercially usable (Emilia-YODAS CC BY 4.0,
128
  LAION's Got Talent, in-house Majestrino), and the architecture/weights carry no
129
  non-commercial restriction. This is the distinguishing feature of this model
130
+ versus the standard VoiceCLAP releases, whose training mix includes CC BY-NC
131
+ corpora.