lostargon commited on
Commit
606fe02
·
verified ·
1 Parent(s): 2f17b92

model card: charts

Browse files
.gitattributes CHANGED
@@ -34,3 +34,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ assets/calibration.png filter=lfs diff=lfs merge=lfs -text
38
+ assets/zero_shot.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -88,6 +88,8 @@ Every domain includes a large share of deliberately hard cases (negations, near-
88
 
89
  All numbers are accuracy on held-out items; **ECE** is expected calibration error (lower is better; 0.01 means the stated confidence is off by about one point on average). *acc@0.9* is accuracy on the subset of answers where the model's confidence is at least 0.9, with *cov@0.9* the share of items that subset covers — the pair that matters when you gate actions on confidence.
90
 
 
 
91
  ### Held-out splits of the training distribution
92
 
93
  | split | items | accuracy | ECE | acc@0.9 | cov@0.9 |
@@ -101,6 +103,8 @@ All numbers are accuracy on held-out items; **ECE** is expected calibration erro
101
 
102
  Item lists follow the open **open-system-one** protocol (2 500 items per dataset, fixed seed), so the reference row is on exactly the same items. Tiny-Jev was additionally fitted on 2 000 items per dataset that are disjoint from the evaluation items — the same allowance that protocol gives its fitted open baselines; the reference model is zero-shot, so read the comparison as *"a 0.6B model plus a small fit set"* versus *"a frontier decision API with no fit set"*.
103
 
 
 
104
  | dataset | options | Tiny-Jev acc | ECE | acc@0.9 | cov@0.9 | reference (zero-shot) |
105
  |---|---|---|---|---|---|---|
106
  | SST-2 | 2 | **90.4** | 0.023 | 96.4 | 77 % | 91.6 |
@@ -111,6 +115,12 @@ Item lists follow the open **open-system-one** protocol (2 500 items per dataset
111
 
112
  † not in the fit set: transfer from the other intent tasks only.
113
 
 
 
 
 
 
 
114
  ### Reading the numbers
115
 
116
  - On its own distribution the model is both accurate and honest: at ≥0.9 confidence it is right 99.8 % of the time and reaches that confidence on 90 % of items.
 
88
 
89
  All numbers are accuracy on held-out items; **ECE** is expected calibration error (lower is better; 0.01 means the stated confidence is off by about one point on average). *acc@0.9* is accuracy on the subset of answers where the model's confidence is at least 0.9, with *cov@0.9* the share of items that subset covers — the pair that matters when you gate actions on confidence.
90
 
91
+ ![Calibration](assets/calibration.png)
92
+
93
  ### Held-out splits of the training distribution
94
 
95
  | split | items | accuracy | ECE | acc@0.9 | cov@0.9 |
 
103
 
104
  Item lists follow the open **open-system-one** protocol (2 500 items per dataset, fixed seed), so the reference row is on exactly the same items. Tiny-Jev was additionally fitted on 2 000 items per dataset that are disjoint from the evaluation items — the same allowance that protocol gives its fitted open baselines; the reference model is zero-shot, so read the comparison as *"a 0.6B model plus a small fit set"* versus *"a frontier decision API with no fit set"*.
105
 
106
+ ![Benchmarks](assets/benchmarks.png)
107
+
108
  | dataset | options | Tiny-Jev acc | ECE | acc@0.9 | cov@0.9 | reference (zero-shot) |
109
  |---|---|---|---|---|---|---|
110
  | SST-2 | 2 | **90.4** | 0.023 | 96.4 | 77 % | 91.6 |
 
115
 
116
  † not in the fit set: transfer from the other intent tasks only.
117
 
118
+ ### Zero-shot multiple choice
119
+
120
+ No training on any of these tasks; 500 items each, options passed as a Choice question. They measure general knowledge and arithmetic, which a model this size only partly has.
121
+
122
+ ![Zero-shot](assets/zero_shot.png)
123
+
124
  ### Reading the numbers
125
 
126
  - On its own distribution the model is both accurate and honest: at ≥0.9 confidence it is right 99.8 % of the time and reaches that confidence on 90 % of items.
assets/benchmarks.png ADDED
assets/calibration.png ADDED

Git LFS Details

  • SHA256: 24fdc62f10f963352701ba9cc28128d25b3002fcff0de58823cc66a34ffd8d14
  • Pointer size: 131 Bytes
  • Size of remote file: 170 kB
assets/zero_shot.png ADDED

Git LFS Details

  • SHA256: edaf711fb1389c26cc4e64178420d951eb51727bfbacd802516a263d31706e61
  • Pointer size: 131 Bytes
  • Size of remote file: 177 kB