Shhhh, don't say anything yet!
BananaMind Base Bench 1.1
Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yes
Category scores
| Category | Elo | Correct | Accuracy | Weighted |
|---|---|---|---|---|
| language_completion | 1345 | 47/50 | 94.00% | 93.95% |
| commonsense | 1084 | 35/50 | 70.00% | 67.51% |
| world_knowledge | 1151 | 38/50 | 76.00% | 75.24% |
| context_tracking | 1002 | 26/50 | 52.00% | 50.64% |
| quantitative | 923 | 17/50 | 34.00% | 33.12% |
| logical_reasoning | 1082 | 25/50 | 50.00% | 47.65% |
| code_completion | 1403 | 43/50 | 86.00% | 85.35% |
BananaMind Base Bench 1.1
Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yesCategory scores
Category Elo Correct Accuracy Weighted language_completion 1345 47/50 94.00% 93.95% commonsense 1084 35/50 70.00% 67.51% world_knowledge 1151 38/50 76.00% 75.24% context_tracking 1002 26/50 52.00% 50.64% quantitative 923 17/50 34.00% 33.12% logical_reasoning 1082 25/50 50.00% 47.65% code_completion 1403 43/50 86.00% 85.35%
Need parameter size first too
BananaMind Base Bench 1.1
Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yesCategory scores
Category Elo Correct Accuracy Weighted language_completion 1345 47/50 94.00% 93.95% commonsense 1084 35/50 70.00% 67.51% world_knowledge 1151 38/50 76.00% 75.24% context_tracking 1002 26/50 52.00% 50.64% quantitative 923 17/50 34.00% 33.12% logical_reasoning 1082 25/50 50.00% 47.65% code_completion 1403 43/50 86.00% 85.35% Need parameter size first too
My bad, it's 90m it's the upcoming palmer-006
BananaMind Base Bench 1.1
Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yesCategory scores
Category Elo Correct Accuracy Weighted language_completion 1345 47/50 94.00% 93.95% commonsense 1084 35/50 70.00% 67.51% world_knowledge 1151 38/50 76.00% 75.24% context_tracking 1002 26/50 52.00% 50.64% quantitative 923 17/50 34.00% 33.12% logical_reasoning 1082 25/50 50.00% 47.65% code_completion 1403 43/50 86.00% 85.35% Need parameter size first too
My bad, it's 90m it's the upcoming palmer-006
Adding right now!
BananaMind Base Bench 1.1
Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yesCategory scores
Category Elo Correct Accuracy Weighted language_completion 1345 47/50 94.00% 93.95% commonsense 1084 35/50 70.00% 67.51% world_knowledge 1151 38/50 76.00% 75.24% context_tracking 1002 26/50 52.00% 50.64% quantitative 923 17/50 34.00% 33.12% logical_reasoning 1082 25/50 50.00% 47.65% code_completion 1403 43/50 86.00% 85.35% Need parameter size first too
My bad, it's 90m it's the upcoming palmer-006
Adding right now!
I need to run my private set of BananaMind Base Bench 1.1 trough it but if the current scores hold up its very good! So sadly cant add right now
I can upload the chekcpoint at BananaMind Model Previewers if it's ok
I can upload the chekcpoint at BananaMind Model Previewers if it's ok
add it at your user for 5 minutes I'll download it then
you need to a bit longer sadly my network very slow
no worry, i gave you access to the gate
done. you can remove it now
Sorry for the delay, we evaluated it on our private set and noticed some accuracy drop. It will take a bit longer.
Its added, though some collumns and the chart are currently excluded as we continue to evaluate it.