Shhhh, don't say anything yet!

#1
by appvoid - opened

BananaMind Base Bench 1.1

Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yes

Category scores

Category Elo Correct Accuracy Weighted
language_completion 1345 47/50 94.00% 93.95%
commonsense 1084 35/50 70.00% 67.51%
world_knowledge 1151 38/50 76.00% 75.24%
context_tracking 1002 26/50 52.00% 50.64%
quantitative 923 17/50 34.00% 33.12%
logical_reasoning 1082 25/50 50.00% 47.65%
code_completion 1403 43/50 86.00% 85.35%
BananaMind AI org

BananaMind Base Bench 1.1

Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yes

Category scores

Category Elo Correct Accuracy Weighted
language_completion 1345 47/50 94.00% 93.95%
commonsense 1084 35/50 70.00% 67.51%
world_knowledge 1151 38/50 76.00% 75.24%
context_tracking 1002 26/50 52.00% 50.64%
quantitative 923 17/50 34.00% 33.12%
logical_reasoning 1082 25/50 50.00% 47.65%
code_completion 1403 43/50 86.00% 85.35%

Need parameter size first too

BananaMind Base Bench 1.1

Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yes

Category scores

Category Elo Correct Accuracy Weighted
language_completion 1345 47/50 94.00% 93.95%
commonsense 1084 35/50 70.00% 67.51%
world_knowledge 1151 38/50 76.00% 75.24%
context_tracking 1002 26/50 52.00% 50.64%
quantitative 923 17/50 34.00% 33.12%
logical_reasoning 1082 25/50 50.00% 47.65%
code_completion 1403 43/50 86.00% 85.35%

Need parameter size first too

My bad, it's 90m it's the upcoming palmer-006

BananaMind AI org

BananaMind Base Bench 1.1

Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yes

Category scores

Category Elo Correct Accuracy Weighted
language_completion 1345 47/50 94.00% 93.95%
commonsense 1084 35/50 70.00% 67.51%
world_knowledge 1151 38/50 76.00% 75.24%
context_tracking 1002 26/50 52.00% 50.64%
quantitative 923 17/50 34.00% 33.12%
logical_reasoning 1082 25/50 50.00% 47.65%
code_completion 1403 43/50 86.00% 85.35%

Need parameter size first too

My bad, it's 90m it's the upcoming palmer-006

Adding right now!

BananaMind AI org

BananaMind Base Bench 1.1

Model: a-cool-mini-language-model
Overall Elo: 1124
Accuracy: 231/350 (66.00%)
Weighted accuracy: 63.68%
Official complete run: yes

Category scores

Category Elo Correct Accuracy Weighted
language_completion 1345 47/50 94.00% 93.95%
commonsense 1084 35/50 70.00% 67.51%
world_knowledge 1151 38/50 76.00% 75.24%
context_tracking 1002 26/50 52.00% 50.64%
quantitative 923 17/50 34.00% 33.12%
logical_reasoning 1082 25/50 50.00% 47.65%
code_completion 1403 43/50 86.00% 85.35%

Need parameter size first too

My bad, it's 90m it's the upcoming palmer-006

Adding right now!

I need to run my private set of BananaMind Base Bench 1.1 trough it but if the current scores hold up its very good! So sadly cant add right now

I can upload the chekcpoint at BananaMind Model Previewers if it's ok

BananaMind AI org

I can upload the chekcpoint at BananaMind Model Previewers if it's ok

add it at your user for 5 minutes I'll download it then

BananaMind AI org

you need to a bit longer sadly my network very slow

no worry, i gave you access to the gate

BananaMind AI org

done. you can remove it now

appvoid changed discussion status to closed
Banaxi-Tech changed discussion status to open
BananaMind AI org

Sorry for the delay, we evaluated it on our private set and noticed some accuracy drop. It will take a bit longer.

BananaMind AI org

Its added, though some collumns and the chart are currently excluded as we continue to evaluate it.

Banaxi-Tech changed discussion status to closed

Sign up or log in to comment