We're excited to release BananaMindBench Leaderboard, our leaderboard for BananaMind Base Bench 1.1. It measures model performance on a variety of different tasks: Language Completion Common sense too World Knowledge Context Tracking Quantitative Logical Reasoning Code Completion Each has a different score and 1 overall score. Submit your own model: BananaMind/BananaMindBench-Leaderboard Check it out: BananaMind/BananaMindBench-Leaderboard
We're excited to release BananaMind Base Bench 1.1 A new benchmark for base language models with 350 text-completion examples across seven categories. Models are scored using continuation likelihood and receive an Overall Elo score. Initial results: BananaMind-2-Medium: 1034 BananaMind-2-Mini: 974 Supra-50M-Base: 973 Supra-1.5-50M-Base-exp: 948 BananaMind-2-Nano: 910 The official script downloads the gated dataset directly from Hugging Face. The dataset is for benchmarking only and may not be used for model training. BananaMind/BananaMind-Base-Bench-1.1
We're excited to announce BananaMind 2V, our small vision model series! These models are NOT released yet. We will release them in mid-august! BananaMind 2V will include: BananaMind 2V 256M, the flagship based on BananaMind 2 Pro (BananaMind 2 Pro is not released yet). BananaMind 2V 100M, our mid model, based on BananaMind 2 Medium. BananaMind 2V 50M, our smallest vision model, based on BananaMind 2 Mini. These are currently unreleased and will release in mid-august. Our training will start after BananaMind 2 Pro has finished training.
We're excited to release BananaMind 2 Medium and BananaMind 2 Medium Chat!
Theyβre both 50M parameter models trained on 50B tokens from FineWeb-Edu, DCLM, Cosmopedia v2, FineMath-4+ and NPSet-2 Python-Edu.
The base model reached 61.86% on PIQA, 43.81% on ARC Easy and 32.43% on HellaSwag. The Chat version was fine-tuned on Smol-SmolTalk and scored 38% overall on our internal instruction benchmark, with 56% on multi-turn, 60% on context recall and 80% on code.