You can now filter models by architecture type. Standard models work out of the box with GGUF, vLLM, and standard runtimes with zero remote code execution.
| Cmp | # โผ | Model | Params | Int Index | HellaSwag | ARC-E | ARC-C | PIQA | ArithMark-3 | Avg | Fit Std Dev | ArithMark-2 | Released |
|---|
Parameters vs. Accuracy
Dashed line: Linear regression fit. Shaded zone: Outperforming parameter expectation.
Top 12 Performers
Top models for the selected benchmark.
| # | Organization | Models Tracked | Mean Fit Std Dev | Mean Score | Top Model |
|---|
ArithMark-3 is now the primary math benchmark for ranking, weighting subject to change. ArithMark-2 remains available in the expanded columns.
UCR's 2.74M SLM uses custom tokenization and digit features to reach 69.4% ArithMark accuracy, setting the leaderboard's #1 score
Built on Axiomic Labs' T-X4 refresh-gated XSA stack, GPT-S2-5M pushes into the sub-10M bracket and lands at the top, dethroning SLM-10M.
Chance-Normalized Intelligence Index Calculation
Standard raw accuracy averages reward models for trivial chance performance on multiple-choice benchmarks. The Intelligence Index normalizes each benchmark against its theoretical random baseline, ensuring that chance performance equals 0 and perfect performance equals 100.
Int Index = (HellaSwag + Combined_ARC + PIQA + 0.65 × ArithMark-3) ÷ 3.65
* Combined ARC is defined as (ARC-Easy + ARC-Challenge) / 2 prior to normalization. If a benchmark component is missing from a model checkpoint, its corresponding weight is deducted from both numerator and denominator. Minimum 2 evaluated tasks required for inclusion.
Add your model
Open a Discussion or a PR on this Space (whichever you're more comfortable with) with your model's benchmark results.
Always report normalized accuracy (acc_norm) when available across evaluation frameworks.
Submissions are independently verified by our team before merging. To qualify, your model must be pretrained by you and have open weights.
Open a Discussion / PR →
Listing Policy & Benchmark Integrity
AtomixLabs evaluates and indexes public open-weight models independently and via community submissions. To preserve benchmark integrity and prevent selective reporting, publicly released models cannot opt out. While creators of active work-in-progress checkpoints may request an exemption, AtomixLabs retains final authority over whether any model or preview is indexed.
Integrity & Neutrality: Any model with confirmed test-set leakage, data contamination, or artificial overfitting will be flagged or removed. Rankings are strictly merit-based: all models stand solely on their verified scores, free from creator disputes or personal politics.