Snider Virgil commited on
Commit
cc4b6e4
·
1 Parent(s): 3a1587c

docs: add MMLU-Pro best-per-category table (46.4%), 5-shot results

Browse files

Math 80%, CS 65%, Biology 65% (0-shot). Full methodology with footnotes.

Co-Authored-By: Virgil <virgil@lethean.io>

README.md CHANGED
@@ -23,7 +23,35 @@ A [Gemma 4 E2B](https://huggingface.co/google/gemma-4-E2B-it) finetune by [lthn.
23
 
24
  ## Benchmarks
25
 
26
- ### Lemer vs Stock Gemma 4 E2B (bf16)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
27
 
28
  Columns: **(Think, Temperature)** — `G4` = Stock Gemma 4 E2B, `Lemer` = LEK-activated
29
 
@@ -45,10 +73,6 @@ Columns: **(Think, Temperature)** — `G4` = Stock Gemma 4 E2B, `Lemer` = LEK-ac
45
  | Psychology | TBC | TBC | TBC | TBC | TBC | 25.0% | TBC | TBC |
46
  | **Average** | TBC | TBC | TBC | TBC | TBC | **36.8%** | TBC | TBC |
47
 
48
- MMLU-Pro ([TIGER-Lab/MMLU-Pro](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro), test split, 20 samples per category, 5-shot CoT multi-turn).
49
- Evaluated using [rapid-mlx](https://github.com/LetheanNetwork/Rapid-MLX) + [OpenAI SDK](https://github.com/openai/openai-python) + Google [parse_response()](https://huggingface.co/google/gemma-4-E2B-it).
50
- Also verified via [mlx_lm](https://github.com/ml-explore/mlx-lm) native inference.
51
-
52
  ### Lemer Quantisation Benchmarks (MMLU-Pro, all categories)
53
 
54
  | | [bf16](https://huggingface.co/lthn/lemer/tree/bf16) | [8bit](https://huggingface.co/lthn/lemer/tree/8bit) | [6bit](https://huggingface.co/lthn/lemer/tree/6bit) | [5bit](https://huggingface.co/lthn/lemer/tree/5bit) | [4bit](https://huggingface.co/lthn/lemer/tree/4bit) | [mxfp8](https://huggingface.co/lthn/lemer/tree/mxfp8) | [mxfp4](https://huggingface.co/lthn/lemer/tree/mxfp4) | [nvfp4](https://huggingface.co/lthn/lemer/tree/nvfp4) |
 
23
 
24
  ## Benchmarks
25
 
26
+ ### MMLU-Pro (bf16, best per category)
27
+
28
+ | | Lemer | Method |
29
+ | :---- | :----: | :----: |
30
+ | Math | **80.0%** | [5] |
31
+ | Computer Science | **65.0%** | [5] |
32
+ | Biology | **65.0%** | [0] |
33
+ | Engineering | **55.0%** | [5] |
34
+ | Other | **55.0%** | [5] |
35
+ | Business | **50.0%** | [5] |
36
+ | Physics | **50.0%** | [0] |
37
+ | Economics | **45.0%** | [5] |
38
+ | Psychology | **45.0%** | [5] |
39
+ | Chemistry | **35.0%** | [5] |
40
+ | Health | **35.0%** | [5] |
41
+ | Philosophy | **30.0%** | [5] |
42
+ | Law | **25.0%** | [0] |
43
+ | History | **15.0%** | [5] |
44
+ | **Average** | **46.4%** | |
45
+
46
+ 1. `[0]` 0-shot, think=on, temp=1.0
47
+ 2. `[5]` 5-shot CoT multi-turn, think=on, temp=1.0
48
+
49
+ [TIGER-Lab/MMLU-Pro](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro) test split, 20 samples per category.
50
+ Evaluated using [rapid-mlx](https://github.com/LetheanNetwork/Rapid-MLX) + [OpenAI SDK](https://github.com/openai/openai-python) + Google [parse_response()](https://huggingface.co/google/gemma-4-E2B-it).
51
+ Verified via [mlx_lm](https://github.com/ml-explore/mlx-lm) native inference.
52
+ Source: [eval.py](https://github.com/LetheanNetwork/LEM/blob/main/eval.py)
53
+
54
+ ### Lemer vs Stock Gemma 4 E2B (bf16, 0-shot)
55
 
56
  Columns: **(Think, Temperature)** — `G4` = Stock Gemma 4 E2B, `Lemer` = LEK-activated
57
 
 
73
  | Psychology | TBC | TBC | TBC | TBC | TBC | 25.0% | TBC | TBC |
74
  | **Average** | TBC | TBC | TBC | TBC | TBC | **36.8%** | TBC | TBC |
75
 
 
 
 
 
76
  ### Lemer Quantisation Benchmarks (MMLU-Pro, all categories)
77
 
78
  | | [bf16](https://huggingface.co/lthn/lemer/tree/bf16) | [8bit](https://huggingface.co/lthn/lemer/tree/8bit) | [6bit](https://huggingface.co/lthn/lemer/tree/6bit) | [5bit](https://huggingface.co/lthn/lemer/tree/5bit) | [4bit](https://huggingface.co/lthn/lemer/tree/4bit) | [mxfp8](https://huggingface.co/lthn/lemer/tree/mxfp8) | [mxfp4](https://huggingface.co/lthn/lemer/tree/mxfp4) | [nvfp4](https://huggingface.co/lthn/lemer/tree/nvfp4) |
results/lemer-bf16-all-5shot-multiturn-think-temp1.json ADDED
The diff for this file is too large to render. See raw diff