nielsr HF Staff commited on
Commit
34ef093
·
verified ·
1 Parent(s): c4c3a8f

Improve model card: Add metadata, paper details, and links

Browse files

This PR significantly enhances the model card by adding essential metadata and detailed information.

Key improvements include:
- Setting `library_name: transformers` to ensure proper integration and enable the "how to use" widget.
- Adding `pipeline_tag: text-generation` for better discoverability on the Hugging Face Hub.
- Specifying `license: apache-2.0` for clarity regarding usage rights.
- Including relevant `tags` such as `mixtral`, `moe`, `reasoning`, and `llm` to improve searchability.
- Providing a direct link to the paper on Hugging Face Papers: [Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks](https://huggingface.co/papers/2508.18672).
- Adding a link to the associated GitHub repository for the code: [https://github.com/rioyokotalab/optimal-sparsity](https://github.com/rioyokotalab/optimal-sparsity).
- Incorporating the paper's abstract to provide immediate context about the model and research findings.

Please review and merge if these improvements align with the repository's goals.

Files changed (1) hide show
  1. README.md +21 -1
README.md CHANGED
@@ -1,3 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ## How to cite
2
 
3
  If you find our work helpful, please feel free to cite the paper.
@@ -12,4 +32,4 @@ If you find our work helpful, please feel free to cite the paper.
12
  primaryClass={cs.LG},
13
  url={https://arxiv.org/abs/2508.18672},
14
  }
15
- ```
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: transformers
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - mixtral
7
+ - moe
8
+ - reasoning
9
+ - llm
10
+ ---
11
+
12
+ # Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
13
+
14
+ This repository contains model checkpoints and resources for the paper [Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks](https://huggingface.co/papers/2508.18672).
15
+
16
+ The associated code and logs are open-source at the [GitHub repository](https://github.com/rioyokotalab/optimal-sparsity).
17
+
18
+ ## Abstract
19
+ Empirical scaling laws have driven the evolution of large language models (LLMs), yet their coefficients shift whenever the model architecture or data pipeline changes. Mixture-of-Experts (MoE) models, now standard in state-of-the-art systems, introduce a new sparsity dimension that current dense-model frontiers overlook. We investigate how MoE sparsity influences two distinct capability regimes: memorization and reasoning. We train families of MoE Transformers that systematically vary total parameters, active parameters, and top-$k$ routing while holding the compute budget fixed. For every model we record pre-training loss, downstream task loss, and task accuracy, allowing us to separate the train-test generalization gap from the loss-accuracy gap. Memorization benchmarks improve monotonically with total parameters, mirroring training loss. By contrast, reasoning performance saturates and can even regress despite continued gains in both total parameters and training loss. Altering top-$k$ alone has little effect when active parameters are constant, and classic hyperparameters such as learning rate and initialization modulate the generalization gap in the same direction as sparsity. Neither post-training reinforcement learning (GRPO) nor extra test-time compute rescues the reasoning deficit of overly sparse models.
20
+
21
  ## How to cite
22
 
23
  If you find our work helpful, please feel free to cite the paper.
 
32
  primaryClass={cs.LG},
33
  url={https://arxiv.org/abs/2508.18672},
34
  }
35
+ ```