Qwen3.5-397B-A17B-REAP-28-NVFP4 / patches /qwen3_next.py.patch.md
rene98c's picture
Upload 2 files
8d74833 verified
|
Raw
History Blame Contribute Delete
1.06 kB

vLLM Patch: qwen3_next.py — Variable expert count per layer

File: vllm/model_executor/models/qwen3_next.py Class: Qwen3NextSparseMoeBlock.__init__ Tested on: vLLM 0.16.1rc1.dev188

What it does

Reads num_experts_per_layer (a list) from the model config to support REAP-pruned models where each layer has a different number of experts. Falls back to config.num_experts for standard models.

Patch

In Qwen3NextSparseMoeBlock.__init__, find where self.n_routed_experts is set (after self.ep_size = ...), and replace the fixed assignment with:

        # Support variable expert counts per layer (REAP pruned models)
        layer_idx = extract_layer_index(prefix)
        if hasattr(config, 'num_experts_per_layer') and config.num_experts_per_layer:
            num_experts = config.num_experts_per_layer[layer_idx]
        else:
            num_experts = config.num_experts
        self.n_routed_experts = num_experts

The extract_layer_index function is already imported from .utils in the stock file.