--- license: other --- # sklearn-pmml-model: Uncontrolled Memory Allocation via `SparseArray/@n` (DoS) **Target:** https://github.com/iamDecode/sklearn-pmml-model (PyPI: `sklearn-pmml-model`) **Platform:** huntr (Model File Vulnerabilities, PMML format) **Class:** CWE-789 (Memory Allocation with Excessive Size Value) / CWE-400 (Uncontrolled Resource Consumption) **Status:** Gated PoC — access granted to protectai-bot for triage. ## Summary `sklearn_pmml_model.base.parse_sparse_array()` builds an in-memory Python list sized directly from the attacker-controlled `n` attribute of a `` element in a PMML file, with no upper bound or sanity check against the document's actual content: ```python # sklearn_pmml_model/base.py def parse_sparse_array(array): ... values = [0] * int(array.get('n')) # <-- unbounded allocation indices = [int(i) - 1 for i in array.find('Indices').text.split(' ')] ... ``` Loading any PMML model that contains a `SupportVectorMachineModel` with a `VectorInstance` whose vector is encoded as a `SparseArray` reaches this line via `sklearn_pmml_model.svm._base.get_vectors()` -> `parse_array()`. This is hit through the library's own public, documented API (`auto_detect_estimator()`, `PMMLSVC`, `PMMLSVR`, `PMMLNuSVC`, `PMMLNuSVR`), which is precisely the workflow the README recommends for loading a "trained model exported to PMML" from any source — i.e., attacker-supplied model files are the intended threat model for this library. A ~1KB PMML file is enough to make the process attempt to allocate a multi-terabyte Python list, which either: - raises `MemoryError` inside `parse_sparse_array` and aborts the caller (denial of service for that request/process), or - (with a smaller-but-still-huge `n`) causes multi-second/minute stalls while CPython allocates and fills the list, tying up the loading thread/process. There is no relationship enforced between `n` and the actual number of `Indices`/`Entries` provided — the allocation happens *before* those are even parsed, so a minimal, otherwise mostly-empty `SparseArray` triggers it. ## Attacker input -> sink chain 1. Attacker supplies a PMML file to any application that loads models via `sklearn_pmml_model.auto_detect.auto_detect_estimator(path)` (the library's documented top-level API) or directly via `sklearn_pmml_model.svm.PMMLSVR/PMMLSVC(pmml=path)`. 2. `auto_detect_estimator` sees ``. - `trigger_svm_api.py` — loads the file via `sklearn_pmml_model.svm.PMMLSVR(pmml=...)`. - `trigger_auto_detect_api.py` — loads the file via the documented top-level `sklearn_pmml_model.auto_detect.auto_detect_estimator(...)`. ### Reproduction (against an unmodified, freshly cloned/installed checkout) ```bash git clone https://github.com/iamDecode/sklearn-pmml-model.git python3 -m venv venv && source venv/bin/activate pip install -e sklearn-pmml-model # Bound virtual memory to 3GB purely so the demo fails fast/safely on the # demonstrator's own machine; the bug itself has no dependency on any limit. (ulimit -v 3000000; python3 trigger_auto_detect_api.py malicious_svr.pmml) ``` ### Observed (real run, this environment, sklearn-pmml-model @ master, April 2026 checkout) ``` Traceback (most recent call last): File ".../trigger_auto_detect_api.py", line 7, in clf = auto_detect_estimator(sys.argv[1]) File ".../sklearn_pmml_model/auto_detect/base.py", line 40, in auto_detect_estimator return auto_detect_regressor(pmml, **kwargs) File ".../sklearn_pmml_model/auto_detect/base.py", line 129, in auto_detect_regressor return parse(f) File ".../sklearn_pmml_model/auto_detect/base.py", line 123, in parse return reg(pmml, **kwargs) File ".../sklearn_pmml_model/svm/_classes.py", line 295, in __init__ PMMLBaseSVM.__init__(self) File ".../sklearn_pmml_model/svm/_base.py", line 56, in __init__ get_vectors(vector_dictionary, s) for s in self.support_ File ".../sklearn_pmml_model/svm/_base.py", line 109, in get_vectors return np.array(parse_array(array)) File ".../sklearn_pmml_model/base.py", line 330, in parse_array return parse_sparse_array(array) File ".../sklearn_pmml_model/base.py", line 369, in parse_sparse_array values = [0] * int(array.get('n')) MemoryError raised as expected via auto_detect_estimator(). Traceback: ``` Without an external memory limit, the same file will make the hosting process attempt to allocate ~8TB (`999999999999 * 8 bytes` for the pointer array alone) until the OS OOM-kills the process or the allocation fails — in either case the process handling the model-load crashes. ## Impact Any service that loads user- or third-party-supplied PMML files with `sklearn-pmml-model` (e.g. a model-serving/inference API, an AutoML pipeline accepting model uploads, a batch scoring job) can be crashed by a single ~1KB file, with no valid support-vector data required. This is a denial-of-service primitive delivered entirely through the model file — the same class of "malicious model file" bug huntr's Model File Vulnerability program targets for other formats (pickle, HDF5/Keras, PyTorch, etc.), here affecting the PMML format via `sklearn-pmml-model`. ## Fix suggestion Bound `n` to a sane maximum (or to the number of `Indices` actually present) before allocating, and/or catch `MemoryError`/`OverflowError` around PMML parsing and re-raise as a controlled `Exception`, e.g.: ```python n = int(array.get('n')) num_indices = len(array.find('Indices').text.split()) if n <= 0 or n > MAX_REASONABLE_SPARSE_ARRAY_SIZE or n < num_indices: raise Exception('Invalid or excessive SparseArray size.') values = [0] * n ``` ## Dedup check - No GitHub issues, PRs, or Security Advisories in `iamDecode/sklearn-pmml-model` reference `SparseArray`, memory exhaustion, or denial of service (checked 2026-07-06 via GitHub search API and the repo's Security Advisories page — zero results). - No CVE/GHSA exists for `sklearn-pmml-model` in the GitHub Advisory Database as of 2026-07-06. ## Reporter Enigma Partners Global — authorized security research, huntr Model File Vulnerability program (target: PMML).