File size: 16,155 Bytes
3750eee
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ef45446
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3750eee
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
 
 
3102fc2
3750eee
 
 
3102fc2
3750eee
 
f9f5a0b
3750eee
 
 
 
 
 
 
 
f9f5a0b
3102fc2
f9f5a0b
3102fc2
f9f5a0b
3102fc2
f9f5a0b
3102fc2
f9f5a0b
3102fc2
f9f5a0b
3102fc2
f9f5a0b
3750eee
 
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
 
 
 
 
3102fc2
3750eee
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
 
f9f5a0b
 
3750eee
 
 
 
 
 
 
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
 
 
 
 
 
 
3102fc2
3750eee
3102fc2
3750eee
3102fc2
3750eee
3102fc2
ef45446
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
---

base_model: openbmb/MiniCPM5-1B
library_name: peft
license: apache-2.0
pipeline_tag: text-generation
language:
- en
datasets:
- Team-ACE/ToolACE
tags:
- base_model:adapter:openbmb/MiniCPM5-1B
- lora
- transformers
- unsloth
- function-calling
- tool-use
- agentic
inference: false
model-index:
- name: MiniCPM5-1B-Agentic-Tooluse-v3
  results:
  - task:
      type: text-generation
      name: Tool calling
    dataset:
      name: External ToolACE-derived first-call evaluation (held-out 300 examples)
      type: Team-ACE/ToolACE
    metrics:
    - type: parseable_rate
      value: 1.0000
      name: Parseable tool-call rate
    - type: valid_name_rate
      value: 0.9867
      name: Valid available-tool name rate
    - type: expected_name_rate
      value: 0.9533
      name: Expected tool-name rate
    - type: args_exact_rate
      value: 0.7467
      name: Exact-arguments rate
    - type: arg_key_overlap
      value: 0.9388
      name: Argument-key overlap
    - type: no_schema_copy_rate
      value: 0.9967
      name: No-schema-copy rate
    - type: no_repetition_rate
      value: 0.3400
      name: No-repetition rate
    - type: stopped_cleanly_rate
      value: 0.0000
      name: Stopped-cleanly rate
---


# MiniCPM5-1B-Agentic-Tooluse

LoRA adapter for `openbmb/MiniCPM5-1B`, fine-tuned on [ToolACE](https://huggingface.co/datasets/Team-ACE/ToolACE)

for single-turn **function calling**: given a conversation and a set of tool schemas, emit the first tool call

with the correct name and correct argument values.

Trained on a single Kaggle T4 with Unsloth + TRL SFT.

## Results

Evaluated on a held-out 300-example test slice drawn from a **seeded shuffle** of ToolACE (see *Split integrity*).

The base-model column is the same model with the same prompt and no adapter.

The **published weights are SFT + GRPO** (see *GRPO / RLVR*). The SFT column is kept because every

negative result below is measured against it.

| metric | v2 (previous release) | SFT retrain (pre-GRPO) | **v3 = SFT + GRPO (published)** |
|---|---|---|---|
| `parseable` — output is a well-formed call | 0.9933 | 1.0000 | **1.0000** |
| `valid_name` — name exists among the offered tools | 0.9700 | 0.9867 | **0.9867** |
| `expected_name` — name matches gold | 0.9067 | 0.9567 | **0.9533** |
| `args_exact`*every* argument value matches gold | 0.6133 | 0.7367 | **0.7467** |
| `arg_key_overlap` — F1 over argument keys | 0.8757 | 0.9422 | **0.9388** |
| **mean of 5** | 0.8718 | 0.9245 | **0.9251** |

**v2 (previous release)** = the previously published SFT adapter. An earlier draft of this card

mislabeled this column "base model (untrained)" -- that was wrong; it is NOT the raw base model.

The real untrained `openbmb/MiniCPM5-1B`, measured on this same test slice, scores `parseable`

0.9333, `valid_name` 0.9133, `expected_name` 0.8867, `args_exact` 0.6300, `arg_key_overlap` 0.8920.

**SFT retrain** = a fresh SFT pass from v2, prior to GRPO. **v3** = what this repo currently serves.

Every "did it improve?" decision in this card is judged against **v2**, not against the untrained

base model — beating an untrained model is not evidence of anything.

GRPO buys +0.0100 on `args_exact`, the metric that matters here, and gives back 0.0034 (one test example

each) on `expected_name` and `arg_key_overlap`. That trade is reported rather than hidden: the mean moves

only +0.0006, so this is a targeted gain on the hardest metric, not a broad improvement.

Four of the five metrics are above 0.80. `args_exact` is not, and the next section explains how much of it is

actually reachable.

### Reproducibility

Two independent training runs were performed. They converged to **identical** `args_exact` (0.7367) despite

different data ordering, and one differing in data composition.

| metric | run 1 | run 2 (composite-oversampled) |
|---|---|---|
| `parseable` | 1.0000 | 1.0000 |
| `valid_name` | 0.9900 | 0.9867 |
| `expected_name` | 0.9567 | 0.9567 |
| `args_exact` | 0.7367 | 0.7367 |
| `arg_key_overlap` | 0.9388 | 0.9422 |

The weights published here are run 2.

## Honest limits of `args_exact`



`args_exact` is strict and all-or-nothing over every argument value. Its measured ceiling on this test slice is

**not 1.0**:

- **9.33% of test cases are unwinnable.** 28 of 300 gold calls contain a *date* that appears nowhere in the

  prompt. There is no anchor "today" to resolve them against — the gold dates span 1990–2027 across 75 distinct

  values, so no single assumed current date recovers them. This caps `args_exact` at **0.9067**.

- The dominant remaining error class is **composite JSON arguments** (~14–20% of argument values are nested

  objects or arrays), where the model must reproduce an entire nested structure exactly.

So 0.7367 sits against a practical ceiling of 0.9067, closing about 42% of the base-model-to-ceiling gap

(0.6133 → 0.7367, out of a possible 0.6133 → 0.9067).

The grader was deliberately **not** loosened. It does normalize formatting-only differences (key order,

whitespace, `70` vs `70.0`) and is guarded at runtime by assertions in both directions: 10 must-differ pairs must

be rejected and 5 formatting-only pairs must be accepted. Relaxing date comparison, or dropping the unwinnable

cases from the denominator, would have raised the headline number without improving the model.

## What did not work

Measured negative and null results, recorded so they need not be re-tried:

- **Oversampling composite-JSON examples** — no effect on `args_exact` (0.7367 in both runs).

- **Thinking mode on** — clearly harmful: validation `args_exact` fell 0.7367 → 0.5267. Consistent with

  [TAFC (arXiv:2601.18282)](https://arxiv.org/abs/2601.18282), which notes over-reasoning degrades simple

  single-parameter function calls.

- **Longer training** — validation plateaued (run 1 at step 600, run 2 at step 750).

- **Self-consistency / majority voting** — not used. [ToolPRM (arXiv:2510.14703)](https://arxiv.org/abs/2510.14703)

  measures majority voting *degrading* argument F1 on function calling (Hammer2.1-3B: 62.83 → 58.27), because

  structured output cannot recover from an early error, so non-greedy sampling ruins whole trajectories.

  Decoding here is **greedy**.

## GRPO / RLVR — partial result

Because SFT was demonstrably saturated (two runs, identical `args_exact`, val plateaued), the next lever tried was

RL with a verifiable reward: GRPO where the reward *is* this repo's own grader, so reward and reported metric

cannot drift apart. Reward design followed [ToolRL (arXiv:2504.13958)](https://arxiv.org/abs/2504.13958) —

fine-grained decomposition (parseable / tool name / argument keys / argument values) rather than all-or-nothing,

correctness weighted 0.90 against format 0.10, and no length reward.

**It works, and the published weights include it.** With a validation ratchet that only keeps a checkpoint

beating the previous best, across several short sessions:

```



[ratchet] step 150   val args_exact 0.7467  (SFT 0.7367)   <-- kept



[ratchet] step  40   val args_exact 0.7533  (best 0.7467)  <-- kept   (higher-exploration config)



[ratchet] step  80   val args_exact 0.7500  (best 0.7533)  <-- rejected



```

The step-40 checkpoint was then measured **once** on the held-out test slice: `args_exact` 0.7367 → **0.7467**.

### Why GRPO stalls here, measured

The first configuration (4 generations, temperature 1.0) gained **zero** test examples over 100 steps. The cause

was instrumented rather than guessed — the fraction of prompt groups whose samples all receive *identical*

reward:

```



[collapse] 18/25 groups (72%) had identical rewards -> zero advantage -> no gradient



[collapse] 37/50 groups (74%) ...



[collapse] 55/75 groups (73%) ...



```

GRPO's advantage is reward minus the group mean, so a group where every sample scores the same contributes

exactly nothing. **~73% of the compute produced no gradient at all.** This is the within-group reward collapse

described by [RC-GRPO (arXiv:2602.03025)](https://www.alphaxiv.org/abs/2602.03025), and it is a direct

consequence of the SFT model already being good: most prompts come back all-right (or all-wrong) across the

whole group.

The DPO run independently measured the same thing from a different angle — 289 of 400 prompts (72%) were

already correct and yielded no preference pair. Two unrelated experiments, the same ~72%.

Two consequences worth carrying forward:

1. **Raising exploration helps.** Going to 8 generations at temperature 1.15 lifted the observed rate from

   +0.0033 to roughly +0.0165 per 100 steps. Sampling is raised for *training only*; evaluation and the

   shipped model remain strictly greedy, so this changes learning rather than the reported metric.

2. **Step count is a poor proxy for progress.** Only ~27% of steps do anything, so "N steps" should be read

   as "~0.27N effective steps" when planning a budget.

## DPO — also tried, also did not beat SFT

[TinyLLM (arXiv:2511.22138)](https://arxiv.org/abs/2511.22138) benchmarks sub-3B models on exactly this task

family and recommends preference optimization over RL for compute-constrained settings ("SFT offers limited

gains"; PPO is "computationally demanding — less ideal for edge deployment"). So DPO was tried, with preference

pairs built free from the model's own errors (`chosen` = gold call, `rejected` = what the model emitted).

**Measured:** val `args_exact` 0.7367 → **0.7333**. Not an improvement, so no test pass was spent and nothing

was published.

Two findings worth recording, both of which make DPO *less* attractive here than the literature implies:

1. **Pair generation dominates the cost.** 400 prompts took **21.5 min** of generation on a T4; DPO training

   itself then took 6.9 min. DPO needs no generation *during* training, but building the dataset is expensive.

2. **Yield is low precisely because the SFT model is good.** 289 of 400 training prompts (72%) were *already



   correct* and produce no pair. Only 111 usable pairs came out — far too few to move a 1B model. Collecting

   thousands of pairs means an hour or more of generation before a single training step.

## Summary of every method tried

| method | val `args_exact` | outcome |
|---|---|---|
| v2 (previous release) | — | test 0.6133 |
| base model (no adapter, untrained) | — | test 0.6300 |
| **v2 = SFT** (2 independent runs) | 0.7367 | superseded by v3 |
| SFT + composite oversampling | 0.7367 | no change |
| thinking-on | 0.6167 / 0.5267 | clearly worse |
| GRPO, lr 2e-5 / β 0.01 | 0.6867 | worse — too aggressive |
| GRPO, lr 5e-6 / β 0.04, 4 gen @ T=1.0 | 0.7467 | better, then stalled (73% dead groups) |
| **GRPO, lr 5e-6 / β 0.04, 8 gen @ T=1.15** | **0.7533** | **published** (test 0.7467) |
| DPO, 111 pairs | 0.7333 | worse |
| constrained / grammar-guided decoding | not run | ruled out on evidence — see below |
| self-consistency / majority voting | not run | ruled out by ToolPRM |

**Constrained decoding was considered and rejected before spending GPU time on it.** Grammar-constrained

decoding improves function calling mainly by eliminating *malformed* calls, but this model is already at

`parseable` 1.0000 and `valid_name` 0.9867 — there are essentially none left to eliminate. Its failures are

semantic (wrong argument *values*), which schema constraints do not fix. Two 2026 results also warn it can

cost accuracy: [Constraint Tax (arXiv:2606.25605)](https://www.alphaxiv.org/abs/2606.25605) finds structured

output constraints suppress tool calling, and [The Format Tax (arXiv:2604.03616)](https://arxiv.org/pdf/2604.03616)

finds format constraints degrade the reasoning the model would otherwise do.

The remaining gap to 0.80 is 0.0533 (16 test examples) against a hard ceiling of 0.9067. GRPO is the only lever

that has moved it, and its effective rate depends far more on how many groups produce a non-zero advantage than

on raw step count.

### A bug worth knowing about

The first long GRPO attempt died ~2.9h in inside the reward function:

```



schema_literal_fix → if v.lower() == ev.lower()



AttributeError: 'int' object has no attribute 'lower'



```

Tool schemas may declare **non-string enums** (`"enum": [1, 2, 3]`); the value was type-guarded but the enum

member was not. The 300-row eval slices contain no integer enum, so this never surfaced in any evaluation — it

took a 3000-row training pool to hit it. Fixed by comparing against `str(ev)` (for string enums `str(ev) is ev`,

so no previously measured number changes), with a regression assert that now fails in seconds on CPU rather than

hours into a GPU run.

## Split integrity

ToolACE's `data.json` is **grouped, not shuffled** — contiguous slices land on very different distributions, so a

naive `select(range(...))` split yields train/test sets that are not comparable. The dataset is therefore shuffled

with a fixed seed (`SPLIT_SEED = 3407`) and filtered to usable rows *before* slicing into test / validation / train.

Checkpoints were selected on **validation only**; the test slice was evaluated once, at the end.

## Training configuration

| | |
|---|---|
| base | `openbmb/MiniCPM5-1B`, 4-bit |
| LoRA | r=32, alpha=64, dropout=0.05 |
| target modules | q, k, v, o, gate, up, down |
| epochs | 3 |
| learning rate | 1e-4, cosine, warmup ratio 0.1 |
| batch size | 16 |
| max seq len | 4096 (prompt cap 1536) |
| train / val / test | 9000 / 300 / 300 |
| hardware | 1× Kaggle T4 |

## Usage

```python



from peft import PeftModel



from transformers import AutoModelForCausalLM, AutoTokenizer



base = "openbmb/MiniCPM5-1B"



tok = AutoTokenizer.from_pretrained(base, trust_remote_code=True)



model = AutoModelForCausalLM.from_pretrained(base, trust_remote_code=True, device_map="auto")



model = PeftModel.from_pretrained(model, "ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse")



prompt = tok.apply_chat_template(



    messages, tools=tools, add_generation_prompt=True,



    enable_thinking=False,          # thinking OFF — see "What did not work"



    tokenize=False,



)



inputs = tok(prompt, return_tensors="pt").to(model.device)



out = model.generate(**inputs, max_new_tokens=256, do_sample=False)   # greedy



```

`enable_thinking=False` and greedy decoding are both load-bearing for the numbers above.

## Framework versions

- PEFT 0.19.1

- transformers 4.57.3

- torch 2.8.0

- TRL 0.24.x, Unsloth

## Base model architecture

MiniCPM5-1B uses a standard `LlamaForCausalLM` architecture:

| Property | Value |
|---|---|
| Parameters (total) | 1,080,632,832 |
| Parameters (non-embedding) | 679,552,512 |
| Architecture | `LlamaForCausalLM` |
| Layers | 24 |
| Attention heads (GQA) | 16 Q / 2 KV |
| Context length | 131,072 tokens |
| Training | SFT → RL (GRPO) fine-tune on [openbmb/MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B) |

## Limitations

## ModelScope

The base model is also available on ModelScope (for users in China and East Asia):

- [OpenBMB/MiniCPM5-1B on ModelScope](https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B)

*(The fine-tuned adapter/GGUF builds are currently HuggingFace-only.)*
## Citation

If you use this model, please cite the base model paper:

```bibtex

@article{minicpm4,

  title   = {MiniCPM4: Ultra-Efficient LLMs on End Devices},

  author  = {MiniCPM Team},

  journal = {arXiv preprint arXiv:2506.07900},

  year    = {2025}

}

```

And the ToolACE dataset used for fine-tuning:

```bibtex

@article{toolace,

  title   = {ToolACE: Winning the Points of LLM Function Calling},

  author  = {Liu, Ying and others},

  journal = {arXiv preprint arXiv:2409.00920},

  year    = {2024}

}

```