--- base_model: meta-llama/Llama-3.2-1B-Instruct base_model_relation: adapter library_name: peft pipeline_tag: text-generation license: llama3.2 language: - en tags: - llama - llama-3.2 - lora - peft - tool-calling - function-calling - agent - code - python - bash - gguf ---
Llama Mini Agent # Llama Mini Agent **A tiny, tool-calling agent fine-tune of Llama-3.2-1B-Instruct.** *(checkpoint: `py-newds` — one revision in a 13-variant training sweep, not the final one)*
--- ## Overview **Llama Mini Agent** is a **LoRA fine-tune of Llama-3.2-1B-Instruct** for lightweight **agentic tool use** — emitting clean function calls and terse **Python** / **Bash** snippets — while staying small enough to run on a laptop or CPU. The adapter is also shipped merged into the base and quantized as ready-to-run **GGUF** files. | | | |---|---| | **Base model** | `unsloth/Llama-3.2-1B-Instruct` (≈1.24B params) | | **Method** | LoRA (rank 32, α 32, dropout 0) | | **Trainable params** | **22.5M** across 224 tensors (≈1.8% of base) | | **Target modules** | `q/k/v/o_proj`, `gate/up/down_proj` (all 32 layers) | | **Task** | Causal LM — tool calling, Python & Bash generation | | **Context** | 128K (inherited from base) | | **Chat / tool format** | Llama 3.2 `ipython` tool-calling template | | **License** | Llama 3.2 Community License | --- ## Why this checkpoint exists This adapter is one point in a **13-repo iteration sweep** (`lora`, `45k`, `filtered`, `nopack`, `2epoc`, `pyclean`, `multiupdate`, `rs-py`, `skew-py`, `anchor`, `nopy`, **`py-newds`**, and an isolated `py-lora` follow-up), run while chasing a specific, recurring problem: joint tool + Python + Bash training reliably *tanked Python* relative to the base model, no matter how the Python slice of the data was cleaned or curated. **The investigation, briefly:** 1. Five prior variants — `pyclean` (prose-stripped Python), `rs-py` (teacher-verified Python), `skew-py` (re-weighted quality mix), `anchor` (verified-only anchor set), and `nopy` (**zero** Python training data) — all landed in the **same 18–26% HumanEval cluster**. 2. `nopy` scoring in that same band with *no Python data at all* was the clincher: the regression wasn't bad Python data, it was **collateral forgetting from the joint tool/bash training pressure itself**. 3. `py-newds` tested the one remaining angle: give Python the same first-class treatment tool and bash already had — a single clean, execution-verified source (`bigcode/self-oss-instruct-sc2-exec-filter-50k`), kept in full (only filter: `ast.parse` validity — no length cap, no class exclusion), rather than the Magicoder-padded or heavily-filtered Python slices used earlier. 4. A later, separate experiment (`llama-3.2-1b-py-lora`) isolated Python entirely —training it *alone*, with no tool/bash in the mix at all — to see whether removing joint-training pressure altogether would beat `py-newds`. Per this project's own eval notebook (`test.ipynb`), that isolated run scored **40/164** on HumanEval — offered here as related context, not as this adapter's own score (see *Performance* below for `py-newds`'s own measured numbers). `py-newds` is therefore best understood as **the "give Python first-class data inside the joint model" answer** in an otherwise well-documented, methodical sweep — not a one-off training run. --- ## Training recipe The joint-training recipe used across this sweep (LoRA hyperparameters below are ground truth from this adapter's own `adapter_config.json`; the SFT settings are the project's standard recipe for joint tool/bash/Python runs): | | | |---|---| | **LoRA** | rank 32, α 32, dropout 0, bias none, all 7 attention+MLP projections, all 32 layers | | **Epochs** | 2 | | **LR / schedule** | 2e-4, cosine, warmup ratio 0.03 | | **Batch** | 2 × grad-accum 8 (effective 16), NEFTune noise α 5 | | **Precision / optim** | bf16, `paged_adamw_8bit` | | **Sequence** | max 1024 tokens, no packing, grouped by length | | **Loss masking** | `train_on_responses_only` — loss computed only on assistant turns, not prompts/tool schemas | | **Data — tool** | Salesforce xLAM function-calling (60k), 15k sampled | | **Data — Python** | `bigcode/self-oss-instruct-sc2-exec-filter-50k`, full/unfiltered except `ast.parse` validity | | **Data — bash** | `westenfelder/NL2SH-ALFA`, deduplicated + per-utility frequency-capped (≤250/utility) to prevent `find`/`grep`/`git`-style over-representation | | **Data — hand-written** | 60 hand-written bash + tool examples, 5× upsampled, covering common sysadmin one-liners and multi-tool-choice scenarios | --- ## Performance Measured by loading the **base model once** and toggling this adapter on/off via `PeftModel.disable_adapter()` — identical tokenizer, chat template, and **greedy** decoding for both, so the LoRA weights are the only variable. Harness ported directly from this project's own `test.ipynb`. | Benchmark | Base `Llama-3.2-1B-Instruct` | **Llama Mini Agent (py-newds)** | |---|---:|---:| | **HumanEval pass@1** (164 problems) | **64/164 (39.0%)** | **55/164 (33.5%)** | | **Tool-call F1** (50 xLAM held-out) | **0.0%** (0/50 valid+correct, 7/50 even valid JSON) | **88.0%** (42/50 exact call match, 50/50 valid JSON) | | **Bash exact match** (50 NL2SH-ALFA held-out) | **0/50 (0%)** | **34/50 (68%)** | | **Bash exact + same-utility**¹ | 0/50 (0%) | **45/50 (90%)** | ¹ *"same-utility"* = the model picked the correct command but different flags/args (e.g. `ls` vs `ls -a`) — counted separately from exact match, not blended into it. 5/50 adapter bash outputs were outright wrong. **Reading this table honestly:** the adapter trades **5.5 points of HumanEval** (64→55, a ~14% relative drop) for **+88 points of tool-call F1** and **+68 points of bash exact-match** — the base model essentially cannot format a tool call or solve a bash task at all in this harness (0% on both), while the adapter is strong on both. The HumanEval drop is the same joint-training tax documented across the whole 13-variant sweep (see *Why this checkpoint exists*), not new information — but it's worth seeing the base model's own Python ceiling (39%) alongside it, since that's the number `py-newds` is giving up part of in exchange for tool/bash competence. For reference only — **not** this adapter's own score — the later Python-isolated `llama-3.2-1b-py-lora` follow-up (Python trained alone, no tool/bash in the mix at all) scored **40/164** on the same HumanEval set per this project's `test.ipynb`. That's *lower* than `py-newds`'s own 55/164 measured here, which is a mildly interesting data point: removing joint-training pressure entirely didn't recover more Python performance than `py-newds` already has. The two runs weren't evaluated back-to-back in one harness, so treat this as a suggestive comparison, not a controlled one. --- ## Qualitative comparison Beyond the aggregate scores, the clearest difference between base and adapter is **tool-calling format and termination behavior**, tested through 🤗 Transformers with context capped at **4096 tokens**; the shipped **Q4_K_M GGUF** was separately smoke-tested through `llama.cpp`. | Behavior | Base `Llama-3.2-1B-Instruct` | **Llama Mini Agent** | |---|---|---| | **Tool-call format** | ❌ Malformed — emits `{"type":"function","function":"…","parameters":…}` (invalid schema) | ✅ Clean `[{"name":…, "arguments":{…}}]` | | **Stops after the call** | ❌ Runs on, re-emitting `assistant` turns and dozens of duplicate calls until the token budget is exhausted | ✅ Emits one call and stops at `<\|eot_id\|>` | | **Verbosity** | Long prose explanations wrapped around every answer | Terse — code / one-liner, little or no filler | | **Bash task correctness** | ❌ `find . -name "*.py" \| wc -l` (counts *files*, not lines) | ✅ `wc -l $(find . -name '*.py' -print)` (counts lines) | | **Latency (greedy, GPU)** | 2.6 – 6.8 s / prompt (keeps generating) | 0.6 – 4.4 s / prompt (stops early) | ### Verbatim test outputs **1. Tool call** — prompt: *"What's the weather like in Paris right now? Use celsius."* with a `get_current_weather(city, unit)` tool. ```text BASE {"type": "function", "function": "get_current_weather", "parameters": {"city": "Paris", "unit": "celsius"}}assistant … (same malformed call repeated ~8× with stray "assistant" headers, never stops) … ADAPTER [{"name": "get_current_weather", "arguments": {"city": "Paris", "unit": "celsius"}}] ← single call, stops ``` **2. Bash** — prompt: *"…count the total number of lines across every .py file in the current directory tree."* ```text BASE find . -type f -name "*.py" | wc -l ← counts files, not lines; +12 lines of (incorrect) explanation ADAPTER wc -l $(find . -name '*.py' -print) ← correct, no filler GGUF wc -l $(find . -name "*.py" -print) ← Q4_K_M, via llama.cpp @ ~208 tok/s ``` **3. Python** — prompt: *"Write a Python function `nth_fib(n)` … iteratively."* — both models produce a plausible iterative function with a docstring; the adapter's is more compact (no surrounding walkthrough). As with any 1B model, verify edge cases (both variants have off-by-one behavior at `n=0/1`). **4. "No tool needed" turn** — prompt: *"Thanks, that's all I needed. Have a good day!"* (tools still in context). The base model hallucinated a long stream of `get_current_weather` calls for random cities; the adapter also emitted a spurious call here rather than a plain sign-off — see *Known quirks*. --- ## Known quirks & caveats - **1B Python ceiling.** As documented above, joint tool/bash/Python training at this scale caps HumanEval well below what an isolated Python fine-tune or the base model's own coding ability might suggest. This is a **known, investigated limitation**, not an unexplained regression — treat Python output as a rough draft, always verify before running. - **`arguments` vs `parameters`.** The model emits calls as `[{"name": …, "arguments": {…}}]`, but the bundled `chat_template.jinja` renders a *prior* assistant tool call as `{"name": …, "parameters": …}` (unwrapped). If you feed multi-turn tool-call history back through the template, normalize the key/shape so training-time and template-time formats match. - **Over-eager tool calls.** With tools in context, a closing pleasantry can still trigger a function call instead of a plain reply. Guard the tool loop on the application side (validate that a call is actually warranted). --- ## Roadmap (not yet implemented) Documented but unexecuted directions from this project's own planning: - **Multi-turn tool data** — current training is single-turn tool calls; agentic loops (call → observe → call again) aren't represented in the data. - **GRPO on Python only, continuing from `py-newds`** — Python has a clean, execution-based reward (the `sc2` test suite); tool-calling doesn't (the "tools" are fictional API schemas, so its only available signal is noisy reference-match against xLAM). Plan: SFT stays as-is for tool/bash, RL nudges Python further, with a KL penalty (and possibly tool/bash examples mixed into RL batches as an anchor) to guard against re-triggering the same collateral-forgetting problem RL would inherit if left unconstrained. - **MTP** (multi-token prediction) as part of a separate, larger 3B project. --- ## Files ``` adapter_model.safetensors # the raw LoRA adapter (22.5M params) adapter_config.json chat_template.jinja tokenizer.json / tokenizer_config.json ...Q3_K_M.gguf ~691 MB # smallest ...Q4_K_M.gguf ~808 MB # recommended balance ...Q5_K_M.gguf ~912 MB ...Q6_K.gguf ~1.02 GB ...Q8_0.gguf ~1.32 GB # near-lossless ...fp16.gguf ~2.48 GB # full precision ``` **Pick a quant:** start with **Q4_K_M** for the best size/quality trade-off. Drop to **Q3_K_M** if memory is tight, or go **Q6_K / Q8_0** for maximum fidelity. --- ## Usage ### llama.cpp ```bash llama-cli -m llama-3.2-1b-tool-py-bash-py-newds.Q4_K_M.gguf --jinja -c 4096 \ -p "Write a bash one-liner to count lines in every .py file." -st ``` ### Ollama ```bash ollama create llama-mini-agent -f Modelfile # FROM ./gguf/...Q4_K_M.gguf ollama run llama-mini-agent ``` ### Transformers (LoRA adapter) ```python from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel base = AutoModelForCausalLM.from_pretrained("unsloth/Llama-3.2-1B-Instruct") model = PeftModel.from_pretrained(base, "adapter") tok = AutoTokenizer.from_pretrained("adapter") msgs = [{"role": "user", "content": "What's the weather in Paris?"}] tools = [{"type": "function", "function": {"name": "get_current_weather", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}] enc = tok.apply_chat_template(msgs, tools=tools, add_generation_prompt=True, return_tensors="pt") print(tok.decode(model.generate(enc, max_new_tokens=128)[0][enc.shape[1]:])) # -> [{"name": "get_current_weather", "arguments": {"city": "Paris"}}] ``` --- ## Intended use & limits - **Good for:** small local agents, function-calling demos, quick Python/Bash helpers, edge / offline deployment. - **Not for:** complex multi-step reasoning, production automation without a human/validation layer, or anything safety-critical. --- ## Reproducing the evaluation Both the aggregate benchmark and the qualitative comparison above load the base model once, wrap it with this adapter, and toggle `model.disable_adapter()` so base and fine-tune share every other variable (tokenizer, template, greedy decoding). HumanEval-164, a 50-example xLAM tool-call F1, and a 50-example NL2SH-ALFA bash suite are ported directly from this project's `test.ipynb`. ---
Fine-tuned with LoRA on Llama-3.2-1B-Instruct · GGUF via llama.cpp · evaluated 26 Jul 2026