Quant review-ish (Apex-I-Compact & Apex-I-Mini) for agentic coding (on rtx 5070+5700x3d+56gb ram)

#2
by sanjxz - opened

I tried this model with iq3_m-code quant from atomic chat and was very disappointed in how it run. it looped, didnt compact, didnt follow instructions mid turn, and so on.

then i tried apex-i-mini, which was a very noticeable step up - loops gone, but it still didnt compact (stuck at 95%) and sometimes ignored instructions. Speeds:
image - barely usable outputs/for desperate.

so i downloaded apex-i-compact and now we are talking - no issues with model at all, everything works as expected, follows steering, no loops, even at q8\q5_1 kv (apex mini was at q8\q8) - and its faster in tg by about 1-2 t/s // i also added dry sampling params with this, maybe i-mini would benefit from them too.

Speeds:

image

image

Template used: https://huggingface.co/sanjxz/Laguna-S-2.1-Agentic-Chat-Template-Jinja

Config: latest clang self compiled mainline llama.cpp >

import subprocess
import sys

# 1. Set environment variables
env = os.environ.copy()
env["GGML_CUDA_NO_PINNED"] = "1"
env["GGML_CUDA_DISABLE_GRAPHS"] = "1"

# 2. Construct command argument list
cmd = [
    r".\llama-server.exe",
    "-m", r"D:\Laguna-S-2.1-APEX-i-compact.gguf",
    "--reasoning-preserve",
    "--reasoning-budget", "-1",
    "-ngl", "999",
    "--n-cpu-moe", "47",
    "--no-mmap",
    "--mlock",
    "-c", "220000",
    "--cache-type-k", "q8_0",
    "--cache-type-v", "q5_1",
    "-np", "1",
    "-fa", "on",
    "-t", "8",
    "-tb", "8",
    "-b", "2560",
    "-ub", "2560",
    "--jinja",
    "-kvu",
    "--temp", "0.7",
    "--top-p", "0.95",
    "--top-k", "20",
    "--min-p", "0.0",
    "--dry-multiplier", "0.8",
    "--dry-base", "1.75",
    "--dry-allowed-length", "3",
    "--dry-penalty-last-n", "-1",
    "--dry-sequence-breaker", r'\n,:,\",*,;,{,}',
    "--samplers", "top_k;top_p;min_p;temperature;dry",
    "--alias", "laguna-s-2.1",
    "--cache-reuse", "256",
    "--cache-ram", "1024",
    "--ctx-checkpoints", "32",
    "--checkpoint-min-step", "2560",
    "--host", "127.0.0.1",
    "--port", "8080",
    "--verbosity", "4",
    "--chat-template-file", r"G:\xlam3\laguna\chat_template_laguna.jinja",
]

# 3. Execute process
if __name__ == "__main__":
    try:
        print("Starting llama-server...")
        subprocess.run(cmd, env=env, check=True)
    except KeyboardInterrupt:
        print("\nServer stopped by user.")
    except Exception as e:
        print(f"\nError running server: {e}")

Overall - great job! This is probably SOTA for ~64gb shmem local agentic coding as of today

Awesome! I'm using Quality on my DGX-spark and it seems to work very well. Glad you found it useful.

Sign up or log in to comment