bankai-v1 / README.md
tgetsov's picture
Add askai-v1 model card
edc6db4 verified
|
Raw
History Blame
3.81 kB
metadata
base_model: Qwen/Qwen3-Coder-Next
library_name: mlx
pipeline_tag: text-generation
license: apache-2.0
tags:
  - mlx
  - lora
  - qwen3-next
  - orchestrator
  - agent

askai-v1

askai-v1 is an MLX LoRA adapter for Qwen/Qwen3-Coder-Next, trained to perform one-step agent orchestration. Given a task, it selects one specialist worker and emits a complete delegated instruction as strict JSON.

The adapter was trained against mlx-community/Qwen3-Coder-Next-4bit. The upstream model is an 80B MoE with approximately 3B active parameters.

Output Contract

{
  "action": "call_agent",
  "agent": "software_engineer",
  "model": "auto",
  "instruction": "A complete execution-ready instruction for the worker",
  "context_refs": ["user_request"],
  "expected_output": "completed_task_with_evidence",
  "budget": {"max_tokens": 4096}
}

Supported worker labels:

  • software_engineer
  • technical_writer
  • marketing_strategist
  • marketing_copywriter
  • sales_specialist
  • customer_support_specialist
  • data_analyst
  • product_designer
  • financial_analyst
  • compliance_specialist

Usage

Install MLX LM on Apple Silicon:

pip install "mlx-lm>=0.31.3"

Generate with the adapter:

mlx_lm.generate \
  --model mlx-community/Qwen3-Coder-Next-4bit \
  --adapter-path tgetsov/askai-v1 \
  --system-prompt "You are AskAI, a concise orchestration model. Select exactly one specialist worker for the user's task and return only the call_agent JSON contract." \
  --prompt "Design and test a GDPR-compliant CRM integration for a real-estate business." \
  --max-tokens 1024 \
  --temp 0

For best consistency, use the complete system prompt from training_config.yaml or the source repository.

Training

  • Training base: mlx-community/Qwen3-Coder-Next-4bit
  • Upstream base: Qwen/Qwen3-Coder-Next
  • Method: MLX QLoRA
  • Trainable parameters: 2.857M, approximately 0.004%
  • Adapted blocks: final 16 of 48
  • Adapted modules: full-attention Q/K/V/O projections and gated-delta input/output projections
  • LoRA rank: 8
  • LoRA scale: 16
  • Optimizer: AdamW
  • Learning rate: 1e-5
  • Updates: 612
  • Gradient accumulation: 2
  • Maximum sequence length: 2048
  • Training tokens: 355,178
  • Peak training memory: 74.4 GB on an Apple M4 Max

The source consisted of 1,080 synthetic task-to-delegated-prompt examples across ten task categories and twenty industries. Grouped splitting held out complete industries and complete (category, subtask) groups:

  • Train: 612
  • Validation: 224
  • Test: 244

Only assistant tokens contributed to the loss.

Evaluation

  • Final sampled validation loss: 1.606
  • Held-out test loss over 32 batches: 1.795
  • Held-out test perplexity: 6.018
  • Behavioral evaluation: one held-out example from each of ten categories
  • Strict JSON validity: 10/10
  • Runtime schema validity: 10/10
  • Worker-routing accuracy: 10/10
  • Complete contract: 10/10

The behavioral sample is small and measures formatting and category routing, not end-to-end quality of worker execution.

Limitations

This first version learns only single-step worker selection and prompt delegation. It does not execute workers, consume observations, choose among specific model providers, verify results, retry failures, call agents in parallel, optimize cost, or perform reinforcement-learned multi-turn orchestration.

Training examples are synthetic and stylistically concentrated. Outputs may be overly verbose and domain-specific legal, financial, medical, or compliance instructions require qualified human review.

Integrity

SHA-256 of adapters.safetensors:

66df3ea1d9a9f0b1c3e07b1620a33dfb0b3bf36b6ea4b1e52fe2fc1bd8da67d9