bankai-v1 / README.md
tgetsov's picture
Add askai-v1 model card
edc6db4 verified
|
Raw
History Blame
3.81 kB
---
base_model: Qwen/Qwen3-Coder-Next
library_name: mlx
pipeline_tag: text-generation
license: apache-2.0
tags:
- mlx
- lora
- qwen3-next
- orchestrator
- agent
---
# askai-v1
`askai-v1` is an MLX LoRA adapter for
[`Qwen/Qwen3-Coder-Next`](https://huggingface.co/Qwen/Qwen3-Coder-Next), trained
to perform one-step agent orchestration. Given a task, it selects one specialist
worker and emits a complete delegated instruction as strict JSON.
The adapter was trained against
[`mlx-community/Qwen3-Coder-Next-4bit`](https://huggingface.co/mlx-community/Qwen3-Coder-Next-4bit).
The upstream model is an 80B MoE with approximately 3B active parameters.
## Output Contract
```json
{
"action": "call_agent",
"agent": "software_engineer",
"model": "auto",
"instruction": "A complete execution-ready instruction for the worker",
"context_refs": ["user_request"],
"expected_output": "completed_task_with_evidence",
"budget": {"max_tokens": 4096}
}
```
Supported worker labels:
- `software_engineer`
- `technical_writer`
- `marketing_strategist`
- `marketing_copywriter`
- `sales_specialist`
- `customer_support_specialist`
- `data_analyst`
- `product_designer`
- `financial_analyst`
- `compliance_specialist`
## Usage
Install MLX LM on Apple Silicon:
```bash
pip install "mlx-lm>=0.31.3"
```
Generate with the adapter:
```bash
mlx_lm.generate \
--model mlx-community/Qwen3-Coder-Next-4bit \
--adapter-path tgetsov/askai-v1 \
--system-prompt "You are AskAI, a concise orchestration model. Select exactly one specialist worker for the user's task and return only the call_agent JSON contract." \
--prompt "Design and test a GDPR-compliant CRM integration for a real-estate business." \
--max-tokens 1024 \
--temp 0
```
For best consistency, use the complete system prompt from `training_config.yaml`
or the source repository.
## Training
- Training base: `mlx-community/Qwen3-Coder-Next-4bit`
- Upstream base: `Qwen/Qwen3-Coder-Next`
- Method: MLX QLoRA
- Trainable parameters: 2.857M, approximately 0.004%
- Adapted blocks: final 16 of 48
- Adapted modules: full-attention Q/K/V/O projections and gated-delta input/output projections
- LoRA rank: 8
- LoRA scale: 16
- Optimizer: AdamW
- Learning rate: `1e-5`
- Updates: 612
- Gradient accumulation: 2
- Maximum sequence length: 2048
- Training tokens: 355,178
- Peak training memory: 74.4 GB on an Apple M4 Max
The source consisted of 1,080 synthetic task-to-delegated-prompt examples
across ten task categories and twenty industries. Grouped splitting held out
complete industries and complete `(category, subtask)` groups:
- Train: 612
- Validation: 224
- Test: 244
Only assistant tokens contributed to the loss.
## Evaluation
- Final sampled validation loss: 1.606
- Held-out test loss over 32 batches: 1.795
- Held-out test perplexity: 6.018
- Behavioral evaluation: one held-out example from each of ten categories
- Strict JSON validity: 10/10
- Runtime schema validity: 10/10
- Worker-routing accuracy: 10/10
- Complete contract: 10/10
The behavioral sample is small and measures formatting and category routing,
not end-to-end quality of worker execution.
## Limitations
This first version learns only single-step worker selection and prompt
delegation. It does not execute workers, consume observations, choose among
specific model providers, verify results, retry failures, call agents in
parallel, optimize cost, or perform reinforcement-learned multi-turn
orchestration.
Training examples are synthetic and stylistically concentrated. Outputs may be
overly verbose and domain-specific legal, financial, medical, or compliance
instructions require qualified human review.
## Integrity
SHA-256 of `adapters.safetensors`:
```text
66df3ea1d9a9f0b1c3e07b1620a33dfb0b3bf36b6ea4b1e52fe2fc1bd8da67d9
```