--- base_model: Qwen/Qwen3-Coder-Next library_name: mlx pipeline_tag: text-generation license: apache-2.0 tags: - mlx - lora - qwen3-next - orchestrator - agent --- # askai-v1 `askai-v1` is an MLX LoRA adapter for [`Qwen/Qwen3-Coder-Next`](https://huggingface.co/Qwen/Qwen3-Coder-Next), trained to perform one-step agent orchestration. Given a task, it selects one specialist worker and emits a complete delegated instruction as strict JSON. The adapter was trained against [`mlx-community/Qwen3-Coder-Next-4bit`](https://huggingface.co/mlx-community/Qwen3-Coder-Next-4bit). The upstream model is an 80B MoE with approximately 3B active parameters. ## Output Contract ```json { "action": "call_agent", "agent": "software_engineer", "model": "auto", "instruction": "A complete execution-ready instruction for the worker", "context_refs": ["user_request"], "expected_output": "completed_task_with_evidence", "budget": {"max_tokens": 4096} } ``` Supported worker labels: - `software_engineer` - `technical_writer` - `marketing_strategist` - `marketing_copywriter` - `sales_specialist` - `customer_support_specialist` - `data_analyst` - `product_designer` - `financial_analyst` - `compliance_specialist` ## Usage Install MLX LM on Apple Silicon: ```bash pip install "mlx-lm>=0.31.3" ``` Generate with the adapter: ```bash mlx_lm.generate \ --model mlx-community/Qwen3-Coder-Next-4bit \ --adapter-path tgetsov/askai-v1 \ --system-prompt "You are AskAI, a concise orchestration model. Select exactly one specialist worker for the user's task and return only the call_agent JSON contract." \ --prompt "Design and test a GDPR-compliant CRM integration for a real-estate business." \ --max-tokens 1024 \ --temp 0 ``` For best consistency, use the complete system prompt from `training_config.yaml` or the source repository. ## Training - Training base: `mlx-community/Qwen3-Coder-Next-4bit` - Upstream base: `Qwen/Qwen3-Coder-Next` - Method: MLX QLoRA - Trainable parameters: 2.857M, approximately 0.004% - Adapted blocks: final 16 of 48 - Adapted modules: full-attention Q/K/V/O projections and gated-delta input/output projections - LoRA rank: 8 - LoRA scale: 16 - Optimizer: AdamW - Learning rate: `1e-5` - Updates: 612 - Gradient accumulation: 2 - Maximum sequence length: 2048 - Training tokens: 355,178 - Peak training memory: 74.4 GB on an Apple M4 Max The source consisted of 1,080 synthetic task-to-delegated-prompt examples across ten task categories and twenty industries. Grouped splitting held out complete industries and complete `(category, subtask)` groups: - Train: 612 - Validation: 224 - Test: 244 Only assistant tokens contributed to the loss. ## Evaluation - Final sampled validation loss: 1.606 - Held-out test loss over 32 batches: 1.795 - Held-out test perplexity: 6.018 - Behavioral evaluation: one held-out example from each of ten categories - Strict JSON validity: 10/10 - Runtime schema validity: 10/10 - Worker-routing accuracy: 10/10 - Complete contract: 10/10 The behavioral sample is small and measures formatting and category routing, not end-to-end quality of worker execution. ## Limitations This first version learns only single-step worker selection and prompt delegation. It does not execute workers, consume observations, choose among specific model providers, verify results, retry failures, call agents in parallel, optimize cost, or perform reinforcement-learned multi-turn orchestration. Training examples are synthetic and stylistically concentrated. Outputs may be overly verbose and domain-specific legal, financial, medical, or compliance instructions require qualified human review. ## Integrity SHA-256 of `adapters.safetensors`: ```text 66df3ea1d9a9f0b1c3e07b1620a33dfb0b3bf36b6ea4b1e52fe2fc1bd8da67d9 ```