{ "schema_version": 1, "target": "MLX-6bit-Group64", "source_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", "validated_at": "2026-08-16T01:26:34.604373+00:00", "structural": { "passed": true, "passes": [ "artifact directory exists", "atomic build completion record", "local SHA-256 manifest", "all build-hashed files present (24)", "all build payload SHA-256 hashes match", "per-artifact quantization manifest", "artifact manifest base model", "artifact manifest source revision", "artifact manifest declares vanilla quantization", "plausible size 21.99 GB in [20, 26]", "required sidecar config.json", "required sidecar tokenizer_config.json", "required sidecar generation_config.json", "required sidecar preprocessor_config.json", "required sidecar video_preprocessor_config.json", "required sidecar chat_template.jinja", "required sidecar tokenizer.json", "required sidecar vocab.json", "required sidecar merges.txt", "required sidecar LICENSE", "chat template byte-identical to source", "generation_config.json semantically intact", "preprocessor_config.json semantically intact", "video_preprocessor_config.json semantically intact", "tokenizer.json byte-identical to source", "vocab.json byte-identical to source", "merges.txt byte-identical to source", "LICENSE byte-identical to source", "tokenizer config preserves chat_template", "tokenizer config preserves eos_token", "tokenizer config preserves pad_token", "tokenizer config preserves additional_special_tokens", "official internal architecture id retained", "MTP layer declaration retained", "vision configuration retained", "image special token id retained", "video special token id retained", "vision-start token id retained", "vision-end token id retained", "text architecture preserves num_hidden_layers", "text architecture preserves hidden_size", "text architecture preserves vocab_size", "text architecture preserves max_position_embeddings", "text architecture preserves full_attention_interval", "text architecture preserves layer_types", "text architecture preserves linear_conv_kernel_dim", "text architecture preserves linear_key_head_dim", "text architecture preserves linear_num_key_heads", "text architecture preserves linear_num_value_heads", "text architecture preserves linear_value_head_dim", "full vision configuration retained", "main safetensors readable", "vision tensors retained (333)", "main MLX bundle does not misplace MTP tensors", "MTP drafter tensors retained (15)", "MTP drafter model type", "all floating tensors finite" ], "failures": [], "warnings": [], "disk_size_bytes": 21989149993 }, "runtime": { "runtime": "mlx-vlm", "model": "/Users/oz/Documents/ChatGPT/qwen3.8-27b/qwen38-27b-quants/quantization-suite/artifacts/MLX-5bit-Group32", "artifact_role": "quantized_candidate", "source_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", "text_tests": [ { "id": "fact", "category": "factual", "passed": true, "failures": [], "text": "The capital of France is Paris.", "finish_reason": "stop", "prompt_tokens": 24, "generation_tokens": 8, "generation_tps": 13.489448408921401, "peak_memory_gb": 23.443856782, "wall_seconds": 1.6853342500980943, "loop_score": 0.0 }, { "id": "math", "category": "mathematical_reasoning", "passed": true, "failures": [], "text": "The user is asking a simple multiplication problem. A box has 6 rows of 7 blue marbles. I need to multiply 6 × 7 to find the total number of blue marbles.\n\n6 × 7 = 42\n\nI'll show a short calculation as requested.\n\n\n**6 × 7 = 42**\n\nThere are **42** blue marbles.", "finish_reason": "stop", "prompt_tokens": 36, "generation_tokens": 85, "generation_tps": 11.825482041537015, "peak_memory_gb": 23.523972168, "wall_seconds": 7.477397124981508, "loop_score": 0.0 }, { "id": "code", "category": "coding", "passed": true, "failures": [], "text": "The user wants a Python function called `clamp(x, low, high)` that returns `x` limited to the inclusive range `[low, high]`. This is a straightforward function:\n- If `x < low`, return `low`\n- If `x > high`, return `high`\n- Otherwise, return `x`\n\nThis can be done with `max(low, min(x, high))` or with if/elif/else. I'll go with the clean one-liner approach using `max` and `min`.\n\n\n```python\ndef clamp(x, low, high):\n return max(low, min(x, high))\n```", "finish_reason": "stop", "prompt_tokens": 67, "generation_tokens": 139, "generation_tps": 11.715345260641968, "peak_memory_gb": 23.620598965, "wall_seconds": 12.2929570840206, "loop_score": 0.0 }, { "id": "multiturn", "category": "multi_turn", "passed": true, "failures": [], "text": "cedar", "finish_reason": "stop", "prompt_tokens": 50, "generation_tokens": 3, "generation_tps": 17.584285524527225, "peak_memory_gb": 23.620598965, "wall_seconds": 0.47414225013926625, "loop_score": 0.0 }, { "id": "thinking_off", "category": "thinking_disabled", "passed": true, "failures": [], "text": "Yes, 17 is prime.\n\n**Reason:** It has no divisors other than 1 and itself (it is not divisible by 2, 3, or any other integer less than its square root, $\\sqrt{17} \\approx 4.12$).", "finish_reason": "stop", "prompt_tokens": 25, "generation_tokens": 61, "generation_tps": 11.757861867844758, "peak_memory_gb": 23.620598965, "wall_seconds": 5.509693667059764, "loop_score": 0.0 }, { "id": "reasoning_low", "category": "reasoning_effort_low", "passed": true, "failures": [], "text": "Compare 3/5 and 5/9.\n\n3/5 = 27/45\n5/9 = 25/45\n\n27/45 > 25/45, so 3/5 is larger.\n\n\n3/5", "finish_reason": "stop", "prompt_tokens": 60, "generation_tokens": 62, "generation_tps": 11.896109264545515, "peak_memory_gb": 23.620598965, "wall_seconds": 5.518751166993752, "loop_score": 0.0 }, { "id": "reasoning_high", "category": "reasoning_effort_high", "passed": true, "failures": [], "text": "We need answer user: Find smallest positive integer divisible by 6,8,15. Explain briefly. Need compute LCM. 6=2*3, 8=2^3, 15=3*5. LCM max powers: 2^3*3*5=120. Need final concise.\n\n\nThe smallest positive integer divisible by 6, 8, and 15 is their least common multiple (LCM).\n\nPrime factorizations:\n- 6 = 2 × 3\n- 8 = 2³\n- 15 = 3 × 5\n\nTake the highest power of each prime:\n- 2³, 3, 5\n\nLCM = 2³ × 3 × 5 = 8 × 3 × 5 = **120**.", "finish_reason": "stop", "prompt_tokens": 73, "generation_tokens": 182, "generation_tps": 11.709988212332219, "peak_memory_gb": 23.639191525, "wall_seconds": 15.973943375051022, "loop_score": 0.0 }, { "id": "long_form", "category": "long_form", "passed": true, "failures": [], "text": "Checksums serve as the foundational integrity mechanism for reproducible machine learning by providing a unique digital fingerprint for every artifact in the release pipeline, including raw datasets, preprocessed features, and model weights. By calculating cryptographic hashes such as SHA-256 for these files, teams can verify that the exact data and code used to train a model remain unchanged during storage, transfer, or deployment. This ensures that if a checksum mismatch is detected at any stage, the system can immediately flag the corruption or tampering, preventing the execution of a model that was trained on different inputs than those documented in the release notes.\n\nImmutable revisions extend this integrity by locking the entire environment and configuration into a specific, unchangeable state, often managed through version control systems or container images. In machine learning, where subtle differences in library versions, random seeds, or hyperparameters can drastically alter outcomes, immutable revisions guarantee that the exact software stack and training parameters are preserved. This eliminates the \"it works on my machine\" problem by ensuring that any engineer or automated system can pull the same revision and reproduce the identical training process, thereby isolating the model’s behavior from external environmental drift.\n\nTogether, these mechanisms create a verifiable audit trail that links the final deployed model back to its specific source data and code version. When a checksum validates the data and an immutable revision secures the code and dependencies, the release becomes a deterministic unit rather than a probabilistic one. This transparency allows data scientists to debug performance regressions by comparing specific revisions, and it enables compliance teams to demonstrate exactly which data and algorithms were used to generate a prediction, which is critical for regulated industries like finance and healthcare.\n\nUltimately, the combination of checksums and immutable revisions transforms machine learning releases from opaque black boxes into reliable, software-engineering-grade products. By enforcing strict consistency across the data, code, and environment, organizations can scale their ML operations with confidence, knowing that a model deployed in production is mathematically identical to the one tested in development. This rigor reduces deployment risks, accelerates the feedback loop for model improvements, and builds trust among stakeholders by ensuring that every prediction is backed by a reproducible and verifiable process.", "finish_reason": "stop", "prompt_tokens": 31, "generation_tokens": 441, "generation_tps": 11.679253224808445, "peak_memory_gb": 23.639191525, "wall_seconds": 38.083526875125244, "loop_score": 0.0 }, { "id": "non_english", "category": "non_english", "passed": true, "failures": [], "text": "太陽光が大気中の分子にぶつかると、波長の短い青い光が他の色よりも強く散乱されます。そのため、私たちの目には空全体が青く見えます。", "finish_reason": "stop", "prompt_tokens": 28, "generation_tokens": 41, "generation_tps": 12.003249511296172, "peak_memory_gb": 23.639191525, "wall_seconds": 3.7392137499991804, "loop_score": 0.0 }, { "id": "json", "category": "structured_json", "passed": true, "failures": [], "text": "```json\n{\n \"name\": \"sample\",\n \"primes\": [2, 3, 5],\n \"active\": true\n}\n```", "finish_reason": "stop", "prompt_tokens": 41, "generation_tokens": 38, "generation_tps": 12.026567006267767, "peak_memory_gb": 23.639191525, "wall_seconds": 3.456523875007406, "loop_score": 0.0 } ], "tool_tests": [ { "id": "one_tool", "passed": true, "failures": [], "calls": [ { "name": "weather_at", "arguments": { "location": "Paris", "units": "celsius" } } ], "raw": "The user is asking to call weather_at for Paris in Celsius units. This is a simple, direct request.\n\n\n\n\n\nParis\n\n\ncelsius\n\n\n", "finish_reason": "stop", "dependent_raw": null }, { "id": "parallel_tools", "passed": true, "failures": [], "calls": [ { "name": "weather_at", "arguments": { "location": "Paris", "units": "celsius" } }, { "name": "weather_at", "arguments": { "location": "Austin", "units": "fahrenheit" } } ], "raw": "The user is asking to call weather_at in parallel for Paris (Celsius) and Austin (Fahrenheit). Since these are independent calls, they can be executed in the same block.\n\n\n\n\n\nParis\n\n\ncelsius\n\n\n\n\n\n\nAustin\n\n\nfahrenheit\n\n\n", "finish_reason": "stop", "dependent_raw": null }, { "id": "sequential_dependency", "passed": true, "failures": [], "calls": [ { "name": "weather_at", "arguments": { "location": "Paris", "units": "celsius" } } ], "raw": "The user is asking to first call weather_at for Paris in Celsius, and then create an event with the title being exactly the returned condition string.\n\nSince the second call depends on the result of the first call, I need to wait for the first result before making the second call.\n\n\n\n\n\nParis\n\n\ncelsius\n\n\n", "finish_reason": "stop", "dependent_raw": "The weather condition is \"sunny\". Next, I need to create an event with the title exactly \"sunny\".\n\n\n\n\n\nsunny\n\n\n" }, { "id": "nested_json", "passed": true, "failures": [], "calls": [ { "name": "add_event", "arguments": { "title": "Release review", "attendees": [ "ana@example.test", "li@example.test" ], "metadata": { "release": { "model": "Qwen3.8-27B", "slot": 21 } } } } ], "raw": "The user is asking to call add_event with specific parameters:\n- title: \"Release review\"\n- attendees: [\"ana@example.test\", \"li@example.test\"]\n- metadata: {\"release\":{\"model\":\"Qwen3.8-27B\",\"slot\":21}}\n\nThis is a straightforward function call request. Let's execute it.\n\n\n\n\n\nRelease review\n\n\n[\"ana@example.test\", \"li@example.test\"]\n\n\n{\"release\": {\"model\": \"Qwen3.8-27B\", \"slot\": 21}}\n\n\n", "finish_reason": "stop", "dependent_raw": null }, { "id": "invalid_recovery", "passed": true, "failures": [], "calls": [ { "name": "weather_at", "arguments": { "location": "Paris", "units": "celsius" } } ], "raw": "The user is asking to call weather_at for Paris with units=kelvin. However, the schema only allows \"celsius\" or \"fahrenheit\" for units. Therefore, kelvin violates the schema. The user is saying that if it violates the schema, recover by using celsius.\n\nSo, I should call weather_at with location=Paris and units=celsius.\n\n\n\n\n\nParis\n\n\ncelsius\n\n\n", "finish_reason": "stop", "dependent_raw": null } ], "vision_tests": [ { "id": "shapes_colors", "passed": true, "missing_patterns": [], "text": "From left to right, the three large shapes are:\n\n1. **Red square** \n2. **Blue circle** \n3. **Green triangle**\n\nThese match the labels beneath each shape in the image: “RED”, “BLUE”, and “GREEN” respectively.", "finish_reason": "stop" }, { "id": "printed_text", "passed": true, "missing_patterns": [], "text": "VISION CHECK 27B", "finish_reason": "stop" }, { "id": "chart", "passed": true, "missing_patterns": [], "text": "Based on the provided image:\n\n- The bar chart has three bars labeled **A**, **B**, and **C**.\n- The numbers printed above each bar are:\n - **A**: 60\n - **B**: 105\n - **C**: 135\n\nThe tallest bar is **C**, as it has the highest value (135) and visually extends the highest in the chart.\n\n✅ **Answer: Bar C is the tallest, and the number printed above it is 135.**", "finish_reason": "stop" } ], "mtp": { "passed": true, "drafter_kind": "mtp", "output_equivalent_temperature_zero": true, "accepted_drafts": 84, "drafted_tokens": 88, "acceptance_rate": 0.9545454545454546, "baseline_tps": 11.889081281356098, "mtp_tps": 13.25660234721037, "speedup": 1.1150232750110602, "measured_improvement": true, "baseline_wall_seconds": 11.084187333006412, "mtp_wall_seconds": 9.887206791900098, "advertise_acceleration": true }, "warnings": [], "phases": [ "mtp", "text", "tools", "vision" ], "validation_inputs": { "prompts_sha256": "136a918e5fee962f2b52f8e520a0275fd5ec5569181e0f5fdf4910ee3c34d528", "tools_sha256": "86ae46ebdbb0324c9672eec87e6b7f7683b0eb9c8c7169dcabdf646cf9bab122", "image_sha256": "0b1ae6badbe19a6049305c36d165a34cdf36df00033842e265339b2b7f057295" }, "metal_memory_policy": { "device": { "device_name": "Apple M5 Pro", "max_recommended_working_set_size": 55662788608, "memory_size": 68719476736, "architecture": "applegpu_g17s", "max_buffer_length": 41747087360, "resource_limit": 499000 }, "cache_limit_bytes": 256000000, "wired_limit_bytes": 54549532835, "previous_cache_limit_bytes": 65283502899, "previous_wired_limit_bytes": 0, "warnings": [] } }, "warnings": [], "overall_passed": true, "runtime_failures": [], "quality": { "schema_version": 1, "comparison_type": "cross-runtime output agreement against pinned BF16 source", "passed": true, "source_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", "embedding_model": { "repo_id": "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2", "revision": "e8f8c211226b894fcb81acc59f3b34ba3efd5f42", "pooling": "attention-mask mean pooling followed by L2 normalization", "maximum_tokens": 256 }, "thresholds": { "mean_semantic_similarity": 0.55, "per_case_severe_regression": 0.25 }, "baseline_valid": true, "candidate_functional": true, "semantic_gate_passed": true, "validation_inputs_match": true, "validation_inputs": { "prompts_sha256": "136a918e5fee962f2b52f8e520a0275fd5ec5569181e0f5fdf4910ee3c34d528", "tools_sha256": "86ae46ebdbb0324c9672eec87e6b7f7683b0eb9c8c7169dcabdf646cf9bab122", "image_sha256": "0b1ae6badbe19a6049305c36d165a34cdf36df00033842e265339b2b7f057295" }, "mean_semantic_similarity": 0.9774516999721528, "exact_matches": 6, "comparisons": [ { "id": "fact", "reference_passed": true, "candidate_passed": true, "exact_match": true, "sequence_agreement": 1.0, "semantic_similarity": 1.0, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "math", "reference_passed": true, "candidate_passed": true, "exact_match": false, "sequence_agreement": 0.5480093676814989, "semantic_similarity": 0.9332988262176514, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "code", "reference_passed": true, "candidate_passed": true, "exact_match": false, "sequence_agreement": 0.32563025210084034, "semantic_similarity": 0.9706727266311646, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "multiturn", "reference_passed": true, "candidate_passed": true, "exact_match": true, "sequence_agreement": 1.0, "semantic_similarity": 1.0, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "thinking_off", "reference_passed": true, "candidate_passed": true, "exact_match": false, "sequence_agreement": 0.9695290858725761, "semantic_similarity": 0.9862716794013977, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "reasoning_low", "reference_passed": true, "candidate_passed": true, "exact_match": true, "sequence_agreement": 1.0, "semantic_similarity": 1.0, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "reasoning_high", "reference_passed": true, "candidate_passed": true, "exact_match": true, "sequence_agreement": 1.0, "semantic_similarity": 1.0, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "long_form", "reference_passed": true, "candidate_passed": true, "exact_match": false, "sequence_agreement": 0.17132798748288675, "semantic_similarity": 0.8842736482620239, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "non_english", "reference_passed": true, "candidate_passed": true, "exact_match": true, "sequence_agreement": 1.0, "semantic_similarity": 1.0000001192092896, "severe_regression": false, "candidate_loop_score": 0.0 }, { "id": "json", "reference_passed": true, "candidate_passed": true, "exact_match": true, "sequence_agreement": 1.0, "semantic_similarity": 1.0, "severe_regression": false, "candidate_loop_score": 0.0 } ], "functional_results": { "reference_text": { "passed": 10, "total": 10 }, "candidate_text": { "passed": 10, "total": 10 }, "reference_tools": { "passed": 5, "total": 5 }, "candidate_tools": { "passed": 5, "total": 5 }, "reference_vision": { "passed": 3, "total": 3 }, "candidate_vision": { "passed": 3, "total": 3 } }, "measurements": { "average_generation_tps": 12.568759032272249, "peak_memory_gb": 23.639191525, "artifact_bytes": 21989153029, "maximum_prompt_tokens_tested": 73, "loop_rate": 0.0 }, "warnings": [ "Semantic similarity is a measured embedding-model proxy, not ground-truth accuracy.", "Raw-logit equality is unavailable across all target runtimes; exact functional gates and output agreement are used for portable release validation.", "Sequence agreement is lexical and is reported diagnostically, not used as semantic accuracy." ] }, "quality_failures": [] }