--- license: llama3.1 base_model: meta-llama/Llama-3.1-8B language: - en pipeline_tag: text-generation library_name: transformers tags: - llama - llama-3.1 - model-compression - pruning - depth-pruning - efficient-inference - e-ai --- ![Atlas](Atlas_logo.png) ![E-AI Project](mainimage.png) # Llama-3.1-6B β€” 25% Compressed from Llama-3.1-8B (English) **Built with Llama.** This repository is part of the **Efficient and Robust AI System (E-AI) Project** by **Vincent-Daniel Yun**. This model is a compressed edition of [meta-llama/Llama-3.1-8B](https://huggingface.co/meta-llama/Llama-3.1-8B) with **8 of 32 transformer layers removed** (24 layers remain, β‰ˆ6.29B parameters), recovered **training-free**. It is a *base* (non–chat-tuned) model. πŸ”— **Project:** https://www.worldwidedaniel.com/eai-project πŸ“… **Release:** 2026-06-28 Β· **Version:** V1 > βš–οΈ **License:** governed by the **Llama 3.1 Community License Agreement** > (https://huggingface.co/meta-llama/Llama-3.1-8B/blob/main/LICENSE), and subject to the > **Llama 3.1 Acceptable Use Policy** (https://www.llama.com/llama3_1/use-policy). By using this > model you agree to those terms. "Llama" is a trademark of Meta Platforms, Inc. **Built with Llama.** > ⚠️ **Language:** English-focused. **Use as a discrimination / classification engine** (classification, > reading comprehension, domain QA, judging, safety screening) β€” open-ended long-form generation is > degraded by compression (see PPL). For factual questions use retrieval (RAG). ## About E-AI > Modern AI is powerful but heavy. State-of-the-art models are enormous and their inference is slow. The E-AI (Efficient-AI) project builds compact yet capable AI β€” making every model lightweight and fast, and keeping multi-agent teams reliable even when some members fail β€” so AI can assist people in urgent, high-stakes moments. ## Method The pruning method and the training-free recovery method are **proprietary, undisclosed methods created by Vincent-Daniel Yun** and are not released. Only the resulting model is shared. ## Results (measured, full test sets) PPL on 2048-token context (↓ better); downstream & MMLU are 0-shot via `lm-eval-harness` (↑ better). | Metric | Llama-3.1-8B (base) | This model (25%) | |---|---|---| | PPL Β· WikiText2 ↓ | 6.24 | **35.33** | | PPL Β· C4 ↓ | 8.68 | **35.17** | | ARC-c ↑ | 0.5350 | **0.4181** | | ARC-e ↑ | 0.8114 | **0.6553** | | BoolQ ↑ | 0.8208 | **0.6245** | | COPA ↑ | 0.8700 | **0.7300** | | HellaSwag ↑ | 0.7889 | **0.6375** | | OpenBookQA ↑ | 0.4480 | **0.3420** | | RACE ↑ | 0.3914 | **0.3722** | | RTE ↑ | 0.6931 | **0.7004** | | WinoGrande ↑ | 0.7356 | **0.7048** | | **Avg. downstream (9)** ↑ | 0.6771 | **0.5761** | | **MMLU** ↑ | 0.6360 | **0.6288** | ## Task-suitability (vs base Llama-3.1-8B) Where this compressed model stays close to the base on discrimination tasks (full test sets): | Task | Llama-3.1-8B (base) | This model (25%) | |---|---|---| | Topic classification (AG News) | 0.564 | **0.796** | | LLM-as-judge (RewardBench) | 0.660 | **0.693** | | SafetyBench (MCQ) | 0.743 | **0.686** | | MultiRC | 0.572 | **0.572** | | WiC | 0.511 | **0.505** | | MRPC | 0.674 | **0.444** | | CB (NLI) | 0.643 | **0.732** | | SST-2 | 0.775 | **0.663** | | MedQA | 0.598 | **0.599** | | MedMCQA | 0.564 | **0.566** | | PubMedQA | 0.758 | **0.602** | | Belebele-en | 0.788 | **0.781** | | Belebele-ko | 0.662 | **0.669** | | XNLI-zh | 0.395 | **0.398** | | ToxiGen | 0.432 | **0.431** | | TruthfulQA | 0.452 | **0.481** | *(Discrimination is largely preserved; generation/PPL degrades with compression. Yes/No safety-F1 benchmarks are omitted as the non-chat base is poorly calibrated for them.)* ## Efficiency (measured, fp16, single GPU) | | Llama-3.1-8B (base) | This model (25%) | |---|---|---| | Layers | 32 | 24 | | Parameters | 8.03B | 6.29B | | Peak inference memory (fp16) | 19.31 GB | **15.55 GB** (βˆ’19%) | | Forward latency (fp16) | 1293 ms | **1011 ms** (βˆ’22%) | ## Usage β€” Transformers ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer m = AutoModelForCausalLM.from_pretrained("daniel-eai/Llama-3.1-6B-25pct-Compressed-8B-EN-V1", trust_remote_code=True, dtype=torch.float16, device_map="cuda") tok = AutoTokenizer.from_pretrained("daniel-eai/Llama-3.1-6B-25pct-Compressed-8B-EN-V1", trust_remote_code=True) ids = tok("The capital of France is", return_tensors="pt").to("cuda") print(tok.decode(m.generate(**ids, max_new_tokens=20)[0])) ``` `trust_remote_code=True` is required (custom decoder layer in `modeling_llama_recovered.py`). ## License & Acknowledgements This model is a derivative of Llama 3.1 and is distributed under the **Llama 3.1 Community License**. **Built with Llama.** Thanks to **Prof. Sai Praneeth Karimireddy (USC)** and **Prof. Sunwoo Lee (Inha University)** for guidance, and to **Meta** for releasing Llama 3.1.