sarthak1 commited on
Commit
87455e8
·
verified ·
1 Parent(s): 68508a1

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ evaluation/similarity_matrix.png filter=lfs diff=lfs merge=lfs -text
37
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ QodoAI Open RAIL++-M License
2
+ Last Updated: February 19, 2025
3
+ Section I: PREAMBLE
4
+ This Open RAIL++-M License applies to Licensor's code embedding and integrity multimodal generative AI models.
5
+ This License strives for both the open and responsible downstream use of the Qodo Model accompanying this License. The openness here means enabling users of the Qodo Models to generally make royalty-free and subscription fee-free use of the applicable Qodo Model, modify and creative Derivatives thereof, while complying with use-based restrictions not permitting the use of the Qodo Model in very specific to ensure responsible use. While Derivatives of the Qodo Model can be released under different licensing terms, the latter will always have to include - at minimum - the same use-based restrictions as the ones in the original License.
6
+ This License governs the use of Qodo Models (and their Derivatives) and is informed by the model card associated with the relevant Qodo Model accompanying this License.
7
+ NOW THEREFORE, You and Licensor agree as follows:
8
+ 1. Definitions
9
+ (a) "License" means these terms and conditions for use, reproduction, and Distribution of the Qodo Models and their Derivatives, as defined herein. For the avoidance of doubt, Data and source code are not licensed under this License.
10
+ (b) "Data" means a collection of information and/or content extracted from the dataset used with the Qodo Model, including to train, pretrain, or otherwise evaluate the Qodo Model. The Data is not licensed under this License.
11
+ (c) "Output" means the results of operating a Qodo Model as embodied in informational content resulting therefrom.
12
+ (d) "Qodo Model" means any accompanying machine-learning based assemblies (including checkpoints), consisting of learnt weights, parameters (including optimizer states), corresponding to the model architecture as embodied in the Complementary Material, that have been trained or tuned, in whole or in part on the Data, using the Complementary Material.
13
+ (e) "Derivatives" means all modifications to the Qodo Model, works based on the Qodo Model, or any other model which is created or initialized by transfer of patterns of the weights, parameters, activations or output of the Qodo Model, to the other model, in order to cause the other model to perform similarly to the Qodo Model, including - but not limited to - distillation methods entailing the use of intermediate data representations or methods based on the generation of synthetic data by the Qodo Model for training the other model.
14
+ (f) "Complementary Material" means the accompanying scripts used to define, run, load, benchmark or evaluate the Qodo Model, and used to prepare data for training or evaluation, if any. This includes any accompanying documentation, tutorials, examples, etc, if any.
15
+ (g) "Distribution" means any transmission, reproduction, publication or other sharing of the Qodo Model or Derivatives of the Qodo Model to a third party, including providing the Qodo Model as a hosted or cloud service made available by electronic or other remote means - e.g. API-based or web-access.
16
+ (h) "Licensor", "Qodo" or "QodoAI" means Codium Ltd. (dba Qodo), or any corporate affiliate of Qodo authorized by it that is granting this License to You.
17
+ (i) "You" (or "Your") means an individual or legal entity exercising permissions granted by this License and/or making use of the Qodo Model for whichever purpose and in any field of use, including commercial usage of the Qodo Model in an end-use application - e.g. chatbot, translator, code generator. If You are accepting this License on behalf of a company or another legal entity, you represent that You have the authority to bind such legal entity to this License.
18
+ (j) "Third Parties" means individuals or legal entities that are not under common control with Licensor or You.
19
+ (k) "Contribution" means any work of authorship, including the original version of the Qodo Model and any modifications or additions to that Model or Derivatives of the Qodo Model thereof, that is intentionally submitted to Licensor for inclusion in the Qodo Model by the copyright owner or by an individual or legal entity authorized to submit on behalf of the copyright owner. For the purposes of this definition, "submitted" means any form of electronic, verbal, or written communication sent to the Licensor or its representatives, including but not limited to communication on electronic mailing lists, source code control systems, and issue tracking systems that are managed by, or on behalf of, the Licensor for the purpose of discussing and improving the Qodo Model, but excluding communication that is conspicuously marked or otherwise designated in writing by the copyright owner as "Not a Contribution."
20
+ (l) "Contributor" means Licensor and any individual or legal entity on behalf of whom a Contribution has been received by Licensor and subsequently incorporated within the Qodo Model.
21
+ Section II: INTELLECTUAL PROPERTY RIGHTS
22
+ Both copyright and patent grants apply to the Qodo Model, Derivatives of the Qodo Model and Complementary Material. The Qodo Model and Derivatives of the Qodo Model are subject to additional terms as described in Section III.
23
+ 2. Grant of Copyright License. Subject to the terms and conditions of this License, each Contributor hereby grants to You a perpetual, worldwide, non-exclusive, no-charge, royalty-free, copyright license to reproduce, prepare, publicly display, publicly perform, sublicense, and distribute the Complementary Material, the Qodo Model, and Derivatives of the Qodo Model.
24
+ 3. Grant of Patent License. Subject to the terms and conditions of this License and where and as applicable, each Contributor hereby grants to You a perpetual, worldwide, non-exclusive, no-charge, royalty-free, (except as stated in this paragraph) patent license to make, have made, use, offer to sell, sell, import, and otherwise transfer the Qodo Model and the Complementary Material, where such license applies only to those patent claims licensable by such Contributor that are necessarily infringed by their Contribution(s) alone or by combination of their Contribution(s) with the Qodo Model to which such Contribution(s) was submitted.
25
+ 4. Termination for IP Claims. If You institute litigation against any entity (including a cross-claim or counterclaim in a lawsuit) alleging that the Qodo Model and/or Complementary Material or a Contribution incorporated within the Qodo Model and/or Complementary Material constitutes direct or contributory intellectual rights (including copyright or patent) infringement, then any licenses granted to You under this License for the Qodo Model and/or Derivatives shall terminate as of the date such litigation is asserted or filed.
26
+ Section III: CONDITIONS OF USAGE, DISTRIBUTION AND REDISTRIBUTION
27
+ 5. Distribution and Redistribution. You may host for Third Parties remote access purposes (e.g. software-as-a-service), reproduce and distribute copies of the Qodo Model or Derivatives of the Qodo Model thereof in any medium, with or without modifications, provided that You meet the following conditions:
28
+ a. Use-based restrictions as referenced in paragraph 5 MUST be included as an enforceable provision by You in any type of legal agreement (e.g. a license, customer or subscription agreement) governing the use and/or Distribution of the Qodo Model or Derivatives of the Qodo Model, and You shall give notice to subsequent users You Distribute to, that the Qodo Model or Derivatives of the Qodo Model are subject to paragraph 5. This provision does not apply to the use of Complementary Material.
29
+ b. You must give any Third Party recipients of the Qodo Model or Derivatives of the Qodo Model a copy of this License;
30
+ c. You must cause any modified files to carry prominent notices stating that You changed the files;
31
+ d. You must retain all copyright, patent, trademark, and attribution notices excluding those notices that do not pertain to any part of the Qodo Model, Derivatives of the Qodo Model.
32
+
33
+ You may add Your own copyright statement to Your modifications and may provide additional or different license terms and conditions - respecting paragraph 4.a. - for use, reproduction, or Distribution of Your modifications, or for any such Derivatives of the Qodo Model as a whole, provided Your use, reproduction, and Distribution of the Qodo Model otherwise complies with the conditions stated in this License.
34
+ 6. Use-based restrictions. The restrictions set forth in Attachment A are considered use-based restrictions. Therefore You cannot use the Qodo Model and the Derivatives of the Qodo Model for the specified restricted uses. You may use the Qodo Model subject to this License, including only for lawful purposes and in accordance with the License. Use may include creating any content with, finetuning, distillation, updating, running, training, evaluating and/or reparametrizing the Qodo Model, provided that if You use the Qodo Model, the Derivatives or Complementary Materials to create, train, fine tune, or otherwise improve an AI model, which is distributed or otherwise made available, You must include "Qodo" at the beginning of any such AI model's name . You shall require all of Your users who use the Qodo Model or a Derivative of the Qodo Model to comply with the terms of this paragraph (paragraph 5).
35
+ 7. The Output You Generate. Except as set forth herein, Licensor claims no rights in the Output You generate using the Qodo Model. You are accountable for the Output you generate and its subsequent uses. No use of the output can contravene any provision as stated in the License.
36
+ Section IV: OTHER PROVISIONS
37
+ 8. Updates and Runtime Restrictions. To the maximum extent permitted by law, Licensor reserves the right to restrict (remotely or otherwise) usage of the Qodo Model in violation of this License, update the Qodo Model through electronic means, or modify the Output of the Qodo Model based on updates. You shall undertake reasonable efforts to use the latest version of the Qodo Model.
38
+ 9. Trademarks and related. Nothing in this License permits You to make use of Licensors' trademarks, trade names, logos or to otherwise suggest endorsement or misrepresent the relationship between the parties; and any rights not expressly granted herein are reserved by the Licensors.
39
+ 10. Disclaimer of Warranty. Unless required by applicable law or agreed to in writing, Licensor provides the Qodo Model and the Complementary Material (and each Contributor provides its Contributions) on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied, including, without limitation, any warranties or conditions of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A PARTICULAR PURPOSE. You are solely responsible for determining the appropriateness of using or redistributing the Qodo Model, Derivatives of the Qodo Model, and the Complementary Material and assume any risks associated with Your exercise of permissions under this License.
40
+ 11. Limitation of Liability. In no event and under no legal theory, whether in tort (including negligence), contract, or otherwise, unless required by applicable law (such as deliberate and grossly negligent acts) or agreed to in writing, shall any Contributor (including Licensor) be liable to You for damages, including any direct, indirect, special, incidental, or consequential damages of any character arising as a result of this License or out of the use or inability to use the Qodo Model and the Complementary Material (including but not limited to damages for loss of goodwill, work stoppage, computer or model failure or malfunction, or any and all other commercial damages or losses), even if such Contributor has been advised of the possibility of such damages. To the extent such limitation of direct damages is restricted under the laws of Your jurisdiction, then in respect of such jurisdiction Contributor's aggregate liability under, arising out of or otherwise in connection with this License shall be capped at $100.
41
+ 12. Indemnity. You will indemnify, defend, and hold harmless Qodo, its affiliates, and their respective shareholders, resellers, vendors, directors, employees and agents from and against all liabilities, damages, and costs (including reasonable attorneys' fees) arising out of any claim, demand, suit or proceeding by a third party arising out of or related to your use of or distribution of the Qodo Model, the Derivatives or Output.
42
+ 13. Accepting Warranty or Additional Liability. While redistributing the Qodo Model, the Derivatives of the Qodo Model and the Complementary Material thereof, You may choose to offer, and charge a fee for, acceptance of support, warranty, indemnity, or other liability obligations and/or rights consistent with this License. However, in accepting such obligations, You may act only on Your own behalf and on Your sole responsibility, not on behalf of any other Contributor, and only if You agree to indemnify, defend, and hold each Contributor harmless for any liability incurred by, or claims asserted against, such Contributor by reason of your accepting any such warranty or additional liability.
43
+ 14. Severability. If any provision of this License is held to be invalid, illegal or unenforceable, the remaining provisions shall be unaffected thereby and remain valid as if such provision had not been set forth herein.
44
+ 15. Amendments. Qodo may update, revise or amend this License, with or without notice, at its sole discretion, including to comply with regulatory requirements and in accordance with developments in the AI field, and will post such updated version on Qodo's website and/or 3rd party code repository website(s) and/or platform(s) where the Qodo Models are made available. Qodo encourages You to regularly review these. The last revision will be reflected in the "Last Updated" heading. Your continued use of the Qodo Models, Complementary Materials and Derivatives following any such amendment will be considered your consent to the updated License terms, and to the extent You disagree with the amendment, You must cease your use and distribution of the Qodo Models, the Complementary Materials and any Derivative thereof.
45
+ END OF TERMS AND CONDITIONS
46
+
47
+
48
+ Attachment A
49
+ Use Restrictions
50
+ You agree not to use the Qodo Model or Derivatives of the Qodo Model:
51
+ (a) In any way that violates any applicable national, federal, state, local or international law or regulation;
52
+ (b) For the purpose of exploiting, harming or attempting to exploit or harm minors in any way;
53
+ (c) To generate or disseminate verifiably false information and/or content with the purpose of harming others;
54
+ (d) To generate or disseminate personal identifiable information that can be used to harm an individual;
55
+ (e) To generate or disseminate information and/or content (e.g. images, code, posts, articles), and place the information and/or content in any context (e.g. bot generating tweets) without expressly and intelligibly disclaiming that the information and/or content is machine generated;
56
+ (f) To defame, disparage or otherwise harass others;
57
+ (g) To impersonate or attempt to impersonate (e.g. deepfakes) others without their consent;
58
+ (h) For fully automated decision making that adversely impacts an individual's legal rights or otherwise creates or modifies a binding, enforceable obligation;
59
+ (i) For any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics, or in a manner that is obscene, violent, hateful, racist or discriminatory, pornographic or is otherwise offensive;
60
+ (j) To exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm;
61
+ (k) For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories;
62
+ (l) To provide medical advice and medical results interpretation;
63
+ (m) To generate or disseminate information for the purpose to be used for administration of justice, law enforcement, immigration or asylum processes, such as predicting an individual will commit fraud/crime commitment (e.g. by text profiling, drawing causal relationships between assertions made in documents, indiscriminate and arbitrarily-targeted use).
64
+ (n) To infringe, violate or misappropriate any patent, trademark, trade secret, copyright, or other intellectual property or other proprietary rights of any other person.
65
+ (o) To transmit or otherwise makes available any malicious code, including any virus, worm, trojan horse, time bomb, web bug, spyware, or any other malicious or harmful computer code, library, file, or program.
66
+
README.md CHANGED
@@ -1,5 +1,155 @@
1
- ---
2
- license: other
3
- license_name: qodoai-open-rail-m
4
- license_link: https://www.qodo.ai/open-rail-m-license/
5
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qodo/Qodo-Embed-1-1.5B
3
+ library_name: model2vec
4
+ license: other
5
+ license_name: qodoai-open-rail-m
6
+ license_link: LICENSE
7
+ model_name: Qodo-Embed-M-1-1.5B-M2V-Distilled
8
+ tags:
9
+ - sentence-transformers
10
+ - sentence-similarity
11
+ - feature-extraction
12
+ - transformers
13
+ - Qwen2
14
+ ---
15
+
16
+ # Qodo-Embed-M-1-1.5B-M2V-Distilled
17
+
18
+ This project optimizes the Qodo-Embed-1-1.5B model using Model2Vec, reducing its size and dramatically improving inference speed while maintaining most of its performance capabilities.
19
+
20
+ ## Overview
21
+
22
+ [Qodo-Embed-1-1.5B](https://huggingface.co/Qodo/Qodo-Embed-1-1.5B) is a state-of-the-art code embedding model designed for retrieval tasks in the software development domain. While powerful, it can be resource-intensive for production use cases.
23
+
24
+ [Model2Vec](https://github.com/MinishLab/model2vec) is a technique to distill large sentence transformer models into small, fast static embedding models. This project applies Model2Vec to create an optimized version of Qodo-Embed-1-1.5B with the following benefits:
25
+
26
+ - **Smaller Size**: Reduces model size by a factor of 25x
27
+ - **Faster Inference**: Up to 112x faster inference
28
+ - **Low Resource Requirements**: Minimal memory footprint and dependencies
29
+ - **Maintains Performance**: Retains most of the original model's capabilities
30
+
31
+ ## Model Information
32
+
33
+ - **Model Name**: Qodo-Embed-M-1-1.5B-M2V-Distilled
34
+ - **Original Model**: [Qodo-Embed-1-1.5B](https://huggingface.co/Qodo/Qodo-Embed-1-1.5B)
35
+ - **Distillation Method**: [Model2Vec](https://github.com/MinishLab/model2vec)
36
+ - **Original Dimensions**: 1536
37
+ - **Distilled Dimensions**: 384
38
+ - **Explained Variance**: ~85%
39
+ - **Size Reduction**: 25.26x (from 5.9GB to 233MB)
40
+ - **Speed Improvement**: 112.14x faster
41
+
42
+ ## Installation
43
+
44
+ First, ensure you have the required dependencies:
45
+
46
+ ```bash
47
+ # Install the base package
48
+ uv add --group model2vec model2vec 'model2vec[distill]' sentence-transformers transformers
49
+
50
+ # Install additional dependencies for evaluation
51
+ uv add --group model2vec matplotlib psutil
52
+ ```
53
+
54
+ ## Usage
55
+
56
+ ### Distillation
57
+
58
+ To create a distilled version of Qodo-Embed-1-1.5B:
59
+
60
+ ```bash
61
+ python models/qodo_embed_m2v/distill.py --pca_dims 384
62
+ ```
63
+
64
+ Options:
65
+ - `--model_name` - Source model name (default: "Qodo/Qodo-Embed-1-1.5B")
66
+ - `--output_dir` - Where to save the distilled model (default: "models/qodo_embed_m2v")
67
+ - `--pca_dims` - Dimensions for PCA reduction; smaller values create faster but less accurate models (default: 384)
68
+ - `--save_to_hub` - Push the model to HuggingFace Hub
69
+ - `--hub_model_id` - Model ID for HuggingFace Hub (required if saving to hub)
70
+ - `--skip_readme` - Skip generating README file (default: True)
71
+
72
+ ### Evaluation
73
+
74
+ To evaluate the distilled model against the original:
75
+
76
+ ```bash
77
+ python models/qodo_embed_m2v/evaluate.py
78
+ ```
79
+
80
+ Options:
81
+ - `--original_model` - Original model name (default: "Qodo/Qodo-Embed-1-1.5B")
82
+ - `--distilled_model` - Path to the distilled model (default: "models/qodo_embed_m2v")
83
+ - `--output_dir` - Where to save evaluation results (default: "models/qodo_embed_m2v/evaluation")
84
+
85
+ ## Example Code
86
+
87
+ ```python
88
+ from model2vec import StaticModel
89
+ from sentence_transformers import SentenceTransformer
90
+ import time
91
+
92
+ # Sample code for embedding
93
+ code_samples = [
94
+ "def process_data_stream(source_iterator):",
95
+ "implement binary search tree",
96
+ "how to handle memory efficient data streaming",
97
+ """class LazyLoader:
98
+ def __init__(self, source):
99
+ self.generator = iter(source)
100
+ self._cache = []"""
101
+ ]
102
+
103
+ # Load original model
104
+ print("Loading original model...")
105
+ original_model = SentenceTransformer("Qodo/Qodo-Embed-1-1.5B")
106
+
107
+ # Load distilled model
108
+ print("Loading distilled model...")
109
+ distilled_model = StaticModel.from_pretrained("models/qodo_embed_m2v")
110
+
111
+ # Compare embedding speed
112
+ print("\nGenerating embeddings with original model...")
113
+ start = time.time()
114
+ original_embeddings = original_model.encode(code_samples)
115
+ original_time = time.time() - start
116
+ print(f"Original model took: {original_time:.4f} seconds")
117
+
118
+ print("\nGenerating embeddings with distilled model...")
119
+ start = time.time()
120
+ distilled_embeddings = distilled_model.encode(code_samples)
121
+ distilled_time = time.time() - start
122
+ print(f"Distilled model took: {distilled_time:.4f} seconds")
123
+ print(f"Speed improvement: {original_time/distilled_time:.2f}x faster")
124
+
125
+ print(f"\nOriginal embedding dimensions: {original_embeddings.shape}")
126
+ print(f"Distilled embedding dimensions: {distilled_embeddings.shape}")
127
+ ```
128
+
129
+ ## Results
130
+
131
+ The distilled model achieves:
132
+
133
+ - 25.26x reduction in model size (from 5.9GB to 233MB)
134
+ - 112.14x increase in inference speed
135
+ - 85.1% explained variance with PCA reduction to 384 dimensions
136
+
137
+ Detailed evaluation results, including similarity plots and performance metrics, are saved to the evaluation output directory.
138
+
139
+ ## Project Structure
140
+
141
+ - `distill.py` - Script to create the distilled model
142
+ - `evaluate.py` - Script to compare performance with the original model
143
+ - `example.py` - Example usage of the distilled model
144
+ - `evaluation/` - Directory containing evaluation results and visualizations
145
+
146
+ ## Acknowledgments
147
+
148
+ This project is built upon the following technologies:
149
+
150
+ - [Qodo-Embed-1-1.5B](https://huggingface.co/Qodo/Qodo-Embed-1-1.5B) - The original code embedding model developed by QodoAI
151
+ - [Model2Vec](https://github.com/MinishLab/model2vec) - The distillation technique used to optimize the model
152
+
153
+ ## License
154
+
155
+ This model is licensed under the [QodoAI-Open-RAIL-M](https://www.qodo.ai/open-rail-m-license/) license, the same as the original Qodo-Embed-1-1.5B model. Any derivative model must include "Qodo" at the beginning of its name per the license requirements.
config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "model2vec",
3
+ "architectures": [
4
+ "StaticModel"
5
+ ],
6
+ "tokenizer_name": "Qodo/Qodo-Embed-1-1.5B",
7
+ "apply_pca": 384,
8
+ "apply_zipf": null,
9
+ "sif_coefficient": 0.0001,
10
+ "hidden_dim": 384,
11
+ "seq_length": 1000000,
12
+ "normalize": true
13
+ }
distill.py ADDED
@@ -0,0 +1,113 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python
2
+ """
3
+ Script to distill Qodo-Embed-1-1.5B using Model2Vec.
4
+
5
+ This script performs the following operations:
6
+ 1. Downloads the Qodo-Embed-1-1.5B model
7
+ 2. Distills it using Model2Vec to create a smaller, faster static model
8
+ 3. Saves the distilled model for further use
9
+ """
10
+
11
+ import argparse
12
+ import logging
13
+ import shutil
14
+ import time
15
+ from pathlib import Path
16
+
17
+ from model2vec.distill import distill
18
+
19
+ # Configure logging
20
+ logging.basicConfig(level=logging.INFO, format="%(asctime)s - %(name)s - %(levelname)s - %(message)s")
21
+ logger = logging.getLogger(__name__)
22
+
23
+
24
+ def main() -> None:
25
+ """Run the distillation process for Qodo-Embed-1-1.5B."""
26
+ parser = argparse.ArgumentParser(description="Distill Qodo-Embed-1-1.5B using Model2Vec")
27
+ parser.add_argument(
28
+ "--model_name", default="Qodo/Qodo-Embed-1-1.5B", help="Model name or path for the source model"
29
+ )
30
+ parser.add_argument("--output_dir", default="models/qodo_embed_m2v", help="Directory to save the distilled model")
31
+ parser.add_argument(
32
+ "--pca_dims", type=int, default=384, help="Dimensions for PCA reduction (smaller = faster but less accurate)"
33
+ )
34
+ parser.add_argument("--save_to_hub", action="store_true", help="Whether to push the model to HuggingFace Hub")
35
+ parser.add_argument("--hub_model_id", default=None, help="Model ID for HuggingFace Hub (if saving to hub)")
36
+ parser.add_argument("--skip_readme", action="store_true", default=True, help="Skip generating the README file")
37
+
38
+ args = parser.parse_args()
39
+
40
+ # Create output directory if it doesn't exist
41
+ output_dir = Path(args.output_dir)
42
+ output_dir.mkdir(parents=True, exist_ok=True)
43
+
44
+ logger.info(f"Starting distillation of {args.model_name}")
45
+ logger.info(f"Distilled model will be saved to {output_dir}")
46
+ logger.info(f"Using PCA dimensions: {args.pca_dims}")
47
+ logger.info(f"Skipping README generation: {args.skip_readme}")
48
+
49
+ # Record start time for benchmarking
50
+ start_time = time.time()
51
+
52
+ # Run the distillation
53
+ try:
54
+ logger.info("Starting Model2Vec distillation...")
55
+ m2v_model = distill(
56
+ model_name=args.model_name,
57
+ pca_dims=args.pca_dims,
58
+ )
59
+
60
+ distill_time = time.time() - start_time
61
+ logger.info(f"Distillation completed in {distill_time:.2f} seconds")
62
+
63
+ # Save the distilled model
64
+ m2v_model.save_pretrained(args.output_dir)
65
+ logger.info(f"Model saved to {args.output_dir}")
66
+
67
+ # Remove README.md if it was created and we want to skip it
68
+ if args.skip_readme and (output_dir / "README.md").exists():
69
+ (output_dir / "README.md").unlink()
70
+ logger.info("Removed auto-generated README.md")
71
+
72
+ # Get model size information
73
+ model_size_mb = sum(
74
+ f.stat().st_size for f in output_dir.glob("**/*") if f.is_file() and f.name != "README.md"
75
+ ) / (1024 * 1024)
76
+ logger.info(f"Distilled model size: {model_size_mb:.2f} MB")
77
+
78
+ # Push to hub if requested
79
+ if args.save_to_hub:
80
+ if args.hub_model_id:
81
+ logger.info(f"Pushing model to HuggingFace Hub as {args.hub_model_id}")
82
+
83
+ # Create a temporary README for Hub upload if needed
84
+ readme_path = output_dir / "README.md"
85
+ had_readme = readme_path.exists()
86
+
87
+ if args.skip_readme and had_readme:
88
+ # Backup the README
89
+ shutil.move(readme_path, output_dir / "README.md.bak")
90
+
91
+ # Push to Hub
92
+ m2v_model.push_to_hub(args.hub_model_id)
93
+
94
+ # Restore state
95
+ if args.skip_readme:
96
+ if had_readme:
97
+ # Restore the backup
98
+ shutil.move(output_dir / "README.md.bak", readme_path)
99
+ elif (output_dir / "README.md").exists():
100
+ # Remove README created during push_to_hub
101
+ (output_dir / "README.md").unlink()
102
+ else:
103
+ logger.error("--hub_model_id must be specified when using --save_to_hub")
104
+
105
+ logger.info("Distillation process completed successfully!")
106
+
107
+ except Exception:
108
+ logger.exception("Error during distillation")
109
+ raise
110
+
111
+
112
+ if __name__ == "__main__":
113
+ main()
evaluate.py ADDED
@@ -0,0 +1,356 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python
2
+ """
3
+ Script to evaluate the performance of the distilled Qodo-Embed model.
4
+
5
+ This script performs the following:
6
+ 1. Loads both the original Qodo-Embed-1-1.5B model and the distilled version
7
+ 2. Compares them on:
8
+ - Embedding similarity
9
+ - Inference speed
10
+ - Memory usage
11
+ 3. Outputs a comprehensive evaluation report
12
+ """
13
+
14
+ import argparse
15
+ import gc
16
+ import logging
17
+ import os
18
+ import time
19
+ from pathlib import Path
20
+ from typing import Any
21
+
22
+ import matplotlib.pyplot as plt
23
+ import numpy as np
24
+ import psutil
25
+ import torch
26
+ from model2vec import StaticModel
27
+ from sentence_transformers import SentenceTransformer
28
+ from sklearn.metrics.pairwise import cosine_similarity
29
+
30
+ # For transformer models
31
+ from transformers.models.auto.modeling_auto import AutoModel
32
+ from transformers.models.auto.tokenization_auto import AutoTokenizer
33
+
34
+ # Configure logging
35
+ logging.basicConfig(level=logging.INFO, format="%(asctime)s - %(name)s - %(levelname)s - %(message)s")
36
+ logger = logging.getLogger(__name__)
37
+
38
+ # Sample texts for evaluation
39
+ SAMPLE_TEXTS = [
40
+ "def process_data_stream(source_iterator):",
41
+ "implement binary search tree",
42
+ "how to handle memory efficient data streaming",
43
+ """class LazyLoader:
44
+ def __init__(self, source):
45
+ self.generator = iter(source)
46
+ self._cache = []""",
47
+ """def dfs_traversal(root):
48
+ if not root:
49
+ return []
50
+ visited = []
51
+ stack = [root]
52
+ while stack:
53
+ node = stack.pop()
54
+ visited.append(node.val)
55
+ if node.right:
56
+ stack.append(node.right)
57
+ if node.left:
58
+ stack.append(node.left)
59
+ return visited""",
60
+ ]
61
+
62
+
63
+ def load_models(original_model_name: str, distilled_model_path: str) -> tuple[Any, StaticModel]:
64
+ """Load both the original and distilled models."""
65
+ logger.info(f"Loading original model: {original_model_name}")
66
+
67
+ try:
68
+ # Try to load as a sentence transformer first
69
+ original_model = SentenceTransformer(original_model_name)
70
+ model_type = "sentence_transformer"
71
+ except Exception:
72
+ # If that fails, try loading as a Hugging Face transformer
73
+ AutoTokenizer.from_pretrained(original_model_name)
74
+ original_model = AutoModel.from_pretrained(original_model_name)
75
+ model_type = "huggingface"
76
+
77
+ logger.info(f"Loading distilled model from: {distilled_model_path}")
78
+ distilled_model = StaticModel.from_pretrained(distilled_model_path)
79
+
80
+ return (original_model, model_type), distilled_model
81
+
82
+
83
+ def measure_memory_usage(model: Any) -> float:
84
+ """Measure memory usage of a model in MB."""
85
+ gc.collect()
86
+ torch.cuda.empty_cache() if torch.cuda.is_available() else None
87
+
88
+ process = psutil.Process(os.getpid())
89
+ memory_before = process.memory_info().rss / (1024 * 1024) # MB
90
+
91
+ # Force model to allocate memory if it hasn't already
92
+ if isinstance(model, StaticModel) or hasattr(model, "encode"):
93
+ _ = model.encode(["Test"])
94
+ else:
95
+ # For HF models, we need to handle differently
96
+ pass
97
+
98
+ gc.collect()
99
+ torch.cuda.empty_cache() if torch.cuda.is_available() else None
100
+
101
+ process = psutil.Process(os.getpid())
102
+ memory_after = process.memory_info().rss / (1024 * 1024) # MB
103
+
104
+ return memory_after - memory_before
105
+
106
+
107
+ def compute_embeddings(
108
+ original_model: Any, original_model_type: str, distilled_model: StaticModel, texts: list[str]
109
+ ) -> tuple[np.ndarray, np.ndarray]:
110
+ """Compute embeddings using both models."""
111
+ # Original model embeddings
112
+ if original_model_type == "sentence_transformer":
113
+ original_embeddings = original_model.encode(texts)
114
+ else:
115
+ # For HF models, we need more custom code
116
+ # Simple mean pooling function for HF models
117
+ def mean_pooling(model_output, attention_mask):
118
+ token_embeddings = model_output[0]
119
+ input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
120
+ return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(
121
+ input_mask_expanded.sum(1), min=1e-9
122
+ )
123
+
124
+ tokenizer = AutoTokenizer.from_pretrained(original_model.config._name_or_path)
125
+ encoded_input = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
126
+
127
+ with torch.no_grad():
128
+ model_output = original_model(**encoded_input)
129
+ original_embeddings = mean_pooling(model_output, encoded_input["attention_mask"]).numpy()
130
+
131
+ # Distilled model embeddings
132
+ distilled_embeddings = distilled_model.encode(texts)
133
+
134
+ return original_embeddings, distilled_embeddings
135
+
136
+
137
+ def measure_inference_speed(model: Any, model_type: str, texts: list[str], n_runs: int = 5) -> float:
138
+ """Measure inference speed in texts/second."""
139
+ # Warmup
140
+ if model_type in {"sentence_transformer", "static_model"}:
141
+ _ = model.encode(texts[:1])
142
+ else:
143
+ # Warmup for HF models
144
+ tokenizer = AutoTokenizer.from_pretrained(model.config._name_or_path)
145
+ encoded_input = tokenizer(texts[:1], padding=True, truncation=True, return_tensors="pt")
146
+ with torch.no_grad():
147
+ _ = model(**encoded_input)
148
+
149
+ # Measure speed
150
+ start_time = time.time()
151
+
152
+ if model_type in {"sentence_transformer", "static_model"}:
153
+ for _ in range(n_runs):
154
+ _ = model.encode(texts)
155
+ else:
156
+ # For HF models
157
+ tokenizer = AutoTokenizer.from_pretrained(model.config._name_or_path)
158
+ for _ in range(n_runs):
159
+ encoded_input = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
160
+ with torch.no_grad():
161
+ _ = model(**encoded_input)
162
+
163
+ total_time = time.time() - start_time
164
+ return (len(texts) * n_runs) / total_time
165
+
166
+
167
+ def compute_cosine_similarity(embeddings1: np.ndarray, embeddings2: np.ndarray) -> np.ndarray:
168
+ """Compute cosine similarity between embeddings, handling different dimensions.
169
+
170
+ For embeddings with different dimensions, we compute similarity by comparing
171
+ how they rank the same texts (semantically equivalent).
172
+ """
173
+ # Ensure embeddings1 and embeddings2 are 2D arrays with shapes (n_samples, n_features)
174
+ if embeddings1.ndim == 1:
175
+ embeddings1 = embeddings1.reshape(1, -1)
176
+ if embeddings2.ndim == 1:
177
+ embeddings2 = embeddings2.reshape(1, -1)
178
+
179
+ # Check and transpose if needed to ensure samples are in rows
180
+ if embeddings2.shape[0] != len(SAMPLE_TEXTS) and embeddings2.shape[1] == len(SAMPLE_TEXTS):
181
+ embeddings2 = embeddings2.T
182
+
183
+ logger.info(f"Embeddings shapes: original={embeddings1.shape}, distilled={embeddings2.shape}")
184
+
185
+ # If dimensions differ, we compute similarity matrix based on how each model ranks text pairs
186
+ # This is a form of semantic similarity evaluation rather than direct vector comparison
187
+ similarity_matrix = np.zeros((len(SAMPLE_TEXTS), len(SAMPLE_TEXTS)))
188
+
189
+ # Compute similarity matrices within each embedding space
190
+ sim1 = cosine_similarity(embeddings1)
191
+ sim2 = cosine_similarity(embeddings2)
192
+
193
+ # The similarity between samples i and j is the correlation between how they rank other samples
194
+ for i in range(len(SAMPLE_TEXTS)):
195
+ for j in range(len(SAMPLE_TEXTS)):
196
+ # For diagonal elements (same sample), use a direct measure of how similar
197
+ # the two models rank that sample against all others
198
+ if i == j:
199
+ # Pearson correlation between the rankings (excluding self-comparison)
200
+ rankings1 = np.delete(sim1[i], i)
201
+ rankings2 = np.delete(sim2[i], i)
202
+ # Higher correlation means the models agree on the semantic similarity
203
+ similarity_matrix[i, j] = np.corrcoef(rankings1, rankings2)[0, 1]
204
+ else:
205
+ # For off-diagonal elements, compare how similarly both models relate samples i and j
206
+ similarity_matrix[i, j] = 1 - abs(sim1[i, j] - sim2[i, j])
207
+
208
+ return similarity_matrix
209
+
210
+
211
+ def format_size(size_bytes: float) -> str:
212
+ """Format size in bytes to human-readable format."""
213
+ for unit in ["B", "KB", "MB", "GB"]:
214
+ if size_bytes < 1024.0:
215
+ return f"{size_bytes:.2f} {unit}"
216
+ size_bytes /= 1024.0
217
+ return f"{size_bytes:.2f} TB"
218
+
219
+
220
+ def plot_comparison(results: dict[str, Any], output_dir: str) -> None:
221
+ """Generate comparison plots and save them."""
222
+ output_path = Path(output_dir)
223
+ output_path.mkdir(exist_ok=True, parents=True)
224
+
225
+ # Speed comparison
226
+ plt.figure(figsize=(10, 6))
227
+ models = ["Original", "Distilled"]
228
+ speeds = [results["original_speed"], results["distilled_speed"]]
229
+ plt.bar(models, speeds, color=["#1f77b4", "#ff7f0e"])
230
+ plt.ylabel("Texts per second")
231
+ plt.title("Inference Speed Comparison")
232
+ plt.savefig(output_path / "speed_comparison.png", dpi=300, bbox_inches="tight")
233
+
234
+ # Memory comparison
235
+ plt.figure(figsize=(10, 6))
236
+ memories = [results["original_memory"], results["distilled_memory"]]
237
+ plt.bar(models, memories, color=["#1f77b4", "#ff7f0e"])
238
+ plt.ylabel("Memory Usage (MB)")
239
+ plt.title("Memory Usage Comparison")
240
+ plt.savefig(output_path / "memory_comparison.png", dpi=300, bbox_inches="tight")
241
+
242
+ # Size comparison
243
+ plt.figure(figsize=(10, 6))
244
+ sizes = [results["original_size"], results["distilled_size"]]
245
+ plt.bar(models, sizes, color=["#1f77b4", "#ff7f0e"])
246
+ plt.ylabel("Model Size (MB)")
247
+ plt.title("Model Size Comparison")
248
+ plt.savefig(output_path / "size_comparison.png", dpi=300, bbox_inches="tight")
249
+
250
+ # Similarity matrix heatmap
251
+ plt.figure(figsize=(8, 6))
252
+ plt.imshow(results["similarity_matrix"], cmap="viridis", interpolation="nearest")
253
+ plt.colorbar(label="Cosine Similarity")
254
+ plt.title("Embedding Similarity Between Original and Distilled Models")
255
+ plt.xticks([])
256
+ plt.yticks(range(len(SAMPLE_TEXTS)), [t[:20] + "..." if len(t) > 20 else t for t in SAMPLE_TEXTS])
257
+ plt.savefig(output_path / "similarity_matrix.png", dpi=300, bbox_inches="tight")
258
+
259
+
260
+ def evaluate_models(original_model_name: str, distilled_model_path: str, output_dir: str):
261
+ """Evaluate the original and distilled models."""
262
+ # Load models
263
+ (original_model, original_model_type), distilled_model = load_models(original_model_name, distilled_model_path)
264
+
265
+ # Measure model sizes
266
+ original_model_size = sum(p.numel() * 4 for p in original_model.parameters()) / (
267
+ 1024 * 1024
268
+ ) # MB (assuming float32)
269
+ distilled_model_size = sum(f.stat().st_size for f in Path(distilled_model_path).glob("**/*") if f.is_file()) / (
270
+ 1024 * 1024
271
+ ) # MB
272
+
273
+ # Measure memory usage
274
+ original_memory = measure_memory_usage(original_model)
275
+ distilled_memory = measure_memory_usage(distilled_model)
276
+
277
+ # Compute embeddings
278
+ original_embeddings, distilled_embeddings = compute_embeddings(
279
+ original_model, original_model_type, distilled_model, SAMPLE_TEXTS
280
+ )
281
+
282
+ # Compute similarity between embeddings
283
+ similarity_matrix = compute_cosine_similarity(original_embeddings, distilled_embeddings)
284
+ similarity_diagonal = np.diag(similarity_matrix)
285
+ avg_similarity = np.mean(similarity_diagonal)
286
+
287
+ # Measure inference speed
288
+ original_speed = measure_inference_speed(original_model, original_model_type, SAMPLE_TEXTS, n_runs=5)
289
+ distilled_speed = measure_inference_speed(distilled_model, "static_model", SAMPLE_TEXTS, n_runs=5)
290
+
291
+ # Collect results
292
+ results = {
293
+ "original_size": original_model_size,
294
+ "distilled_size": distilled_model_size,
295
+ "original_memory": original_memory,
296
+ "distilled_memory": distilled_memory,
297
+ "similarity_matrix": similarity_matrix,
298
+ "avg_similarity": avg_similarity,
299
+ "original_speed": original_speed,
300
+ "distilled_speed": distilled_speed,
301
+ "speed_improvement": distilled_speed / original_speed if original_speed > 0 else float("inf"),
302
+ "size_reduction": original_model_size / distilled_model_size if distilled_model_size > 0 else float("inf"),
303
+ "memory_reduction": original_memory / distilled_memory if distilled_memory > 0 else float("inf"),
304
+ }
305
+
306
+ # Generate plots
307
+ plot_comparison(results, output_dir)
308
+
309
+ # Print results
310
+ logger.info("\n" + "=" * 50)
311
+ logger.info("Model Evaluation Results")
312
+ logger.info("=" * 50)
313
+ logger.info(f"Original Model Size: {results['original_size']:.2f} MB")
314
+ logger.info(f"Distilled Model Size: {results['distilled_size']:.2f} MB")
315
+ logger.info(f"Size Reduction Factor: {results['size_reduction']:.2f}x")
316
+ logger.info("\n")
317
+ logger.info(f"Original Model Memory: {results['original_memory']:.2f} MB")
318
+ logger.info(f"Distilled Model Memory: {results['distilled_memory']:.2f} MB")
319
+ logger.info(f"Memory Reduction Factor: {results['memory_reduction']:.2f}x")
320
+ logger.info("\n")
321
+ logger.info(f"Original Model Speed: {results['original_speed']:.2f} texts/second")
322
+ logger.info(f"Distilled Model Speed: {results['distilled_speed']:.2f} texts/second")
323
+ logger.info(f"Speed Improvement Factor: {results['speed_improvement']:.2f}x")
324
+ logger.info("\n")
325
+ logger.info(f"Average Embedding Similarity: {results['avg_similarity']:.4f}")
326
+ logger.info("=" * 50)
327
+
328
+ return results
329
+
330
+
331
+ def main() -> None:
332
+ """Run the evaluation process."""
333
+ parser = argparse.ArgumentParser(description="Evaluate the distilled model against the original")
334
+ parser.add_argument("--original_model", default="Qodo/Qodo-Embed-1-1.5B", help="Original model name or path")
335
+ parser.add_argument("--distilled_model", default="models/qodo_embed_m2v", help="Path to the distilled model")
336
+ parser.add_argument(
337
+ "--output_dir", default="models/qodo_embed_m2v/evaluation", help="Directory to save evaluation results"
338
+ )
339
+
340
+ args = parser.parse_args()
341
+
342
+ # Create output directory
343
+ output_dir = Path(args.output_dir)
344
+ output_dir.mkdir(parents=True, exist_ok=True)
345
+
346
+ # Run evaluation
347
+ try:
348
+ evaluate_models(args.original_model, args.distilled_model, args.output_dir)
349
+ logger.info(f"Evaluation completed. Results saved to {args.output_dir}")
350
+ except Exception:
351
+ logger.exception("Error during evaluation")
352
+ raise
353
+
354
+
355
+ if __name__ == "__main__":
356
+ main()
evaluation/memory_comparison.png ADDED
evaluation/similarity_matrix.png ADDED

Git LFS Details

  • SHA256: 4285ab87463ab9dc1d217805061e50125da7c118b078d64f0287d6eded3ac641
  • Pointer size: 131 Bytes
  • Size of remote file: 112 kB
evaluation/size_comparison.png ADDED
evaluation/speed_comparison.png ADDED
example.py ADDED
@@ -0,0 +1,335 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python
2
+ """
3
+ Example script demonstrating how to use the distilled Qodo-Embed-M-1-1.5B-M2V-Distilled model.
4
+
5
+ This script shows:
6
+ 1. Loading both original and distilled models
7
+ 2. Running performance benchmarks
8
+ 3. Example code search and similarity scenarios
9
+ """
10
+
11
+ import argparse
12
+ import logging
13
+ import time
14
+
15
+ import numpy as np
16
+ from model2vec import StaticModel
17
+ from sentence_transformers import SentenceTransformer
18
+
19
+ # Configure logging
20
+ logging.basicConfig(level=logging.INFO, format="%(asctime)s - %(name)s - %(levelname)s - %(message)s")
21
+ logger = logging.getLogger(__name__)
22
+
23
+ # Sample code snippets for demonstration
24
+ CODE_SAMPLES = [
25
+ "def process_data_stream(source_iterator):",
26
+ "implement binary search tree",
27
+ "how to handle memory efficient data streaming",
28
+ """class LazyLoader:
29
+ def __init__(self, source):
30
+ self.generator = iter(source)
31
+ self._cache = []""",
32
+ """def dfs_traversal(root):
33
+ if not root:
34
+ return []
35
+ visited = []
36
+ stack = [root]
37
+ while stack:
38
+ node = stack.pop()
39
+ visited.append(node.val)
40
+ if node.right:
41
+ stack.append(node.right)
42
+ if node.left:
43
+ stack.append(node.left)
44
+ return visited""",
45
+ """def quick_sort(arr):
46
+ if len(arr) <= 1:
47
+ return arr
48
+ pivot = arr[len(arr) // 2]
49
+ left = [x for x in arr if x < pivot]
50
+ middle = [x for x in arr if x == pivot]
51
+ right = [x for x in arr if x > pivot]
52
+ return quick_sort(left) + middle + quick_sort(right)""",
53
+ """def fibonacci(n, memo={}):
54
+ if n in memo:
55
+ return memo[n]
56
+ if n <= 1:
57
+ return n
58
+ memo[n] = fibonacci(n-1, memo) + fibonacci(n-2, memo)
59
+ return memo[n]""",
60
+ ]
61
+
62
+ # Example code database for search scenarios
63
+ CODE_DATABASE = [
64
+ """def binary_search(arr, target):
65
+ left, right = 0, len(arr) - 1
66
+ while left <= right:
67
+ mid = (left + right) // 2
68
+ if arr[mid] == target:
69
+ return mid
70
+ elif arr[mid] < target:
71
+ left = mid + 1
72
+ else:
73
+ right = mid - 1
74
+ return -1""",
75
+ """class BinarySearchTree:
76
+ def __init__(self, value=None):
77
+ self.value = value
78
+ self.left = None
79
+ self.right = None
80
+
81
+ def insert(self, value):
82
+ if self.value is None:
83
+ self.value = value
84
+ return
85
+ if value < self.value:
86
+ if self.left is None:
87
+ self.left = BinarySearchTree(value)
88
+ else:
89
+ self.left.insert(value)
90
+ else:
91
+ if self.right is None:
92
+ self.right = BinarySearchTree(value)
93
+ else:
94
+ self.right.insert(value)""",
95
+ """def process_stream(source):
96
+ buffer = []
97
+ for item in source:
98
+ if len(buffer) >= 1000:
99
+ yield buffer
100
+ buffer = []
101
+ buffer.append(item)
102
+ if buffer: # Don't forget the remainder
103
+ yield buffer""",
104
+ """class StreamProcessor:
105
+ def __init__(self, chunk_size=1000):
106
+ self.chunk_size = chunk_size
107
+
108
+ def process(self, data_stream):
109
+ chunks = []
110
+ current_chunk = []
111
+ for item in data_stream:
112
+ current_chunk.append(item)
113
+ if len(current_chunk) >= self.chunk_size:
114
+ chunks.append(current_chunk)
115
+ current_chunk = []
116
+ if current_chunk:
117
+ chunks.append(current_chunk)
118
+ return chunks""",
119
+ """def memory_efficient_generator(large_file_path):
120
+ with open(large_file_path, 'r') as f:
121
+ for line in f:
122
+ yield process_line(line)
123
+
124
+ def process_line(line):
125
+ # Process the line somehow
126
+ return line.strip().upper()""",
127
+ """class CachedLoader:
128
+ def __init__(self, datasource, cache_size=100):
129
+ self.datasource = datasource
130
+ self.cache_size = cache_size
131
+ self.cache = {}
132
+
133
+ def get(self, key):
134
+ if key in self.cache:
135
+ return self.cache[key]
136
+
137
+ value = self.datasource.fetch(key)
138
+
139
+ # Add to cache, potentially evicting oldest item
140
+ if len(self.cache) >= self.cache_size:
141
+ oldest_key = next(iter(self.cache))
142
+ del self.cache[oldest_key]
143
+
144
+ self.cache[key] = value
145
+ return value""",
146
+ """def breadth_first_search(root):
147
+ if not root:
148
+ return []
149
+
150
+ result = []
151
+ queue = [root]
152
+
153
+ while queue:
154
+ node = queue.pop(0)
155
+ result.append(node.value)
156
+
157
+ if node.left:
158
+ queue.append(node.left)
159
+ if node.right:
160
+ queue.append(node.right)
161
+
162
+ return result""",
163
+ ]
164
+
165
+
166
+ def cosine_similarity(v1, v2):
167
+ """Compute cosine similarity between two vectors."""
168
+ dot_product = np.dot(v1, v2)
169
+ norm_v1 = np.linalg.norm(v1)
170
+ norm_v2 = np.linalg.norm(v2)
171
+ return dot_product / (norm_v1 * norm_v2)
172
+
173
+
174
+ def calculate_semantic_similarity(emb1, emb2, samples):
175
+ """Calculate semantic similarity between embeddings of different dimensions.
176
+
177
+ Instead of direct vector comparison, measure how similarly they rank the same samples.
178
+ """
179
+ # Calculate similarity matrices within each embedding space
180
+ sim1 = np.zeros((len(samples), len(samples)))
181
+ sim2 = np.zeros((len(samples), len(samples)))
182
+
183
+ # Calculate pairwise similarities within each embedding space
184
+ for i in range(len(samples)):
185
+ for j in range(len(samples)):
186
+ e1_i, e1_j = emb1[i], emb1[j]
187
+ e2_i, e2_j = emb2[i], emb2[j]
188
+
189
+ sim1[i, j] = np.dot(e1_i, e1_j) / (np.linalg.norm(e1_i) * np.linalg.norm(e1_j))
190
+ sim2[i, j] = np.dot(e2_i, e2_j) / (np.linalg.norm(e2_i) * np.linalg.norm(e2_j))
191
+
192
+ # Calculate correlation between similarity matrices
193
+ similarities = []
194
+ for i in range(len(samples)):
195
+ # Get rankings for this sample (exclude self-comparison)
196
+ rankings1 = np.delete(sim1[i], i)
197
+ rankings2 = np.delete(sim2[i], i)
198
+
199
+ # Calculate correlation
200
+ corr = np.corrcoef(rankings1, rankings2)[0, 1]
201
+ similarities.append(corr)
202
+
203
+ return similarities
204
+
205
+
206
+ def run_speed_benchmark(original_model, distilled_model, samples):
207
+ """Run speed benchmark comparing original and distilled models."""
208
+ logger.info("\n" + "=" * 50)
209
+ logger.info("SPEED BENCHMARK")
210
+ logger.info("=" * 50)
211
+
212
+ # Warmup
213
+ _ = original_model.encode(samples[:1])
214
+ _ = distilled_model.encode(samples[:1])
215
+
216
+ # Test original model
217
+ logger.info("Testing original model...")
218
+ start_time = time.time()
219
+ original_embeddings = original_model.encode(samples)
220
+ original_time = time.time() - start_time
221
+ logger.info(f"Original model took: {original_time:.4f} seconds for {len(samples)} samples")
222
+ logger.info(f"Speed: {len(samples) / original_time:.2f} samples/second")
223
+
224
+ # Test distilled model
225
+ logger.info("\nTesting distilled model...")
226
+ start_time = time.time()
227
+ distilled_embeddings = distilled_model.encode(samples)
228
+ distilled_time = time.time() - start_time
229
+ logger.info(f"Distilled model took: {distilled_time:.4f} seconds for {len(samples)} samples")
230
+ logger.info(f"Speed: {len(samples) / distilled_time:.2f} samples/second")
231
+
232
+ # Compare
233
+ speedup = original_time / distilled_time
234
+ logger.info(f"\nSpeedup factor: {speedup:.2f}x")
235
+
236
+ # Check embedding dimensions
237
+ logger.info(f"\nOriginal embedding dimensions: {original_embeddings.shape}")
238
+ logger.info(f"Distilled embedding dimensions: {distilled_embeddings.shape}")
239
+
240
+ return original_embeddings, distilled_embeddings
241
+
242
+
243
+ def demonstrate_similarity(original_embeddings, distilled_embeddings, samples):
244
+ """Demonstrate similarity between original and distilled embeddings."""
245
+ logger.info("\n" + "=" * 50)
246
+ logger.info("EMBEDDING SIMILARITY")
247
+ logger.info("=" * 50)
248
+
249
+ # Compare embeddings semantic similarity
250
+ similarities = calculate_semantic_similarity(original_embeddings, distilled_embeddings, samples)
251
+
252
+ for i, sim in enumerate(similarities):
253
+ logger.info(f"Sample {i + 1}: Semantic Similarity = {sim:.4f}")
254
+
255
+ avg_similarity = np.mean(similarities)
256
+ logger.info(f"\nAverage semantic similarity: {avg_similarity:.4f}")
257
+ logger.info("Note: Similarity is measured by how similarly both models rank the relationships between samples")
258
+
259
+ return avg_similarity
260
+
261
+
262
+ def demonstrate_code_search(distilled_model, query, code_database) -> None:
263
+ """Demonstrate code search functionality with the distilled model."""
264
+ logger.info("\n" + "=" * 50)
265
+ logger.info("CODE SEARCH DEMO")
266
+ logger.info("=" * 50)
267
+
268
+ logger.info(f"Query: '{query}'")
269
+
270
+ # Get query embedding
271
+ query_embedding = distilled_model.encode(query)
272
+
273
+ # Get database embeddings
274
+ database_embeddings = distilled_model.encode(code_database)
275
+
276
+ # Calculate similarities
277
+ similarities = []
278
+ for i, code_embedding in enumerate(database_embeddings):
279
+ sim = cosine_similarity(query_embedding, code_embedding)
280
+ similarities.append((i, sim))
281
+
282
+ # Sort by similarity
283
+ similarities.sort(key=lambda x: x[1], reverse=True)
284
+
285
+ # Display results
286
+ logger.info("\nTop 3 results:")
287
+ for i, (idx, sim) in enumerate(similarities[:3]):
288
+ code_snippet = code_database[idx]
289
+ if len(code_snippet) > 100:
290
+ code_snippet = code_snippet[:100] + "..."
291
+ logger.info(f"{i + 1}. Similarity: {sim:.4f}")
292
+ logger.info(f" {code_snippet}")
293
+ logger.info("")
294
+
295
+
296
+ def main() -> None:
297
+ """Run the example and benchmarks."""
298
+ parser = argparse.ArgumentParser(description="Example usage of distilled Qodo-Embed model")
299
+ parser.add_argument("--original_model", default="Qodo/Qodo-Embed-1-1.5B", help="Original model name or path")
300
+ parser.add_argument("--distilled_model", default="models/qodo_embed_m2v", help="Path to the distilled model")
301
+
302
+ args = parser.parse_args()
303
+
304
+ logger.info(f"Using original model: {args.original_model}")
305
+ logger.info(f"Using distilled model: {args.distilled_model}")
306
+
307
+ # Load models
308
+ logger.info("\nLoading original model...")
309
+ original_model = SentenceTransformer(args.original_model)
310
+
311
+ logger.info("Loading distilled model (Qodo-Embed-M-1-1.5B-M2V-Distilled)...")
312
+ distilled_model = StaticModel.from_pretrained(args.distilled_model)
313
+
314
+ # Run benchmarks
315
+ original_embeddings, distilled_embeddings = run_speed_benchmark(original_model, distilled_model, CODE_SAMPLES)
316
+
317
+ # Compare embedding similarity
318
+ demonstrate_similarity(original_embeddings, distilled_embeddings, CODE_SAMPLES)
319
+
320
+ # Demonstrate code search
321
+ for query in ["binary search implementation", "memory efficient streaming", "tree traversal algorithm"]:
322
+ demonstrate_code_search(distilled_model, query, CODE_DATABASE)
323
+
324
+ logger.info("\n" + "=" * 50)
325
+ logger.info("Model Summary: Qodo-Embed-M-1-1.5B-M2V-Distilled")
326
+ logger.info("=" * 50)
327
+ logger.info("• 25.26x smaller than the original model (233MB vs 5.9GB)")
328
+ logger.info("• 112.14x faster inference speed")
329
+ logger.info("• Preserves 85.1% of the original model's explained variance")
330
+ logger.info("• Same semantic search capabilities in a vastly more efficient package")
331
+ logger.info("=" * 50)
332
+
333
+
334
+ if __name__ == "__main__":
335
+ main()
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ceabf8fc5d30c1899511f7bc3b70895f575354327cebf74b6f21726bb97c63eb
3
+ size 232928344
modules.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "idx": 0,
4
+ "name": "0",
5
+ "path": ".",
6
+ "type": "sentence_transformers.models.StaticEmbedding"
7
+ },
8
+ {
9
+ "idx": 1,
10
+ "name": "1",
11
+ "path": "1_Normalize",
12
+ "type": "sentence_transformers.models.Normalize"
13
+ }
14
+ ]
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d8372feaa064372d176aff57e8f1e64f194814bb074519104f64c66a2825f091
3
+ size 11419037