Instructions to use senior5207/munche-768-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use senior5207/munche-768-GGUF with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("senior5207/munche-768-GGUF") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - llama-cpp-python
How to use senior5207/munche-768-GGUF with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="senior5207/munche-768-GGUF", filename="munche-768-f16.gguf", )
output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use senior5207/munche-768-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf senior5207/munche-768-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf senior5207/munche-768-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf senior5207/munche-768-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf senior5207/munche-768-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf senior5207/munche-768-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf senior5207/munche-768-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf senior5207/munche-768-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf senior5207/munche-768-GGUF:Q4_K_M
Use Docker
docker model run hf.co/senior5207/munche-768-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use senior5207/munche-768-GGUF with Ollama:
ollama run hf.co/senior5207/munche-768-GGUF:Q4_K_M
- Unsloth Studio
How to use senior5207/munche-768-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for senior5207/munche-768-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for senior5207/munche-768-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for senior5207/munche-768-GGUF to start chatting
- Atomic Chat new
- Docker Model Runner
How to use senior5207/munche-768-GGUF with Docker Model Runner:
docker model run hf.co/senior5207/munche-768-GGUF:Q4_K_M
- Lemonade
How to use senior5207/munche-768-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull senior5207/munche-768-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.munche-768-GGUF-Q4_K_M
List all available models
lemonade list
output = llm(
"Once upon a time,",
max_tokens=512,
echo=True
)
print(output)Munche-768 GGUF
Baragi-AI/Munche-768 ์ llama.cpp
์์ ์ธ ์ ์๋๋ก GGUF ๋ก ๋ณํํ ๊ฒ์
๋๋ค. ํ๊ตญ์ด ์์ค์ ๋ฌธ์ฒด ์ ์ฌ๋๋ฅผ ๋ํ๋ด๋
768์ฐจ์ ์๋ฒ ๋ฉ ๋ชจ๋ธ์
๋๋ค.
์๋ณธ์ google/embeddinggemma-300m
์ ๋ํ LoRA ์ด๋ํฐ์ด๋ฏ๋ก, ๋ฒ ์ด์ค์ ๋ณํฉํ ๋ค ๋ณํํ์ต๋๋ค. SentenceTransformer
ํ์ดํ๋ผ์ธ์ Dense ํ๋ก์ ์
(768โ3072โ768)๊ณผ mean pooling ๋ GGUF ์ ํฌํจ๋์ด
์์ด, ๋ณ๋ ํ์ฒ๋ฆฌ ์์ด ์๋ณธ๊ณผ ๊ฐ์ ์๋ฒ ๋ฉ์ด ๋์ต๋๋ค.
ํ์ผ
| ํ์ผ | ํฌ๊ธฐ | ์ต์ ์ฝ์ฌ์ธ ์ผ์น๋ | ์ ์ฌ๋ ์ต๋ ์ค์ฐจ |
|---|---|---|---|
munche-768-f32.gguf |
1.18 GB | 1.000000 | 0.000054 |
munche-768-f16.gguf |
593 MB | 0.999999 | 0.000121 |
munche-768-q8_0.gguf |
318 MB | 0.999437 | 0.002015 |
munche-768-q4_k_m.gguf |
228 MB | 0.990537 | 0.010029 |
F16 ์ ๊ถ์ฅํฉ๋๋ค. ์๋ณธ ๊ฐ์ค์น๋ F32 ์ด์ง๋ง F16 ๊ณผ์ ์ฐจ์ด๊ฐ ์ธก์ ๋ ธ์ด์ฆ ์์ค์ด๊ณ , ์ฉ๋์ ์ ๋ฐ์ ๋๋ค. F32 ๋ ์ฐธ์กฐ์ฉ์ผ๋ก ํจ๊ป ์ฌ๋ ค๋ก๋๋ค.
์ฉ๋์ด ์ค์ํ๋ฉด Q8_0 ์ด ๋ฌด๋ํฉ๋๋ค. Q4_K_M ์ ๋ฌธ์ฅ ๊ฐ ์ ์ฌ๋๊ฐ ์ต๋ 0.01 ๊น์ง ํ๋ค๋ฆฌ๋ฏ๋ก, ๋ฏธ์ธํ ๋ฌธ์ฒด ์ฐจ์ด๋ฅผ ๋ค๋ฃจ๋ ์ด ๋ชจ๋ธ์ ์ฉ๋์์๋ ์์๊ฐ ๋ค์งํ ์ ์์ต๋๋ค.
์ธก์ ๋ฐฉ๋ฒ: ๋ฌธ์ฒด๊ฐ ๋ค๋ฅธ ํ๊ตญ์ด ๋ฌธ์ฅ 5๊ฐ๋ฅผ ์๋ณธ PyTorch ๋ชจ๋ธ๊ณผ ๊ฐ GGUF ๋ก ์ธ์ฝ๋ฉํด, ๊ฐ์ ๋ฌธ์ฅ๋ผ๋ฆฌ์ ์ฝ์ฌ์ธ ์ ์ฌ๋(์ต์ ๊ฐ)์ ๋ฌธ์ฅ ๊ฐ ์ ์ฌ๋ ํ๋ ฌ์ ์ต๋ ์ ๋ ์ค์ฐจ๋ฅผ ๋น๊ตํ์ต๋๋ค.
์ฌ์ฉ๋ฒ
llama-server -m munche-768-f16.gguf --embeddings --pooling mean -c 2048 -ub 2048 -b 2048
-ub ์ -b ๋ฅผ 2048 ๋ก ์ง์ ํด์ผ ํฉ๋๋ค. ์๋ตํ๋ฉด llama.cpp ๊ฐ ๋ฐฐ์น ํฌ๊ธฐ๋ฅผ 512 ๋ก
๋ฎ์ถฐ์ ๊ธด ์
๋ ฅ์ด ์๋ฆฝ๋๋ค.
import numpy as np
import requests
texts = [
"๊ทธ๋ ์ฐฝ๋ฐ์ ์ค๋ ๋ฐ๋ผ๋ณด์๋ค. ๋น์๋ฆฌ๊ฐ ๋ฐฉ ์์ ๊ฐ๋ ์ฑ์ ๋ค.",
"์ผ, ๊ทธ๊ฑฐ ์ง์ง์ผ? ๋ง๋ ์ ๋ผ. ๋ ์ด์ ๊ฑ ๋ดค๋๋ฐ ์๋ฌด ๋ง๋ ์์๊ฑฐ๋ .",
]
response = requests.post(
"http://127.0.0.1:8080/v1/embeddings",
json={"input": texts, "model": "munche-768"},
)
rows = sorted(response.json()["data"], key=lambda r: r["index"])
embeddings = np.array([r["embedding"] for r in rows])
print(embeddings.shape) # (2, 768)
print(embeddings @ embeddings.T) # ์ฝ์ฌ์ธ ์ ์ฌ๋
--pooling mean ์ผ๋ก ๋์ฐ๋ฉด llama.cpp ๊ฐ L2 ์ ๊ทํ๊น์ง ๋ง์น ๋ฒกํฐ๋ฅผ ๋ฐํํ๋ฏ๋ก,
์ฝ์ฌ์ธ ์ ์ฌ๋๋ ๋ด์ ๋ง์ผ๋ก ๊ณ์ฐํ ์ ์์ต๋๋ค. ๋ค๋ฅธ pooling ์ต์
์ ์ฐ๊ฑฐ๋ ๊ฐ์
์ง์ ๋ค๋ฃฐ ๋๋ norm ์ ํ์ธํ์ธ์.
์ฃผ์์ฌํญ
ํ๋กฌํํธ ํ๋ฆฌํฝ์ค๋ ํฌํจ๋์ง ์์ต๋๋ค. EmbeddingGemma ๊ณ์ด์
task: search result | query: ๊ฐ์ ํ๋ฆฌํฝ์ค๋ฅผ ๋ถ์ฌ ์ฐ๋๋ก ์ค๊ณ๋์ด ์๋๋ฐ, ์ด
๊ท์น์ GGUF ์ ๋ค์ด๊ฐ์ง ์์ต๋๋ค. ์๋ณธ SentenceTransformer ์ encode_query() /
encode_document() ์ ๋์ผํ ๊ฒฐ๊ณผ๊ฐ ํ์ํ๋ค๋ฉด ํธ์ถํ๋ ์ชฝ์์ ํ๋ฆฌํฝ์ค๋ฅผ ์ง์
๋ถ์ฌ์ผ ํฉ๋๋ค. ์ ํ์ ์ผ์น๋๋ ์์ชฝ ๋ชจ๋ ํ๋ฆฌํฝ์ค ์์ด ์ธก์ ํ ๊ฐ์
๋๋ค.
์ต๋ ์ ๋ ฅ ๊ธธ์ด๋ 2,048 ํ ํฐ์ ๋๋ค.
๋ณํ ๋ฐฉ๋ฒ
google/embeddinggemma-300m ์ LoRA ์ด๋ํฐ๋ฅผ ๋ณํฉํ ๋ค llama.cpp ๋ก ๋ณํํ์ต๋๋ค.
python convert_hf_to_gguf.py munche-768-merged \
--outfile munche-768-f32.gguf \
--outtype f32 \
--sentence-transformers-dense-modules
llama-quantize munche-768-f32.gguf munche-768-q8_0.gguf Q8_0
--sentence-transformers-dense-modules ๊ฐ ์์ผ๋ฉด Dense ๋ ์ด์ด๊ฐ ๋น ์ ธ์, ์ฐจ์์
768 ๋ก ๊ฐ์ง๋ง ์๋ณธ๊ณผ ๋ค๋ฅธ ์๋ฒ ๋ฉ์ด ๋์ต๋๋ค.
๋ณํ ์ ์์๋ ์ ์ด ๋ ๊ฐ์ง ์์ต๋๋ค.
SentenceTransformer.save()๋tokenizer.model์ ์ ์ฅํ์ง ์์ต๋๋ค. ์ด ํ์ผ์ด ์์ผ๋ฉด ๋ณํ๊ธฐ๊ฐ sentencepiece ๋์ BPE ๊ฒฝ๋ก๋ฅผ ํ๊ณ , embeddinggemma ์ pre-tokenizer ํด์๊ฐ ๋ฑ๋ก๋์ด ์์ง ์์ ์คํจํฉ๋๋ค. ๋ฒ ์ด์ค ๋ฆฌํฌ์์ ํจ๊ป ๋ณต์ฌํด์ผ ํฉ๋๋ค.- ์๋ณธ ์ด๋ํฐ๋ ํ
์ ํค์
base_model.model.์ ๋์ฌ์.default๊ฐ ๋น ์ ธ ์์ด,PeftModel.from_pretrained()๋ก ๋ก๋ํ๋ฉด LoRA ๊ฐ ์ ์ฉ๋์ง ์์ ์ฑ ๊ฒฝ๊ณ ๋ง ์ถ๋ ฅ๋ฉ๋๋ค. ํค๋ฅผ ๊ต์ ํด ๋ณํฉํ์ต๋๋ค.
๋ผ์ด์ ์ค
์๋ณธ Munche-768 ๊ณผ ๋์ผํ๊ฒ Gemma Terms of Use ๋ฅผ ๋ฐ๋ฆ ๋๋ค. EmbeddingGemma ํ์๋ฌผ์ด๋ฏ๋ก ์ฌ์ฉ ์ ์ฝ๊ด์ ํ์ธํ์๊ธฐ ๋ฐ๋๋๋ค.
- ์๋ณธ ๋ชจ๋ธ: Baragi-AI/Munche-768 (Baragi AI)
- ๋ฒ ์ด์ค ๋ชจ๋ธ: google/embeddinggemma-300m (Google)
์ด ๋ฆฌํฌ๋ ํ์ ๋ณํ๋ง ์ํํ์ผ๋ฉฐ, ๋ชจ๋ธ ๊ฐ์ค์น์ ์ฑ๋ฅ์ ์๋ณธ์ ๋ฐ๋ฆ ๋๋ค. ํ์ต ๋ฐ์ดํฐ, ํ๊ฐ ๊ฒฐ๊ณผ, ํ๊ณ์ ์ ์๋ณธ ๋ชจ๋ธ ์นด๋๋ฅผ ์ฐธ๊ณ ํ์ธ์.
์ธ์ฉ
@software{munche768,
title = {Munche-768: Korean Fiction Style Embedding Model},
author = {Baragi AI},
year = {2026},
url = {https://huggingface.co/Baragi-AI/Munche-768}
}
- Downloads last month
- -
4-bit
8-bit
16-bit
32-bit
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="senior5207/munche-768-GGUF", filename="", )