Instructions to use flyingfishinwater/good_and_small_models with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use flyingfishinwater/good_and_small_models with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf flyingfishinwater/good_and_small_models:Q4_K_M # Run inference directly in the terminal: llama cli -hf flyingfishinwater/good_and_small_models:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf flyingfishinwater/good_and_small_models:Q4_K_M # Run inference directly in the terminal: llama cli -hf flyingfishinwater/good_and_small_models:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf flyingfishinwater/good_and_small_models:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf flyingfishinwater/good_and_small_models:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf flyingfishinwater/good_and_small_models:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf flyingfishinwater/good_and_small_models:Q4_K_M
Use Docker
docker model run hf.co/flyingfishinwater/good_and_small_models:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use flyingfishinwater/good_and_small_models with Ollama:
ollama run hf.co/flyingfishinwater/good_and_small_models:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use flyingfishinwater/good_and_small_models with Docker Model Runner:
docker model run hf.co/flyingfishinwater/good_and_small_models:Q4_K_M
- Lemonade
How to use flyingfishinwater/good_and_small_models with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull flyingfishinwater/good_and_small_models:Q4_K_M
Run and chat with the model
lemonade run user.good_and_small_models-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Upload models.json
Browse files- models.json +3 -3
models.json
CHANGED
|
@@ -6,7 +6,7 @@
|
|
| 6 |
"model_url": "https://",
|
| 7 |
"model_info_url": "https://huggingface.co/princeton-nlp/Sheared-LLaMA-1.3B",
|
| 8 |
"model_avatar": "ava0",
|
| 9 |
-
"model_intention": "It's good for talking and casual writing.",
|
| 10 |
"model_license": "license_llama2.txt",
|
| 11 |
"model_license_info": "Meta Llama 2 Community License Agreement",
|
| 12 |
"model_license_url": "https://ai.meta.com/llama/license/",
|
|
@@ -322,7 +322,7 @@
|
|
| 322 |
"file_size": 2218,
|
| 323 |
"context" : 4096,
|
| 324 |
"temp" : 0.6,
|
| 325 |
-
"prompt_format" : "<|im_start|>
|
| 326 |
"top_k" : 5,
|
| 327 |
"top_p" : 0.9,
|
| 328 |
"model_inference" : "llama",
|
|
@@ -353,7 +353,7 @@
|
|
| 353 |
"model_description": "This model is based on Mistral-7b-v0.2 with 16k context lengths. It's a uncensored model and supports a variety of instruction, conversational, and coding skills.",
|
| 354 |
"developer": "Eric Hartford and Cognitive Computations",
|
| 355 |
"developer_url": "https://erichartford.com/",
|
| 356 |
-
"file_size":
|
| 357 |
"context" : 16384,
|
| 358 |
"temp" : 0.6,
|
| 359 |
"prompt_format" : "<|im_start|>user\n{{prompt}}\n<|im_end|>\n<|im_start|>assistant\n",
|
|
|
|
| 6 |
"model_url": "https://",
|
| 7 |
"model_info_url": "https://huggingface.co/princeton-nlp/Sheared-LLaMA-1.3B",
|
| 8 |
"model_avatar": "ava0",
|
| 9 |
+
"model_intention": "It's good for talking and casual writing. Most devices can run it well.",
|
| 10 |
"model_license": "license_llama2.txt",
|
| 11 |
"model_license_info": "Meta Llama 2 Community License Agreement",
|
| 12 |
"model_license_url": "https://ai.meta.com/llama/license/",
|
|
|
|
| 322 |
"file_size": 2218,
|
| 323 |
"context" : 4096,
|
| 324 |
"temp" : 0.6,
|
| 325 |
+
"prompt_format" : "<|im_start|>user\n{{prompt}}\n<|im_end|>\n<|im_start|>assistant\n",
|
| 326 |
"top_k" : 5,
|
| 327 |
"top_p" : 0.9,
|
| 328 |
"model_inference" : "llama",
|
|
|
|
| 353 |
"model_description": "This model is based on Mistral-7b-v0.2 with 16k context lengths. It's a uncensored model and supports a variety of instruction, conversational, and coding skills.",
|
| 354 |
"developer": "Eric Hartford and Cognitive Computations",
|
| 355 |
"developer_url": "https://erichartford.com/",
|
| 356 |
+
"file_size": 2728,
|
| 357 |
"context" : 16384,
|
| 358 |
"temp" : 0.6,
|
| 359 |
"prompt_format" : "<|im_start|>user\n{{prompt}}\n<|im_end|>\n<|im_start|>assistant\n",
|