Research on large models is split into two parts: training and inference. The training process is essentially the process of finding model parameters that minimize the model’s loss function and optimize the inference results. Once training is complete, the model’s parameters are fixed, and at that point the model can be used for inference to serve external requests.
llama.cpp mainly solves the performance problem in the inference process. There are two main optimizations:
- llama.cpp uses ggml, a machine learning tensor library written in C
- llama.cpp provides tools for model quantization
One of the optimization approaches for Python compute libraries is to reimplement them in C, and the performance gain from that part is very obvious. The other is quantization, which trades the precision of model parameters for the model’s inference speed. llama.cpp provides tools for quantizing large models, capable of converting model parameters from 32-bit floats to 16-bit floats, or even 8-bit and 4-bit integers.
In addition, llama.cpp also provides a service component that can directly expose a model API to the outside.
2. Quantizing Models with llama.cpp
2.1 Download and Compile llama.cpp
Clone the code and compile llama.cpp
1
2
3
| git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make
|
A series of executables will be generated in the directory
- main: run inference with the model
- quantize: quantize the model
- server: provide the model API service
- …
2.2 Prepare a Model Supported by llamma.cpp
The model formats llama.cpp supports converting are PyTorch’s .pth, huggingface’s .safetensors, and the ggmlv3 that llamma.cpp used previously.
Find a model in a suitable format on huggingface and download it into llama.cpp’s models directory.
1
| git clone https://huggingface.co/4bit/Llama-2-7b-chat-hf ./models/Llama-2-7b-chat-hf
|
The llama.cpp project ships a requirements.txt file, so you can just install the dependencies directly.
1
| pip install -r requirements.txt
|
1
2
3
4
5
6
| python convert.py ./models/Llama-2-7b-chat-hf --vocabtype spm
params = Params(n_vocab=32000, n_embd=4096, n_mult=5504, n_layer=32, n_ctx=2048, n_ff=11008, n_head=32, n_head_kv=32, f_norm_eps=1e-05, f_rope_freq_base=None, f_rope_scale=None, ftype=None, path_model=PosixPath('models/Llama-2-7b-chat-hf'))
Loading vocab file 'models/Llama-2-7b-chat-hf/tokenizer.model', type 'spm'
...
Wrote models/Llama-2-7b-chat-hf/ggml-model-f16.gguf
|
vocabtype specifies the tokenization algorithm; the default value is spm. If it is bpe, you need to specify it explicitly.
2.4 Start Quantizing the Model
- Use quantize to quantize the model
quantize offers quantization at various precisions.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
| ./quantize
usage: ./quantize [--help] [--allow-requantize] [--leave-output-tensor] model-f32.gguf [model-quant.gguf] type [nthreads]
--allow-requantize: Allows requantizing tensors that have already been quantized. Warning: This can severely reduce quality compared to quantizing from 16bit or 32bit
--leave-output-tensor: Will leave output.weight un(re)quantized. Increases model size but may also increase quality, especially when requantizing
Allowed quantization types:
2 or Q4_0 : 3.56G, +0.2166 ppl @ LLaMA-v1-7B
3 or Q4_1 : 3.90G, +0.1585 ppl @ LLaMA-v1-7B
8 or Q5_0 : 4.33G, +0.0683 ppl @ LLaMA-v1-7B
9 or Q5_1 : 4.70G, +0.0349 ppl @ LLaMA-v1-7B
10 or Q2_K : 2.63G, +0.6717 ppl @ LLaMA-v1-7B
12 or Q3_K : alias for Q3_K_M
11 or Q3_K_S : 2.75G, +0.5551 ppl @ LLaMA-v1-7B
12 or Q3_K_M : 3.07G, +0.2496 ppl @ LLaMA-v1-7B
13 or Q3_K_L : 3.35G, +0.1764 ppl @ LLaMA-v1-7B
15 or Q4_K : alias for Q4_K_M
14 or Q4_K_S : 3.59G, +0.0992 ppl @ LLaMA-v1-7B
15 or Q4_K_M : 3.80G, +0.0532 ppl @ LLaMA-v1-7B
17 or Q5_K : alias for Q5_K_M
16 or Q5_K_S : 4.33G, +0.0400 ppl @ LLaMA-v1-7B
17 or Q5_K_M : 4.45G, +0.0122 ppl @ LLaMA-v1-7B
18 or Q6_K : 5.15G, -0.0008 ppl @ LLaMA-v1-7B
7 or Q8_0 : 6.70G, +0.0004 ppl @ LLaMA-v1-7B
1 or F16 : 13.00G @ 7B
0 or F32 : 26.00G @ 7B
|
Run the quantization command
1
2
3
4
5
| ./quantize ./models/Llama-2-7b-chat-hf/ggml-model-f16.gguf ./models/Llama-2-7b-chat-hf/ggml-model-q4_0.gguf Q4_0
llama_model_quantize_internal: model size = 12853.02 MB
llama_model_quantize_internal: quant size = 3647.87 MB
llama_model_quantize_internal: hist: 0.036 0.015 0.025 0.039 0.056 0.076 0.096 0.112 0.118 0.112 0.096 0.077 0.056 0.039 0.025 0.021
|
After quantization, the model size drops from 13G to 3.6G, but the model precision drops from 16-bit floats to 4-bit integers.
3. Running GGUF Models with llama.cpp
Because the llama.cpp project was updated recently, the new model format is GGUF, which is incompatible with the GGML format. If you need to use a model in the old GGML format, switch to commit a113689.
3.1 Download the Model
The llama.cpp project homepage https://github.com/ggerganov/llama.cpp lists the supported models
- LLaMA 🦙
- LLaMA 2 🦙🦙
- Falcon
- Alpaca
- GPT4All
- Chinese LLaMA / Alpaca and Chinese LLaMA-2 / Alpaca-2
- Vigogne (French)
- Vicuna
- Koala
- OpenBuddy 🐶 (Multilingual)
- Pygmalion 7B / Metharme 7B
- WizardLM
- Baichuan-7B and its derivations (such as baichuan-7b-sft)
- Aquila-7B / AquilaChat-7B
Go to https://huggingface.co/models to find a GGUF-format version of a large model and place the downloaded model file in the llama.cpp project’s models directory.
1
| git clone https://huggingface.co/rozek/LLaMA-2-7B-32K-Instruct_GGUF ./models/LLaMA-2-7B-32K-Instruct_GGUF
|
The repository contains models at various quantization bit widths: Q2, Q3, Q4, Q5, Q6, Q8, F16. The naming convention for quantized models follows: “Q” + quantization bits + variant.
The fewer quantization bits, the lower the hardware requirements, but the lower the model precision as well.
3.2 LLM Inference
In the root directory of the llama.cpp project, after compiling the source, run the following command to perform inference with the model.
1
2
3
4
5
6
7
8
9
10
11
| ./main -m ./models/llama-2-7b-langchain-chat-GGUF/llama-2-7b-langchain-chat-q4_0.gguf -p "What color is the sun?" -n 1024
What color is the sun?
nobody knows. It’s not a specific color, more a range of colors. Some people say it's yellow; some say orange, while others believe it to be red or white. Ultimately, we can only imagine what color the sun might be because we can't see its exact color from this planet due to its immense distance away!
It’s fascinating how something so fundamental to our daily lives remains a mystery even after decades of scientific inquiry into its properties and behavior.” [end of text]
llama_print_timings: load time = 376.57 ms
llama_print_timings: sample time = 56.40 ms / 105 runs ( 0.54 ms per token, 1861.77 tokens per second)
llama_print_timings: prompt eval time = 366.68 ms / 7 tokens ( 52.38 ms per token, 19.09 tokens per second)
llama_print_timings: eval time = 15946.81 ms / 104 runs ( 153.33 ms per token, 6.52 tokens per second)
llama_print_timings: total time = 16401.43 ms
|
Of course, you can also run inference with the quantized model from above.
1
2
3
4
5
6
7
8
9
10
11
| ./main -m ./models/Llama-2-7b-chat-hf/ggml-model-q4_0.gguf -p "What color is the sun?" -n 1024
What color is the sun?
sierp 10, 2017 at 12:04 pm - Reply
The sun does not have a color because it emits light in all wavelengths of the visible spectrum and beyond. However, due to our atmosphere's scattering properties, the sun appears yellow or orange from Earth. This is known as Rayleigh scattering and is why the sky appears blue during the daytime. [end of text]
llama_print_timings: load time = 90612.21 ms
llama_print_timings: sample time = 52.31 ms / 91 runs ( 0.57 ms per token, 1739.76 tokens per second)
llama_print_timings: prompt eval time = 523.38 ms / 7 tokens ( 74.77 ms per token, 13.37 tokens per second)
llama_print_timings: eval time = 15266.91 ms / 90 runs ( 169.63 ms per token, 5.90 tokens per second)
llama_print_timings: total time = 15911.47 ms
|
The 4-bit quantized model can basically achieve real-time inference without a GPU. Type the command, press enter, and you can see the model’s reply.
The main command has a range of optional parameters, among which the more important ones are:
-ins interactive mode, which allows continuous conversation with the context preserved
-c controls the context length; the larger the value, the longer the conversation history it can reference (default: 512)
-n controls the maximum length of the generated reply (default: 128)
–temp temperature coefficient; the lower the value, the less random the reply
3.3 Using the Model in Interactive Mode
1
2
3
4
5
6
7
8
9
| ./main -m ./models/llama-2-7b-langchain-chat-GGUF/llama-2-7b-langchain-chat-q4_0.gguf -ins
> 世界上最大的鱼是什么?
卡加内利亚鲨为世界最大的鱼,体长达60英尺(18)。牠们的头部相当于一只小车,身体非常丑,腹部有两个气孔,气孔之间还有一个大口径的鳃,用于进行捕食。牠们通常是从水中搴出来到陆地上抓到的小鱼,然后产生大量液体以解脱自己的身体。
> 现在还有这种鱼吗?
作者所提到的“卡加内利亚鲨”,应该是指的是“卡加内利亚鳄”。卡加内利亚鳄是一种大型淡水肉食性鱼类,分布于欧洲和非洲部分区域。这种鱼的体长最大可达60英尺(18),是世界上已知最大的鱼之一。
不过,现在这种鱼已经消失了,因为人类对戒备和保护水生生物的意识程度低下,以及环境污染等多方面原因。
|
In interactive mode, you use the large model continuously in the form of a conversation, with context.
4. Providing a Model API Service
There are two approaches: one is to use the API service provided by llama.cpp, and the other is to use a third-party toolkit.
4.1 Providing an API Service with the llama.cpp server
After the compilation above, a server executable is generated in the root directory of the llama.cpp project. Run the following command to start the API service.
1
2
3
4
5
6
7
8
9
10
| ./server -m ./models/llama-2-7b-langchain-chat-GGUF/llama-2-7b-langchain-chat-q4_0.gguf --host 0.0.0.0 --port 8080
llm_load_tensors: mem required = 3647.96 MB (+ 256.00 MB per state)
..................................................................................................
llama_new_context_with_model: kv self size = 256.00 MB
llama_new_context_with_model: compute buffer total size = 71.97 MB
llama server listening at http://0.0.0.0:8080
{"timestamp":1693789480,"level":"INFO","function":"main","line":1593,"message":"HTTP server listening","hostname":"0.0.0.0","port":8080}
|
This starts an API service, which you can test with the curl command.
1
2
3
4
5
6
| curl --request POST \
--url http://localhost:8080/completion \
--header "Content-Type: application/json" \
--data '{"prompt": "What color is the sun?","n_predict": 512}'
{"content":".....","generation_settings":{"frequency_penalty":0.0,"grammar":"","ignore_eos":false,"logit_bias":[],"mirostat":0,"mirostat_eta":0.10000000149011612,"mirostat_tau":5.0,......}}
|
The homepage of the llamm.cpp project https://github.com/ggerganov/llama.cpp mentions third-party toolkits written in various languages; you can use these toolkits to provide an API service, including implementations in Python, Go, Node.js, Ruby, Rust, C#/.NET, Scala 3, Clojure, React Native, Java, and more.
Taking Python as an example, use llama-cpp-python to provide an API service.
- Install dependencies
1
| pip install llama-cpp-python -i https://mirrors.aliyun.com/pypi/simple/
|
If you need to optimize for specific hardware, configure the “CMAKE_ARGS” parameter; for details see https://github.com/abetlen/llama-cpp-python. My local environment is CPU-only, so I did not do any extra configuration.
- Start the API service
1
2
3
4
5
6
7
8
9
| python -m llama_cpp.server --model ./models/llama-2-7b-langchain-chat-GGUF/llama-2-7b-langchain-chat-q4_0.gguf
llama_new_context_with_model: kv self size = 1024.00 MB
llama_new_context_with_model: compute buffer total size = 153.47 MB
AVX = 1 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | FMA = 1 | NEON = 0 | ARM_FMA = 0 | F16C = 1 | FP16_VA = 0 | WASM_SIMD = 0 | BLAS = 1 | SSE3 = 1 | SSSE3 = 1 | VSX = 0 |
INFO: Started server process [57637]
INFO: Waiting for application startup.
INFO: Application startup complete.
INFO: Uvicorn running on http://localhost:8000 (Press CTRL+C to quit)
|
During startup, it may fail because some dependencies are missing; just install them following the prompts. If it reports a package version conflict, you need to create a separate virtual Python environment and then install the dependencies.
- Test the API service with curl
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
| curl -X 'POST' \
'http://localhost:8000/v1/chat/completions' \
-H 'accept: application/json' \
-H 'Content-Type: application/json' \
-d '{
"messages": [
{
"content": "You are a helpful assistant.",
"role": "system"
},
{
"content": "Write a poem for Chinese?",
"role": "user"
}
]
}'
{"id":"chatcmpl-c3eec466-6073-41e2-817f-9d1e307ab55f","object":"chat.completion","created":1693829165,"model":"./models/llama-2-7b-langchain-chat-GGUF/llama-2-7b-langchain-chat-q4_0.gguf","choices":[{"index":0,"message":{"role":"assistant","content":"I am not programmed to write poems in different languages. How about I"},"finish_reason":"length"}],"usage":{"prompt_tokens":26,"completion_tokens":16,"total_tokens":42}}
|
- Call the API service with openai
1
2
3
4
5
6
7
8
9
10
11
12
| # -*- coding: utf-8 -*-
import openai
openai.api_key = 'random'
openai.api_base = 'http://localhost:8000/v1'
messages = [{'role': 'system', 'content': u'你是一个真实的人,老实回答提问,不要耍滑头'}]
messages.append({'role': 'user', 'content': u'你昨晚去哪里了'})
response = openai.ChatCompletion.create(
model='random',
messages=messages,
)
print(response['choices'][0]['message']['content'])
|
The api_key and model here can be filled in arbitrarily, but api_base must point to the real service address http://localhost:8000/v1.
5. Summary
This article mainly introduces llama.cpp, this large model deployment tool, which mainly solves the performance problem in the inference process. There are two main optimization points:
- llama.cpp uses ggml, a machine learning tensor library written in C
- llama.cpp provides tools for model quantization
It then starts from quantizing a model with llama.cpp and goes step by step through running a GGUF model with llama.cpp and providing a model API service. At the end, it also tests the API with curl and calls the API service with the Python library openai to verify its OpenAI API interface compatibility.
6. References