-
-
Notifications
You must be signed in to change notification settings - Fork 5
Expand file tree
/
Copy path.env.example
More file actions
40 lines (36 loc) · 1.64 KB
/
Copy path.env.example
File metadata and controls
40 lines (36 loc) · 1.64 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
# Machine specific environment variables
#
# Use it from the repository root like:
#
# . .\load_env.ps1; llama-server
#
# Note: These variables can be overwritten with
# command line options (e.g. --models-preset).
# llama.cpp runtime configuration
#
# @see https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
LLAMA_ARG_HOST=0.0.0.0
LLAMA_ARG_PORT=8080
LLAMA_ARG_MODELS_DIR=D:\AI\LLM\gguf
LLAMA_ARG_MODELS_PRESET=presets\models_24GB_VRAM.ini
# CUDA driver configuration
#
# @see https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/environment-variables.html
CUDA_DEVICE_ORDER=PCI_BUS_ID
# Only models_16GB_8GB_VRAM.ini pins its devices. The other two presets leave
# split-mode and tensor-split unset, so on a multi-GPU host every entry spreads
# across all visible cards and the tier name becomes a floor rather than a cap
# — measured at 56.5 t/s on one 16 GB card against 36.5 t/s spread over a
# 16 GB plus an 8 GB one. Uncomment and set this when running either of those
# two presets here. Take the UUID from `nvidia-smi -L`; an index would follow
# CUDA_DEVICE_ORDER above rather than the nvidia-smi order.
#
# docs/presets.md -> Device pinning and multi-GPU
#CUDA_VISIBLE_DEVICES=GPU-00000000-0000-0000-0000-000000000000
# Needs an R615 or newer driver. Pairing CUDA_SCALE_LAUNCH_QUEUES with CUDA
# graphs, which llama.cpp enables by default, produced Xid 32 errors and
# application crashes on graphs with long chains of kernel nodes until the
# R615 branch fixed it (CUDA 13.4 release notes, resolved issue 5686696).
#
# @see https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html
CUDA_SCALE_LAUNCH_QUEUES=4x