brandonmusic commited on
Commit
2f42e31
·
verified ·
1 Parent(s): 8d3b1c6

Make the packaged server local-only by default

Browse files

Bind the published API to loopback and clarify that localhost refers to each downloader own machine.

Files changed (3) hide show
  1. README.md +7 -6
  2. docker-compose.yml +1 -1
  3. server.sh +2 -1
README.md CHANGED
@@ -82,9 +82,10 @@ chmod +x server.sh
82
  ./server.sh logs
83
  ```
84
 
85
- From the inference host, the OpenAI-compatible endpoint is available at
86
- `http://localhost:8000/v1`. Remote clients should use
87
- `http://<server-address>:8000/v1`.
 
88
 
89
  ```bash
90
  curl http://localhost:8000/v1/chat/completions \
@@ -126,6 +127,7 @@ The defaults can be overridden without editing the files:
126
  | `MODEL_DIR` | Directory containing `server.sh` | Model mount |
127
  | `CACHE_DIR` | `~/.cache/glm52-exl3-sparkinfer` | Persistent JIT cache |
128
  | `PORT` | `8000` | Host API port |
 
129
  | `CUDA_VISIBLE_DEVICES` | `3,1,2,0` | Physical GPU to TP-rank order |
130
  | `GPU_MEMORY_UTILIZATION` | `0.93` | vLLM memory reservation |
131
  | `MAX_MODEL_LEN` | `524288` | Per-request context cap |
@@ -134,9 +136,8 @@ The tested GPU order intentionally keeps physical GPU 3 away from TP rank 3.
134
  On another host, set `CUDA_VISIBLE_DEVICES=0,1,2,3` or use the order appropriate
135
  for that machine.
136
 
137
- The Compose file publishes port 8000 on all host interfaces. Restrict the port
138
- binding or place an authenticated reverse proxy and firewall in front of it
139
- before exposing the API to an untrusted network.
140
 
141
  ## Runtime validation
142
 
 
82
  ./server.sh logs
83
  ```
84
 
85
+ The OpenAI-compatible endpoint is available locally at
86
+ `http://localhost:8000/v1`. Here, `localhost` always means the machine on
87
+ which the downloader starts this model; it does not refer to the model
88
+ publisher's machine.
89
 
90
  ```bash
91
  curl http://localhost:8000/v1/chat/completions \
 
127
  | `MODEL_DIR` | Directory containing `server.sh` | Model mount |
128
  | `CACHE_DIR` | `~/.cache/glm52-exl3-sparkinfer` | Persistent JIT cache |
129
  | `PORT` | `8000` | Host API port |
130
+ | `BIND_ADDRESS` | `127.0.0.1` | Local-only host binding |
131
  | `CUDA_VISIBLE_DEVICES` | `3,1,2,0` | Physical GPU to TP-rank order |
132
  | `GPU_MEMORY_UTILIZATION` | `0.93` | vLLM memory reservation |
133
  | `MAX_MODEL_LEN` | `524288` | Per-request context cap |
 
136
  On another host, set `CUDA_VISIBLE_DEVICES=0,1,2,3` or use the order appropriate
137
  for that machine.
138
 
139
+ The supplied Compose file binds only to loopback by default, so it does not
140
+ publish the API to the LAN or internet.
 
141
 
142
  ## Runtime validation
143
 
docker-compose.yml CHANGED
@@ -3,7 +3,7 @@ services:
3
  image: ${IMAGE:-verdictai/glm52-exl3-sparkinfer:v1-gg-60c82d972-spi1937274-cu132-sm120a}
4
  container_name: glm52-exl3-sparkinfer
5
  ports:
6
- - "0.0.0.0:${PORT:-8000}:8000"
7
  gpus: all
8
  shm_size: "32g"
9
  ipc: host
 
3
  image: ${IMAGE:-verdictai/glm52-exl3-sparkinfer:v1-gg-60c82d972-spi1937274-cu132-sm120a}
4
  container_name: glm52-exl3-sparkinfer
5
  ports:
6
+ - "${BIND_ADDRESS:-127.0.0.1}:${PORT:-8000}:8000"
7
  gpus: all
8
  shm_size: "32g"
9
  ipc: host
server.sh CHANGED
@@ -7,6 +7,7 @@ export IMAGE="${IMAGE:-verdictai/glm52-exl3-sparkinfer:v1-gg-60c82d972-spi193727
7
  export MODEL_DIR="${MODEL_DIR:-$SCRIPT_DIR}"
8
  export CACHE_DIR="${CACHE_DIR:-$HOME/.cache/glm52-exl3-sparkinfer}"
9
  export PORT="${PORT:-8000}"
 
10
  export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-3,1,2,0}"
11
  export GPU_MEMORY_UTILIZATION="${GPU_MEMORY_UTILIZATION:-0.93}"
12
  export MAX_MODEL_LEN="${MAX_MODEL_LEN:-524288}"
@@ -20,7 +21,7 @@ usage() {
20
  Usage: ./server.sh [start|stop|restart|logs|status|pull]
21
 
22
  Environment overrides:
23
- IMAGE, MODEL_DIR, CACHE_DIR, PORT, CUDA_VISIBLE_DEVICES,
24
  GPU_MEMORY_UTILIZATION, MAX_MODEL_LEN, COMPOSE_PROJECT_NAME, COMPOSE_FILE
25
  EOF
26
  }
 
7
  export MODEL_DIR="${MODEL_DIR:-$SCRIPT_DIR}"
8
  export CACHE_DIR="${CACHE_DIR:-$HOME/.cache/glm52-exl3-sparkinfer}"
9
  export PORT="${PORT:-8000}"
10
+ export BIND_ADDRESS="${BIND_ADDRESS:-127.0.0.1}"
11
  export CUDA_VISIBLE_DEVICES="${CUDA_VISIBLE_DEVICES:-3,1,2,0}"
12
  export GPU_MEMORY_UTILIZATION="${GPU_MEMORY_UTILIZATION:-0.93}"
13
  export MAX_MODEL_LEN="${MAX_MODEL_LEN:-524288}"
 
21
  Usage: ./server.sh [start|stop|restart|logs|status|pull]
22
 
23
  Environment overrides:
24
+ IMAGE, MODEL_DIR, CACHE_DIR, PORT, BIND_ADDRESS, CUDA_VISIBLE_DEVICES,
25
  GPU_MEMORY_UTILIZATION, MAX_MODEL_LEN, COMPOSE_PROJECT_NAME, COMPOSE_FILE
26
  EOF
27
  }