Skip to content

Commit bab0ebe

Browse files
committed
docs(local-models): add vllm hidden-state serving notes
1 parent a10f972 commit bab0ebe

2 files changed

Lines changed: 101 additions & 1 deletion

File tree

‎docs/vllm-serve-hidden-state.md‎

Lines changed: 100 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,100 @@
1+
# vLLM Hidden-State Serving for Local Models
2+
3+
This page captures operational notes for serving a local vLLM model with hidden-state extraction enabled. The hidden-state connector writes prefill activations to `.safetensors` files and returns the actual file path in `kv_transfer_params.hidden_states_path`.
4+
5+
Docker is not required by the protocol. Use Docker when you want a reproducible CUDA/vLLM runtime; use `vllm serve` directly when the local Python environment has a vLLM build that includes `extract_hidden_states` and `ExampleHiddenStatesConnector`.
6+
7+
## Docker launch
8+
9+
Pick one filesystem path for hidden states and mount it into the container. The container path used in `shared_storage_path` must be the same path clients pass as `kv_transfer_params.hidden_states_path`.
10+
11+
```bash
12+
export HIDDEN_STATES_DIR=/tmp/vllm-hidden-states
13+
export HF_CACHE_DIR=/tmp/vllm-hf-cache
14+
mkdir -p "${HIDDEN_STATES_DIR}" "${HF_CACHE_DIR}"
15+
16+
docker run -d --name vllm_qwen35 \
17+
--gpus all \
18+
-p 0.0.0.0:8000:8000 \
19+
-v "${HF_CACHE_DIR}:/root/.cache/huggingface" \
20+
-v "${HIDDEN_STATES_DIR}:${HIDDEN_STATES_DIR}" \
21+
vllm/vllm-openai:latest-cu129 \
22+
Qwen/Qwen3.6-35B-A3B \
23+
--tensor-parallel-size 8 \
24+
--max-model-len 32768 \
25+
--reasoning-parser qwen3 \
26+
--enable-auto-tool-choice \
27+
--tool-call-parser hermes \
28+
--no-enable-chunked-prefill \
29+
--speculative-config '{"method":"extract_hidden_states","num_speculative_tokens":1,"draft_model_config":{"hf_config":{"eagle_aux_hidden_state_layer_ids":[39]}}}' \
30+
--kv-transfer-config '{"kv_connector":"ExampleHiddenStatesConnector","kv_role":"kv_producer","kv_connector_extra_config":{"shared_storage_path":"/tmp/vllm-hidden-states"}}'
31+
```
32+
33+
For `Qwen/Qwen3.6-35B-A3B`, layer `39` is the last hidden-state layer. To capture multiple layers, add each layer id to `eagle_aux_hidden_state_layer_ids`, for example `[0,1,2,39]`. Capturing all layers can make each probe much larger and may require a lower `--max-model-len` to leave enough KV-cache memory.
34+
35+
## Direct vLLM CLI launch
36+
37+
The direct CLI form serves the same model without Docker. There is no volume mount; `shared_storage_path` is a host path and clients must be able to read that same path.
38+
39+
```bash
40+
export HIDDEN_STATES_DIR=/tmp/vllm-hidden-states
41+
mkdir -p "${HIDDEN_STATES_DIR}"
42+
43+
vllm serve Qwen/Qwen3.6-35B-A3B \
44+
--host 0.0.0.0 \
45+
--port 8000 \
46+
--tensor-parallel-size 8 \
47+
--max-model-len 32768 \
48+
--reasoning-parser qwen3 \
49+
--enable-auto-tool-choice \
50+
--tool-call-parser hermes \
51+
--no-enable-chunked-prefill \
52+
--speculative-config '{"method":"extract_hidden_states","num_speculative_tokens":1,"draft_model_config":{"hf_config":{"eagle_aux_hidden_state_layer_ids":[39]}}}' \
53+
--kv-transfer-config '{"kv_connector":"ExampleHiddenStatesConnector","kv_role":"kv_producer","kv_connector_extra_config":{"shared_storage_path":"/tmp/vllm-hidden-states"}}'
54+
```
55+
56+
Use the direct CLI only after confirming your installed vLLM accepts both `--speculative-config '{"method":"extract_hidden_states",...}'` and `--kv-transfer-config '{"kv_connector":"ExampleHiddenStatesConnector",...}'`. If those flags fail, use the known container image or install a vLLM build that contains the connector.
57+
58+
## Verify one hidden-state file
59+
60+
Send one Chat Completions request with `max_tokens=1`. The probe should return a `kv_transfer_params.hidden_states_path` value that points at a `.safetensors` file.
61+
62+
```bash
63+
curl http://localhost:8000/v1/chat/completions \
64+
-H "Content-Type: application/json" \
65+
-d '{
66+
"model": "Qwen/Qwen3.6-35B-A3B",
67+
"messages": [{"role": "user", "content": "Return one short sentence."}],
68+
"max_tokens": 1,
69+
"kv_transfer_params": {
70+
"hidden_states_path": "/tmp/vllm-hidden-states",
71+
"include_output_tokens": false
72+
}
73+
}'
74+
```
75+
76+
Read the path from the response rather than assuming a filename. vLLM may choose the concrete safetensors file name.
77+
78+
```bash
79+
uv run python - <<'PY_INNER'
80+
from pathlib import Path
81+
from safetensors import safe_open
82+
83+
path = Path("/tmp/vllm-hidden-states")
84+
files = sorted(path.glob("*.safetensors"), key=lambda item: item.stat().st_mtime)
85+
if not files:
86+
raise SystemExit("no safetensors files written")
87+
88+
with safe_open(files[-1], framework="numpy") as handle:
89+
for key in handle.keys():
90+
tensor = handle.get_tensor(key)
91+
print(files[-1], key, tensor.shape, tensor.dtype)
92+
PY_INNER
93+
```
94+
95+
## Troubleshooting
96+
97+
- `probe response missing kv_transfer_params`: the server is not running with `ExampleHiddenStatesConnector`, or the request did not include `kv_transfer_params`.
98+
- `no safetensors files written`: check that `shared_storage_path` exists and is writable by the vLLM process.
99+
- Context-length startup errors from vLLM: lower `--max-model-len`, reduce the number of captured layers, or increase available GPU memory.
100+
- Hidden-state extraction does not work with chunked prefill; keep `--no-enable-chunked-prefill` in the launch command.

‎mkdocs.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -39,7 +39,7 @@ nav:
3939
exclude_docs: |
4040
internal/**
4141
research/**
42-
local_models.md
42+
vllm-serve-hidden-state.md
4343
4444
theme:
4545
name: material

0 commit comments

Comments
 (0)