vLLM Invocation

VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 VLLM_WORKER_MULTIPROC_METHOD=spawn vllm serve --model /llm/models/Qwen/Qwen3.8-27B-FP8 --dtype=float16 --enforce-eager --port 8000 --host 0.0.0.0 --trust-remote-code --disable-sliding-window --gpu-memory-util=0.90 --max-num-batched-tokens=8192 --max-model-len=192608 --block-size 64 -tp=2 --speculative-config '{"method":"mtp","num_speculative_tokens":2}' --enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 --max-num-seqs=4 --kv-cache-dtype fp8 

Popular posts from this blog

Create a new repo

Change Fedora Silverblue automatic updates to check

Add VMs to Fedora Atomic