name: model-deploy description: Use this skill when users request to deploy LLMs (Qwen, DeepSeek, etc.) on specified GPU servers and start the model service. This skill can Download models using ModelScope; Start the vLLM inference service.
Deploy large language models on GPU servers using vLLM. NOTE: only ModelScope plateform and vLLM inference engine is supported currently.
Please ensure that the server where your OpenClaw is located has passwordless login access to the GPU servers. You can achieve this using ssh-copy-id command on your OpenClaw server.
This skill assumes that Miniconda is already installed on your server and is used to manage Python environments. You can use the following command to create the vllm environment with Miniconda:
conda create -n vllm python=3.10 -y
conda activate vllm
pip install vllm
On the ModelScope platform, models are uniquely identified by <MODEL_ORG>/<MODEL_NAME>. For example, for Qwen/Qwen3.5-0.8B, MODEL_ORG is Qwen and MODEL_NAME is Qwen3.5-0.8B.
To deploy Qwen-Family models, use the deployment script scripts/deploy.sh. The usage of the script is as follows:
访问小葱技能站7w4.net,解锁更多实用的AI技能插件。
Usage: [ENV_VARS] deploy.sh <model_name>
Example:
PORT=8001 \
GPU_COUNT=4 \
./deploy.sh Qwen3.5-0.8B
Environment Variables:
ENV_NAME conda environment name (default: vllm)
PORT service port (default: 8000)
GPU_COUNT number of GPUs for tensor parallelism (default: 1)
PROXY proxy address (default: http://{proxyaddress}:{port})
MODEL_BASE_PATH local path to store models (default: /home/work/models)
| Variable | Description | Default |
|---|---|---|
| MODEL_ORG | model organization | Qwen |
| MODEL_NAME | model name | Qwen3.5-0.8B |
| ENV_NAME | conda environment | vllm |
| PORT | model service port | 8000 |
| GPU_COUNT | number of GPUs for tensor parallelism | 1 |
| PROXY | proxy address | http://{proxyaddress}:{port} |
| MODEL_BASE_PATH | local storage path for models | /home/work/models |
Extract required information from the user request: model name (MODEL_NAME), model organization (MODEL_ORG), target server address (TARGET_HOST), deployment user (TARGET_USER), and other necessary parameters.
Copy ./skills/model-deploy/scripts/deploy.sh to the specified path on the target server, e.g., $HOME/wangwei1237.
ssh ${TARGET_USER}@${TARGET_HOST} "cd $HOME/wangwei1237 && PORT=8001 && ./deploy.sh Qwen3.5-0.8B"
curl -X POST http://127.0.0.1:8001/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "你好"}],
"max_tokens": 512
}'
ssh ${TARGET_USER}@${TARGET_HOST} "${CMD}"netstat -tlnp | grep <port>pip install vllm --upgradeexport https_proxy="http://{proxyaddress}:{port}"nvidia-smi, find a suitable GPU index GPU_FAN, set export CUDA_VISIBLE_DEVICES=$GPU_FAN to specify the GPU, then rerun the deployment script.这是一个用于在 GPU 服务器上部署 AI 大模型的工具。文档写得清楚明白,操作步骤一目了然,脚本功能基本完整,错误提示也比较友好。主要问题是支持的模型种类有限(主要是 Qwen 系列),部署后缺少自动检查是否成功的功能,而且需要用户自己配置网络代理等细节。对于有明确部署需求的场景,这个工具可以胜任;但对于新手或需要部署多种模型的用户来说,可能还需要进一步完善。