ops: add isolated ASR deployment and safe GPU test handoff

This commit is contained in:
2026-10-08 12:42:01 +08:00
parent fea33285e8
commit 8bb4bd07ae
14 changed files with 530 additions and 0 deletions
+116
View File
@@ -0,0 +1,116 @@
# Qwen3-ASR isolated preparation package
This package prepares a **separate** ASR image and immutable model snapshot. It does
not change, replace, stop, or restart the existing `qwen36-35b` deployment.
No GPU process is started by `download_model.py`, `build-image.sh`, or `init-env.py`.
## Fixed identities
- Base image ID: `sha256:251eba5cc7c12fed0b75da22a9240e582b1c9e39f6fbc064f86781b963bd814f`
- Base: vLLM 0.24.0, torch 2.11.0+cu130, transformers 5.12.1, numpy 2.2.6.
- Official model: `Qwen/Qwen3-ASR-1.7B`.
- Official ModelScope snapshot: `a04930dbe5419bfee073f7cade734f572689a3a8`.
- Weights total 4,698,521,512 bytes; do not use the earlier ~3.5GB estimate.
- New service root: `/home/www/qwen-vllm/asr`.
- Model root: `/home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B`.
- `DOWNLOAD_MANIFEST.json` preserves official SHA256 and size for every payload.
- Runtime image is pinned to the built immutable image ID in private `.env`.
The existing base contains the Qwen3ASR model registry and transcription implementation
but lacks audio extras. The independent thin image adds only `av`, `scipy`,
`soundfile`, and `soxr` with explicit versions and `--no-deps`; it preserves torch,
transformers, numpy and the original image bytes. CPU preflight performs actual WAV
encode/decode, resampling, and Whisper feature extraction plus static route checks.
This is **not** GPU inference or endpoint acceptance.
## Prepare without GPU execution
```sh
cd /home/www/qwen-vllm/asr
python3 download_model.py /home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B
./build-image.sh
python3 init-env.py
docker compose --env-file .env -f compose.yaml config --quiet
```
`init-env.py` creates a random API key in a new mode-0600 `.env`, refuses to overwrite
an existing file, and never prints the key. Do not publish `.env`, full compose
config output, container environment, or API Authorization headers.
## Start only after the project controller has released GPU1
**The controller must first verify that Qwen is healthy on GPU0 alone, GPU1 has no
competing process, and the shared-card guard permits ASR. Preparation is not this gate.**
```sh
cd /home/www/qwen-vllm/asr
docker compose --env-file .env -f compose.yaml up -d --no-build --pull never asr
```
Host GPU1 alone is exposed via Docker DeviceIDs `["1"]`; it appears as logical
CUDA device 0 inside the one-GPU container. Host port is only `127.0.0.1:8001`.
`restart: unless-stopped` supports reboot persistence without reactivating an
explicitly stopped service. Initial `gpu_memory_utilization=0.15` is adjustable
in `.env`; do not increase without measuring allocation and the shared-card budget.
`max-num-seqs=1`, `max-model-len=32768`, eager execution and max batched tokens 4096
bound the initial test workload. The model native text context is 65536. The installed
vLLM audio encoder maps about 13 audio tokens/second: 745 seconds is about 9685
audio tokens before output, so 4096 is not safe for an unsplit 7–12 minute input.
A 32768-token BF16 KV cache is approximately 3.5 GiB from the model configuration
(28 layers, 8 KV heads, head dimension 128); this is an estimate, not measured VRAM.
Keep the 0.15 allocation cap and verify actual startup/long-audio behavior. The
OpenAI transcription implementation has internal clip splitting based on the
feature extractor's 30-second clip limit, but the application should still segment
recordings into at most 5-minute bounded jobs with overlap and ordered merging.
An hour-long recording must be split; this package does not claim a one-hour
single-request acceptance. Long calls need a separate chunking/quality gate.
No forced aligner is loaded; word timestamps are not promised.
## Verify after controller start
```sh
cd /home/www/qwen-vllm/asr
python3 probe.py /absolute/path/to/non-sensitive-speech.wav
```
The probe requires `/health` HTTP 200, rejects unauthenticated `/v1/models` with
401, confirms the served model with authentication, and requires nonempty text
from an actual multipart `/v1/audio/transcriptions` request. The authenticated
probe reads the key locally and does not print it. A public or synthetic fixture is
required; never use patient recordings for this gate. Recognition accuracy needs
an expected-text comparison/listening in addition to a successful HTTP response.
## Stop / rollback only the new ASR service
```sh
cd /home/www/qwen-vllm/asr
docker compose --env-file .env -f compose.yaml stop asr
```
This releases ASR's GPU allocation while retaining model/image/config artifacts.
Do not start ComfyUI on GPU1 until the parent-controlled shared-card guard has
confirmed ASR stopped and memory released. ComfyUI is not deployed by this package.
Official references:
- https://github.com/QwenLM/Qwen3-ASR#deployment-with-vllm
- https://modelscope.cn/models/Qwen/Qwen3-ASR-1.7B
- https://huggingface.co/Qwen/Qwen3-ASR-1.7B
## GPU 1 与未来 ComfyUI 测试的互斥使用
在 ai 服务器运行:
```sh
python3 /home/www/qwen-vllm/asr/gpu1-mode.py status
# 等待 ASR 在途/排队请求结束并停止它,确认 GPU 1 空闲;不会启动或停止 ComfyUI。
python3 /home/www/qwen-vllm/asr/gpu1-mode.py test
# 测试实例完全退出、GPU 1 空闲后恢复 ASR,并等待健康检查。
python3 /home/www/qwen-vllm/asr/gpu1-mode.py asr
```
控制器遇到未知 GPU 1 进程会拒绝操作,不强杀。ComfyUI 测试实例尚未创建;未来应显式绑定 GPU 1,不修改当前 GPU 2/3 上的服务。请通过此控制器切换,而非同时启动两个占卡服务。`unless-stopped` 保留手工停止状态,新服务同时设置 600 秒引擎优雅关闭和 660 秒容器停止宽限。
Qwen 单卡切换的原始配置、候选配置、合成文本/长文本/图片检查与回滚控制器保存在 ai 的 `/home/www/qwen-vllm/rollouts/20261008-single-gpu/`。需要回退双卡时,先释放 GPU 1,再运行该目录的 `deploy.py rollback`;它拒绝覆盖不认识的配置或驱逐 GPU 1 上的其他进程。此操作会重启 Qwen,需维护窗口。源码回滚和服务器运行状态回滚不是同一件事。
控制器无 GPU 单元检查:`python3 deployment/followup-audio-asr/test_gpu1_mode.py`。真实启动、转录和卡交接结果请以本次部署证据为准,不把 CPU-only 准备检查当 GPU 运行成功。