ops: add isolated ASR deployment and safe GPU test handoff
This commit is contained in:
@@ -0,0 +1,116 @@
|
||||
# Qwen3-ASR isolated preparation package
|
||||
|
||||
This package prepares a **separate** ASR image and immutable model snapshot. It does
|
||||
not change, replace, stop, or restart the existing `qwen36-35b` deployment.
|
||||
No GPU process is started by `download_model.py`, `build-image.sh`, or `init-env.py`.
|
||||
|
||||
## Fixed identities
|
||||
|
||||
- Base image ID: `sha256:251eba5cc7c12fed0b75da22a9240e582b1c9e39f6fbc064f86781b963bd814f`
|
||||
- Base: vLLM 0.24.0, torch 2.11.0+cu130, transformers 5.12.1, numpy 2.2.6.
|
||||
- Official model: `Qwen/Qwen3-ASR-1.7B`.
|
||||
- Official ModelScope snapshot: `a04930dbe5419bfee073f7cade734f572689a3a8`.
|
||||
- Weights total 4,698,521,512 bytes; do not use the earlier ~3.5GB estimate.
|
||||
- New service root: `/home/www/qwen-vllm/asr`.
|
||||
- Model root: `/home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B`.
|
||||
- `DOWNLOAD_MANIFEST.json` preserves official SHA256 and size for every payload.
|
||||
- Runtime image is pinned to the built immutable image ID in private `.env`.
|
||||
|
||||
The existing base contains the Qwen3ASR model registry and transcription implementation
|
||||
but lacks audio extras. The independent thin image adds only `av`, `scipy`,
|
||||
`soundfile`, and `soxr` with explicit versions and `--no-deps`; it preserves torch,
|
||||
transformers, numpy and the original image bytes. CPU preflight performs actual WAV
|
||||
encode/decode, resampling, and Whisper feature extraction plus static route checks.
|
||||
This is **not** GPU inference or endpoint acceptance.
|
||||
|
||||
## Prepare without GPU execution
|
||||
|
||||
```sh
|
||||
cd /home/www/qwen-vllm/asr
|
||||
python3 download_model.py /home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B
|
||||
./build-image.sh
|
||||
python3 init-env.py
|
||||
docker compose --env-file .env -f compose.yaml config --quiet
|
||||
```
|
||||
|
||||
`init-env.py` creates a random API key in a new mode-0600 `.env`, refuses to overwrite
|
||||
an existing file, and never prints the key. Do not publish `.env`, full compose
|
||||
config output, container environment, or API Authorization headers.
|
||||
|
||||
## Start only after the project controller has released GPU1
|
||||
|
||||
**The controller must first verify that Qwen is healthy on GPU0 alone, GPU1 has no
|
||||
competing process, and the shared-card guard permits ASR. Preparation is not this gate.**
|
||||
|
||||
```sh
|
||||
cd /home/www/qwen-vllm/asr
|
||||
docker compose --env-file .env -f compose.yaml up -d --no-build --pull never asr
|
||||
```
|
||||
|
||||
Host GPU1 alone is exposed via Docker DeviceIDs `["1"]`; it appears as logical
|
||||
CUDA device 0 inside the one-GPU container. Host port is only `127.0.0.1:8001`.
|
||||
`restart: unless-stopped` supports reboot persistence without reactivating an
|
||||
explicitly stopped service. Initial `gpu_memory_utilization=0.15` is adjustable
|
||||
in `.env`; do not increase without measuring allocation and the shared-card budget.
|
||||
`max-num-seqs=1`, `max-model-len=32768`, eager execution and max batched tokens 4096
|
||||
bound the initial test workload. The model native text context is 65536. The installed
|
||||
vLLM audio encoder maps about 13 audio tokens/second: 745 seconds is about 9685
|
||||
audio tokens before output, so 4096 is not safe for an unsplit 7–12 minute input.
|
||||
A 32768-token BF16 KV cache is approximately 3.5 GiB from the model configuration
|
||||
(28 layers, 8 KV heads, head dimension 128); this is an estimate, not measured VRAM.
|
||||
Keep the 0.15 allocation cap and verify actual startup/long-audio behavior. The
|
||||
OpenAI transcription implementation has internal clip splitting based on the
|
||||
feature extractor's 30-second clip limit, but the application should still segment
|
||||
recordings into at most 5-minute bounded jobs with overlap and ordered merging.
|
||||
An hour-long recording must be split; this package does not claim a one-hour
|
||||
single-request acceptance. Long calls need a separate chunking/quality gate.
|
||||
No forced aligner is loaded; word timestamps are not promised.
|
||||
|
||||
## Verify after controller start
|
||||
|
||||
```sh
|
||||
cd /home/www/qwen-vllm/asr
|
||||
python3 probe.py /absolute/path/to/non-sensitive-speech.wav
|
||||
```
|
||||
|
||||
The probe requires `/health` HTTP 200, rejects unauthenticated `/v1/models` with
|
||||
401, confirms the served model with authentication, and requires nonempty text
|
||||
from an actual multipart `/v1/audio/transcriptions` request. The authenticated
|
||||
probe reads the key locally and does not print it. A public or synthetic fixture is
|
||||
required; never use patient recordings for this gate. Recognition accuracy needs
|
||||
an expected-text comparison/listening in addition to a successful HTTP response.
|
||||
|
||||
## Stop / rollback only the new ASR service
|
||||
|
||||
```sh
|
||||
cd /home/www/qwen-vllm/asr
|
||||
docker compose --env-file .env -f compose.yaml stop asr
|
||||
```
|
||||
|
||||
This releases ASR's GPU allocation while retaining model/image/config artifacts.
|
||||
Do not start ComfyUI on GPU1 until the parent-controlled shared-card guard has
|
||||
confirmed ASR stopped and memory released. ComfyUI is not deployed by this package.
|
||||
|
||||
Official references:
|
||||
- https://github.com/QwenLM/Qwen3-ASR#deployment-with-vllm
|
||||
- https://modelscope.cn/models/Qwen/Qwen3-ASR-1.7B
|
||||
- https://huggingface.co/Qwen/Qwen3-ASR-1.7B
|
||||
|
||||
|
||||
## GPU 1 与未来 ComfyUI 测试的互斥使用
|
||||
|
||||
在 ai 服务器运行:
|
||||
|
||||
```sh
|
||||
python3 /home/www/qwen-vllm/asr/gpu1-mode.py status
|
||||
# 等待 ASR 在途/排队请求结束并停止它,确认 GPU 1 空闲;不会启动或停止 ComfyUI。
|
||||
python3 /home/www/qwen-vllm/asr/gpu1-mode.py test
|
||||
# 测试实例完全退出、GPU 1 空闲后恢复 ASR,并等待健康检查。
|
||||
python3 /home/www/qwen-vllm/asr/gpu1-mode.py asr
|
||||
```
|
||||
|
||||
控制器遇到未知 GPU 1 进程会拒绝操作,不强杀。ComfyUI 测试实例尚未创建;未来应显式绑定 GPU 1,不修改当前 GPU 2/3 上的服务。请通过此控制器切换,而非同时启动两个占卡服务。`unless-stopped` 保留手工停止状态,新服务同时设置 600 秒引擎优雅关闭和 660 秒容器停止宽限。
|
||||
|
||||
Qwen 单卡切换的原始配置、候选配置、合成文本/长文本/图片检查与回滚控制器保存在 ai 的 `/home/www/qwen-vllm/rollouts/20261008-single-gpu/`。需要回退双卡时,先释放 GPU 1,再运行该目录的 `deploy.py rollback`;它拒绝覆盖不认识的配置或驱逐 GPU 1 上的其他进程。此操作会重启 Qwen,需维护窗口。源码回滚和服务器运行状态回滚不是同一件事。
|
||||
|
||||
控制器无 GPU 单元检查:`python3 deployment/followup-audio-asr/test_gpu1_mode.py`。真实启动、转录和卡交接结果请以本次部署证据为准,不把 CPU-only 准备检查当 GPU 运行成功。
|
||||
Reference in New Issue
Block a user