Files
zyt/deployment/followup-audio-asr
..

Qwen3-ASR isolated preparation package

This package prepares a separate ASR image and immutable model snapshot. It does not change, replace, stop, or restart the existing qwen36-35b deployment. No GPU process is started by download_model.py, build-image.sh, or init-env.py.

Fixed identities

  • Base image ID: sha256:251eba5cc7c12fed0b75da22a9240e582b1c9e39f6fbc064f86781b963bd814f
  • Base: vLLM 0.24.0, torch 2.11.0+cu130, transformers 5.12.1, numpy 2.2.6.
  • Official model: Qwen/Qwen3-ASR-1.7B.
  • Official ModelScope snapshot: a04930dbe5419bfee073f7cade734f572689a3a8.
  • Weights total 4,698,521,512 bytes; do not use the earlier ~3.5GB estimate.
  • New service root: /home/www/qwen-vllm/asr.
  • Model root: /home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B.
  • DOWNLOAD_MANIFEST.json preserves official SHA256 and size for every payload.
  • Runtime image is pinned to the built immutable image ID in private .env.

The existing base contains the Qwen3ASR model registry and transcription implementation but lacks audio extras. The independent thin image adds only av, scipy, soundfile, and soxr with explicit versions and --no-deps; it preserves torch, transformers, numpy and the original image bytes. CPU preflight performs actual WAV encode/decode, resampling, and Whisper feature extraction plus static route checks. This is not GPU inference or endpoint acceptance.

Prepare without GPU execution

cd /home/www/qwen-vllm/asr
python3 download_model.py /home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B
./build-image.sh
python3 init-env.py
docker compose --env-file .env -f compose.yaml config --quiet

init-env.py creates a random API key in a new mode-0600 .env, refuses to overwrite an existing file, and never prints the key. Do not publish .env, full compose config output, container environment, or API Authorization headers.

Start only after the project controller has released GPU1

The controller must first verify that Qwen is healthy on GPU0 alone, GPU1 has no competing process, and the shared-card guard permits ASR. Preparation is not this gate.

cd /home/www/qwen-vllm/asr
docker compose --env-file .env -f compose.yaml up -d --no-build --pull never asr

Host GPU1 alone is exposed via Docker DeviceIDs ["1"]; it appears as logical CUDA device 0 inside the one-GPU container. Host port is only 127.0.0.1:8001. restart: unless-stopped supports reboot persistence without reactivating an explicitly stopped service. Initial gpu_memory_utilization=0.15 is adjustable in .env; do not increase without measuring allocation and the shared-card budget. max-num-seqs=1, max-model-len=32768, eager execution and max batched tokens 4096 bound the initial test workload. The model native text context is 65536. The installed vLLM audio encoder maps about 13 audio tokens/second: 745 seconds is about 9685 audio tokens before output, so 4096 is not safe for an unsplit 7–12 minute input. A 32768-token BF16 KV cache is approximately 3.5 GiB from the model configuration (28 layers, 8 KV heads, head dimension 128); this is an estimate, not measured VRAM. Keep the 0.15 allocation cap and verify actual startup/long-audio behavior. The OpenAI transcription implementation has internal clip splitting based on the feature extractor's 30-second clip limit, but the application should still segment recordings into at most 5-minute bounded jobs with overlap and ordered merging. An hour-long recording must be split; this package does not claim a one-hour single-request acceptance. Long calls need a separate chunking/quality gate. No forced aligner is loaded; word timestamps are not promised.

Verify after controller start

cd /home/www/qwen-vllm/asr
python3 probe.py /absolute/path/to/non-sensitive-speech.wav

The probe requires /health HTTP 200, rejects unauthenticated /v1/models with 401, confirms the served model with authentication, and requires nonempty text from an actual multipart /v1/audio/transcriptions request. The authenticated probe reads the key locally and does not print it. A public or synthetic fixture is required; never use patient recordings for this gate. Recognition accuracy needs an expected-text comparison/listening in addition to a successful HTTP response.

Stop / rollback only the new ASR service

cd /home/www/qwen-vllm/asr
docker compose --env-file .env -f compose.yaml stop asr

This releases ASR's GPU allocation while retaining model/image/config artifacts. Do not start ComfyUI on GPU1 until the parent-controlled shared-card guard has confirmed ASR stopped and memory released. ComfyUI is not deployed by this package.

Official references:

GPU 1 与未来 ComfyUI 测试的互斥使用

在 ai 服务器运行:

python3 /home/www/qwen-vllm/asr/gpu1-mode.py status
# 等待 ASR 在途/排队请求结束并停止它,确认 GPU 1 空闲;不会启动或停止 ComfyUI。
python3 /home/www/qwen-vllm/asr/gpu1-mode.py test
# 测试实例完全退出、GPU 1 空闲后恢复 ASR,并等待健康检查。
python3 /home/www/qwen-vllm/asr/gpu1-mode.py asr

控制器遇到未知 GPU 1 进程会拒绝操作,不强杀。ComfyUI 测试实例尚未创建;未来应显式绑定 GPU 1,不修改当前 GPU 2/3 上的服务。请通过此控制器切换,而非同时启动两个占卡服务。unless-stopped 保留手工停止状态,新服务同时设置 600 秒引擎优雅关闭和 660 秒容器停止宽限。

Qwen 单卡切换的原始配置、候选配置、合成文本/长文本/图片检查与回滚控制器保存在 ai 的 /home/www/qwen-vllm/rollouts/20261008-single-gpu/。需要回退双卡时,先释放 GPU 1,再运行该目录的 deploy.py rollback;它拒绝覆盖不认识的配置或驱逐 GPU 1 上的其他进程。此操作会重启 Qwen,需维护窗口。源码回滚和服务器运行状态回滚不是同一件事。

控制器无 GPU 单元检查:python3 deployment/followup-audio-asr/test_gpu1_mode.py。真实启动、转录和卡交接结果请以本次部署证据为准,不把 CPU-only 准备检查当 GPU 运行成功。