Qwen3-ASR isolated preparation package
This package prepares a separate ASR image and immutable model snapshot. It does
not change, replace, stop, or restart the existing qwen36-35b deployment.
No GPU process is started by download_model.py, build-image.sh, or init-env.py.
Fixed identities
- Base image ID:
sha256:251eba5cc7c12fed0b75da22a9240e582b1c9e39f6fbc064f86781b963bd814f - Base: vLLM 0.24.0, torch 2.11.0+cu130, transformers 5.12.1, numpy 2.2.6.
- Official model:
Qwen/Qwen3-ASR-1.7B. - Official ModelScope snapshot:
a04930dbe5419bfee073f7cade734f572689a3a8. - Weights total 4,698,521,512 bytes; do not use the earlier ~3.5GB estimate.
- New service root:
/home/www/qwen-vllm/asr. - Model root:
/home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B. DOWNLOAD_MANIFEST.jsonpreserves official SHA256 and size for every payload.- Runtime image is pinned to the built immutable image ID in private
.env.
The existing base contains the Qwen3ASR model registry and transcription implementation
but lacks audio extras. The independent thin image adds only av, scipy,
soundfile, and soxr with explicit versions and --no-deps; it preserves torch,
transformers, numpy and the original image bytes. CPU preflight performs actual WAV
encode/decode, resampling, and Whisper feature extraction plus static route checks.
This is not GPU inference or endpoint acceptance.
Prepare without GPU execution
cd /home/www/qwen-vllm/asr
python3 download_model.py /home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B
./build-image.sh
python3 init-env.py
docker compose --env-file .env -f compose.yaml config --quiet
init-env.py creates a random API key in a new mode-0600 .env, refuses to overwrite
an existing file, and never prints the key. Do not publish .env, full compose
config output, container environment, or API Authorization headers.
Start only after the project controller has released GPU1
The controller must first verify that Qwen is healthy on GPU0 alone, GPU1 has no competing process, and the shared-card guard permits ASR. Preparation is not this gate.
cd /home/www/qwen-vllm/asr
docker compose --env-file .env -f compose.yaml up -d --no-build --pull never asr
Host GPU1 alone is exposed via Docker DeviceIDs ["1"]; it appears as logical
CUDA device 0 inside the one-GPU container. Host port is only 127.0.0.1:8001.
restart: unless-stopped supports reboot persistence without reactivating an
explicitly stopped service. Initial gpu_memory_utilization=0.15 is adjustable
in .env; do not increase without measuring allocation and the shared-card budget.
max-num-seqs=1, max-model-len=32768, eager execution and max batched tokens 4096
bound the initial test workload. The model native text context is 65536. The installed
vLLM audio encoder maps about 13 audio tokens/second: 745 seconds is about 9685
audio tokens before output, so 4096 is not safe for an unsplit 7–12 minute input.
A 32768-token BF16 KV cache is approximately 3.5 GiB from the model configuration
(28 layers, 8 KV heads, head dimension 128); this is an estimate, not measured VRAM.
Keep the 0.15 allocation cap and verify actual startup/long-audio behavior. The
OpenAI transcription implementation has internal clip splitting based on the
feature extractor's 30-second clip limit, but the application should still segment
recordings into at most 5-minute bounded jobs with overlap and ordered merging.
An hour-long recording must be split; this package does not claim a one-hour
single-request acceptance. Long calls need a separate chunking/quality gate.
No forced aligner is loaded; word timestamps are not promised.
Verify after controller start
cd /home/www/qwen-vllm/asr
python3 probe.py /absolute/path/to/non-sensitive-speech.wav
The probe requires /health HTTP 200, rejects unauthenticated /v1/models with
401, confirms the served model with authentication, and requires nonempty text
from an actual multipart /v1/audio/transcriptions request. The authenticated
probe reads the key locally and does not print it. A public or synthetic fixture is
required; never use patient recordings for this gate. Recognition accuracy needs
an expected-text comparison/listening in addition to a successful HTTP response.
Stop / rollback only the new ASR service
cd /home/www/qwen-vllm/asr
docker compose --env-file .env -f compose.yaml stop asr
This releases ASR's GPU allocation while retaining model/image/config artifacts. Do not start ComfyUI on GPU1 until the parent-controlled shared-card guard has confirmed ASR stopped and memory released. ComfyUI is not deployed by this package.
Official references:
- https://github.com/QwenLM/Qwen3-ASR#deployment-with-vllm
- https://modelscope.cn/models/Qwen/Qwen3-ASR-1.7B
- https://huggingface.co/Qwen/Qwen3-ASR-1.7B
GPU 1 与未来 ComfyUI 测试的互斥使用
在 ai 服务器运行:
python3 /home/www/qwen-vllm/asr/gpu1-mode.py status
# 等待 ASR 在途/排队请求结束并停止它,确认 GPU 1 空闲;不会启动或停止 ComfyUI。
python3 /home/www/qwen-vllm/asr/gpu1-mode.py test
# 测试实例完全退出、GPU 1 空闲后恢复 ASR,并等待健康检查。
python3 /home/www/qwen-vllm/asr/gpu1-mode.py asr
控制器遇到未知 GPU 1 进程会拒绝操作,不强杀。ComfyUI 测试实例尚未创建;未来应显式绑定 GPU 1,不修改当前 GPU 2/3 上的服务。请通过此控制器切换,而非同时启动两个占卡服务。unless-stopped 保留手工停止状态,新服务同时设置 600 秒引擎优雅关闭和 660 秒容器停止宽限。
Qwen 单卡切换的原始配置、候选配置、合成文本/长文本/图片检查与回滚控制器保存在 ai 的 /home/www/qwen-vllm/rollouts/20261008-single-gpu/。需要回退双卡时,先释放 GPU 1,再运行该目录的 deploy.py rollback;它拒绝覆盖不认识的配置或驱逐 GPU 1 上的其他进程。此操作会重启 Qwen,需维护窗口。源码回滚和服务器运行状态回滚不是同一件事。
控制器无 GPU 单元检查:python3 deployment/followup-audio-asr/test_gpu1_mode.py。真实启动、转录和卡交接结果请以本次部署证据为准,不把 CPU-only 准备检查当 GPU 运行成功。