# Qwen3-ASR isolated preparation package This package prepares a **separate** ASR image and immutable model snapshot. It does not change, replace, stop, or restart the existing `qwen36-35b` deployment. No GPU process is started by `download_model.py`, `build-image.sh`, or `init-env.py`. ## Fixed identities - Base image ID: `sha256:251eba5cc7c12fed0b75da22a9240e582b1c9e39f6fbc064f86781b963bd814f` - Base: vLLM 0.24.0, torch 2.11.0+cu130, transformers 5.12.1, numpy 2.2.6. - Official model: `Qwen/Qwen3-ASR-1.7B`. - Official ModelScope snapshot: `a04930dbe5419bfee073f7cade734f572689a3a8`. - Weights total 4,698,521,512 bytes; do not use the earlier ~3.5GB estimate. - New service root: `/home/www/qwen-vllm/asr`. - Model root: `/home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B`. - `DOWNLOAD_MANIFEST.json` preserves official SHA256 and size for every payload. - Runtime image is pinned to the built immutable image ID in private `.env`. The existing base contains the Qwen3ASR model registry and transcription implementation but lacks audio extras. The independent thin image adds only `av`, `scipy`, `soundfile`, and `soxr` with explicit versions and `--no-deps`; it preserves torch, transformers, numpy and the original image bytes. CPU preflight performs actual WAV encode/decode, resampling, and Whisper feature extraction plus static route checks. This is **not** GPU inference or endpoint acceptance. ## Prepare without GPU execution ```sh cd /home/www/qwen-vllm/asr python3 download_model.py /home/www/qwen-vllm/asr/models/Qwen3-ASR-1.7B ./build-image.sh python3 init-env.py docker compose --env-file .env -f compose.yaml config --quiet ``` `init-env.py` creates a random API key in a new mode-0600 `.env`, refuses to overwrite an existing file, and never prints the key. Do not publish `.env`, full compose config output, container environment, or API Authorization headers. ## Start only after the project controller has released GPU1 **The controller must first verify that Qwen is healthy on GPU0 alone, GPU1 has no competing process, and the shared-card guard permits ASR. Preparation is not this gate.** ```sh cd /home/www/qwen-vllm/asr docker compose --env-file .env -f compose.yaml up -d --no-build --pull never asr ``` Host GPU1 alone is exposed via Docker DeviceIDs `["1"]`; it appears as logical CUDA device 0 inside the one-GPU container. Host port is only `127.0.0.1:8001`. `restart: unless-stopped` supports reboot persistence without reactivating an explicitly stopped service. Initial `gpu_memory_utilization=0.15` is adjustable in `.env`; do not increase without measuring allocation and the shared-card budget. `max-num-seqs=1`, `max-model-len=32768`, eager execution and max batched tokens 4096 bound the initial test workload. The model native text context is 65536. The installed vLLM audio encoder maps about 13 audio tokens/second: 745 seconds is about 9685 audio tokens before output, so 4096 is not safe for an unsplit 7–12 minute input. A 32768-token BF16 KV cache is approximately 3.5 GiB from the model configuration (28 layers, 8 KV heads, head dimension 128); this is an estimate, not measured VRAM. Keep the 0.15 allocation cap and verify actual startup/long-audio behavior. The OpenAI transcription implementation has internal clip splitting based on the feature extractor's 30-second clip limit, but the application should still segment recordings into at most 5-minute bounded jobs with overlap and ordered merging. An hour-long recording must be split; this package does not claim a one-hour single-request acceptance. Long calls need a separate chunking/quality gate. No forced aligner is loaded; word timestamps are not promised. ## Verify after controller start ```sh cd /home/www/qwen-vllm/asr python3 probe.py /absolute/path/to/non-sensitive-speech.wav ``` The probe requires `/health` HTTP 200, rejects unauthenticated `/v1/models` with 401, confirms the served model with authentication, and requires nonempty text from an actual multipart `/v1/audio/transcriptions` request. The authenticated probe reads the key locally and does not print it. A public or synthetic fixture is required; never use patient recordings for this gate. Recognition accuracy needs an expected-text comparison/listening in addition to a successful HTTP response. ## Stop / rollback only the new ASR service ```sh cd /home/www/qwen-vllm/asr docker compose --env-file .env -f compose.yaml stop asr ``` This releases ASR's GPU allocation while retaining model/image/config artifacts. Do not start ComfyUI on GPU1 until the parent-controlled shared-card guard has confirmed ASR stopped and memory released. ComfyUI is not deployed by this package. Official references: - https://github.com/QwenLM/Qwen3-ASR#deployment-with-vllm - https://modelscope.cn/models/Qwen/Qwen3-ASR-1.7B - https://huggingface.co/Qwen/Qwen3-ASR-1.7B ## GPU 1 与未来 ComfyUI 测试的互斥使用 在 ai 服务器运行: ```sh python3 /home/www/qwen-vllm/asr/gpu1-mode.py status # 等待 ASR 在途/排队请求结束并停止它,确认 GPU 1 空闲;不会启动或停止 ComfyUI。 python3 /home/www/qwen-vllm/asr/gpu1-mode.py test # 测试实例完全退出、GPU 1 空闲后恢复 ASR,并等待健康检查。 python3 /home/www/qwen-vllm/asr/gpu1-mode.py asr ``` 控制器遇到未知 GPU 1 进程会拒绝操作,不强杀。ComfyUI 测试实例尚未创建;未来应显式绑定 GPU 1,不修改当前 GPU 2/3 上的服务。请通过此控制器切换,而非同时启动两个占卡服务。`unless-stopped` 保留手工停止状态,新服务同时设置 600 秒引擎优雅关闭和 660 秒容器停止宽限。 Qwen 单卡切换的原始配置、候选配置、合成文本/长文本/图片检查与回滚控制器保存在 ai 的 `/home/www/qwen-vllm/rollouts/20261008-single-gpu/`。需要回退双卡时,先释放 GPU 1,再运行该目录的 `deploy.py rollback`;它拒绝覆盖不认识的配置或驱逐 GPU 1 上的其他进程。此操作会重启 Qwen,需维护窗口。源码回滚和服务器运行状态回滚不是同一件事。 控制器无 GPU 单元检查:`python3 deployment/followup-audio-asr/test_gpu1_mode.py`。真实启动、转录和卡交接结果请以本次部署证据为准,不把 CPU-only 准备检查当 GPU 运行成功。