# Qwen vision verification — 2026-09-11 Qwen3.8 Flash Next now accepts the existing PNG/JPEG attachments. Its optional vision encoder has a separate Model Manager entry; downloading, validating or deleting it does not alter the text-model artifact set. The existing keep-vision- weights-loaded preference also applies to Qwen. Application preprocessing, loading and inference are Rust, using the existing Metal runtime kernels. The four pinned vision files (897,900,287 bytes total) were downloaded and SHA-256 verified in `~/Library/Application Support/de.rfc1437.ds4server/models/qwen3.8-flash-next`. The artifact revision is `74559cdf34fbfc0b593de72d17e93f37fd4f9ea7` of `Youssofal/Qwen3.8-Flash-Next-MTPLX-Bare-Speed`; the text manifest remains unchanged. The video processor configuration is part of that artifact set; this change implements still images. ## Grounding check Every primary description run used exactly `Describe this image`. The supplied 1024×1024 PNG was copied unchanged to `local-eval-results/qwen-vision-input/image.png`. Its SHA-256 is `b7568d90f4df6180d9af14a824dd553cd995457561d29167354c1ba66b728347`. Only the image bytes and prompt enter the model; neither the source filename nor the neutral filename is included in its text input. Rust/Metal generation succeeded both with MTP and with ordinary autoregressive decoding. Qwen described a golden-tan cartoon llama/alpaca, large eyes, upright ears, an open smiling mouth, mountains, a sunset and a grainy poster texture. These details are visible in the supplied image. The direct cold-session result begins: > This is a stylized, cartoon-style illustration of a llama (or alpaca) shown from the neck up, set against a sunset landscape. Controls used the same prompt: | Input / session | Observed result | | --- | --- | | No image, Rust and MTPLX | Reports no attached image and requests one | | Solid blue image, same dimensions, after the animal image | Describes a uniform blue field; zero cached prompt tokens | | Same image repeated | Same description; all 1,069 prompt tokens reused | | Saved checkpoint, reset, restore, follow-up | Correct animal description; 1,446 cached tokens out of 1,460 | This demonstrates image-dependent descriptions for these inputs, not a general guarantee against hallucinations. ## Oracle and reproducibility The oracle is local MTPLX reference `e652d55` with MLX 0.32.2. Python scripts under `tools/qwen-vision*-reference.py` run only that reference, never the application. The Rust tower matches its exported values exactly at patch embedding, position embedding, rotary positions, blocks 0 and 26, and the final merger. Both the 1024×1024 input (2,621,440 final values) and a small non-square fixture (168,960 final values) had zero differing values. CPU resize/preprocessing also matches three Pillow/MTPLX golden hashes. This is exact encoder agreement; full generated token-sequence parity is not claimed. Local evidence is retained under `local-eval-results/`: - `qwen-vision-rust-mtp.jsonl`, `qwen-vision-rust-ar.jsonl`, `qwen-vision-rust-no-image.jsonl`: complete application runs. - `qwen-vision-lifecycle.jsonl`: cold, repeat, restored and changed-image runs. - `qwen-vision-chat-reference.jsonl`: oracle image/no-image runs. - `qwen-vision-image/`, `qwen-vision-small/`: exported oracle arrays. - `qwen-vision-small-rust.log`: small-fixture exact comparison. Example application invocation (empty YAML config avoids an unrelated system prompt): ```sh target/release/ds4-server model-eval \ --model qwen3.8-flash-next --config /tmp/qwen-vision-config.yaml \ --prompt 'Describe this image' \ --image-file local-eval-results/qwen-vision-input/image.png \ --context 8192 --max-tokens 1024 --reasoning low \ --temperature 0 --top-p 0.95 --seed 1 --acceleration on \ --prefill-chunk 2048 --warmup off --canary on --max-memory-gib 108 ``` The ignored GPU tests `qwen_vision_tower_matches_mtplx_image` and `qwen_vision_chat_checkpoint_preserves_image_identity` are runnable with `DS4_QWEN38_ARTIFACTS` pointing to the model directory and respectively `DS4_QWEN_VISION_REFERENCE` pointing to exported arrays or `DS4_QWEN_VISION_IMAGE` pointing to the neutral input. Run one GPU model process at a time under `test-supervisor` with an appropriate memory limit. Final checks: `cargo fmt --all -- --check`, Clippy with all targets/features and warnings denied, `RUST_TEST_THREADS=1 cargo test --all-features` (315 passed, 204 opt-in tests ignored), and `make bundle` all succeeded. The two encoder comparisons and image lifecycle test were additionally executed explicitly with GPU access. The updated, signed application is `target/release/DS4Server.app`.