4.6 KiB
Qwen vision verification — 2026-09-11
Qwen3.8 Flash Next now accepts the existing PNG/JPEG attachments. Its optional vision encoder has a separate Model Manager entry; downloading, validating or deleting it does not alter the text-model artifact set. The existing keep-vision- weights-loaded preference also applies to Qwen. Application preprocessing, loading and inference are Rust, using the existing Metal runtime kernels.
The four pinned vision files (897,900,287 bytes total) were downloaded and SHA-256
verified in ~/Library/Application Support/de.rfc1437.ds4server/models/qwen3.8-flash-next.
The artifact revision is 74559cdf34fbfc0b593de72d17e93f37fd4f9ea7 of
Youssofal/Qwen3.8-Flash-Next-MTPLX-Bare-Speed; the text manifest remains unchanged.
The video processor configuration is part of that artifact set; this change
implements still images.
Grounding check
Every primary description run used exactly Describe this image. The supplied
1024×1024 PNG was copied unchanged to
local-eval-results/qwen-vision-input/image.png. Its SHA-256 is
b7568d90f4df6180d9af14a824dd553cd995457561d29167354c1ba66b728347.
Only the image bytes and prompt enter the model; neither the source filename nor
the neutral filename is included in its text input.
Rust/Metal generation succeeded both with MTP and with ordinary autoregressive decoding. Qwen described a golden-tan cartoon llama/alpaca, large eyes, upright ears, an open smiling mouth, mountains, a sunset and a grainy poster texture. These details are visible in the supplied image. The direct cold-session result begins:
This is a stylized, cartoon-style illustration of a llama (or alpaca) shown from the neck up, set against a sunset landscape.
Controls used the same prompt:
| Input / session | Observed result |
|---|---|
| No image, Rust and MTPLX | Reports no attached image and requests one |
| Solid blue image, same dimensions, after the animal image | Describes a uniform blue field; zero cached prompt tokens |
| Same image repeated | Same description; all 1,069 prompt tokens reused |
| Saved checkpoint, reset, restore, follow-up | Correct animal description; 1,446 cached tokens out of 1,460 |
This demonstrates image-dependent descriptions for these inputs, not a general guarantee against hallucinations.
Oracle and reproducibility
The oracle is local MTPLX reference e652d55 with MLX 0.32.2. Python scripts under
tools/qwen-vision*-reference.py run only that reference, never the application.
The Rust tower matches its exported values exactly at patch embedding, position
embedding, rotary positions, blocks 0 and 26, and the final merger. Both the
1024×1024 input (2,621,440 final values) and a small non-square fixture (168,960
final values) had zero differing values. CPU resize/preprocessing also matches
three Pillow/MTPLX golden hashes. This is exact encoder agreement; full generated
token-sequence parity is not claimed.
Local evidence is retained under local-eval-results/:
qwen-vision-rust-mtp.jsonl,qwen-vision-rust-ar.jsonl,qwen-vision-rust-no-image.jsonl: complete application runs.qwen-vision-lifecycle.jsonl: cold, repeat, restored and changed-image runs.qwen-vision-chat-reference.jsonl: oracle image/no-image runs.qwen-vision-image/,qwen-vision-small/: exported oracle arrays.qwen-vision-small-rust.log: small-fixture exact comparison.
Example application invocation (empty YAML config avoids an unrelated system prompt):
target/release/ds4-server model-eval \
--model qwen3.8-flash-next --config /tmp/qwen-vision-config.yaml \
--prompt 'Describe this image' \
--image-file local-eval-results/qwen-vision-input/image.png \
--context 8192 --max-tokens 1024 --reasoning low \
--temperature 0 --top-p 0.95 --seed 1 --acceleration on \
--prefill-chunk 2048 --warmup off --canary on --max-memory-gib 108
The ignored GPU tests qwen_vision_tower_matches_mtplx_image and
qwen_vision_chat_checkpoint_preserves_image_identity are runnable with
DS4_QWEN38_ARTIFACTS pointing to the model directory and respectively
DS4_QWEN_VISION_REFERENCE pointing to exported arrays or
DS4_QWEN_VISION_IMAGE pointing to the neutral input. Run one GPU model process
at a time under test-supervisor with an appropriate memory limit.
Final checks: cargo fmt --all -- --check, Clippy with all targets/features and
warnings denied, RUST_TEST_THREADS=1 cargo test --all-features (315 passed,
204 opt-in tests ignored), and make bundle all succeeded. The two encoder
comparisons and image lifecycle test were additionally executed explicitly with
GPU access. The updated, signed application is target/release/DS4Server.app.