Files
DS4Server/docs/QWEN38_PERFORMANCE.md
2026-09-04 23:27:09 +02:00

7.7 KiB

Qwen3.8 Flash Next performance and parity receipt

Issue #99 optimizes the pinned Youssofal/Qwen3.8-Flash-Next-MTPLX-Bare-Speed text model without changing its artifact precision, sampler, or token behavior. All application and inference work remains Rust; native additions are confined to the existing Objective-C Metal bridge and Qwen .metal kernels. Integrated MTP is the Qwen production default; an explicitly stored user opt-out remains authoritative.

Result

The acceptance B1 median is 60.825779 tok/s MTP. A later B2 confirmation, after adding complete repair telemetry and under the then-current machine load, measured 58.019436 tok/s. Both sequences produced the same 64-token SHA-256 46ce63601c44d80b5ef8297eaa49a636f6680040b916fe30cf7a154c05daef72, with 44 accepted drafts from 60 proposals in 20 verifier passes. The fresh matched MTPLX median was 59.751170 tok/s, putting B1 1.8% above and B2 2.9% below the oracle; both are inside the issue's 5% parity band. Raw per-run data is in docs/qwen38-issue-99.csv.

Plain AR reached a 50.671077 tok/s median with stable token hash 15373689bcb1c667123accfcfbe49736466f28ce7beb89d2cbe5e4120416056d. The published 47 tok/s MTPLX figure is contextual rather than a fresh matched AR baseline: after the valid MTP oracle run, the installed MTPLX package and checked-out source drifted apart. The current CLI rejects this Flash Next pack as unsupported and the source path fails at the missing DecoderLayer.input_layernorm, so no invalid number is presented as a fresh AR comparison.

The published 75.9 tok/s MTP result is from a different MTPLX runtime/workload. With identical pinned artifacts, 4,096 context, prompt, sampling, seed, and 64-token measurement window, the current usable oracle result is 59.751170 tok/s. The remaining gap is therefore not explained by quantization drift or a different output distribution.

Maximum fan control could not be established because neither ThermalForge nor TG Pro is installed; MTPLX reports its fan boost as unavailable. Both compared engines ran under the same no-boost constraint. No system utility was installed or substituted for the missing control.

Fixed workload

  • Hardware: MacBook Pro Mac17,7, Apple M5 Max, 18 CPU cores, 128 GiB unified memory.
  • OS: macOS 26.6.2, build 25G83, arm64.
  • Model revision: 74559cdf34fbfc0b593de72d17e93f37fd4f9ea7.
  • Model manifest contract SHA-256: 6f1172de47fa30b9602e13fc7ad14e578a813bb5b1fab8f0ef320041aac6c19a.
  • Tensor inventory SHA-256: b5731e6febcf865d276a0e7b144da02375f2d7e3129594f271ee9c6c351f4c8f.
  • config.json SHA-256: 826302af9ff3293cfc9975c9a5a96546fa0aa9f0808505adfc30a62fef806965.
  • mtplx_runtime.json SHA-256: 381c7666872f878542494a48a874bb0784c72a955d18cc077324d3c9e74b8cdb.
  • Prompt: Write a concise Rust implementation of binary search and explain its invariants.
  • Prompt tokens: 26; context: 4,096; generated tokens: 64.
  • Sampling: temperature 1.0, top-p 0.95, top-k 20, min-p 0, seed 12,345.
  • One process, three runs: first run cold, following runs warm.

The exact DS4Server commands were:

DS4SERVER_QWEN38_SOURCE=/private/tmp/qwen38-bare-source \
DS4SERVER_QWEN38_MEASURED=64 DS4SERVER_QWEN38_RUNS=3 \
cargo test --release qwen_m5_performance_gate -- --ignored --nocapture

DS4SERVER_QWEN38_SOURCE=/private/tmp/qwen38-bare-source \
DS4SERVER_QWEN38_MTP=1 DS4SERVER_QWEN38_MEASURED=64 \
DS4SERVER_QWEN38_RUNS=3 \
cargo test --release qwen_m5_performance_gate -- --ignored --nocapture

DS4SERVER_QWEN38_SOURCE=/private/tmp/qwen38-bare-source \
cargo test --release qwen_mtp_matches_target_and_restores_its_cache \
  -- --ignored --nocapture

The MTPLX comparison used its normal serve path with the same local Bare Speed directory, generation mode, prompt, context, sampler, seed, and single request stream. A later retry is not a valid benchmark because the local MTPLX source/package compatibility gate fails as described above.

MTPLX logic review

The production path now follows the MTPLX/MLX graph rather than retaining the temporary Qwen experiments used while earlier issues established correctness:

  • PLE uses the artifact's 4-bit, group-32 rows, streams them with pread, keeps a bounded hot cache, and publishes hit, miss, eviction, byte, and read-time counters. The warm MTP runs had 2,080 hits, no misses, and no reads.
  • GDN uses fused input projection, MTPLX-equivalent conv+SiLU+L2 normalization, BF16 rounding points, decay/delta recurrence, gated norm, and direct verifier capture buffers. Rejected rows restore or advance authoritative recurrence rather than inventing state.
  • QSA uses the MTPLX qsa_flash_skip row ownership, online-softmax merge, and native 256-thread/eight-SIMD layout. Long-context state remains linear and admitted through 262,144 tokens without a gather-copy fallback.
  • MoE uses the exact router ownership and merged gate/up projection. The custom verifier MoE experiment and NAX path were removed because they are not the oracle's promoted path.
  • MTP uses the same checkpoint, history, proposal, verifier, rejection-repair, full-accept bonus, and NumPy-compatible PCG64 sampling rules. The exact cache restoration test passes after the final telemetry change.
  • Sparse top-k sampling stays on device for the supported official sampler and uses the existing exact CPU path only when the requested sampler is outside that kernel's contract. Full production logits are not copied to the CPU.
  • Capability fallbacks remain exact implementations, not alternate Qwen behavior. Experiments that changed tokens (normal Metal math, wider Tensor QMM, and recurrence-only repair) were removed rather than left dormant.

The 4-bit/group-64 core, declared 8-bit router/QSA matrices, BF16 recurrent and normalization tensors, 4-bit/group-32 PLE, and published 2/4/8-bit MTP tensors are consumed as stored. No blanket up-conversion or replacement quantization is used. Peak RSS in the final MTP sequence was 52.61 GB, down from the original 102.65 GB path; the 105.44 GB artifact size includes a 32.00 GB file-backed PLE table that is not resident wholesale.

A/B evidence

The optimization work used exact hashes or tensor comparisons before promotion:

Area Before After Decision
Resident memory 102.65 GB RSS 52.61 GB RSS Reused staging and streamed PLE promoted
Router 49.7 ms GPU 18.7 ms GPU, 23.4 ms wall Row-owned reduced router promoted
16K QSA AR 2.6 tok/s 45.7 tok/s MTPLX flash-skip rewrite promoted
Shared verifier gate/up 54.7 tok/s 58.3 tok/s Exact fused projection promoted
Direct GDN verifier capture 58.5/58.7 tok/s 60.8 tok/s Direct capture buffers promoted

The direct-capture B1 acceptance was followed by the B2 confirmation in the CSV; both retained the same token hash. Normal Metal math, Tensor QMM at verifier width, and shortened recurrence repair failed exact-oracle checks and were removed. Custom MoE verify, fused GDN output projection, and async AR were not promoted because the MTPLX evidence did not show a win and DS4Server profiles did not justify overriding that result.

Context and verification

The long-context gate exercised 2,048, 16,384, 65,536, 131,072, and the native 262,144-token maximum. AR remained 43.488 tok/s at the maximum, and live state grew from 58,195,968 bytes at 2K to 7,449,083,904 bytes at 262K without collapse or over-admission. The CSV records every context result and state size.

Final issue verification comprises the exact MTP cache/token oracle, affine QMV/QMM MLX reference comparisons, focused Qwen tests, the long-context state gate, the release performance gate, formatting, Clippy with warnings denied, the macOS application bundle, and the full all-features Rust test suite.