# Apple Metal parity for issue #78 ## GLM 5.3 Flash parity target - Previous DS4 oracle: `8db89fe083ae4d17c9a2428ccd29803d3ae8f577` - Current DS4 oracle: `ec7642cdd9ec81d01ad4b1fd8f8a3d1511533748` - Main artifact: `GLM-5.3-Flash-Q2.gguf`, SHA-256 `e81fd6241c6e55a64e1e14e47a3eab61a173fa8d7e4b5c1d1848827119705b32` - Vision artifact: `GLM-5.3-Flash-Vision-Encoder.gguf`, SHA-256 `ae23e14c6979e889051b2e4a39351abcdafb161e18e606fae4d8c40095a4bf3a` - Machine: MacBook Pro `Mac17,7`, Apple M5 Max, 128 GiB unified memory, resident Q2 weights The applicable commits after the previous oracle are: | DS4 commit(s) | Scope carried into DS4Server | | --- | --- | | `147109a` | GLM 5.3 shape, GGUF binding, KDA/DSA attention, hyperconnections, routed MoE, BF16 tensors, output head, integrated MTP, and Metal kernels | | `4771329`, `9f95d9f` | Vision encoder, image preprocessing, prompt overlays, and compact-prefill correctness | | `b0c31af` | Pooled indexer cache, bounded attention memory, and batched prefill | | `747eafc`, `1df8ff0` | Fixed-image encoder oracle, replay/logit checks, and multimodal quality gates | | `8adb221` | Directional attention and FFN steering at the DS4 residual points | | `639d4eb` | Continued-prefill cache and position behavior | | `6cf658a` | BF16 decode projection schedule and Metal kernel specialization | | `01f3316`, `81c8c39`, `5da3206` | Current Metal compact-cache ABI, Q2/full-checkpoint binding, and GLM streaming compatibility | CUDA- and ROCm-only commits in the range are excluded. The Rust release gates compare three-run medians against this exact DS4 build, require identical greedy token hashes, and fail only when Rust is more than 2% slower than DS4; there is no upper bound when Rust is faster. Text measures an 8192-token frontier plus steady decode, both with and without integrated MTP. Vision compares the fixed encoder embedding (cosine at least 0.995, mean absolute error at most 0.001, maximum error at most 0.06), encoder latency, multimodal prefill, steady decode, image-conditioned logits, and replay. ### GLM 5.3 Flash acceptance results The text fixture is `ds4.c`, rendered as one no-system chat turn (`-sys ''` in DS4), truncated to exactly 8192 tokens, with a 32768-token allocation and 64 greedy output tokens. Same-process three-run medians are used because both DS4 and DS4Server exhibit a load-order/thermal artifact with this 96 GB model. | Capability | Metric | DS4 | DS4Server | Result | | --- | --- | ---: | ---: | --- | | Text | Prefill tok/s | 288.91 | 295.21 | Rust 2.18% faster | | Text | Decode tok/s | 22.58 | 22.78 | Rust 0.89% faster | | Integrated MTP | Decode tok/s | 26.55 | 26.80 | Rust 0.93% faster | | Vision encoder | Fixed-image seconds | 2.6255 | 2.4354 | Rust 7.24% faster | | Vision prompt | Prefill tok/s | 185.81 | 201.51 | Rust 8.45% faster | | Vision prompt | Decode tok/s | 26.40 | 27.21 | Rust 3.09% faster | Text and MTP produced the same token SHA-256, `0b61e613708d8367fc87f5a32fbfd4ba098b40ee920645b1e38381e867131108`. The 1094-token fixed-image prompt produced the same vision token SHA-256, `424fdadc10c0a2bd95b405a3985151e84aa844a209d2a2ca72cd01c39dfa53db`. The Rust vision embedding compared with the DS4 oracle at cosine `0.999958726`, mean absolute error `0.000036702`, and maximum error `0.006153620`. The checkpoint advertises a 1048576-token maximum. On this 128 GiB machine, 32768 is the default manageable context and 50000 is the extended option: the 50000-token run used 3.51 GiB of live runtime tensors in addition to 90.61 GiB of mapped model wrappers and retained the exact output hash. Larger allocations leave too little operating-system and application headroom for a reliable default. ## Frozen comparison - DS4 oracle: `8db89fe083ae4d17c9a2428ccd29803d3ae8f577` - DS4Server pre-issue baseline: `79468c65b68c28e1539acf0a2a14bd42e000e365` - DeepSeek V4 Flash 0731 IQ2XXS/w2Q2K model SHA-256: `ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0` - DSpark support-model SHA-256: `7e319924541db3f7a163ed7e11d7532a70d48228ab59d36cb81e1d4511885360` - Machine: MacBook Pro `Mac17,7`, Apple M5 Max, 128 GiB unified memory, macOS 26.6.2, power 100% - Context and prefill allocation: 4096 tokens; quality, steering, speculative decoding, and DSpark disabled for the throughput gate - SSD mode: 4096 cached experts, 4096 popularity-preloaded experts, ordinary (not cold) streaming - Prompt: `Count from one to two hundred, spelling out every number.` rendered to 16 tokens; 128 greedy non-EOS tokens are measured The longer prompt is deliberate: all 128 generated tokens are reachable before EOS and are identical between DS4 and DS4Server in each mode. A short-response fixture reached EOS early and made a forced post-EOS speed run compare unreachable, mode-dependent states. ## DS4 performance-change matrix The audited range is `efdadd41e20134af4f3381e1ed90e96fe4faef6f..8db89fe083ae4d17c9a2428ccd29803d3ae8f577`. Documentation-only, parser/server-only, CUDA, ROCm, tensor-parallel, model download, and merge commits do not affect the Rust-owned single-session Metal graph. The applicable performance/runtime commits are accounted for below. | DS4 commit(s) | DS4 call site or behavior | DS4Server result | | --- | --- | --- | | `3ce6777`, `8e0f997` | Native streamed-expert cache pinning and safe fallback | Active in the byte-matched native Metal cache; Rust preserves the cache budget and fallback. | | `df04887` | `ds4_gpu_attention_indexed_mixed_batch_heads_tensor` | Active through `encode_layer_with_cache_rows`; the long-context fixture crosses the indexed boundary. | | `74e8f11`, `4591cb1`, `a51e6ec`, `f284165`, `af80694`, `7fb2830`, `0e89a0e`, `769a8ba`, `8d683d6` | DSpark block commit, deterministic scheduling, acceptance replay/direct commit, greedy identity, and exact/opportunistic sampling | Active in `Dspark::propose` and the verifier commit paths; greedy and exact stochastic fixtures cover the frozen support model. | | `96c3ba4`, `d69a017`, `81f8534`, `ac5cb2c`, `48392e3`, `222b2cb`, `532ec8b` | Metal 4 Q4 projection, compact routed-MoE, aligned Q8, indexed-attention, and routed prefill kernels | Active through the existing `encode_batch_layer` Q4/indexed/routed entry points; no second prefill graph was added. | | `24903f3` | `ds4_gpu_head_rms_norm_rope_tail_tensor` | Existing Rust Q-head norm/RoPE fusion is retained and used. | | `4893e0c`, `d14ce35` | Mapped-prefill expert seeding and consistent expert/non-expert maps | Rust seeds with `ds4_gpu_stream_expert_cache_seed_experts_gpu_copy`, falls back to CPU/pread, and restores one cached static decode map with the same non-routed spans. | | `0dd0d36` | Release speed regression gate | Rust has an ignored release gate with retained raw CSV, DS4 reference threshold, baseline threshold, p50/p95, cache metrics, and token hashes. | | `4e401a1`, `0e125e6` | Automatic pre-M5/M5 admission for exact decode fusions | Rust binds both device predicates and mirrors per-generation rollback variables without a UI preference. | | `71d8c2a` | Q/KV pair plus ratio-4/128 compressor projection/store | `try_qkv_pair_compressor_fusion` calls the native combined entry point only for DS4's resident full-decode shapes. | | `b64d3e0` | Emit-path compressor finalizer | `update_compression_pair` defers both updates only when the resident finalizer is eligible, then requires exact success. | | `3826f52`, `ad4d05c` | Exact compressor softmax/product and ratio-4 pool | Native kernels are active; Rust's compressor ABI now passes the upstream `decode_one_token` and `defer_finalize` flags exactly. | | `92d83e6` | HC norm/mix producer clustering | `decode_hc_mix` uses split-norm producer fusion, with the established norm/mix/split sequence as fallback. | | `1be41c8` | Packed exact split-K attention reduction | Active in the vendored native attention kernel reached by Rust's decode attention calls. | | `0e8fb54`, `e6a6ab5` | Router/select, shared gate/up, QKV norm/KV RoPE/store, parallel full FFN, inverse-attention RoPE, and pre-M5 ports | All applicable entry points are called with DS4 shape, tensor, quality, steering, profile, SSD, and device guards. Parallel FFN aborts on partial failure; inverse RoPE falls back when the attention kernel does not consume it. | | `0658dac` | Long-context prefill/decode correctness | The 4096-crossing ignored fixture covers compressed/indexed state and the next decode token. | | `023614e`, `0ad494e`, `42033ee` | Remove stream-512 experiment; make exact paths automatic; verify DFlash pipeline | No stream-512 or opt-in experiment remains. Feature rollbacks retain identical token hashes. | | `24fa85e` | Keep successfully loaded disk KV checkpoints | Existing Rust checkpoint load/round-trip behavior is verified; only invalid checkpoints follow deletion policy. | | `7694112` | Reuse a shorter exact GLM live prefix | `live_prefix_rewind_target` rewinds to `prompt_len - 1`, clears pending speculation, and reevaluates the last prompt token. | | `5186e2c` | Keep Metal decode changes portable | Device checks and unavailable returns preserve non-M5 and unfused execution. | The Metal MXFP4 commits (`725b084`, `1f862bd`, `7bec128`, `bcd1e57`, `f75a065`, and their merge/format work) are intentionally excluded: MXFP4 is issue #79. `ds4_gpu_set_decode_pipeline_fast_lookup` is reset to disabled, matching DS4 because that cache is enabled only for the pre-M5 MXFP4 path. ## Results Three implementations were run sequentially, three times per mode, with a five-second cooldown. Medians are from [`metal-parity-issue-78.csv`](metal-parity-issue-78.csv). | Implementation | Mode | Prefill tok/s | TTFT ms | First decode ms | Stable tok/s | p50 ms | p95 ms | | --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | | DS4Server baseline | resident | 69.63 | 256.07 | 26.37 | 39.53 | 24.88 | 26.61 | | DS4 | resident | 56.49 | 305.94 | 22.71 | 45.60 | 21.91 | 22.32 | | DS4Server issue #78 | resident | 69.76 | 252.01 | 22.66 | 45.60 | 21.94 | 22.33 | | DS4Server baseline | SSD | 6.81 | 2476.94 | 125.30 | 8.63 | 114.93 | 124.28 | | DS4 | SSD | 11.88 | 1393.54 | 46.75 | 26.11 | 35.62 | 46.86 | | DS4Server issue #78 | SSD | 11.49 | 1440.79 | 48.84 | 25.07 | 37.70 | 48.85 | TTFT is the per-run prompt prefill duration plus first decode latency; its median is reported separately from the first decode kernel latency. Resident stable decode improves 15.4% over the baseline and matches DS4. SSD stable decode improves 190.3% over the baseline and is 4.0% below DS4, inside the required 5% boundary. SSD cache accounting is identical to DS4 for the measured token sequence: 4096 entries, 32,958 hits, 4,194 misses, and 29,684,662,272 pread bytes. The resident and SSD token SHA-256 values are, respectively, `ed5461476a82824c3e0440e943a0b2f6b8a8f19dc26a550a1bc255003851ab8d` and `2bc33a04bb88f9f3c56f8bae0124b7bf25fd2f32aee81a1bdaf6e979448da052`; DS4 emitted the same 128-token sequences in the corresponding modes. Representative `/usr/bin/time -l` measurements report: | Implementation | Mode | Maximum RSS bytes | Peak footprint bytes | Process swaps | | --- | --- | ---: | ---: | ---: | | DS4 | resident | 78,790,656 | 4,746,511,440 | 0 | | DS4Server issue #78 | resident | 129,024,000 | 104,071,816 | 0 | | DS4 | SSD | 29,066,280,960 | 29,595,620,664 | 0 | | DS4Server issue #78 | SSD | 29,098,393,600 | 104,350,344 | 0 | The post-run system memory-pressure reading was 92% free. Every measured process reported zero swaps; the post-run host counters were 0 swap-ins and 20 historical swap-outs. ## Reproduction DS4 uses the exact 16 Rust-rendered prompt token IDs with its speed-benchmark format. DS4Server uses the installed-model fixture: ```sh DS4_REFERENCE_TPS=45.60 \ DS4SERVER_BASELINE_TPS=39.526849 \ cargo test --release --all-features flash_0731_m5_decode_performance_gate \ -- --ignored --nocapture --test-threads=1 DS4SERVER_BENCH_SSD=1 \ DS4_REFERENCE_TPS=26.11 \ DS4SERVER_BASELINE_TPS=8.634659 \ cargo test --release --all-features flash_0731_m5_decode_performance_gate \ -- --ignored --nocapture --test-threads=1 ``` The gate fails unless throughput improves over the frozen DS4Server baseline and remains at least 95% of DS4. `DS4SERVER_BENCH_TOKENS=1` prints the prompt and generated sequences for direct DS4 comparison.