Align DeepSeek and GLM execution with DS4

This commit is contained in:
Georg Bauer
2026-09-11 17:18:45 +02:00
parent 02db0968ae
commit 48c2f751b4
27 changed files with 4517 additions and 400 deletions
+585
View File
@@ -0,0 +1,585 @@
# GLM execution and responsiveness follow-up — 2026-09-11
## User acceptance and scope
- Qwen is confirmed good in normal interactive use.
- GLM decode is now also confirmed good interactively. GLM prefill remains
usable, but feels less smooth than Qwen and affects other applications.
This is not a claim of a complete freeze or a new confirmed beachball.
- The user authorized resuming the outstanding work and explicitly authorized
compiling/running antirez/ds4 as a standalone, supervised reference benchmark.
No DS4 C objects are linked into DS4Server or its application bundle.
- Qwen's golden master remains MTPLX; DeepSeek/GLM remain antirez/ds4.
The accepted 2.6% Qwen Summary AR exception is not a general tolerance.
## Implemented execution changes
Both scalar GLM loops now flush periodically every four completed layers,
excluding the final layer and SSD expert streaming. Previously they flushed
only once at layer four. This follows the active indexed DS4 graph, including
scalar MTP fallback/rejection replay. The reference's dynamic per-layer mapping
fallback must not be confused with Rust's static non-expert decode mapping:
`glm_streaming_model_spans` retains non-expert tensors, the configured resident
expert prefix, and incompatible expert layouts; selected experts are loaded
through the existing native cache. No new per-layer SSD waits were introduced.
Low-memory dynamic mapping fallback parity is not established by this patch.
GLM 5.3 prefill progress now advances at existing completed GPU drains and after
the final output evaluation, not after every submitted layer. Chunk selection,
prefill flush/drain placement, Metal kernels, sampling and power policy are
unchanged. This corrects progress accounting; it does not by itself fix the
remaining prefill smoothness issue.
Targeted checks passed: periodic/final/SSD decode boundary test; existing
prefill boundary test; live two-row verifier acceptance, rejection, rewind,
scalar fallback, recurrent state, unused HC workspace guards and lifetime
counters. The live test additionally verifies completed-prefill progress points.
## Standalone reference and instrumentation
`tools/ds4-session-reference.rs` is a separate Rust benchmark driver for the
unchanged public DS4 engine/session interface. It is deliberately not a Cargo
target and is never included in the app. Its build script verifies reference
commit `ec7642cdd9ec81d01ad4b1fd8f8a3d1511533748`, the pinned header hash,
unchanged tracked engine sources and current reference objects. The arm64 ABI
layout is checked against Clang's record layout (engine options 280 bytes,
distributed offset 152, TP offset 216).
The public DS4 CLI/API maps `low` to `high`. The driver instead constructs the
Low system prefix through the public chat API and passes tokens to the original
session implementation. It also reproduces the UI's separate system-prefix
prefill (9 tokens for GLM), and retains generated token history in one session.
`glm53_reference_prompt_tokens_match_shared_runtime` checks every token of all
four prompt streams, including continued turns, against the production Rust
tokenizer. The first complete matched-bootstrap AR reference passed this check.
Build from the DS4Server checkout, with already-built reference objects:
```sh
bash tools/build-ds4-session-reference.sh /Users/gb/Projects/ds4 /absolute/path/reference
```
Run the resulting binary from the reference checkout under `test-supervisor`:
```text
test-supervisor 114688 30 45 --command /absolute/path/reference MODEL_GGUF glm on README_PATH
```
The driver uses power100, Low, context32768, temperature0.6, top-p0.95,
top-k0, min-p0, seed42, SSD streaming off and graph-selected GLM chunks.
Warmup is a separate session (up to32 tokens), followed by README Summary,
lighthouse Story, and Python `is_prime` in one ongoing chat, to natural EOS.
There is no total-runtime watchdog. Missing models fail; no downloads occur.
DeepSeek is supported by the driver's `deepseek` family argument, with the
installed DSpark support GGUF required for acceleration-on; it has not yet been
validated by this follow-up's GLM runs.
Optional `DS4_REFERENCE_CANARY=/absolute/path/ds4-server` starts the **same native
probe and monitor implementation** as the UI/harness in a separate process.
The `gpu-canary` CLI accepts phase labels on stdin and ends on EOF. A readiness
handshake waits for the first successful probe before model loading. Its
readiness/stall/sample reports are not model-progress watchdog heartbeats.
Clean throughput runs leave this variable unset. An external observer and an
in-process observer must not be treated as identical OS scheduling conditions.
`DS4_REFERENCE_IN_PROCESS_CANARY=1` instead enables a native probe thread inside
the reference process. It links the same `native/metal/ds4_canary.m` used by
DS4Server, with the same4096-byte blit, separate queue and100ms cadence, rather
than duplicating a Metal implementation. This bridge contains no model code.
Its Rust monitor logs per-sample phases/timing but is not the UI event loop.
The optional external observer can also be enabled simultaneously. Neither is
enabled for clean reference throughput. `reference --canary-self-test` exercises
readiness, two phases and clean shutdown without loading a model.
The initial external-probe integration test exposed a startup race: a phase
could end before the executable initialized. It was not worked around by
loosening the assertion; the adapter now requires a readiness handshake. The
initial diagnostic without that handshake is retained, not a full-startup proof.
## Measurements and limitations
All raw receipts, outputs and failed attempts are retained under
`local-eval-results/glm-scheduling-20260911.7VSaeW/` (ignored, local evidence).
The before binary is the secured `02db096` implementation, SHA256
`ab289f067a7a01c22113eec76aa896638d83e392242192e1440d14ed11524d5c`.
Fresh complete AR runs (after, then before), same output/reasoning/tokens/EOS:
| Turn | Tokens | Before decode t/s | After decode t/s | Before prefill ms | After prefill ms |
| --- | ---: | ---: | ---: | ---: | ---: |
| Summary |640|24.467|25.403|7561|6034|
| Story |1102|23.703|24.135|262|255|
| Python |200|24.726|25.401|305|302|
These are single sequential pairs, not drift-controlled medians. The large
summary-prefill difference cannot be attributed to a decode-only flush change.
No hard decode regression was observed; full performance parity is not proven.
The matched-bootstrap original DS4 AR reference completed naturally at
26.618/25.123/25.825 decode t/s, with779/1047/198 output tokens. Its generated
text differs from Rust despite matching initial prompt tokens/settings; later
contexts therefore also differ. This is not an exact-output performance pair.
Reference prefill timers measure session sync; Rust's current GLM `prefill_ms`
still includes the observed UI phase. Do not silently equate those intervals.
The new Rust MTP run retains the previous625/1026/196 completion tokens and
identical output/reasoning. Its draft acceptance fractions are241/385 (62.6%),
342/685 (49.9%) and95/102 (93.1%). Python therefore has the expected higher
acceptance; MTP's benefit is workload-dependent, not uniformly absent.
The fresh MTP before/after pair also preserves every output/reasoning token
and natural completion:
| Turn | Tokens | Before decode t/s | After decode t/s | Before prefill ms | After prefill ms |
| --- | ---: | ---: | ---: | ---: | ---: |
| Summary |625|19.988|23.154|9303|6190|
| Story |1026|17.728|19.295|306|272|
| Python |196|27.843|29.506|347|322|
The same sequential-run/drift limitation applies. This establishes no observed
hard regression, not a controlled causal speedup or reference-parity acceptance.
### Canary placement and timestamp attribution
The full `after-mtp-dual-canary` run had simultaneous internal/external probes.
The internal probe recorded867 successful samples, with a prefill maximum
of490.740ms and decode maximum3.180ms. The external probe was ready before model
launch and continued until after termination:3004 successful samples, overall
maximum3.331ms (startup), and1.572ms while labelled `preparing` across the model
lifetime. That external label is deliberately not turn/phase attribution.
Both probes stopped cleanly; no sample failed or reached2s. This is diagnostic
evidence, not a clean throughput run or a compositor-frame test.
Thus the previous `completed_ms` cannot be interpreted as a measured systemwide
GPU blockade. It includes host-side waiting and completion delivery. The optional
shared native probe now also records commit-to-GPU-start (`gpu_wait_ms`),
GPU-start-to-end (`gpu_interval_ms`), and GPU-end-to-host-return (`host_return_ms`).
Metal's GPU timestamps use system mach time; the probe uses `mach_absolute_time`
and the native timebase for those differences, not `CLOCK_MONOTONIC`. Missing or
inconsistent timestamps remain null, not zero. The GPU interval includes possible
GPU scheduling/preemption, not exclusively active blit execution. See Apple's
[GPUStartTime documentation](https://developer.apple.com/documentation/metal/mtlcommandbuffer/gpustarttime).
The model-free Metal integration check verifies phase coverage, valid nonnegative
intervals and their bounds against wall completion; the synthetic unit check
retains null timing when unavailable. The existing UI uses the same enhanced
native probe, but its stats panel still displays the existing wall latency fields.
Neither model work nor disabled-canary execution invokes the new timestamp work.
The first full Rust timestamp run (`after-mtp-timeline`) preserved every MTP
output/reasoning token and EOS. All902 samples had valid Metal timestamps and
none failed. Its worst prefill sample was300.003ms:299.892ms before GPU start,
0.001917ms GPU interval, and0.105958ms after GPU end. The decode maximum was
3.733ms. This directly rules out delayed host return as the dominant cause of
that prefill sample; the queued probe waits for GPU execution. It does not show
that a different application's rendering queue is delayed by the same amount.
The subsequent extraction into the shared native object changes no probe work:
direct `[cb commit]` replaces the wrapper whose model-queue-only hook never
applied to this separate canary queue. Both native bindings pass their model-free
checks after extraction. Final current product binary SHA256:
`50b4b4abbbdc45ff600c1f46d0bec611879249ac8e4d8291d22d656b9c6e9a5d`;
standalone reference binary:
`ccc7a8a774cb1c202add6dba60b04dffe3597822b15a34e22c7e4a5574b50adf`;
shared probe source:
`dd3abc34088ee27ba0759f01a291b9b714114420295252d63e85fd6f326fddab`.
Answer correctness is checked separately from natural termination. Rust AR,
Rust MTP and reference AR passed their generated assertions plus5011 `is_prime`
cases (-10 through5000). The preliminary reference MTP output passed its own
five assertions but failed347 additional cases, first at49: it omits the
`i + 2` divisor test. This is a failed generated Python answer, not by itself
evidence of an engine defect. It must not be reported as a successful code
benchmark merely because EOS was reached. Details are in `python-check.json`.
### Reference clean MTP and record integrity
The clean `reference-mtp-clean` run (both canaries disabled) completed all turns
with the same593/872/167 tokens, text, stop tokens and failed Python answer as
the diagnostic reference run. Its prefill times were5462.306/351.803/434.414ms;
decode21.483/16.754/26.143t/s. The preceding in-process diagnostic measured
23.623/19.287/30.440t/s. This spread must not be disguised as a port speedup or
accepted2% parity: it is one sequential comparison with different probe state,
not controlled repeated clean medians. Canary-on throughput is not the baseline.
The first internal reference run reported788 successful probes, but only787
were independently parseable: a watchdog resource record interrupted one
canary JSON record at a pipe-read boundary. That failed record is preserved in
`reference-mtp-inline.stderr.log`, not silently counted as missing/zero latency.
The supervisor now forwards complete lines in one locked stream write (with a
64KiB cap for newline-free output), while watchdog progress still consumes every
incoming chunk immediately. EOF flushes partial output. A split-record regression
test and all existing memory/start/continuation/long-run watchdog tests pass.
The reference diagnostic is repeated as `reference-mtp-inline-records` for a
fully parseable receipt; the earlier run is retained as the failure evidence.
That repeated reference run completed with **771/771 parseable, successful,
fully timestamped samples** and identical593/872/167 generated tokens/text/EOS.
Prefill p95/max was233.106/264.292ms (48 samples); decode p95/max was
0.226/19.338ms (713 samples). The worst prefill probe waited264.195ms before
GPU start, ran over0.001750ms, and returned to the host0.092083ms after GPU end.
No probe reached2s. The reference's prefill samples also include its short
warmup; the worst sample occurred during the measured summary prefill.
Its diagnostic throughput was23.893/19.177/32.090t/s, not the clean baseline.
| Matched native in-process probe | Prefill p95 ms | Prefill max ms | Decode max ms |
| --- | ---: | ---: | ---: |
| DS4Server, `after-mtp-timeline` |289.978|300.003|3.733|
| Original DS4, `reference-mtp-inline-records` |233.106|264.292|19.338|
These sequential diagnostics reproduce the same GPU-start-wait phenomenon in
the golden master. They do not excuse the remaining Rust prefill cost, establish
statistical latency equivalence, or measure another application's compositor.
Moving inference to another thread cannot by itself reproduce the independent
process's scheduling conditions; process isolation is a distinct architectural
option, not implemented or declared proven as a UI fix here.
Verification at this checkpoint: release all-target/all-feature build; release
all-target/all-feature Clippy with warnings denied; rustfmt and diff checks;
11 model-eval unit tests; both model-free native probe bindings; all4 supervisor
tests; earlier live GLM verifier/progress/HC guards and the full AR/MTP chats.
The updated supervisor fixes measurement transport, not inference scheduling.
## Sampling versus model execution — continued investigation
The prior follow-up made concrete progress (execution fixes plus a fair native
in-process latency reference), but did not establish the full three-model,
AR/speculative2% goal. This continuation addresses the different GLM outputs
before treating their different ongoing histories as matched performance work.
`reference --sampler-fixture` runs the original public `ds4_sample_logits` without
loading a model or using Metal. The checked-in
`tests/fixtures/ds4-sampling-ec7642c.json` contains64 cases: four vocabulary sizes,
eight temperature/top-k/top-p/min-p settings, seeds0/42,32 consecutive tokens
per case and the final RNG state. The original Rust test failed40 of64 cases.
The shared DS4/GLM sampler now preserves the first argmax tie and original
negative sentinel, skips RNG consumption for greedy/all-invalid and the DS4
full-vocabulary min-p fallback, and preserves seed0 until the original RNG's
zero-state substitution. Qwen's independent MTPLX sampler is untouched.
All64 oracle cases and the16 enabled sampling tests pass. Crucially, the positive
temperature/top-p benchmark cases at seed42 already passed before the fix:
these edge corrections are not the explanation for the observed GLM chat gap.
Optional `DS4_REFERENCE_LOGITS_TRACE` records the first32 summary logit rows
through the public original session API. It requires AR mode, creates a new
file rather than overwriting one, and does not change generated tokens or RNG.
The full `reference-ar-logits` chat retained exactly the779/1047/198 tokens,
text and stop tokens of `reference-ar-bootstrap`. Its timings are diagnostic,
not a clean performance baseline. The binary trace contains19,824,640 bytes
(32 rows of154,880 little-endian floats). Its path is serialized as an OsString
and decoded losslessly by the replay test.
`glm53_reference_logits_replay_separates_sampling_from_execution` first samples
those original C-produced rows through the production Rust sampler: **all32
tokens match**. It then opens the installed GLM at Power100/context32768,
prefills the same9-token bootstrap and exact summary suffix, and advances only
with reference-selected tokens. Thus histories never diverge during comparison.
On the Rust-generated rows the test **fails at step17**, choosing906 instead of
the reference320. Already the first post-prefill row has max absolute difference
5.722162 and RMS difference0.851167. All32 per-step row errors are retained in
`logits-replay.stderr.log`; the watched test terminates normally with failure
status in9s. This is a new, deliberately retained red parity test, not a passed
live validation or a speed result. No DS4/GLM/Metal/CPU diagnostic override was
present in the parent environment.
The next localization belongs in the model execution path: compare existing
original DS4 per-layer tensor dumps with the corresponding Rust HC/KDA/DSA/FFN
stages, starting at the first bootstrap/prefill block. Do not explain this away
as stochastic output variation or hide it with a lower chunk/power setting.
No speculative numerical tolerance or new scheduling workaround was applied.
## Root cause: GLM 5.2 chunk boundary applied to GLM 5.3
The active original indexed GLM 5.3 path deliberately keeps full2048-token
chunks across both the old2048 indexer threshold and the4096/8192 dense-attention
threshold. Rust was still applying the GLM 5.2 top-k boundary: after the9-token
bootstrap it evaluated2039 tokens, whereas DS4 evaluated2048. This changes the
recurrent prefill computation, despite identical total prompt tokens.
The original layer0 bootstrap `attn_out` and `ffn_out` dumps matched Rust
bit-for-bit. The original position9 dumps contain2048*4096 floats, establishing
the actual chunk geometry rather than inferring it from configuration.
Detailed HC dump hooks elsewhere in DS4 belong to an inactive dense path and
were not used as evidence for the active indexed execution.
Rust now retains complete GLM 5.3 chunks and splits only the attention slices
at the dense/sparse boundary, as DS4 does. This also removes the incorrect
whole-pair sparse override for a two-row verifier crossing that boundary.
GLM 5.2 retains its old top-k splitting. Unit checks cover both families and
the4096/8192 attention transitions. No smaller chunk, delay, or power reduction
was introduced.
After this correction, the same fixed-history replay is green: **all32 full
154880-value logit rows are bit-identical** to the original trace (max absolute
and RMS error both0), and all sampled tokens agree. This run had no stage
instrumentation enabled. Evidence is retained under
`local-eval-results/glm-stage-20260911.rwQBaJ/mixed-replay.*.log`.
The earlier red replay remains historical evidence, not the current result.
The optional Rust stage reader exists only under `cfg(test)` and validates
tensor geometry before comparing values; it adds no production GPU drains.
The live verifier at frontier4095/context32768 passed across the4096 boundary,
including acceptance, rejection, rewind to either retained frontier, scalar
fallback and recurrent-state restoration (`mixed-boundary.*.log`,70.64s).
The32-row replay alone is not a complete performance or output-parity claim.
The initial source-only note about one-token suffixes was incomplete: the
shared UI/headless consumer already routes one-token extensions through scalar
execution. The actual remaining crossover was two/three-token extensions;
see the subsequent common-consumer correction below.
### Complete chats after the chunk correction
Fresh clean runs used the same ongoing workload, Power100/Low, native EOS,
separate warmup and no active canary. Rust executable SHA256:
`4b23c04325c931854b98c23bd2c98df8a5c2362927aa9b1faed65019d07fd40d`.
The original reference retained its prior tokens/text/stops exactly.
| Mode / turn | Rust tokens | DS4 tokens | Rust decode t/s | DS4 decode t/s | Output + thinking identical |
| --- | ---: | ---: | ---: | ---: | --- |
| AR Summary |779|779|24.512|20.955|yes|
| AR Story |1047|1047|22.956|19.744|yes|
| AR Python |198|198|23.436|20.650|yes|
| MTP Summary |593|593|20.235|21.427|yes|
| MTP Story |905|872|18.039|18.007|no|
| MTP Python |169|167|29.443|29.511|no|
All six Rust turns and six reference turns ended naturally. AR prompt/cached
counts also match exactly. Receipts: `clean-comparison.json`,
`mixed-ar-output-check.json`, `mixed-mtp-output-check.json` in the stage evidence
directory. The AR reference was materially slower than earlier clean runs;
these sequential pairs are not a controlled speedup or2% acceptance claim.
MTP Summary is about5.6% slower in Rust in this pair; the later MTP throughput
numbers do not compare identical histories. Prefill UI-phase and original
session-sync timers still have different boundaries (raw values in the receipt).
Control-loop maxima of4059ms are not GPU canary or compositor measurements.
### Second root cause: MTP stop token retained in the ongoing frontier
Although MTP Summary text/thinking and593 emitted tokens match, Rust starts
Story with3229 cached tokens and3249 prompt tokens; DS4 uses3228/3248.
The shared Rust generation consumer returned on an MTP stop token without
rewinding the already evaluated block. Both normal and raw original DS4 agent
consumers call `ds4_session_rewind(block_start + ti)` at that point. The standalone
reference's stop handling therefore agrees with its real agent, not just an
arbitrary benchmark convention.
The shared UI/headless consumer now calls `rewind_speculative_output`, a thin
GLM adapter over the existing two-row rollback, to keep exactly
`prompt_tokens + emitted_tokens` before retaining the chat.
This restores the saved two-row KDA state and replays the retained row; it does
not merely truncate IDs or re-render generated text. Invalid frontiers fail
explicitly. Both sampled and greedy generation use this consumer. Qwen's own
whole-turn controller is unchanged. Other model-specific speculative stop
contracts are not claimed validated by this GLM change.
The live verifier regression now exercises that same consumer rollback path.
`align_prompt` is intentionally not used: it retains one fewer token to force
logit recomputation during prompt synchronization, which is a different contract.
The full post-frontier-fix MTP measurement (`frontier-mtp.*.log`) now matches
the original for **all three turns**: text, thinking, emitted token count,
prompt count, cached frontier and natural stop. Emitted counts are593/872/167;
Story starts at3228 cached/3248 prompt tokens, Python at4120/4145. The executable
SHA256 is `b964336d64fbb90b3a9ca595a4705eda02e7afe9c39aedb4ea775e0d52fcf20e`.
`frontier-mtp-output-check.json` has three entries with every equality true;
the checked `jq -e` assertion requires all three entries and all five properties.
Decode rates are23.880/19.357/31.643t/s, versus21.427/18.007/29.511 in the directly
preceding clean original MTP run. This is one sequential pair, not repeated2%
acceptance. The Python answer is now exactly the reference's previously checked
incorrect answer (first counterexample49); matching the oracle does not waive
the independent generated-code quality failure.
### Final regression and responsiveness diagnostics
The final strict replay passes with bit-equal logits at all32 steps. The two
original layer0/position9 stage tensors each contain8388608 floats and also
match bit-for-bit (`final-replay.*.log`,10.21s). The updated live verifier at4095
passes through the same rollback entrypoint used by the consumer, including
invalid/unchanged-frontier checks, rejection and both retained rows
(`final-boundary.*.log`,82.27s). The five enabled GLM unit tests pass.
`final-canary` retained identical full MTP output/frontiers. Its in-memory
summary reports840 samples, no failures, prefill p95/max395.611/483.166ms,
decode max3.820ms and no sample crossing the configured2s threshold. However,
strict raw-log parsing found an interleaved canary/resource JSON record: the
model-eval parent inherited the child's stderr, and both processes serialized
JSON fragments to that descriptor. This raw file is retained as a **failed
record-integrity diagnostic**, not silently filtered into a complete sample set.
The model-eval supervisor now pipes child stderr and forwards complete lines
under the parent's shared stderr lock, the same lock used by resource samples.
Diagnostics do not refresh inference progress deadlines. Reader failures are
reported on join. This fixes the app harness counterpart of the earlier
standalone watchdog forwarding issue; it changes measurement transport, not
GPU scheduling or the UI inference graph.
The directly following original DS4 in-process probe run
(`final-reference-canary`) has841/841 parseable samples, no failures, unchanged
reference tokens/text/stops, prefill p95/max373.220/388.090ms and decode max4.370ms.
The worst prefill sample spent387.964ms before GPU start,0.002875ms over its
GPU interval and0.122ms returning to the host. Thus substantial prefill queue
waiting still occurs in the original oracle; the larger Rust spike is not
declared equivalent or explained away.
The repeated Rust run after the forwarding correction (`final-canary-records`)
completed the entire chat in83.744s and preserved all output/frontier fields.
Every JSON record beginning with `{` in stderr was parsed with `fromjson`
(no error suppression): **776/776 canary records and82/82 resource records**
match the independently reported totals. The checked receipt is
`final-canary-records-check.json`. There are no probe failures or observed2s
threshold crossings. Prefill p95/max is119.507/247.181ms; decode max1.854ms.
The worst sample waits247.062ms before GPU start, spans0.001750ms on the GPU,
and returns after0.115458ms. This lower maximum is not attributed to the
transport-only fix: the prior483ms Rust and388ms original spikes remain recorded,
and scheduling/throughput variability still requires repeated paired testing.
The optional probe remains off by default; no negligible-overhead claim is made.
Final source verification: release all-target/all-feature build and Clippy
with warnings denied; rustfmt/diff checks; five GLM unit tests; eleven
model-eval unit tests; sixteen sampling tests including the64-case original
sampler fixture; strict live logits/stage and consumer-rollback boundary tests.
The final CLI SHA256 is
`cbe04f8ce8f8d2fcb6c82b97c3d85b7bed561418893621a6a653d344d1aa6d85`.
The previously good bundle remains unchanged at SHA256
`ea4d555c2faf0940d9cbcf76d8638ca614a9cb2c6b034e3b2f80aeef86b0b339`.
## Common prompt timing and DS4 CPU sampling follow-up
Evidence for this continuation is under
`local-eval-results/glm-paired-20260911.eClfCS/`. The preceding goal turn made
verified progress (chunk scheduling and stop-token frontier fixes); it did not
establish the full six-cell performance goal.
The shared consumer now uses DS4's GLM5.3 resumed-prefill crossover of2 tokens,
not the generic4-token threshold. DS4 explicitly documents this choice as
measured on M5 Max/GB10 (`ds4.c:36784`). One-token continuations were already
scalar; cold/vision paths and the separate MTPLX whole-turn controller are
unchanged. The enabled crossover test covers GLM5.3 versus GLM5.2/DeepSeek.
DeepSeek/GLM now publish the existing `PromptTiming` at the shared prompt-
evaluation boundary: after restoration/bootstrap, around actual suffix execution
including its progress callbacks, before decode/checkpoint storage. Exact cache
hits report zero evaluated work. Separately unmeasured restore/history components
are `null`, not fabricated zeros; Qwen continues reporting the same measured
numeric values through `Some`. The new metric test and existing Qwen progress/
decode-timer test pass. The ordinary UI-prefill timer remains separately visible.
A fresh clean AR pair kept all three outputs/thinking/token counts/frontiers
identical. Rust's engine-prefill times were5235.183/269.769/320.767ms, original
DS4 session-sync8160.750/495.857/479.216ms; Rust decode24.811/23.114/23.410t/s
versus16.633/17.333/19.054. These large sequential-run differences are not a
controlled speedup or a completed repeat matrix (`baseline-ar-comparison.json`).
The reference driver now additionally queries and checks actual engine power100
after load, rather than only recording its requested options.
The CPU sampler still differed algorithmically: Rust sorted the full vocabulary
and drew from renormalized probabilities, while DS4 first tries a512-candidate
heap and draws from raw retained weights. A CPU-only replay uses the existing32
full logit rows, one32-draw warmup and16 measured batches (512 draws). The same
small runner serves the independent original public `ds4_sample_logits` and the
production Rust sampler. It loads no model and performs no Metal work; both are
supervised with1GiB memory/start30s/idle30s limits. The original public function
allocates a scratch buffer per call, unlike its session API, so its microbenchmark
is not an exact measure of session-sampler overhead.
Before alignment Rust took2.645ms/draw versus original0.834ms, with all512 tokens
equal. The aligned Rust path initially measured0.401ms/draw with the same512
tokens (`sampler-{before,after}-rust.json`, `sampler-reference.stdout.log`).
It uses stdlib `BinaryHeap`, DS4's logit/index tie order, bounded-nucleus fallback
without advancing RNG, original raw cumulative sampling, full-vocabulary/min-p
fallback and the original expf-verified log-space rejection boundary. Top-k
retains the original1024 cap. Separate distribution materialization for
speculative correction and Qwen's MTPLX sampler are untouched.
All64 original sampler fixture cases and17 enabled sampling tests pass, as does
the added missing-mass/near-one fallback, RNG and signed-zero tie check. Release
all-target/all-feature build and warnings-denied Clippy pass. The new executable
SHA256 is `ece6aed3601fb402e6dba6ac2e289d6e0c2dc86663600c3d4b1a4cc07e8fb42c`.
The first full post-sampler AR and MTP pairs both preserve all three outputs,
thinking, completion/prompt/cached counts and natural stops. The independently
queried reference engine reports power100. Receipts are
`sampler-{ar,mtp}-comparison.json`; these are single pairs, not the repeat matrix.
| Mode / turn | Rust / original engine-prefill ms | Rust / original decode t/s |
| --- | ---: | ---: |
| AR Summary | 5208.845 / 5548.775 | 26.129 / 24.711 |
| AR Story | 267.263 / 287.268 | 24.553 / 23.728 |
| AR Python | 319.127 / 341.355 | 24.964 / 24.520 |
| MTP Summary | 6574.400 / 5427.866 | 23.935 / 23.628 |
| MTP Story | 273.240 / 278.956 | 19.419 / 19.089 |
| MTP Python | 315.279 / 360.621 | 33.140 / 31.919 |
The Summary MTP prefill regression in this pair remains visible despite the
slightly faster Rust decode. Reversed-order repetitions are needed to distinguish
run variability from a repeatable graph cost. AR before/after the sampler keeps
the entire chat output identical and improves decode by5.310/6.225/6.638% in this
one sequential comparison (`sampler-ar-before-after.json`); no controlled causal
end-to-end percentage is inferred from that pair alone.
MTP is not universally beneficial in the original either: its Story decode is
19.089t/s versus23.728 AR, while Python is31.919 versus24.520. Rust's full MTP
cycle receipts show228/366,289/584 and82/86 accepted drafts respectively
(62.3%,49.5%,95.3%). The corresponding complete decode-loop time per cycle is
67.69/76.89/58.60ms. At1.62/1.49/1.94 emitted tokens per cycle, the Python case
amortizes the extra draft/verification work much better. These are whole-cycle
averages, not isolated kernel timings: the existing `verifier_ms` includes other
cycle work and must not be presented as an exclusive verification stage.
AR and MTP have different natural histories, so their t/s comparison is not a
matched-token microbenchmark. The previously recorded Python correctness failure
also remains open even though both implementations produce the same code.
### Reversed-order pairs: acceptance still fails
Both modes were repeated in original-then-Rust order, serially without builds
or canary probes. All twelve measured answers in these four processes again
match text/thinking/counts/cache frontiers and end naturally; all watchdogs
exit successfully. No slow run was discarded (`repeat2-*-comparison.json`).
| Mode / turn | Rust / original engine-prefill ms | Rust / original decode t/s |
| --- | ---: | ---: |
| AR Summary | 6738.280 / 5283.041 | 23.569 / 25.535 |
| AR Story | 309.647 / 281.591 | 21.743 / 24.186 |
| AR Python | 387.381 / 329.667 | 21.107 / 24.898 |
| MTP Summary | 9240.178 / 8835.102 | 17.706 / 17.678 |
| MTP Story | 358.905 / 402.447 | 15.146 / 14.581 |
| MTP Python | 403.605 / 491.547 | 25.843 / 23.655 |
AR decode now misses by7.70/10.10/15.23%; MTP Summary prefill misses by4.38%.
The subsequent original MTP run is itself much slower than its first run.
This excludes neither a Rust scheduling difference nor changing device clocks;
it does preclude a pass based on the favorable first pair or a selected median.
The required third pair and full six-cell acceptance remain outstanding.
Rust AR emits exactly9445/12612/2424 command buffers in both repetitions, with
the same outputs, but its GPU timestamp-interval sums increase from
34536/42061/8084ms to39285/47601/9590ms (`ar-drift-comparison.json`). Those sums
are `GPUEndTime - GPUStartTime` and may include preemption; they are not exclusive
kernel or hardware-clock measurements. The slowdown is not explained by changed
token counts or extra command buffers, and is not declared thermal throttling.
During the sequence, a read-only process snapshot showed only the intended
reference model process. macOS reported no recorded thermal/performance warning
and normal VM pressure (1), which does not exclude frequency changes. The
AGX PerformanceStatistics snapshot exposes utilization but no frequency field.
Hardware was freshly checked: Apple M5 Max,128GiB,18 logical CPUs.
## Remaining acceptance
- Compare repeated clean throughput pairs; GLM ongoing histories now match in
both modes, but sequential run variability does not establish2% performance parity.
- Localize the remaining GLM prefill cost against original DS4's active indexed
path, now that in-process GPU-start waiting is observable on both sides.
The engine-prefill timer is now exposed separately from UI-phase timing;
use that aligned boundary in the paired comparisons.
No chunk reduction or extra waits are justified by these measurements alone.
- Verify the corrected short-extension crossover live where needed, and other
model-specific speculative stop contracts; the recorded GLM workload does not
cover every possible interaction. Full chats pass after CPU-sampler alignment;
repeated timing acceptance remains separate.
- Validate the remaining SSD expert-streaming cases separately from resident
scheduling. This is unrelated to replacing DS4 KV checkpoint persistence.
- Complete the DeepSeek AR/DSpark reference cells and Qwen residual performance
analysis. Interactive confirmations are not a substitute for the six-cell
numerical acceptance matrix.
No bundle replacement, commit or push has been performed by this follow-up so far.
All processes have terminated. The subsequent DeepSeek comparison and its
separate bootstrap/DSpark findings are recorded in
[DeepSeek follow-up](deepseek-reference-followup-20260911.md).