586 lines
35 KiB
Markdown
586 lines
35 KiB
Markdown
# GLM execution and responsiveness follow-up — 2026-09-11
|
||
|
||
## User acceptance and scope
|
||
|
||
- Qwen is confirmed good in normal interactive use.
|
||
- GLM decode is now also confirmed good interactively. GLM prefill remains
|
||
usable, but feels less smooth than Qwen and affects other applications.
|
||
This is not a claim of a complete freeze or a new confirmed beachball.
|
||
- The user authorized resuming the outstanding work and explicitly authorized
|
||
compiling/running antirez/ds4 as a standalone, supervised reference benchmark.
|
||
No DS4 C objects are linked into DS4Server or its application bundle.
|
||
- Qwen's golden master remains MTPLX; DeepSeek/GLM remain antirez/ds4.
|
||
The accepted 2.6% Qwen Summary AR exception is not a general tolerance.
|
||
|
||
## Implemented execution changes
|
||
|
||
Both scalar GLM loops now flush periodically every four completed layers,
|
||
excluding the final layer and SSD expert streaming. Previously they flushed
|
||
only once at layer four. This follows the active indexed DS4 graph, including
|
||
scalar MTP fallback/rejection replay. The reference's dynamic per-layer mapping
|
||
fallback must not be confused with Rust's static non-expert decode mapping:
|
||
`glm_streaming_model_spans` retains non-expert tensors, the configured resident
|
||
expert prefix, and incompatible expert layouts; selected experts are loaded
|
||
through the existing native cache. No new per-layer SSD waits were introduced.
|
||
Low-memory dynamic mapping fallback parity is not established by this patch.
|
||
|
||
GLM 5.3 prefill progress now advances at existing completed GPU drains and after
|
||
the final output evaluation, not after every submitted layer. Chunk selection,
|
||
prefill flush/drain placement, Metal kernels, sampling and power policy are
|
||
unchanged. This corrects progress accounting; it does not by itself fix the
|
||
remaining prefill smoothness issue.
|
||
|
||
Targeted checks passed: periodic/final/SSD decode boundary test; existing
|
||
prefill boundary test; live two-row verifier acceptance, rejection, rewind,
|
||
scalar fallback, recurrent state, unused HC workspace guards and lifetime
|
||
counters. The live test additionally verifies completed-prefill progress points.
|
||
|
||
## Standalone reference and instrumentation
|
||
|
||
`tools/ds4-session-reference.rs` is a separate Rust benchmark driver for the
|
||
unchanged public DS4 engine/session interface. It is deliberately not a Cargo
|
||
target and is never included in the app. Its build script verifies reference
|
||
commit `ec7642cdd9ec81d01ad4b1fd8f8a3d1511533748`, the pinned header hash,
|
||
unchanged tracked engine sources and current reference objects. The arm64 ABI
|
||
layout is checked against Clang's record layout (engine options 280 bytes,
|
||
distributed offset 152, TP offset 216).
|
||
|
||
The public DS4 CLI/API maps `low` to `high`. The driver instead constructs the
|
||
Low system prefix through the public chat API and passes tokens to the original
|
||
session implementation. It also reproduces the UI's separate system-prefix
|
||
prefill (9 tokens for GLM), and retains generated token history in one session.
|
||
`glm53_reference_prompt_tokens_match_shared_runtime` checks every token of all
|
||
four prompt streams, including continued turns, against the production Rust
|
||
tokenizer. The first complete matched-bootstrap AR reference passed this check.
|
||
|
||
Build from the DS4Server checkout, with already-built reference objects:
|
||
|
||
```sh
|
||
bash tools/build-ds4-session-reference.sh /Users/gb/Projects/ds4 /absolute/path/reference
|
||
```
|
||
|
||
Run the resulting binary from the reference checkout under `test-supervisor`:
|
||
|
||
```text
|
||
test-supervisor 114688 30 45 --command /absolute/path/reference MODEL_GGUF glm on README_PATH
|
||
```
|
||
|
||
The driver uses power100, Low, context32768, temperature0.6, top-p0.95,
|
||
top-k0, min-p0, seed42, SSD streaming off and graph-selected GLM chunks.
|
||
Warmup is a separate session (up to32 tokens), followed by README Summary,
|
||
lighthouse Story, and Python `is_prime` in one ongoing chat, to natural EOS.
|
||
There is no total-runtime watchdog. Missing models fail; no downloads occur.
|
||
DeepSeek is supported by the driver's `deepseek` family argument, with the
|
||
installed DSpark support GGUF required for acceleration-on; it has not yet been
|
||
validated by this follow-up's GLM runs.
|
||
|
||
Optional `DS4_REFERENCE_CANARY=/absolute/path/ds4-server` starts the **same native
|
||
probe and monitor implementation** as the UI/harness in a separate process.
|
||
The `gpu-canary` CLI accepts phase labels on stdin and ends on EOF. A readiness
|
||
handshake waits for the first successful probe before model loading. Its
|
||
readiness/stall/sample reports are not model-progress watchdog heartbeats.
|
||
Clean throughput runs leave this variable unset. An external observer and an
|
||
in-process observer must not be treated as identical OS scheduling conditions.
|
||
|
||
`DS4_REFERENCE_IN_PROCESS_CANARY=1` instead enables a native probe thread inside
|
||
the reference process. It links the same `native/metal/ds4_canary.m` used by
|
||
DS4Server, with the same4096-byte blit, separate queue and100ms cadence, rather
|
||
than duplicating a Metal implementation. This bridge contains no model code.
|
||
Its Rust monitor logs per-sample phases/timing but is not the UI event loop.
|
||
The optional external observer can also be enabled simultaneously. Neither is
|
||
enabled for clean reference throughput. `reference --canary-self-test` exercises
|
||
readiness, two phases and clean shutdown without loading a model.
|
||
|
||
The initial external-probe integration test exposed a startup race: a phase
|
||
could end before the executable initialized. It was not worked around by
|
||
loosening the assertion; the adapter now requires a readiness handshake. The
|
||
initial diagnostic without that handshake is retained, not a full-startup proof.
|
||
|
||
## Measurements and limitations
|
||
|
||
All raw receipts, outputs and failed attempts are retained under
|
||
`local-eval-results/glm-scheduling-20260911.7VSaeW/` (ignored, local evidence).
|
||
The before binary is the secured `02db096` implementation, SHA256
|
||
`ab289f067a7a01c22113eec76aa896638d83e392242192e1440d14ed11524d5c`.
|
||
|
||
Fresh complete AR runs (after, then before), same output/reasoning/tokens/EOS:
|
||
|
||
| Turn | Tokens | Before decode t/s | After decode t/s | Before prefill ms | After prefill ms |
|
||
| --- | ---: | ---: | ---: | ---: | ---: |
|
||
| Summary |640|24.467|25.403|7561|6034|
|
||
| Story |1102|23.703|24.135|262|255|
|
||
| Python |200|24.726|25.401|305|302|
|
||
|
||
These are single sequential pairs, not drift-controlled medians. The large
|
||
summary-prefill difference cannot be attributed to a decode-only flush change.
|
||
No hard decode regression was observed; full performance parity is not proven.
|
||
|
||
The matched-bootstrap original DS4 AR reference completed naturally at
|
||
26.618/25.123/25.825 decode t/s, with779/1047/198 output tokens. Its generated
|
||
text differs from Rust despite matching initial prompt tokens/settings; later
|
||
contexts therefore also differ. This is not an exact-output performance pair.
|
||
Reference prefill timers measure session sync; Rust's current GLM `prefill_ms`
|
||
still includes the observed UI phase. Do not silently equate those intervals.
|
||
|
||
The new Rust MTP run retains the previous625/1026/196 completion tokens and
|
||
identical output/reasoning. Its draft acceptance fractions are241/385 (62.6%),
|
||
342/685 (49.9%) and95/102 (93.1%). Python therefore has the expected higher
|
||
acceptance; MTP's benefit is workload-dependent, not uniformly absent.
|
||
|
||
The fresh MTP before/after pair also preserves every output/reasoning token
|
||
and natural completion:
|
||
|
||
| Turn | Tokens | Before decode t/s | After decode t/s | Before prefill ms | After prefill ms |
|
||
| --- | ---: | ---: | ---: | ---: | ---: |
|
||
| Summary |625|19.988|23.154|9303|6190|
|
||
| Story |1026|17.728|19.295|306|272|
|
||
| Python |196|27.843|29.506|347|322|
|
||
|
||
The same sequential-run/drift limitation applies. This establishes no observed
|
||
hard regression, not a controlled causal speedup or reference-parity acceptance.
|
||
|
||
### Canary placement and timestamp attribution
|
||
|
||
The full `after-mtp-dual-canary` run had simultaneous internal/external probes.
|
||
The internal probe recorded867 successful samples, with a prefill maximum
|
||
of490.740ms and decode maximum3.180ms. The external probe was ready before model
|
||
launch and continued until after termination:3004 successful samples, overall
|
||
maximum3.331ms (startup), and1.572ms while labelled `preparing` across the model
|
||
lifetime. That external label is deliberately not turn/phase attribution.
|
||
Both probes stopped cleanly; no sample failed or reached2s. This is diagnostic
|
||
evidence, not a clean throughput run or a compositor-frame test.
|
||
|
||
Thus the previous `completed_ms` cannot be interpreted as a measured systemwide
|
||
GPU blockade. It includes host-side waiting and completion delivery. The optional
|
||
shared native probe now also records commit-to-GPU-start (`gpu_wait_ms`),
|
||
GPU-start-to-end (`gpu_interval_ms`), and GPU-end-to-host-return (`host_return_ms`).
|
||
Metal's GPU timestamps use system mach time; the probe uses `mach_absolute_time`
|
||
and the native timebase for those differences, not `CLOCK_MONOTONIC`. Missing or
|
||
inconsistent timestamps remain null, not zero. The GPU interval includes possible
|
||
GPU scheduling/preemption, not exclusively active blit execution. See Apple's
|
||
[GPUStartTime documentation](https://developer.apple.com/documentation/metal/mtlcommandbuffer/gpustarttime).
|
||
|
||
The model-free Metal integration check verifies phase coverage, valid nonnegative
|
||
intervals and their bounds against wall completion; the synthetic unit check
|
||
retains null timing when unavailable. The existing UI uses the same enhanced
|
||
native probe, but its stats panel still displays the existing wall latency fields.
|
||
Neither model work nor disabled-canary execution invokes the new timestamp work.
|
||
|
||
The first full Rust timestamp run (`after-mtp-timeline`) preserved every MTP
|
||
output/reasoning token and EOS. All902 samples had valid Metal timestamps and
|
||
none failed. Its worst prefill sample was300.003ms:299.892ms before GPU start,
|
||
0.001917ms GPU interval, and0.105958ms after GPU end. The decode maximum was
|
||
3.733ms. This directly rules out delayed host return as the dominant cause of
|
||
that prefill sample; the queued probe waits for GPU execution. It does not show
|
||
that a different application's rendering queue is delayed by the same amount.
|
||
|
||
The subsequent extraction into the shared native object changes no probe work:
|
||
direct `[cb commit]` replaces the wrapper whose model-queue-only hook never
|
||
applied to this separate canary queue. Both native bindings pass their model-free
|
||
checks after extraction. Final current product binary SHA256:
|
||
`50b4b4abbbdc45ff600c1f46d0bec611879249ac8e4d8291d22d656b9c6e9a5d`;
|
||
standalone reference binary:
|
||
`ccc7a8a774cb1c202add6dba60b04dffe3597822b15a34e22c7e4a5574b50adf`;
|
||
shared probe source:
|
||
`dd3abc34088ee27ba0759f01a291b9b714114420295252d63e85fd6f326fddab`.
|
||
|
||
Answer correctness is checked separately from natural termination. Rust AR,
|
||
Rust MTP and reference AR passed their generated assertions plus5011 `is_prime`
|
||
cases (-10 through5000). The preliminary reference MTP output passed its own
|
||
five assertions but failed347 additional cases, first at49: it omits the
|
||
`i + 2` divisor test. This is a failed generated Python answer, not by itself
|
||
evidence of an engine defect. It must not be reported as a successful code
|
||
benchmark merely because EOS was reached. Details are in `python-check.json`.
|
||
|
||
### Reference clean MTP and record integrity
|
||
|
||
The clean `reference-mtp-clean` run (both canaries disabled) completed all turns
|
||
with the same593/872/167 tokens, text, stop tokens and failed Python answer as
|
||
the diagnostic reference run. Its prefill times were5462.306/351.803/434.414ms;
|
||
decode21.483/16.754/26.143t/s. The preceding in-process diagnostic measured
|
||
23.623/19.287/30.440t/s. This spread must not be disguised as a port speedup or
|
||
accepted2% parity: it is one sequential comparison with different probe state,
|
||
not controlled repeated clean medians. Canary-on throughput is not the baseline.
|
||
|
||
The first internal reference run reported788 successful probes, but only787
|
||
were independently parseable: a watchdog resource record interrupted one
|
||
canary JSON record at a pipe-read boundary. That failed record is preserved in
|
||
`reference-mtp-inline.stderr.log`, not silently counted as missing/zero latency.
|
||
The supervisor now forwards complete lines in one locked stream write (with a
|
||
64KiB cap for newline-free output), while watchdog progress still consumes every
|
||
incoming chunk immediately. EOF flushes partial output. A split-record regression
|
||
test and all existing memory/start/continuation/long-run watchdog tests pass.
|
||
The reference diagnostic is repeated as `reference-mtp-inline-records` for a
|
||
fully parseable receipt; the earlier run is retained as the failure evidence.
|
||
|
||
That repeated reference run completed with **771/771 parseable, successful,
|
||
fully timestamped samples** and identical593/872/167 generated tokens/text/EOS.
|
||
Prefill p95/max was233.106/264.292ms (48 samples); decode p95/max was
|
||
0.226/19.338ms (713 samples). The worst prefill probe waited264.195ms before
|
||
GPU start, ran over0.001750ms, and returned to the host0.092083ms after GPU end.
|
||
No probe reached2s. The reference's prefill samples also include its short
|
||
warmup; the worst sample occurred during the measured summary prefill.
|
||
Its diagnostic throughput was23.893/19.177/32.090t/s, not the clean baseline.
|
||
|
||
| Matched native in-process probe | Prefill p95 ms | Prefill max ms | Decode max ms |
|
||
| --- | ---: | ---: | ---: |
|
||
| DS4Server, `after-mtp-timeline` |289.978|300.003|3.733|
|
||
| Original DS4, `reference-mtp-inline-records` |233.106|264.292|19.338|
|
||
|
||
These sequential diagnostics reproduce the same GPU-start-wait phenomenon in
|
||
the golden master. They do not excuse the remaining Rust prefill cost, establish
|
||
statistical latency equivalence, or measure another application's compositor.
|
||
Moving inference to another thread cannot by itself reproduce the independent
|
||
process's scheduling conditions; process isolation is a distinct architectural
|
||
option, not implemented or declared proven as a UI fix here.
|
||
|
||
Verification at this checkpoint: release all-target/all-feature build; release
|
||
all-target/all-feature Clippy with warnings denied; rustfmt and diff checks;
|
||
11 model-eval unit tests; both model-free native probe bindings; all4 supervisor
|
||
tests; earlier live GLM verifier/progress/HC guards and the full AR/MTP chats.
|
||
The updated supervisor fixes measurement transport, not inference scheduling.
|
||
|
||
## Sampling versus model execution — continued investigation
|
||
|
||
The prior follow-up made concrete progress (execution fixes plus a fair native
|
||
in-process latency reference), but did not establish the full three-model,
|
||
AR/speculative2% goal. This continuation addresses the different GLM outputs
|
||
before treating their different ongoing histories as matched performance work.
|
||
|
||
`reference --sampler-fixture` runs the original public `ds4_sample_logits` without
|
||
loading a model or using Metal. The checked-in
|
||
`tests/fixtures/ds4-sampling-ec7642c.json` contains64 cases: four vocabulary sizes,
|
||
eight temperature/top-k/top-p/min-p settings, seeds0/42,32 consecutive tokens
|
||
per case and the final RNG state. The original Rust test failed40 of64 cases.
|
||
The shared DS4/GLM sampler now preserves the first argmax tie and original
|
||
negative sentinel, skips RNG consumption for greedy/all-invalid and the DS4
|
||
full-vocabulary min-p fallback, and preserves seed0 until the original RNG's
|
||
zero-state substitution. Qwen's independent MTPLX sampler is untouched.
|
||
All64 oracle cases and the16 enabled sampling tests pass. Crucially, the positive
|
||
temperature/top-p benchmark cases at seed42 already passed before the fix:
|
||
these edge corrections are not the explanation for the observed GLM chat gap.
|
||
|
||
Optional `DS4_REFERENCE_LOGITS_TRACE` records the first32 summary logit rows
|
||
through the public original session API. It requires AR mode, creates a new
|
||
file rather than overwriting one, and does not change generated tokens or RNG.
|
||
The full `reference-ar-logits` chat retained exactly the779/1047/198 tokens,
|
||
text and stop tokens of `reference-ar-bootstrap`. Its timings are diagnostic,
|
||
not a clean performance baseline. The binary trace contains19,824,640 bytes
|
||
(32 rows of154,880 little-endian floats). Its path is serialized as an OsString
|
||
and decoded losslessly by the replay test.
|
||
|
||
`glm53_reference_logits_replay_separates_sampling_from_execution` first samples
|
||
those original C-produced rows through the production Rust sampler: **all32
|
||
tokens match**. It then opens the installed GLM at Power100/context32768,
|
||
prefills the same9-token bootstrap and exact summary suffix, and advances only
|
||
with reference-selected tokens. Thus histories never diverge during comparison.
|
||
On the Rust-generated rows the test **fails at step17**, choosing906 instead of
|
||
the reference320. Already the first post-prefill row has max absolute difference
|
||
5.722162 and RMS difference0.851167. All32 per-step row errors are retained in
|
||
`logits-replay.stderr.log`; the watched test terminates normally with failure
|
||
status in9s. This is a new, deliberately retained red parity test, not a passed
|
||
live validation or a speed result. No DS4/GLM/Metal/CPU diagnostic override was
|
||
present in the parent environment.
|
||
|
||
The next localization belongs in the model execution path: compare existing
|
||
original DS4 per-layer tensor dumps with the corresponding Rust HC/KDA/DSA/FFN
|
||
stages, starting at the first bootstrap/prefill block. Do not explain this away
|
||
as stochastic output variation or hide it with a lower chunk/power setting.
|
||
No speculative numerical tolerance or new scheduling workaround was applied.
|
||
|
||
## Root cause: GLM 5.2 chunk boundary applied to GLM 5.3
|
||
|
||
The active original indexed GLM 5.3 path deliberately keeps full2048-token
|
||
chunks across both the old2048 indexer threshold and the4096/8192 dense-attention
|
||
threshold. Rust was still applying the GLM 5.2 top-k boundary: after the9-token
|
||
bootstrap it evaluated2039 tokens, whereas DS4 evaluated2048. This changes the
|
||
recurrent prefill computation, despite identical total prompt tokens.
|
||
|
||
The original layer0 bootstrap `attn_out` and `ffn_out` dumps matched Rust
|
||
bit-for-bit. The original position9 dumps contain2048*4096 floats, establishing
|
||
the actual chunk geometry rather than inferring it from configuration.
|
||
Detailed HC dump hooks elsewhere in DS4 belong to an inactive dense path and
|
||
were not used as evidence for the active indexed execution.
|
||
|
||
Rust now retains complete GLM 5.3 chunks and splits only the attention slices
|
||
at the dense/sparse boundary, as DS4 does. This also removes the incorrect
|
||
whole-pair sparse override for a two-row verifier crossing that boundary.
|
||
GLM 5.2 retains its old top-k splitting. Unit checks cover both families and
|
||
the4096/8192 attention transitions. No smaller chunk, delay, or power reduction
|
||
was introduced.
|
||
|
||
After this correction, the same fixed-history replay is green: **all32 full
|
||
154880-value logit rows are bit-identical** to the original trace (max absolute
|
||
and RMS error both0), and all sampled tokens agree. This run had no stage
|
||
instrumentation enabled. Evidence is retained under
|
||
`local-eval-results/glm-stage-20260911.rwQBaJ/mixed-replay.*.log`.
|
||
The earlier red replay remains historical evidence, not the current result.
|
||
The optional Rust stage reader exists only under `cfg(test)` and validates
|
||
tensor geometry before comparing values; it adds no production GPU drains.
|
||
|
||
The live verifier at frontier4095/context32768 passed across the4096 boundary,
|
||
including acceptance, rejection, rewind to either retained frontier, scalar
|
||
fallback and recurrent-state restoration (`mixed-boundary.*.log`,70.64s).
|
||
The32-row replay alone is not a complete performance or output-parity claim.
|
||
The initial source-only note about one-token suffixes was incomplete: the
|
||
shared UI/headless consumer already routes one-token extensions through scalar
|
||
execution. The actual remaining crossover was two/three-token extensions;
|
||
see the subsequent common-consumer correction below.
|
||
|
||
### Complete chats after the chunk correction
|
||
|
||
Fresh clean runs used the same ongoing workload, Power100/Low, native EOS,
|
||
separate warmup and no active canary. Rust executable SHA256:
|
||
`4b23c04325c931854b98c23bd2c98df8a5c2362927aa9b1faed65019d07fd40d`.
|
||
The original reference retained its prior tokens/text/stops exactly.
|
||
|
||
| Mode / turn | Rust tokens | DS4 tokens | Rust decode t/s | DS4 decode t/s | Output + thinking identical |
|
||
| --- | ---: | ---: | ---: | ---: | --- |
|
||
| AR Summary |779|779|24.512|20.955|yes|
|
||
| AR Story |1047|1047|22.956|19.744|yes|
|
||
| AR Python |198|198|23.436|20.650|yes|
|
||
| MTP Summary |593|593|20.235|21.427|yes|
|
||
| MTP Story |905|872|18.039|18.007|no|
|
||
| MTP Python |169|167|29.443|29.511|no|
|
||
|
||
All six Rust turns and six reference turns ended naturally. AR prompt/cached
|
||
counts also match exactly. Receipts: `clean-comparison.json`,
|
||
`mixed-ar-output-check.json`, `mixed-mtp-output-check.json` in the stage evidence
|
||
directory. The AR reference was materially slower than earlier clean runs;
|
||
these sequential pairs are not a controlled speedup or2% acceptance claim.
|
||
MTP Summary is about5.6% slower in Rust in this pair; the later MTP throughput
|
||
numbers do not compare identical histories. Prefill UI-phase and original
|
||
session-sync timers still have different boundaries (raw values in the receipt).
|
||
Control-loop maxima of40–59ms are not GPU canary or compositor measurements.
|
||
|
||
### Second root cause: MTP stop token retained in the ongoing frontier
|
||
|
||
Although MTP Summary text/thinking and593 emitted tokens match, Rust starts
|
||
Story with3229 cached tokens and3249 prompt tokens; DS4 uses3228/3248.
|
||
The shared Rust generation consumer returned on an MTP stop token without
|
||
rewinding the already evaluated block. Both normal and raw original DS4 agent
|
||
consumers call `ds4_session_rewind(block_start + ti)` at that point. The standalone
|
||
reference's stop handling therefore agrees with its real agent, not just an
|
||
arbitrary benchmark convention.
|
||
|
||
The shared UI/headless consumer now calls `rewind_speculative_output`, a thin
|
||
GLM adapter over the existing two-row rollback, to keep exactly
|
||
`prompt_tokens + emitted_tokens` before retaining the chat.
|
||
This restores the saved two-row KDA state and replays the retained row; it does
|
||
not merely truncate IDs or re-render generated text. Invalid frontiers fail
|
||
explicitly. Both sampled and greedy generation use this consumer. Qwen's own
|
||
whole-turn controller is unchanged. Other model-specific speculative stop
|
||
contracts are not claimed validated by this GLM change.
|
||
The live verifier regression now exercises that same consumer rollback path.
|
||
`align_prompt` is intentionally not used: it retains one fewer token to force
|
||
logit recomputation during prompt synchronization, which is a different contract.
|
||
The full post-frontier-fix MTP measurement (`frontier-mtp.*.log`) now matches
|
||
the original for **all three turns**: text, thinking, emitted token count,
|
||
prompt count, cached frontier and natural stop. Emitted counts are593/872/167;
|
||
Story starts at3228 cached/3248 prompt tokens, Python at4120/4145. The executable
|
||
SHA256 is `b964336d64fbb90b3a9ca595a4705eda02e7afe9c39aedb4ea775e0d52fcf20e`.
|
||
`frontier-mtp-output-check.json` has three entries with every equality true;
|
||
the checked `jq -e` assertion requires all three entries and all five properties.
|
||
Decode rates are23.880/19.357/31.643t/s, versus21.427/18.007/29.511 in the directly
|
||
preceding clean original MTP run. This is one sequential pair, not repeated2%
|
||
acceptance. The Python answer is now exactly the reference's previously checked
|
||
incorrect answer (first counterexample49); matching the oracle does not waive
|
||
the independent generated-code quality failure.
|
||
|
||
### Final regression and responsiveness diagnostics
|
||
|
||
The final strict replay passes with bit-equal logits at all32 steps. The two
|
||
original layer0/position9 stage tensors each contain8388608 floats and also
|
||
match bit-for-bit (`final-replay.*.log`,10.21s). The updated live verifier at4095
|
||
passes through the same rollback entrypoint used by the consumer, including
|
||
invalid/unchanged-frontier checks, rejection and both retained rows
|
||
(`final-boundary.*.log`,82.27s). The five enabled GLM unit tests pass.
|
||
|
||
`final-canary` retained identical full MTP output/frontiers. Its in-memory
|
||
summary reports840 samples, no failures, prefill p95/max395.611/483.166ms,
|
||
decode max3.820ms and no sample crossing the configured2s threshold. However,
|
||
strict raw-log parsing found an interleaved canary/resource JSON record: the
|
||
model-eval parent inherited the child's stderr, and both processes serialized
|
||
JSON fragments to that descriptor. This raw file is retained as a **failed
|
||
record-integrity diagnostic**, not silently filtered into a complete sample set.
|
||
|
||
The model-eval supervisor now pipes child stderr and forwards complete lines
|
||
under the parent's shared stderr lock, the same lock used by resource samples.
|
||
Diagnostics do not refresh inference progress deadlines. Reader failures are
|
||
reported on join. This fixes the app harness counterpart of the earlier
|
||
standalone watchdog forwarding issue; it changes measurement transport, not
|
||
GPU scheduling or the UI inference graph.
|
||
|
||
The directly following original DS4 in-process probe run
|
||
(`final-reference-canary`) has841/841 parseable samples, no failures, unchanged
|
||
reference tokens/text/stops, prefill p95/max373.220/388.090ms and decode max4.370ms.
|
||
The worst prefill sample spent387.964ms before GPU start,0.002875ms over its
|
||
GPU interval and0.122ms returning to the host. Thus substantial prefill queue
|
||
waiting still occurs in the original oracle; the larger Rust spike is not
|
||
declared equivalent or explained away.
|
||
|
||
The repeated Rust run after the forwarding correction (`final-canary-records`)
|
||
completed the entire chat in83.744s and preserved all output/frontier fields.
|
||
Every JSON record beginning with `{` in stderr was parsed with `fromjson`
|
||
(no error suppression): **776/776 canary records and82/82 resource records**
|
||
match the independently reported totals. The checked receipt is
|
||
`final-canary-records-check.json`. There are no probe failures or observed2s
|
||
threshold crossings. Prefill p95/max is119.507/247.181ms; decode max1.854ms.
|
||
The worst sample waits247.062ms before GPU start, spans0.001750ms on the GPU,
|
||
and returns after0.115458ms. This lower maximum is not attributed to the
|
||
transport-only fix: the prior483ms Rust and388ms original spikes remain recorded,
|
||
and scheduling/throughput variability still requires repeated paired testing.
|
||
The optional probe remains off by default; no negligible-overhead claim is made.
|
||
|
||
Final source verification: release all-target/all-feature build and Clippy
|
||
with warnings denied; rustfmt/diff checks; five GLM unit tests; eleven
|
||
model-eval unit tests; sixteen sampling tests including the64-case original
|
||
sampler fixture; strict live logits/stage and consumer-rollback boundary tests.
|
||
The final CLI SHA256 is
|
||
`cbe04f8ce8f8d2fcb6c82b97c3d85b7bed561418893621a6a653d344d1aa6d85`.
|
||
The previously good bundle remains unchanged at SHA256
|
||
`ea4d555c2faf0940d9cbcf76d8638ca614a9cb2c6b034e3b2f80aeef86b0b339`.
|
||
|
||
## Common prompt timing and DS4 CPU sampling follow-up
|
||
|
||
Evidence for this continuation is under
|
||
`local-eval-results/glm-paired-20260911.eClfCS/`. The preceding goal turn made
|
||
verified progress (chunk scheduling and stop-token frontier fixes); it did not
|
||
establish the full six-cell performance goal.
|
||
|
||
The shared consumer now uses DS4's GLM5.3 resumed-prefill crossover of2 tokens,
|
||
not the generic4-token threshold. DS4 explicitly documents this choice as
|
||
measured on M5 Max/GB10 (`ds4.c:36784`). One-token continuations were already
|
||
scalar; cold/vision paths and the separate MTPLX whole-turn controller are
|
||
unchanged. The enabled crossover test covers GLM5.3 versus GLM5.2/DeepSeek.
|
||
|
||
DeepSeek/GLM now publish the existing `PromptTiming` at the shared prompt-
|
||
evaluation boundary: after restoration/bootstrap, around actual suffix execution
|
||
including its progress callbacks, before decode/checkpoint storage. Exact cache
|
||
hits report zero evaluated work. Separately unmeasured restore/history components
|
||
are `null`, not fabricated zeros; Qwen continues reporting the same measured
|
||
numeric values through `Some`. The new metric test and existing Qwen progress/
|
||
decode-timer test pass. The ordinary UI-prefill timer remains separately visible.
|
||
|
||
A fresh clean AR pair kept all three outputs/thinking/token counts/frontiers
|
||
identical. Rust's engine-prefill times were5235.183/269.769/320.767ms, original
|
||
DS4 session-sync8160.750/495.857/479.216ms; Rust decode24.811/23.114/23.410t/s
|
||
versus16.633/17.333/19.054. These large sequential-run differences are not a
|
||
controlled speedup or a completed repeat matrix (`baseline-ar-comparison.json`).
|
||
The reference driver now additionally queries and checks actual engine power100
|
||
after load, rather than only recording its requested options.
|
||
|
||
The CPU sampler still differed algorithmically: Rust sorted the full vocabulary
|
||
and drew from renormalized probabilities, while DS4 first tries a512-candidate
|
||
heap and draws from raw retained weights. A CPU-only replay uses the existing32
|
||
full logit rows, one32-draw warmup and16 measured batches (512 draws). The same
|
||
small runner serves the independent original public `ds4_sample_logits` and the
|
||
production Rust sampler. It loads no model and performs no Metal work; both are
|
||
supervised with1GiB memory/start30s/idle30s limits. The original public function
|
||
allocates a scratch buffer per call, unlike its session API, so its microbenchmark
|
||
is not an exact measure of session-sampler overhead.
|
||
|
||
Before alignment Rust took2.645ms/draw versus original0.834ms, with all512 tokens
|
||
equal. The aligned Rust path initially measured0.401ms/draw with the same512
|
||
tokens (`sampler-{before,after}-rust.json`, `sampler-reference.stdout.log`).
|
||
It uses stdlib `BinaryHeap`, DS4's logit/index tie order, bounded-nucleus fallback
|
||
without advancing RNG, original raw cumulative sampling, full-vocabulary/min-p
|
||
fallback and the original expf-verified log-space rejection boundary. Top-k
|
||
retains the original1024 cap. Separate distribution materialization for
|
||
speculative correction and Qwen's MTPLX sampler are untouched.
|
||
|
||
All64 original sampler fixture cases and17 enabled sampling tests pass, as does
|
||
the added missing-mass/near-one fallback, RNG and signed-zero tie check. Release
|
||
all-target/all-feature build and warnings-denied Clippy pass. The new executable
|
||
SHA256 is `ece6aed3601fb402e6dba6ac2e289d6e0c2dc86663600c3d4b1a4cc07e8fb42c`.
|
||
The first full post-sampler AR and MTP pairs both preserve all three outputs,
|
||
thinking, completion/prompt/cached counts and natural stops. The independently
|
||
queried reference engine reports power100. Receipts are
|
||
`sampler-{ar,mtp}-comparison.json`; these are single pairs, not the repeat matrix.
|
||
|
||
| Mode / turn | Rust / original engine-prefill ms | Rust / original decode t/s |
|
||
| --- | ---: | ---: |
|
||
| AR Summary | 5208.845 / 5548.775 | 26.129 / 24.711 |
|
||
| AR Story | 267.263 / 287.268 | 24.553 / 23.728 |
|
||
| AR Python | 319.127 / 341.355 | 24.964 / 24.520 |
|
||
| MTP Summary | 6574.400 / 5427.866 | 23.935 / 23.628 |
|
||
| MTP Story | 273.240 / 278.956 | 19.419 / 19.089 |
|
||
| MTP Python | 315.279 / 360.621 | 33.140 / 31.919 |
|
||
|
||
The Summary MTP prefill regression in this pair remains visible despite the
|
||
slightly faster Rust decode. Reversed-order repetitions are needed to distinguish
|
||
run variability from a repeatable graph cost. AR before/after the sampler keeps
|
||
the entire chat output identical and improves decode by5.310/6.225/6.638% in this
|
||
one sequential comparison (`sampler-ar-before-after.json`); no controlled causal
|
||
end-to-end percentage is inferred from that pair alone.
|
||
|
||
MTP is not universally beneficial in the original either: its Story decode is
|
||
19.089t/s versus23.728 AR, while Python is31.919 versus24.520. Rust's full MTP
|
||
cycle receipts show228/366,289/584 and82/86 accepted drafts respectively
|
||
(62.3%,49.5%,95.3%). The corresponding complete decode-loop time per cycle is
|
||
67.69/76.89/58.60ms. At1.62/1.49/1.94 emitted tokens per cycle, the Python case
|
||
amortizes the extra draft/verification work much better. These are whole-cycle
|
||
averages, not isolated kernel timings: the existing `verifier_ms` includes other
|
||
cycle work and must not be presented as an exclusive verification stage.
|
||
AR and MTP have different natural histories, so their t/s comparison is not a
|
||
matched-token microbenchmark. The previously recorded Python correctness failure
|
||
also remains open even though both implementations produce the same code.
|
||
|
||
### Reversed-order pairs: acceptance still fails
|
||
|
||
Both modes were repeated in original-then-Rust order, serially without builds
|
||
or canary probes. All twelve measured answers in these four processes again
|
||
match text/thinking/counts/cache frontiers and end naturally; all watchdogs
|
||
exit successfully. No slow run was discarded (`repeat2-*-comparison.json`).
|
||
|
||
| Mode / turn | Rust / original engine-prefill ms | Rust / original decode t/s |
|
||
| --- | ---: | ---: |
|
||
| AR Summary | 6738.280 / 5283.041 | 23.569 / 25.535 |
|
||
| AR Story | 309.647 / 281.591 | 21.743 / 24.186 |
|
||
| AR Python | 387.381 / 329.667 | 21.107 / 24.898 |
|
||
| MTP Summary | 9240.178 / 8835.102 | 17.706 / 17.678 |
|
||
| MTP Story | 358.905 / 402.447 | 15.146 / 14.581 |
|
||
| MTP Python | 403.605 / 491.547 | 25.843 / 23.655 |
|
||
|
||
AR decode now misses by7.70/10.10/15.23%; MTP Summary prefill misses by4.38%.
|
||
The subsequent original MTP run is itself much slower than its first run.
|
||
This excludes neither a Rust scheduling difference nor changing device clocks;
|
||
it does preclude a pass based on the favorable first pair or a selected median.
|
||
The required third pair and full six-cell acceptance remain outstanding.
|
||
|
||
Rust AR emits exactly9445/12612/2424 command buffers in both repetitions, with
|
||
the same outputs, but its GPU timestamp-interval sums increase from
|
||
34536/42061/8084ms to39285/47601/9590ms (`ar-drift-comparison.json`). Those sums
|
||
are `GPUEndTime - GPUStartTime` and may include preemption; they are not exclusive
|
||
kernel or hardware-clock measurements. The slowdown is not explained by changed
|
||
token counts or extra command buffers, and is not declared thermal throttling.
|
||
During the sequence, a read-only process snapshot showed only the intended
|
||
reference model process. macOS reported no recorded thermal/performance warning
|
||
and normal VM pressure (1), which does not exclude frequency changes. The
|
||
AGX PerformanceStatistics snapshot exposes utilization but no frequency field.
|
||
Hardware was freshly checked: Apple M5 Max,128GiB,18 logical CPUs.
|
||
|
||
## Remaining acceptance
|
||
|
||
- Compare repeated clean throughput pairs; GLM ongoing histories now match in
|
||
both modes, but sequential run variability does not establish2% performance parity.
|
||
- Localize the remaining GLM prefill cost against original DS4's active indexed
|
||
path, now that in-process GPU-start waiting is observable on both sides.
|
||
The engine-prefill timer is now exposed separately from UI-phase timing;
|
||
use that aligned boundary in the paired comparisons.
|
||
No chunk reduction or extra waits are justified by these measurements alone.
|
||
- Verify the corrected short-extension crossover live where needed, and other
|
||
model-specific speculative stop contracts; the recorded GLM workload does not
|
||
cover every possible interaction. Full chats pass after CPU-sampler alignment;
|
||
repeated timing acceptance remains separate.
|
||
- Validate the remaining SSD expert-streaming cases separately from resident
|
||
scheduling. This is unrelated to replacing DS4 KV checkpoint persistence.
|
||
- Complete the DeepSeek AR/DSpark reference cells and Qwen residual performance
|
||
analysis. Interactive confirmations are not a substitute for the six-cell
|
||
numerical acceptance matrix.
|
||
|
||
No bundle replacement, commit or push has been performed by this follow-up so far.
|
||
All processes have terminated. The subsequent DeepSeek comparison and its
|
||
separate bootstrap/DSpark findings are recorded in
|
||
[DeepSeek follow-up](deepseek-reference-followup-20260911.md).
|