520 lines
34 KiB
Markdown
520 lines
34 KiB
Markdown
# DeepSeek standalone-reference follow-up — 2026-09-11
|
||
|
||
The original antirez/ds4 process is used only under the user's explicit
|
||
reference-benchmark authorization. No original C engine is linked into DS4Server.
|
||
Evidence: `local-eval-results/deepseek-paired-20260911.nyUCsL/`.
|
||
|
||
**Measurement qualification (user clarification, 2026-09-11):** The user was
|
||
watching videos in parallel, using part of the GPU. The exact overlap with each
|
||
earlier run is not recorded. The current sequential comparisons are therefore
|
||
not controlled idle-device performance acceptance or causal before/after proof.
|
||
Their outputs and functional checks remain useful; timing receipts are retained,
|
||
not discarded. Do not attribute every gap to video playback or claim thermal
|
||
throttling from this information. Further throughput acceptance runs are deferred
|
||
until competing GPU activity can be controlled; code work and functional tests
|
||
can continue. This does not waive the2% requirement.
|
||
|
||
## Initial pair is not performance-parity evidence
|
||
|
||
Both AR processes finished the full Summary README → lighthouse Story → Python
|
||
`is_prime` conversation after a separate32-token OK warmup. Installed0731 GGUF
|
||
only, power100, Low, temperature0.6/top-p0.95/min-p0/top-k0/seed42, context32768,
|
||
quality/SSD/canary/DSpark off. Startup30s/progress45s,112GiB process-memory guards,
|
||
no total timeout. No builds or other model tests overlapped the GPU runs.
|
||
|
||
Two mismatches invalidate this initial pair (`ar-{reference,rust}.*`):
|
||
|
||
- The standalone driver requested2048 prefill rows, whereas the UI/harness
|
||
requested automatic (0), which the DeepSeek executor resolves to4096 at this
|
||
context. Original DS4's automatic setting also resolves to4096. The reference
|
||
driver now requests0 and records the public session prefill-cap value instead
|
||
of silently assuming the same chunk for all models. GLM's automatic request
|
||
remains unchanged. No product chunk reduction or special harness path.
|
||
- The same first prompt contains2741 tokens in the original but2742 in the
|
||
production runtime. The direct Rust tokenizer and all three original
|
||
continuation fixtures agree exactly; the mismatch occurs later, in the shared
|
||
cached-prompt renderer. A valid system-bootstrap tag with empty chat history
|
||
took the continuation branch and prepended an EOS before the first user turn.
|
||
Both UI and headless generation use this function.
|
||
|
||
## Root-cause correction and regression
|
||
|
||
`render_text_prompt` now requires nonempty history for the cached continuation
|
||
branch. A first user turn uses canonical full rendering, while `align_prompt`
|
||
still reuses its prepared system prefix. Real ongoing histories retain the
|
||
existing append-only behavior. No artificial wait, power change, kernel change
|
||
or KV-persistence rewrite was introduced.
|
||
|
||
The CPU-only `ds4_chat_matches_original_session_tokens` regression loads only
|
||
GGUF metadata/tokenizer and original recorded token IDs. It first checks the
|
||
three-turn tokenizer contract, then invokes the same shared renderer with the
|
||
actual system-bootstrap frontier and tag. Before the fix it fails with:
|
||
|
||
```text
|
||
shared bootstrap renderer differs: 2742/2741 tokens;
|
||
starts [0, 1, 128803, 45585]/[0, 128803, 45585, 260]
|
||
```
|
||
|
||
`token-before.*` is the passing standalone-tokenizer diagnostic;
|
||
`bootstrap-before.*` is the deliberately failing shared-runtime regression.
|
||
Neither file is overwritten or treated as a successful full parity run.
|
||
The test also accepts the saved GLM fixture to cover the other shared consumer.
|
||
|
||
Reproduction after building release tests:
|
||
|
||
```sh
|
||
DS4SERVER_CHAT_REFERENCE=local-eval-results/deepseek-paired-20260911.nyUCsL/chat-token-fixture.json \
|
||
target/release/test-supervisor 2048 30 30 \
|
||
target/release/deps/ds4_server-6141856e5c0fa6df \
|
||
ds4_chat_matches_original_session_tokens
|
||
```
|
||
|
||
The fixed renderer passes both saved DeepSeek and GLM three-turn fixtures
|
||
(`{chat,glm-chat}-token-fixture-fixed.*`,0.04/0.05s). Release all-target/all-feature
|
||
build, warnings-denied Clippy, rustfmt/diff checks,17 enabled sampling tests and
|
||
two enabled tokenizer tests pass. The model-dependent tests remain explicitly
|
||
ignored by default and were invoked with the recorded local fixtures above.
|
||
CLI SHA256: `9fb13c7f9049c3b5059b4fe79b68bf997dbe4e1bfdc4c631c130595cab8c359a`.
|
||
Standalone driver SHA256:
|
||
`ad7a4b94a5918505c58130d336d026e2c5b9640aab717ef185808bebd7536a25`.
|
||
|
||
## Corrected complete AR pair
|
||
|
||
`fixed-ar-{reference,rust}.*` and `fixed-ar-comparison.json` contain the new
|
||
original-then-Rust pair. The original session confirms prefill-cap4096 and
|
||
engine power100. Every text/thinking/token/cache comparison passes; all turns
|
||
end naturally. Counts are826/1444/199, prompts2741/3587/5057 and cached1/3567/5031.
|
||
|
||
| Turn | Rust / original engine-prefill ms | Rust / original decode t/s |
|
||
| --- | ---: | ---: |
|
||
| Summary | 6561.316 / 5425.137 | 36.916 / 37.813 |
|
||
| Story | 286.826 / 290.380 | 32.712 / 34.246 |
|
||
| Python | 334.486 / 369.018 | 32.020 / 32.734 |
|
||
|
||
This is functional agreement for the complete AR workload, **not** performance
|
||
acceptance: all three decode ratios are below98% in this single pair, and the
|
||
Summary prefill is slower. Different natural outputs in the invalid initial
|
||
pair must not be used as before/after throughput evidence. No GUI/canary run was
|
||
performed in this clean series.
|
||
|
||
## DSpark pair: functional comparison fails
|
||
|
||
Both full processes finished naturally and both watchdogs exited successfully.
|
||
The final comparison exits1 deliberately (`dspark-comparison.json`); this is
|
||
not an inference crash. Rust confirms opportunistic sampling, confidence0.6,
|
||
strict/exact off; the original loads the installed three-stage/block5 support
|
||
model and enables direct verifier commits.
|
||
|
||
The first prompt/cache counts agree (2741/1), but its output already diverges:
|
||
647 Rust tokens versus1208 original tokens. Thus later speed ratios compare
|
||
different histories and **cannot establish matched-work performance parity**.
|
||
Counts in Story/Python are1675/195 versus1455/347. The raw rates are retained
|
||
in the comparison file, not promoted to valid performance results.
|
||
|
||
An independent frontier defect is visible even within Rust's own history:
|
||
after2741 prompt +647 output tokens, the next cache is3389 instead of3388;
|
||
after3409 prompt +1675 output tokens, it is5085 instead of5084. A speculatively
|
||
returned EOS remains committed. The shared consumer currently rewinds only
|
||
GLM, while original `ds4_session_rewind` also truncates the DeepSeek logical
|
||
frontier and invalidates DSpark capture. This is separate from the now-fixed
|
||
first-prompt bootstrap EOS and from the first-answer decode divergence.
|
||
|
||
### Next cohesive DSpark work unit
|
||
|
||
1. Locate the first divergent cycle with identical recorded target history:
|
||
compare proposal IDs/confidence, GPU row argmax, accepted prefixes and
|
||
compression/DSpark capture state against original DS4. Cover full acceptance,
|
||
partial acceptance and rejection before declaring the full implementation
|
||
equivalent. Existing target-owned self-tests alone are not an oracle proof.
|
||
2. Match the consumer's EOS frontier and capture invalidation using DS4's
|
||
DeepSeek contract, including an accepted-EOS transition into the next user
|
||
turn. Do not transplant GLM's KDA rollback or merely hide the extra count.
|
||
3. Remove verified extra verifier work as part of that same comparison:
|
||
Rust `eval_batch_inner(collect_tops)` reads/clones every row's full vocabulary
|
||
and computes argmax on the CPU. Original opportunistic verification passes
|
||
`row_logits=NULL`, obtains GPU row tops, then reads only the committed row
|
||
(`ds4.c:69112–69124,69153–69159,69210–69216`). Preserve the separate exact-
|
||
sampling requirements and verify row/stop decisions before timing the change.
|
||
Also verify per-turn counter scope across warmup/resident switching before
|
||
using cycle/acceptance ratios as exclusive stage measurements.
|
||
|
||
The four reviewed AR/DSpark Python answers (both implementations) each pass
|
||
their five generated assertions and5011 independent cases in isolated Python
|
||
with restricted builtins (`python-quality.jsonl`). No model-generated tools
|
||
were invoked. This quality check does not waive DSpark's token/state mismatch.
|
||
|
||
The six-cell performance goal, repeated timing acceptance and actual UI
|
||
responsiveness acceptance remain open. The user's working bundle is unchanged.
|
||
All model/reference/canary processes from this follow-up have terminated.
|
||
|
||
## DSpark cycle/state audit (second work unit)
|
||
|
||
Evidence: `local-eval-results/ds4-cycle-audit-20260911.vfsUc8/`. The original
|
||
source and installed artifacts remain the same. `DS4_SPEC_CYCLE_TRACE=1` is an
|
||
opt-in diagnostic in the shared production consumer and standalone reference
|
||
driver; it records the sampled first token, returned IDs, generated count and
|
||
committed position. It is off by default. Original `DS4_DSPARK_SPEC_LOG=1`
|
||
adds draft lengths, full/partial acceptance and scheduler decisions without
|
||
changing the graph. These traced runs are **not clean throughput acceptance**.
|
||
|
||
The initial complete traced pair reproduces the failure. In the first Summary
|
||
cycle, both targets sample671; Rust proposes12275 and emits only671, while DS4
|
||
proposes3967 and emits `[671,3967]`. The original warmup first differs in the
|
||
fifth returned cycle, after earlier partial accepts. This is not evidence of
|
||
a context-size or power-setting cause.
|
||
|
||
### Consolidated findings and implementation tasks
|
||
|
||
This table scopes the inspected path to the installed0731 DSpark support,
|
||
resident M5 Max execution and default opportunistic sampling. It does not claim
|
||
completion for arbitrary SSD/strict/exact modes or the six-cell performance goal.
|
||
Original line references below are for pinned `ds4.c` at `ec7642c`.
|
||
|
||
| Area | Original behavior / Rust discrepancy | Current action |
|
||
| --- | --- | --- |
|
||
| Initial support KV | `32627–32821`: target HC expansion, per-stage HC mixing and attention normalization precede KV projection. Rust projected stage-0 normalized hidden directly to KV. | Implemented the missing operations using existing Metal calls and batch scratch. First Summary proposal now matches. |
|
||
| Single accepted draft | `36205` and `69112`: even one draft uses the target batch verifier. Rust used scalar decode, with different arithmetic and no captured suffix for the next support cache. | Removed the resident single-draft scalar branch. First two Summary cycles then match. |
|
||
| Verifier capture | `27951–28053`: capture contains the preceding seed row plus all verified target rows. Rust retained only the suffix. | Added seed-row capture with the original slot stride and sufficient workspace. |
|
||
| Cache timing and gaps | `66665–66732`: seed from the captured batch only when its end equals the next proposal position; otherwise crop/reset the absolute cache window. Ring maintenance does not create or bridge a window (`33329`, `27712`). Rust eagerly seeded each batch, extended windows after failed proposals and tracked only modulo positions. | Added deferred seed ownership and absolute window start; no false append after rejected proposals. Replaced batch workspaces are released before allocation of the next batch. |
|
||
| Partial acceptance | `69221–69240`: ordinary partial commits invalidate row and batch captures, preserve compressor-prefix state, and do not claim support KV rows. Rust retained the captures and extended support counters; its old seed scratch also overwrote captured hidden rows. | Invalidate the deferred capture, preserve the original cache frontier, and stop overwriting hidden capture storage. Three and then32 recorded Summary cycles pass. |
|
||
| Scheduler | `53500`, `65215`: reset per request. `53666`: a window pause replaces, rather than maximizes with, the cold no-draft pause. | Reset in shared prompt alignment; regression covers7→4 and clearing prior-request counters. |
|
||
| EOS consumer | `73690`: logical DeepSeek rewind plus capture invalidation, not GLM KDA rollback. | Shared consumer now invokes the model-specific DeepSeek rewind; invalid forward rewind is rejected before mutation. Full next-turn verification still required. |
|
||
| Q8 confidence arithmetic | `34350` calls `matvec_any`, which dispatches type8 to activation-quantized Q8 matvec (`8144–8170`). The installed confidence tensor is Q8_0 `[4352,1]`; Rust instead used the dequantized-weight × F32-input reference calculation. | Quantize activations for the confidence dot too, and match the two alternating four-lane ARM SDOT/FMA accumulators (`7510`) in both confidence and Markov scoring. CPU regression distinguishes the quantized result from the old F32-input dot. |
|
||
| Confidence/output head | `66780–66835`: check first confidence before doing the vocabulary projection; read later rows only as required. Rust calculated all base logits and read all rows before the check. | Implemented the early gate and per-needed-row readback. Confidence-disabled execution keeps hidden/head submission fused. Full1755-cycle comparison remains exact; clean timing acceptance is separate. |
|
||
| Verifier output | `36325–36431`, `69153`, `69210`: fuse head into the layer command sequence, obtain GPU row tops, read only the committed logits row. Rust drained before the head and cloned all vocabulary rows for CPU argmax. | Implemented fused submission, existing GPU argmax/top-k reductions and committed-row readback for the ordinary verifier. Exact sampling retains all target distributions. Full1755-cycle default-mode comparison remains exact. |
|
||
| CPU Markov workers | `33889–33963`: persistent helper pool and fused Q8 argmax; `1912` defaults to min(online CPUs,12), including the caller. Rust created scoped OS threads on every dense argmax, using all18 logical CPUs here. | Implemented persistent Rust workers with the existing Q8 arithmetic, identical contiguous row partition and ordered first-tie reduction. The caller executes slot0; default12 total threads and positive integer `DS4_THREADS` overrides up to32. All1755 original cycles remain exact; timing acceptance remains separate. |
|
||
| Verifier allocations | Original graph retains verifier scratch and GPU frontier/prefix buffers; Rust constructed `BatchScratch` and snapshot buffers for every verifier. | Reuse baseline and high-water prefix buffers, plus one verifier batch per existing padded row shape. Scratch is recycled only after delayed seeding or capture invalidation. Extended32-cycle oracle regression proves native-buffer identity reuse and byte-exact rollback; full-chat and timing evidence below. |
|
||
| Acceptance telemetry | Rust's executor-wide cycle counts and resident support counters have different ownership scopes. | **Open:** reconcile request/session counters before deriving acceptance ratios or exclusive stage percentages. |
|
||
|
||
Excluded after checking the guards: seed-batch fusion defaults to the ROCm
|
||
gfx1151 path, not this Mac (`53412`); the Markov GPU branch is under
|
||
`#ifndef __APPLE__` (`34266`); adaptive extra decode splits are pre-M5-only.
|
||
The output-head padding to8 rows is present in DS4's helper (`26086`) and is
|
||
not an unmatched Rust optimization. None of these were blindly enabled/removed.
|
||
|
||
### Runnable regression evidence
|
||
|
||
`dspark_matches_original_summary_cycles` loads only the installed GGUFs and
|
||
the recorded original JSONL, prepares the actual bootstrap boundary, samples
|
||
with the shared DS4 sampler and verifies each cycle's IDs and position. It is
|
||
ignored by default and explicitly supervised when run:
|
||
|
||
```sh
|
||
DS4SERVER_DSPARK_REFERENCE=local-eval-results/ds4-cycle-audit-20260911.vfsUc8/before-reference.jsonl \
|
||
target/release/test-supervisor 114688 30 45 \
|
||
target/release/deps/ds4_server-6141856e5c0fa6df \
|
||
dspark_matches_original_summary_cycles
|
||
```
|
||
|
||
- `cycle-test-before.*`: fails at generated2, Rust `[10059]` versus original
|
||
`[10059,260,13672,294,270]`.
|
||
- `cycle-test-batch.*`: after single-row batch verification, fails at generated7,
|
||
Rust `[4496,3051,943,30941,22]` versus original `[4496,3051,943,30941]`.
|
||
- `cycle-test-capture.*`: all three unchanged expected cycles pass (7.28s).
|
||
- `cycle-test-32.*`:32 original cycles, prior-request scheduler reset and
|
||
safe logical rewind pass (9.32s). These are functional diagnostics, not
|
||
stories truncated for a throughput comparison.
|
||
|
||
The original full pair, intermediate cache-only full run and failing receipts
|
||
are retained. No original C implementation was added to the application; no
|
||
model download, bundle replacement, commit or push was performed.
|
||
|
||
### Confidence follow-up from the full chat
|
||
|
||
The subsequent `fixed-rust.*` full run matches451 complete returned cycles
|
||
(including warmup), then differs at Summary generated1018. The first divergent
|
||
decision is one cycle earlier: at position3759, original DS4 proposes `[588,699]`
|
||
and rejects the first draft, while Rust's confidence logit0.40281284 falls below
|
||
the0.6 sigmoid threshold and suppresses drafting entirely. The different
|
||
no-draft scheduler decision then changes the returned cycles. All turns still
|
||
end naturally, with counts1194/1424/296 and **no extra cached EOS**; this is
|
||
progress, not a matched-output performance pair (`fixed-cycle-comparison.json`).
|
||
|
||
Read-only inspection of the installed support GGUF confirms the confidence
|
||
head type8/Q8_0,4352 inputs. Correcting the activation quantization yields
|
||
confidence0.4063788 at that exact position and the original two draft IDs.
|
||
The stable Rust SDOT intrinsic is unavailable in this toolchain; the two SDOT
|
||
instructions are isolated in guarded Rust inline assembly, with stdlib NEON
|
||
FMA/reduction and the existing scalar fallback. No C host code or dependency
|
||
was added. Six focused DSpark CPU/layout tests pass, including the new
|
||
quantized-confidence test. Release all-target/all-feature build, Clippy with
|
||
warnings denied, rustfmt and diff checks pass after the correction.
|
||
|
||
Checkpoint-load invalidation now clears the new deferred capture and absolute
|
||
cache metadata through the same `reset_cache` helper as session reset. This
|
||
does not change the checkpoint format or disk KV policy.
|
||
|
||
### Completed functional comparison, performance still open
|
||
|
||
`quantized-rust.*` completes all three turns to natural EOS. All1755 returned
|
||
cycles match the original recording exactly, including warmup, IDs and committed
|
||
positions (`quantized-cycle-comparison.json`). All three text, thinking,
|
||
completion, prompt and cached-token comparisons pass
|
||
(`quantized-result-comparison.json`):
|
||
|
||
| Turn | Completion tokens | Prompt / cached | Rust / original prefill ms | Rust / original decode t/s | Decode gap |
|
||
| --- | ---: | ---: | ---: | ---: | ---: |
|
||
| Summary | 1208 | 2741 /1 | 5562.979 /5519.685 | 36.128 /36.580 | −1.24% |
|
||
| Story | 1455 | 3969 /3949 | 261.651 /281.080 | 30.366 /31.321 | −3.05% |
|
||
| Python | 347 | 5450 /5424 | 323.474 /344.868 | 33.936 /34.777 | −2.42% |
|
||
|
||
These are **diagnostic** timings with cycle logging, not repeat-median clean
|
||
performance acceptance. Story and Python still miss even the single-pair2%
|
||
threshold; the missing early-confidence gate, verifier submission/readback,
|
||
allocation and worker-lifetime tasks above remain explicit work. No UI canary
|
||
or real GUI event-loop acceptance is claimed for this series.
|
||
|
||
The Story has a coherent title, narrative and ending, not a clarification or
|
||
broken output. The Python answer is byte-identical to the previous original
|
||
DSpark answer (`python-reference-unchanged.json`), whose five generated asserts
|
||
and5011 independent cases passed in `deepseek-paired-20260911.nyUCsL/python-quality.jsonl`.
|
||
Six DSpark CPU/layout tests,17 shared sampling tests, the separate heap-fallback
|
||
test and two tokenizer tests pass; model-dependent tests are not silently
|
||
counted as run. The explicit32-cycle live regression passed before the final
|
||
confidence correction, and the final full1755-cycle harness/oracle comparison
|
||
validates the production path after it.
|
||
|
||
Final CLI SHA256:
|
||
`2d15458d12a14457519a22232061ce76182123b24f23d1ae595df424c36963e5`.
|
||
Traced standalone reference SHA256:
|
||
`395cc58f8ea59afa0a6e22f8d46b3fb3f129bc367b8ee3ee1a0dd4bc0f01bd4d`.
|
||
The working app bundle remains
|
||
`ea4d555c2faf0940d9cbcf76d8638ca614a9cb2c6b034e3b2f80aeef86b0b339`.
|
||
|
||
### Early confidence and lazy verifier output
|
||
|
||
The shared Rust executor now follows the original early confidence gate before
|
||
the draft vocabulary projection. Hidden and logits rows are read only when the
|
||
proposal loop needs them; disabling confidence keeps hidden/head encoding fused.
|
||
The ordinary verifier keeps its layer commands alive through the output head
|
||
and the existing GPU argmax/top-k reduction. It reads only the committed
|
||
distribution; exact stochastic sampling retains all row distributions.
|
||
No prefill-cap, power, artificial pause, kernel math or KV persistence change.
|
||
|
||
`lazy-head-rust.*`, `lazy-head-cycle-comparison.json` and
|
||
`lazy-head-result-comparison.json` establish the unchanged complete1755-cycle
|
||
recording and all three answers/thinking/token/cache/stop boundaries. Seven
|
||
focused CPU/layout tests pass (two installed-model tests remain ignored by
|
||
default), including the committed-row boundary regression. Release all-target,
|
||
all-feature build and warnings-denied Clippy pass. This full default-mode
|
||
comparison is not a new exact-stochastic oracle claim.
|
||
|
||
CLI SHA256: `de10b537c2cf2bfbb680ebc350e810305def3ca2ba59f0bea89f1d8c980699cb`.
|
||
The original driver and working app bundle hashes above are unchanged.
|
||
`run-clean-pairs.sh` records serial three-turn timing pairs without cycle logging,
|
||
original proposal logging or canary, with alternating reference-first/Rust-first
|
||
order. Both workers retain startup/progress/memory supervision, no total timeout.
|
||
The first postprocessing command used the wrong reference event name (`result`
|
||
instead of `reference_result`) and failed after both workers had finished; the
|
||
comparison was corrected against the unchanged receipts, with no inference rerun.
|
||
|
||
All six processes completed successfully, with natural EOS for all nine Rust
|
||
answers and matching reference content, thinking, prompt, cache and completion
|
||
counts (1208/1455/347 in every pair). `clean-{1,2,3}-comparison.json` and
|
||
`clean-pairs-summary.json` retain all results; none were discarded:
|
||
|
||
| Pair / order | Turn | Rust / DS4 prefill ms | Rust / DS4 decode t/s | Decode delta |
|
||
| --- | --- | ---: | ---: | ---: |
|
||
| 1 DS4→Rust | Summary | 6396.150 /5435.701 | 34.878 /37.298 | −6.49% |
|
||
| 1 | Story | 262.284 /289.008 | 28.623 /31.716 | −9.75% |
|
||
| 1 | Python | 350.909 /332.242 | 30.372 /35.511 | −14.47% |
|
||
| 2 Rust→DS4 | Summary | 7905.861 /7910.498 | 29.783 /28.784 | +3.47% |
|
||
| 2 | Story | 307.278 /376.104 | 25.918 /24.574 | +5.47% |
|
||
| 2 | Python | 363.522 /488.866 | 29.186 /27.959 | +4.39% |
|
||
| 3 DS4→Rust | Summary | 8389.686 /8292.276 | 27.523 /27.801 | −1.00% |
|
||
| 3 | Story | 301.627 /383.962 | 24.399 /24.151 | +1.02% |
|
||
| 3 | Python | 379.914 /458.934 | 27.995 /27.801 | +0.70% |
|
||
|
||
Pair3 is within2% in decode, with Summary prefill1.17% slower and the two short
|
||
continuation prefills faster. **It is not sufficient acceptance for this series:**
|
||
DS4 itself slows from37.30 to27.80 Summary t/s (about25.5%), and pair order changes
|
||
the sign of the relative gap. Median aggregation cannot establish a causal code
|
||
speedup in this nonstationary series. Do not compare the clean first pair against
|
||
the earlier logged pair as a before/after regression claim. Remaining worker and
|
||
allocation discrepancies still need their reference-aligned implementation and
|
||
an appropriately controlled follow-up.
|
||
|
||
Read-only system checks during pair2 reported no recorded thermal/performance
|
||
warning and25% system memory free; our benchmark processes ran strictly serially.
|
||
These observations do not prove stable GPU clocks or exclude other system GPU
|
||
activity. No application was stopped or configuration changed in response.
|
||
Rust supervisor totals were107.691/121.168/129.178s, maximum observed progress
|
||
gaps6.396/7.905/8.389s in prefill, supervisor lag60/59/59ms and exit0 throughout.
|
||
Canary was off: these are not GPU-canary or actual GUI event-loop measurements.
|
||
Original supervisor totals were96.236/124.665/127.935s, all with `error:null`.
|
||
|
||
The existing ignored `flash_0731_runs_exact_sampled_dspark` initially failed its
|
||
draft-count assertion: its four-token allowance takes the scheduler's existing
|
||
`max_tokens <10` no-draft branch. The test now allows16 tokens and explicitly
|
||
disables the confidence gate to exercise drafting independently of this short
|
||
fixture. It retains the original assertions and additionally invokes two-row
|
||
verification, checking both complete finite distributions, the GPU top ID
|
||
against CPU argmax, and the retained last-row logits. The supervised rerun
|
||
passes in1.60s (`lazy-head-exact-fixed.*`); the initial failure is retained in
|
||
`lazy-head-exact-test.*`. This is a regression check, not a new exact-sampling
|
||
reference/performance claim. No production setting was changed for this test.
|
||
Final warnings-denied Clippy, rustfmt and diff checks pass after this test-only
|
||
edit. No model or benchmark process remains from this work unit.
|
||
|
||
### Persistent CPU Markov workers
|
||
|
||
`src/engine/metal/markov.rs` replaces per-draft scoped thread creation with a
|
||
model-owned Rust worker pool. Original source contracts are `ds4.c:1912–1947`
|
||
(default min(online,12), caller plus helpers), `1971–2005` (contiguous partitions,
|
||
serial execution below512 rows), and `33859–33963` (ordered first-tie Q8 argmax).
|
||
Positive integer `DS4_THREADS` settings are supported up to the original32-thread
|
||
limit. No CUDA/non-Apple Markov branch is enabled.
|
||
|
||
The GGUF mapping is shared through `Arc<Mmap>` without remapping or copying the
|
||
weights. Worker inputs own their temporary data: the full logits `Vec` is moved
|
||
into shared read-only ownership and returned after all dispatched jobs finish.
|
||
There are no borrowed raw pointers between worker lifetimes. Workers release
|
||
their input before signalling completion; results are drained on failure too,
|
||
and pool destruction closes and joins every helper. Existing dot/quantization
|
||
functions are reused without arithmetic changes. No dependency or native host
|
||
code was added.
|
||
|
||
`markov-rust.*` completes the full warmup and ongoing three-turn chat. All1755
|
||
original cycles and all answer/thinking/token/cache/EOS checks remain exact
|
||
(`markov-cycle-comparison.json`, `markov-result-comparison.json`). Two focused
|
||
Markov tests and seven GGUF tests pass, including thread and logits-buffer
|
||
reuse, ordered ties, shape validation and mapping bounds. Release all-target/
|
||
all-feature build and warnings-denied Clippy pass. The full diagnostic run is
|
||
not throughput acceptance. CLI SHA256:
|
||
`0e24c6e8102c2035cc89e93f77eede4235f5aa4611c36ce5fb225770641485f2`.
|
||
|
||
The remaining allocation task is concrete: `BatchScratch::allocate` reserves
|
||
workspace based on both row count and context position; `snapshot_spec_frontier`
|
||
allocates compressor/indexer/target snapshots every verifier cycle. Reuse must
|
||
preserve delayed support seeding, partial acceptance, error rollback and padded
|
||
output-head behavior. Merely retaining an arbitrary previous batch is not a
|
||
safe implementation of the original persistent scratch contract.
|
||
|
||
Two clean full pairs (`clean-markov-{1,2}-*`) ran DS4→Rust→Rust→DS4, power100,
|
||
same installed files/settings/warmup/ongoing chat, trace and canary off. All
|
||
content/thinking/prompt/cache/token checks pass; both workers in both pairs exit
|
||
successfully, without downloads or overlapping model processes:
|
||
|
||
| Pair | Turn | Rust / DS4 prefill ms | Rust / DS4 decode t/s | Decode delta |
|
||
| --- | --- | ---: | ---: | ---: |
|
||
| 1 | Summary | 6686.486 /5445.567 | 32.046 /36.968 | −13.32% |
|
||
| 1 | Story | 296.457 /280.371 | 26.201 /31.162 | −15.92% |
|
||
| 1 | Python | 376.741 /362.034 | 28.251 /34.122 | −17.21% |
|
||
| 2 | Summary | 8407.889 /8535.688 | 26.901 /27.488 | −2.14% |
|
||
| 2 | Story | 332.174 /383.535 | 23.301 /24.322 | −4.20% |
|
||
| 2 | Python | 398.755 /466.153 | 26.975 /28.215 | −4.40% |
|
||
|
||
The second pair still misses decode parity. DS4 Summary throughput again falls
|
||
by about25.6% over the series; no systemwide throttling cause is asserted. A
|
||
single `sudo -n powermetrics` query failed immediately because a password was
|
||
required (`markov-power-sample.txt`); it started no sampler. Rust supervisor
|
||
totals117.095/132.871s versus DS4 totals98.054/127.837s include different
|
||
frontend/finalization work and are not pure decode. Rust maximum progress gaps
|
||
were6.686/8.407s and supervisor lag59/58ms. No canary/UI-responsiveness claim.
|
||
|
||
To separate dispatch cost from GPU drift, the ignored CPU-only
|
||
`installed_markov_worker_dispatch` test uses installed Markov W1 row671 and the
|
||
actual W2 mapping with a fixed logits row. It compares per-call18-thread spawning
|
||
against persistent12 and18 workers, holding row arithmetic/input ownership
|
||
constant. Four alternating orders,128 calls per mode per round, all1536 argmax
|
||
results identical. Median times per128 calls:
|
||
|
||
| Dispatch | Median ms | ms/call |
|
||
| --- | ---: | ---: |
|
||
| Scoped18 | 40.356 | 0.3153 |
|
||
| Persistent12 (production default) | 36.077 | 0.2819 |
|
||
| Persistent18 (diagnostic only) | 30.496 | 0.2382 |
|
||
|
||
The default pool reduces this isolated dispatch/calculation time by10.6%, about
|
||
0.034ms per call; it does not explain seconds of whole-chat difference. This
|
||
diagnostic is neither a C Markov microbenchmark nor end-to-end parity evidence.
|
||
Production retains the reference's default12-thread policy. The same runnable
|
||
worker regression also closes one helper and checks that all other jobs drain,
|
||
an error is returned and the caller's original logits buffer is preserved.
|
||
Receipts: `markov-dispatch-test.*`, `markov-worker-failure-tests.txt`.
|
||
|
||
Final verification after the additional test-only coverage: seven enabled
|
||
DSpark tests pass; the explicitly supervised exact-sampling test passes in1.35s;
|
||
warnings-denied Clippy, rustfmt and diff checks pass. All processes from this
|
||
work unit have finished. No app bundle replacement, commit or push was made.
|
||
The full six-cell performance goal remains open.
|
||
|
||
### Persistent verifier buffers: functional checks pass, timing remains unaccepted
|
||
|
||
The next implementation reuses baseline compressor/indexer snapshots, high-water
|
||
prefix snapshots and one verifier workspace per existing padded row shape.
|
||
Deferred seed ownership determines when a batch can be recycled; full-prefill
|
||
workspaces are not retained. Position-sensitive verifier storage reserves the
|
||
session context. Existing output-head row padding, sampling, power and command
|
||
submission boundaries are unchanged. Error rollback keeps its previous behavior;
|
||
an error may drop scratch and require allocation on the next attempt.
|
||
|
||
Evidence remains in `local-eval-results/ds4-cycle-audit-20260911.vfsUc8/`.
|
||
CLI SHA256: `4f3643bd854cad1d317fdd6bb93649327902b19cb3cd7b9cfff24a760d57e880`.
|
||
The prior CLI is retained as `before-verifier-reuse-ds4-server` for a later
|
||
controlled comparison; merely comparing older sequential runs is insufficient.
|
||
|
||
- `reuse-cycle-test.txt`: the32-cycle original oracle passes, now also checking
|
||
native buffer identities for both padded verifier shapes and byte-exact
|
||
compressor/indexer rollback after another target step.
|
||
- `reuse-cycle-comparison.json` and `reuse-result-comparison.json`: the complete
|
||
traced chat matches all1755 original cycles and all text/thinking/token/cache
|
||
checks. Traced timing is diagnostic only.
|
||
- `reuse-exact-test.txt`: supervised exact-sampling regression passes after the
|
||
reuse changes, including complete target-distribution readback.
|
||
- `reuse-build.txt`, `reuse-clippy.txt`, `reuse-dspark-tests.txt`: release build,
|
||
warnings-denied Clippy and seven enabled DSpark tests pass.
|
||
- GPU allocation at each measured turn end is95,039,750,144bytes, approximately
|
||
73MiB more retained than the prior Summary endpoint. This is retained scratch,
|
||
not proof of a throughput improvement.
|
||
|
||
Two complete pairs ran in DS4→Rust→Rust→DS4 order, with trace/canary off and no
|
||
overlapping model processes or builds. The user subsequently confirmed concurrent
|
||
video playback; "clean" in these artifact filenames means instrumentation off,
|
||
**not** an idle GPU. Both pairs preserve all outputs and exit successfully.
|
||
|
||
| Pair | Turn | Rust / DS4 prefill ms | Rust / DS4 decode t/s | Decode delta |
|
||
| --- | --- | ---: | ---: | ---: |
|
||
| 1 | Summary | 6679.941 /5718.646 | 34.204 /38.273 | −10.63% |
|
||
| 1 | Story | 259.070 /273.257 | 28.306 /33.200 | −14.74% |
|
||
| 1 | Python | 355.567 /333.715 | 29.343 /36.744 | −20.14% |
|
||
| 2 | Summary | 8968.416 /10283.116 | 26.678 /24.830 | +7.44% |
|
||
| 2 | Story | 324.111 /445.256 | 21.926 /22.225 | −1.35% |
|
||
| 2 | Python | 365.097 /560.790 | 23.092 /26.814 | −13.88% |
|
||
|
||
Receipts: `clean-reuse-{1,2}-{comparison.json,rust.jsonl,reference.jsonl}` and
|
||
matching stderr. Rust supervisor totals110.153/140.256s versus original
|
||
93.671/141.037s include frontend/finalization differences, not just decode.
|
||
Rust progress gaps6.679/8.968s and supervisor lag45/45ms are not GPU-canary or
|
||
GUI-eventloop delays. No new responsiveness acceptance, app bundle replacement,
|
||
commit or push. Remaining work includes telemetry ownership reconciliation,
|
||
reference code/cost audit and the controlled full six-cell comparison.
|
||
|
||
### Interim checkpoint requested by the user
|
||
|
||
The current work unit is being closed for commit/push, not declared full parity.
|
||
The commit-gate test run exposed an already-committed stale Qwen source-inventory
|
||
assertion:22 runtime units were expected although both HEAD's Metal export and
|
||
the pinned generator contain26. The read-only command below verifies the entire
|
||
export byte for byte against pinned sources (17 custom bodies,26 runtime units,
|
||
dynamic QSA sources). Only the stale expected count is corrected; no kernel,
|
||
hash, fixture or per-body assertion is changed or removed.
|
||
|
||
```sh
|
||
python3 tools/mtplx-kernel-source.py local-eval-results/mtplx-reference-e652d55 --gated-delta-source local-eval-results/mtplx-reference-env-0.32.2/lib/python3.12/site-packages/mlx_lm/models/gated_delta.py --check
|
||
```
|
||
|
||
The app bundle is rebuilt as required by the commit gates. This supersedes the
|
||
earlier statements that this follow-up had not yet replaced the bundle; no GUI
|
||
or new throughput series is launched. Unrelated `tools/__pycache__/` files are
|
||
left untouched and excluded from the commit.
|
||
|
||
Commit verification: `cargo fmt --all -- --check`, warnings-denied all-target/
|
||
all-feature Clippy, `make bundle`, and `cargo test --all-features` pass. The full
|
||
suite reports303 main tests plus4 supervisor and4 integration tests passed,
|
||
zero failures,202 explicitly ignored main tests. The separately supervised
|
||
32-cycle and exact-sampling results above remain distinct from these normal
|
||
gates. `codesign --verify --deep --strict` and the bundle's `model-eval --help`
|
||
also pass without opening the GUI. Final suite receipt:
|
||
`local-eval-results/ds4-cycle-audit-20260911.vfsUc8/interim-commit-tests.txt`.
|