593 lines
36 KiB
Markdown
593 lines
36 KiB
Markdown
# DS4Server implementation plan
|
|
|
|
Bundle/application identifier: `DS4Server.rfc1437.de`
|
|
|
|
The GUI is a Rust/Iced reimplementation of the user-facing model server and
|
|
coding-agent workflows in `../ds4`. The existing engine remains the behavioral
|
|
reference: `ds4.c`/`ds4.h` define model and session behavior,
|
|
`ds4_server.c` defines the compatible API, and `ds4_agent.c` defines agent
|
|
sessions, tools, and turn orchestration.
|
|
|
|
## Current status
|
|
|
|
| Area | Status | Available now | Still open |
|
|
| --- | --- | --- | --- |
|
|
| App shell | Implemented | Native multi-window Iced app, Codex-inspired layout, Chat/Stats segmented workspace, native Application/Edit/Window menus, Preferences panel, Model Manager, and streaming composer | Native app release packaging |
|
|
| Projects and sessions | Implemented with transcripts and KV resume | Native folder picker, project-name dialog, create/select/delete projects and sessions, durable structured chat history, and per-session KV checkpoints | Rename/archive and model binding |
|
|
| Persistence | Implemented for local chat | SQLite in Application Support, Diesel ORM, embedded migrations, constrained defaults, streamed message updates, reopen tests, context/speed state, and local KV rehydration | Downloaded-artifact metadata and session-runtime migrations; HTTP chats deliberately never enter SQLite |
|
|
| Model runtime | DeepSeek Flash decode and batched prefill implemented | Rust-owned mmap/model/session graph, local Metal kernels, full configured sparse context, exact absolute-position chunk scheduling, cold and resumed layer-major prefill (including unaligned compressor/indexer frontiers), DS4 sampling, one process-wide UI/HTTP owner, cancellation, idle unloading, prefix continuation, KV checkpoint save/load, and lock-free runtime metrics | Remaining KV continuation/rewrite parity, DSpark/SSD/steering, GLM/Pro execution, and sampled per-kernel GPU timing if coarse phase metrics prove insufficient |
|
|
| Local endpoint | Exact parity in progress | Configurable localhost listener, verified-on-disk model discovery, transient Chat/Completions/Messages/Responses routes, native JSON/SSE envelopes, exact tool schema/call/result replay, live tool continuation, and differential C/Rust fixtures for tool use, cache accounting, reasoning, cold 6k-token prefill, resumed unaligned prefill, and real Pi tool execution | Remaining `ds4_server.c` parser/tool recovery, KV-store policy, disconnect/cancellation, scheduling/lifecycle, and error/context-limit fixtures; enable/disable controls |
|
|
| Agent | Solid local chat implemented | Multi-turn role rendering, structured reasoning disclosure, Markdown, automatic scrolling, context/speed display, durable transcript rehydration, KV-backed resume, and a live Stats dashboard | Tools, approvals, and compaction |
|
|
| A2UI | Future extension | bDS2 provides the reference structured render-tool contract | Native inline surfaces for local chat after core chat and tools are stable |
|
|
| Dev brain | Not started | No vault access | Obsidian vault selection, durable access, and agent knowledge/memory tools |
|
|
| Release | Not started | Development binary builds and tests | `.app` packaging, signing, entitlements, and notarization |
|
|
|
|
The application manages durable project/session metadata and model preferences.
|
|
Its Model Manager downloads, validates, and deletes managed GGUF artifacts
|
|
without blocking the UI. Local chat now lazily acquires the Rust DeepSeek Flash
|
|
executor, streams generated text, supports cancellation and multi-turn role
|
|
replay, persists reasoning and answers while they stream, rehydrates selected
|
|
sessions, and unloads model resources after the configured idle timeout. The
|
|
current Rust graph supports the full configured context with ratio-4 sparse
|
|
attention, fixed-size cold/resumed batch chunks, unaligned compressor/indexer
|
|
continuation, and durable local KV checkpoints. The shared HTTP service exposes
|
|
every `ds4_server.c` route and passes the initial differential corpus for native
|
|
tool use/continuation, cache accounting, reasoning, and Pi execution. Full
|
|
reference parity remains in progress until the remaining parser, KV-store,
|
|
cancellation, scheduling, and error fixtures are ported.
|
|
|
|
## Phase 0 — project and session shell
|
|
|
|
Status: **implemented**.
|
|
|
|
- Show projects in a sidebar with their sessions nested underneath.
|
|
- Create projects with a native macOS folder picker and a project-name dialog.
|
|
- Create, select, and delete projects and sessions.
|
|
- Use a Codex-inspired two-pane layout with embedded, action-specific SVG icons.
|
|
- Persist projects and sessions in SQLite under
|
|
`~/Library/Application Support/DS4Server.rfc1437.de/`, using Diesel ORM
|
|
queries and embedded Diesel migrations.
|
|
- Never delete workspace contents when removing app metadata.
|
|
|
|
Exit criterion: restart the app and see the same project/session tree.
|
|
|
|
## Phase 1 — model library and on-demand loading
|
|
|
|
Status: **partially implemented**. The typed preference/download/intake path and
|
|
a DeepSeek Flash Rust/Metal vertical are implemented, including full configured
|
|
context, ratio-4 sparse indexed attention, prefix continuation, KV checkpoint
|
|
save/load, DS4 sampling, cancellation, idle unload, and shared HTTP acquisition.
|
|
Other model families and optional execution modes remain open.
|
|
|
|
1. **Implemented:** the typed model/runtime preference contract described below
|
|
is complete before chat, the HTTP endpoint, or the engine lifecycle.
|
|
2. **Implemented for DeepSeek Flash:** mmap-backed loading, tokenizer/prompt
|
|
rendering, Rust-owned session creation/evaluation/sampling, prefix sync, and
|
|
KV save/load are complete.
|
|
3. **Implemented:** retain the tested model restrictions from DwarfStar. Main
|
|
and DSpark artifacts reject invalid GGUF versions, catalog mismatches, and
|
|
incompatible metadata, tensor shapes, offsets, or quantization before
|
|
promotion and again before model use.
|
|
4. **Partially implemented:** reuse the existing `.metal` kernels and preserve mmap-backed resident
|
|
loading plus the explicit SSD-streaming path. Keep Objective-C only at the
|
|
Metal interop boundary.
|
|
5. Keep exactly one model resident process-wide. Local chat and every endpoint
|
|
request must use the same owner and queue; model selection may replace the
|
|
resident model only after active work reaches a safe boundary. Multiple
|
|
simultaneously loaded models are out of scope because DS4Server targets
|
|
models large enough to consume most available memory.
|
|
6. **Implemented for local chat:** move engine work off the UI thread and support cooperative cancellation at
|
|
the safe session boundaries already defined by `ds4_session_sync`.
|
|
|
|
### DS4 KV-cache port and shared session core
|
|
|
|
KV-cache behavior is part of the DS4 engine contract, not a new application
|
|
cache to redesign. Port the existing `ds4_session` and `ds4_kvstore` behavior
|
|
to Rust while preserving its semantics and on-disk compatibility:
|
|
|
|
- A live engine session owns one mutable inference timeline: exact rendered
|
|
tokens, logits, and the DS4-specific KV graph state. Callers provide the full
|
|
token prefix and the Rust port of `ds4_session_sync` must retain the longest
|
|
compatible prefix, evaluate only its suffix, or safely rewind, invalidate,
|
|
and rebuild exactly when DS4 does.
|
|
- Carry over DS4's safe cancellation/checkpoint boundaries, common-prefix
|
|
matching, canonical rewrite and replay, snapshots, payload staging,
|
|
save/load, strip-to-transcript behavior, and atomic replacement. The engine
|
|
continues to own the opaque KV payload format; UI, HTTP, and persistence code
|
|
must not reinterpret tensor state.
|
|
- Port `ds4_kvstore` policy rather than replacing it with a generic cache:
|
|
exact token and rendered-text prefix lookup, cold and continued aligned
|
|
checkpoints, model/quantization/context/payload-ABI compatibility checks,
|
|
tool-call replay metadata, disk-budget enforcement, hit tracking, and
|
|
eviction behavior.
|
|
- Use one process-wide Rust model owner and KV checkpoint implementation for
|
|
local chat and the HTTP server. Model access is serialized so two callers
|
|
never mutate the single resident inference graph concurrently.
|
|
- Application projects and sessions belong exclusively to local chat. HTTP
|
|
requests are transient and must never create projects, sessions, messages,
|
|
transcript rows, request-identity mappings, or any other durable chat record.
|
|
The outside client owns conversation identity and resends the history needed
|
|
for each request, exactly as `ds4_server.c` expects.
|
|
- Transient HTTP requests may read and write engine-owned KV cache files for
|
|
compatible-prefix reuse. Those opaque cache files are the only server-side
|
|
persistence permitted for external conversations and do not establish an
|
|
application session or recoverable transcript.
|
|
- Store engine-owned checkpoints under Application Support in `kv-cache/`.
|
|
Local chats use their application session ID; HTTP uses content-addressed
|
|
entries under `kv-cache/http/`. The large opaque payload is never stored in
|
|
SQLite, and HTTP cache entries do not create application sessions.
|
|
- Cache sharing does not merge conversations or expose transcript text. It
|
|
shares only the DS4 implementation and compatible opaque prefix checkpoints;
|
|
the request body remains the source of truth for external conversation state.
|
|
- Engine unload must first leave every reusable live timeline in a valid
|
|
persisted checkpoint. Reload restores a compatible checkpoint immediately;
|
|
if model identity, quantization, context size, payload ABI, or rendered
|
|
history is incompatible, rebuild from the transcript and replace the stale
|
|
checkpoint only after the new state is valid.
|
|
|
|
### Lazy lifecycle
|
|
|
|
- Start the application and the local HTTP endpoint with no model loaded.
|
|
- Acquire the configured model only when generation starts, whether initiated
|
|
by a local chat or an OpenAI-compatible endpoint request. Concurrent first
|
|
requests share the same load instead of opening the model more than once.
|
|
- Expose `Unloaded → Downloading → Loading → Ready → Unloading` plus failure
|
|
states and progress in the GUI. A request that triggered loading waits for
|
|
the shared load and receives a clear error if download or initialization
|
|
fails.
|
|
- Treat active local generations, endpoint requests, prefills, and tool turns
|
|
as model activity. Start the idle timer only when all activity has stopped.
|
|
- Unload the engine after the configured idle timeout. New work cancels a
|
|
pending unload; work arriving after unload transparently loads the model
|
|
again. Persisted transcripts and KV files remain available when model memory
|
|
is released.
|
|
- Changing the selected model prevents new work from using the old model and
|
|
unloads it as soon as current work reaches a safe boundary.
|
|
|
|
### Preferences foundation (implemented)
|
|
|
|
- Provide a Preferences panel from the sidebar and the standard
|
|
macOS `Command-,` shortcut.
|
|
- Let the user choose the active model. Default to `deepseek-v4-flash`.
|
|
- Let the user enable or disable DSpark for the selected model. Disable the
|
|
control for models that do not support DSpark.
|
|
- Let the user configure the inactivity timeout; default to 10 minutes.
|
|
- Let the user configure the localhost endpoint port; default to 4000 and apply
|
|
a changed port when preferences are saved.
|
|
- Persist preferences in the application SQLite database. Changing preferences
|
|
never loads a model by itself.
|
|
|
|
### DS4 model and generation preferences (implemented)
|
|
|
|
The Preferences panel must expose every persistent model, runtime, and
|
|
generation parameter accepted by the `ds4` command line. `../ds4/ds4_cli.c`,
|
|
`ds4.h`, and the shared GPU/distributed parsers are the source of truth; use the
|
|
same defaults, ranges, units, dependencies, and model-family overrides rather
|
|
than inventing application-specific behavior.
|
|
|
|
- **Model and capacity:** selected main GGUF (`--model`), optional MTP/DSpark
|
|
support GGUF (`--mtp`), context tokens (`--ctx`), maximum generated tokens
|
|
(`--tokens`), and system prompt (`--system`). Managed catalog selections map
|
|
to their installed files instead of requiring users to type paths.
|
|
- **Sampling and reasoning:** temperature (`--temp`), top-p, min-p, optional
|
|
seed, and one reasoning mode representing `--think`, `--think-max`, or
|
|
`--nothink`. Preserve DS4's explicit-vs-default distinction because GLM
|
|
applies model-family sampling defaults only when the user has not overridden
|
|
them.
|
|
- **Execution:** the macOS app uses Metal exclusively; omit the unnecessary CPU
|
|
backend and backend selector. Persist CPU helper threads for host-side work,
|
|
GPU power percentage, prefill chunk, exact/quality kernels, and warm weights.
|
|
- **Speculative decoding:** MTP draft-token count and margin, GLM MTP, GLM MTP
|
|
timing, DSpark enablement, DSpark confidence threshold, and DSpark strict
|
|
target-only decode. Enabling a dependent control must enable or require its
|
|
support artifact exactly as the CLI does.
|
|
- **SSD streaming:** enablement, cold start, expert count or GiB cache budget,
|
|
fully resident GLM layers, and explicit expert preload count. Retain the
|
|
distinction between unset/automatic and an explicit zero where DS4 does.
|
|
- **GPU placement:** per-device VRAM budgets, device indices, and CUDA tensor
|
|
parallelism when available. Hide or disable these controls in a Metal-only or
|
|
CPU-only build rather than accepting values the engine cannot honor.
|
|
- **Directional steering:** direction-vector file, FFN scale, and attention
|
|
scale, including the CLI rule that a file with no explicit scale defaults the
|
|
FFN scale to 1.
|
|
- **Advanced diagnostics that change model execution:** simulated used memory
|
|
and routed expert profile output. Keep these in an Advanced section with
|
|
clear diagnostic labels, but persist and pass them through like their CLI
|
|
equivalents.
|
|
|
|
Prompt sources, raw one-shot mode, inspect/dump/perplexity/imatrix/self-test
|
|
actions, and distributed coordinator/worker network topology are CLI workflows,
|
|
not persistent model parameters; they belong to their corresponding chat,
|
|
diagnostic, or deployment surfaces rather than Preferences. Any future DS4 CLI
|
|
option must be classified explicitly as either a preference or one of these
|
|
non-preference actions.
|
|
|
|
Implement this before chat as one typed settings path shared by the local GUI,
|
|
the lazy engine lifecycle, and HTTP defaults:
|
|
|
|
- Persist unset/automatic values separately from explicit values and provide a
|
|
Reset to DS4 defaults action. Validate before saving with the same accepted
|
|
ranges and cross-field rules as the CLI.
|
|
- Mark settings as load-time or turn-time. A load-time change unloads/reloads at
|
|
the next safe boundary; sampling, seed, token budget, and reasoning changes
|
|
apply to the next turn without duplicating the engine.
|
|
- Local chat uses these values directly. HTTP requests use them as defaults and
|
|
may override only request-scoped generation fields supported by the API; both
|
|
paths still resolve through the same effective-settings builder.
|
|
- Add a focused parity test that inventories every option branch in the DS4 CLI
|
|
model/runtime, sampling, and steering groups and fails until each option is
|
|
mapped to a typed preference or explicitly classified as a non-preference
|
|
action.
|
|
|
|
Exit criterion: save, restart, and recover every supported DS4 model/runtime
|
|
and generation value; show the same effective defaults and validation as the
|
|
CLI; and construct one engine configuration plus one turn configuration without
|
|
chat or HTTP code re-parsing preference fields.
|
|
|
|
### Runtime observability (implemented)
|
|
|
|
- Keep instrumentation in the shared model/runtime and HTTP paths so local chat
|
|
and endpoint work report through the same counters.
|
|
- Publish hot-path state with relaxed atomics only. Generation and prefill never
|
|
wait for the UI, allocate telemetry events, write logs, or scan the KV cache.
|
|
- Sample snapshots in the UI every 200 ms and retain only a bounded in-memory
|
|
history for prefill, decode, request-rate, and KV read/write-rate graphs.
|
|
- Show model phase and lifecycle, prefill/decode throughput, context occupancy,
|
|
request queue and source, exact memory/disk hits, misses, invalid entries,
|
|
prefix reuse, checkpoint storage and I/O, endpoint traffic and latency,
|
|
errors, and model mapping details on the Stats tab.
|
|
- Derive KV file sizes when checkpoints are already being saved and scan the
|
|
cache directory only once at startup. Add sampled Metal command-buffer or
|
|
per-kernel timings later only if coarse phase counters cannot diagnose a
|
|
measured engine bottleneck.
|
|
|
|
Exit criterion: model generation and endpoint traffic visibly update the Stats
|
|
dashboard without introducing locks, telemetry I/O, or unbounded history on the
|
|
prefill/decode path.
|
|
|
|
### Targeted downloads and storage
|
|
|
|
- Preferences only selects the model/runtime configuration. A separate Model
|
|
Manager window, reachable from the native Window menu, lists every managed
|
|
main and DSpark artifact with its on-disk size and state, and provides
|
|
Download/Resume, Validate, Stop, and confirmed Delete actions.
|
|
- A dedicated worker thread performs HTTP and verification work so the Model
|
|
Manager and the rest of the application stay interactive.
|
|
- Use an in-process Rust HTTP client with Rustls rather than invoking a system
|
|
downloader. Stream response chunks directly into a `.part` file and use a
|
|
Rust SHA-256 implementation for verification.
|
|
- Report aggregate bytes present, total bytes, bytes missing, percentage,
|
|
smoothed transfer rate, current artifact/verification phase, and estimated
|
|
time remaining. Keep the detailed display in Model Manager and a persistent
|
|
status bar with a Stop action while the transfer runs.
|
|
- During SHA-256 verification, show a dedicated progress bar based on bytes
|
|
read and hashed rather than the already-complete file size. Update the bytes
|
|
verified, verification throughput, percentage, and estimated time remaining
|
|
throughout the pass so large models never appear stalled at 100%.
|
|
- Stop is non-destructive: close the active transfer and retain its `.part`
|
|
file. Starting the download again in the same process or after quitting and
|
|
relaunching must derive progress from that file and issue an HTTP range
|
|
request from its exact byte length. If the server ignores the range request,
|
|
safely restart that artifact from byte zero rather than appending corrupt
|
|
data.
|
|
- Promote a `.part` file only after its exact size and SHA-256 are verified. A
|
|
stopped verification keeps the complete partial file so the next run can
|
|
verify and promote it without downloading it again.
|
|
- Store managed model artifacts under
|
|
`~/Library/Application Support/DS4Server.rfc1437.de/models/<model-id>/` and
|
|
mmap/load them from there. Do not place managed models in the app bundle or a
|
|
project directory.
|
|
- Download only the artifact explicitly requested in Model Manager; never
|
|
prefetch other catalog models or optional artifacts.
|
|
- Remove checksum-invalid partial files without touching a valid model. Keep
|
|
partial files after cancellation, application exit, network failure, or an
|
|
interrupted verification so they remain restart-resumable.
|
|
- Delete model artifacts only through an explicit, confirmed Model Manager
|
|
action; changing Preferences must never remove files.
|
|
|
|
Exit criterion: select a supported DeepSeek V4 or GLM 5.2 GGUF, load it without
|
|
blocking the UI only when a chat begins, run a deterministic prompt, unload it
|
|
after the configured idle timeout, and pass the matching `../ds4` token-output
|
|
fixture. Repeat through the local endpoint and verify both paths share the same
|
|
engine lifecycle.
|
|
|
|
## Phase 2 — exact `ds4_server.c` HTTP parity
|
|
|
|
Status: **exact parity in progress**. The shared single-model runtime,
|
|
verified-model discovery, configurable localhost port, all four generation
|
|
routes, native JSON/SSE response envelopes, exact batched prefill, tool replay,
|
|
and live continuation exist. Differential C/Rust fixtures cover tools,
|
|
reasoning, usage/cache accounting, and real Pi execution. Full parser/tool
|
|
recovery, KV-store, cancellation, scheduling, lifecycle, and error fixture
|
|
parity remain open, so no route is considered complete yet.
|
|
|
|
`../ds4/ds4_server.c` is normative for every externally observable HTTP
|
|
behavior. “Compatible” approximations are not acceptable: routes, accepted
|
|
fields, defaults, aliases, prompt rendering, status codes, headers, JSON/SSE
|
|
shapes and ordering, five-second prefill keepalives, reasoning/tool translation,
|
|
usage, finish reasons, errors, cancellation, and cache continuation must match
|
|
the C server exactly. Port its fixtures alongside each Rust slice and keep every
|
|
previous parity fixture passing.
|
|
|
|
This requirement covers the complete communication behavior, not only the HTTP
|
|
wire format. The Rust request path must follow the same request parsing and
|
|
validation, model aliases, prompt/token construction, reasoning mode, sampling,
|
|
stop matching, tool schema/call/result translation, live tool continuation,
|
|
canonical replay, KV lookup/rewrite/checkpoint policy, scheduling, disconnect
|
|
cancellation, usage accounting, finish decisions, and model lifecycle invoked
|
|
by `ds4_server.c`. `ds4.c` is normative for engine operations reached from the
|
|
server. A route is complete only when its differential request corpus passes
|
|
against the C server; partial route support stays labeled incomplete.
|
|
|
|
1. **Partially implemented:** add a localhost-only HTTP service whose lifecycle is independent of the
|
|
engine. It can listen while the model is unloaded and acquires the configured
|
|
model only for inference. It listens on configurable port 4000 by default;
|
|
start/stop controls remain open. Never expose a LAN listener without an
|
|
explicit setting.
|
|
2. **Partially implemented:** expose the complete DwarfStar HTTP surface from
|
|
`ds4_server.c` in this
|
|
order, preserving its request parsing, prompt formatting, reasoning fields,
|
|
streaming events, errors, usage accounting, stop behavior, tool replay, and
|
|
cancellation semantics rather than inventing an app-specific protocol:
|
|
- `GET /v1/models`
|
|
- `GET /v1/models/{known-alias}`
|
|
- `POST /v1/messages`
|
|
- `POST /v1/chat/completions`
|
|
- `POST /v1/responses`
|
|
- `POST /v1/completions`
|
|
- `OPTIONS` with the same CORS behavior when enabled
|
|
|
|
`GET /v1/models` scans managed artifacts without forcing a model to load and
|
|
reports only main model files that are fully downloaded and have a matching
|
|
on-disk verification marker. Missing, partial, unverified, or checksum-invalid
|
|
artifacts are never advertised.
|
|
3. **Partially implemented:** streaming SSE, reasoning output, usage accounting,
|
|
sampling parameters, tool schemas, and automatic/disabled tool choice work;
|
|
exact disconnect cancellation and remaining forced-choice errors are open.
|
|
4. **Partially implemented:** port exact sampled tool-call replay and deterministic canonicalization from
|
|
`ds4_server.c` so a client's normalized JSON does not destroy KV-prefix
|
|
reuse.
|
|
5. **Partially implemented:** route API work through the single process-wide model owner and engine-owned
|
|
KV store defined in Phase 1. External requests are always stateless at the
|
|
application layer: never write their messages to SQLite and never bind a
|
|
Responses `conversation`, request identifier, or client-supplied session key
|
|
to a local project/session. The client must replay its protocol history;
|
|
DS4-compatible prefix lookup may accelerate that replay using opaque KV cache
|
|
files. Add resident batching only if one-model serialized execution is later
|
|
proven insufficient.
|
|
6. **Implemented:** display endpoint address, model, active requests, token
|
|
rates, cache behavior, and errors in the GUI Stats dashboard.
|
|
|
|
Exit criterion: the complete `ds4_server.c` fixture suite and a differential
|
|
black-box corpus pass against the Rust app for Models, Anthropic Messages, Chat
|
|
Completions, Responses, Completions, and OPTIONS, in streaming and non-streaming
|
|
forms where supported. The corpus covers valid requests, aliases/defaults,
|
|
context limits, tools and multi-turn continuation, reasoning, cache reuse,
|
|
cancellation/disconnects, queueing, errors, usage, and finish reasons. Verified
|
|
installed models are listed without loading one; missing/unverified models are
|
|
absent; local and HTTP generations cannot load two models concurrently;
|
|
external calls leave no project, session, message, or transcript rows behind;
|
|
and compatible KV cache files remain reusable.
|
|
|
|
## Phase 3 — durable agent sessions
|
|
|
|
Status: **partially implemented for local chat only**. Project/session metadata,
|
|
structured transcripts, context statistics, and per-session KV checkpoints are
|
|
durable. Model binding and archive behavior remain open. This phase never
|
|
applies to external endpoint conversations.
|
|
|
|
Extend each metadata-only session with transcript and shared-checkpoint
|
|
metadata:
|
|
|
|
```text
|
|
Application Support/DS4Server.rfc1437.de/
|
|
data.sqlite3
|
|
models/<model-id>/
|
|
main.gguf
|
|
dspark.gguf # only when enabled
|
|
kv-cache/
|
|
<local-session-id>.bin
|
|
http/<conversation-tag>.bin
|
|
```
|
|
|
|
- **Implemented:** migrated message tables persist local user text, reasoning
|
|
state, assistant output, context usage, and generation speed while generation
|
|
streams. Local session KV payloads remain engine-owned files written with
|
|
atomic replacement; SQLite never stores opaque tensor data.
|
|
- Record model identity, context size, rendered token history, title,
|
|
timestamps, and working directory.
|
|
- Restore compatible KV immediately. If KV is absent or incompatible, rebuild
|
|
from the transcript and show prefill progress.
|
|
- Add strip/archive behavior equivalent to the native agent: retain the
|
|
transcript and clear its checkpoint reference. Remove the cache file only
|
|
when it is not reusable by another live/session prefix; otherwise normal DS4
|
|
disk-budget eviction reclaims it.
|
|
- Version formats before the first release and add migrations only when a real
|
|
format change exists.
|
|
|
|
Exit criterion: quit during normal use, relaunch, select a session, and resume
|
|
the same conversation without prefill when its KV payload is compatible.
|
|
|
|
## Phase 4 — agent chat and tools
|
|
|
|
Status: **basic local chat implemented**. The conversation area streams
|
|
multi-turn DeepSeek Flash output, persists and rehydrates structured reasoning
|
|
and answers, and supports Stop; tool turns and the coding-agent runtime remain
|
|
open.
|
|
|
|
Every chat feature must persist its semantic state and rehydrate it when a
|
|
session is reopened in the same implementation slice. In-memory-only chat
|
|
features are incomplete by definition; disclosure state and other purely
|
|
ephemeral presentation preferences are the only exception.
|
|
|
|
1. **Partially implemented:** build the classic chat surface. Transcript,
|
|
structured reasoning disclosure, composer, streamed output, and Stop are
|
|
complete; tool-call cards, queued input, prefill progress, and generation
|
|
statistics remain. Local turns must call the same session registry and KV
|
|
path used by the HTTP endpoint.
|
|
2. Port the native agent turn loop from `ds4_agent.c`, including model-specific
|
|
system prompts and DeepSeek DSML/GLM tool-call parsing.
|
|
3. Implement the local tool set with the project directory as its boundary:
|
|
- shell command execution
|
|
- file read and continuation
|
|
- file write
|
|
- anchored edit
|
|
- text/file search
|
|
- opt-in Dev Brain knowledge lookup and memory storage
|
|
4. Stream assistant tokens and partial tool calls into Iced while inference and
|
|
tools run on workers. Preserve cooperative interruption and a valid KV
|
|
prefix at every stop point.
|
|
5. Require confirmation for commands or writes outside the project boundary,
|
|
destructive actions, and network tools. Add web search/page visiting only
|
|
after the local tool loop is stable.
|
|
6. Port context compaction and system-prompt reinjection after ordinary
|
|
multi-turn operation is correct.
|
|
|
|
### Native macOS menus (required with chat)
|
|
|
|
Status: **partially implemented**. Application, Edit, and Window menus are
|
|
installed; Cut, Copy, Paste, and Select All now bridge into the focused Iced
|
|
text field. File/View/Help menus, dynamic enabled state, the remaining Edit
|
|
actions, and selectable transcript text remain open.
|
|
|
|
The application must install and maintain a complete native macOS menu bar;
|
|
the application-name menu alone is not sufficient. Use platform menu roles,
|
|
ordering, selectors, and standard key equivalents rather than drawing menus
|
|
inside the Iced window.
|
|
|
|
- Provide the standard **Application**, **File**, **Edit**, **View**,
|
|
**Window**, and **Help** menus. Include the normal macOS Application actions
|
|
(About, Settings, Services, Hide, Hide Others, Show All, and Quit) and Window
|
|
actions (Minimize, Zoom, Bring All to Front) with native behavior.
|
|
- Make File actions operate on real application concepts: New Session, Add/Open
|
|
Project, and Close Window. Make View actions cover actual available views,
|
|
including sidebar visibility and Enter Full Screen. Help opens the local help
|
|
or documentation surface. Do not add permanently inert placeholder items.
|
|
- The Edit menu is part of the chat acceptance criteria. Undo, Redo, Cut, Copy,
|
|
Paste, Paste and Match Style where supported, Delete, and Select All must act
|
|
on the current first responder. They must work in the chat composer and every
|
|
other editable text field; Copy and Select All must also work for selectable
|
|
transcript text.
|
|
- Prefer native first-responder selectors for native controls. Where an Iced
|
|
widget owns focus, bridge the same menu command into that focused widget so
|
|
menu clicks and `Command-Z`, `Shift-Command-Z`, `Command-X`, `Command-C`,
|
|
`Command-V`, and `Command-A` behave identically.
|
|
- Validate menu enabled state whenever focus or application state changes:
|
|
editing commands reflect the focused selection/clipboard and session/project
|
|
commands reflect whether their action is currently possible. Menu actions
|
|
must use the same application messages as toolbar/sidebar actions instead of
|
|
duplicating business logic.
|
|
- Preserve conventional macOS labels, grouping, keyboard shortcuts, and
|
|
accessibility metadata. Keep platform-specific menu integration behind the
|
|
macOS application shell so the chat and session core remain platform-neutral.
|
|
|
|
Exit criterion: with the composer focused, every standard Edit command works
|
|
from both the menu and its keyboard shortcut, including undo/redo across
|
|
multiple edits; transcript text can be copied; and Application, File, View,
|
|
Window, and Help expose functional, correctly enabled native actions.
|
|
|
|
### Dev Brain — Obsidian knowledge and memory
|
|
|
|
- Add an optional Dev Brain setting in Preferences. Use the native macOS folder
|
|
picker to select an existing Obsidian vault, allow the vault to be changed or
|
|
disconnected, and persist its canonical path plus a security-scoped bookmark
|
|
in SQLite so sandboxed builds can regain access after restart.
|
|
- Keep the vault global to the application rather than tied to one project or
|
|
session. The selected vault remains in its original location; never copy it
|
|
into Application Support or a project directory.
|
|
- Add one local `dev_brain` agent tool with the minimum useful operations:
|
|
search Markdown notes, read a selected note, and create or append a memory
|
|
note. Return vault-relative paths with results so the agent can cite where
|
|
knowledge came from.
|
|
- Restrict every operation to the explicitly selected vault. Resolve and check
|
|
paths before access, ignore `.obsidian` and hidden files, and do not provide
|
|
deletion or arbitrary overwrite operations initially.
|
|
- Let the agent retrieve relevant global knowledge during a turn and record
|
|
durable conclusions after a turn. Vault access is opt-in, does not preload the
|
|
model, and must not block sessions when no vault is configured or available.
|
|
- Reuse the existing local file/search implementation underneath the new vault
|
|
boundary instead of building a second filesystem stack.
|
|
|
|
Exit criterion: create a project session, ask the model to inspect and edit a
|
|
small fixture repository, review each tool result in the GUI, restart, and
|
|
continue from the persisted session. With Dev Brain enabled, store a memory in
|
|
one session and retrieve it with a cited vault-relative note path from a session
|
|
in another project after restarting the app.
|
|
|
|
## Phase 5 — A2UI rich local-chat surfaces
|
|
|
|
Status: **future extension; not part of the core milestone**. Start only after
|
|
ordinary local chat, structured tool calls, transcript persistence, and
|
|
cancellation are stable. Use the A2UI implementation in `../bDS2` as the
|
|
behavioral reference, especially its render-tool schemas, typed surface
|
|
normalization, inline message placement, chart geometry, and interaction state.
|
|
|
|
- Expose typed `render_card`, `render_chart`, `render_form`, `render_list`,
|
|
`render_metric`, `render_mindmap`, `render_table`, and `render_tabs` tools to
|
|
the local agent. Keep these tools presentation-only; they return structured
|
|
data and do not perform project mutations themselves.
|
|
- Support the bDS2 chart set initially: bar, stacked bar, line, area, pie,
|
|
donut, and heatmap. Render charts and mind maps with native Iced primitives
|
|
and SVG geometry inside the assistant message rather than adding a web view.
|
|
- Treat tool arguments as untrusted data. Validate type, nesting depth, row and
|
|
series counts, numeric values, labels, action names, and payload size before
|
|
rendering. Never accept model-authored HTML, JavaScript, styles, file URLs, or
|
|
arbitrary commands.
|
|
- Attach each validated surface to the assistant message/tool call that created
|
|
it and persist the versioned structured payload with the transcript. A
|
|
restarted session must reconstruct the same surface without asking the model
|
|
to emit it again.
|
|
- Build a surface only after its tool arguments are complete and valid while
|
|
allowing ordinary assistant text to continue streaming. Malformed or unknown
|
|
render payloads fall back to a bounded JSON/text representation and must not
|
|
fail the surrounding chat turn.
|
|
- Preserve local interaction state for forms and tabs. Dispatch form submits
|
|
and card actions through a small allowlist of application actions; any action
|
|
that mutates files, settings, or external state still follows the normal
|
|
approval path.
|
|
- Provide keyboard navigation, readable labels, theme-aware colors, and a text
|
|
or tabular fallback for chart data so the information remains accessible and
|
|
copyable.
|
|
- Keep A2UI local-chat-only initially. Do not extend the OpenAI-compatible HTTP
|
|
surface or invent a general plugin/component protocol until a real client
|
|
requires it.
|
|
|
|
Exit criterion: ask the local agent to present fixture data as each supported
|
|
surface type; verify a chart and nested tab surface render correctly, an
|
|
interactive form returns validated values, malformed payloads degrade safely,
|
|
and all surfaces survive application restart with their message history.
|
|
|
|
## Verification and packaging
|
|
|
|
Status: **partially implemented**. Formatting, Clippy, build, and unit-test
|
|
commit gates are documented in `AGENTS.md`; database migration/CRUD constraints
|
|
and the preferences shortcut have unit coverage. Model fixtures, integration
|
|
tests, and release packaging remain open.
|
|
|
|
- Keep unit coverage narrow: persistence round trips, format compatibility,
|
|
DS4 KV prefix/rewrite/store/load parity, cross-surface session binding,
|
|
request parsing, and tool-boundary checks.
|
|
- Reuse `../ds4` fixtures for prompt rendering, sampling, server streaming, and
|
|
KV correctness; run live Metal tests only when a supported GGUF is available.
|
|
- Preserve the DwarfStar/llama.cpp notices required by `../ds4/LICENSE` for
|
|
reused or adapted source and kernels.
|
|
- Produce a signed `.app` with the exact identifier above, then add hardened
|
|
runtime, entitlements, notarization, and update delivery as release work.
|
|
|
|
## Recommended next milestones
|
|
|
|
1. Finish the Flash session core as one larger slice: port the ratio-4 indexer
|
|
and indexed attention for full configured context, then add prefix sync and
|
|
compatible KV checkpoint save/load. **Implemented.**
|
|
2. Complete exact `ds4_server.c` communication parity as one tracked program,
|
|
beginning with exact layer-major/chunked/resumed prefill and KV continuation,
|
|
then porting every parser/route/state machine and its fixtures. Keep HTTP
|
|
conversation state transient except for opaque KV cache files. **In
|
|
progress; no route is considered complete until differential tests pass.**
|
|
3. Local transcript/KV persistence and the conversation UI are implemented;
|
|
port the agent tools after the shared endpoint lifecycle is stable.
|
|
4. Add the opt-in Dev Brain tool after local file/search boundaries and agent
|
|
approvals are stable.
|
|
5. Add bDS2-style A2UI render tools and native inline surfaces after the core
|
|
local chat/tool loop is complete.
|