3.8 KiB
Grid-agent LLM transport and tool loop
The grid agent talks to one operator-supplied OpenAI-compatible chat-completion
endpoint. llm.endpoint_url is the complete request URL and llm.api_key is
the bearer credential. The client sends POST to that exact URL. It does not
append a path, select a provider, discover models, send a model field, or retry
against another service. Endpoint routing and model selection remain operator
responsibilities.
Compatibility envelope
Requests contain only the normalized messages and tools members. Message
roles map to system, user, assistant, and tool. Content is an array of
text or image_url parts. Tool definitions use the standard function name,
description, and JSON-schema parameters envelope. Tool observations carry the
original tool_call_id. The client accepts the first response choice, optional
text, function tool calls, and optional token usage. Unknown response members
are ignored; missing required members, unsupported empty choices, malformed
JSON, duplicate call IDs, and oversized values return typed errors.
The HTTP client uses Rustls and performs no automatic content decompression
because no compression feature is enabled. Redirect following is disabled, so
a bearer credential can never be forwarded to either a same-origin or
cross-origin redirect target. Connect, whole-request, response-idle, pool-idle,
and total elapsed timeouts are independently bounded. Prompt bytes, response
bytes, concurrent requests, retry count, and retry delay also have validated
hard ceilings. Only transport failures, timeouts, and HTTP 408, 425, 429, 500,
502, 503, or 504 are retryable. Retry-After seconds are honored only up to the
configured delay ceiling; otherwise bounded deterministic jitter is used.
Tool-loop safety
ToolLoop validates every registered schema before use. A proposed call must
have a unique session call ID, a registered name, valid JSON arguments, and
arguments matching that registered schema before it reaches ToolExecutor.
Unknown names, malformed arguments, and schema mismatches become bounded tool
observations for the next model turn and never call the executor. Tool
executions are sequential, making the simultaneous execution ceiling one.
Turn count, calls per turn and session, history messages, history bytes, wall-clock time, and collected usage records are bounded. When history exceeds its configured envelope, an injected deterministic summarizer may compact it. A failed summarizer inserts a fixed bounded truncation marker and retains recent context. An ambiguous mutating result terminates the loop immediately; it is never retried or turned into another model request.
Shutdown, operator cancellation, disconnect, session expiry, and avatar-session
replacement use cancellation plus SessionGeneration. Superseding a generation
cancels its in-flight HTTP request or executor and prevents a late result from
starting another action. Request and correlation IDs are deterministic and the
completion exposes token usage, latency, and attempt count. The final returned
message is the model-authored action summary. The wire mapping has no reasoning
or chain-of-thought field and does not retain or expose hidden reasoning.
Focused verification
The llm_transport integration suite runs only against bounded loopback fake
endpoints. It covers exact URL and authorization behavior, secret redaction,
fragmented responses, images and tool schemas, malformed and oversized input,
redirect refusal, transient-only retry, cancellation and timeout races,
concurrency, complete tool round trips, invalid-call observations, endless
loops, duplicate IDs, ambiguous mutation, supersession, and safe history
compaction.
cargo test --locked -p metacrate-grid-agent --test llm_transport
cargo clippy --locked -p metacrate-grid-agent --all-targets -- -D warnings