Stabilize OpenSim agent runtime and Mentra integration
Some checks failed
CI / rust-skia (Rust only) (push) Has been cancelled
CI / required (push) Has been cancelled

This commit is contained in:
2026-08-22 10:44:09 +02:00
parent 2f5f03ac6f
commit 0dd2ca5824
28 changed files with 1765 additions and 2308 deletions

View File

@@ -31,8 +31,8 @@ observable receivers.
| Interaction ingress / observations | `InteractionHandle` | 8,192 each hard, lower queue configuration | bounded admission; lifecycle uses nonblocking watch state |
| Per-avatar interaction FIFO | `InteractionCoordinator` | 4,096 senders / 64 messages each hard, lower `interaction` limits | fair ready queue, one active request per avatar/channel |
| Interaction inference / outbound | generation tasks and channel rate limiters | 64 concurrent hard; 1,023 bytes per grid part | timeout/cancellation fencing; independent public/IM pacing |
| LLM request slots | shared `LlmClient` semaphore | 256 hard / configured concurrent requests | async acquire or cancellation |
| Reasoning/tool session | `ToolLoop` caller | 32 turns / 256 calls hard, with lower configured limits | total timeout, cancellation, or supersession |
| LLM conversation runtime | avatar-scoped Mentra agent | configured model/tool budgets and 2-minute default wall time | SSE cancellation, compaction, or supersession |
| Autonomous approval review | fresh volatile Mentra runtime | one model request, no tools/history/memory | exact ALLOW/DENY; every error fails closed |
| Policy tools / approvals / schedules | `PolicyGateway` mutex | 64 tools / 4,096 approval records / 1,024 scheduler grants hard | deny before opaque authorization |
| Principal/global resource budget | `PolicyGateway` time window | validated calls, zero L$, upload, inventory, movement, and build ceilings | atomic charge or stable denial |
| Authorized avatars | immutable `AgentConfig` set | 1,024 hard ceiling, lower configured limit | malformed, nil, duplicate, and wildcard input rejected |
@@ -84,10 +84,10 @@ cleanup, fencing old events and late LLM/tool results. See
non-forgeable `AuthorizedAction` produced by `PolicyGateway`. Raw calls and
caller-created decisions are not accepted. This issue supplies no live
mutation implementation.
- LLM traffic crosses one exact configured URL through `LlmClient`. Redirects
are refused, response bodies are bounded while streaming, bearer secrets are
redacted, and provider/model discovery does not exist. Proposed calls cross
`ToolExecutor` only after registered-name and schema validation.
- LLM traffic crosses the configured provider base through Mentra's Responses
SSE runtime. MetaCrate supplies only endpoint, credential, model, permitted
tools, and run bounds. Proposed calls cross `PolicyToolExecutor` only after
Mentra schema handling and MetaCrate's independent policy evaluation.
- Conversation context is keyed by immutable avatar UUID and either public chat
or direct IM. Group channels are not representable. The LLM projection can
retrieve only one exact key, and recovered/untrusted summaries remain user-role
@@ -113,11 +113,12 @@ typed config/events/policy boundaries live-grid feature boundary
| |
+-----------------> libremetaverse-types +--> libremetaverse::GridClient
avatar session -> bounded ToolLoop -> exact-endpoint LlmClient
avatar session -> persistent Mentra agent -> Responses SSE / compaction / memory
|
+-> PolicyToolExecutor -> PolicyGateway
|
+-> AuthorizedToolBackend
| |
| +-> AuthorizedToolBackend
+-> one-shot Mentra safety reviewer when required
```
The package has no build script or direct native dependency. The focused

View File

@@ -1,68 +1,91 @@
# Grid-agent LLM transport and tool loop
# Grid-agent LLM runtime
The grid agent talks to one operator-supplied OpenAI-compatible chat-completion
endpoint. `llm.endpoint_url` is the complete request URL and `llm.api_key` is
the bearer credential. The client sends `POST` to that exact URL. It does not
append a path, select a provider, discover models, send a model field, or retry
against another service. Endpoint routing and model selection remain operator
responsibilities.
MetaCrate delegates the conversation runtime to Mentra. `llm.endpoint_url` is
the provider base URL, `llm.api_key` is its bearer credential, and `llm.model`
is the model ID. Mentra owns the Responses request path, HTTP/SSE streaming,
message history, tool-call rounds, provider errors, cancellation, compaction,
and persisted runtime state. MetaCrate does not implement a second HTTP client
or parse SSE itself.
## Compatibility envelope
For an OpenAI-compatible provider, configure the base that Mentra can extend
with `v1/responses`. For example, a proxy base ending in `/go` is valid when
its Responses endpoint is `/go/v1/responses`; do not include that final path in
MetaCrate configuration. Endpoint behavior, redirects, response transport, and
provider retry semantics are Mentra responsibilities.
Requests contain only the normalized `messages` and `tools` members. Message
roles map to `system`, `user`, `assistant`, and `tool`. Content is an array of
`text` or `image_url` parts. Tool definitions use the standard function name,
description, and JSON-schema parameters envelope. Tool observations carry the
original `tool_call_id`. The client accepts the first response choice, optional
text, function tool calls, and optional token usage. Unknown response members
are ignored; missing required members, unsupported empty choices, malformed
JSON, duplicate call IDs, and oversized values return typed errors.
## Conversation and memory
The HTTP client uses Rustls and performs no automatic content decompression
because no compression feature is enabled. Redirect following is disabled, so
a bearer credential can never be forwarded to either a same-origin or
cross-origin redirect target. Connect, whole-request, response-idle, pool-idle,
and total elapsed timeouts are independently bounded. Prompt bytes, response
bytes, concurrent requests, retry count, and retry delay also have validated
hard ceilings. Only transport failures, timeouts, and HTTP 408, 425, 429, 500,
502, 503, or 504 are retryable. `Retry-After` seconds are honored only up to the
configured delay ceiling; otherwise bounded deterministic jitter is used.
The main system prompt is intentionally minimal:
## Tool-loop safety
> You are in an OpenSim virtual world. Use the available tools to act there.
`ToolLoop` validates every registered schema before use. A proposed call must
have a unique session call ID, a registered name, valid JSON arguments, and
arguments matching that registered schema before it reaches `ToolExecutor`.
Unknown names, malformed arguments, and schema mismatches become bounded tool
observations for the next model turn and never call the executor. Tool
executions are sequential, making the simultaneous execution ceiling one.
It contains no authorization claims. MetaCrate derives authority from the
authenticated sender UUID and channel, then gives Mentra only the permitted
tool profile. Every tool call is checked again by `PolicyGateway`; prompt text
cannot grant authority.
Turn count, calls per turn and session, history messages, history bytes,
wall-clock time, and collected usage records are bounded. When history exceeds its
configured envelope, an injected deterministic summarizer may compact it. A
failed summarizer inserts a fixed bounded truncation marker and retains recent
context. An ambiguous mutating result terminates the loop immediately; it is
never retried or turned into another model request.
Authorized IM agents have a stable avatar-scoped Mentra identity, so history,
compaction, and memory survive IM session changes and process restarts. Public
and unprivileged conversations remain session-scoped and receive neither
durable memory tools nor privileged grid tools. Authorized agents receive
Mentra's `memory_search`, `memory_pin`, and `memory_forget` tools. Automatic
memory injection is disabled because Mentra 0.18.3 appends recalled memory as a
new user turn after the current request; explicit memory tools preserve durable
learning without displacing the command being handled.
Shutdown, operator cancellation, disconnect, session expiry, and avatar-session
replacement use cancellation plus `SessionGeneration`. Superseding a generation
cancels its in-flight HTTP request or executor and prevents a late result from
starting another action. Request and correlation IDs are deterministic and the
completion exposes token usage, latency, and attempt count. The final returned
message is the model-authored action summary. The wire mapping has no reasoning
or chain-of-thought field and does not retain or expose hidden reasoning.
Runtime records use `HybridRuntimeStore`: conversation/runtime state is stored
in `runtime.sqlite`, with the associated Mentra memory store and transcript,
task, team, and workspace paths under the configured state directory.
## Focused verification
## Tools and autonomous safety review
The `llm_transport` integration suite runs only against bounded loopback fake
endpoints. It covers exact URL and authorization behavior, secret redaction,
fragmented responses, images and tool schemas, malformed and oversized input,
redirect refusal, transient-only retry, cancellation and timeout races,
concurrency, complete tool round trips, invalid-call observations, endless
loops, duplicate IDs, ambiguous mutation, supersession, and safe history
compaction.
Mentra receives the JSON schemas registered by MetaCrate. It validates and
orchestrates model tool calls; `PolicyToolExecutor` is the only bridge to grid
backends. The gateway independently binds authorization to the authenticated
principal, origin, exact canonical arguments, estimated resource cost, expiry,
and single execution.
An action above its approval-free risk threshold does not wait for a human
operator. MetaCrate starts a separate one-shot Mentra agent backed by a fresh
volatile store. That reviewer has no tools, history, or memory and receives
only the proposed tool, exact arguments, and resource estimate as untrusted
data. Only an explicit `ALLOW` verdict grants the existing single-use bound
approval; `DENY`, transport failure, an invalid verdict, timeout, or
cancellation fails closed. Human control-plane decisions remain an emergency
facility, not a requirement for autonomous operation.
## Images
Vision captures are real software-rendered viewport images. MetaCrate renders
the current scene, encodes the final frame directly as JPEG, and attaches it as
a Mentra image content block. The default 320x180 frame, quality, entity,
triangle, texture-fetch, decoded-pixel, byte, rate, and concurrency limits are
validated before runtime. JPEG avoids the much larger PNG payloads.
The current renderer handles legacy prim geometry, approximate texture colors,
simple avatars, and flat terrain/water. Viewer-grade mesh/sculpt rendering,
full UV materials, lighting, authoritative terrain, and model-controlled camera
tools are tracked separately.
## Bounds and cancellation
MetaCrate sets generous but finite run budgets on Mentra: model rounds, tool
calls, stored history, and total wall-clock time. The default interaction/model
window is two minutes. Grid disconnect, session replacement, expiry, operator
cancel, and shutdown bridge the generation cancellation token into Mentra.
Non-idempotent ambiguous mutations stop immediately and are never retried.
The provider proxy used in live verification could not combine replayed input
with `previous_response_id`, so MetaCrate selects Mentra's full-history
Responses mode for that provider. Mentra still owns the transport and history.
## Verification
The focused suite uses loopback Responses/SSE endpoints and covers streamed
text, images, tool rounds, compaction, persisted memory across restart, and
schema validation:
```sh
cargo test --locked -p metacrate-grid-agent --test llm_transport
cargo clippy --locked -p metacrate-grid-agent --all-targets -- -D warnings
cargo clippy --locked -p metacrate-grid-agent --all-targets --all-features -- -D warnings
```

View File

@@ -55,19 +55,20 @@ deployments should use absolute paths.
## Endpoint, grid identity, and authority
`llm.endpoint_url` is the exact OpenAI-compatible chat-completions URL. It is
not a base URL: MetaCrate does not append a path, discover models, select a
provider, or rewrite query parameters. `llm.api_key` is sent as the bearer key
only to that exact origin; redirects are refused. Set `llm.model` when the
endpoint requires a model name. These values live in the private `config.yml`.
`llm.endpoint_url` is the OpenAI-compatible provider base URL. Mentra owns the
Responses path, SSE stream, endpoint behavior, and conversation runtime. For a
proxy whose Responses endpoint is `/go/v1/responses`, configure the base ending
in `/go`, not the final request path. Set `llm.model` to the model ID and keep
`llm.api_key` in the private `config.yml`.
Live modes require `grid.login_url`, `grid.avatar_name`, and `grid.password`.
`authorized_avatar_uuids` contains exact
grid UUIDs, never display names. Text claiming an authorized identity grants no
authority. Public chat can request bounded informational work and safe public
LSL delivery; movement, teleport, building, roaming changes, and administration
remain policy-gated and require an authenticated authorized IM, operator action,
or a narrowly bound scheduler grant as documented in the policy matrix.
remain policy-gated and require an authenticated authorized IM or a narrowly
bound scheduler grant. Above-threshold actions receive an isolated one-shot LLM
safety review; they do not depend on a continuously present human operator.
## Configuration contract and migration
@@ -90,8 +91,8 @@ the operator's retention policy. JSON remains readable only for migration.
Mode, endpoints, credentials, authorization UUIDs, TLS, storage, queue/resource
limits, reconnect policy, behavior, and interaction settings are restart-only.
Runtime control can pause/resume autonomy, toggle the roaming job, decide an
existing approval, cancel an active action, expire/delete conversation state,
Runtime control can pause/resume autonomy, toggle the roaming job, make an
emergency decision on an existing approval, cancel an active action, expire/delete conversation state,
inject an operator message, reconnect, or shut down; it does not silently
rewrite the configuration. A future reloadable field must be explicitly added
to the versioned control/config contract.
@@ -203,8 +204,8 @@ start. Never run two service generations against one writable data directory.
- Authentication blocked: pause retries, verify login URL/avatar and rotate the
password file; never paste it into logs. Force reconnect after correction.
- LLM unavailable/rate limited: autonomy remains bounded; verify the exact URL,
firewall/DNS, and key. Multimodal rejection falls back to the textual scene
- LLM unavailable/rate limited: autonomy remains bounded; verify the provider
base URL, firewall/DNS, key, and model compatibility. Multimodal rejection falls back to the textual scene
summary without resending the large image.
- Maintenance/disconnect: allow generation fencing and bounded backoff. Stale
inference/mutation results are discarded; do not bypass reconnect controls.
@@ -224,12 +225,12 @@ start. Never run two service generations against one writable data directory.
The example records all current queue, message, conversation, tool, behavior,
reconnect, and interaction defaults. Important defaults include 256 grid events,
32 control commands, 512 observations, four concurrent inference requests,
16 tool calls, 512 active senders/sessions, a 1 MiB transport body, and bounded
10-second shutdown. Vision defaults to a 320x180 synthetic image with bounded
entities, triangles, texture work, PNG bytes, time, and concurrency.
16 tool calls, 512 active senders/sessions, a two-minute model window, and bounded
10-second shutdown. Vision defaults to a 320x180 software-rendered JPEG with bounded
entities, triangles, texture work, JPEG bytes, time, and concurrency.
Unsupported by design: arbitrary raw packets or agent-control flags, arbitrary
shell/subprocess execution, provider SDK/model discovery, remote plaintext
shell/subprocess execution, provider/model discovery outside Mentra, remote plaintext
control, unauthenticated mutation, unrestricted walking/teleport/touch/follow,
automatic config migration, persistence of viewport pixels, framebuffer/screen
capture, and treating untrusted grid/LLM text as instructions or authority.