Agent runtime: majordomo in-process, Ollama Cloud config, chat endpoint (#56) #70

Merged
steve merged 2 commits from feat/agent-runtime into main 2026-07-21 06:22:10 +00:00
2 Commits
Author SHA1 Message Date
steveandClaude Opus 4.8 91c6d66fa2 Address Gadfly review on the agent runtime
Build image / build-and-push (push) Successful in 7s
The first finding breaks this PR's central promise and is the reason the review
was worth running.

The run context carries a timeout. When it fired, WithChangeSet's recovery path
tried to record what had already committed using that SAME dead context — which
fails, losing the history for changes that really happened. So a timed-out turn
left its partial work un-undoable, while the user-facing message cheerfully said
"anything I'd already changed is in History". That message was a lie in exactly
the case it was written for.

Both recovery paths (WithChangeSet and RevertChangeSet) now commit with
context.WithoutCancel. The commonest reason those paths run at all is a
cancelled or timed-out context, so using it to write the record of what it did
was self-defeating. There's a test that cancels mid-turn and asserts the partial
work is recorded, marked partial, and revertible.

turnSummary sliced bytes, so a message whose 120th byte fell inside a multibyte
character stored invalid UTF-8 in the history summary. Not hypothetical for text
people type. Trimmed by runes now, with a test using emoji.

RecordAgentExchange's failure was logged and swallowed, directly under a comment
claiming the user would see that their turn wasn't saved. The turn itself
succeeded, so Done still goes out — but with a warning saying the exchange
wasn't saved and won't survive a reload, because a clean "done" followed by a
conversation that has forgotten it is the quieter lie. It also now records with
a detached context, since the commonest reason that write fails is the client
having gone away, and the exchange is worth keeping either way.

AgentRunID was declared, documented as an executus join, and never set by
anything. It's set now, from a per-run id that's also logged at run start, so a
row in the history list has a thread back to the run that produced it — and the
comment describes that rather than a dependency this repo doesn't have.

Turn.History was populated and never read, left over from the client-held
history design that persistence replaced. Removed.

The SSE plumbing moved out of the handler into openEventStream, so agentChat
reads like the other decode/call/encode handlers in the package.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01H3zbym8Doka2d7D48maSgZ
2026-07-21 02:21:36 -04:00
steveandClaude Opus 4.8 3f3a5b057c Agent runtime: majordomo in-process, Ollama Cloud config, chat endpoint (#56)
Build image / build-and-push (push) Successful in 25s
Gadfly review (reusable) / review (pull_request) Canceled after 5m56s
Adversarial Review (Gadfly) / review (pull_request) Canceled after 5m56s
Everything below the run loop already existed. This is the thing that runs a
model.

The build tag is gone, deliberately. internal/agent's doc comment promised two
separations — cmd/pansy not importing the package, and the tool wiring behind
//go:build majordomo — and both have been rewritten rather than left as a stale
aspiration. A tag that keeps the agent out of the binary only earns its keep if
you would ever ship a build without the agent, and the agent is the point;
keeping it meant an untagged CI that never compiled the code that matters.
majordomo is a real dependency now, resolved from the Gitea instance as a
pseudo-version with no replace directive, so the Docker build (which has no
sibling checkout) resolves it the same way this machine does. It is stdlib-first
and pure Go, so CGO_ENABLED=0 and the single static binary survive.

A TURN IS ONE CHANGE SET. That is the whole reason acting without a confirmation
prompt is defensible: "empty the garlic bed and plant cucumbers" is one object
edit and a dozen planting inserts, and it has to undo as one action rather than
thirteen. The scope is opened even for a turn that turns out to be a question,
because a change set with no revisions is never written — so asking costs
nothing and history isn't littered with empty entries.

The model spec goes to majordomo.Parse verbatim. That grammar, including
comma-separated failover chains, is majordomo's; re-implementing any of it here
would only mean two places to update when it grows. The key needs a bridge
though: majordomo's ollama-cloud preset reads OLLAMA_API_KEY while pansy (like
gadfly) is configured with OLLAMA_CLOUD_API_KEY, so the provider is registered
explicitly on a private registry rather than depending on ambient environment.

Runs are bounded by a step cap, a timeout and majordomo's loop guards. This is
loop safety, not cost control — pansy is a personal tool and spend caps are
explicitly not a v2 concern. A capped run does NOT fail: it kept whatever it
managed to do, that work is recorded and undoable, and the reply says it stopped
early rather than going silent.

The chat endpoint streams. A turn that clears a bed and replants it makes a
dozen tool calls over tens of seconds, and without streaming that is a long
silence followed by everything at once — which reads as a hang, and defeats a
design that rests on watching the canvas change as it happens.

Conversations persist per (user, garden). Client-held history would be lost on a
refresh, which is exactly when someone reloads to check whether the agent's
change landed. Only the user/assistant TEXT is stored, not the model's full
transcript: continuity needs what was said and what came back, and replaying a
stored tool call would replay a decision made against a garden that has since
moved on. It also keeps majordomo's message shape out of the schema.

An instance with no key starts, serves the app, and doesn't advertise the agent
— the routes aren't registered at all, the same shape as OIDC 404ing when
unconfigured. A configured-but-unresolvable model logs and disables the
assistant rather than refusing to boot: a garden planner that won't start
because of a chat feature is worse than one without chat.

Tool refusals reach the model as tool results it can explain, not 500s. The ACL
story only works if it can narrate the refusal.

Closes #56

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01H3zbym8Doka2d7D48maSgZ
2026-07-21 02:15:22 -04:00