Files
pansy/internal/agent/runtime.go
T
steveandClaude Opus 4.8 41c28592a2
Gadfly review (reusable) / review (pull_request) Failing after 30s
Adversarial Review (Gadfly) / review (pull_request) Failing after 30s
Build image / build-and-push (push) Successful in 14s
Admin-gated Settings: runtime model selection, enforce is_admin (#79)
Moves the agent model out of env-only config into an admin-editable Settings
section, and enforces is_admin for the first time — it has been in the schema
since migration 0001, plumbed to the client, and checked nowhere.

Backend:
- Migration 0010: instance_settings, a single-row (CHECK id=1) table — pansy's
  first instance-level state. Holds agent_model ('' = inherit env) and
  agent_enabled (NULL = inherit env), version-guarded like every mutable row.
  SECRETS STAY IN ENV: OLLAMA_CLOUD_API_KEY is never stored here.
- requireAdmin at the service seam (authoritative) plus a cheap middleware
  early-403. Non-admin gets 403, not 404 — settings existence isn't masked.
- EffectiveAgent resolves DB-over-env (model, enabled); key always from env.
- The live Runner is hot-swapped, not built once. agentHolder holds it behind
  an atomic.Pointer; the chat routes are now registered UNCONDITIONALLY and
  nil-check agent.get(), so a settings change turns the assistant on/off/onto a
  new model with no restart and no race against in-flight readers. /capabilities
  reads the pointer, so it reports what's live, not what booted.
- internal/agentmodel is a new leaf package holding the one place that knows how
  to turn a spec into a model. Both agent (to run) and service (to validate a
  spec before storing it) import it; it can't live in agent, which imports
  service. Settings PATCH validates the spec via Parse, so a typo is a 400 now
  rather than a broken assistant on the next turn.

Frontend:
- /settings route (admin guard), a Settings page (model field, tri-state
  enabled, live status), nav link shown only to admins.
- useCapabilities drops staleTime:Infinity — the assistant can now change under
  a running page — and the settings save invalidates it.

Contract change: chat routes always exist, so "assistant off" is a runtime 503
+ capabilities:false, not a missing route. Updated the test that asserted the
old shape.

Verified live against the built binary: disable flips capabilities to false and
logs it; re-enable with a new model swaps it back; a bad spec is rejected 400;
the setting persists across a restart. Swap is race-clean under `go test -race`.

Docs: README (precedence + key-stays-in-env), DESIGN (decision + routes),
CLAUDE (don't re-add conditional route registration; key never in the DB).

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01H3zbym8Doka2d7D48maSgZ
2026-07-21 20:27:36 -04:00

231 lines
8.4 KiB
Go

package agent
import (
"context"
"crypto/rand"
"encoding/hex"
"errors"
"fmt"
"log/slog"
"strings"
"time"
"gitea.stevedudenhoeffer.com/steve/majordomo/agent"
"gitea.stevedudenhoeffer.com/steve/majordomo/llm"
"gitea.stevedudenhoeffer.com/steve/pansy/internal/agentmodel"
"gitea.stevedudenhoeffer.com/steve/pansy/internal/domain"
"gitea.stevedudenhoeffer.com/steve/pansy/internal/service"
)
// Loop bounds. These exist so a stuck run terminates, NOT to control spend —
// pansy is a personal tool and multi-tenant cost control is explicitly not a
// concern here. A model that gets wedged calling describe_garden forever should
// stop on its own, and a hung upstream shouldn't hold a connection open all day.
const (
maxSteps = 24
// A turn that legitimately fills several beds does a lot of round trips, so
// this is generous; it is a backstop, not a budget.
runTimeout = 4 * time.Minute
// Successive all-error steps, and identical repeated calls, that end a run.
maxConsecutiveToolErrors = 4
maxSameCallRepeats = 3
)
// Runner drives a model over pansy's toolbox. One per process; Run is safe to
// call concurrently.
type Runner struct {
svc *service.Service
model llm.Model
}
// NewRunner resolves modelSpec against pansy's registry and returns a Runner, or
// an error if the assistant can't be offered. Callers should treat an error as
// "no assistant" rather than a startup failure — an instance with no key must
// still serve the app.
//
// It takes the key and spec explicitly rather than a *config.Config so the same
// constructor serves both boot (from env) and a runtime settings change (from
// the DB) — the Runner has no idea which one configured it.
func NewRunner(svc *service.Service, apiKey, modelSpec string) (*Runner, error) {
if apiKey == "" || strings.TrimSpace(modelSpec) == "" {
return nil, errors.New("agent: not configured")
}
model, err := agentmodel.Resolve(apiKey, modelSpec)
if err != nil {
return nil, err
}
return &Runner{svc: svc, model: model}, nil
}
// Turn is the outcome of one exchange.
type Turn struct {
// Reply is what to show the user.
Reply string `json:"reply"`
// ChangeSetID is the change set this turn produced, if it changed anything —
// the handle the UI needs to offer Undo.
ChangeSetID *int64 `json:"changeSetId,omitempty"`
// Steps is how many model round trips it took.
Steps int `json:"steps"`
// Truncated is set when the run hit its step cap rather than finishing.
Truncated bool `json:"truncated,omitempty"`
}
// Run executes one turn against a garden, as actorID.
//
// The whole turn runs inside ONE change set, so everything the model did undoes
// together. That is what makes acting without a confirmation prompt defensible.
// The scope is opened even for a turn that turns out to be a question — a change
// set with no revisions is never written, so asking costs nothing.
func (r *Runner) Run(ctx context.Context, actorID, gardenID int64, message string, history []llm.Message, onStep func(agent.Step)) (*Turn, error) {
message = strings.TrimSpace(message)
if message == "" {
return nil, domain.ErrInvalidInput
}
ctx, cancel := context.WithTimeout(ctx, runTimeout)
defer cancel()
garden, err := r.svc.GetGarden(ctx, actorID, gardenID)
if err != nil {
return nil, err
}
// An id for this run, stamped on the change set so a row in the history list
// can be matched to the log lines that produced it. Without it, "the agent
// did something odd on Tuesday" has no thread back to what it was thinking.
runID := newRunID()
slog.Info("agent: run start", "run", runID, "garden", gardenID, "actor", actorID)
var (
result *agent.Result
runErr error
truncErr bool
)
changeSet, err := r.svc.WithChangeSet(ctx, actorID, gardenID, service.ChangeSetOptions{
Source: domain.SourceAgent,
Summary: turnSummary(message),
AgentRunID: &runID,
}, func(ctx context.Context) error {
box := NewToolbox(r.svc, actorID)
a := agent.New(r.model, systemPrompt(garden),
agent.WithMaxSteps(maxSteps),
agent.WithToolErrorLimits(maxConsecutiveToolErrors, maxSameCallRepeats),
)
a.AddToolbox(box)
opts := []agent.RunOption{agent.WithHistory(history)}
if onStep != nil {
opts = append(opts, agent.OnStep(onStep))
}
result, runErr = a.Run(ctx, message, opts...)
// A run that ran out of steps still DID things, and those things must be
// recorded and undoable. So the loop-guard errors don't fail the scope —
// they're reported to the user instead.
if runErr != nil && isLoopLimit(runErr) {
truncErr = true
return nil
}
return runErr
})
if err != nil {
// WithChangeSet records whatever committed before failing, so the partial
// work is undoable even though the turn errored.
return nil, err
}
turn := &Turn{Truncated: truncErr}
if changeSet != nil {
turn.ChangeSetID = &changeSet.ID
}
if result != nil {
turn.Reply = result.Output
turn.Steps = len(result.Steps)
}
if turn.Reply == "" {
turn.Reply = fallbackReply(turn)
}
return turn, nil
}
// isLoopLimit reports whether an error is one of majordomo's loop guards firing
// rather than a genuine failure. Those runs have a partial result worth keeping.
func isLoopLimit(err error) bool {
return errors.Is(err, agent.ErrMaxSteps) || errors.Is(err, agent.ErrToolLoop)
}
// fallbackReply covers a run that finished with no text — a model that made its
// last tool call and then stopped. Silence reads as a failure, so say what
// happened.
func fallbackReply(t *Turn) string {
switch {
case t.Truncated:
return "I stopped partway through — that turned into more steps than I should take in one go. " +
"Have a look at what changed, and tell me what to do next."
case t.ChangeSetID != nil:
return "Done — have a look at the canvas."
default:
return "I didn't change anything."
}
}
// newRunID returns a short random identifier for one run.
func newRunID() string {
var b [8]byte
if _, err := rand.Read(b[:]); err != nil {
// The id is for correlating logs, not for security. A clock-based
// fallback is worse than random and better than an empty string.
return fmt.Sprintf("t%d", time.Now().UnixNano())
}
return hex.EncodeToString(b[:])
}
// turnSummary is what the history list shows for this turn. The user's own words
// are the most useful label available, trimmed to fit a list row.
//
// Trimmed by RUNES, not bytes: slicing a byte offset would cut a multibyte
// character in half and store invalid UTF-8 in the summary — which is not a
// hypothetical for text people type.
func turnSummary(message string) string {
const max = 120
s := strings.Join(strings.Fields(message), " ")
runes := []rune(s)
if len(runes) > max {
s = strings.TrimSpace(string(runes[:max])) + "…"
}
return s
}
// systemPrompt gives the model the conventions it cannot infer.
//
// The compass convention in particular is not guessable: -y is north because
// screen y grows downward, and a model that assumes otherwise plants the south
// half when asked for the north one.
func systemPrompt(g *domain.Garden) string {
units := "metric — all measurements are centimeters"
if g.UnitPref == domain.UnitImperial {
units = "imperial for display, but every measurement you send or receive is in CENTIMETERS"
}
return fmt.Sprintf(`You are pansy's garden assistant. You help plan and edit a real garden by calling tools.
The garden you are working on is %q (id %d), %.0f x %.0f cm. The user's units are %s.
Conventions you cannot guess and must not assume:
- Positions in a garden are centimeters from its top-left corner: x grows east, y grows SOUTH.
- Inside an object (a bed), positions are relative to that object's CENTER, and -y is NORTH.
So the north half of a bed is negative y. Getting this backwards plants the wrong end.
- Objects and plantings are version-guarded. Use the version from describe_garden when editing.
How to work:
- Start from describe_garden to see what is actually there. Do not guess ids.
- Use find_plant to turn a plant name into an id. If it returns several candidates,
pick the one that matches what the user said, or ask them which they meant.
- To replant a bed with something else: clear_object, then fill_region with region "all".
- When a tool refuses (for example, the user only has view access to this garden),
explain what happened in plain words. Do not retry it.
When you are done, say briefly what you changed — the user is watching the canvas
and wants to know what to look at. If you changed nothing, say that too.`,
g.Name, g.ID, g.WidthCM, g.HeightCM, units)
}