T

steve 27f196d333 feat: add Ollama target client, model poller, and native passthrough

Phase 2 of foreman: the daemon now acts as a transparent Ollama proxy.

- internal/ollama: Client interface and HTTP implementation for chat
  (streaming + non-streaming), embed, tags, ps with auth forwarding,
  NDJSON streaming via bufio.Scanner, and connection vs HTTP error
  classification via custom error types.
- internal/ollama: ModelInventory with background poller for /api/tags
  and /api/ps, degraded mode on target unreachable with model retention,
  automatic recovery on reconnect.
- internal/server: Passthrough routes (/api/chat, /api/tags, /api/ps,
  /api/embed, /api/embeddings) with model validation, chat serialization
  gate (capacity-1 channel), concurrent embedding bypass (ADR-0013),
  NDJSON streaming with per-chunk flush, and degraded health reporting.
- cmd/foreman: Full serve wiring with Ollama client, poller goroutine,
  embedder warmup (keep_alive:-1), and signal-based shutdown.

The Mac is now usable as a go-llm target through foreman.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-05-23 18:07:33 -04:00

.gitea/workflows

feat: scaffold project with config, store, health endpoint, CI, and Dockerfile

2026-05-23 17:58:36 -04:00

cmd/foreman

feat: add Ollama target client, model poller, and native passthrough

2026-05-23 18:07:33 -04:00

docs/adr

initial commit

2026-05-23 16:41:20 -04:00

internal

feat: add Ollama target client, model poller, and native passthrough

2026-05-23 18:07:33 -04:00

prompts

add initial prompts

2026-05-23 16:51:19 -04:00

.env.example

feat: scaffold project with config, store, health endpoint, CI, and Dockerfile

2026-05-23 17:58:36 -04:00

.gitignore

feat: scaffold project with config, store, health endpoint, CI, and Dockerfile

2026-05-23 17:58:36 -04:00

CLAUDE.md

initial commit

2026-05-23 16:41:20 -04:00

Dockerfile

feat: scaffold project with config, store, health endpoint, CI, and Dockerfile

2026-05-23 17:58:36 -04:00

go.mod

feat: scaffold project with config, store, health endpoint, CI, and Dockerfile

2026-05-23 17:58:36 -04:00

go.sum

feat: scaffold project with config, store, health endpoint, CI, and Dockerfile

2026-05-23 17:58:36 -04:00

progress.md

feat: add Ollama target client, model poller, and native passthrough

2026-05-23 18:07:33 -04:00

README.md

feat: scaffold project with config, store, health endpoint, CI, and Dockerfile

2026-05-23 17:58:36 -04:00

README.md

foreman

A small, always-on Go daemon that fronts one Ollama target. It turns a single Ollama instance into a queued, observable job endpoint: it polls the target's installed models, serializes work through the target (managing model swaps), assigns every job an ID, and reports progress via webhooks.

On the wire it speaks native Ollama, so it doubles as a drop-in go-llm target.

Quickstart

# Set the required Ollama target URL
export FOREMAN_OLLAMA_URL=http://mac.tail:11434

# Run directly
go run ./cmd/foreman serve

# Or build and run
go build -o foreman ./cmd/foreman
./foreman serve

Docker

docker build -t foreman .
docker run -e FOREMAN_OLLAMA_URL=http://mac.tail:11434 -p 8080:8080 foreman

Configuration

All configuration is via environment variables, namespaced under FOREMAN_*. See .env.example for the full list.

Variable	Default	Description
`FOREMAN_ADDR`	`:8080`	Listen address
`FOREMAN_OLLAMA_URL`	(required)	Ollama target base URL
`FOREMAN_OLLAMA_TOKEN`	(empty)	Bearer token sent to the target
`FOREMAN_TOKEN`	(empty)	Bearer token callers must present
`FOREMAN_EMBED_MODEL`	(empty)	Always-resident embedder model
`FOREMAN_DB_PATH`	`foreman.db`	SQLite database path
`FOREMAN_POLL_INTERVAL`	`30s`	Target model poll interval
`FOREMAN_WEBHOOK_SECRET`	(empty)	HMAC key for webhook signing

Health check

curl http://localhost:8080/healthz
# {"status":"ok","degraded":false}

Architecture

See docs/adr/ for design decisions. Key points:

One daemon per Ollama target (ADR-0001)
SQLite-backed durable job queue in WAL mode (ADR-0008)
Single worker loop with drain-by-model scheduling (ADR-0009)
Native Ollama passthrough + async /jobs surface (ADR-0003, ADR-0004)
Embeddings bypass the queue entirely (ADR-0013)

Description

🪓 Small always-on Go daemon that fronts one Ollama target — turns it into a queued, observable job endpoint (model-swap serialization, job IDs, progress webhooks). Speaks native Ollama on the wire, so it's a drop-in target for any Ollama client.

Readme MIT 244 KiB