The GPU is a size-1 resource, so a single long job monopolises the box for its
whole duration and every interactive request queues behind it. Callers can now
declare intent with an X-LlamaSwap-Priority header and the serial scheduler
dispatches by score instead of by arrival.
- X-LlamaSwap-Priority: signed integer, 0 default, absent/unparseable means 0.
interactive/normal/batch aliases resolve to +100/0/-100. Values are not
clamped: the caller composes band and any per-user offset itself.
- serial dispatch score = priority + swap affinity + aging. Bands sit 100 apart
so a small caller offset orders work inside a band without crossing one;
aging is unbounded so low-priority work cannot starve.
- routing.scheduler.settings.serial.{agingDivisor,swapAffinityBonus}, defaulting
to 60s/point and +10. swapAffinityBonus is capped at 99 so it can never
promote a request into the next band.
- fifo adds the header to its per-model priority, so the header is not silently
ignored under that scheduler.
- /metrics exports per-band queue depth, oldest wait and dispatch counts, plus
counters for how often aging or swap affinity changed the pick. Each request
records its priority, band, score and queue wait in the activity log.
Note swapAffinityBonus defaults to 10, so equal-priority requests for the
already-loaded model now run before older requests that need a swap. Set it to
0 for the previous strict arrival order.
fixes #9
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01WUyhZBgv8BBCC5MduX88gE
58 lines
2.0 KiB
Go
58 lines
2.0 KiB
Go
package router
|
|
|
|
import (
|
|
"net/http"
|
|
"time"
|
|
|
|
"github.com/mostlygeek/llama-swap/internal/logmon"
|
|
"github.com/mostlygeek/llama-swap/internal/process"
|
|
"github.com/mostlygeek/llama-swap/internal/router/scheduler"
|
|
"github.com/mostlygeek/llama-swap/internal/shared"
|
|
)
|
|
|
|
var (
|
|
ErrNoRouterFound = shared.ErrNoRouterFound
|
|
ErrNoPeerModelFound = shared.ErrNoPeerModelFound
|
|
ErrNoLocalModelFound = shared.ErrNoLocalModelFound
|
|
)
|
|
|
|
type Router interface {
|
|
// Shutdown blocks until the router has shutdown returning nil
|
|
// when the router has shutdown successfully.
|
|
//
|
|
// timeout controls how long to wait for inflight requests to finish. After
|
|
// the timeout all inflight requests will be cancelled.
|
|
Shutdown(timeout time.Duration) error
|
|
|
|
// ServeHTTP implements the http.Handler and requests coming in will
|
|
// trigger any model swapping and routing logic.
|
|
ServeHTTP(http.ResponseWriter, *http.Request)
|
|
|
|
// Handles reports whether this router can serve requests for the given model.
|
|
Handles(model string) bool
|
|
}
|
|
|
|
// LocalRouter is a Router backed by local processes whose state can be
|
|
// inspected and which can be individually stopped. Peer routers, which only
|
|
// forward to remote hosts, do not implement it.
|
|
type LocalRouter interface {
|
|
Router
|
|
|
|
// RunningModels returns the current state of every process that is not
|
|
// stopped or shut down, keyed by model ID.
|
|
RunningModels() map[string]process.ProcessState
|
|
|
|
// Unload stops the named models, or every running model when none are
|
|
// named. It blocks until each targeted process has stopped.
|
|
Unload(timeout time.Duration, models ...string)
|
|
|
|
// ProcessLogger returns the log monitor for the named model's process.
|
|
// modelID must be a real (non-alias) config key. Returns false when the
|
|
// model is not known to this router.
|
|
ProcessLogger(modelID string) (*logmon.Monitor, bool)
|
|
|
|
// SchedulerStats returns a snapshot of the scheduler's queue for metrics.
|
|
// ok is false when the configured scheduler does not report stats.
|
|
SchedulerStats() (scheduler.QueueStats, bool)
|
|
}
|