# Request priority A GPU is a size-1 resource: one job at a time. Without priorities, a single long job monopolises the box for its whole duration and every interactive request queues behind it. A 16-minute render blocking a chat completion is uncomfortable; a 12-hour batch render blocking everything for a day is not workable. Callers declare intent with a header, and the scheduler dispatches by score. ## The header ``` X-LlamaSwap-Priority: 0 # normal — the default X-LlamaSwap-Priority: 100 # interactive, a human is waiting X-LlamaSwap-Priority: -100 # batch, nobody is watching ``` Priority is a **signed integer**; higher is more urgent. An absent or unparseable header means `0`. Values are not clamped — llama-swap trusts the caller. Named aliases resolve to numbers for convenience, but **the number is the interface**: callers may send any value. | alias | value | | ------------- | -----: | | `interactive` | `100` | | `normal` | `0` | | `batch` | `-100` | Priority is a property of the *caller's intent*, not of the model — the same model serves both an interactive request and a batch job — which is why it travels as a header rather than as per-model config. It also leaves the OpenAI-compatible request body untouched. ```bash curl http://localhost:8080/v1/chat/completions \ -H 'Content-Type: application/json' \ -H 'X-LlamaSwap-Priority: interactive' \ -d '{"model":"qwen3","messages":[{"role":"user","content":"hi"}]}' ``` ## Bands and offsets The three anchors sit 100 apart. That spacing is deliberate: it leaves room for a caller to add a small per-user offset without ever crossing a band. | tier | offset | | ---- | -----: | | free | `+0` | | pro | `+1` | | max | `+2` | A max member's batch job is `-100 + 2 = -98`: ahead of pro's `-99` and free's `-100`, and still **below** any normal request at `0`. A tier bonus breaks ties *inside* a band and can never promote across one. > **Invariant:** band spacing (100) must stay far larger than the largest offset > a caller adds plus `swapAffinityBonus`. Otherwise a max member's batch work > could outrank a free member's interactive request. ## How the serial scheduler dispatches Scheduling is **non-preemptive**. A generation cannot be interrupted mid-sample without discarding the work, so priority applies at *dispatch* time: when a job finishes, the best-scoring waiting request goes next. ``` score = X-LlamaSwap-Priority + swap affinity + aging ``` | term | range | crosses bands? | | --------------------- | ------- | ----------------------- | | **request priority** | any | — it *is* the band | | **swap affinity** | `0..99` | **never** | | **aging** | `0..∞` | **yes, deliberately** | Equal scores keep arrival order. **Swap affinity** is a bounded bonus for a request that can run without a model swap. Cold loads are expensive, so this keeps a run of same-model requests together instead of forcing a reload on every dispatch. Because it is bounded well below the band spacing, it can only break near-ties — it can never promote batch work above interactive. **Aging** is `waited_seconds / agingDivisor`, and is unbounded on purpose: a batch job that has waited long enough *should* eventually beat normal traffic, or it starves. At the default divisor of 60 (one point per minute), a batch job at `-100` overtakes normal traffic after ~100 minutes of waiting. That asymmetry is the design: **bounded offsets express policy, unbounded aging prevents starvation.** ### What priority does not fix Because dispatch is non-preemptive, an interactive request can still wait **up to one full job duration**. The complementary lever is client-side: a batch producer should submit **one unit at a time** and re-queue, so gaps are frequent. For a long render that means one shot per submission, not one 12-hour chain. Priority is also not fair-share. If one caller submits 200 batch jobs they all sit at the same priority, and llama-swap will work through them in order. Per-user fairness belongs in the client that submits the work. ## Configuration ```yaml routing: scheduler: use: serial settings: serial: # Seconds of waiting per priority point. 0 disables aging. agingDivisor: 60 # Bonus for a request needing no model swap. 0..99; 0 disables it. swapAffinityBonus: 10 ``` Setting `swapAffinityBonus: 0` gives strict priority-then-arrival order — with uniform priorities that is exact FIFO. The `fifo` scheduler honours the header too, adding it to the model's configured priority: a model at `10` with a `batch` header queues at `-90`. ## Observability `GET /metrics` exports the queue in Prometheus text format, labelled by band (`interactive`, `normal`, `batch`): | metric | type | meaning | | --------------------------------------------- | ------- | ---------------------------------------------------- | | `llamaswap_scheduler_queue_depth` | gauge | requests waiting | | `llamaswap_scheduler_queue_oldest_wait_seconds`| gauge | how long the longest-waiting request has waited | | `llamaswap_scheduler_dispatched_total` | counter | requests dispatched | | `llamaswap_scheduler_reorders_total` | counter | dispatches where `aging` or `swap_affinity` changed the pick | The reorder counters are what make `agingDivisor` and `swapAffinityBonus` tunable rather than guesswork: they say how often each term actually changed which request went next. Every dispatched request also records its decision in the activity log: | key | meaning | | ---------------------- | ------------------------------------------ | | `serial_priority` | the caller's priority | | `serial_band` | which band it fell in | | `serial_score` | the score it won with | | `serial_queue_wait_ms` | how long it waited before dispatch | That is what answers "why did my request take 20 minutes" after the fact.