llama-swap

Author	SHA1	Message	Date
Marcus	5b4beaceef	fix: ?no-history flag and improve /logs monitoring docs (#721 ) - improve logging documentation - small tweaks for edge case issues in upstream and log requests	2026-04-30 00:50:36 -07:00
Benson Wong	fd3c28ffc5	Refactor Activity Page (#710 ) - inference handles to store an activity record for all inference endpoints - add path, status code, and content type to Activities page - toggle on/off columns no Activities page - add configurable capture level for inference endpoints so large binary blobs are not stored in memory - store captures in compressed binary format v209	2026-04-28 20:33:03 -07:00
Quentin Machu	a846c4f18c	config: remove hard cap on macro length (#718 ) Remove macro value limit of 1024 characters	2026-04-28 13:32:54 -07:00
Marcus	5bae33a769	ui-svelte: default theme to user preferred color scheme (#712 ) Simple, if not set is localStorage use whatever the user's preferred color scheme is to start.	2026-04-27 06:44:22 -07:00
Benson Wong	8f4ff01f93	ui-svelte: make it easier to toggle panels in logs view	2026-04-26 22:12:43 -07:00
Benson Wong	e8d4384cd2	ui-svelte: support reasoning and reasoning_content (#708 ) Support `reasoning` v1/chat/completion delta that vLLM uses. v208	2026-04-26 13:11:48 -07:00
Benson Wong	ce28485be2	ui-svelte: add prompt processing histogram (#705 ) Activities page shows histograms for prompt processing and token generation times. Fix: #691 Fix: #703	2026-04-25 16:13:07 -07:00
Damir	3cd7837b1f	fix: support architecture-specific download URLs in install script (#698 ) Just a small fix to include proper llama-swap binary when building the arm64 architecture.	2026-04-23 18:05:33 -07:00
Benson Wong	0b31ccacc1	ui-svelte: fix histogram calculation (#695 ) - Fix the histogram calculation to use server provided generation tokens/second. - Move histogram to Activities page where it can exist with the rest of the token metrics Fixes #681 v206 v207	2026-04-22 23:42:39 -07:00
Bryan Gahagan	5938dbee8f	Push unified docker images on scheduled runs (#694 ) Fixes #693	2026-04-22 20:46:51 -07:00
Benson Wong	66639e83f7	proxy: replace fsnotify with stat-poll watcher and add SIGHUP reload (#685 ) The fsnotify-based config watcher does not work reliably when the config file is bind-mounted into a Docker container as an individual file, and mishandles k8s ConfigMap projections (atomically swapped symlinks). Replace it with a small os.Stat-polling watcher and add SIGHUP as an explicit reload signal. - new proxy/configwatcher package: 2s os.Stat poller, follows symlinks, fires on mtime/size change and on missing -> present transitions - SIGHUP triggers reload unconditionally (works without --watch-config) via the same ConfigFileChangedEvent pipeline so the UI sees identical state transitions - watcher goroutine now exits cleanly on shutdown via a context - drop github.com/fsnotify/fsnotify dependency fixes #682 v205	2026-04-21 23:21:48 -07:00
Benson Wong	625b296720	docker/unified: add uv via pip install (#681 ) Install uv after the cpp tool binaries are copied and before the llama-swap binary, enabling `uv run` usage for Python-based inference backends like vLLM. - add python3-pip to runtime apt installs - add `pip install uv --break-system-packages` after cpp installs fixes #628 Co-authored-by: Claude <noreply@anthropic.com>	2026-04-20 20:55:51 -07:00
Benson Wong	231e62291c	proxy: fix matrix race and process stop bug (#677 ) - matrix.go change logic to consider any proxy.Process not in StateStopped or StateShutdown - process.StopImmediately, and Stop() which called it had a subtle bug where it only handled state transitions from StateReady to StateStopping. StateStarting -> StateStopping was ignored completely. fix: #670 v204	2026-04-20 00:21:11 -07:00
Benson Wong	57ac666598	.github/workflows: tweak push ghcr conditional (#676 )	2026-04-19 13:56:26 -07:00
Benson Wong	69728301f5	.github/workflows: add toggle for pushing unified images to github (#672 ) Add ability to dispatch (manually run) unified container builds in github without push to ghcr.io.	2026-04-19 10:10:48 -07:00
Benson Wong	c176fa70f1	docker/unified: add spirv-headers to fix vulkan build (#669 )	2026-04-18 12:18:10 -07:00
Benson Wong	5e3c646829	proxy: compress captures with zstd (#668 ) The previous captures were saved uncompressed in memory. In agentic workflows there can be many turns with each request containing the previous context in the body with a lot of redundant data. Use zstd to compress the request and response data before keeping a copy of memory. Results: - Average Percentage Saved: 73.19% - Average Compression Factor: ~6.77:1 v203	2026-04-17 23:29:37 -07:00
Benson Wong	c3f0d43e6e	proxy: fix race conditions during swap (#667 ) I pointed Opus 4.7 (high effort) at proxy.ProcessGroup to identify any race conditions in the swapping code. It found a race condition where there is a small window in the fast path for routing a request to a loaded model. There is a very small window where: - model M1 is loaded and ready for requests - a request, R1, for M1 comes in - a request, R2, for M2 comes in almost immediately after - R1 acquires the lock, sees M1 is loaded (fast path), releases the lock `[race window]` and the request is ready to be forwarded - the race window occurs between the release of the lock and the request being forwarded - the lock is released so requests can be handled concurrently - R2 comes in within the `[race window]`, acquires the lock, triggers a model swap to M2. stopping M1 - R1 is forwarded to a model that is unloaded or in the process of shutting down creating an error response In deployed systems the race window is very small and doesn't happen often. However with #635 and PR #656 I though this deserved a bit more attention. It is not concluded that this race is the cause of #635 but the race is likely to happen more often under sustained or high load. AI Note: Opus 4.7 x-high effort took about an hour to write the original patch. With the pattern discovered the fix to matrix.go was very quick. GLM 5.1 using the previous established patterns was able to easily write the fix for ProcessGroup.StopProcesses(). Supersedes: #656 Updates: #277, #635	2026-04-17 21:23:17 -07:00
Benson Wong	f6cf9f5844	proxy: Refactor tests (#660 ) - use YAML for test configurations - remove most uses of simple-responder, opting to use process.testHandler Fixes #655	2026-04-16 22:47:42 -07:00
Benson Wong	121fd93ad8	Makefile: restore linux arm64 targets Fix #641	2026-04-14 22:05:39 -07:00
Benson Wong	17233e9278	docs: update configuration.md for matrix v202	2026-04-14 22:01:03 -07:00
Benson Wong	4866d16c3e	README.md: update to use matrix instead of groups	2026-04-14 21:57:49 -07:00
Benson Wong	35193f82f1	proxy: add swap matrix with solver-based model swapping (#646 ) Add a new swap matrix to supersede groups for running concurrent models. The matrix uses a solver that picks the lowest cost evictions to make a requested model available. This simple approach along with a very basic DSL grammar can enable very complex swapping scenarios. - add DSL parser for set expressions with & (AND), \| (OR), (), +ref - add MatrixConfig structs, validation, and topological sort for +ref - add MatrixSolver with cost-minimizing swap decisions - add Matrix runtime integrating solver with Process lifecycle - integrate matrix into ProxyManager with if-branches at all endpoints - update config.example.yaml and config-schema.json with matrix schema - config enforces groups XOR matrix (cannot use both) fixes #643	2026-04-14 21:55:30 -07:00
Benson Wong	40e39f7a86	ui-svelte: fix security issues (#649 )	2026-04-12 16:21:31 -07:00
Benson Wong	a9d840ffd7	proxy,proxy/config: restore timeouts to pre PR 619 (#648 ) Reset the default ResponseHeader timeout to 0 (no timeout) which was set to 60 seconds in PR #619. Fixes #647 v201	2026-04-11 20:42:13 -07:00
Benson Wong	7b2b82777f	docker/unified: derive rootless image from root container (#644 ) Build the root image once, then derive the rootless variant from it using a small inline Dockerfile that adds the non-root user and chowns the writable directories. This halves the number of CI jobs (4 → 2) and eliminates the redundant full CUDA compilation for the rootless variant. - remove RUN_UID build arg from build-image.sh - derive rootless image inline after root build completes - collapse variant matrix out of unified-docker.yml - push both root and rootless tags in a single CI job Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-10 22:59:54 -07:00
Benson Wong	d87f0ce2c5	docker/unified: publish rootless image variant (#630 ) v200	2026-04-07 03:05:53 -07:00
Leoy	06bc6a614c	proxy: preserve wall-clock duration in metrics (#629 ) Keep request duration from being underreported when upstream timings only cover part of the full request lifecycle. - compare wall-clock and upstream timing durations - keep token and throughput values from timings - add regression coverage for underreported timings fixes #602	2026-04-07 01:52:41 -07:00
Ron M	a37b4866d8	proxy: add configurable HTTP timeouts for models and peers (#619 ) Add configurable HTTP timeout settings to both models and peers to support installations that requires longer timeouts than the current hardcoded defaults. Closes #618	2026-04-06 19:30:27 +08:00
Benson Wong	981910d734	ci: validate config.example.yaml against config-schema.json (#627 ) Extend the existing config-schema workflow to also validate config.example.yaml against config-schema.json using check-jsonschema. - add config.example.yaml to PR and push path triggers - install check-jsonschema via pip - run validation of config.example.yaml against schema https://claude.ai/code/session_01Y1oqwE6mwNs9UTJgZRgXtG --------- Co-authored-by: Claude <noreply@anthropic.com>	2026-04-05 15:17:57 +08:00
Benson Wong	a185efe37e	docker: make CMAKE_CUDA_ARCHITECTURES configurable via build arg (#625 ) Expose CMAKE_CUDA_ARCHITECTURES as a Docker build ARG so users can customize CUDA architectures via --build-arg without editing the Dockerfile. - convert hardcoded ENV to ARG with default, feeding into ENV - replace silent fallback defaults (:-) in scripts with :? guards to fail fast if the env var is missing - add usage example to Dockerfile header Follow up to: #624 https://claude.ai/code/session_01EWiUe7jNABX7Uz95dUGJqK Co-authored-by: Claude <noreply@anthropic.com>	2026-04-04 08:49:59 +08:00
Benson Wong	1dd1aadf93	docker/unified: add ik_llama.cpp to CUDA container (#620 )	2026-04-03 15:16:30 +08:00
Benson Wong	955900972a	add /sdapi to list of supported endpoints	2026-04-01 12:01:38 +08:00
Benson Wong	c2c8cfaf81	docker/unified: build llama.cpp with static libraries (#616 )	2026-04-01 03:38:07 +08:00
Benson Wong	1e440770ea	ci: fix matrix exclude for scheduled docker workflow (#610 )	2026-03-29 20:04:28 +09:00
Benson Wong	c794273c83	docker/unified,.github: fix unified build (#606 )	2026-03-27 10:31:12 +09:00
dependabot[bot]	6574a52cbb	build(deps): bump picomatch from 4.0.3 to 4.0.4 in /ui-svelte (#605 )	2026-03-26 22:28:24 +09:00
Benson Wong	8fabc75634	docker/unified: vulkan build fixes (#600 ) multiple fixes to vulkan build: - use ubuntu 26.04 to be compatible with AMD 395+ (Strix halo) hardware - add home directory in container - fix stable-diffusion install to actually enable vulkan --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> v199	2026-03-25 23:26:13 +09:00
Benson Wong	e5e7391b6d	.github,docker/unified: include vulkan build (#599 ) Update docker/unified scripts to support building both cuda and vulkan unified images.	2026-03-25 06:58:28 +09:00
Benson Wong	2c282dccad	.github,docker/unified: improve caching and fix bugs (#598 ) - set up a GHA scheduled job to build the container nightly - enabling pushing a llama-swap:unified and a llama-swap:unified-Y-M-D image to ghcr.io - tidy up Dockerfile to use a non-root user and llama-swap as an entry point	2026-03-23 22:24:40 +09:00
Benson Wong	916d13f5bd	.github/workflows,docker/unified: add cuda based unified container (#597 ) Add Docker build scripts for a unified cuda docker container with llama-server, stable-diffusion.cpp, whisper.cpp.	2026-03-22 21:11:54 +09:00
Benson Wong	a3725e7d09	Update go.mod to 1.26.1 (#593 )	2026-03-20 16:09:58 +09:00
Benson Wong	15bd55d3a9	proxy, ui-svelte: add /sdapi/v1 endpoint support (#587 ) Add proxy routes for stable-diffusion.cpp's /sdapi/v1/txt2img, /sdapi/v1/img2img, and /sdapi/v1/loras endpoints. POST endpoints use proxyInferenceHandler (model in JSON body), GET /loras uses proxyGETModelHandler (model in query param). Update the image playground with a dual-mode UI supporting both OpenAI and SDAPI backends. In SDAPI mode, loras are fetched first to prime the server-side cache, and all txt2img parameters are exposed (negative prompt, steps, cfg_scale, seed, batch_size, clip_skip, sampler, scheduler, lora selection with multipliers). - Add 3 sdapi route registrations in proxymanager.go - Add sdApi.ts client with generateSdImage and fetchSdLoras - Add SDAPI types (SdApiTxt2ImgRequest, SdApiResponse, etc.) - Add /sdapi to vite dev proxy config - Add backend tests for sdapi routing - Support batch image display in gallery grid https://claude.ai/code/session_0186MGX6NXdHVBTv2KH45fqn --------- Co-authored-by: Claude <noreply@anthropic.com>	2026-03-19 22:08:31 +09:00
Benson Wong	c3c258a55d	proxy: fix metrics capture for v1/responses (#586 ) properly parse anthropic compatible usage data from streaming responses. closes: #577 v198	2026-03-13 16:50:12 -07:00
Benson Wong	29a38fde0d	ui-svelte: upgrade to vite 8 (#585 ) Upgrade vite and related dependencies to take advantage of Vite 8's improved build times via Rolldown and Oxc. - vite: ^6.3.5 → ^8.0.0 - @sveltejs/vite-plugin-svelte: ^5.0.3 → ^7.0.0 - svelte: ^5.19.0 → ^5.46.4 - vite-plugin-compression2: ^2.4.0 → ^2.5.1 - vitest: ^4.0.18 → ^4.1.0 --------- Co-authored-by: Claude <noreply@anthropic.com>	2026-03-13 08:45:59 -07:00
tesuri	d569681daa	Change model sorting to natural order (#582 ) Use natural sorting for model names. Previously the model list was sorted lexicographically, which resulted in unintuitive ordering when numbers were included in the name. Example: Before qwen3.5:2B qwen3.5:35B-3AB qwen3.5:9B After qwen3.5:2B qwen3.5:9B qwen3.5:35B-3AB This change sorts models using natural order so numeric parts are compared numerically.	2026-03-12 07:49:34 -07:00
Benson Wong	24efdb76b1	config: add macro support for name and description fields (#578 ) Extend macro substitution to the name and description fields of ModelConfig, matching the behavior already present for cmd, proxy, checkEndpoint, and filters. - substitute global/model macros (including MODEL_ID) in name and description - substitute PORT macro in name and description when allocated - validate no unknown macros remain in name and description after substitution - add tests for macro substitution, MODEL_ID, and unknown macro error	2026-03-10 08:27:05 -07:00
Benson Wong	cc77139ff8	proxy,proxy/config: add global TTL feature (#554 ) Add a new configuration parameter globalTTL that all models will inherit. The default value is 0 which matches the currently functionality to never automatically unload a model. The model.ttl's default has changed to -1, which means use the global TTL value. Any model.ttl >=0 is now value with 0 meaning never unload. This allows a model to override a globalTTL > 0 and be configured to never unload. Fixes #459 Closes #512 v197	2026-03-01 21:02:12 -08:00
Benson Wong	390a35bf93	ui-svelte: add copy button to markdown code blocks (#537 ) Add a copy-to-clipboard button that appears on hover for each code block rendered in the chat interface assistant messages. - Svelte action `codeBlockCopy` injects a button into every `<pre>` element - MutationObserver reattaches buttons as streaming content arrives - Button shows a check icon for 2 seconds after a successful copy - Uses clipboard API with execCommand fallback for non-secure contexts - CSS hides button by default and reveals it on pre:hover https://claude.ai/code/session_01PTA5ao5YQuFAS6a9juLeZW --------- Co-authored-by: Claude <noreply@anthropic.com>	2026-03-01 09:48:56 -08:00
pdscomp	181f71ca11	.github,docker: add cuda13 architecture support (#551 ) Add `cuda13` as a supported build architecture, targeting the `ghcr.io/ggml-org/llama.cpp:server-cuda13` upstream base image. The `server-cuda13` image ships with CUDA 13 libraries, providing improved performance on recent NVIDIA hardware compared to the existing `server-cuda` (CUDA 12) image. Users with newer GPUs (e.g., RTX 50-series) benefit from reduced model load latency and higher token throughput. - Add `cuda13` to the allowed architectures list in `docker/build-container.sh` - Add `cuda13` to the CI matrix in `.github/workflows/containers.yml` so the container is built and pushed automatically	2026-03-01 09:37:08 -08:00

1 2 3 4 5 ...

496 Commits