llama-swap

Author	SHA1	Message	Date
Benson Wong	49546e2cf2	ui: fix text size svg v196	2026-02-27 23:47:52 -08:00
Benson Wong	2c078964f4	Update README with additional images Added new images for model loading and real-time log streaming sections.	2026-02-27 23:45:40 -08:00
Benson Wong	175bb36fb1	Revise README description for clarity and detail Updated description to clarify compatibility and usage.	2026-02-27 23:42:40 -08:00
Benson Wong	aedb640471	Enhance web UI section in README Updated README to enhance the description of the web interface and added details about features like token metrics, request inspection, model management, and real-time log streaming.	2026-02-27 23:40:31 -08:00
Benson Wong	2f377f6dc6	ui: add OGG audio format support to transcription playground (#544 )	2026-02-26 19:48:19 -08:00
Benson Wong	64e4c79fc3	ui: add Rerank tab to playground (#536 ) Add a new Rerank tab to the playground that lets users test /v1/rerank endpoints. Supports a visual table editor and a JSON editor mode that stay in sync when toggling between them. - add rerankApi.ts with typed wrapper for /v1/rerank - add RerankInterface.svelte with query input, sortable document table, color-coded scores, auto-add row, cancel/clear, and token usage - add rerankLoading store to playgroundActivity derived store - register Rerank tab in Playground.svelte Updates #481 v195	2026-02-21 21:59:14 -08:00
Benson Wong	19fb5f35e9	proxy: implement setParamsByID filter (#535 ) Add setParamsByID filter that applies different request parameters based on the requested model ID, enabling per-alias behaviour for a single loaded model. - add SetParamsByID field to Filters struct and SanitizedSetParamsByID method - substitute ${MODEL_ID} and other macros in setParamsByID keys and values - validate no unknown macros remain in keys or values after substitution - apply setParamsByID in proxyInferenceHandler after setParams (can override it) - update config-schema.json with setParamsByID definition - update UI to show aliases and make them selectable in the Playground closes #534 v194	2026-02-19 22:21:10 -08:00
Benson Wong	b45102bde8	ui: smart auto-scroll in LogPanel (#530 ) Pause auto-scroll when the user scrolls up to review logs, and resume when they scroll back to the bottom. - add `userScrolledUp` state variable - add `handleScroll` to detect scroll position with 40px threshold - guard the auto-scroll effect with `!userScrolledUp` closes #529 v193	2026-02-18 19:47:37 -08:00
Brian Mendonca	1688bdd1e9	proxy, ui: add pending requests count to the main dashboard (#516 ) add a real time counter of pending (inflight) requests to the UI. v192	2026-02-16 09:41:15 -08:00
Benson Wong	d33d51fa75	.coderabbit.yaml,AGENTS.md: small tweaks v191	2026-02-15 21:31:30 -08:00
Benson Wong	e3bf065574	ui: persist playground state across route navigation (#525 ) - Keep Playground component mounted when navigating away, preserving streaming/generating state - Add animated gradient effect on Playground nav link when activity is in progress	2026-02-15 21:30:52 -08:00
Benson Wong	3e52144058	ui-svelte: incremental rendering of chat messages in the Playground (#520 ) add incremental rendering to Playground > Chat	2026-02-15 11:00:44 -08:00
Benson Wong	d5e52d7d00	build: disable provenance attestations in container builds (#523 ) ## Summary - Add `--provenance=false` to docker build commands in `build-container.sh` - BuildKit attestation manifests are stored as untagged images in GHCR, and the `delete-untagged-containers` cleanup job deletes them, breaking the manifest list and causing `manifest unknown` errors on pull - ref: https://github.com/actions/delete-package-versions/issues/162	2026-02-14 10:23:08 -08:00
Benson Wong	17e5263a76	.github/workflows: fix expired token in publishing images (#522 ) Fixes: #517	2026-02-14 10:06:05 -08:00
Benson Wong	8d6d949ec3	proxy: support timings for /infill from llama-server (#510 ) fixes: #463	2026-02-07 17:16:27 -08:00
Benson Wong	b5fde8eb6d	proxy,ui-svelte: add request/response capturing (#508 ) Add saving request and response headers and bodies that go through llama-swap in memory. - captureBuffer added to configuration. Captures are enabled by default. - 5MB of memory is allocated for req/response captures in a ring buffer. Setting captureBuffer to 0 will disable captures. - UI elements to view captured data added to Activity page. Includes some QOL features like json formatting and recombining SSE chat streams - capture saving is done at the byte level and has minimal impact on llama-swap performance Fixes #464 Ref #503	2026-02-07 15:40:01 -08:00
Nuno	7eef5defb8	docs: add stable-diffusion.cpp references (#506 ) Signed-off-by: rare-magma <rare-magma@posteo.eu>	2026-02-04 20:20:39 -08:00
Benson Wong	bc01e6f539	build: add stable-diffusion server to musa and vulkan container images (#504 ) Add sd-server from stable-diffusion.cpp docker image for vulkan and musa containers. closes #450 v189	2026-02-01 16:17:26 -08:00
Benson Wong	0462e3dc3f	Reorganize UI controls and improve form interactions (#500 ) Reorganizes control placement in the playground interfaces and improves form interactions for better UX, particularly on mobile devices. ## Key Changes - AudioInterface & ImageInterface: Moved "Clear" buttons from the top control bar into the action button group below the form inputs for better visual hierarchy and logical grouping - ImageInterface: - Added prompt clearing to the `clearImage()` function so the input field is reset when clearing generated images - Updated Clear button disabled state to also check if prompt is empty, allowing users to clear an empty prompt - Added responsive flex styling (`flex-1 md:flex-none`) to the Clear button for better mobile layout - ExpandableTextarea: - Imported `untrack` from Svelte to properly handle reactive dependencies - Wrapped `expandedValue.length` in `untrack()` to prevent unnecessary reactivity when setting cursor position - Improved button visibility on mobile by changing opacity from `opacity-0` to `opacity-60` with `md:opacity-0` breakpoint, making the expand button more discoverable on touch devices ## Implementation Details The `untrack()` usage in ExpandableTextarea ensures that reading the text length doesn't create a reactive dependency, preventing potential infinite loops while still allowing the effect to run when `isExpanded` changes.	2026-02-01 15:18:22 -08:00
Benson Wong	7b20fc011b	Add path filters to CI workflows and create UI test workflow (#501 ) * .github/workflows: add UI tests and path-filter Go CI Add ui-tests.yml workflow to run svelte type checking and vitest on push/PR to main when ui-svelte/ files change. - Add path filters to go-ci.yml and go-ci-windows.yml to skip Go tests when only non-backend files change - Filter on */.go, go.mod, go.sum, and Makefile https://claude.ai/code/session_01E6acq54D8JjuE7pczxPGT7 * ui-svelte: remove unused declarations in SpeechInterface Remove unused `generatedText` state and `clearAudio` function that caused svelte-check errors. https://claude.ai/code/session_01E6acq54D8JjuE7pczxPGT7 * .github/workflows: update Node.js to v24 Node 23 is end-of-life; bump to 24 in ui-tests.yml and release.yml. https://claude.ai/code/session_01E6acq54D8JjuE7pczxPGT7 --------- Co-authored-by: Claude <noreply@anthropic.com>	2026-02-01 15:11:49 -08:00
Benson Wong	20738f3623	proxy,ui-svelte: replace old UI with svelte+playground Replace the legacy React UI with the new Svelte-based one. Introduce a Playground in the UI to quickly test out text, image, text to speech and speech to text models behind llama-swap. Key Changes New Svelte UI (ui-svelte/) - Multi-tab Playground with Chat, Image Generation, Audio Transcription, and Speech interfaces - Chat: message editing/regeneration, markdown rendering with LaTeX math support, image attachments, code syntax highlighting - Image: size selector, download/fullscreen viewing - Audio: transcription with peer support - Speech: voice caching with manual refresh, download button - Responsive mobile layout with collapsible navigation - XSS fixes and accessibility improvements Proxy Improvements - Add gzip/brotli compression for UI static assets (proxy/ui_compress.go) - Add GET /v1/audio/voices?model={model} endpoint for voice listing - Add peer support for /v1/audio/transcriptions v188	2026-01-31 22:49:13 -08:00
Benson Wong	cdea7d16bd	proxy/config: skip env macros in YAML comment lines (#496 ) Fix a bug where ${env.macro_not_exist} in comments would trigger a non-substituted macro error. fixes #495	2026-01-30 20:10:29 -08:00
Benson Wong	5de387dbf9	ui: fix node-tar vulnerability v187	2026-01-28 21:40:18 -08:00
Benson Wong	6f8e7ccb57	.github/workflows: switch release.yml to build ui-svelte	2026-01-28 21:39:10 -08:00
Benson Wong	4384315b44	ui-svelte: add Svelte port of React UI (#487 ) Trying out svelte for the UI. The port was done by Claude Code on the iOS app w/ Opus 4.5. --- * ui: add Svelte port of React UI Port the React-based UI to Svelte 5 with the following changes: - Create new ui-svelte directory with complete Svelte 5 implementation - Use Svelte stores instead of React contexts for state management - Implement custom ResizablePanels component to replace react-resizable-panels - Port all pages: LogViewer, Models, Activity - Port all components: Header, ConnectionStatus, LogPanel, ModelsPanel, etc. - Use svelte-spa-router for client-side routing - Same build output directory (proxy/ui_dist) and base path (/ui/) - Tailwind CSS 4 with same theme configuration https://claude.ai/code/session_01F3xXLYsd62gePVSFv7aboP * ui-svelte: simplify state management - Remove redundant state syncing pattern in LogPanel and ModelsPanel - Use store values directly with $ syntax instead of manual subscriptions - Consolidate duplicate title sync logic in App.svelte - Use existing syncTitleToDocument() from theme.ts https://claude.ai/code/session_01F3xXLYsd62gePVSFv7aboP * ui-svelte: use idiomatic Svelte 5 patterns - Use $effect for document side effects (theme, title) instead of store subscriptions - Use class: directive for active nav links in Header - Remove SSR guards (unnecessary for client-only SPA) - Remove leaked subscription in syncThemeToDocument - Simplify theme.ts by removing sync functions https://claude.ai/code/session_01F3xXLYsd62gePVSFv7aboP * ui-svelte: fix build warnings and improve accessibility Fix Svelte build warnings and add proper accessibility support to interactive components. - add aria-labels to buttons for screen readers - implement keyboard navigation for resizable separator - suppress intentional state initialization warnings - update Makefile to use ui-svelte build directory - add peer:true to package-lock.json dependencies * ui-svelte: reorganize navigation and add log view toggle Make Models the default landing page and add view mode toggle to the Logs page with persistent state. - set Models as default route at / - move Logs to /logs route - reorder navigation: Models, Activity, Logs - add view toggle with three modes: Panels, Proxy only, Upstream only - fix horizontal overflow with width constraints	2026-01-28 21:37:29 -08:00
Benson Wong	6439ab1515	ui: add peer:true in package-lock.json v186	2026-01-22 08:43:36 -08:00
dependabot[bot]	f94226122c	build(deps-dev): bump tar from 7.5.3 to 7.5.6 in /ui (#477 ) Bumps [tar](https://github.com/isaacs/node-tar) from 7.5.3 to 7.5.6. - [Release notes](https://github.com/isaacs/node-tar/releases) - [Changelog](https://github.com/isaacs/node-tar/blob/main/CHANGELOG.md) - [Commits](https://github.com/isaacs/node-tar/compare/v7.5.3...v7.5.6) --- updated-dependencies: - dependency-name: tar dependency-version: 7.5.6 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2026-01-21 22:55:02 -08:00
Ryan Voots	7493618fdc	Add count_tokens api proxying (#476 )	2026-01-20 09:34:42 -08:00
Benson Wong	205efd40a1	proxy: extend /running endpoint with additional process data (#474 ) Extend the /running endpoint to return more details about running processes beyond just model and state. - add cmd field to show the command being executed - add proxy field to show the proxy URL - add ttl (UnloadAfter) for automatic unloading configuration - add name and description for model metadata - update tests to verify new fields are returned correctly fixes #471	2026-01-19 17:37:00 -08:00
Benson Wong	14207f8492	ui: npm security update	2026-01-18 21:56:32 -08:00
Benson Wong	4e850c2834	config: refactor macro substitution in configuration (#470 ) This commit simplifies substitution of environment variables into the configuration. There was a lot of repetitive code substituting ${env.VAR_NAME} into different fields after the configuration was parsed into a config.Config. This refactor uses a string substitution of env vars into the YAML config before it is fully parsed. This eliminates a lot of logic while maintaining backwards compatibility.	2026-01-18 21:52:34 -08:00
Benson Wong	75fced579e	config: support macros in peer apiKey and filters (#469 ) * config: support environment variable macros in peer apiKeys Add ${env.VAR_NAME} substitution for peer apiKey fields, consistent with existing env macro support for model fields and global apiKeys. - Add env macro substitution for peers.{name}.apiKey in LoadConfigFromReader - Add tests for peer apiKey env substitution - Update config.example.yaml to show env macro usage * config: support macros in peer apiKey and filters Extend macro substitution to peer configuration fields: - peers.{name}.apiKey supports both global macros and env macros - peers.{name}.filters.stripParams supports both macro types - peers.{name}.filters.setParams supports both macro types Also renamed validateMetadataForUnknownMacros to validateNestedForUnknownMacros for reuse across model metadata and peer filters validation. v185	2026-01-16 23:10:50 -08:00
Benson Wong	b73f367f22	config-schema.json,config.example.yaml: Update examples and schema v184	2026-01-16 22:43:25 -08:00
Benson Wong	8f2137c72b	config: support environment variable macros in apiKeys (#467 ) Add substituteEnvMacros support for apiKeys configuration field, allowing API keys to be loaded from environment variables using the ${env.VAR_NAME} syntax. - Apply env macro substitution before validation - Add tests for env macro substitution in apiKeys	2026-01-16 22:41:14 -08:00
Benson Wong	124007cc98	config: add environment variable macros (#466 ) * config: add environment variable macros Add support for ${env.VAR_NAME} syntax to pull values from system environment variables during config loading. - env macros processed before regular macros (allows macros to reference env vars) - works in cmd, cmdStop, proxy, checkEndpoint, filters.stripParams, metadata - returns error if env var is not set - add comprehensive tests fixes #462 * docs: add env macro example to config.example.yaml	2026-01-16 22:25:20 -08:00
Benson Wong	eb5bfff0b0	proxy: unify filtering for local models and peers This unifies the filtering capabilities for models and peers - stripParams: removes params in the request - setParams: sets params in the request fixes #453	2026-01-15 18:59:43 -08:00
Benson Wong	3edb180c08	ci: free up disk space before ROCm container build (#460 )	2026-01-14 22:03:42 -08:00
Benson Wong	66d555e625	Improve container build reliability (#457 ) * docker: add .env usage in build-container.sh * .github,docker: add rocm, improve logging * .github,CLAUDE.md: fix workflow and update guidelines Update containers workflow to only push images when triggered manually or on schedule, not on workflow file changes. - add push trigger for workflow file changes in containers.yml - update push condition to skip on regular push events - update CLAUDE.md commit message guidelines * docker: remove comma in build-container.sh * .github,docker: improve container build workflow Add pagination support for fetching llama.cpp tags and improve debugging. - add build-container.sh to workflow trigger paths - implement fetch_llama_tag() with pagination support - replace .env with local testing instructions - add DEBUG_ABORT_BUILD flag for testing	2026-01-10 22:14:33 -08:00
Benson Wong	4f863fd9fc	CLAUDE.md: tweak instructions	2026-01-09 21:42:06 -08:00
Benson Wong	267c030457	ui: update react-router-dom to 7.12.0 (#456 ) Update react-router-dom from 7.6.2 to 7.12.0 to address security vulnerability. - Updated dependency in package.json - Regenerated package-lock.json - Verified build passes successfully - Confirmed 0 vulnerabilities with npm audit Co-authored-by: Claude <noreply@anthropic.com> v183	2026-01-08 16:13:09 -08:00
Benson Wong	c19309fe7e	CLAUDE.md: small instruction tweaks	2026-01-07 21:34:23 -08:00
Benson Wong	4413881b2d	proxy: actually add /v1/responses endpoint (#449 ) ref: #448 v182	2026-01-01 13:35:45 -08:00
Benson Wong	8df5e8563b	proxy: add /v1/responses and /v1/audio/voices endpoints (#448 ) Updates #433 Fixes #442 #226 v181	2026-01-01 12:52:12 -08:00
Benson Wong	7931212d3e	proxy: add v1/images/edits API endpoint (#447 ) Updates #433	2026-01-01 12:43:06 -08:00
Benson Wong	3dc36032fb	proxy: skip very slow tests in -short test mode (#446 ) * proxy: skip very slow tests in -short test mode * CLAUDE.md: update testing instructions	2025-12-31 14:08:56 -08:00
Benson Wong	addb98646f	proxy: add support for basic authorization (#445 ) Fixes #444 where the UI with api keys did not work. The choice to use http basic authorization is for simple, automatic browser support. No changes to the UI were necessary. Just use an API key as the password, no user name is required.	2025-12-31 13:42:35 -08:00
Benson Wong	37d74efc2d	proxy: add /v1/images/generations (#443 ) Add support for the /v1/images/generations endpoint Updates #433 Closes #191 v180	2025-12-30 21:04:58 -08:00
Benson Wong	22e098ac8b	Add Peer Model Support (#438 ) This PR allows a single llama-swap to be the central proxy for models served by other inference servers. The peer servers can be another llama-swap or any API that supports the /v1/* inference endpoint. Updates: #433, #299 Closes: #296 v179	2025-12-27 20:18:06 -08:00
Benson Wong	9864f9f517	.coderabbit.yaml: disable annoying features	2025-12-23 23:53:06 -08:00
Benson Wong	53b32f3601	proxy: add API key support (#436 ) Add configuration support for api keys that are enforced by llama-swap. Keys are stripped before sending them to upstream servers. Updates: #433, #50 and #251	2025-12-23 23:39:33 -08:00

1 2 3 4 5 ...

396 Commits