llama-swap

Author	SHA1	Message	Date
krzychdre	b2fcc2daa1	ui-svelte: fix cached tokens total counting -1 sentinel (#760 ) The backend uses cache_tokens=-1 as a sentinel for endpoints that don't report cache stats (embeddings, vLLM). The activity table correctly renders these as "-", but the totals widget summed the sentinels directly, so each such request subtracted 1 from the displayed total. - clamp cache_tokens with Math.max(0, ...) when reducing	2026-05-15 14:42:44 -07:00
Benson Wong	fd3c28ffc5	Refactor Activity Page (#710 ) - inference handles to store an activity record for all inference endpoints - add path, status code, and content type to Activities page - toggle on/off columns no Activities page - add configurable capture level for inference endpoints so large binary blobs are not stored in memory - store captures in compressed binary format	2026-04-28 20:33:03 -07:00
Benson Wong	ce28485be2	ui-svelte: add prompt processing histogram (#705 ) Activities page shows histograms for prompt processing and token generation times. Fix: #691 Fix: #703	2026-04-25 16:13:07 -07:00
Benson Wong	0b31ccacc1	ui-svelte: fix histogram calculation (#695 ) - Fix the histogram calculation to use server provided generation tokens/second. - Move histogram to Activities page where it can exist with the rest of the token metrics Fixes #681	2026-04-22 23:42:39 -07:00