Add missing object type to /v1/models endpoint (#154 )

Update README.md [skip ci]
Readme updates [skip ci]
2025-06-02 09:25:45 -07:00 · 2025-05-30 09:38:45 -07:00 · 2025-05-30 09:19:08 -07:00
2 changed files with 15 additions and 17 deletions
@@ -40,7 +40,7 @@ In the most basic configuration llama-swap handles one model at a time. For more

 ## config.yaml

-llama-swap's configuration is purposefully simple.
+llama-swap's configuration is purposefully simple:

 ```yaml
 models:
@@ -57,29 +57,28 @@ models:
      --port ${PORT}
 ```

-But also very powerful:
+.. but also supports many advanced features:

- ⚡ `groups` to run multiple models at once
- ⚡ `macros` for reusable snippets
- ⚡ `ttl` to automatically unload models
- ⚡ `aliases` to use familiar model names (e.g., "gpt-4o-mini")
- ⚡ `env` variables to pass custom environment to inference servers
- ⚡ `useModelName` to override model names sent to upstream servers
- ⚡ `healthCheckTimeout` to control model startup wait times
- ⚡ `${PORT}` automatic port variables for dynamic port assignment
- ⚡ Docker/podman compatible
+- `groups` to run multiple models at once
+- `macros` for reusable snippets
+- `ttl` to automatically unload models
+- `aliases` to use familiar model names (e.g., "gpt-4o-mini")
+- `env` variables to pass custom environment to inference servers
+- `useModelName` to override model names sent to upstream servers
+- `healthCheckTimeout` to control model startup wait times
+- `${PORT}` automatic port variables for dynamic port assignment
+- `cmdStop` for to gracefully stop Docker/Podman containers

-Check the [wiki](https://github.com/mostlygeek/llama-swap/wiki/Configuration) full documentation.
+Check the [configuration documentation](https://github.com/mostlygeek/llama-swap/wiki/Configuration) in the wiki for all options.

 ## Docker Install ([download images](https://github.com/mostlygeek/llama-swap/pkgs/container/llama-swap))

 Docker is the quickest way to try out llama-swap:

 ```shell
-# use CPU inference
+# use CPU inference comes with the example config above
 $ docker run -it --rm -p 9292:8080 ghcr.io/mostlygeek/llama-swap:cpu

-
 # qwen2.5 0.5B
 $ curl -s http://localhost:9292/v1/chat/completions \
    -H "Content-Type: application/json" \
@@ -87,7 +86,6 @@ $ curl -s http://localhost:9292/v1/chat/completions \
    -d '{"model":"qwen2.5","messages": [{"role": "user","content": "tell me a joke"}]}' | \
    jq -r '.choices[0].message.content'

-
 # SmolLM2 135M
 $ curl -s http://localhost:9292/v1/chat/completions \
    -H "Content-Type: application/json" \
@@ -97,7 +95,7 @@ $ curl -s http://localhost:9292/v1/chat/completions \
 ```

 <details>
-<summary>Docker images are nightly ...</summary>
+<summary>Docker images are built nightly for cuda, intel, vulcan, etc ...</summary>

 They include:

@@ -291,7 +291,7 @@ func (pm *ProxyManager) listModelsHandler(c *gin.Context) {
 	}

 	// Encode the data as JSON and write it to the response writer
-	if err := json.NewEncoder(c.Writer).Encode(map[string]interface{}{"data": data}); err != nil {
+	if err := json.NewEncoder(c.Writer).Encode(map[string]interface{}{"object": "list", "data": data}); err != nil {
 		pm.sendErrorResponse(c, http.StatusInternalServerError, fmt.Sprintf("error encoding JSON %s", err.Error()))
 		return
 	}
Author	SHA1	Message	Date
Daniel Hofer	a84098d3b4	Add missing object type to /v1/models endpoint (#154 )	2025-06-02 09:25:45 -07:00
Benson Wong	4d02ccd26a	Update README.md [skip ci]	2025-05-30 09:38:45 -07:00
Benson Wong	dfd47eeac4	Readme updates [skip ci]	2025-05-30 09:19:08 -07:00