Fix panic when requesting non-members of profiles

A panic occurs when a request for an invalid profile:model pair is made. The edge case is that the profile exists and the model exists but they're not configured as a pair. This adds an additional check to make sure the profile:model pair is valid before attempting to swap the model.
Update README.md
2025-01-16 12:06:38 -08:00 · 2025-01-13 22:37:30 -08:00 · 2025-01-12 19:48:35 -08:00
3 changed files with 64 additions and 3 deletions
@@ -5,10 +5,12 @@
 # Introduction
 llama-swap is a light weight, transparent proxy server that provides automatic model swapping to llama.cpp's server.

-Written in golang, it is very easy to install (single binary with no dependancies) and configure (single yaml file). Download a pre-built [release](https://github.com/mostlygeek/llama-swap/releases) or built it yourself from source with `make clean all`.
+Written in golang, it is very easy to install (single binary with no dependancies) and configure (single yaml file). 
+
+Download a pre-built [release](https://github.com/mostlygeek/llama-swap/releases) or build it yourself from source with `make clean all`.

 ## How does it work?
-When a request is made to an OpenAI compatible endpoints, lama-swap will extract the `model` value load the appropriate server configuration to serve it. If a server is already running it will stop it and start a new one. This is where the "swap" part comes in. The upstream server is automatically swapped to the correct one to serve the request.
+When a request is made to an OpenAI compatible endpoint, lama-swap will extract the `model` value and load the appropriate server configuration to serve it. If a server is already running it will stop it and start the correct one. This is where the "swap" part comes in. The upstream server is automatically swapped to the correct one to serve the request.

 In the most basic configuration llama-swap handles one model at a time. For more advanced use cases, the `profiles` feature can load multiple models at the same time. You have complete control over how your system resources are used.

@@ -26,7 +28,7 @@ Any OpenAI compatible server would work. llama-swap was originally designed for
  - `v1/chat/completions`
  - `v1/embeddings`
  - `v1/rerank`
-  - `v1/audio/speech`
+  - `v1/audio/speech` ([#36](https://github.com/mostlygeek/llama-swap/issues/36))
 - ✅ Multiple GPU support
 - ✅ Run multiple models at once with `profiles`
 - ✅ Remote log monitoring at `/log`
@@ -202,6 +202,21 @@ func (pm *ProxyManager) swapModel(requestedModel string) (*Process, error) {
 		return nil, fmt.Errorf("could not find modelID for %s", requestedModel)
 	}

+	// check if model is part of the profile
+	if profileName != "" {
+		found := false
+		for _, item := range pm.config.Profiles[profileName] {
+			if item == realModelName {
+				found = true
+				break
+			}
+		}
+
+		if !found {
+			return nil, fmt.Errorf("model %s part of profile %s", realModelName, profileName)
+		}
+	}
+
 	// exit early when already running, otherwise stop everything and swap
 	requestedProcessKey := ProcessKeyName(profileName, realModelName)

@@ -210,3 +210,47 @@ func TestProxyManager_ListModelsHandler(t *testing.T) {
 	// Ensure all expected models were returned
 	assert.Empty(t, expectedModels, "not all expected models were returned")
 }
+
+func TestProxyManager_ProfileNonMember(t *testing.T) {
+
+	model1 := "path1/model1"
+	model2 := "path2/model2"
+
+	profileMemberName := ProcessKeyName("test", model1)
+	profileNonMemberName := ProcessKeyName("test", model2)
+
+	config := &Config{
+		HealthCheckTimeout: 15,
+		Models: map[string]ModelConfig{
+			model1: getTestSimpleResponderConfig("model1"),
+			model2: getTestSimpleResponderConfig("model2"),
+		},
+		Profiles: map[string][]string{
+			"test": {model1},
+		},
+	}
+
+	proxy := New(config)
+	defer proxy.StopProcesses()
+
+	// actual member of profile
+	{
+		reqBody := fmt.Sprintf(`{"model":"%s"}`, profileMemberName)
+		req := httptest.NewRequest("POST", "/v1/chat/completions", bytes.NewBufferString(reqBody))
+		w := httptest.NewRecorder()
+
+		proxy.HandlerFunc(w, req)
+		assert.Equal(t, http.StatusOK, w.Code)
+		assert.Contains(t, w.Body.String(), "model1")
+	}
+
+	// actual model, but non-member will 404
+	{
+		reqBody := fmt.Sprintf(`{"model":"%s"}`, profileNonMemberName)
+		req := httptest.NewRequest("POST", "/v1/chat/completions", bytes.NewBufferString(reqBody))
+		w := httptest.NewRecorder()
+
+		proxy.HandlerFunc(w, req)
+		assert.Equal(t, http.StatusNotFound, w.Code)
+	}
+}