feat(videogen): LastImage — pin the trailing keyframe (first-last-frame-to-video) #26

Merged
steve merged 4 commits from feat/videogen-last-frame into main 2026-08-08 07:10:27 +00:00
3 changed files with 29 additions and 12 deletions
Showing only changes of commit 588e092465 - Show all commits
+18 -5
View File
@@ -272,17 +272,30 @@ tr, err := tm.Transcribe(ctx, audio.TranscriptionRequest{
voices, err := ls.ListVoices(ctx, "kokoro") // []string of voice ids
```
## Video: text-to-video + image-to-video
## Video: text-to-video, image-to-video, first-last-frame
Video generation lives in the `videogen` package (ADR-0019), mirroring
imagegen/audio: one small `Model` contract, zero values mean backend
defaults, bytes in/out. Text-to-video and image-to-video are one surface
a nil `InitImage` is a pure text prompt; setting it conditions generation
on that frame (hybrid checkpoints like Wan 2.2 TI2V serve both). First
backend: llama-swap (blocking `/v1/videos/sync`, vLLM-Omni style — the
defaults, bytes in/out. All modes are one surface, selected by which
keyframes are set rather than by a mode flag:
| `InitImage` | `LastImage` | mode |
|---|---|---|
| nil | nil | text-to-video |
| set | nil | image-to-video (hybrid checkpoints like Wan 2.2 TI2V serve both) |
| set | set | first-last-frame — both ends pinned |
| nil | set | pin the destination, model invents the approach |
First backend: llama-swap (blocking `/v1/videos/sync`, vLLM-Omni style — the
response body is the encoded clip, so `Result` carries a single `Video`).
Generation runs for minutes; bound the call with a context deadline.
**`LastImage` support is per-model and cannot be detected.** A backend that
does not understand a trailing keyframe ignores the part and returns an
ordinary clip — indistinguishable from success. There is no capability bit,
because the contract has no way to learn one, so a caller depending on the
pin must establish support out of band.
```go
vm, _ := ls.VideoModel("videogen-wan22-5b")
res, err := vm.Generate(ctx, videogen.Request{Prompt: "a cat surfing"},
+6 -3
View File
@@ -38,10 +38,13 @@ type videoModel struct {
// bound the call with a context deadline.
//
// Parameter names follow vLLM-Omni's videos API (num_frames, fps,
// num_inference_steps, guidance_scale); the conditioning frame is sent as an
// `input_reference` file part, following OpenAI's videos API. Upstreams
// num_inference_steps, guidance_scale); the leading conditioning frame is sent
// as an `input_reference` file part, following OpenAI's videos API, and a
// trailing keyframe (Request.LastImage) as `input_reference_last`. Upstreams
// ignore fields they don't understand, and optional fields stay off the wire
// entirely so the model's own defaults apply.
// entirely so the model's own defaults apply — which is also why a backend
// without first-last-frame support returns an ordinary clip here rather than
// an error.
func (m *videoModel) Generate(ctx context.Context, req videogen.Request, opts ...videogen.Option) (*videogen.Result, error) {
req = req.Apply(opts...)
if strings.TrimSpace(req.Prompt) == "" {
1
+5 -4
View File
@@ -59,10 +59,11 @@ type Request struct {
//
// Support is per-model and NOT advertised anywhere in this contract: a
// backend that does not understand a trailing keyframe ignores it and
// returns an ordinary clip, which is indistinguishable from success. A
// caller that needs to know whether the pin took effect must establish
// that out of band — see the note on LastImage support in
// provider/llamaswap.
// returns an ordinary clip, which is indistinguishable from success.
// There is no capability bit to consult, because the contract has no way
// to learn one. A caller that needs to know whether the pin actually took
// effect must establish that out of band — by configuration it controls,
// not by inspecting the result.
LastImage *Image
// Size is the requested resolution, e.g. "1280x704"; "" = backend default.