feat(videogen): LastImage — pin the trailing keyframe (first-last-frame-to-video) #26
@@ -272,17 +272,30 @@ tr, err := tm.Transcribe(ctx, audio.TranscriptionRequest{
|
||||
voices, err := ls.ListVoices(ctx, "kokoro") // []string of voice ids
|
||||
```
|
||||
|
||||
## Video: text-to-video + image-to-video
|
||||
## Video: text-to-video, image-to-video, first-last-frame
|
||||
|
||||
Video generation lives in the `videogen` package (ADR-0019), mirroring
|
||||
imagegen/audio: one small `Model` contract, zero values mean backend
|
||||
defaults, bytes in/out. Text-to-video and image-to-video are one surface —
|
||||
a nil `InitImage` is a pure text prompt; setting it conditions generation
|
||||
on that frame (hybrid checkpoints like Wan 2.2 TI2V serve both). First
|
||||
backend: llama-swap (blocking `/v1/videos/sync`, vLLM-Omni style — the
|
||||
defaults, bytes in/out. All modes are one surface, selected by which
|
||||
keyframes are set rather than by a mode flag:
|
||||
|
||||
| `InitImage` | `LastImage` | mode |
|
||||
|---|---|---|
|
||||
| nil | nil | text-to-video |
|
||||
| set | nil | image-to-video (hybrid checkpoints like Wan 2.2 TI2V serve both) |
|
||||
| set | set | first-last-frame — both ends pinned |
|
||||
| nil | set | pin the destination, model invents the approach |
|
||||
|
||||
First backend: llama-swap (blocking `/v1/videos/sync`, vLLM-Omni style — the
|
||||
response body is the encoded clip, so `Result` carries a single `Video`).
|
||||
Generation runs for minutes; bound the call with a context deadline.
|
||||
|
||||
**`LastImage` support is per-model and cannot be detected.** A backend that
|
||||
does not understand a trailing keyframe ignores the part and returns an
|
||||
ordinary clip — indistinguishable from success. There is no capability bit,
|
||||
because the contract has no way to learn one, so a caller depending on the
|
||||
pin must establish support out of band.
|
||||
|
||||
```go
|
||||
vm, _ := ls.VideoModel("videogen-wan22-5b")
|
||||
res, err := vm.Generate(ctx, videogen.Request{Prompt: "a cat surfing"},
|
||||
|
||||
@@ -38,10 +38,13 @@ type videoModel struct {
|
||||
// bound the call with a context deadline.
|
||||
//
|
||||
// Parameter names follow vLLM-Omni's videos API (num_frames, fps,
|
||||
// num_inference_steps, guidance_scale); the conditioning frame is sent as an
|
||||
// `input_reference` file part, following OpenAI's videos API. Upstreams
|
||||
// num_inference_steps, guidance_scale); the leading conditioning frame is sent
|
||||
// as an `input_reference` file part, following OpenAI's videos API, and a
|
||||
// trailing keyframe (Request.LastImage) as `input_reference_last`. Upstreams
|
||||
// ignore fields they don't understand, and optional fields stay off the wire
|
||||
// entirely so the model's own defaults apply.
|
||||
// entirely so the model's own defaults apply — which is also why a backend
|
||||
// without first-last-frame support returns an ordinary clip here rather than
|
||||
// an error.
|
||||
func (m *videoModel) Generate(ctx context.Context, req videogen.Request, opts ...videogen.Option) (*videogen.Result, error) {
|
||||
req = req.Apply(opts...)
|
||||
if strings.TrimSpace(req.Prompt) == "" {
|
||||
|
||||
@@ -59,10 +59,11 @@ type Request struct {
|
||||
//
|
||||
// Support is per-model and NOT advertised anywhere in this contract: a
|
||||
// backend that does not understand a trailing keyframe ignores it and
|
||||
// returns an ordinary clip, which is indistinguishable from success. A
|
||||
// caller that needs to know whether the pin took effect must establish
|
||||
// that out of band — see the note on LastImage support in
|
||||
// provider/llamaswap.
|
||||
// returns an ordinary clip, which is indistinguishable from success.
|
||||
// There is no capability bit to consult, because the contract has no way
|
||||
// to learn one. A caller that needs to know whether the pin actually took
|
||||
// effect must establish that out of band — by configuration it controls,
|
||||
// not by inspecting the result.
|
||||
LastImage *Image
|
||||
|
||||
// Size is the requested resolution, e.g. "1280x704"; "" = backend default.
|
||||
|
||||
Reference in New Issue
Block a user