Commit Graph
2 Commits
Author SHA1 Message Date
steve f16ac83790 fix(gif): correct the streaming threshold and three encode-path hazards
This repo has no CI and no adversarial reviewer, so these come from reading the
harness back against the production sandbox contract rather than from a bot.

The size estimate was wrong by ~1000x in the wrong direction. `width * height *
9e-5` was never MiB — for a 640x480 frame it yields 27.6, so a trivial 2-second
GIF "projected" 800+ MiB and every render would have tripped the streaming
branch and come back as an MP4. The bench missed it because it bind-mounted
/workspace from the host disk, where statvfs reports hundreds of GB and the
threshold never fires; production mounts a 64 MiB tmpfs, where it always would
have. Replaced with a named PNG_BYTES_PER_PIXEL = 0.29, measured (86 KiB per
busy 640x480 frame). Verified against the real tmpfs: a 6s piece now estimates
7.6 MiB against a 38 MiB budget and stays a GIF, while a 90s piece streams.

Three smaller ones on the streaming path:

- ffmpeg's stderr was a pipe nobody drained until the encode finished, so a
  chatty encoder could fill the 64 KiB buffer, block, stop reading stdin, and
  deadlock a long render halfway. It writes to a file now.
- _close_pipe assumed a live pipe; if ffmpeg had already died, closing stdin
  raised BrokenPipeError and buried the real exit code and log.
- A failed encode could leave a half-written /workspace/final.* behind, which
  still comes back as a files_out file_id — a broken artifact that looks
  deliverable is worse than none, so it's removed on the way out.
2026-07-24 19:05:19 -04:00
steve 7906a1ab4c feat(gif): staging rails + bundled render harness
Two changes, both driven by measurement against 128 production gifsmith runs
and a local replica of the agent loop (docs/audits/2026-07-24-gifsmith-quality.md
in mort).

Staging rules. The renders that "worked" still read badly, always the same four
ways: the subject drawn 5-10% of the frame height under an empty sky, long runs
of pixel-identical frames, captions narrating action that was never drawn, and
flat untextured shapes. The recipe never asked for anything else. SKILL.md now
states four checkable rules — fill the frame, every frame moves, show it don't
caption it, give it depth — plus easing/anticipation/squash-and-stretch, and the
self-critique makes a "no" on depiction, subject size or caption legibility a
mandatory re-render instead of a suggestion.

Render harness. 26% of production code_exec calls end in a traceback, and every
crash class lives in plumbing the agent re-derives each run: scratch dir, frame
loop, numbering, canvas size, contact sheet, encode. scripts/gifsmith.py owns all
of it and the agent writes only @anim.scene drawing functions. A scene that
raises is now skipped while the others still render and ship, the canvas is
size-locked so ffmpeg cannot fail on jittering frames, and the drawing vocabulary
is re-exported so a missing import cannot NameError.

The harness also unblocks the multi-minute pieces this skill advertises but could
never produce: /tmp is a 16 MiB tmpfs (~13s of frames) and /workspace 64 MiB
(~51s), which is why "No space left on device" is 15% of all production crashes.
Anything past the GIF length threshold is an MP4, and an MP4 encodes in one
streaming pass, so those frames now go straight into ffmpeg's stdin and never
touch disk — verified at 1350 frames / 90 s with 0 MiB on disk.

encode.py is removed; the harness supersedes it.
2026-07-24 18:49:11 -04:00