PaperEdits lab

AI video benchmarks with the messy parts left in.

Methods, failures, limitations, and raw evidence for every test. We publish useful findings—even when there is no clean winner.

Published reports

September 2026 · Exploratory

Gemini agentic video understanding for long-form editing

Gemini 3.7 Flash agentic versus static processing across six synthetic 10-minute videos: five valid matched pairs, ten planned metrics, and one rejected output.

Freeze first

Prompts, ground truth, schemas, and scoring rules are committed before model calls.

Show attrition

Malformed and failed runs remain visible instead of disappearing from the denominator.

Separate evidence

Automated scores, human judgments, and unmeasured questions are clearly labeled.