Freeze first
Prompts, ground truth, schemas, and scoring rules are committed before model calls.
PaperEdits lab
Methods, failures, limitations, and raw evidence for every test. We publish useful findings—even when there is no clean winner.
Gemini 3.7 Flash agentic versus static processing across six synthetic 10-minute videos: five valid matched pairs, ten planned metrics, and one rejected output.
Prompts, ground truth, schemas, and scoring rules are committed before model calls.
Malformed and failed runs remain visible instead of disappearing from the denominator.
Automated scores, human judgments, and unmeasured questions are clearly labeled.