watch
Watch a video (URL or local file) and answer questions about it. Samples frames with ffmpeg, pulls a timestamped transcript from captions or local speech-to-text, and hands both to the agent so answers are grounded in what is actually on screen and said. Triggers: 'watch this video', 'what happens at', 'summarize this video', 'analyze this YouTube link', 'what does this clip show', 'review this screen recording', 'what went wrong in this bug repro', 'read the slides in this talk', 'transcribe this video', 'what did they say about', 'check this demo video', 'look at this loom', 'what is on screen at', 'youtube link'. Self-improving: every run ends with a gotcha review.