vision
Analyze images, audio, video, or PDFs using Google Gemini via Vertex AI (GCP credits). Generic multimodal router. Use when the user explicitly says "analyze with gemini", "use gemini vision", "run through gemini", "transcribe with gemini", "describe with gemini", or invokes /vision. Works on any media file the caller supplies — photos, screenshots, diagrams, UI, products, art, audio, video, documents. NOT for: live streams, real-time analysis, or when the host model's own vision/audio is acceptable.
Pinned to revision 4137d29feb8f, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/vision/SKILL.md
- skills/vision/references/file-type-notes.md
- skills/vision/references/prompt-patterns.md
Every link opens the file at its source, pinned to the revision this page describes.