Delegate image understanding tasks to any configurable OpenAI-compatible vision model from Claude Code.
Give text-only Claude Code models the ability to see and understand images, screenshots, and diagrams by delegating to OpenAI-compatible vision models.
A focused workflow with measurable value.
Stars
2
Forks
0
Updated
Sep 6, 2026
See and understand images when you (the current model) have no native vision. Use this WHENEVER you need to look at, read, describe, OCR, or reason about the contents of an image, screenshot, photo, diagram, chart, UI mockup, or scanned page — including when the user references a local image file or an image URL and you cannot view it yourself.
Process local image files or remote URLs using a configurable vision model to perform OCR, diagnose errors, read charts, or inspect UI mockups and diagrams.
Extracting text and preserving layouts from scanned documents or receipts
Diagnosing error messages and application crashes from screenshots
Converting visual charts and diagrams into structured markdown tables
Analyzing UI mockups and design assets against expected implementation outputs
Explaining software architecture diagrams from remote image URLs
npx skills add kotot/vision@visionTopics and capabilities