V

Delegate image understanding tasks to any configurable OpenAI-compatible vision model from Claude Code.

Vision

Give text-only Claude Code models the ability to see and understand images, screenshots, and diagrams by delegating to OpenAI-compatible vision models.

Image AnalysisAgent WorkflowTool Orchestration
2 Stars0 ForksUpdated Sep 6, 2026View source

Skill impact

A focused workflow with measurable value.

Stars

2

Forks

0

Updated

Sep 6, 2026

About Vision

See and understand images when you (the current model) have no native vision. Use this WHENEVER you need to look at, read, describe, OCR, or reason about the contents of an image, screenshot, photo, diagram, chart, UI mockup, or scanned page — including when the user references a local image file or an image URL and you cannot view it yourself.

What Vision Can Help You Do

Process local image files or remote URLs using a configurable vision model to perform OCR, diagnose errors, read charts, or inspect UI mockups and diagrams.

Process local image files and remote URLs to extract visual content
Transcribe text and layouts accurately using OCR from images and receipts
Diagnose application errors and trace issues from debugging screenshots
Convert charts and visual diagrams into structured markdown tables and text

Common use cases

01

Extracting text and preserving layouts from scanned documents or receipts

02

Diagnosing error messages and application crashes from screenshots

03

Converting visual charts and diagrams into structured markdown tables

04

Analyzing UI mockups and design assets against expected implementation outputs

05

Explaining software architecture diagrams from remote image URLs

Quick start

Install command

npx skills add kotot/vision@vision

Tags

Topics and capabilities

#Media Analysis#Workflow Automation#Agent Skills