Vision (describe_image)
The describe_image tool lets the agent analyze images you attach or reference by path.
Requirements
- Code or Office mode (when vision feature is enabled)
VISION_API_KEYor provider config supporting image models (seeconfig.toml)- Image within size limits enforced by the runtime
Typical uses
- Screenshot UI bugs โ describe layout and suggest fixes
- Diagram or whiteboard photo โ extract structure
- Scan a chart in
inbox/for an Office summary
How to invoke
- Drag/drop or paste an image into chat, or
- Place an image under workspace and ask the agent to read it
The model receives a text description from the vision backend, then continues reasoning with other tools.
Privacy
Images are sent to your configured vision provider โ treat sensitive screenshots accordingly. Do not attach credentials or personal ID photos unless you accept that risk.
Related: File tools ยท API key