Skip to main content

VisionTool

Description

This tool is used to extract text from images. When passed to the agent it will extract the text from the image and then use it to generate a response, report or any other output. The URL or the PATH of the image should be passed to the Agent. You can also ask a custom query about the image and pick a complexity_level that automatically selects the model best suited for the request: When an explicit llm or model is provided to the tool, it takes precedence over the complexity-based model selection.

Installation

Install the crewai_tools package

Usage

In order to use the VisionTool, the OpenAI API key should be set in the environment variable OPENAI_API_KEY.
Code

Arguments

The VisionTool accepts the following arguments: