
Connect to VLM for document processing, OCR, and document extraction.
Overview
The VLM integration on StackAI gives workflows access to vision-language models for image and document understanding. Use it to caption images, extract data from screenshots, read charts, and run multimodal reasoning inside agents.
Top Use Cases
Screenshot and UI understanding
Use VLM on StackAI to interpret screenshots inside agentic workflows.
Visual document QA
Answer questions about complex documents with visual layout using VLM.
Image upload
Run StackAI workflows that upload files and verify uploads.
