Skip to content

Text & analysis ​

30 text models that read an image or video and return text — captioning, OCR, classification, Q&A, and summarization. Unlike every other mode, these LLMs analyze media instead of generating it.

Quick start ​

They run through gen-ai describe (not generate), and print the answer to stdout:

bash
# describe an image (Claude Sonnet is the default)
gen-ai describe -i photo.jpg

# ask a specific question
gen-ai describe -i receipt.jpg -p "extract the total and tax"

# summarize a video (auto-routes to a video-capable model)
gen-ai describe --video clip.mp4 -p "summarize what happens"

From an MCP client, the same models are reachable through picsart_generate with an imageUrls (or videoUrl) input.

Input types ​

TypeMeaningModels
i2timage → textClaude Fable / Opus / Sonnet / Haiku, GPT, Gemini Flash
v2tvideo (or image) → textGemini 3 Pro

Only Gemini 3 Pro accepts video, so gen-ai describe --video … auto-selects it unless you force another model with -m.

Providers ​

ProviderModelsHighlights
AnthropicClaude Fable, Opus, Sonnet, HaikuOpus for hard reasoning; Haiku for high-volume
OpenAIGPT-6 Astra, GPT-5.6 family, GPT-5 and GPT-4 familiesStrong general image understanding
GoogleGemini 3 Pro, Gemini 3.8 / 3.7 / 3.6 Flash, 3.5 Flash Lite, 2.5 FlashPro reads video; Flash tiers handle image analysis

Common parameters ​

ParamCLI flagNotes
prompt-pThe question or instruction (optional — defaults to "describe this")
imageUrls-iImage(s) to analyze
videoUrl--videoVideo to analyze (Gemini 3 Pro only)
thinking--thinkingReasoning depth, where the model supports it

These models return text, so the CLI prints the answer to stdout and skips download / Drive save. Add -q to drop the model/time header (printed to stderr), or --json for { text, model, durationMs }.

Built on @picsart/ai-sdk · gen-ai CLI · Picsart MCP · Media Studio · Skills