Skip to content

Changelog

Notable changes to the Picsart model catalog and these docs. Newest first. The catalog is served by @picsart/ai-sdk; the same models are reachable from the CLI, the MCP server, and Skills.

2026-09-19

SDK 6.2.3 catalog refresh — 201 models across 31 providers. Compared with the previously published catalog: 26 additions, 6 removals; Meta joins and Reve leaves the active provider list. The SDK upgrade from 6.0.0 adds 11 text models and changes result metadata and Seedance color-depth handling. See the complete update and migration guide.

2026-09-01

Catalog refresh — now 181 models across 31 providers (@picsart/ai-sdk 5.24.0). These docs were last generated from 5.19.0, so this refresh spans 5.20–5.24: three video/image models added, one renamed, none retired.

New

  • Flux Video Upscale (video, v2v — added in SDK 5.20.0) — upscales an existing clip toward 4K at 1.5x3x, in either precise (source-faithful) or creative (detail-enhancing) mode. Source clips up to 20 seconds and 2K. See Flux.
  • MiniMax H3 Max (video, t2v — added in SDK 5.22.0) — the top tier of the H3 family: text-to-video or start/end-frame interpolation with a promptExpansionMode pass, up to 15s at 768p. Image-to-video is available as an edit workflow. See MiniMax.
  • Qwen 3.0 Pro (image, t2i — added in SDK 5.22.1) — the GA flagship Qwen-Image 3.0 Pro, covering text-to-image and image editing with prompt-rewrite modes (promptExtendMode) and an optional enableThinking pass. See Qwen.

Changed

  • Hailuo 03 is now MiniMax H3 (SDK 5.21.1) — the catalog id changed from hailuo-03 to minimax-h3. This moves a public identifier: ?model=hailuo-03 deeplinks and any saved reference to the old id no longer resolve. The workflow (minimax/v2/video-generation) and pricing key (minimax-h3) are unchanged, and no parameters moved.
  • Gemini Omni 1.2 Flash is now displayed as Gemini Omni 1.1 Flash (SDK 5.21.2) — this resolves the id/name mismatch flagged in the 2026-08-27 entry above. The id was always gemini-omni-1.1-flash-preview; the display name has been corrected to agree with it rather than the id being changed, so nothing matching on the id is affected.

2026-08-27

Catalog refresh — now 178 models across 31 providers (@picsart/ai-sdk 5.19.0). These docs were last generated from 5.16.0, so this refresh spans 5.17–5.19: two video models added, none retired.

New

  • Wan 3.0 Prime (video, t2v — added in SDK 5.18.0) — the same all-in-one model as Wan 3.0, up to 7x faster. Takes image, video and audio references plus start/end frames; 5/10/15/30s at up to 1080P, adaptive aspect ratio, native audio, and an optional --enable-thinking pass. See Wan.
  • Gemini Omni 1.2 Flash (video, t2v — added in SDK 5.19.0) — the more capable of the two Gemini Omni entries. Over the original it adds a resolution param at all (360p · 720p · 1080p · 4k, default 720p), start/end frame interpolation, up to 5 reference images and 3 reference videos, clip extension from a source video under 30s, and every integer duration 3–10s rather than a fixed set. See Google.

NOTE

Gemini Omni 1.2 Flash's id is gemini-omni-1.1-flash-preview — the id says 1.1, the display name says 1.2. That mismatch comes from the upstream catalog, so anything matching on the id must use the 1.1 string. -preview in the id is likewise historical; the model is a production release.

Changed

  • Grok Imagine 1.5 is now classified t2v (was i2v), as of SDK 5.18.0 — its image input is optional, so it generates from a prompt alone. The CLI and MCP call is unchanged; only the input-type label moved. See Grok.

Docs

  • The provider-page generator now updates the frontmatter model count on multi-mode provider pages. Its regex matched only single-word mode lists, so slash-joined descriptions ("17 image/video/audio/text model(s)") silently kept a stale count — Google had been carrying one across refreshes.
  • The same generator no longer appends a blank line at EOF on every run for pages whose ## Parameters is the last section, so repeat refreshes are now byte-idempotent.

2026-08-25

Catalog refresh — now 176 models across 31 providers (@picsart/ai-sdk 5.16.0). Eleven models added, nine retired.

New

  • Wan 3.0 — an all-in-one video model taking text, image/video/audio references, and start/end frames, with adaptive ratio, intelligent duration, and audio. See Wan.
  • Picsart Effects (image + video)picsart-flow and picsart-flow-video apply curated Picsart presets as multi-step Magic Flow pipelines in one call. Preset ids come from a runtime catalog; list them with gen-ai models info picsart-flow --json. See Picsart.
  • Recraft V4 Styles — four style-reference models (Styles, Styles Vector, Styles Pro, Styles Pro Vector) with 10K-character prompts. See Recraft.
  • Grok Imagine 2.0 — sharper detail with a low/medium quality tier. See Grok.
  • MiniMax Music v3 — text-to-music with vocals or instrumentals from a style prompt and optional lyrics, with configurable audio encoding. See MiniMax.
  • Gemini 3.7 Flash — low-latency multimodal text generation for gen-ai describe. See Google.
  • ByteDance Video Enhance — now live, replacing the Video Upscaler: denoise, colour-correct and super-resolve footage up to 8K, with frame-rate conversion. See ByteDance.

Deprecated / retired

  • Pika is retired. pika-2.2, pika-2.2-scenes and pika-2.2-frames are gone and Pika is no longer a provider. The provider page is kept as a pointer to replacements — Seedance 2.5 covers multi-reference composition and start/end-frame morphs.
  • Google Imagen 4.0 familyimagen-4.0, imagen-4.0-ultra and imagen-4.0-fast removed; use the Nano Banana (Gemini Image) models instead.
  • Also removedbytedance-video-upscaler (superseded by Video Enhance), kling-v2-new-image, and qwen-image-edit-plus (Qwen 2 Pro now covers reference-guided editing from up to 3 images).

Docs

  • Provider pages now render catalog-kind parameters (voice, avatar, and effect-preset pickers) instead of leaving the Values cell blank. These lists are fetched at runtime, so the docs point at gen-ai models info <id> --json rather than snapshotting a list that would rot.
  • The count drift guard now also covers reference/providers/index.md and README.md, and recognises the "N AI model providers" phrasing.
  • The wiki generator keeps pages for retired providers, so a tombstoned page like Pika exists on both the Pages site and the Wiki instead of silently vanishing from one.

2026-08-07

Catalog refresh — now 174 production models across 32 providers (@picsart/ai-sdk 3.35.6).

New

  • Seedance 2.5 — ByteDance's next-generation video model, with companion Seedance 2.5 Video Edit (edit an existing clip from a prompt) and Seedance 2.5 Video Extend (continue a clip past its original end). See Seedance.
  • Seedream 4.7 — text-to-image up to 4K with multi-image reference input and batch counts up to 10. See Seedream.
  • Flux 3 Video — Black Forest Labs' first video model on Picsart: text-to-video and image-to-video up to 20 seconds at FHD, with native audio generation and a fast --draft mode. This makes Flux a multi-mode provider. See Flux.
  • HeyGen Video Avatar — pick a HeyGen avatar and voice, supply a script, and get a lip-synced presenter video up to 4K. Unlike Talking Photo it needs no input image. See HeyGen.

2026-07-31

AI Playground · July 31 Release — catalog refreshed to 168 production models across 32 providers (@picsart/ai-sdk 3.30.0).

New

  • Hailuo 03 — MiniMax text-to-video and image-to-video with start/end frames, multimodal references, 5–15 second clips, and output up to 2K. See MiniMax.
  • Video Enhance — ByteDance video enhancement can denoise, color-correct, super-resolve existing footage up to 8K, and convert frame rate. The SDK definition is gated until its worker rollout completes.
  • Ideogram P-Image — out of preview and generally available, with four speed/quality tiers and output up to 2K. See Ideogram.
  • HEIC support — iPhone and Windows HEIC/HEIF photos now upload with automatic conversion.
  • Catalog additions — Seed Audio and Seed Audio Multilingual add named voices and voice cloning, while Gemini 3.6 Flash and Gemini 3.5 Flash Lite expand text/image analysis. See Seed Audio and Google.

Improvements

  • Empty boards — redesigned with a card-stack hero and a responsive “Need a spark?” inspiration strip.
  • Explore — the Picsart Effects strip now serves the creator and business catalog: about 950 presets across 16 categories.
  • Topaz catalog — Topaz Video Upscale is available with Proteus, Artemis, Nyx, Gaia, and Starlight enhancement families; eight legacy image-enhancement IDs are consolidated as presets on topaz-upscale-image.

2026-07-17

Catalog refresh — now 169 models across 31 providers (@picsart/ai-sdk 3.17.0).

  • New modelsSeedream 5.0 Pro (text-to-image, see Seedream), Nano Banana 2 Lite (gemini-3.1-flash-lite-image, see Google), ElevenLabs Music v2 (full music tracks, see ElevenLabs), and Picsart Angle Change (picsart-qwen-image-edit-angle, image editing, see Picsart).
  • Recraft V4 & V4.1 image-to-image. All 12 V4/V4.1 models (base, Pro, Utility, Utility Pro, and their vector variants) now take an optional source image (-i) with an adjustable image weight (--weight, 0–100, default 80). See Recraft.
  • Happy Horse 1.1 — model ids renamed from happyhorse-1.5-* to happyhorse-1.1-* (same models; update any saved commands).
  • Parameter updatesGrok Imagine Video 1.5 adds 1080p; HeyGen Talking Photo prompt limit raised to 5000 chars; Ideogram Character now takes a required prompt; Seedance 2.0 reference images/videos must be at least 0.4MP; the ByteDance Video Upscaler now rejects sources at or above 1080p up front (it only upscales sub-1080p video).

2026-06-26

Catalog refresh — now 165 models across 31 providers (@picsart/ai-sdk 3.10.0). A big lineup update since the 3.6.2 catalog:

  • Seedance 2.0 family — the full lineup. Seedance 2.0 and Seedance 2.0 Fast (text-to-video with keyframe control and native audio), each with a Video Edit and Video Extend variant, plus the new Seedance 2.0 Mini (T2V + edit + extend). See Seedance.
  • Text & analysis models (new mode) — analyze images and video with LLMs via gen-ai describe: Anthropic (new provider) Claude Opus 4.8 / Sonnet 4.6 / Haiku 4.5, OpenAI GPT-5.5, and Google Gemini 3 Pro (the only one that reads video). See Text & analysis.
  • Google — added the Imagen 4.0 family (4.0 / Ultra / Fast) and Gemini Omni.
  • Happy Horse — added Happy Horse 1.5 (text-to-video and ref-to-video).
  • Deprecated / retired — removed from the catalog: GPT Image 1, OpenAI TTS-1 / TTS-1 HD, the Kling image & multi-image family, LTX Pro / Fast / Retake, Wan 2.6 (T2V / R2V / Image), Seedream 4.0, Runway Gen-3 Turbo & Aleph, Recraft v2 (+ Vector), Qwen v1, and Seedance 1.5 Pro / I2V. These stay resolvable for historical jobs and pricing but are no longer offered for new generations.

New CLI command — gen-ai describe. Analyze an image or video with an LLM (Claude, GPT, or Gemini) and get a text answer — caption, OCR, classify, or summarize a clip. The prompt is optional, video auto-routes to a video-capable model, and output goes to stdout for easy piping. See the CLI Quickstart.


2026-06-19

Try in Playground links. Every model now has a direct Try in Playground link — on each provider page's parameter section and on every card in the interactive Model Catalog — that opens the web Playground with the model preselected.


2026-06-18

Catalog — now 176 models across 30 providers (@picsart/ai-sdk 3.6.2). 36 models added since the previous catalog:

  • Kling — Kling V3 Turbo, plus the Kling image & multi-image family (V2 / V2.1 / V1.5 Image, Multi-Image, Multi-Image V2.1).
  • PixVerse (new provider) — V6 and C1 lines: text-to-video, image-to-video, and Fusion reference-to-video.
  • LTX — LTX Pro, LTX Fast, LTX Retake.
  • Wan — Wan 2.6 (text-to-video), Wan 2.6 Ref-to-Video, Wan 2.6 Image.
  • Luma — Ray 3.2, Ray 3.2 Edit, Ray 3.2 Reframe.
  • OpenAI — GPT Image 1, plus TTS-1 and TTS-1 HD (text-to-speech).
  • Seedance — Seedance 1.5 Pro, Seedance I2V.
  • Runway — Gen-3 Alpha Turbo, Aleph 2.
  • Recraft — Recraft 20B and Recraft 20B Vector.
  • Ideogram — Ideogram 4.0. · Seedream — Seedream 4.0. · Qwen — Qwen. · Grok — Grok Imagine 1.5.
  • Async (new provider) — Async Flash v1.0 (text-to-speech, 100+ voices).
  • Picsart — Remove Background now runs on picsart-sod-v8-2 (replaces the previous picsart-remove-bg).

Docs — Public documentation site launched: CLI, MCP, and Skills guides, a searchable Model Catalog, and a reference page per provider with parameters and CLI + MCP examples.


Looking for the full catalog right now? Browse the Model Catalog or the Providers grid — both are generated from the live catalog.

Built on @picsart/ai-sdk · gen-ai CLI · Picsart MCP · Media Studio · Skills