Model Directory

97 AI Models One Platform

Access the best AI models for image generation, video creation and audio synthesis — all through a single API with pay-as-you-go pricing.

Image Models

35 models available

Nano Banana Pro

ImageUltra

State-of-the-art image generation and editing — Gemini-class

~15 secQuality 96/100

Kling O1

ImageUltra

Kling O1 reasoning model with enhanced cinematic quality

~90sQuality 96/100

GPT-Image 1.5

ImageUltra

GPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail

~15sQuality 95/100

Nano Banana 2

ImagePremium

Latest Gemini 3.1 Flash — fast, cheap, with web search and thinking modes

Quality 95/100

GPT-Image 2.0

ImageUltra

Next-gen image model — near-perfect text rendering (>99%), enhanced photorealism, UI/screenshot-grade detail, and superior multilingual text (JP/KR/HI/BN). Supports optional reference-image-guided generation.

Quality 95/100

GPT-Image 2.0 Edit

ImageUltra

Precise single-image editing with detail preservation. Built on GPT-Image 2.0's natively multimodal architecture.

Quality 94/100

FLUX Kontext Pro

ImagePremium

Advanced image editing and transformation

~10 secQuality 92/100

Cosmos 3 Super

ImagePremium

NVIDIA's flagship photorealistic text-to-image model with optional agentic refinement — generates multiple candidate images per round and rewrites the prompt between rounds for stronger prompt adherence. Top-end physics-aware photorealism, structured-JSON prompt expansion supported.

~12sQuality 90/100

Seedream 5.0 Pro

ImagePremium

Premium image model for precise text-to-image and multi-reference editing — product visuals, design assets, dense layouts, and commercial creative. Edit mode takes up to 10 reference images.

~15sQuality 90/100

Wan 2.7 Image Pro

ImagePremium

Wan 2.7 image model, higher quality tier — unified text-to-image and image editing.

Quality 88/100

Ideogram V4

ImagePremium

Ideogram's V4 — best-in-class in-image text rendering, sharper typography, stronger prompt adherence, and three rendering speeds (TURBO / BALANCED / QUALITY). Native multi-aspect support, no fine-tuning needed for posters, logos, thumbnails, ads.

~6sQuality 88/100

Seedream 5.0 Pro Layer Decomposition

NEW
ImagePremium

Split an image into editable layers — background, subject, and each key visual element as its own transparent PNG. Built on Seedream 5.0 Pro's image understanding, so structure and detail survive the separation. The number of layers depends on the image.

~25sQuality 88/100

Grok Imagine Image 2.0

NEW
ImageStandard

xAI's Grok Imagine Image 2.0 — fast text-to-image with a 2K option, a low/medium quality dial, and thirteen aspect ratios from ultrawide to ultratall. Up to four images per request.

~10sQuality 88/100

MAI-Image-2.5

ImageStandard

Microsoft's MAI-Image-2.5 — general-purpose text-to-image with strong photorealism, prompt adherence, and in-image text rendering. Supports up to 4 images per request and seven aspect ratios.

~8sQuality 84/100

Grok Imagine — Image

ImageStandard

Grok Imagine image variant. Text-to-image plus image-to-image variation editing with strong prompt adherence. Sibling of Grok Imagine Video.

~10sQuality 82/100

Wan 2.7 Image

ImageStandard

Wan 2.7 image model — unified text-to-image and image editing.

Quality 80/100

Z-Image

ImageBudget

Tongyi-MAI's lightweight 6B image model — photorealistic output, sub-second generation, and notably accurate bilingual (English + Chinese) text rendering. Apache-2.0 licensed.

~2sQuality 78/100

Nano Banana 2 Lite

ImageStandard

Fast, cost-efficient Gemini image model — rapid ideation, high-throughput generation, and lightweight editing. Built for concept exploration, social visuals, product drafts, and scaled production.

~15sQuality 78/100

Image Translator

ImageStandard

Translate the text inside an image into another language and get back a clean, localized version — great for localizing posters, menus, screenshots, and product shots.

~15sQuality 78/100

Nano Banana

ImageStandard

Nano Banana standard image generation

Quality 70/100

Nano Banana Edit

ImageStandard

Nano Banana image editing

Quality 70/100

FLUX 2 Pro Edit

ImageStandard

FLUX 2 Pro image-to-image editing

Quality 70/100

FLUX 2 Flex

ImageStandard

FLUX 2 Flex fast text-to-image generation

Quality 70/100

Seedream 3.0

ImageStandard

Seedream 3.0 text-to-image

Quality 70/100

Seedream 4.0

ImageStandard

Seedream 4.0 text-to-image

Quality 70/100

Seedream 4.5

ImageStandard

Seedream 4.5 text-to-image with improved quality

Quality 70/100

Seedream 5.0 Lite

ImageStandard

Seedream 5.0 Lite fast text-to-image

Quality 70/100

Ideogram V3

ImageStandard

Ideogram V3 text-to-image with excellent typography

Quality 70/100

Ideogram V3 Edit

ImageStandard

Ideogram V3 image editing with text support

Quality 70/100

GPT-4o Image

ImageStandard

GPT-4o multimodal image generation

Quality 70/100

Qwen Image

ImageStandard

Qwen text-to-image generation

Quality 70/100

Qwen 2 Image

ImageStandard

Qwen 2 improved text-to-image

Quality 70/100

Topaz Upscale

ImageStandard

Topaz AI image upscaling with detail enhancement

Quality 70/100

Recraft Remove BG

ImageStandard

Recraft AI background removal

Quality 70/100

Recraft Upscale

ImageStandard

Recraft AI crisp image upscaling

Quality 70/100

Video Models

44 models available

Sora 2 Pro

VideoUltra

Sora 2 with enhanced quality and longer durations

~90sQuality 99/100

Sora 2

VideoUltra

State-of-the-art video generation model with exceptional quality and understanding

~60sQuality 98/100

Veo 3.1

VideoUltra

Premium video generation model with native audio

~60 secQuality 98/100

Seedance 2.5

NEW
VideoUltra

ByteDance's newest video model. Generates up to 30 seconds in a single shot with synchronized audio — dialogue, music, and effects in the same pass — from a prompt, a start frame, or up to 50 image, video, and audio references.

~3-6 minQuality 97/100

Kling 3.0

VideoUltra

Latest Kling flagship with best-in-class motion and cinematic quality

~60sQuality 96/100

MiniMax H3

NEW
VideoUltra

MiniMax's next-generation multimodal video model (Hailuo-03). Generates up to 2K video with native stereo audio from a text prompt or a reference image, with strong instruction following, motion, and readable on-screen text.

~90sQuality 95/100

Gemini Omni Flash

VideoPremium

Google's any-input video model. Create 4-10 second clips with native audio from a prompt, an image, three reference images, or an existing video — then reshape scenes with plain language. Grounded in Gemini's real-world knowledge for coherent physics and motion.

~2-4 minQuality 95/100

Kling 2.1 Master

VideoUltra

Latest Kling with best-in-class motion and quality

~60sQuality 94/100

Grok Imagine Video 1.5

VideoUltra

Grok Imagine Video 1.5 — image-to-video with synchronized audio, strong prompt adherence, and consistent visual quality. Currently ranked #1 on the Arena blind test leaderboard for image-to-video, surpassing Seedance 2.

~60sQuality 94/100

Veo 3.1 Fast

VideoPremium

Faster, more cost-effective version of Veo 3.1

~30 secQuality 93/100

Avatar X

NEW
VideoPremium

Mirage's most advanced avatar model, with industry-leading identity preservation and expressivity. Type a script and a lifelike stock presenter delivers it — or drive any face from a photo, a clip, or your own voice track, up to 3 minutes in one take.

~1-3 minQuality 93/100

FLUX 3

NEW
VideoPremium

Black Forest Labs' FLUX 3 — text-to-video with synchronized audio generated by the model itself. Up to 20 seconds at 1080p, no reference image required.

~2-4 minQuality 93/100

Kling 2.0 Master

VideoUltra

Kling 2.0 with significantly improved quality and motion

~60sQuality 92/100

Wan 2.7

VideoPremium

Wan 2.7 — text-to-video and image-to-video (first-frame / first-last-frame / continuation).

Quality 92/100

Kling 1.5 Pro

VideoPremium

High quality video generation

~60 secQuality 90/100

Runway Gen 4.5

VideoPremium

Newer Runway Gen 4.5 — text-to-video with optional reference image, 5/10s.

Quality 90/100

Seedance 2.0

VideoPremium

Seedance 2.0 — text-to-video and image-to-video.

Quality 90/100

Runway Gen-3 Turbo

VideoPremium

Fast video generation from Runway

~20sQuality 88/100

Luma Ray 2

VideoPremium

Cinematic video generation

~90 secQuality 88/100

Sync LipSync

VideoPremium

Professional lip sync

~30 secQuality 88/100

Kling 3.0 Motion Control

VideoPremium

Reference-image + reference-video motion transfer. Pick whose pose drives the output via the character orientation toggle.

~90 secQuality 88/100

Kling 3.0 Turbo

VideoPremium

Speed-optimised Kling 3.0 — faster generation at high visual quality. Text-to-video (up to 2,500-char prompts), first-frame image-to-video, and multi-shot storyboarding (up to 6 prompted shots in one clip). Standard tier is 720p, Pro tier is 1080p.

~60sQuality 88/100

Kling 1.6 Pro

VideoPremium

Latest Kling 1.x with enhanced quality

~50sQuality 87/100

Luma Dream Machine

VideoPremium

High-quality video generation from Luma AI

~40sQuality 86/100

Happy Horse 1.1

VideoPremium

Upgraded cinematic short-video model — smoother motion, stronger instruction following, native audio behaviour, and multilingual lip-sync. Text-to-video, first-frame image-to-video, and reference-to-video (up to 9 reference images).

~90sQuality 86/100

Vidu 2

VideoPremium

High-quality video generation with strong motion

~50sQuality 85/100

Volcengine Lip Sync

VideoStandard

Video-to-video lip sync — aligns mouth movements in an existing video with target audio. Multilingual dubbing supported.

~60sQuality 85/100

Happy Horse

VideoPremium

Cinematic short-video model — superior motion, film-grade lighting, strong prompt adherence, and multi-shot storytelling. Single text-to-video + image-to-video endpoint.

~90sQuality 84/100

Video Upscaler

VideoStandard

Enlarge and sharpen a video to a higher resolution. Choose a scale factor from 1x to 8x (default 2x).

~1 minQuality 84/100

Pika 2

VideoPremium

Pika's latest video generation model

~30sQuality 83/100

MiniMax Video 01

VideoStandard

Fast video generation from images

~45 secQuality 82/100

Video Background Removal

VideoStandard

Remove the background from a video and get a clean subject cutout, ready to composite onto any scene. Optional audio passthrough and background color.

~30sQuality 82/100

LatentSync

VideoStandard

Latent-space lip sync model

~45sQuality 80/100

Seedance 2.0 Mini

VideoStandard

Fast, budget-friendly video model — text-to-video and image-to-video (first frame, last frame, or reference) with optional synced audio. Great for quick drafts and high-volume social clips.

~30sQuality 80/100

Video Translator

VideoStandard

Translate short videos — transcribes the speech and generates dubbed audio plus subtitles in your target language. Up to 6 minutes per clip.

~1-2 minQuality 80/100

Music Video

VideoStandard

Render a shareable video for a track, with optional artist and site credits.

~1-2 minQuality 80/100

Seedance 2.0 Fast

VideoStandard

Faster, cheaper variant of Seedance 2.0 — text-to-video and image-to-video.

Quality 78/100

Kling 1.0 Standard

VideoStandard

Standard-tier video generation model

~30sQuality 75/100

Veo 3.1 Lite

VideoStandard

Veo 3.1 Lite — budget-friendly video generation with text-to-video, image-to-video, and reference-to-video modes

Quality 70/100

Hailuo 2.3 Pro

VideoStandard

MiniMax Hailuo 2.3 Pro high-quality video generation

Quality 70/100

Hailuo 2.3 Standard

VideoStandard

MiniMax Hailuo 2.3 Standard video generation

Quality 70/100

Hailuo 02 Pro

VideoStandard

MiniMax Hailuo 02 Pro video generation

Quality 70/100

Wan 2.6 Flash

VideoStandard

Wan 2.6 Flash fast image-to-video

Quality 70/100

Topaz Video Upscale

VideoStandard

Topaz AI video upscaling

Quality 70/100

Audio Models

16 models available

ElevenLabs Multilingual V2

AudioPremium

High-quality multilingual text-to-speech

~4sQuality 95/100

OpenAI TTS HD

AudioPremium

Premium text-to-speech model

~3 secQuality 92/100

Suno Music

AudioPremium

AI music generation — create songs, instrumentals, and soundtracks

~60sQuality 90/100

ElevenLabs Turbo V2.5

AudioStandard

Fast text-to-speech with good quality

~2sQuality 88/100

OpenAI TTS

AudioStandard

Text-to-speech model

~3sQuality 85/100

F5-TTS

AudioStandard

High quality text-to-speech

~5 secQuality 85/100

Gemini 3.1 Flash TTS

AudioStandard

Gemini 3.1 Flash text-to-speech. 30 voice presets, 80+ languages, inline expressive tags ([sigh], [laughing]), style instructions up to 4K chars, and multi-speaker support.

Quality 80/100

Vocal Separation

AudioStandard

Split a track into vocal and instrumental stems, or into up to twelve individual instrument stems.

~1-2 minQuality 80/100

Extend Music

AudioStandard

Continue an existing track from any point, keeping its style.

~1-3 minQuality 80/100

Add Vocals

AudioStandard

Sing over an instrumental you upload.

~1-3 minQuality 80/100

Add Instrumental

AudioStandard

Build a backing track under vocals you upload.

~1-3 minQuality 80/100

Replace Section

AudioStandard

Re-generate one time range of a track, leaving the rest untouched.

~1-3 minQuality 80/100

Boost Music Style

AudioStandard

Expand a short style description into a richer prompt for better results.

~5 sQuality 80/100

Convert to WAV

AudioStandard

Export a track as lossless WAV.

~10 sQuality 80/100

Music Cover

AudioStandard

Generate a new interpretation of an existing track.

~1-2 minQuality 80/100

Generate MIDI

AudioStandard

Transcribe separated stems into MIDI note data.

~1-2 minQuality 80/100

3D Models

2 models available