HomeModels

AI Models

Discover, compare and find the best AI models for your creative projects

Filter
172 models found
FLUX Fill icon

FLUX Fill

Black Forest Labs|12B

FLUX Fill is the specialized inpainting and outpainting model within the FLUX model family developed by Black Forest Labs, designed for professional-grade region editing, content filling, and image extension. Built on the 12-billion parameter Diffusion Transformer architecture that powers all FLUX models, FLUX Fill takes an input image along with a binary mask indicating the region to be modified and generates seamlessly blended content that matches the surrounding context in style, lighting, perspective, and detail level. The model excels at both inpainting tasks where masked areas within an image are filled with contextually appropriate content and outpainting tasks where image boundaries are extended to create larger compositions. FLUX Fill leverages the superior prompt adherence of the FLUX architecture, allowing users to guide the generation with text descriptions of what should appear in the masked region, providing precise creative control over the output. The model handles complex scenarios including filling regions that span multiple materials and textures, maintaining structural continuity of architectural elements, and generating photorealistic human features in masked face areas. As a proprietary model, FLUX Fill is accessible through Black Forest Labs' API and partner platforms including Replicate and fal.ai, with usage-based pricing. Professional photographers use FLUX Fill for removing unwanted elements and extending compositions, e-commerce teams employ it for product background replacement, digital artists leverage it for creative compositing, and marketing professionals use it for adapting images to different aspect ratios and formats without losing content quality.

Proprietary
4.7
RVC v2 icon

RVC v2

RVC Project|40M

RVC v2 (Retrieval-based Voice Conversion v2) is an open-source AI model for real-time voice conversion that transforms one person's voice into another person's voice while preserving the original speech content, intonation patterns, and emotional expressiveness. Built on a VITS architecture enhanced with a retrieval-based approach, the model with approximately 40 million parameters uses a feature index to find and match the closest vocal characteristics from the target speaker's training data, resulting in highly natural and artifact-free voice transformations. RVC v2 requires only 10 to 20 minutes of clean audio from the target speaker to train a voice model, making it one of the most accessible voice cloning solutions available. The model operates in real-time with latencies suitable for live streaming and voice chat applications, processing audio at faster than real-time speeds on modern consumer GPUs. Key improvements in v2 over the original version include reduced breathiness artifacts, better pitch tracking with the RMVPE algorithm, enhanced consonant clarity, and support for 48kHz output quality. Released under the MIT license, RVC v2 has become the most widely used open-source voice conversion tool with an extensive community providing pre-trained voice models, training guides, and integration plugins. Common applications include content creation with character voices, music cover generation in different vocal styles, voice privacy and anonymization, accessibility tools for speech-impaired users, and creative audio production. The model integrates with OBS, Discord, and various DAW software for streamlined production workflows.

Open Source
4.4
Kling 2.0 icon

Kling 2.0

Kuaishou Technology|undisclosed

Kling 2.0 is Kuaishou Technology's latest video generation model, released in January 2025, representing a major upgrade in video quality, motion realism, and generation capabilities over its predecessor Kling 1.5. The model generates video clips at up to 1080p resolution with dramatically improved physical simulation, human motion accuracy, and scene consistency. Kling 2.0 introduces a Master Mode for highest-quality cinematic generation with enhanced attention to lighting, depth of field, and camera cinematography. The model supports both text-to-video and image-to-video generation with clip durations up to 10 seconds in standard mode and 5 seconds in Master Mode. Notable improvements include better hand rendering, more natural facial expressions, smoother camera movements, and more physically accurate object interactions. The model processes complex scene descriptions with multiple subjects and dynamic interactions, generating videos where physical laws are more consistently maintained. Available through the Kling AI web platform and mobile app, the model offers daily free generations with premium plans for higher quality, longer clips, and commercial usage. Kling 2.0 competes with Runway Gen-3, Sora, and Veo 2 as one of the leading AI video generation models.

Proprietary
4.7
Leonardo AI icon

Leonardo AI

Leonardo AI|N/A

Leonardo AI is a comprehensive AI image generation platform that offers multiple fine-tuned models optimized for specific creative domains including game assets, character design, concept art, and product photography. Unlike single-model solutions, Leonardo provides a suite of specialized models such as Leonardo Diffusion XL, Leonardo Vision XL, and DreamShaper that users can select based on their specific needs. The platform features an intuitive web interface with built-in tools for real-time canvas editing, AI-powered image guidance, texture generation for 3D assets, and motion generation capabilities. Leonardo's model training pipeline allows users to create custom fine-tuned models using their own datasets, enabling brand-specific or style-specific image generation with as few as 10 training images. The platform particularly excels in game development workflows, offering dedicated models for generating consistent game environments, characters, items, and UI elements. It supports ControlNet-style image conditioning, inpainting, outpainting, and prompt enhancement features. Leonardo AI operates on a freemium model with daily token allocations for free users and premium subscription tiers for higher volume needs. Game developers, indie studios, concept artists, e-commerce businesses, and social media content creators form its primary user base. The API access enables integration into production pipelines for automated content generation at scale. Leonardo AI positions itself as an all-in-one creative platform rather than just a model, differentiating through its combination of multiple specialized models, training capabilities, and integrated editing tools.

Proprietary
4.5
Kling Image-to-Video icon

Kling Image-to-Video

Kuaishou|N/A

Kling Image-to-Video is the image animation mode of Kuaishou's Kling video generation platform, designed to create video content from reference images with natural motion, temporal coherence, and high visual fidelity. Released in June 2024 as part of the Kling 1.5 suite, this capability allows users to provide a still image as a starting frame and generate video sequences that animate the scene with contextually appropriate motion. The model leverages Kling's transformer-based architecture to understand spatial composition, depth relationships, and semantic content of the input image, then generates plausible temporal evolution maintaining consistency with the source. Kling Image-to-Video demonstrates strength in animating human subjects with realistic facial expressions, body movements, and clothing dynamics, as well as generating environmental motion such as wind effects, water flow, and atmospheric changes. The model supports various output durations and resolutions for different creative and commercial applications from short social media animations to longer-form content. Users can provide optional text prompts alongside the reference image to guide the direction of generated motion, offering additional creative control. The model handles diverse input types including photographs, digital artwork, illustrations, and rendered scenes, applying motion patterns respecting the visual style and physical properties of the source. As a proprietary service, Kling Image-to-Video is accessible through Kuaishou's platform and through fal.ai and Replicate, enabling integration into custom creative tools and production pipelines for professional content creators.

Proprietary
4.6
Ideogram 2.0 icon

Ideogram 2.0

Ideogram|N/A

Ideogram 2 is a text-to-image generation model developed by Ideogram AI that has established itself as the industry benchmark for typography and text rendering within AI-generated images. While most image generation models struggle with producing legible, accurately spelled text, Ideogram 2 consistently generates high-quality typography that integrates naturally into images across diverse contexts including posters, logos, book covers, and social media graphics. The model builds upon the success of its predecessor with enhanced photorealistic capabilities, improved compositional accuracy, and better understanding of complex multi-element prompts. Ideogram 2 supports multiple artistic styles ranging from photorealism and 3D rendering to illustration, anime, and graphic design aesthetics. The model is accessible through the Ideogram web platform and API, offering both free and premium subscription tiers. Its architecture incorporates specialized attention mechanisms for text positioning and rendering that go beyond standard diffusion model capabilities. Graphic designers, social media managers, marketing professionals, and small business owners particularly value Ideogram 2 for creating branded content, promotional materials, and designs that require integrated typography without post-processing in external tools. The model also performs well in general image generation tasks, producing detailed and coherent images across various subjects and styles. Its unique strength in text rendering fills a critical gap in the AI image generation landscape that competitors have not yet matched consistently.

Proprietary
4.7
IP-Adapter icon

IP-Adapter

Tencent|22M

IP-Adapter is an image prompt adapter developed by Tencent AI Lab that enables image-guided generation for text-to-image diffusion models without requiring any fine-tuning of the base model. The adapter works by extracting visual features from reference images using a CLIP image encoder and injecting these features into the diffusion model's cross-attention layers through a decoupled attention mechanism. This allows users to provide reference images as visual prompts alongside text prompts, guiding the generation process to produce images that share stylistic elements, compositional features, or visual characteristics with the reference while still following the text description. IP-Adapter supports multiple modes of operation including style transfer, where the generated image adopts the artistic style of the reference, and content transfer, where specific subjects or elements from the reference appear in the output. The adapter is lightweight, adding minimal computational overhead to the base model's inference process. It can be combined with other control mechanisms like ControlNet for multi-modal conditioning, enabling sophisticated workflows where pose, style, and content can each be controlled independently. IP-Adapter is open-source and available for various Stable Diffusion versions including SD 1.5 and SDXL. It integrates with ComfyUI and Automatic1111 through community extensions. Digital artists, product designers, brand managers, and content creators who need to maintain visual consistency across generated images or transfer specific aesthetic qualities from reference material particularly benefit from IP-Adapter's capabilities.

Open Source
4.6
Veo 2 icon

Veo 2

Google DeepMind|N/A

Veo 2 is Google DeepMind's most advanced video generation model, capable of producing high-quality video content with up to 4K resolution, representing the cutting edge of AI-powered video synthesis. Released in December 2024, Veo 2 builds upon Google's extensive research in video understanding, delivering significant improvements in visual fidelity, motion realism, temporal coherence, and prompt comprehension. The model supports both text-to-video and image-to-video modes, interpreting detailed descriptions to create sequences that accurately reflect specified scenes, characters, actions, and atmospheric conditions. Veo 2 demonstrates exceptional understanding of real-world physics, generating videos with realistic lighting, shadows, reflections, and material properties. The model handles complex cinematic concepts including depth of field, camera movements like dolly shots and crane movements, and advanced compositional techniques, enabling footage that rivals professional cinematography. Veo 2 excels at maintaining character consistency across extended sequences, generating natural human motion and facial expressions, and producing content in diverse styles from photorealistic footage to animation and artistic interpretations. The model supports longer video sequences compared to most competitors, with improved temporal stability that reduces flickering and morphing artifacts. As a proprietary model, Veo 2 is currently available through limited access channels within Google's ecosystem, with plans for broader integration into Google products. The model represents Google's strategic positioning in the competitive AI video generation landscape alongside OpenAI's Sora and Runway's Gen-3 Alpha.

Proprietary
4.8
Upscayl icon

Upscayl

Upscayl Team|N/A

Upscayl is a free and open-source desktop application for AI-powered image upscaling, built on top of Real-ESRGAN and other super-resolution models. Developed by Nayam Amarshe and TGS963, Upscayl provides a user-friendly graphical interface that makes advanced AI image upscaling accessible to non-technical users on Windows, macOS, and Linux platforms. The application wraps multiple AI upscaling models in an Electron-based desktop app, allowing users to enhance image resolution with just a few clicks without any command-line knowledge or Python environment setup. Upscayl includes several pre-installed upscaling models optimized for different content types including general photography, digital art, anime, and sharpening, with each model producing different aesthetic characteristics suited to its target content. Users can select upscaling factors of 2x, 3x, or 4x and process individual images or entire folders through batch processing. The application supports common image formats including PNG, JPG, and WebP, and provides options for output format and quality settings. Upscayl also supports custom model loading, allowing users to import additional NCNN-compatible upscaling models from the community. Released under the AGPL-3.0 license, Upscayl is fully open source with its code available on GitHub and has accumulated a large community of users and contributors. The application runs entirely locally with no internet connection required, ensuring privacy for sensitive images. Upscayl is particularly popular among photographers, graphic designers, content creators, and hobbyists who need a simple, free solution for enhancing image quality without subscriptions or cloud processing dependencies.

Open Source
4.5
SD Inpainting icon

SD Inpainting

Stability AI|1B

Stable Diffusion Inpainting is a specialized variant of Stability AI's Stable Diffusion model fine-tuned specifically for image inpainting tasks, enabling users to fill masked regions of an image with contextually coherent content guided by text prompts. Released in 2022, the model builds upon the latent diffusion architecture but extends it with additional input channels for mask-aware processing, where the original image, mask, and masked image are fed as extra channels to the U-Net. The v1.5 inpainting model was trained on 595K curated inpainting examples in collaboration with RunwayML, while community-developed SDXL variants have since extended capabilities with higher resolution output. Common applications include removing unwanted objects from photographs, completing damaged image regions, modifying content such as adding elements to scenes, and cleaning watermarks or text overlays. Professional use cases span photography post-production, advertising visual preparation, real estate staging, product photography background replacement, and digital art workflows. The model is accessible through popular open-source interfaces including AUTOMATIC1111 WebUI, ComfyUI, InvokeAI, and the Hugging Face Diffusers library. Users can create masks manually with brush tools or automatically through segmentation models like SAM. ControlNet integration adds additional control layers for more precise output guidance. Released under the CreativeML Open RAIL-M license, the model runs on GPUs with 8GB VRAM and supports optimizations like xFormers for reduced memory usage, making it one of the most widely adopted open-source inpainting solutions available.

Open Source
4.4
BRIA RMBG icon

BRIA RMBG

BRIA AI|N/A

BRIA RMBG is a state-of-the-art background removal model developed by BRIA AI, an Israeli startup specializing in responsible and commercially licensed generative AI. The model delivers exceptional accuracy in separating foreground subjects from backgrounds, handling complex scenarios including fine hair details, transparent objects, intricate edges, smoke, and glass with remarkable precision. BRIA RMBG is built on a proprietary architecture trained on exclusively licensed and ethically sourced data, ensuring full commercial safety and IP compliance that distinguishes it from models trained on scraped internet data. It produces high-quality alpha mattes preserving fine edge details and natural transparency gradients for clean cutouts suitable for professional workflows. Available in versions including RMBG 1.4 and RMBG 2.0, the model consistently ranks among top performers on background removal benchmarks including DIS5K and HRS10K datasets. BRIA RMBG is accessible through Hugging Face with a permissive license for research and commercial use, and through BRIA's commercial API for scalable cloud processing. Integration options include Python SDK, REST API, and popular image processing pipeline compatibility. Applications span e-commerce product photography, graphic design compositing, video conferencing virtual backgrounds, automotive and real estate photography, social media content creation, and document digitization. The model processes images in milliseconds on modern GPUs, suitable for real-time and high-volume batch processing. BRIA RMBG has established itself as one of the most commercially trusted and technically advanced background removal solutions available.

Open Source
4.7
This Person Does Not Exist icon

This Person Does Not Exist

Philip Wang|N/A

This Person Does Not Exist is a web-based demonstration created by Uber software engineer Philip Wang that generates photorealistic portraits of entirely fictional people using NVIDIA's StyleGAN technology. Launched in February 2019, the website became a viral sensation by producing a new AI-generated human face each time the page is refreshed, showcasing the capability of generative adversarial networks to synthesize convincing portraits indistinguishable from real photographs. The underlying model was trained on the FFHQ dataset containing 70,000 high-resolution photographs of real human faces, learning to generate novel facial compositions with realistic skin textures, hair patterns, lighting, eye reflections, and natural asymmetries. The generated faces span diverse demographics including various ages, ethnicities, and genders, demonstrating the model's understanding of facial diversity. While outputs are convincing at first glance, careful examination occasionally reveals telltale artifacts such as asymmetric earrings, distorted backgrounds, or inconsistencies in hair at image edges. The project serves multiple purposes beyond demonstration: it has been widely used in discussions about deepfake technology and media literacy, serves as a privacy-preserving source of placeholder portraits for design mockups and UI prototyping, and provides stock-photo-like imagery without licensing concerns. The website itself is proprietary, though the underlying StyleGAN architecture is open source. This Person Does Not Exist remains one of the most recognized public demonstrations of GAN capabilities and continues to spark conversations about AI-generated media authenticity and digital trust in an era of increasingly sophisticated synthetic content.

Proprietary
4.3
OpenPose icon

OpenPose

CMU|25M

OpenPose is the pioneering real-time multi-person pose estimation system developed at Carnegie Mellon University that simultaneously detects body, face, hand, and foot keypoints of multiple people in images and videos. As the first open-source system to achieve real-time multi-person pose detection, OpenPose has become a foundational tool in computer vision research and creative AI applications. Built on a CNN (Convolutional Neural Network) architecture with approximately 25 million parameters, the model uses Part Affinity Fields (PAFs) to associate detected body parts with the correct individuals in crowded scenes, enabling accurate pose estimation even when people overlap or partially occlude each other. OpenPose detects up to 135 keypoints per person covering the full body skeleton with 25 points, each hand with 21 points, and the face with 70 points, providing comprehensive pose information for detailed motion analysis. The system processes both images and video streams, delivering real-time performance on modern GPUs that makes it suitable for interactive applications. OpenPose has been extensively integrated into AI image generation workflows, particularly as the standard pose extraction method for ControlNet conditioning in Stable Diffusion and FLUX-based generation pipelines. Released under a custom non-commercial license, the source code is available on GitHub and has accumulated one of the highest star counts among computer vision repositories. Key applications include motion capture for animation and gaming, fitness and rehabilitation tracking, sports biomechanics analysis, sign language recognition, dance analysis, human-computer interaction research, and providing pose conditioning for AI image generation tools.

Open Source
4.3
Depth Anything v2 icon

Depth Anything v2

TikTok / ByteDance|25M-335M

Depth Anything v2 is a state-of-the-art monocular depth estimation model developed by TikTok and ByteDance researchers as a significant upgrade to the original Depth Anything. The model extracts precise depth maps from single RGB images without requiring stereo pairs or specialized depth sensors. Built on a DINOv2 vision foundation model backbone combined with a DPT (Dense Prediction Transformer) decoder head, Depth Anything v2 achieves remarkable improvements in fine-grained detail preservation and edge sharpness compared to its predecessor. The model comes in three scale variants ranging from 25 million to 335 million parameters, offering flexible trade-offs between accuracy and inference speed for different deployment scenarios. A key innovation in v2 is the use of large-scale synthetic training data generated from precise depth sensors combined with pseudo-labeled real images, which significantly reduces the noise and artifacts common in earlier monocular depth models. The model produces both relative and metric depth estimates, making it suitable for diverse applications from 3D scene reconstruction and augmented reality to autonomous navigation and robotics. Released under the Apache 2.0 license, it is fully open source and available through Hugging Face with pre-trained checkpoints. Depth Anything v2 integrates naturally with creative AI workflows including ControlNet depth conditioning for Stable Diffusion and FLUX, enabling artists and developers to generate depth-aware compositions. It also supports video depth estimation with temporal consistency, making it valuable for visual effects production and spatial computing applications.

Open Source
4.6
FLUX.1 [pro] Ultra icon

FLUX.1 [pro] Ultra

Black Forest Labs|12B+

FLUX.1 [pro] Ultra is Black Forest Labs' highest-quality image generation model, released in November 2024, delivering native 4-megapixel resolution images with exceptional detail, photorealism, and artistic versatility. The model generates images at up to 2048x2048 pixels natively without upscaling, producing outputs suitable for professional print, large-format displays, and commercial photography applications. FLUX.1 [pro] Ultra achieves the highest quality scores in the FLUX family, surpassing even FLUX.1 Pro in benchmark evaluations. The model demonstrates exceptional performance in complex scene composition, lighting accuracy, material rendering, and text generation within images. Available exclusively through API access on platforms including Replicate and fal.ai, FLUX.1 [pro] Ultra serves professional photographers, advertising agencies, design studios, and enterprise content teams requiring the highest possible AI image generation quality.

Proprietary
4.9
Wan Video 2.1 icon

Wan Video 2.1

Alibaba|14B

Wan Video 2.1 is Alibaba's open-source video generation model combining high visual quality with controllable generation capabilities, making it one of the most capable freely available video synthesis solutions. Built on a diffusion transformer architecture, it supports text-to-video and image-to-video generation with enhanced temporal consistency, smooth motion, and improved visual fidelity compared to earlier open-source video models. Wan Video 2.1 introduces controllability features allowing users to guide generation through conditioning signals beyond text prompts, including motion control, camera trajectory specification, and reference image styling, providing creative control approaching proprietary solutions. The model handles diverse content from realistic human motion to natural landscapes, architectural environments, and stylized artistic content with consistent quality. Multiple model variants with different parameter counts are available for various hardware capabilities, from lightweight versions for consumer GPUs to full-scale models for maximum quality. The Apache 2.0 open-source license encourages community extensions, custom fine-tuning, and integration into creative pipelines. Wan Video 2.1 runs locally without cloud dependencies, ensuring data privacy and eliminating subscription costs. Applications include social media content creation, advertising video production, film concept visualization, educational materials, and creative experimentation. The model is available through Hugging Face with documentation and integration with ComfyUI and Diffusers. Wan Video 2.1 positions Alibaba as a major contributor to the open-source video generation ecosystem, providing a competitive alternative to proprietary models from Runway, Google, and OpenAI.

Open Source
4.5
LivePortrait icon

LivePortrait

Kuaishou|Unknown

LivePortrait is an efficient AI portrait animation model developed by Kuaishou Technology that generates expressive and lifelike facial animations from a single static portrait photograph. The model takes a source portrait image and a driving video containing facial movements, then transfers the expressions, head rotations, eye movements, and mouth gestures from the video onto the portrait while maintaining the original person's identity and appearance. Built on an implicit keypoint detection architecture with warping-based rendering, LivePortrait achieves real-time inference speeds that make it practical for interactive applications and live content creation. The model introduces stitching and retargeting modules that prevent common artifacts in portrait animation such as face boundary distortion, neck disconnection, and unnatural eye movements, producing seamless results that preserve the natural appearance of the subject. LivePortrait handles diverse portrait types including photographs, paintings, illustrations, and even cartoon characters, adapting its animation approach to different artistic styles. The model supports fine-grained control over individual facial action units, allowing selective animation of specific facial features like eyebrow raises, eye blinks, or smile intensity independently. Released under the MIT license, LivePortrait is fully open source and has been integrated into ComfyUI and other creative tools. Common applications include creating animated avatars for social media and messaging, producing animated portrait NFTs, generating facial animations for virtual presenters and digital humans, creating engaging content from historical photographs, and building interactive portrait experiences for museums and exhibitions.

Open Source
4.5
Imagen 3 icon

Imagen 3

Google DeepMind|undisclosed

Imagen 3 is Google DeepMind's most advanced text-to-image generation model, representing a significant leap in photorealistic image quality, prompt understanding, and visual detail compared to its predecessors. Released in August 2024 through Google's Vertex AI platform and ImageFX interface, Imagen 3 generates images with exceptional photographic quality, accurate lighting, natural skin textures, and precise spatial relationships. The model demonstrates remarkable improvement in text rendering within images, accurately generating legible text on signs, labels, and surfaces. Imagen 3 excels at understanding complex compositional prompts, correctly interpreting spatial relationships like 'next to,' 'behind,' and 'above' with higher accuracy than competing models. The model incorporates Google's SynthID digital watermarking technology that embeds invisible identifiers into generated images for provenance tracking. Available through Google Cloud's Vertex AI API and the consumer-facing ImageFX web application, Imagen 3 serves both enterprise developers and creative professionals. The model supports various aspect ratios and generates images up to 1024x1024 pixels natively, with upscaling capabilities for higher resolutions. Safety features include built-in content filters and responsible AI guardrails designed to prevent harmful content generation. Imagen 3 competes directly with DALL-E 3, Midjourney v6, and FLUX.1 Pro in the premium image generation segment, with particular strengths in photorealism and compositional accuracy.

Proprietary
4.7
Nano Banana Pro icon

Nano Banana Pro

Google DeepMind|undisclosed

Nano Banana Pro (technically Gemini 3 Pro Image) is Google DeepMind's premium image generation model released in November 2025, offering studio-grade visual quality with up to 4K resolution output. It builds upon the original Nano Banana with dramatically improved capabilities including multi-language readable text rendering for infographics and marketing materials, face consistency for up to 5 different people across multiple images, advanced creative controls for camera angles, lighting, depth of field, and color grading, and web search grounding for real-time data integration. The model features a 'thinking mode' for complex prompts that require multi-step reasoning, and supports localized image editing to modify specific regions without affecting the rest. Scoring 9.0+ out of 10 for photorealism and ranking as the second-best model for text generation after GPT Image, it delivers professional-grade results suitable for commercial design, marketing, and brand content creation. Available through the Gemini app (free with watermark), Google AI Studio, Gemini API, and Google Ads, with pricing starting at $0.134 per 1K image through the API.

Proprietary
4.6
AnimateDiff icon

AnimateDiff

Yuwei Guo|N/A

AnimateDiff is a motion module framework developed by Yuwei Guo that transforms any personalized text-to-image diffusion model into a video generator by inserting learnable temporal attention layers into the existing architecture. Released in July 2023, AnimateDiff introduced a groundbreaking approach by decoupling motion learning from visual appearance learning, allowing users to leverage the vast ecosystem of fine-tuned Stable Diffusion models and LoRA adaptations for video creation without retraining. The core innovation is a plug-and-play motion module that learns general motion patterns from video data and can be inserted into any Stable Diffusion checkpoint to animate its outputs while preserving visual style and quality. The motion module consists of temporal transformer blocks with self-attention across frames, generating temporally coherent sequences with natural object movement. AnimateDiff supports both SD 1.5 and SDXL base models with optimized motion module versions for each architecture. The framework enables generation of animated GIFs and short video loops with customizable frame counts, frame rates, and motion intensities. Users can combine AnimateDiff with ControlNet for pose-guided animation, IP-Adapter for reference-based motion, and various LoRA models for style-specific video generation. Common applications include animated artwork, social media content, game asset animation, product visualization, and creative storytelling. Available under the Apache 2.0 license, AnimateDiff is accessible on Hugging Face, Replicate, and fal.ai, with extensive community support through ComfyUI workflows and Automatic1111 extensions. The framework has become one of the most influential open-source video generation approaches, enabling creators to produce stylized animated content with unprecedented flexibility.

Open Source
4.5
XTTS v2 icon

XTTS v2

Coqui AI|467M

XTTS v2 (Cross-lingual Text-to-Speech v2) is a multilingual voice cloning and text-to-speech model developed by Coqui AI that can replicate any person's voice from just a 6-second audio sample and synthesize speech in 17 supported languages. Built on a GPT-like autoregressive architecture paired with a HiFi-GAN vocoder, XTTS v2 with 467 million parameters produces natural-sounding speech with realistic prosody, intonation, and emotional expressiveness. The model's cross-lingual capability allows a voice cloned from an English sample to speak fluently in French, Spanish, German, Turkish, and other supported languages while maintaining the original speaker's vocal characteristics. XTTS v2 achieves this through a language-agnostic speaker embedding space that separates voice identity from linguistic content. The synthesis quality approaches human-level naturalness for many languages, with particularly strong performance in English, Spanish, and Portuguese. The model supports streaming inference for real-time applications, generating speech with latencies suitable for conversational AI and interactive voice assistants. Released under the MPL-2.0 license, XTTS v2 is open source and can be deployed locally for privacy-sensitive applications. Common use cases include creating multilingual audiobook narrations, localizing video content with consistent voice identity, building accessible text-to-speech interfaces, developing custom voice assistants, podcast production, and e-learning content creation. The model provides a Python API and can be fine-tuned on additional voice data for improved quality with specific speakers or specialized domains.

Open Source
4.5
IP-Adapter FaceID icon

IP-Adapter FaceID

Tencent|22M (adapter)

IP-Adapter FaceID is a specialized adapter module developed by Tencent AI Lab that injects facial identity information into the diffusion image generation process, enabling the creation of new images that faithfully preserve a specific person's facial features. Unlike traditional face-swapping approaches, IP-Adapter FaceID extracts face recognition feature vectors from the InsightFace library and feeds them into the diffusion model through cross-attention layers, allowing the model to generate diverse scenes, styles, and compositions while maintaining consistent facial identity. With only approximately 22 million adapter parameters layered on top of existing Stable Diffusion models, FaceID achieves remarkable identity preservation without requiring per-subject fine-tuning or multiple reference images. A single clear face photo is sufficient to generate the person in various artistic styles, different clothing, diverse environments, and novel poses. The adapter supports both SDXL and SD 1.5 base models and can be combined with other ControlNet adapters for additional control over pose, depth, and composition. IP-Adapter FaceID Plus variants incorporate additional CLIP image features alongside face embeddings for improved likeness and detail preservation. Released under the Apache 2.0 license, the model is fully open source and widely integrated into ComfyUI workflows and the Diffusers library. Common applications include personalized avatar creation, custom portrait generation in various artistic styles, character consistency in storytelling and comic creation, personalized marketing content, and social media content creation where maintaining a recognizable likeness across multiple generated images is essential.

Open Source
4.5
FLUX Redux icon

FLUX Redux

Black Forest Labs|12B

FLUX Redux is the specialized image variation model within the FLUX model family developed by Black Forest Labs, designed for generating creative variations of reference images while preserving their core style, color palette, and compositional essence. Built on the 12-billion parameter Diffusion Transformer architecture, FLUX Redux takes a reference image as input and produces new images that maintain the visual DNA of the original while introducing controlled variations in content, composition, or perspective. The model captures high-level stylistic attributes including artistic technique, color harmony, lighting mood, and textural qualities, then applies them to generate fresh compositions that feel aesthetically consistent with the source material. FLUX Redux can be combined with text prompts to guide the direction of variation, allowing users to request specific changes like 'same style but with a mountain landscape' or 'similar color palette with an urban scene.' This makes it particularly powerful for brand consistency workflows where marketing teams need multiple visuals sharing a unified aesthetic. The model also supports image-to-image workflows where the reference serves as a strong stylistic prior while text prompts define new content. As a proprietary model, FLUX Redux is accessible through Black Forest Labs' API and partner platforms including Replicate and fal.ai with usage-based pricing. Key applications include generating cohesive visual content series for social media campaigns, creating style-consistent variations for A/B testing in advertising, producing product imagery in consistent brand aesthetics, and creative exploration where artists iterate on a visual direction without starting from scratch.

Proprietary
4.6
DWPose icon

DWPose

IDEA Research|100M

DWPose is a state-of-the-art whole-body pose estimation model developed by IDEA Research that detects body keypoints, hand gestures, and facial landmarks within a single unified framework. Built on an RTMPose-based architecture combining CNN and Transformer components, DWPose achieves superior accuracy compared to OpenPose and other traditional pose estimation systems while maintaining fast inference speeds. The model with approximately 100 million parameters simultaneously estimates 133 keypoints covering the full body skeleton, both hands with individual finger joints, and 68 facial landmarks, providing comprehensive pose information in a single forward pass. DWPose has become the preferred pose estimation backbone for ControlNet-based image generation workflows, where extracted pose data guides diffusion models like Stable Diffusion and FLUX to generate images matching specific body positions and gestures. The model handles multiple persons in a single frame, works reliably across diverse body types, clothing styles, and partial occlusions, and maintains accuracy even in challenging scenarios with overlapping limbs or unusual poses. Released under the Apache 2.0 license, DWPose is fully open source and integrates seamlessly with ComfyUI, the Diffusers library, and custom animation pipelines. Beyond AI image generation, it serves applications in motion capture for game development, fitness tracking applications, sign language recognition, dance choreography analysis, and sports biomechanics research. The model runs efficiently on consumer hardware and supports real-time processing for interactive applications requiring immediate pose feedback.

Open Source
4.5

172 models found · Page 3 / 8