HomeModels

AI Models

Discover, compare and find the best AI models for your creative projects

Filter
172 models found
StyleGAN3 icon

StyleGAN3

NVIDIA|N/A

StyleGAN3 is the third generation of NVIDIA's groundbreaking StyleGAN series of generative adversarial networks, designed to produce high-quality, photorealistic images with unprecedented control over visual attributes. Presented at NeurIPS 2021, StyleGAN3 addresses a fundamental limitation of its predecessors by eliminating texture sticking artifacts that occurred during continuous transformations and animations. Previous GAN architectures suffered from features that appeared fixed to pixel coordinates rather than moving naturally with objects, creating noticeable visual glitches during interpolation. StyleGAN3 solves this through alias-free generation using continuous signal processing principles, ensuring that fine details move smoothly and naturally with the underlying content. The architecture introduces rotation and translation equivariance, meaning generated features transform correctly and consistently when the image undergoes geometric transformations. This makes StyleGAN3 particularly suited for video generation, animation, and any application requiring smooth transitions between generated frames. The model supports configurable output resolutions and maintains the style mixing capabilities from earlier versions, allowing granular control over coarse features like pose and face shape independently from fine details like hair texture and skin quality. StyleGAN3 has been trained on various domains including human faces (FFHQ dataset), animal faces (AFHQv2), and other image categories. The model is fully open source under a custom NVIDIA license permitting research and commercial use, with official PyTorch implementations available on GitHub. It continues to serve as a benchmark reference for unconditional image generation quality and has influenced numerous subsequent GAN architectures and diffusion model designs in the generative AI landscape.

Open Source
4.5
PhotoMaker icon

PhotoMaker

Tencent|N/A

PhotoMaker is a personalized photo generation model developed by TencentARC that creates realistic and diverse human portraits from reference images using a novel Stacked ID Embedding approach. Unlike traditional fine-tuning methods such as DreamBooth that require lengthy training processes, PhotoMaker achieves identity-preserving generation in seconds by extracting and stacking embeddings from multiple reference photos through CLIP and specialized identity encoders. Built on the SDXL pipeline, the model injects identity representations via modified cross-attention layers, enabling high-quality outputs that maintain facial features while allowing creative freedom in style, pose, and setting variations. PhotoMaker supports identity mixing, allowing users to blend features from multiple people to create unique composite faces with adjustable contribution weights. The model excels in personalized portrait generation, identity-consistent story illustration for comics and visual novels, virtual try-on applications, and advertising content creation. PhotoMaker V2 brought significant improvements in identity preservation accuracy, natural generation quality, and text alignment, particularly in challenging scenarios like extreme pose changes and age transformations. As an open-source model released under the Apache 2.0 license, PhotoMaker is freely available on Hugging Face with community integrations in ComfyUI and other popular creative tools. It requires only one to four reference images to produce compelling results, making it one of the most accessible and efficient identity-preserving generation solutions available for both individual creators and professional production workflows.

Open Source
4.5
Minimax Video-01 icon

Minimax Video-01

MiniMax|undisclosed

Minimax Video-01 is MiniMax's flagship video generation model that powers the Hailuo AI platform, capable of generating high-quality video clips from text descriptions and images. Released in September 2024, the model quickly gained attention for producing remarkably natural motion, cinematic camera movements, and consistent character depiction across video frames. Video-01 generates clips up to 6 seconds at 720p resolution with smooth 25fps playback. The model demonstrates particular strength in realistic human movement, facial expressions, and environmental dynamics like water flow, fire, and wind effects. Unlike many competitors that produce visually impressive but physically implausible motion, Video-01 maintains strong physical consistency throughout generated clips. The model supports both text-to-video and image-to-video generation modes, allowing users to animate still images with natural motion while preserving the original image's style and composition. MiniMax's approach combines a large-scale transformer architecture with temporal attention mechanisms to ensure frame-to-frame coherence. The model is accessible through the Hailuo AI web platform with a freemium model offering limited free generations and paid plans for higher volume usage. Video-01 competes with Runway Gen-3, Kling 1.5, and Luma Dream Machine in the consumer video generation space, with particular advantages in natural motion quality and free-tier accessibility.

Proprietary
4.6
F5-TTS icon

F5-TTS

SWivid|335M

F5-TTS is an open-source text-to-speech model developed by SWivid that achieves fast and high-quality speech synthesis through a novel flow matching approach. The model uses a non-autoregressive architecture based on flow matching, learning smooth transformation paths between noise and target speech distributions, enabling efficient single-pass generation significantly faster than autoregressive TTS methods while maintaining comparable quality. F5-TTS supports voice cloning from short reference audio, allowing speech generation in a target speaker's voice from just a few seconds of sample audio. It reproduces vocal characteristics including timbre, pitch range, speaking rhythm, and accent with notable accuracy. A key advantage is inference speed, delivering real-time or faster-than-real-time synthesis on modern GPUs, suitable for interactive and latency-sensitive applications. The model generates speech with natural prosody, appropriate emotional expression, and contextually aware pausing and emphasis patterns. F5-TTS handles multiple languages and produces output at high sample rates suitable for professional audio production. The architecture's simplicity compared to complex multi-stage TTS pipelines makes it easier to train, fine-tune, and deploy in production environments. Released under an open-source license, F5-TTS provides a free alternative to commercial TTS services for research and production use cases. Common applications include voiceover generation, audiobook narration, accessibility tools, virtual assistant voices, podcast production, and automated voice generation for applications requiring personalized speech. Available through Hugging Face with Python integration and ONNX export for cross-platform deployment.

Open Source
4.4
SDXL Turbo icon

SDXL Turbo

Stability AI|6.6B

SDXL Turbo is a real-time image generation model developed by Stability AI that achieves near-instantaneous image creation by requiring only a single diffusion step instead of the typical 20 to 50 steps used by standard Stable Diffusion models. Built using Adversarial Diffusion Distillation technology, SDXL Turbo distills the knowledge of the full SDXL model into a streamlined variant capable of generating 512x512 images in under one second on modern GPUs. This dramatic speed improvement opens up entirely new use cases for diffusion models, including real-time interactive image generation where users see results update live as they type or modify prompts. The model maintains surprisingly good image quality for its speed, though it naturally trades some fine detail and resolution compared to multi-step SDXL generation. SDXL Turbo is particularly effective for rapid prototyping, live creative exploration, and applications where responsiveness is more important than maximum image quality. Released as open-source, the model is available on Hugging Face and integrates with the Diffusers library, ComfyUI, and other popular interfaces. It runs efficiently on consumer GPUs with as little as 6GB VRAM. Developers building interactive AI applications, creative tools with real-time previews, and educational platforms particularly benefit from SDXL Turbo's instant generation capability. While not suitable for final production-quality output, it serves as an invaluable tool for creative ideation and real-time visual feedback in design workflows.

Open Source
4.3
Wan Video icon

Wan Video

Alibaba|14B

Wan Video is an open-source video generation suite developed by Alibaba that offers multiple model sizes for text-to-video generation, providing scalable options from lightweight variants for rapid experimentation to large-scale models for production-quality output. Released in February 2025, Wan Video represents Alibaba's significant contribution to the open-source video generation ecosystem, with the largest variant featuring 14 billion parameters making it one of the most powerful freely available video generation models. Built on a transformer-based architecture that processes text prompts through advanced language understanding modules, it generates temporally coherent video sequences through latent diffusion. Wan Video supports multiple output resolutions and aspect ratios for different platforms and use cases. The model demonstrates strong capabilities in generating diverse video content including realistic human subjects with natural motion, environmental scenes with dynamic elements, creative animations, and stylized artistic interpretations. The multi-size approach allows users to choose appropriate trade-offs between quality and computational requirements, with smaller variants enabling consumer-grade hardware deployment while larger variants deliver state-of-the-art quality. Wan Video incorporates advanced temporal modeling techniques maintaining consistency across frames, reducing common artifacts such as flickering, morphing, and identity drift. Available under the Apache 2.0 license, the suite is accessible on Hugging Face and through fal.ai and Replicate. The release includes comprehensive documentation and training code, enabling the research community to study and build upon Alibaba's advances for both academic and commercial applications.

Open Source
4.5
SUPIR icon

SUPIR

Tencent ARC|N/A

SUPIR is an advanced AI image restoration and upscaling model developed by Tencent ARC researchers in 2024 that harnesses the generative power of SDXL, a large-scale Stable Diffusion model, for photo-realistic image enhancement. SUPIR stands for Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration in the Wild. The model introduces a degradation-aware encoder that analyzes the specific types of quality loss present in an input image and generates intelligent text prompts to guide the restoration process, effectively telling the diffusion model what kind of content needs to be restored and how. This intelligent prompting approach enables SUPIR to produce remarkably detailed and natural-looking upscaled results that go beyond simple pixel interpolation to generate semantically meaningful detail. The model leverages the vast visual knowledge embedded in SDXL's pre-trained weights to synthesize realistic textures, facial features, text, and fine patterns during upscaling. SUPIR excels particularly at restoring severely degraded images where traditional upscaling methods fail, including old photographs, heavily compressed web images, and low-resolution captures. The model supports high upscaling factors while maintaining coherent content and natural appearance. Released under a research-only license, SUPIR is open source with code and weights available on GitHub. While computationally intensive due to its SDXL backbone, the model produces results that represent the current frontier of AI-powered image restoration quality. SUPIR is particularly valuable for professional photographers restoring archival images, forensic analysts enhancing surveillance footage, and digital artists who need maximum quality from limited source material.

Open Source
4.6
DALL-E Inpainting icon

DALL-E Inpainting

OpenAI|N/A

DALL-E Inpainting is OpenAI's proprietary image editing capability that allows users to modify specific regions of existing images through natural language prompts, available through both the DALL-E web interface and the OpenAI API. Building on the DALL-E image generation architecture, the inpainting feature enables users to select rectangular or custom-shaped regions of an image and describe what should appear in the masked area, with the AI generating contextually appropriate content that blends with the surrounding image. The system understands complex spatial relationships, lighting conditions, and artistic styles to produce edits that maintain visual coherence with the original image. Key capabilities include adding new objects to scenes, replacing backgrounds, modifying clothing or accessories on people, changing weather conditions or time of day in landscapes, and removing unwanted elements. The API provides programmatic access for building automated editing pipelines and integrating inpainting into custom applications, with options for controlling output resolution and the number of generated variations. Unlike open-source alternatives, DALL-E Inpainting operates entirely in the cloud with no local GPU requirements, making it accessible to users without specialized hardware. The model benefits from OpenAI's continuous improvements and safety filters that prevent generation of harmful content. Commercial usage is permitted under OpenAI's terms of service, with generated images belonging to the user. While it requires a paid API subscription or credits-based usage, its ease of integration, consistent quality, and the backing of OpenAI's infrastructure make it a reliable choice for developers and businesses requiring scalable AI-powered image editing capabilities.

Proprietary
4.5
TRELLIS icon

TRELLIS

Microsoft Research|Unknown

TRELLIS is a revolutionary AI model developed by Microsoft Research that generates high-quality 3D assets from text descriptions or single 2D images using a novel Structured Latent Diffusion architecture. Released in December 2024, TRELLIS represents a fundamental advancement in 3D content generation by operating in a structured latent space that encodes geometry, texture, and material properties simultaneously rather than treating them as separate stages. The model produces complete 3D meshes with detailed PBR (Physically Based Rendering) textures, enabling direct use in game engines, 3D rendering pipelines, and AR/VR applications without extensive manual post-processing. TRELLIS supports both text-to-3D generation where users describe desired objects in natural language and image-to-3D reconstruction where a single photograph is converted into a full 3D model with inferred geometry from occluded viewpoints. The structured latent representation ensures geometric consistency and prevents the common artifacts seen in other 3D generation approaches such as floating geometry, texture seams, and unrealistic proportions. TRELLIS outputs standard 3D formats including GLB and OBJ with UV-mapped textures, making integration with professional tools like Blender, Unity, and Unreal Engine straightforward. Released under the MIT license, the model is fully open source and available on GitHub. Key applications include rapid 3D asset prototyping for game development, architectural visualization, product design mockups, virtual staging for real estate, educational 3D content creation, and metaverse asset generation. The model particularly benefits indie developers and small studios who lack resources for traditional 3D modeling workflows.

Open Source
4.5
InstructPix2Pix v2 icon

InstructPix2Pix v2

UC Berkeley|1.5B

InstructPix2Pix v2 is an advanced diffusion model developed at UC Berkeley that edits images based on natural language instructions, building upon the success of the original InstructPix2Pix by Tim Brooks and collaborators. The model takes an input image and a text instruction such as 'make it sunset' or 'turn the cat into a dog' and generates the edited result while preserving unrelated parts of the image. Built on a Stable Diffusion backbone with instruction tuning, the v2 version introduces significant improvements in instruction comprehension, output quality, and editing precision compared to its predecessor. The architecture learns to follow complex multi-step instructions and handles nuanced editing requests including style changes, object modifications, color adjustments, weather transformations, and compositional alterations. Unlike mask-based editing approaches, InstructPix2Pix v2 requires no manual region selection as it automatically identifies which parts of the image to modify based on the text instruction. The model with approximately 1.5 billion parameters runs efficiently on consumer GPUs with 8GB or more VRAM. Released under the MIT license, it is fully open source and has been integrated into popular creative tools and workflows including ComfyUI and the Diffusers library. Professional photographers, digital artists, e-commerce teams, and content creators use InstructPix2Pix v2 for rapid iterative editing, product photo enhancement, creative experimentation, and batch processing of visual content where traditional manual editing would be time-prohibitive.

Open Source
4.4
Mochi 1 Preview icon

Mochi 1 Preview

Genmo|10B

Mochi 1 Preview is an open-source text-to-video AI model developed by Genmo that sets a new standard for motion quality and physical realism in generated video content. With 10 billion parameters built on an Asymmetric Diffusion Transformer architecture, Mochi 1 Preview produces videos with remarkably natural and physically plausible motion that distinguishes it from competing models. The asymmetric architecture processes spatial and temporal information through dedicated pathways optimized for their respective characteristics, resulting in videos where objects move with realistic momentum, gravity, and interaction dynamics. Mochi 1 Preview generates 480p resolution videos at 30 frames per second with smooth, continuous motion free from the temporal flickering and object morphing artifacts common in earlier video generation models. The model demonstrates strong understanding of real-world physics including fluid dynamics, rigid body interactions, and natural phenomena like fire, smoke, and water, producing content that feels grounded in physical reality. Mochi 1 Preview responds well to detailed text prompts describing camera movements, scene transitions, and specific motion choreography, giving creators meaningful control over the generated output. Released under the Apache 2.0 license, the model is fully open source and represents one of the strongest open alternatives to proprietary video generation services. It is available through Hugging Face and supported by cloud inference providers for accessible deployment. Key applications include creating concept videos for film and advertising pre-production, generating social media video content, producing animated product demonstrations, creating visual references for motion design projects, and prototyping video ideas before committing to expensive live-action production.

Open Source
4.3
RealVisXL icon

RealVisXL

SG161222|6.6B

RealVisXL is a specialized SDXL fine-tuned model created by SG_161222, purpose-built for generating ultra-photorealistic images that are often indistinguishable from professional photography. The model has been meticulously fine-tuned from the Stable Diffusion XL base with a focus on photographic accuracy, natural skin textures, realistic lighting, and true-to-life color reproduction. RealVisXL excels at portrait photography, product photography, architectural visualization, and landscape imagery, consistently producing results with the quality and feel of images captured by professional cameras. Its training emphasizes natural-looking outputs without the artificial smoothness or oversaturation commonly seen in standard AI-generated images. The model handles diverse photographic scenarios including studio lighting, outdoor natural light, golden hour, and night photography with remarkable authenticity. Available on CivitAI and compatible with all SDXL-supporting interfaces including ComfyUI and Automatic1111, RealVisXL has become one of the go-to models for users who need photographic realism above all else. It requires 8GB or more VRAM and supports all standard SDXL features including img2img, inpainting, ControlNet conditioning, and various LoRA combinations. Photographers seeking AI-assisted compositing, e-commerce businesses needing product imagery, real estate professionals requiring architectural previews, and content creators producing stock-photo-quality images all rely on RealVisXL. The model demonstrates that targeted fine-tuning of foundation models can achieve specialized excellence that surpasses the base model's capabilities in specific domains.

Open Source
4.5
Playground v3 icon

Playground v3

Playground AI|N/A

Playground v3 is a creative AI image generation model developed by Playground AI, specifically designed for graphic design and mixed-media content creation rather than purely photorealistic output. The model distinguishes itself through superior color palette handling, typographic awareness, and the ability to generate design-ready compositions that feel intentionally crafted rather than randomly generated. Playground v3 excels at creating social media graphics, marketing banners, poster designs, and brand materials with cohesive visual hierarchies. Built on a proprietary architecture that emphasizes aesthetic control and design principles, the model understands concepts like visual balance, contrast, and focal point placement in ways that general-purpose image generators typically do not. It supports a wide range of design styles including minimalist, maximalist, retro, modern, and editorial aesthetics. The model is accessible through the Playground AI web platform, which provides an intuitive canvas-based interface for iterative design work alongside inpainting and outpainting capabilities. Playground v3 also offers an API for developers building design automation tools and content creation pipelines. Graphic designers, social media managers, content creators, and marketing teams use it as a rapid ideation and production tool, significantly reducing the time from concept to finished design. While it may not match the photorealistic fidelity of models like Midjourney v6 or FLUX.1 [pro], its design-oriented approach makes it uniquely valuable for commercial visual content that prioritizes intentional composition and brand alignment over raw photographic realism.

Proprietary
4.5
Meshy icon

Meshy

Meshy AI|N/A

Meshy is a proprietary AI-powered 3D generation platform developed by Meshy AI that creates detailed, production-ready 3D models from text descriptions and images. The platform combines text-to-3D and image-to-3D capabilities with advanced AI texturing features, positioning itself as a comprehensive solution for rapid 3D content creation. Meshy uses a transformer-based architecture that generates textured 3D meshes with PBR-compatible materials, making outputs directly usable in game engines like Unity and Unreal Engine without additional processing. The platform offers multiple generation modes including text-to-3D for creating objects from written descriptions, image-to-3D for converting photographs into 3D models, and AI texturing for applying realistic materials to existing untextured meshes. Generated models include proper UV mapping, normal maps, and physically based rendering materials suitable for professional workflows. Meshy provides both a web-based interface and an API for programmatic access, making it accessible to individual artists and scalable for enterprise pipelines. The platform is particularly popular among game developers, animation studios, and AR/VR content creators who need to produce large volumes of 3D assets efficiently. As a proprietary commercial service launched in 2023, Meshy operates on a subscription model with free tier access for limited generations. The platform continuously updates its models to improve output quality, topology optimization, and texture fidelity, competing directly with other AI 3D generation services in the rapidly evolving market.

Proprietary
4.4
Stable Audio icon

Stable Audio

Stability AI|N/A

Stable Audio is Stability AI's commercial text-to-audio generation model that produces high-quality music and sound effects from natural language descriptions. Built on a latent diffusion architecture adapted for audio, Stable Audio represents a significant advancement in AI-generated audio quality, producing outputs with professional-grade clarity and musical coherence. The model uses a variational autoencoder to compress audio spectrograms into a compact latent space, then applies a diffusion process conditioned on text embeddings to generate audio in that latent space, which is decoded back into high-fidelity waveforms. Stable Audio supports generation of music tracks and sound effects up to 90 seconds in duration at 44.1 kHz stereo quality, making it suitable for professional audio production workflows. The model was trained on a licensed music dataset from AudioSparx, addressing copyright concerns that affect many competing models. Users can specify genre, mood, tempo, instrumentation, and other musical attributes through natural language prompts, and the model produces coherent compositions that follow the described characteristics. Stable Audio also supports audio-to-audio workflows where an input audio clip is used as a starting point for generation. Released under the Stability AI Community License, the model is available for non-commercial research use with commercial access through the Stable Audio API and web platform. Stable Audio is particularly valued by content creators, video producers, podcasters, and game developers who need high-quality, original audio content generated quickly without licensing complications.

Open Source
4.4
CogVideoX icon

CogVideoX

Tsinghua & ZhipuAI|5B

CogVideoX is an open-source video generation model jointly developed by Tsinghua University and ZhipuAI that utilizes an expert transformer architecture to produce high-quality videos from text descriptions. Released in August 2024, CogVideoX represents a significant advancement in open-source video generation, offering capabilities that approach proprietary models while remaining freely available for research. Built on a 5 billion parameter transformer architecture that processes text and visual tokens through specialized expert layers, it enables efficient computation while maintaining high output quality. CogVideoX employs a 3D causal VAE for video encoding and decoding, capturing both spatial and temporal information in a unified latent space, resulting in videos with smooth motion transitions and consistent visual coherence. The model supports variable-length video generation and multiple resolution outputs, providing flexibility for different use cases. CogVideoX demonstrates strong performance in generating videos with accurate motion dynamics, scene transitions, and visual storytelling elements, handling both simple prompts and complex narrative scenarios. The training approach incorporates progressive resolution scaling and temporal consistency losses that maintain stable generation quality across different durations. Available under the Apache 2.0 license on Hugging Face, CogVideoX can be accessed through fal.ai and Replicate, and can be run locally with sufficient GPU resources. The model has been well-received in the research community as a strong open-source baseline for video generation, enabling academic studies and commercial applications that require transparent, modifiable video generation capabilities without proprietary API constraints.

Open Source
4.3
Img2Img SDXL icon

Img2Img SDXL

Stability AI|6.6B

Img2Img SDXL is the image-to-image pipeline of Stability AI's Stable Diffusion XL model, enabling users to transform existing images through style conversion, enhancement, and creative modification while maintaining structural coherence with the original input. Built on SDXL's 6.6 billion parameter latent diffusion architecture with dual text encoders, the img2img pipeline takes an input image along with a text prompt and denoising strength parameter to produce variations ranging from subtle refinements to dramatic transformations. The denoising strength controls how much the model departs from the original image, with lower values preserving more of the source composition. The SDXL base produces high-resolution 1024x1024 outputs natively without quality degradation seen in earlier Stable Diffusion versions. Key capabilities include artistic style transfer where photographs can be converted into paintings or illustrations, image enhancement, concept iteration where designers rapidly explore variations of an existing visual, and creative compositing where elements are reimagined within new contexts. The pipeline supports ControlNet integration for precise structural guidance, LoRA models for style customization, and various schedulers for fine-tuning the generation process. Released under the CreativeML Open RAIL-M license, Img2Img SDXL is available through Stability AI's platform, fal.ai, Replicate, and Hugging Face, and can be run locally with a minimum of 8GB VRAM. It serves as an essential tool for designers, digital artists, and creative professionals who need to iterate quickly on visual concepts while maintaining specific compositional elements from their source material.

Open Source
4.4
Kokoro TTS icon

Kokoro TTS

Kokoro Team|82M

Kokoro TTS is a lightweight and fast open-source text-to-speech model designed to deliver natural-sounding speech with high-quality prosody while maintaining minimal computational overhead. Built on a StyleTTS-inspired architecture, the model achieves an impressive balance between output quality and efficiency, producing expressive speech with natural rhythm, intonation, and stress placement that rivals larger and more expensive models. Kokoro TTS is optimized for edge deployment and real-time applications where low latency and small model footprint are critical, running efficiently on CPUs without GPU acceleration while maintaining production-quality output. It supports multiple voices and speaking styles with controllable parameters for speech rate, pitch, and expressiveness. Its compact architecture enables deployment in resource-constrained environments including mobile devices, embedded systems, IoT devices, and web browsers through WebAssembly, opening speech synthesis capabilities where larger models would be impractical. Kokoro TTS produces clean audio with minimal artifacts, appropriate breathing patterns, and natural sentence-level prosody that avoids the robotic quality common in lightweight TTS solutions. The model is fully open source with permissive licensing for personal and commercial use, providing a free alternative to paid TTS API services. Common applications include voice interfaces for applications, accessibility features for reading text aloud, educational tools, smart home device voice output, chatbot responses, notification systems, and scenarios requiring high-quality speech synthesis without significant computational resources. Available through Python packages and Hugging Face, Kokoro TTS integrates easily into applications and supports batch processing for offline audio generation.

Open Source
4.3
BiRefNet icon

BiRefNet

ZhengPeng7|N/A

BiRefNet (Bilateral Reference Network) is an advanced open-source segmentation model developed by ZhengPeng7 for high-resolution dichotomous image segmentation, precisely separating foreground objects from backgrounds with pixel-level accuracy at fine structural details. The model introduces a bilateral reference framework leveraging both global semantic information and local detail features through a dual-branch architecture, enabling superior edge quality compared to traditional segmentation approaches. BiRefNet processes images through a backbone encoder to extract multi-scale features, then applies bilateral reference modules that cross-reference global context with local boundary information to produce crisp segmentation masks with clean edges around complex structures like hair strands, lace patterns, chain links, and transparent materials. The model achieves state-of-the-art results on multiple benchmarks including DIS5K, demonstrating strength in handling objects with intricate boundaries that challenge conventional models. BiRefNet has gained significant popularity as a background removal solution due to its exceptional edge quality, outperforming many dedicated background removal tools on challenging images. It supports high-resolution input processing and produces alpha mattes suitable for professional compositing. Available through Hugging Face with multiple model variants optimized for different quality-speed tradeoffs, BiRefNet integrates easily into Python-based pipelines and has been adopted by several popular AI platforms. Common applications include precision background removal for product photography, fine-grained object isolation for graphic design, medical image segmentation, and creating high-quality cutouts for visual effects. Released under an open-source license, BiRefNet provides a free and technically sophisticated alternative to commercial segmentation services.

Open Source
4.5
Stable Diffusion 3.5 Medium icon

Stable Diffusion 3.5 Medium

Stability AI|2.5B

Stable Diffusion 3.5 Medium is Stability AI's optimized open-source text-to-image model with 2.5 billion parameters, released in October 2024. Designed to run efficiently on consumer hardware, the model generates high-quality images at resolutions from 0.25MP to 2MP without requiring the powerful GPUs needed by larger models. SD 3.5 Medium delivers quality that punches well above its weight class, producing detailed images with good prompt adherence, accurate text rendering, and natural compositions. The model uses the Multimodal Diffusion Transformer (MMDiT) architecture and supports customization through LoRA fine-tuning and ControlNet integration. Released under the Stability AI Community License for non-commercial use and a separate commercial license, it is freely downloadable from Hugging Face. SD 3.5 Medium is particularly valuable for developers and researchers who need a capable image generation model that can run locally without enterprise-grade hardware, making it accessible for prototyping, education, and personal creative projects.

Open Source
4.4
ProPainter icon

ProPainter

S-Lab|Unknown

ProPainter is an advanced deep learning model developed by S-Lab at Nanyang Technological University for video inpainting and object removal with exceptional temporal consistency. The model employs a dual-domain propagation architecture combined with Transformer-based attention to fill in masked or removed regions across video frames while maintaining seamless visual continuity. ProPainter takes a video and a binary mask indicating regions to be removed or filled, then generates the completed video with content that naturally blends with surrounding pixels and remains consistent across frames. The dual-domain approach propagates information in both spatial and temporal dimensions, using optical flow-guided warping to transfer texture details from neighboring frames and Transformer attention to synthesize content for regions with no visible reference. This combination allows ProPainter to handle challenging scenarios including large masked areas, fast camera motion, and complex scene dynamics that cause previous methods to produce flickering or ghosting artifacts. The model achieves state-of-the-art results on standard video inpainting benchmarks including DAVIS and YouTube-VOS, significantly outperforming previous approaches in both quantitative metrics and perceptual quality. Released under the S-Lab license, the model is open source for research purposes. Practical applications include removing unwanted objects or people from video footage, restoring damaged or corrupted video content, removing watermarks, creating clean background plates for visual effects compositing, and video-based content moderation. ProPainter integrates with standard video processing pipelines and can process videos at practical speeds on modern GPUs.

Open Source
4.4
Playground v4 icon

Playground v4

Playground AI|undisclosed

Playground v4 is Playground AI's fourth-generation image generation model, released in late 2024, designed specifically to excel at graphic design tasks alongside photorealistic image generation. The model features an innovative design-first approach that understands layout, typography placement, color theory, and brand consistency at a fundamental level. Playground v4 generates images with exceptional aesthetic quality, clean compositions, and professional design sensibility that makes outputs immediately usable in real-world design workflows. The model supports a unique canvas-based interface that allows combining multiple generations, text overlays, and design elements in a single workspace. Playground v4 competes with Midjourney in artistic quality while offering a more accessible, design-oriented user experience. The model handles photorealism, illustrations, graphic design, product photography, and social media content with consistent quality. Available through the Playground web platform with a freemium model offering daily free generations, it serves designers, content creators, and marketers who need production-ready visual content.

Proprietary
4.5
Stable Point Aware 3D (SPA3D) icon

Stable Point Aware 3D (SPA3D)

Stability AI|Unknown

Stable Point Aware 3D (SPA3D) is an advanced feed-forward 3D reconstruction model developed by Stability AI that generates high-quality textured 3D meshes from a single input image in seconds. Unlike iterative optimization-based approaches that require minutes of processing, SPA3D uses a direct feed-forward architecture that predicts 3D geometry and texture in a single pass, making it practical for interactive workflows and production pipelines. The model employs point cloud alignment techniques that significantly improve geometric consistency compared to other single-view reconstruction methods, ensuring that generated 3D models maintain accurate proportions and structural integrity from multiple viewpoints. SPA3D produces industry-standard mesh outputs with clean topology and UV-mapped textures, enabling direct import into 3D software including Blender, Unity, Unreal Engine, and professional CAD tools. The model handles diverse object categories from organic shapes like characters and animals to hard-surface objects like furniture and vehicles, adapting its reconstruction approach to the structural characteristics of each input. Released under the Stability AI Community License, the model is open source for personal and commercial use with revenue-based restrictions. Key applications include rapid 3D asset creation for game development, augmented reality content production, 3D printing preparation, virtual product photography, architectural visualization, and e-commerce 3D product displays. SPA3D is particularly valuable for creative professionals who need quick 3D mockups from concept sketches or photographs without investing hours in manual modeling. The model runs on consumer GPUs and is available through cloud APIs for scalable deployment.

Open Source
4.3
Mochi 1 icon

Mochi 1

Genmo|10B

Mochi 1 is an open-source video generation model developed by Genmo that delivers high motion fidelity and temporal consistency, establishing itself as one of the most capable freely available video generation models. Released in October 2024 with 10 billion parameters, Mochi 1 produces clips with remarkably smooth motion, consistent character appearances, and natural scene dynamics that rival some proprietary alternatives. Built on a transformer architecture that processes text prompts through a language encoder and generates video through iterative denoising, it features architectural innovations focused on maintaining temporal coherence across extended frame sequences. Mochi 1 demonstrates strong capabilities in generating realistic human motion, facial expressions, camera movements, and physical interactions between objects, areas where many competing open-source models produce noticeable artifacts. The model supports text-to-video generation with detailed prompt interpretation, producing clips that accurately reflect specified scenes, actions, and styles. At 10 billion parameters, it is one of the largest open-source video generation models, and this scale contributes to superior ability to capture complex visual details and maintain consistency throughout sequences. The model handles diverse visual styles including photorealistic content, stylized animation, and artistic interpretations. Available under the Apache 2.0 license, Mochi 1 is accessible on Hugging Face and through fal.ai and Replicate, enabling both research and commercial applications. The model has received particular praise for its motion quality, setting a new standard for open-source video generation and providing a compelling alternative for developers who need capable video generation without the constraints and costs of proprietary API services.

Open Source
4.4

172 models found · Page 5 / 8