HomeModels

AI Models

Discover, compare and find the best AI models for your creative projects

Filter
172 models found
PixArt-Sigma icon

PixArt-Sigma

PixArt|900M

PixArt-Sigma is a highly efficient transformer-based text-to-image model developed by the PixArt research team, capable of generating images at resolutions up to 4K directly without requiring separate upscaling steps. Built on a Diffusion Transformer architecture, the model achieves quality comparable to much larger models while using significantly fewer computational resources and training costs. PixArt-Sigma represents the evolution of the PixArt series, incorporating improvements in token compression and attention mechanisms that enable native high-resolution generation. The model supports flexible aspect ratios and can produce images from 512x512 up to 4096x4096 pixels, making it particularly valuable for print design and large-format digital display applications. Its training efficiency is a standout feature, having been developed with a fraction of the computational budget required by comparable models like DALL-E 2 or Imagen. PixArt-Sigma uses a T5 text encoder for prompt understanding, providing strong semantic comprehension across diverse text inputs. Released as open-source, the model is available on Hugging Face and compatible with the Diffusers library for easy integration into existing workflows. It runs on consumer GPUs with moderate VRAM requirements, making it accessible to individual creators and small studios. AI researchers, digital artists, and developers interested in efficient high-resolution image generation use PixArt-Sigma for projects ranging from academic research to commercial content creation. Its efficiency-focused design philosophy makes it an important contribution to sustainable AI development.

Open Source
4.3
Shap-E icon

Shap-E

OpenAI|N/A

Shap-E is a 3D generation model developed by OpenAI that creates 3D objects directly from text descriptions or input images by generating the parameters of implicit neural representations. Unlike its predecessor Point-E which produces point clouds, Shap-E generates Neural Radiance Fields (NeRF) and textured meshes that can be directly rendered and used in 3D applications. The model employs a two-stage training approach where an encoder first learns to map 3D assets to implicit function parameters, then a conditional diffusion model learns to generate those parameters from text or image inputs. This architecture enables fast generation times of just a few seconds on a modern GPU. Shap-E supports both text-to-3D and image-to-3D workflows, making it versatile for different creative pipelines. The generated 3D objects include color and texture information, producing more complete results than geometry-only approaches. Released under the MIT license in May 2023, the model is fully open source with pre-trained weights available on GitHub. While the output quality may not match optimization-heavy methods like DreamFusion that take minutes per object, Shap-E offers a practical balance between speed and quality for rapid prototyping and concept exploration. The model is particularly useful for game developers, 3D artists, and researchers who need quick 3D visualizations from text prompts. As one of OpenAI's contributions to open-source 3D AI research, Shap-E has influenced subsequent work in fast feed-forward 3D generation approaches.

Open Source
4.0
Instant Style icon

Instant Style

InstantX Team|N/A

Instant Style is a style transfer model developed by the InstantX Team that applies the artistic style of a reference image to generated content while preserving the original content structure and semantics. Released in April 2024, the model introduces a Decoupled Style Adapter architecture built on IP-Adapter, which separates style information from content information to enable clean style injection without contaminating the subject matter of the generated image. This decoupling is achieved through specialized attention mechanisms that process style features independently from content features, allowing the model to capture color palettes, brushwork patterns, texture characteristics, and overall aesthetic qualities from the reference while maintaining compositional integrity. Instant Style works within the Stable Diffusion ecosystem, making it compatible with existing SDXL checkpoints, LoRA models, and ControlNet conditions for maximum creative flexibility. The model requires only a single reference image to extract style information, with no fine-tuning needed, enabling instant style application in real-time workflows. Key applications include artistic content creation, brand-consistent visual asset generation, game art production with unified aesthetic styles, illustration series maintaining visual coherence, and rapid prototyping of visual concepts in different artistic treatments. Available as an open-source project under the Apache 2.0 license on Hugging Face, Instant Style can also be accessed through Replicate and fal.ai. The model represents a significant advancement in controllable style transfer, offering superior content preservation compared to earlier approaches that often distorted subject matter when applying strong stylistic transformations.

Open Source
4.3
Stable Audio 2.0 icon

Stable Audio 2.0

Stability AI|undisclosed

Stable Audio 2.0 is Stability AI's latest music and sound generation model, released in April 2024, capable of producing high-quality stereo audio up to 3 minutes in length at 44.1kHz from text prompts. The model generates full musical tracks with coherent song structures including intros, verses, choruses, and outros, as well as sound effects and ambient soundscapes. A key innovation in Stable Audio 2.0 is audio-to-audio generation, enabling users to transform uploaded audio samples into new compositions while maintaining structural elements from the original. The model was trained on a licensed dataset from AudioSparx, ensuring commercial safety for generated content. Available through the Stable Audio web platform and API, the model serves music producers, content creators, game developers, and filmmakers who need custom audio content. The open-source variant is available under the Stability AI Community License for non-commercial research use.

Open Source
4.4
MusicLM icon

MusicLM

Google|N/A

MusicLM is a text-to-music generation model developed by Google Research that generates high-fidelity music from text descriptions at 24 kHz. Published in January 2023 alongside a research paper, MusicLM was one of the first models to demonstrate that AI could generate coherent, high-quality music spanning multiple minutes from natural language descriptions alone. The model employs a hierarchical sequence-to-sequence architecture combining SoundStream for audio tokenization and w2v-BERT for audio representation learning, generating music tokens at multiple temporal resolutions that are then decoded into waveforms. MusicLM can produce music in diverse genres and styles based on text prompts describing instruments, tempo, mood, and musical characteristics, maintaining musical coherence and structural consistency across extended durations. The model also supports melody conditioning where users can hum or whistle a melody that guides the generated output, enabling more intuitive music creation workflows. MusicLM generates audio with rich timbral quality and natural-sounding dynamics that represent a significant improvement over earlier text-to-music approaches. As a proprietary Google model, MusicLM is not open source and was initially accessible only through the AI Test Kitchen experimental platform before being integrated into broader Google services. While newer models like MusicGen and Suno have since achieved wider adoption, MusicLM remains historically significant as a pioneering demonstration of high-quality text-to-music generation. The model influenced subsequent research and commercial developments in the AI music generation space and helped establish text-to-music as a viable and rapidly advancing field of AI research.

Proprietary
4.3
AudioLDM 2 icon

AudioLDM 2

CUHK & Surrey|N/A

AudioLDM 2 is a unified audio generation framework developed by researchers at the Chinese University of Hong Kong and the University of Surrey, capable of producing music, sound effects, and speech from text descriptions within a single model. Building on the original AudioLDM, version 2 introduces a universal audio representation called Language of Audio that bridges the gap between different audio types by encoding them into a shared semantic space. The model combines a GPT-2 language model for understanding text inputs with an AudioMAE encoder for audio conditioning, feeding into a latent diffusion model that generates audio spectrograms which are converted to waveforms. This architecture enables AudioLDM 2 to handle diverse audio generation tasks without requiring separate specialized models for each audio type. The model achieves competitive performance across multiple benchmarks including text-to-music, text-to-sound-effects, and text-to-speech evaluations. AudioLDM 2 generates audio at up to 48 kHz with good perceptual quality for both musical and non-musical content. Released in August 2023 under a research license, the model is open source with code and pre-trained weights available on GitHub and Hugging Face. AudioLDM 2 supports audio inpainting, style transfer, and super-resolution in addition to text-conditioned generation. The model is particularly relevant for researchers studying unified audio generation, content creators needing diverse audio types from a single tool, and developers building comprehensive audio generation systems. Its unified approach to handling speech, music, and environmental sounds makes it a versatile foundation for multi-purpose audio applications.

Open Source
4.2
Stable Cascade icon

Stable Cascade

Stability AI|5.1B

Stable Cascade is an efficient three-stage image generation model developed by Stability AI, built upon the Wuerstchen architecture that operates in a highly compressed latent space for dramatically improved training and inference efficiency. The model uses a cascaded pipeline consisting of three stages: Stage C generates a compact 24x24 latent representation, Stage B decodes this to a 256x256 latent image, and Stage A produces the final high-resolution output. This extreme compression in the initial stage allows Stable Cascade to be trained and run with significantly less computational resources than comparable quality models while maintaining impressive image quality. The architecture achieves approximately 16x compression ratio compared to standard latent diffusion models, making it one of the most resource-efficient high-quality image generators available. Stable Cascade supports text-to-image generation, image-to-image transformation, inpainting, and ControlNet-style conditioning. Its modular three-stage design allows researchers to experiment with and improve individual stages independently. Released under an open-source license, the model is available on Hugging Face and compatible with the Diffusers library. It runs effectively on consumer GPUs with modest VRAM requirements, typically 8GB or more. AI researchers studying efficient generative architectures and developers building resource-constrained applications particularly value Stable Cascade's approach to maximizing quality per compute unit. While it has been somewhat overshadowed by the release of FLUX.1, its architectural innovations in latent space compression represent important research contributions to the field of efficient image generation.

Open Source
4.2
PowerPaint icon

PowerPaint

Tencent ARC|N/A

PowerPaint is a versatile open-source inpainting model developed by researchers at Tsinghua University and HKUST under the Tencent ARC umbrella, introducing the innovative concept of learnable task prompts that enable multiple inpainting functions within a single unified model. Rather than requiring separate specialized models for each editing task, PowerPaint uses learnable task vectors that activate different behaviors within shared model weights, supporting four distinct modes: text-guided object insertion, object removal, shape-guided inpainting, and image outpainting. Built upon a Stable Diffusion backbone enriched with a ControlNet-like control mechanism, the model allows users to describe desired content through text prompts for contextual generation, cleanly remove objects while preserving surrounding textures, generate content within specific mask shapes, or extend images beyond their original boundaries. This multi-task flexibility eliminates the need to switch between different tools or models during editing workflows. In benchmark evaluations, PowerPaint achieves competitive results against separately optimized task-specific models, with its object removal quality rivaling specialized models like LaMa and MAT. Applications span photography editing, graphic design mockups, e-commerce product image preparation, digital art canvas extension, and social media content adaptation for different platform dimensions. The model is PyTorch-based and publicly available through Hugging Face with a Gradio demo interface and Diffusers library integration. GPU requirements are similar to standard Stable Diffusion models with 8GB or more VRAM recommended. PowerPaint has established a new paradigm in multi-task inpainting and continues to inspire research in unified visual editing systems.

Open Source
4.3
T2I-Adapter icon

T2I-Adapter

Tencent ARC|77M

T2I-Adapter is a lightweight conditioning framework for text-to-image diffusion models developed by Tencent ARC Lab that provides structural control over generated images through various guidance signals including sketch, depth, segmentation, color, and style inputs. Unlike ControlNet which adds substantial computational overhead by creating full copies of the encoder, T2I-Adapter uses a compact adapter architecture that achieves similar conditioning capabilities with significantly less memory usage and faster inference times. The adapter extracts multi-scale features from conditioning images and injects them into the diffusion model's intermediate feature maps, guiding the generation process to follow the desired spatial structure while maintaining the model's creative freedom in unspecified areas. T2I-Adapter supports multiple conditioning types that can be combined for complex multi-condition generation, allowing users to specify both structural layout and stylistic direction simultaneously. Each adapter type is trained independently and can be mixed and matched at inference time, providing flexible compositional control. The framework is particularly effective for professional workflows requiring consistent spatial layouts across multiple variations, such as architectural visualization, product design iteration, and character sheet generation. T2I-Adapter is open-source and available for Stable Diffusion 1.5 and SDXL on Hugging Face, compatible with the Diffusers library and ComfyUI. Its lightweight nature makes it especially valuable for deployment on resource-constrained hardware and for applications requiring real-time or near-real-time conditioning. Designers, architects, product developers, and animation studios use T2I-Adapter for production workflows where precise structural guidance is needed without the computational cost of heavier control solutions.

Open Source
4.2
Hunyuan-DiT icon

Hunyuan-DiT

Tencent|1.5B

Hunyuan-DiT is a bilingual text-to-image diffusion transformer model developed by Tencent, featuring a Diffusion Transformer architecture designed for high-quality image generation with native Chinese and English language understanding. The model employs a transformer-based diffusion approach that replaces the traditional U-Net backbone used in earlier diffusion models with a more scalable and efficient transformer architecture. Hunyuan-DiT uses a bilingual CLIP text encoder combined with a multilingual T5 encoder to process prompts in both Chinese and English with deep semantic understanding. The model generates high-resolution images with strong compositional accuracy, detailed textures, and faithful prompt adherence across various artistic styles including photorealism, traditional Chinese painting, modern illustration, and digital art. Its training dataset includes extensive Chinese cultural content, enabling it to accurately render Chinese characters, traditional artistic motifs, architectural elements, and cultural scenes that most Western-trained models cannot handle properly. Hunyuan-DiT supports controllable generation through various conditioning mechanisms and can produce images at multiple resolutions and aspect ratios. Released as open-source under a permissive license, the model is available on Hugging Face and GitHub with full training and inference code. It requires GPUs with 11GB or more VRAM for efficient operation. Chinese technology companies, digital content creators in Chinese-speaking markets, researchers in multilingual AI, and artists exploring cross-cultural visual creation form its primary user base. Hunyuan-DiT represents Tencent's significant contribution to the open-source image generation ecosystem and advances the state of bilingual visual AI.

Open Source
4.2
Unique3D icon

Unique3D

Tencent|N/A

Unique3D is a high-quality single-image 3D reconstruction model developed by Tencent that produces detailed, well-textured 3D meshes from single input images through a multi-stage pipeline combining multi-view generation, geometry reconstruction, and texture refinement. The model is designed to produce production-quality 3D assets with sharp textures and clean geometry that can be directly used in professional 3D applications. Unique3D employs a multi-level upscale refinement strategy where the initial 3D reconstruction is progressively enhanced at multiple resolution levels, resulting in significantly finer surface details and texture quality compared to single-pass methods. The pipeline first generates consistent multi-view images using a diffusion model, then reconstructs an initial 3D mesh, and finally applies iterative upscaling and refinement to both geometry and texture. This approach produces meshes with crisp texture details and well-defined geometric features even for complex objects with intricate patterns or fine structures. Released under the Apache 2.0 license in May 2024, Unique3D is fully open source with code and pre-trained weights available on GitHub. The model handles a variety of object types including characters, animals, manufactured products, and artistic objects. Output meshes include high-resolution texture maps and proper UV coordinates compatible with standard 3D software. Unique3D is particularly suited for professional workflows in game development, animation, product visualization, and digital content creation where the quality of 3D assets directly impacts the final output. The multi-level refinement approach represents an important contribution to achieving production-grade quality in AI-generated 3D content.

Open Source
4.3
LGM icon

LGM

Peking University|N/A

LGM (Large Gaussian Model) is a 3D generation model developed by researchers at Peking University that produces high-quality 3D objects from single images or text prompts in approximately five seconds using 3D Gaussian Splatting representation. Released in 2024 under the MIT license, LGM combines multi-view image generation with Gaussian-based 3D reconstruction in an end-to-end framework. The model first generates multiple consistent views of the target object using a multi-view diffusion backbone, then a U-Net-based Gaussian decoder predicts 3D Gaussian parameters from these views to construct the full 3D representation. Unlike mesh-based approaches, the Gaussian Splatting output enables real-time rendering with high visual quality including accurate lighting, transparency, and reflective surface effects. LGM supports resolutions up to 512 pixels for the generated views and produces detailed 3D content with clean geometry and vivid textures. The model can be used for both image-to-3D conversion from photographs and text-to-3D generation when paired with a text-to-image model as a front end. As an open-source project with code and pre-trained weights available on GitHub, LGM is accessible to researchers and developers for both academic study and practical applications. The model is particularly suited for interactive 3D visualization, virtual reality content, game asset prototyping, and any scenario where real-time rendering of generated 3D content is required. LGM demonstrates that Gaussian Splatting provides a compelling alternative to traditional mesh representations for AI-generated 3D content.

Open Source
4.2
Pix2Pix icon

Pix2Pix

UC Berkeley|54M

Pix2Pix is a pioneering image-to-image translation framework developed at UC Berkeley that introduced the concept of using conditional generative adversarial networks for paired image translation tasks. Published in November 2017 as part of the landmark paper "Image-to-Image Translation with Conditional Adversarial Networks," Pix2Pix demonstrated that a single general-purpose architecture could learn mappings between different visual domains when provided with paired training examples. The architecture consists of a U-Net-based generator that preserves spatial information through skip connections and a PatchGAN discriminator that evaluates image quality at the patch level rather than globally, enabling the model to capture fine-grained texture details while maintaining structural coherence. With approximately 54 million parameters, Pix2Pix is relatively lightweight compared to modern diffusion models, enabling fast inference and efficient training. The model excels at diverse translation tasks including converting semantic label maps to photorealistic scenes, transforming architectural facades from sketches, colorizing black-and-white photographs, converting edge maps to realistic images, and translating satellite imagery to street maps. The BSD-licensed open-source implementation has become one of the most influential works in generative AI, establishing fundamental principles that influenced subsequent models like CycleGAN, SPADE, and modern diffusion-based image editing approaches. Despite being superseded by newer techniques in terms of raw output quality, Pix2Pix remains widely used in educational contexts, rapid prototyping, and applications where paired training data is available and deterministic translation behavior is desired. Available on Hugging Face and Replicate, the model continues to serve as a foundational reference for understanding conditional image generation and adversarial training dynamics.

Open Source
4.0
Kandinsky 3.0 icon

Kandinsky 3.0

Sber AI|11.9B

Kandinsky 3 is an open-source text-to-image generation model developed by Sber AI and the AI Forever research team, named after the famous abstract painter Wassily Kandinsky. The model stands out for its strong multilingual prompt understanding, particularly excelling in Russian and English language inputs while also supporting other languages. Built on a latent diffusion architecture with approximately 3 billion parameters, Kandinsky 3 incorporates a large language model backbone for text encoding that provides more nuanced semantic understanding than traditional CLIP-based approaches. The model generates high-quality images at 1024x1024 resolution across diverse styles including photorealism, digital art, anime, and traditional painting aesthetics. Its training data is notably diverse in cultural representation, producing images that reflect a broader global perspective compared to predominantly Western-trained models. Kandinsky 3 supports img2img generation, inpainting, and various conditioning methods for controlled output. Released under an open-source license, the model is freely available on Hugging Face and can be deployed locally on GPUs with 8GB or more VRAM. It integrates with the Diffusers library for easy implementation in Python-based workflows. AI researchers, digital artists, and developers in Russian-speaking communities particularly value Kandinsky 3, though its multilingual capabilities make it useful worldwide. The model also serves as a foundation for academic research in multimodal AI and cross-lingual image generation, contributing valuable diversity to the open-source image generation ecosystem.

Open Source
4.2
StyleDrop icon

StyleDrop

Google|N/A

StyleDrop is a method developed by Google Research for fine-tuning text-to-image generation models to faithfully capture and reproduce a specific visual style from as few as one or two reference images. Unlike general text-to-image models that generate images in varied or generic styles, StyleDrop enables precise style control by efficiently adapting model parameters through adapter tuning, requiring only a handful of style exemplars rather than large datasets. The method was demonstrated primarily on Google's Muse model, a masked generative transformer architecture, and achieves remarkable style fidelity across diverse artistic styles including flat illustrations, oil paintings, watercolors, 3D renders, pixel art, and abstract compositions. StyleDrop works by training lightweight adapter parameters that capture style-specific features such as color palettes, brush stroke patterns, texture characteristics, and compositional tendencies from the reference images. During inference, these adapters guide the generation process to produce new images with arbitrary content while consistently maintaining the learned stylistic qualities. An optional iterative training procedure with human or CLIP-based feedback further refines style accuracy. This approach is particularly valuable for brand identity applications where visual consistency across multiple generated assets is essential, as well as for artists wanting to maintain a signature style across AI-generated works. The method outperforms DreamBooth and textual inversion on style-specific generation benchmarks while requiring fewer training images and less computation. While StyleDrop itself is not open source, its concepts have influenced subsequent open-source style adaptation techniques in the Stable Diffusion ecosystem including LoRA and IP-Adapter approaches.

Proprietary
4.3
I2VGen-XL icon

I2VGen-XL

Alibaba DAMO|N/A

I2VGen-XL is a high-quality image-to-video generation model developed by Alibaba DAMO Academy that produces video content with strong semantic and temporal coherence from single input images. Released in November 2023, I2VGen-XL employs a cascaded architecture decomposing video generation into two stages: a base stage generating low-resolution video with correct semantic content and motion patterns, followed by a refinement stage that upscales and enhances visual quality for the final output. This two-stage approach lets the model first focus on understanding content and motion dynamics before applying detailed visual refinement, resulting in videos maintaining both semantic accuracy and visual quality. The model demonstrates strong capabilities in preserving the identity and visual characteristics of the input image while generating plausible temporal evolution, making it effective where maintaining visual consistency with source material is critical. I2VGen-XL handles diverse input types including photographs of people, animals, landscapes, and artistic compositions, applying contextually appropriate motion patterns respecting physical properties and spatial relationships in the original image. The model generates videos with smooth frame transitions, consistent lighting, and natural motion dynamics avoiding artifacts common in earlier approaches. Key use cases include animated product showcases, dynamic content from stock photography, animating concept art and design mockups, and social media content with engaging visual motion. Available under the Apache 2.0 license, I2VGen-XL is accessible on Hugging Face and Replicate, offering a capable open-source solution for image-to-video generation that balances quality with computational efficiency.

Open Source
4.1
Wonder3D icon

Wonder3D

Tsinghua University|N/A

Wonder3D is a single-image 3D reconstruction model developed by researchers at Tsinghua University that generates both multi-view color images and corresponding normal maps from a single input image for high-quality 3D mesh reconstruction. Accepted at CVPR 2024, Wonder3D introduces a cross-domain diffusion approach that simultaneously produces RGB color views and geometric normal maps, ensuring that the generated views are both visually consistent and geometrically accurate. This dual-output strategy provides significantly richer information for downstream 3D reconstruction compared to methods that generate only color images. The model uses a multi-view cross-domain attention mechanism that enforces consistency between the color and normal map domains during the diffusion process, resulting in coherent multi-view outputs that faithfully represent the 3D structure of the input object. Wonder3D can reconstruct a complete textured 3D mesh from a single photograph in approximately two to three minutes. The output meshes feature clean geometry with well-defined surface details, making them suitable for use in professional 3D workflows. Released under the Apache 2.0 license, the model is fully open source with code and pre-trained weights available on GitHub. Wonder3D handles diverse object categories including characters, animals, furniture, and manufactured objects with consistent quality. The model is particularly valuable for applications in game development, animation, product visualization, and virtual reality where high-quality 3D assets need to be created from limited reference imagery. Its cross-domain approach has influenced subsequent research in multi-view generation for 3D reconstruction.

Open Source
4.1
One-2-3-45 icon

One-2-3-45

UC San Diego|N/A

One-2-3-45 is a single-image 3D reconstruction system developed by researchers at UC San Diego that generates textured 3D meshes from a single input image through a two-stage pipeline combining multi-view generation with sparse-view 3D reconstruction. The name reflects the core process: from one image, generate two to three to four to five views, then reconstruct a complete 3D object. In the first stage, a fine-tuned Zero123 model generates multiple novel views of the object from different angles based on the single input photograph. In the second stage, these generated multi-view images are fed into a cost-volume-based sparse-view reconstruction network that produces a textured 3D mesh with consistent geometry. Released in June 2023 under the MIT license, One-2-3-45 was among the first systems to demonstrate that combining 2D diffusion models with 3D reconstruction could produce reasonable 3D assets in under a minute. The model handles a variety of object types including everyday items, animals, vehicles, and artistic objects. Unlike optimization-based approaches like DreamFusion that require per-object optimization taking tens of minutes, One-2-3-45 runs in a feed-forward manner making it significantly faster. The output meshes include color and texture information and can be exported for use in standard 3D applications. As a fully open-source project with code available on GitHub, it has served as an influential reference for subsequent research in single-image 3D generation. The system is particularly useful for researchers and developers exploring rapid 3D content creation from limited input data.

Open Source
4.0
Rodin Gen-1 icon

Rodin Gen-1

Microsoft|N/A

Rodin Gen-1 is a 3D generation model developed by Microsoft Research that creates detailed, high-quality 3D models and digital avatars from text descriptions and images. The model represents Microsoft's significant entry into the AI-powered 3D content creation space, leveraging the company's extensive research in computer vision and generative AI. Rodin Gen-1 uses a diffusion-based architecture that generates 3D representations through a denoising process operating in a learned latent space, producing results with fine geometric details and realistic surface textures. The model is particularly specialized in generating 3D digital avatars with accurate facial features, hair, clothing, and accessories from textual descriptions, making it highly relevant for gaming, virtual reality, and metaverse applications. Beyond avatars, Rodin Gen-1 can generate general 3D objects and scenes with consistent quality across different categories. The generation process produces textured meshes with proper topology suitable for animation and rigging workflows. Microsoft has positioned Rodin Gen-1 as a research contribution, releasing it under a research-only license that permits academic use but restricts commercial deployment. The model builds on Microsoft's broader 3D AI research portfolio and demonstrates how large-scale generative models can be effectively applied to 3D content creation. Rodin Gen-1 is particularly noteworthy for its avatar generation quality, achieving results that approach the fidelity of manually crafted 3D characters while requiring only a text prompt as input, significantly reducing the time and expertise traditionally needed for professional 3D character creation.

Proprietary
4.2
OpenLRM icon

OpenLRM

Zexiang Xu|N/A

OpenLRM is an open-source implementation of the Large Reconstruction Model architecture for single-image 3D reconstruction, developed by Zexiang Xu and collaborators. The project provides a fully open and reproducible implementation of the LRM approach, which uses a transformer-based architecture to predict 3D representations from single input images in a feed-forward manner. OpenLRM processes an input image through a pre-trained vision encoder like DINOv2, then feeds the resulting features into a transformer decoder that generates a triplane-based neural radiance field representation, which can be rendered from novel viewpoints or converted to a textured 3D mesh. The entire reconstruction takes only a few seconds on a modern GPU, making it practical for interactive applications and batch processing workflows. Released under the Apache 2.0 license in December 2023, OpenLRM fills a critical gap in the 3D AI research community by providing an accessible reference implementation that researchers can study, modify, and build upon. The model supports various output formats and can be integrated into existing 3D pipelines for applications ranging from game development to e-commerce product visualization. OpenLRM handles diverse object categories including furniture, vehicles, characters, and everyday items with reasonable geometric fidelity. Pre-trained model weights are available on Hugging Face for immediate use. As one of the foundational open-source projects in feed-forward 3D reconstruction, OpenLRM has directly influenced and enabled numerous downstream projects and research efforts in the rapidly evolving single-image 3D generation space.

Open Source
4.1
ModelScope T2V icon

ModelScope T2V

Alibaba DAMO|1.7B

ModelScope T2V is an early open-source text-to-video generation model developed by Alibaba DAMO Academy that pioneered accessible video generation research by making a functional text-to-video pipeline freely available. Released in March 2023, ModelScope T2V was among the first open-source models to demonstrate practical text-to-video capabilities, establishing an important baseline for subsequent developments. Built on a 1.7 billion parameter diffusion architecture, it extends latent diffusion to the temporal domain, incorporating temporal convolution and attention layers for generating short video clips from text descriptions. The architecture processes text prompts through a CLIP encoder and generates video through a modified U-Net with temporal dimensions, producing clips with basic motion coherence and prompt alignment. While output quality is modest compared to recent models like Sora or Runway Gen-3, ModelScope T2V played a crucial historical role in democratizing video generation technology by providing the first truly accessible open-source implementation that researchers could experiment with, modify, and build upon. The model supports generation of short clips at moderate resolutions, handling simple scene descriptions with recognizable subjects and basic motion patterns. Common use cases include research experimentation, educational demonstrations of video generation concepts, rapid prototyping, and serving as a baseline for training more advanced models. Available under the Apache 2.0 license on Hugging Face and Replicate, ModelScope T2V remains relevant as a lightweight, resource-efficient option for scenarios where state-of-the-art quality is not required but functional video generation capability is needed with minimal computational overhead.

Open Source
3.8
Era3D icon

Era3D

Alibaba|N/A

Era3D is a multi-view generation model developed by Alibaba that produces high-resolution, camera-aware multi-view images and normal maps from single input images for 3D reconstruction. The model introduces two key innovations that address common limitations in multi-view generation: a focal length estimation module that adapts to the camera perspective of the input image, and an efficient row-wise attention mechanism that enables generation at higher resolutions than competing methods while using less GPU memory. Era3D generates six consistent views along with corresponding normal maps at 512x512 resolution, providing rich geometric information for downstream 3D mesh reconstruction. The camera-aware design means the model can handle input images taken from different perspectives and focal lengths without degradation in output quality, a significant improvement over methods that assume a fixed camera model. The row-wise attention mechanism replaces the computationally expensive full cross-view attention with a more efficient alternative that processes attention along horizontal rows, reducing memory requirements while maintaining view consistency. Released in May 2024 under the Apache 2.0 license, Era3D is fully open source with code and pre-trained weights available on GitHub. The model demonstrates strong performance across diverse object categories and produces clean multi-view outputs suitable for high-quality 3D reconstruction. Era3D is particularly valuable for professional 3D content creation workflows where input images come from varied sources with different camera characteristics, and where high-resolution multi-view generation is essential for capturing fine details in the final 3D models.

Open Source
4.2
SyncDreamer icon

SyncDreamer

Tsinghua University|N/A

SyncDreamer is a multi-view generation and 3D reconstruction model developed by researchers at Tsinghua University that generates synchronized, 3D-consistent views of objects from single input images. Released in 2023 under the Apache 2.0 license, SyncDreamer introduces a synchronized multi-view diffusion approach that generates multiple views simultaneously while enforcing 3D consistency through a novel attention mechanism. Unlike sequential view generation methods that often produce inconsistent results between views, SyncDreamer's synchronized generation process ensures that all output views share coherent geometry, lighting, and appearance. The model uses a modified diffusion architecture with a 3D-aware feature attention module that allows information to flow between different viewpoint predictions during the denoising process. This cross-view communication enables the model to maintain spatial consistency across all generated views. The output multi-view images can be used with standard multi-view reconstruction methods like NeuS or NeRF to produce high-quality textured 3D meshes. SyncDreamer generates 16 evenly spaced views around the object, providing comprehensive coverage for accurate 3D reconstruction. The model handles a variety of object categories including animals, vehicles, furniture, and artistic objects with good consistency. As a fully open-source project with code and weights available on GitHub, SyncDreamer has become an important reference in the multi-view generation literature. The model is particularly relevant for researchers working on 3D generation pipelines and for applications in game development, product visualization, and virtual reality content creation where converting single images to 3D assets is a common requirement.

Open Source
4.0
ProGAN icon

ProGAN

NVIDIA|N/A

ProGAN (Progressive Growing of GANs) is a generative adversarial network architecture developed by NVIDIA researchers Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen, introduced in 2017, that pioneered progressively growing both generator and discriminator networks during training to produce high-resolution face images. Instead of training at the target resolution directly, ProGAN starts at 4x4 pixels and incrementally adds layers handling progressively higher resolutions, smoothly fading in each detail level. This progressive strategy stabilizes training by learning large-scale structure before fine details, reduces training time compared to full-resolution training from scratch, and enables much higher resolution output than previously possible with GANs. ProGAN was the first GAN to convincingly generate 1024x1024 photorealistic face images, a milestone that captured widespread attention. The model was trained on CelebA-HQ, a high-quality celebrity faces dataset curated for this research. Beyond faces, ProGAN successfully generated high-resolution images of bedrooms, cars, and other categories, demonstrating versatility. The architecture introduced minibatch standard deviation for output diversity and equalized learning rate for training stability. ProGAN is fully open source with official TensorFlow implementations and community PyTorch ports. While subsequent architectures like StyleGAN built upon ProGAN's progressive training foundation to achieve higher quality and controllability, ProGAN remains a landmark contribution that changed how high-resolution GANs are trained and inspired an entire generation of improved generative models.

Open Source
4.0

172 models found · Page 7 / 8