ホームモデル

AIモデル

AIモデルを名前順に探し、公開資料に基づく機能を比較できます

160 件のモデル
Ideogram 2.0 icon

Ideogram 2.0

Ideogram|N/A

Ideogram 2 is a text-to-image generation model developed by Ideogram AI that has established itself as the industry benchmark for typography and text rendering within AI-generated images. While most image generation models struggle with producing legible, accurately spelled text, Ideogram 2 consistently generates high-quality typography that integrates naturally into images across diverse contexts including posters, logos, book covers, and social media graphics. The model builds upon the success of its predecessor with enhanced photorealistic capabilities, improved compositional accuracy, and better understanding of complex multi-element prompts. Ideogram 2 supports multiple artistic styles ranging from photorealism and 3D rendering to illustration, anime, and graphic design aesthetics. The model is accessible through the Ideogram web platform and API, offering both free and premium subscription tiers. Its architecture incorporates specialized attention mechanisms for text positioning and rendering that go beyond standard diffusion model capabilities. Graphic designers, social media managers, marketing professionals, and small business owners particularly value Ideogram 2 for creating branded content, promotional materials, and designs that require integrated typography without post-processing in external tools. The model also performs well in general image generation tasks, producing detailed and coherent images across various subjects and styles. Its unique strength in text rendering fills a critical gap in the AI image generation landscape that competitors have not yet matched consistently.

詳細版は英語で提供されています。

プロプライエタリ
Ideogram 3.0 icon

Ideogram 3.0

Ideogram AI|undisclosed

Ideogram 3.0 is the third generation of Ideogram AI's text-to-image model, released in February 2025, building upon its predecessor's renowned text rendering capabilities while adding significant improvements in overall image quality, photorealism, and style diversity. The model continues to lead the industry in generating accurate, stylistically consistent text within images, a capability that was Ideogram's original breakthrough. Version 3.0 adds dramatic improvements in photorealistic quality, pushing the model into direct competition with Midjourney and FLUX for artistic and creative imagery while maintaining its text rendering supremacy. The model features enhanced prompt understanding with better compositional accuracy, improved human anatomy rendering, and more natural lighting effects. Ideogram 3.0 introduces a magic prompt enhancement feature that automatically optimizes user prompts for better results. Available through the Ideogram web platform and API, the model offers multiple style presets including photo, design, 3D render, and painting modes. A freemium model provides limited daily generations, while paid plans offer increased quotas, higher resolution output, and commercial licensing.

詳細版は英語で提供されています。

プロプライエタリ
IDM-VTON icon

IDM-VTON

Yisol Studio|1B+

IDM-VTON (Improving Diffusion Models for Virtual Try-On) is a groundbreaking diffusion-based model developed by Yisol Studio that enables highly realistic virtual clothing try-on by combining a person's photograph with a garment image. The model uses a sophisticated two-stage architecture built on Stable Diffusion with specialized garment encoding that captures clothing details including texture, pattern, fabric drape, and structural elements with exceptional fidelity. Given a person image and a flat-lay or mannequin clothing photo, IDM-VTON generates a photorealistic visualization of the person wearing the garment while preserving their body shape, skin tone, pose, and background context. The model handles diverse clothing types from casual wear to formal attire, accessories, and layered outfits with remarkable accuracy. With over one billion parameters, IDM-VTON achieves state-of-the-art results on standard virtual try-on benchmarks, producing outputs that are often indistinguishable from real photographs. The garment encoding module specifically preserves fine details such as logos, text, buttons, and stitching patterns that previous models often blurred or lost. Released under the CC BY-NC-SA 4.0 license for research and non-commercial use, the model has been widely adopted by fashion technology startups, e-commerce platforms, and creative agencies. Applications include online shopping virtual try-on experiences, fashion design prototyping, social media content creation, and catalog generation without physical photo shoots. The model integrates with popular inference frameworks and can be deployed through cloud APIs for scalable production use.

詳細版は英語で提供されています。

公開ウェイト
Imagen 2 icon

Imagen 2

Google|N/A

Imagen 2 is Google DeepMind's advanced text-to-image generation model that combines cutting-edge diffusion model architecture with Google's deep expertise in natural language processing for superior prompt understanding and image quality. The model generates highly detailed and photorealistic images with exceptional accuracy in text rendering within images, a capability that has been a persistent challenge for most competing models. Imagen 2 leverages Google's proprietary large language model technology for text encoding, providing nuanced understanding of complex prompts including spatial relationships, attributes, and abstract concepts. The model is available through Google's Vertex AI platform and is integrated into Google's consumer products including Gemini, making it accessible to both developers and general users. Imagen 2 supports multiple output formats and resolutions, with strong performance across photorealistic, artistic, and illustrative styles. Google has implemented comprehensive safety measures including SynthID watermarking that embeds invisible identifying metadata into generated images for provenance tracking. The model also features robust content filtering aligned with Google's responsible AI principles. Enterprise customers, marketing teams, application developers building on Google Cloud, and Google Workspace users benefit from Imagen 2's tight integration with the Google ecosystem. While access is more restricted than open-source alternatives, its quality, safety features, and enterprise support make it a compelling choice for businesses already invested in Google's cloud infrastructure. Imagen 2 represents Google's commitment to making AI image generation both powerful and responsible.

詳細版は英語で提供されています。

プロプライエタリ
Imagen 3 icon

Imagen 3

Google DeepMind|undisclosed

Imagen 3 is Google DeepMind's most advanced text-to-image generation model, representing a significant leap in photorealistic image quality, prompt understanding, and visual detail compared to its predecessors. Released in August 2024 through Google's Vertex AI platform and ImageFX interface, Imagen 3 generates images with exceptional photographic quality, accurate lighting, natural skin textures, and precise spatial relationships. The model demonstrates remarkable improvement in text rendering within images, accurately generating legible text on signs, labels, and surfaces. Imagen 3 excels at understanding complex compositional prompts, correctly interpreting spatial relationships like 'next to,' 'behind,' and 'above' with higher accuracy than competing models. The model incorporates Google's SynthID digital watermarking technology that embeds invisible identifiers into generated images for provenance tracking. Available through Google Cloud's Vertex AI API and the consumer-facing ImageFX web application, Imagen 3 serves both enterprise developers and creative professionals. The model supports various aspect ratios and generates images up to 1024x1024 pixels natively, with upscaling capabilities for higher resolutions. Safety features include built-in content filters and responsible AI guardrails designed to prevent harmful content generation. Imagen 3 competes directly with DALL-E 3, Midjourney v6, and FLUX.1 Pro in the premium image generation segment, with particular strengths in photorealism and compositional accuracy.

詳細版は英語で提供されています。

プロプライエタリ
Img2Img SDXL icon

Img2Img SDXL

Stability AI|6.6B

Img2Img SDXL is the image-to-image pipeline of Stability AI's Stable Diffusion XL model, enabling users to transform existing images through style conversion, enhancement, and creative modification while maintaining structural coherence with the original input. Built on SDXL's 6.6 billion parameter latent diffusion architecture with dual text encoders, the img2img pipeline takes an input image along with a text prompt and denoising strength parameter to produce variations ranging from subtle refinements to dramatic transformations. The denoising strength controls how much the model departs from the original image, with lower values preserving more of the source composition. The SDXL base produces high-resolution 1024x1024 outputs natively without quality degradation seen in earlier Stable Diffusion versions. Key capabilities include artistic style transfer where photographs can be converted into paintings or illustrations, image enhancement, concept iteration where designers rapidly explore variations of an existing visual, and creative compositing where elements are reimagined within new contexts. The pipeline supports ControlNet integration for precise structural guidance, LoRA models for style customization, and various schedulers for fine-tuning the generation process. Released under the CreativeML Open RAIL-M license, Img2Img SDXL is available through Stability AI's platform, fal.ai, Replicate, and Hugging Face, and can be run locally with a minimum of 8GB VRAM. It serves as an essential tool for designers, digital artists, and creative professionals who need to iterate quickly on visual concepts while maintaining specific compositional elements from their source material.

詳細版は英語で提供されています。

公開ウェイト
Instant Style icon

Instant Style

InstantX Team|N/A

Instant Style is a style transfer model developed by the InstantX Team that applies the artistic style of a reference image to generated content while preserving the original content structure and semantics. Released in April 2024, the model introduces a Decoupled Style Adapter architecture built on IP-Adapter, which separates style information from content information to enable clean style injection without contaminating the subject matter of the generated image. This decoupling is achieved through specialized attention mechanisms that process style features independently from content features, allowing the model to capture color palettes, brushwork patterns, texture characteristics, and overall aesthetic qualities from the reference while maintaining compositional integrity. Instant Style works within the Stable Diffusion ecosystem, making it compatible with existing SDXL checkpoints, LoRA models, and ControlNet conditions for maximum creative flexibility. The model requires only a single reference image to extract style information, with no fine-tuning needed, enabling instant style application in real-time workflows. Key applications include artistic content creation, brand-consistent visual asset generation, game art production with unified aesthetic styles, illustration series maintaining visual coherence, and rapid prototyping of visual concepts in different artistic treatments. Available as an open-source project under the Apache 2.0 license on Hugging Face, Instant Style can also be accessed through Replicate and fal.ai. The model represents a significant advancement in controllable style transfer, offering superior content preservation compared to earlier approaches that often distorted subject matter when applying strong stylistic transformations.

詳細版は英語で提供されています。

公開ウェイト
InstantID icon

InstantID

InstantX Team|N/A

InstantID is a zero-shot identity-preserving image generation framework developed by InstantX Team that can generate images of a specific person in various styles, poses, and contexts using only a single reference photograph. Unlike traditional face-swapping or personalization methods that require multiple reference images or time-consuming fine-tuning, InstantID achieves accurate identity preservation from just one facial photograph through an innovative architecture combining a face encoder, IP-Adapter, and ControlNet for facial landmark guidance. The system extracts detailed facial identity features from the reference image and injects them into the generation process, ensuring that the generated person maintains recognizable facial features, proportions, and characteristics across diverse output scenarios. InstantID supports various creative applications including generating portraits in different artistic styles, placing the person in imagined scenes or contexts, creating profile pictures and avatars, and producing marketing materials featuring consistent character representations. The model works with Stable Diffusion XL as its base and is open-source, available on GitHub and Hugging Face for local deployment. It integrates with ComfyUI through community-developed nodes and can be accessed through cloud APIs. Portrait photographers, social media content creators, marketing teams creating personalized campaigns, game developers designing character variants, and digital artists exploring identity-based creative work all use InstantID. The framework has influenced subsequent identity-preservation models and remains one of the most effective solutions for single-image identity transfer in the open-source ecosystem.

詳細版は英語で提供されています。

公開ウェイト
InstantMesh icon

InstantMesh

Tencent|N/A

InstantMesh is a feed-forward 3D mesh generation model developed by Tencent that creates high-quality textured 3D meshes from single input images through a multi-view generation and sparse-view reconstruction pipeline. Released in April 2024 under the Apache 2.0 license, InstantMesh combines a multi-view diffusion model with a large reconstruction model to achieve both speed and quality in single-image 3D reconstruction. The pipeline first generates multiple consistent views of the input object using a fine-tuned multi-view diffusion model, then feeds these views into a transformer-based reconstruction network that predicts a triplane neural representation, which is finally converted to a textured mesh. This two-stage approach produces significantly higher quality results than single-stage methods while maintaining generation times of just a few seconds. InstantMesh supports both text-to-3D workflows when combined with an image generation model and direct image-to-3D conversion from photographs or artwork. The output meshes include detailed geometry and texture maps compatible with standard 3D software and game engines. The model handles a wide variety of object types including characters, vehicles, furniture, and organic shapes with good geometric fidelity. As an open-source project with code and weights available on GitHub and Hugging Face, InstantMesh has become a popular choice for developers building 3D asset generation pipelines. It is particularly useful for game development, e-commerce product visualization, and rapid prototyping scenarios where fast turnaround and reasonable quality are both important requirements.

詳細版は英語で提供されています。

公開ウェイト
InstructPix2Pix icon

InstructPix2Pix

Tim Brooks|1B

InstructPix2Pix is an innovative image editing model developed by researchers at UC Berkeley that enables users to edit images using natural language instructions without requiring manual masks, sketches, or reference images. The model was trained on a dataset of paired image edits generated by combining GPT-3's language capabilities with Stable Diffusion's image generation, learning to translate text-based editing instructions into precise visual modifications. Users can provide an input image along with a text instruction such as 'make it snowy,' 'turn the cat into a dog,' or 'add dramatic sunset lighting,' and InstructPix2Pix applies the requested changes while preserving the overall structure and unaffected elements of the original image. The model operates in a single forward pass, making edits quickly without iterative optimization. It handles a wide range of editing operations including style transfer, object replacement, lighting changes, season and weather modifications, material changes, and artistic transformations. InstructPix2Pix is built on the Stable Diffusion architecture and is open-source, available on Hugging Face with integration into the Diffusers library. It runs on consumer GPUs with 6GB or more VRAM. Photographers, digital artists, content creators, and developers building image editing applications use InstructPix2Pix for rapid creative editing workflows. While it may not match the precision of manual editing in complex scenarios, its natural language interface makes sophisticated image edits accessible to users without any image editing expertise.

詳細版は英語で提供されています。

公開ウェイト
InstructPix2Pix v2 icon

InstructPix2Pix v2

UC Berkeley|1.5B

InstructPix2Pix v2 is an advanced diffusion model developed at UC Berkeley that edits images based on natural language instructions, building upon the success of the original InstructPix2Pix by Tim Brooks and collaborators. The model takes an input image and a text instruction such as 'make it sunset' or 'turn the cat into a dog' and generates the edited result while preserving unrelated parts of the image. Built on a Stable Diffusion backbone with instruction tuning, the v2 version introduces significant improvements in instruction comprehension, output quality, and editing precision compared to its predecessor. The architecture learns to follow complex multi-step instructions and handles nuanced editing requests including style changes, object modifications, color adjustments, weather transformations, and compositional alterations. Unlike mask-based editing approaches, InstructPix2Pix v2 requires no manual region selection as it automatically identifies which parts of the image to modify based on the text instruction. The model with approximately 1.5 billion parameters runs efficiently on consumer GPUs with 8GB or more VRAM. Released under the MIT license, it is fully open source and has been integrated into popular creative tools and workflows including ComfyUI and the Diffusers library. Professional photographers, digital artists, e-commerce teams, and content creators use InstructPix2Pix v2 for rapid iterative editing, product photo enhancement, creative experimentation, and batch processing of visual content where traditional manual editing would be time-prohibitive.

詳細版は英語で提供されています。

公開ウェイト
IP-Adapter icon

IP-Adapter

Tencent|22M

IP-Adapter is an image prompt adapter developed by Tencent AI Lab that enables image-guided generation for text-to-image diffusion models without requiring any fine-tuning of the base model. The adapter works by extracting visual features from reference images using a CLIP image encoder and injecting these features into the diffusion model's cross-attention layers through a decoupled attention mechanism. This allows users to provide reference images as visual prompts alongside text prompts, guiding the generation process to produce images that share stylistic elements, compositional features, or visual characteristics with the reference while still following the text description. IP-Adapter supports multiple modes of operation including style transfer, where the generated image adopts the artistic style of the reference, and content transfer, where specific subjects or elements from the reference appear in the output. The adapter is lightweight, adding minimal computational overhead to the base model's inference process. It can be combined with other control mechanisms like ControlNet for multi-modal conditioning, enabling sophisticated workflows where pose, style, and content can each be controlled independently. IP-Adapter is open-source and available for various Stable Diffusion versions including SD 1.5 and SDXL. It integrates with ComfyUI and Automatic1111 through community extensions. Digital artists, product designers, brand managers, and content creators who need to maintain visual consistency across generated images or transfer specific aesthetic qualities from reference material particularly benefit from IP-Adapter's capabilities.

詳細版は英語で提供されています。

公開ウェイト
IP-Adapter FaceID icon

IP-Adapter FaceID

Tencent|22M (adapter)

IP-Adapter FaceID is a specialized adapter module developed by Tencent AI Lab that injects facial identity information into the diffusion image generation process, enabling the creation of new images that faithfully preserve a specific person's facial features. Unlike traditional face-swapping approaches, IP-Adapter FaceID extracts face recognition feature vectors from the InsightFace library and feeds them into the diffusion model through cross-attention layers, allowing the model to generate diverse scenes, styles, and compositions while maintaining consistent facial identity. With only approximately 22 million adapter parameters layered on top of existing Stable Diffusion models, FaceID achieves remarkable identity preservation without requiring per-subject fine-tuning or multiple reference images. A single clear face photo is sufficient to generate the person in various artistic styles, different clothing, diverse environments, and novel poses. The adapter supports both SDXL and SD 1.5 base models and can be combined with other ControlNet adapters for additional control over pose, depth, and composition. IP-Adapter FaceID Plus variants incorporate additional CLIP image features alongside face embeddings for improved likeness and detail preservation. Released under the Apache 2.0 license, the model is fully open source and widely integrated into ComfyUI workflows and the Diffusers library. Common applications include personalized avatar creation, custom portrait generation in various artistic styles, character consistency in storytelling and comic creation, personalized marketing content, and social media content creation where maintaining a recognizable likeness across multiple generated images is essential.

詳細版は英語で提供されています。

公開ウェイト
IP-Adapter Style icon

IP-Adapter Style

Tencent|N/A

IP-Adapter Style is a specialized variant of Tencent's IP-Adapter framework focused on artistic style transfer within diffusion model image generation pipelines. Unlike the standard IP-Adapter which transfers both content and style from reference images, the Style variant extracts and applies only stylistic qualities such as color palettes, brush stroke patterns, texture characteristics, and artistic mood while allowing the text prompt to control content and subject matter. The model encodes style reference images through a CLIP image encoder and injects extracted style features into the cross-attention layers of Stable Diffusion models through decoupled attention mechanisms separating style from content. This zero-shot approach requires no fine-tuning on the target style, making it immediately usable with any reference image. Users adjust style influence strength through a weight parameter, enabling precise control over how strongly the reference style affects output while maintaining prompt adherence. IP-Adapter Style is compatible with both SD 1.5 and SDXL architectures and integrates seamlessly with ComfyUI and Diffusers workflows. It can be combined with ControlNet for structural guidance and works alongside LoRA models for further customization. Common applications include maintaining visual consistency across illustration series, applying specific artistic aesthetics to generated images, brand identity-consistent content creation, and exploring creative style variations. The model is open source under Apache 2.0, lightweight to deploy, and has become a standard tool in AI art workflows for style-controlled image creation.

詳細版は英語で提供されています。

公開ウェイト
Kandinsky 3.0 icon

Kandinsky 3.0

Sber AI|11.9B

Kandinsky 3 is an open-source text-to-image generation model developed by Sber AI and the AI Forever research team, named after the famous abstract painter Wassily Kandinsky. The model stands out for its strong multilingual prompt understanding, particularly excelling in Russian and English language inputs while also supporting other languages. Built on a latent diffusion architecture with approximately 3 billion parameters, Kandinsky 3 incorporates a large language model backbone for text encoding that provides more nuanced semantic understanding than traditional CLIP-based approaches. The model generates high-quality images at 1024x1024 resolution across diverse styles including photorealism, digital art, anime, and traditional painting aesthetics. Its training data is notably diverse in cultural representation, producing images that reflect a broader global perspective compared to predominantly Western-trained models. Kandinsky 3 supports img2img generation, inpainting, and various conditioning methods for controlled output. Released under an open-source license, the model is freely available on Hugging Face and can be deployed locally on GPUs with 8GB or more VRAM. It integrates with the Diffusers library for easy implementation in Python-based workflows. AI researchers, digital artists, and developers in Russian-speaking communities particularly value Kandinsky 3, though its multilingual capabilities make it useful worldwide. The model also serves as a foundation for academic research in multimodal AI and cross-lingual image generation, contributing valuable diversity to the open-source image generation ecosystem.

詳細版は英語で提供されています。

公開ウェイト
Kandinsky 3.1 icon

Kandinsky 3.1

Sber AI|12B

Kandinsky 3.1 is an advanced text-to-image AI model developed by Sber AI, Russia's largest technology company, named after the pioneering abstract artist Wassily Kandinsky. With 12 billion parameters built on a diffusion architecture, the model represents a significant improvement over Kandinsky 3.0 with enhanced image quality, faster generation speeds, and better prompt adherence. Kandinsky 3.1 particularly excels at rendering Cyrillic text within images and understanding Russian language prompts with native fluency, while also supporting English and other languages effectively. The model employs a cascaded generation pipeline that first produces images at lower resolution then upscales them with a separate super-resolution module, resulting in highly detailed outputs. Kandinsky 3.1 achieves competitive results on standard image generation benchmarks, producing photorealistic imagery, digital art, and illustrations across diverse styles. The architecture features improved text encoding that better captures semantic nuances and spatial relationships described in prompts. Released under the Apache 2.0 license, the model is fully open source and available on Hugging Face for download and local deployment. It integrates with the Diffusers library and can be customized through fine-tuning for domain-specific applications. Common use cases include marketing content creation for Russian-speaking markets, editorial illustration, concept art, product visualization, and educational material generation. The model is also available through Sber's cloud API for developers who prefer managed infrastructure, making it accessible for both individual creators and enterprise teams building AI-powered visual content pipelines.

詳細版は英語で提供されています。

公開ウェイト
Kling 2.0 icon

Kling 2.0

Kuaishou Technology|undisclosed

Kling 2.0 is Kuaishou Technology's latest video generation model, released in January 2025, representing a major upgrade in video quality, motion realism, and generation capabilities over its predecessor Kling 1.5. The model generates video clips at up to 1080p resolution with dramatically improved physical simulation, human motion accuracy, and scene consistency. Kling 2.0 introduces a Master Mode for highest-quality cinematic generation with enhanced attention to lighting, depth of field, and camera cinematography. The model supports both text-to-video and image-to-video generation with clip durations up to 10 seconds in standard mode and 5 seconds in Master Mode. Notable improvements include better hand rendering, more natural facial expressions, smoother camera movements, and more physically accurate object interactions. The model processes complex scene descriptions with multiple subjects and dynamic interactions, generating videos where physical laws are more consistently maintained. Available through the Kling AI web platform and mobile app, the model offers daily free generations with premium plans for higher quality, longer clips, and commercial usage. Kling 2.0 competes with Runway Gen-3, Sora, and Veo 2 as one of the leading AI video generation models.

詳細版は英語で提供されています。

プロプライエタリ
Kling 3.0 icon

Kling 3.0

Kuaishou|Unknown

Kling 3.0 is Kuaishou's third-generation AI video generation model delivering cinematic quality output with support for longer video durations than most competitors. Developed by the AI team behind China's popular Kuaishou short-video platform, Kling 3.0 produces videos with impressive visual fidelity, realistic motion dynamics, and strong temporal coherence across extended clips. The model supports text-to-video and image-to-video generation, enabling creation from textual descriptions or animating static images with natural motion and camera movements. Its long-form video capability is a notable differentiator, allowing clips significantly longer than the few-second outputs typical of many competitors, making it suitable for narrative content and complete scene generation. The model handles complex scenarios including multi-character interactions, dynamic camera movements, environmental effects, and realistic physics simulation with consistent quality. It demonstrates particular strength in generating human motion, facial expressions, and hand gestures with reduced artifacts compared to earlier video models. The underlying architecture employs advanced diffusion transformer techniques with specialized temporal modeling maintaining coherence over longer time horizons. Kling 3.0 is accessible through Kuaishou's Kling AI platform and API with free-tier and premium options. Use cases include social media content creation, advertising video production, entertainment previsualization, educational content, and creative storytelling. With its combination of visual quality, motion realism, and extended duration support, Kling 3.0 has established itself as one of the leading video generation models, competing directly with Runway, Google, and OpenAI offerings.

詳細版は英語で提供されています。

プロプライエタリ
Kling Image-to-Video icon

Kling Image-to-Video

Kuaishou|N/A

Kling Image-to-Video is the image animation mode of Kuaishou's Kling video generation platform, designed to create video content from reference images with natural motion, temporal coherence, and high visual fidelity. Released in June 2024 as part of the Kling 1.5 suite, this capability allows users to provide a still image as a starting frame and generate video sequences that animate the scene with contextually appropriate motion. The model leverages Kling's transformer-based architecture to understand spatial composition, depth relationships, and semantic content of the input image, then generates plausible temporal evolution maintaining consistency with the source. Kling Image-to-Video demonstrates strength in animating human subjects with realistic facial expressions, body movements, and clothing dynamics, as well as generating environmental motion such as wind effects, water flow, and atmospheric changes. The model supports various output durations and resolutions for different creative and commercial applications from short social media animations to longer-form content. Users can provide optional text prompts alongside the reference image to guide the direction of generated motion, offering additional creative control. The model handles diverse input types including photographs, digital artwork, illustrations, and rendered scenes, applying motion patterns respecting the visual style and physical properties of the source. As a proprietary service, Kling Image-to-Video is accessible through Kuaishou's platform and through fal.ai and Replicate, enabling integration into custom creative tools and production pipelines for professional content creators.

詳細版は英語で提供されています。

プロプライエタリ
Kokoro TTS icon

Kokoro TTS

Kokoro Team|82M

Kokoro TTS is a lightweight and fast open-source text-to-speech model designed to deliver natural-sounding speech with high-quality prosody while maintaining minimal computational overhead. Built on a StyleTTS-inspired architecture, the model achieves an impressive balance between output quality and efficiency, producing expressive speech with natural rhythm, intonation, and stress placement that rivals larger and more expensive models. Kokoro TTS is optimized for edge deployment and real-time applications where low latency and small model footprint are critical, running efficiently on CPUs without GPU acceleration while maintaining production-quality output. It supports multiple voices and speaking styles with controllable parameters for speech rate, pitch, and expressiveness. Its compact architecture enables deployment in resource-constrained environments including mobile devices, embedded systems, IoT devices, and web browsers through WebAssembly, opening speech synthesis capabilities where larger models would be impractical. Kokoro TTS produces clean audio with minimal artifacts, appropriate breathing patterns, and natural sentence-level prosody that avoids the robotic quality common in lightweight TTS solutions. The model is fully open source with permissive licensing for personal and commercial use, providing a free alternative to paid TTS API services. Common applications include voice interfaces for applications, accessibility features for reading text aloud, educational tools, smart home device voice output, chatbot responses, notification systems, and scenarios requiring high-quality speech synthesis without significant computational resources. Available through Python packages and Hugging Face, Kokoro TTS integrates easily into applications and supports batch processing for offline audio generation.

詳細版は英語で提供されています。

公開ウェイト
Kolors icon

Kolors

Kuaishou|8B

Kolors is a bilingual text-to-image generation model developed by Kuaishou Technology, designed with native understanding of both Chinese and English languages for prompt-driven image creation. The model is built on a large-scale diffusion architecture trained on billions of image-text pairs with particular emphasis on Chinese cultural content, visual aesthetics, and linguistic nuances that Western-trained models often miss. Kolors demonstrates strong capabilities in generating images that accurately reflect Chinese artistic traditions, cultural symbols, calligraphy, and modern Chinese design aesthetics alongside standard Western visual concepts. The model achieves competitive image quality with good prompt adherence, accurate color reproduction, and detailed rendering across photorealistic, illustrative, and artistic styles. Its bilingual architecture processes Chinese and English prompts with equal proficiency, making it particularly valuable for creators producing content for Chinese-speaking audiences or cross-cultural projects. Kolors supports text-to-image generation at various resolutions and aspect ratios. Released as open-source by Kuaishou, the model is available on Hugging Face and compatible with the Diffusers library for integration into Python-based workflows. It runs on GPUs with 8GB or more VRAM and can be deployed locally or accessed through various cloud platforms. Chinese content creators, international marketing teams targeting Chinese markets, digital artists interested in Chinese aesthetics, and AI researchers studying multilingual visual generation form its primary user base. Kolors fills an important gap in the image generation landscape by providing high-quality bilingual capabilities with cultural awareness.

詳細版は英語で提供されています。

公開ウェイト
Lama Cleaner icon

Lama Cleaner

Sanster|N/A

Lama Cleaner is an open-source image inpainting tool built around the LaMa (Large Mask Inpainting) model, designed for removing unwanted objects, watermarks, text overlays, and blemishes from photographs with minimal effort. Developed by Sanster as an accessible desktop application, it provides a user-friendly brush-based interface where users simply paint over the area they want removed, and the AI fills the region with contextually appropriate content that blends seamlessly with the surrounding image. The underlying LaMa model uses a fast Fourier convolution-based architecture that excels at handling large masked areas, a common weakness in traditional inpainting approaches. Unlike many AI tools that require cloud processing, Lama Cleaner runs entirely locally on the user's machine, ensuring privacy and eliminating subscription costs. The tool supports multiple inpainting backends beyond LaMa, including LDM, ZITS, MAT, and Stable Diffusion-based models, giving users flexibility to choose the best engine for their specific task. It handles various image formats and can process both photographs and illustrations effectively. Common use cases include cleaning up travel photos by removing tourists, erasing power lines or signage from architectural shots, removing date stamps from scanned photographs, and eliminating skin blemishes in portraits. The tool is available as a Python package installable via pip and also offers a web-based interface for browser access. Its combination of powerful AI-driven inpainting, local processing, and zero cost makes it an essential utility for photographers, designers, and content creators who need quick object removal capabilities.

詳細版は英語で提供されています。

公開ウェイト
Leonardo AI icon

Leonardo AI

Leonardo AI|N/A

Leonardo AI is a comprehensive AI image generation platform that offers multiple fine-tuned models optimized for specific creative domains including game assets, character design, concept art, and product photography. Unlike single-model solutions, Leonardo provides a suite of specialized models such as Leonardo Diffusion XL, Leonardo Vision XL, and DreamShaper that users can select based on their specific needs. The platform features an intuitive web interface with built-in tools for real-time canvas editing, AI-powered image guidance, texture generation for 3D assets, and motion generation capabilities. Leonardo's model training pipeline allows users to create custom fine-tuned models using their own datasets, enabling brand-specific or style-specific image generation with as few as 10 training images. The platform particularly excels in game development workflows, offering dedicated models for generating consistent game environments, characters, items, and UI elements. It supports ControlNet-style image conditioning, inpainting, outpainting, and prompt enhancement features. Leonardo AI operates on a freemium model with daily token allocations for free users and premium subscription tiers for higher volume needs. Game developers, indie studios, concept artists, e-commerce businesses, and social media content creators form its primary user base. The API access enables integration into production pipelines for automated content generation at scale. Leonardo AI positions itself as an all-in-one creative platform rather than just a model, differentiating through its combination of multiple specialized models, training capabilities, and integrated editing tools.

詳細版は英語で提供されています。

プロプライエタリ
LGM icon

LGM

Peking University|N/A

LGM (Large Gaussian Model) is a 3D generation model developed by researchers at Peking University that produces high-quality 3D objects from single images or text prompts in approximately five seconds using 3D Gaussian Splatting representation. Released in 2024 under the MIT license, LGM combines multi-view image generation with Gaussian-based 3D reconstruction in an end-to-end framework. The model first generates multiple consistent views of the target object using a multi-view diffusion backbone, then a U-Net-based Gaussian decoder predicts 3D Gaussian parameters from these views to construct the full 3D representation. Unlike mesh-based approaches, the Gaussian Splatting output enables real-time rendering with high visual quality including accurate lighting, transparency, and reflective surface effects. LGM supports resolutions up to 512 pixels for the generated views and produces detailed 3D content with clean geometry and vivid textures. The model can be used for both image-to-3D conversion from photographs and text-to-3D generation when paired with a text-to-image model as a front end. As an open-source project with code and pre-trained weights available on GitHub, LGM is accessible to researchers and developers for both academic study and practical applications. The model is particularly suited for interactive 3D visualization, virtual reality content, game asset prototyping, and any scenario where real-time rendering of generated 3D content is required. LGM demonstrates that Gaussian Splatting provides a compelling alternative to traditional mesh representations for AI-generated 3D content.

詳細版は英語で提供されています。

公開ウェイト

160 件のモデル · Sayfa 3 / 7