Modèles d’IA
Parcourez les modèles d’IA par ordre alphabétique et comparez leurs capacités documentées
PhotoMaker
PhotoMaker is a personalized photo generation model developed by TencentARC that creates realistic and diverse human portraits from reference images using a novel Stacked ID Embedding approach. Unlike traditional fine-tuning methods such as DreamBooth that require lengthy training processes, PhotoMaker achieves identity-preserving generation in seconds by extracting and stacking embeddings from multiple reference photos through CLIP and specialized identity encoders. Built on the SDXL pipeline, the model injects identity representations via modified cross-attention layers, enabling high-quality outputs that maintain facial features while allowing creative freedom in style, pose, and setting variations. PhotoMaker supports identity mixing, allowing users to blend features from multiple people to create unique composite faces with adjustable contribution weights. The model excels in personalized portrait generation, identity-consistent story illustration for comics and visual novels, virtual try-on applications, and advertising content creation. PhotoMaker V2 brought significant improvements in identity preservation accuracy, natural generation quality, and text alignment, particularly in challenging scenarios like extreme pose changes and age transformations. As an open-source model released under the Apache 2.0 license, PhotoMaker is freely available on Hugging Face with community integrations in ComfyUI and other popular creative tools. It requires only one to four reference images to produce compelling results, making it one of the most accessible and efficient identity-preserving generation solutions available for both individual creators and professional production workflows.
La version détaillée est disponible en anglais.
Pika 1.0
Pika 1.0 is a creative video generation platform developed by Pika Labs that combines powerful AI video synthesis with intuitive editing tools, making professional-quality video creation accessible to users without technical expertise. Released in December 2023, Pika emerged from Stanford research to become one of the most user-friendly video generation platforms available, offering both text-to-video and image-to-video capabilities through a streamlined web interface. The model generates short video clips from natural language descriptions, interpreting creative prompts to produce content with coherent motion, consistent lighting, and visually appealing compositions. Pika distinguishes itself through its integrated editing toolkit, which includes features like motion control for directing movement within specific regions of the frame, video extension for lengthening existing clips, and re-styling capabilities that allow users to transform the visual aesthetic of generated or uploaded content. The platform supports lip-sync functionality for adding speech to generated characters and offers expand-canvas features for changing aspect ratios or extending the visual boundaries of video content. Pika handles diverse creative styles including cinematic footage, animation, 3D renders, and stylized artistic content, with particular strength in producing visually polished short-form content suitable for social media and marketing. The model operates as a proprietary cloud-based service with freemium pricing, offering limited free generations alongside paid subscription tiers for professional users. Pika has gained significant traction among content creators, social media managers, and marketing teams who need to produce engaging video content rapidly without access to traditional video production resources or extensive AI expertise.
La version détaillée est disponible en anglais.
Pika Image-to-Video
Pika Image-to-Video is the image animation feature of Pika Labs' creative video platform that transforms still images into dynamic video content using creative motion effects and intuitive controls. Released in December 2023 as part of Pika 1.0, this capability allows users to upload any image and generate video sequences where the scene comes to life with AI-inferred motion, offering a simple yet powerful approach to creating animated content from static visuals. The model analyzes the input image to understand spatial composition, subject matter, and depth relationships, then applies contextually appropriate motion patterns while maintaining visual integrity of the source. Pika's image-to-video feature distinguishes itself through creative motion effects beyond simple camera movements, including adding specific motion to selected regions, modifying visual style during animation, and applying dramatic cinematic effects. The platform supports expand canvas for changing animation framing, lip sync for adding speech to character portraits, and motion control brushes for directing specific motion patterns. The model handles diverse input types including photographs, illustrations, digital art, memes, and design mockups, making it accessible for social media content creation, marketing materials, and artistic experimentation. The diffusion-based architecture produces smooth temporal transitions and consistent visual quality throughout sequences. As a proprietary feature within Pika's platform, Image-to-Video is available through freemium pricing with limited free generations and paid tiers for professional users requiring higher volume output and advanced controls for content production.
La version détaillée est disponible en anglais.
Pix2Pix
Pix2Pix is a pioneering image-to-image translation framework developed at UC Berkeley that introduced the concept of using conditional generative adversarial networks for paired image translation tasks. Published in November 2017 as part of the landmark paper "Image-to-Image Translation with Conditional Adversarial Networks," Pix2Pix demonstrated that a single general-purpose architecture could learn mappings between different visual domains when provided with paired training examples. The architecture consists of a U-Net-based generator that preserves spatial information through skip connections and a PatchGAN discriminator that evaluates image quality at the patch level rather than globally, enabling the model to capture fine-grained texture details while maintaining structural coherence. With approximately 54 million parameters, Pix2Pix is relatively lightweight compared to modern diffusion models, enabling fast inference and efficient training. The model excels at diverse translation tasks including converting semantic label maps to photorealistic scenes, transforming architectural facades from sketches, colorizing black-and-white photographs, converting edge maps to realistic images, and translating satellite imagery to street maps. The BSD-licensed open-source implementation has become one of the most influential works in generative AI, establishing fundamental principles that influenced subsequent models like CycleGAN, SPADE, and modern diffusion-based image editing approaches. Despite being superseded by newer techniques in terms of raw output quality, Pix2Pix remains widely used in educational contexts, rapid prototyping, and applications where paired training data is available and deterministic translation behavior is desired. Available on Hugging Face and Replicate, the model continues to serve as a foundational reference for understanding conditional image generation and adversarial training dynamics.
La version détaillée est disponible en anglais.
PixArt-Sigma
PixArt-Sigma is a highly efficient transformer-based text-to-image model developed by the PixArt research team, capable of generating images at resolutions up to 4K directly without requiring separate upscaling steps. Built on a Diffusion Transformer architecture, the model achieves quality comparable to much larger models while using significantly fewer computational resources and training costs. PixArt-Sigma represents the evolution of the PixArt series, incorporating improvements in token compression and attention mechanisms that enable native high-resolution generation. The model supports flexible aspect ratios and can produce images from 512x512 up to 4096x4096 pixels, making it particularly valuable for print design and large-format digital display applications. Its training efficiency is a standout feature, having been developed with a fraction of the computational budget required by comparable models like DALL-E 2 or Imagen. PixArt-Sigma uses a T5 text encoder for prompt understanding, providing strong semantic comprehension across diverse text inputs. Released as open-source, the model is available on Hugging Face and compatible with the Diffusers library for easy integration into existing workflows. It runs on consumer GPUs with moderate VRAM requirements, making it accessible to individual creators and small studios. AI researchers, digital artists, and developers interested in efficient high-resolution image generation use PixArt-Sigma for projects ranging from academic research to commercial content creation. Its efficiency-focused design philosophy makes it an important contribution to sustainable AI development.
La version détaillée est disponible en anglais.
Playground v3
Playground v3 is a creative AI image generation model developed by Playground AI, specifically designed for graphic design and mixed-media content creation rather than purely photorealistic output. The model distinguishes itself through superior color palette handling, typographic awareness, and the ability to generate design-ready compositions that feel intentionally crafted rather than randomly generated. Playground v3 excels at creating social media graphics, marketing banners, poster designs, and brand materials with cohesive visual hierarchies. Built on a proprietary architecture that emphasizes aesthetic control and design principles, the model understands concepts like visual balance, contrast, and focal point placement in ways that general-purpose image generators typically do not. It supports a wide range of design styles including minimalist, maximalist, retro, modern, and editorial aesthetics. The model is accessible through the Playground AI web platform, which provides an intuitive canvas-based interface for iterative design work alongside inpainting and outpainting capabilities. Playground v3 also offers an API for developers building design automation tools and content creation pipelines. Graphic designers, social media managers, content creators, and marketing teams use it as a rapid ideation and production tool, significantly reducing the time from concept to finished design. While it may not match the photorealistic fidelity of models like Midjourney v6 or FLUX.1 [pro], its design-oriented approach makes it uniquely valuable for commercial visual content that prioritizes intentional composition and brand alignment over raw photographic realism.
La version détaillée est disponible en anglais.
Playground v4
Playground v4 is Playground AI's fourth-generation image generation model, released in late 2024, designed specifically to excel at graphic design tasks alongside photorealistic image generation. The model features an innovative design-first approach that understands layout, typography placement, color theory, and brand consistency at a fundamental level. Playground v4 generates images with exceptional aesthetic quality, clean compositions, and professional design sensibility that makes outputs immediately usable in real-world design workflows. The model supports a unique canvas-based interface that allows combining multiple generations, text overlays, and design elements in a single workspace. Playground v4 competes with Midjourney in artistic quality while offering a more accessible, design-oriented user experience. The model handles photorealism, illustrations, graphic design, product photography, and social media content with consistent quality. Available through the Playground web platform with a freemium model offering daily free generations, it serves designers, content creators, and marketers who need production-ready visual content.
La version détaillée est disponible en anglais.
PowerPaint
PowerPaint is a versatile open-source inpainting model developed by researchers at Tsinghua University and HKUST under the Tencent ARC umbrella, introducing the innovative concept of learnable task prompts that enable multiple inpainting functions within a single unified model. Rather than requiring separate specialized models for each editing task, PowerPaint uses learnable task vectors that activate different behaviors within shared model weights, supporting four distinct modes: text-guided object insertion, object removal, shape-guided inpainting, and image outpainting. Built upon a Stable Diffusion backbone enriched with a ControlNet-like control mechanism, the model allows users to describe desired content through text prompts for contextual generation, cleanly remove objects while preserving surrounding textures, generate content within specific mask shapes, or extend images beyond their original boundaries. This multi-task flexibility eliminates the need to switch between different tools or models during editing workflows. In benchmark evaluations, PowerPaint achieves competitive results against separately optimized task-specific models, with its object removal quality rivaling specialized models like LaMa and MAT. Applications span photography editing, graphic design mockups, e-commerce product image preparation, digital art canvas extension, and social media content adaptation for different platform dimensions. The model is PyTorch-based and publicly available through Hugging Face with a Gradio demo interface and Diffusers library integration. GPU requirements are similar to standard Stable Diffusion models with 8GB or more VRAM recommended. PowerPaint has established a new paradigm in multi-task inpainting and continues to inspire research in unified visual editing systems.
La version détaillée est disponible en anglais.
ProGAN
ProGAN (Progressive Growing of GANs) is a generative adversarial network architecture developed by NVIDIA researchers Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen, introduced in 2017, that pioneered progressively growing both generator and discriminator networks during training to produce high-resolution face images. Instead of training at the target resolution directly, ProGAN starts at 4x4 pixels and incrementally adds layers handling progressively higher resolutions, smoothly fading in each detail level. This progressive strategy stabilizes training by learning large-scale structure before fine details, reduces training time compared to full-resolution training from scratch, and enables much higher resolution output than previously possible with GANs. ProGAN was the first GAN to convincingly generate 1024x1024 photorealistic face images, a milestone that captured widespread attention. The model was trained on CelebA-HQ, a high-quality celebrity faces dataset curated for this research. Beyond faces, ProGAN successfully generated high-resolution images of bedrooms, cars, and other categories, demonstrating versatility. The architecture introduced minibatch standard deviation for output diversity and equalized learning rate for training stability. ProGAN is fully open source with official TensorFlow implementations and community PyTorch ports. While subsequent architectures like StyleGAN built upon ProGAN's progressive training foundation to achieve higher quality and controllability, ProGAN remains a landmark contribution that changed how high-resolution GANs are trained and inspired an entire generation of improved generative models.
La version détaillée est disponible en anglais.
ProPainter : retouche vidéo, masques et licence
Évaluez ProPainter sur votre plan : préparation des masques, droits d’utilisation, réglages mémoire et contrôle du mouvement, du son et du fichier livré.
PuLID
PuLID is an identity-preserving image generation model developed by ByteDance that introduces a Pure and Lightning ID customization approach for creating personalized portraits with exceptional speed and fidelity. Released in April 2024, PuLID addresses the core challenge of maintaining a person's identity features across different generated images without requiring lengthy fine-tuning processes. The model achieves this through a novel contrastive alignment loss and accurate ID loss mechanism that works directly with pre-trained diffusion models, specifically integrating with SDXL and FLUX architectures. PuLID's key innovation lies in its ability to decouple identity features from other image attributes such as pose, expression, and background, enabling highly controllable generation where the subject's identity remains consistent while all other aspects can be freely modified. The model processes reference images through an InsightFace-based identity encoder to extract robust facial feature representations, which are then injected into the generation pipeline through specialized adapter layers. This approach enables real-time personalization without any per-subject training, making it significantly faster than alternatives like DreamBooth or textual inversion. PuLID excels in applications including personalized avatar creation, social media content generation, virtual try-on scenarios, and identity-consistent multi-scene illustration. As an open-source project released under the Apache 2.0 license, PuLID is available on Hugging Face and supported through platforms like fal.ai, offering both researchers and creators a powerful tool for identity-preserving image generation with minimal computational overhead.
La version détaillée est disponible en anglais.
Real-ESRGAN
Real-ESRGAN is an open-source image upscaling and restoration model developed by Xintao Wang and collaborators at Tencent ARC Lab that enhances low-resolution, degraded, or compressed images to high-resolution outputs with remarkable detail recovery. Released in 2021 under the BSD license, Real-ESRGAN builds on the original ESRGAN architecture by introducing a high-order degradation modeling approach that simulates the complex, unpredictable quality loss found in real-world images, including compression artifacts, noise, blur, and downsampling. The model uses a U-Net architecture with Residual-in-Residual Dense Blocks as its generator network, trained with a combination of perceptual loss, GAN loss, and pixel loss to produce sharp, natural-looking upscaled results. Real-ESRGAN supports upscaling factors of 2x, 4x, and higher, and includes specialized model variants for anime and illustration content alongside the general-purpose photographic model. The model handles real-world degradations far better than its predecessor ESRGAN, which was trained only on synthetic degradation patterns. Real-ESRGAN has become one of the most widely deployed AI upscaling solutions, integrated into numerous applications including desktop tools, web services, mobile apps, and professional image editing workflows. The model runs efficiently on both CPU and GPU, with the lighter RealESRGAN-x4plus-anime variant optimized for consumer hardware. As a fully open-source project available on GitHub with pre-trained weights, it serves as the backbone for popular tools like Upscayl and various ComfyUI nodes. Real-ESRGAN is essential for photographers, content creators, game developers, and anyone who needs to enhance image resolution while preserving natural appearance and adding realistic detail.
La version détaillée est disponible en anglais.
RealVisXL
RealVisXL is a specialized SDXL fine-tuned model created by SG_161222, purpose-built for generating ultra-photorealistic images that are often indistinguishable from professional photography. The model has been meticulously fine-tuned from the Stable Diffusion XL base with a focus on photographic accuracy, natural skin textures, realistic lighting, and true-to-life color reproduction. RealVisXL excels at portrait photography, product photography, architectural visualization, and landscape imagery, consistently producing results with the quality and feel of images captured by professional cameras. Its training emphasizes natural-looking outputs without the artificial smoothness or oversaturation commonly seen in standard AI-generated images. The model handles diverse photographic scenarios including studio lighting, outdoor natural light, golden hour, and night photography with remarkable authenticity. Available on CivitAI and compatible with all SDXL-supporting interfaces including ComfyUI and Automatic1111, RealVisXL has become one of the go-to models for users who need photographic realism above all else. It requires 8GB or more VRAM and supports all standard SDXL features including img2img, inpainting, ControlNet conditioning, and various LoRA combinations. Photographers seeking AI-assisted compositing, e-commerce businesses needing product imagery, real estate professionals requiring architectural previews, and content creators producing stock-photo-quality images all rely on RealVisXL. The model demonstrates that targeted fine-tuning of foundation models can achieve specialized excellence that surpasses the base model's capabilities in specific domains.
La version détaillée est disponible en anglais.
Recraft V3
Recraft V3 is a state-of-the-art text-to-image generation model developed by Recraft AI that achieved the highest-ever ELO score on the Artificial Analysis Image Arena, surpassing all existing models including Midjourney v6 and FLUX.1 Pro at the time of its release in October 2024. The model excels at generating images with precise text rendering, brand-consistent design elements, and production-ready vector graphics. Recraft V3 uniquely supports multiple output formats including raster images, SVG vector graphics, and illustrations with transparent backgrounds, making it particularly valuable for professional designers and brand teams. The model demonstrates exceptional understanding of design principles, accurately following complex layout instructions, maintaining typographic consistency, and generating visually balanced compositions. Available through Recraft's web platform, API, and integrated into design workflows, the model offers style control through reference images and detailed style descriptions. Recraft V3 generates images at resolutions up to 4096x4096 pixels with support for various aspect ratios. The model's vector output capability sets it apart from all competitors, producing clean, scalable SVG files directly from text prompts without post-processing. Professional designers, marketing teams, brand agencies, and e-commerce businesses use Recraft V3 for logo concepts, social media graphics, product mockups, and editorial illustrations where design precision and brand consistency are critical requirements.
La version détaillée est disponible en anglais.
Recraft V4 Styles
Recraft V4 Styles uses reference images to guide a reusable visual style. This source-based evaluation explains how to plan a consistent campaign set and distinguish its model variants.
La version détaillée est disponible en anglais.
RemBG
RemBG is a popular open-source tool developed by Daniel Gatis for automatic background removal from images, providing a simple and efficient solution for isolating foreground subjects without manual selection or professional editing skills. The tool leverages multiple pre-trained segmentation models including U2-Net, IS-Net, SAM, and specialized variants optimized for different use cases such as general objects, human subjects, anime characters, and clothing items. RemBG processes images through semantic segmentation to identify foreground elements and generates precise alpha matte masks that cleanly separate subjects from backgrounds, producing transparent PNG outputs ready for immediate use. The tool excels at handling complex edge cases including wispy hair, translucent fabrics, intricate jewelry, and objects with irregular boundaries. RemBG is available as a Python library via pip, a command-line interface for batch processing, and through API integrations for production deployment. It processes images locally without sending data to external servers, making it suitable for privacy-sensitive applications. Common use cases include e-commerce product photography, social media content creation, passport photo processing, graphic design compositing, real estate photography, and marketing materials. The tool supports JPEG, PNG, and WebP formats and handles both single images and batch directory operations. RemBG has become one of the most starred background removal repositories on GitHub with millions of downloads, and its models are integrated into numerous other AI tools. Released under the MIT license, it provides a free and commercially viable alternative to paid background removal services.
La version détaillée est disponible en anglais.
Riffusion
Riffusion is an innovative AI music generation model that takes a unique approach to audio synthesis by generating spectrograms as images using a fine-tuned version of Stable Diffusion v1.5. Created as a side project by Seth Forsyth and Hayk Martiros in late 2022, Riffusion demonstrated that image diffusion models could be repurposed for audio generation by training on spectrogram representations of music. The model generates mel spectrograms conditioned on text prompts describing musical genres, instruments, moods, and styles, which are then converted back to audio waveforms using the Griffin-Lim algorithm or neural vocoders. This image-based approach to music generation was groundbreaking at the time of release, showing that the powerful generative capabilities of Stable Diffusion could transfer to the audio domain. Riffusion can produce short music clips in various styles including rock, jazz, electronic, classical, and ambient, with real-time interpolation between different prompts enabling smooth musical transitions. The model has approximately 1 billion parameters inherited from its Stable Diffusion base. Released under the MIT license, Riffusion is fully open source with the fine-tuned model weights, training code, and an interactive web application available on GitHub. While newer purpose-built music generation models like MusicGen and Suno have surpassed Riffusion in output quality and duration, the model remains historically significant as a proof of concept that sparked widespread interest in AI music generation. Riffusion continues to be used by hobbyists and researchers exploring the intersection of image generation and audio synthesis.
La version détaillée est disponible en anglais.
Rodin Gen-1
Rodin Gen-1 is a 3D generation model developed by Microsoft Research that creates detailed, high-quality 3D models and digital avatars from text descriptions and images. The model represents Microsoft's significant entry into the AI-powered 3D content creation space, leveraging the company's extensive research in computer vision and generative AI. Rodin Gen-1 uses a diffusion-based architecture that generates 3D representations through a denoising process operating in a learned latent space, producing results with fine geometric details and realistic surface textures. The model is particularly specialized in generating 3D digital avatars with accurate facial features, hair, clothing, and accessories from textual descriptions, making it highly relevant for gaming, virtual reality, and metaverse applications. Beyond avatars, Rodin Gen-1 can generate general 3D objects and scenes with consistent quality across different categories. The generation process produces textured meshes with proper topology suitable for animation and rigging workflows. Microsoft has positioned Rodin Gen-1 as a research contribution, releasing it under a research-only license that permits academic use but restricts commercial deployment. The model builds on Microsoft's broader 3D AI research portfolio and demonstrates how large-scale generative models can be effectively applied to 3D content creation. Rodin Gen-1 is particularly noteworthy for its avatar generation quality, achieving results that approach the fidelity of manually crafted 3D characters while requiring only a text prompt as input, significantly reducing the time and expertise traditionally needed for professional 3D character creation.
La version détaillée est disponible en anglais.
Runway Gen-4 Turbo
Runway Gen-4 Turbo is Runway's fastest and most advanced video generation model, producing high-quality AI-generated video with significantly improved speed, visual fidelity, and motion coherence compared to predecessors. The model generates videos from text descriptions and image inputs with enhanced temporal consistency, producing smooth natural-looking motion that maintains subject integrity throughout clips. Gen-4 Turbo features substantially faster inference than previous Runway models, making it practical for iterative creative workflows where rapid feedback is essential. It handles diverse content types including human figures with realistic body mechanics, natural environments with dynamic elements, architectural scenes with accurate perspective, and abstract artistic compositions. Multiple generation modes are supported: text-to-video for creating clips from descriptions, image-to-video for animating still images, and video-to-video for style transformations on existing footage. The architecture builds on Runway's years of video diffusion research, incorporating temporal attention mechanisms and motion modeling for physically plausible results. Gen-4 Turbo is available through Runway's web platform and API with integration options for creative applications. Professional use cases include commercial content creation, social media video production, music video concepts, film previsualization, product advertising, and motion design. The model operates on a credit-based pricing system within Runway's subscription tiers. Gen-4 Turbo solidifies Runway's position as a leading AI video generation platform, offering professional-grade tools enabling creators to produce compelling video content without traditional production infrastructure.
La version détaillée est disponible en anglais.
RVC v2 : installation, écoute, latence et droits
Évaluer RVC v2 à partir des sources : chemins actuels, provenance checkpoint-index, écoute contrôlée, latence et droits vocaux.
SD Inpainting
Stable Diffusion Inpainting is a specialized variant of Stability AI's Stable Diffusion model fine-tuned specifically for image inpainting tasks, enabling users to fill masked regions of an image with contextually coherent content guided by text prompts. Released in 2022, the model builds upon the latent diffusion architecture but extends it with additional input channels for mask-aware processing, where the original image, mask, and masked image are fed as extra channels to the U-Net. The v1.5 inpainting model was trained on 595K curated inpainting examples in collaboration with RunwayML, while community-developed SDXL variants have since extended capabilities with higher resolution output. Common applications include removing unwanted objects from photographs, completing damaged image regions, modifying content such as adding elements to scenes, and cleaning watermarks or text overlays. Professional use cases span photography post-production, advertising visual preparation, real estate staging, product photography background replacement, and digital art workflows. The model is accessible through popular open-source interfaces including AUTOMATIC1111 WebUI, ComfyUI, InvokeAI, and the Hugging Face Diffusers library. Users can create masks manually with brush tools or automatically through segmentation models like SAM. ControlNet integration adds additional control layers for more precise output guidance. Released under the CreativeML Open RAIL-M license, the model runs on GPUs with 8GB VRAM and supports optimizations like xFormers for reduced memory usage, making it one of the most widely adopted open-source inpainting solutions available.
La version détaillée est disponible en anglais.
SDXL Turbo
SDXL Turbo is a real-time image generation model developed by Stability AI that achieves near-instantaneous image creation by requiring only a single diffusion step instead of the typical 20 to 50 steps used by standard Stable Diffusion models. Built using Adversarial Diffusion Distillation technology, SDXL Turbo distills the knowledge of the full SDXL model into a streamlined variant capable of generating 512x512 images in under one second on modern GPUs. This dramatic speed improvement opens up entirely new use cases for diffusion models, including real-time interactive image generation where users see results update live as they type or modify prompts. The model maintains surprisingly good image quality for its speed, though it naturally trades some fine detail and resolution compared to multi-step SDXL generation. SDXL Turbo is particularly effective for rapid prototyping, live creative exploration, and applications where responsiveness is more important than maximum image quality. Released as open-source, the model is available on Hugging Face and integrates with the Diffusers library, ComfyUI, and other popular interfaces. It runs efficiently on consumer GPUs with as little as 6GB VRAM. Developers building interactive AI applications, creative tools with real-time previews, and educational platforms particularly benefit from SDXL Turbo's instant generation capability. While not suitable for final production-quality output, it serves as an invaluable tool for creative ideation and real-time visual feedback in design workflows.
La version détaillée est disponible en anglais.
Seedance 2.5
Seedance 2.5 generates and edits video using text and audiovisual references. Assess shot continuity, product identity and the limits of your chosen service before delivery.
La version détaillée est disponible en anglais.
Segment Anything (SAM)
Segment Anything Model (SAM) is Meta AI's foundation model for promptable image segmentation, designed to segment any object in any image based on input prompts including points, bounding boxes, masks, or text descriptions. Released in April 2023 alongside the SA-1B dataset containing over 1 billion masks from 11 million images, SAM creates a general-purpose segmentation model that handles diverse tasks without task-specific fine-tuning. The architecture consists of three components: a Vision Transformer image encoder that processes input images into embeddings, a flexible prompt encoder handling different prompt types, and a lightweight mask decoder producing segmentation masks in real-time. SAM's zero-shot transfer capability means it can segment objects never seen during training, making it applicable across visual domains from medical imaging to satellite photography to creative content editing. The model supports automatic mask generation for segmenting everything in an image, interactive point-based segmentation for precise object selection, and box-prompted segmentation for region targeting. SAM has spawned derivative works including SAM 2 with video support, EfficientSAM for edge deployment, and FastSAM for faster inference. Practical applications span background removal, medical image annotation, autonomous driving perception, agricultural monitoring, GIS mapping, and interactive editing tools. SAM is fully open source under Apache 2.0 with PyTorch implementations, and models and dataset are freely available through Meta's repositories. It has become one of the most influential computer vision models, fundamentally changing how segmentation tasks are approached across industries.
La version détaillée est disponible en anglais.
160 modèles trouvés · Sayfa 5 / 7