Modèles d’IA
Parcourez les modèles d’IA par ordre alphabétique et comparez leurs capacités documentées
Udio
Udio is an AI music generation platform developed by former Google DeepMind researchers that creates high-quality songs with vocals, lyrics, and instrumentals from text prompts. Launched in April 2024, Udio quickly gained attention for producing remarkably realistic and musically coherent outputs that rival professional studio recordings in audio fidelity. The platform uses a proprietary transformer-based architecture that generates all aspects of a musical composition including vocal performances, instrumental arrangements, harmonies, and production effects in a unified process. Udio supports an extensive range of musical genres and styles from mainstream pop and rock to niche genres like lo-fi, synthwave, Afrobeat, and traditional folk music from various cultures. Generated songs feature studio-quality audio at high sample rates with realistic vocal timbres, proper musical dynamics, and professional-sounding mixing and mastering. The platform allows users to provide custom lyrics, specify song structure, and control various musical parameters through text descriptions. Udio also supports audio extensions where users can generate additional sections to extend existing songs, enabling the creation of full-length tracks through iterative generation. The platform operates on a freemium model with free daily generations and paid subscription tiers for commercial use and higher generation limits. Udio is particularly notable for its vocal quality, which includes natural-sounding vibrato, breath sounds, and emotional expressiveness that many competing platforms struggle to achieve. The platform is popular among content creators, independent musicians exploring AI-assisted composition, marketing teams needing original music, and hobbyists who want to create professional-sounding songs without musical training or expensive production equipment.
La version détaillée est disponible en anglais.
Udio v1.5
Udio v1.5 is the updated version of Udio's AI music generation platform, released in August 2024, delivering substantial improvements in audio fidelity, instrument separation, and genre accuracy over the original Udio v1. The model generates full songs from text prompts describing genre, mood, instrumentation, and lyrical content with notably higher production quality than its predecessor. Udio v1.5 is particularly praised for its instrumental detail, producing recordings where individual instruments are clearly distinguishable with natural timbres and realistic playing techniques. The model excels at accurately reproducing genre-specific production aesthetics, from the warm analog saturation of classic rock to the crisp digital precision of modern electronic music. Songs can be generated up to 2 minutes in length with options to extend sections. The model supports custom lyric input, vocal style control, and instrumental-only generation. Udio v1.5 demonstrates strong capabilities in complex musical genres including jazz with appropriate improvisation patterns, classical with correct orchestration, and electronic music with sophisticated sound design. Available through Udio's web platform with a freemium model offering limited free generations, the platform competes directly with Suno as the other leading AI music generation service, with Udio generally preferred for instrumental quality and genre precision.
La version détaillée est disponible en anglais.
Unique3D
Unique3D generates a concept mesh from one image using predicted views, normals and ISOMER. Check framing, hidden geometry and asset cleanup before treating the result as usable.
La version détaillée est disponible en anglais.
Upscayl
Upscayl is a free and open-source desktop application for AI-powered image upscaling, built on top of Real-ESRGAN and other super-resolution models. Developed by Nayam Amarshe and TGS963, Upscayl provides a user-friendly graphical interface that makes advanced AI image upscaling accessible to non-technical users on Windows, macOS, and Linux platforms. The application wraps multiple AI upscaling models in an Electron-based desktop app, allowing users to enhance image resolution with just a few clicks without any command-line knowledge or Python environment setup. Upscayl includes several pre-installed upscaling models optimized for different content types including general photography, digital art, anime, and sharpening, with each model producing different aesthetic characteristics suited to its target content. Users can select upscaling factors of 2x, 3x, or 4x and process individual images or entire folders through batch processing. The application supports common image formats including PNG, JPG, and WebP, and provides options for output format and quality settings. Upscayl also supports custom model loading, allowing users to import additional NCNN-compatible upscaling models from the community. Released under the AGPL-3.0 license, Upscayl is fully open source with its code available on GitHub and has accumulated a large community of users and contributors. The application runs entirely locally with no internet connection required, ensuring privacy for sensitive images. Upscayl is particularly popular among photographers, graphic designers, content creators, and hobbyists who need a simple, free solution for enhancing image quality without subscriptions or cloud processing dependencies.
La version détaillée est disponible en anglais.
VALL-E
VALL-E is a neural codec language model for text-to-speech synthesis developed by Microsoft Research, introduced in January 2023. Unlike traditional TTS systems that use mel spectrograms and vocoders, VALL-E treats text-to-speech as a conditional language modeling task, generating discrete audio codec codes from text input conditioned on a short audio prompt. The model uses a combination of autoregressive and non-autoregressive transformer decoders operating on EnCodec audio tokens to synthesize speech that preserves the speaker's voice characteristics, emotional tone, and acoustic environment from just a 3-second reference audio sample. This approach enables remarkable zero-shot voice cloning capabilities where the model can generate speech in any voice after hearing only a brief sample, without requiring speaker-specific fine-tuning. VALL-E was trained on 60,000 hours of English speech data from the LibriLight dataset, giving it exposure to a vast diversity of speakers, accents, and speaking styles. The generated speech maintains natural prosody, appropriate pausing, and emotional expressiveness that closely matches the reference speaker's characteristics. VALL-E represents a paradigm shift in TTS technology by demonstrating that language modeling approaches can effectively solve speech synthesis when paired with neural audio codecs. Released under a research-only license, the model is not available for commercial use, reflecting Microsoft's cautious approach given potential misuse concerns. VALL-E has significantly influenced subsequent research in zero-shot TTS, with its architecture inspiring numerous follow-up models. The model is particularly relevant for researchers studying speech synthesis, voice conversion, and the application of language modeling techniques to audio generation tasks.
La version détaillée est disponible en anglais.
Veo 2
Veo 2 is Google DeepMind's most advanced video generation model, capable of producing high-quality video content with up to 4K resolution, representing the cutting edge of AI-powered video synthesis. Released in December 2024, Veo 2 builds upon Google's extensive research in video understanding, delivering significant improvements in visual fidelity, motion realism, temporal coherence, and prompt comprehension. The model supports both text-to-video and image-to-video modes, interpreting detailed descriptions to create sequences that accurately reflect specified scenes, characters, actions, and atmospheric conditions. Veo 2 demonstrates exceptional understanding of real-world physics, generating videos with realistic lighting, shadows, reflections, and material properties. The model handles complex cinematic concepts including depth of field, camera movements like dolly shots and crane movements, and advanced compositional techniques, enabling footage that rivals professional cinematography. Veo 2 excels at maintaining character consistency across extended sequences, generating natural human motion and facial expressions, and producing content in diverse styles from photorealistic footage to animation and artistic interpretations. The model supports longer video sequences compared to most competitors, with improved temporal stability that reduces flickering and morphing artifacts. As a proprietary model, Veo 2 is currently available through limited access channels within Google's ecosystem, with plans for broader integration into Google products. The model represents Google's strategic positioning in the competitive AI video generation landscape alongside OpenAI's Sora and Runway's Gen-3 Alpha.
La version détaillée est disponible en anglais.
Wan 3.0
Wan 3.0 generates video from text and reference materials through Alibaba Cloud Model Studio. Use this source-based evaluation to plan a product-video trial and check its limits.
La version détaillée est disponible en anglais.
Wan Video
Wan Video is an open-source video generation suite developed by Alibaba that offers multiple model sizes for text-to-video generation, providing scalable options from lightweight variants for rapid experimentation to large-scale models for production-quality output. Released in February 2025, Wan Video represents Alibaba's significant contribution to the open-source video generation ecosystem, with the largest variant featuring 14 billion parameters making it one of the most powerful freely available video generation models. Built on a transformer-based architecture that processes text prompts through advanced language understanding modules, it generates temporally coherent video sequences through latent diffusion. Wan Video supports multiple output resolutions and aspect ratios for different platforms and use cases. The model demonstrates strong capabilities in generating diverse video content including realistic human subjects with natural motion, environmental scenes with dynamic elements, creative animations, and stylized artistic interpretations. The multi-size approach allows users to choose appropriate trade-offs between quality and computational requirements, with smaller variants enabling consumer-grade hardware deployment while larger variants deliver state-of-the-art quality. Wan Video incorporates advanced temporal modeling techniques maintaining consistency across frames, reducing common artifacts such as flickering, morphing, and identity drift. Available under the Apache 2.0 license, the suite is accessible on Hugging Face and through fal.ai and Replicate. The release includes comprehensive documentation and training code, enabling the research community to study and build upon Alibaba's advances for both academic and commercial applications.
La version détaillée est disponible en anglais.
Wan Video 2.1
Wan Video 2.1 is Alibaba's open-source video generation model combining high visual quality with controllable generation capabilities, making it one of the most capable freely available video synthesis solutions. Built on a diffusion transformer architecture, it supports text-to-video and image-to-video generation with enhanced temporal consistency, smooth motion, and improved visual fidelity compared to earlier open-source video models. Wan Video 2.1 introduces controllability features allowing users to guide generation through conditioning signals beyond text prompts, including motion control, camera trajectory specification, and reference image styling, providing creative control approaching proprietary solutions. The model handles diverse content from realistic human motion to natural landscapes, architectural environments, and stylized artistic content with consistent quality. Multiple model variants with different parameter counts are available for various hardware capabilities, from lightweight versions for consumer GPUs to full-scale models for maximum quality. The Apache 2.0 open-source license encourages community extensions, custom fine-tuning, and integration into creative pipelines. Wan Video 2.1 runs locally without cloud dependencies, ensuring data privacy and eliminating subscription costs. Applications include social media content creation, advertising video production, film concept visualization, educational materials, and creative experimentation. The model is available through Hugging Face with documentation and integration with ComfyUI and Diffusers. Wan Video 2.1 positions Alibaba as a major contributor to the open-source video generation ecosystem, providing a competitive alternative to proprietary models from Runway, Google, and OpenAI.
La version détaillée est disponible en anglais.
Wav2Lip
Wav2Lip is a deep learning model developed by researchers at IIIT Hyderabad that generates perfectly synchronized lip movements from any audio recording, representing a breakthrough in visual speech synthesis. The model takes a face video and an audio track as input, then produces realistic lip movements that precisely match the spoken content while preserving the original facial identity, expressions, and head movements. Built on a GAN (Generative Adversarial Network) architecture, Wav2Lip employs a pre-trained lip-sync discriminator that ensures the generated mouth movements are perceptually indistinguishable from real speech. This discriminator evaluates sync quality at a fine-grained level, resulting in significantly more accurate lip synchronization than previous approaches. The model works with any face regardless of identity, ethnicity, or language, and handles various audio types including speech, singing, and dubbed content. Wav2Lip operates on pre-recorded videos as well as static images which it animates with speech-driven lip movements. Released under the Apache 2.0 license, it is fully open source and has been widely adopted by the content creation community. Common applications include dubbing foreign language films, creating multilingual video content, animating avatars and virtual characters, producing educational materials with synthetic presenters, and accessibility applications for hearing-impaired users. The model can process videos at reasonable speeds on consumer GPUs and integrates with popular video editing pipelines for professional production workflows.
La version détaillée est disponible en anglais.
Whisper Large v3
Whisper Large v3 is the most advanced multilingual automatic speech recognition model developed by OpenAI, featuring 1.55 billion parameters trained on over 680,000 hours of diverse audio data spanning more than 100 languages. Built on an Encoder-Decoder Transformer architecture, the model takes raw audio waveforms as input and outputs accurate text transcriptions with punctuation, capitalization, and speaker-appropriate formatting. Whisper Large v3 achieves near-human accuracy for English transcription and delivers strong performance across dozens of languages including low-resource languages that other ASR systems struggle with. The model supports both transcription of speech in the source language and direct translation to English, enabling cross-lingual content accessibility from a single model. Key improvements in v3 over previous versions include expanded language coverage, reduced hallucination on silent or noisy audio segments, better handling of accented speech, and improved timestamp accuracy for subtitle generation. Whisper Large v3 processes audio in 30-second chunks with a sliding window approach, handling recordings of any length from brief voice messages to multi-hour lectures and podcasts. Released under the MIT license, the model is fully open source and has become the gold standard for open ASR systems. It is available through Hugging Face, integrates with the Transformers library, and can be accelerated with frameworks like faster-whisper and whisper.cpp for real-time processing. Common applications include meeting transcription, podcast and video captioning, voice-to-text input, medical dictation, legal transcription, accessibility services for hearing-impaired users, content indexing for search, and building voice-controlled applications across multilingual markets.
La version détaillée est disponible en anglais.
Wonder3D
Wonder3D predicts paired color and normal views from one image, then reconstructs a mesh. Inspect masks, camera assumptions and implementation version before judging a result.
La version détaillée est disponible en anglais.
Wuerstchen
Wuerstchen is a highly efficient text-to-image generation model developed by researchers at Stability AI that introduces a novel three-stage architecture operating in an extremely compressed latent space, achieving dramatic improvements in both training and inference efficiency. The model's key innovation is its use of a 42x compression ratio in its latent space, far exceeding the 8x compression used by standard latent diffusion models like Stable Diffusion. This extreme compression is achieved through a hierarchical approach where Stage C works with tiny 24x24 latent representations, Stage B decodes these to intermediate resolution, and Stage A produces the final output. Despite this aggressive compression, Wuerstchen maintains image quality competitive with much more computationally expensive models. The architecture enables training on consumer hardware and significantly faster inference times compared to models of similar output quality. Wuerstchen can generate a 1024x1024 image using substantially less memory and compute than SDXL while maintaining comparable quality. The model served as the architectural foundation for Stable Cascade, validating its design principles for broader deployment. Released as open-source, Wuerstchen is available on Hugging Face and compatible with the Diffusers library. AI researchers studying efficient generative model architectures, developers building resource-constrained applications, and academic institutions with limited GPU access particularly value Wuerstchen. The model demonstrates that extreme latent space compression can be a viable path toward democratizing high-quality image generation by making it accessible on less powerful hardware.
La version détaillée est disponible en anglais.
XTTS v2
XTTS v2 (Cross-lingual Text-to-Speech v2) is a multilingual voice cloning and text-to-speech model developed by Coqui AI that can replicate any person's voice from just a 6-second audio sample and synthesize speech in 17 supported languages. Built on a GPT-like autoregressive architecture paired with a HiFi-GAN vocoder, XTTS v2 with 467 million parameters produces natural-sounding speech with realistic prosody, intonation, and emotional expressiveness. The model's cross-lingual capability allows a voice cloned from an English sample to speak fluently in French, Spanish, German, Turkish, and other supported languages while maintaining the original speaker's vocal characteristics. XTTS v2 achieves this through a language-agnostic speaker embedding space that separates voice identity from linguistic content. The synthesis quality approaches human-level naturalness for many languages, with particularly strong performance in English, Spanish, and Portuguese. The model supports streaming inference for real-time applications, generating speech with latencies suitable for conversational AI and interactive voice assistants. Released under the MPL-2.0 license, XTTS v2 is open source and can be deployed locally for privacy-sensitive applications. Common use cases include creating multilingual audiobook narrations, localizing video content with consistent voice identity, building accessible text-to-speech interfaces, developing custom voice assistants, podcast production, and e-learning content creation. The model provides a Python API and can be fine-tuned on additional voice data for improved quality with specific speakers or specialized domains.
La version détaillée est disponible en anglais.
YOLOv10
YOLOv10 is the tenth major iteration of the YOLO (You Only Look Once) real-time object detection series, developed by researchers at Tsinghua University. The model introduces a fundamentally redesigned NMS-free (Non-Maximum Suppression free) architecture that eliminates the post-processing bottleneck present in all previous YOLO versions, enabling true end-to-end object detection with consistent latency. YOLOv10 employs a dual-assignment training strategy that combines one-to-many and one-to-one label assignments during training, achieving rich supervision signals while maintaining efficient inference without redundant predictions. Built on a CSPNet backbone with enhanced feature aggregation, the model comes in six scale variants ranging from Nano (8M parameters) to Extra-Large (68M parameters), allowing deployment across edge devices, mobile platforms, and high-performance servers. Each variant is optimized for its target hardware profile, delivering the best accuracy-latency trade-off in its class. YOLOv10 achieves state-of-the-art performance on the COCO benchmark, outperforming previous YOLO versions and competing models like RT-DETR with significantly lower computational cost. Released under the AGPL-3.0 license, the model is open source and integrates seamlessly with the Ultralytics ecosystem for training, validation, and deployment. Common applications include autonomous driving perception, industrial quality inspection, security surveillance, retail analytics, robotics, and drone-based monitoring. The model supports ONNX and TensorRT export for optimized production deployment.
La version détaillée est disponible en anglais.
Zero123++
Zero123++ is a multi-view image generation model developed by Stability AI that generates six consistent canonical views of an object from a single input image. Released in 2023 under the Apache 2.0 license, the model extends the original Zero123 approach with significantly improved view consistency and serves as a critical component in modern 3D reconstruction pipelines. Zero123++ takes a single photograph or rendered image of an object and produces six evenly spaced views covering the full 360-degree range around the object, all maintaining consistent geometry, lighting, and appearance. The model is built on a fine-tuned Stable Diffusion backbone with specialized conditioning mechanisms that ensure multi-view coherence. Unlike the original Zero123 which generates views independently and often produces inconsistent results, Zero123++ generates all six views simultaneously in a single diffusion process, dramatically improving 3D consistency. The generated multi-view images serve as input for downstream 3D reconstruction methods like NeRF, Gaussian Splatting, or direct mesh reconstruction, enabling high-quality 3D model creation from a single photograph. Zero123++ is fully open source with pre-trained weights available on Hugging Face, making it accessible to researchers and developers building 3D generation systems. The model has become a foundational component in many state-of-the-art 3D generation pipelines and is widely used in academic research. It is particularly valuable for applications in game development, product visualization, and virtual reality where converting 2D images to 3D assets is a frequent workflow requirement.
La version détaillée est disponible en anglais.
160 modèles trouvés · Sayfa 7 / 7