ホームモデル

AIモデル

AIモデルを名前順に探し、公開資料に基づく機能を比較できます

160 件のモデル
LivePortrait icon

LivePortrait

Kuaishou|Unknown

LivePortrait is an efficient AI portrait animation model developed by Kuaishou Technology that generates expressive and lifelike facial animations from a single static portrait photograph. The model takes a source portrait image and a driving video containing facial movements, then transfers the expressions, head rotations, eye movements, and mouth gestures from the video onto the portrait while maintaining the original person's identity and appearance. Built on an implicit keypoint detection architecture with warping-based rendering, LivePortrait achieves real-time inference speeds that make it practical for interactive applications and live content creation. The model introduces stitching and retargeting modules that prevent common artifacts in portrait animation such as face boundary distortion, neck disconnection, and unnatural eye movements, producing seamless results that preserve the natural appearance of the subject. LivePortrait handles diverse portrait types including photographs, paintings, illustrations, and even cartoon characters, adapting its animation approach to different artistic styles. The model supports fine-grained control over individual facial action units, allowing selective animation of specific facial features like eyebrow raises, eye blinks, or smile intensity independently. Released under the MIT license, LivePortrait is fully open source and has been integrated into ComfyUI and other creative tools. Common applications include creating animated avatars for social media and messaging, producing animated portrait NFTs, generating facial animations for virtual presenters and digital humans, creating engaging content from historical photographs, and building interactive portrait experiences for museums and exhibitions.

詳細版は英語で提供されています。

公開ウェイト
LTX Video icon

LTX Video

Lightricks|N/A

LTX Video is a real-time video generation model developed by Lightricks that produces 768x512 resolution videos at 24 frames per second, emphasizing generation speed and efficiency without sacrificing visual quality. Released in November 2024, LTX Video is built on a transformer-based architecture optimized for rapid inference, capable of generating video content faster than many competing models, making it suitable for interactive applications requiring quick iteration. The model supports text-to-video generation, interpreting natural language descriptions to produce short clips with coherent motion, consistent scene dynamics, and visually appealing quality. LTX Video's architecture incorporates efficient attention mechanisms and optimized latent space operations that reduce computational requirements while maintaining quality for professional creative applications. The model demonstrates competence in generating diverse content types including human subjects with natural motion, environmental scenes with dynamic elements, abstract visual content, and stylized artistic interpretations. LTX Video supports integration with existing creative workflows through API availability and compatibility with popular development frameworks. The emphasis on real-time performance makes it valuable for interactive content creation tools, live preview systems, and prototype generation where extended wait times would disrupt creative flow. Available under the Apache 2.0 license, LTX Video is accessible on Hugging Face and through fal.ai and Replicate, enabling both local deployment and cloud-based integration. Lightricks' background as a creative tools company is reflected in the model's focus on practical usability, with optimizations targeted at content creators and designers who prioritize workflow efficiency alongside output quality.

詳細版は英語で提供されています。

公開ウェイト
Luma Dream Machine icon

Luma Dream Machine

Luma AI|N/A

Luma Dream Machine is a fast video generation model developed by Luma AI that creates realistic five-second video clips from text prompts or reference images with impressive speed and visual quality. Released in June 2024, Dream Machine leverages a transformer-based architecture trained on large-scale video data to produce clips with natural motion dynamics, consistent character appearances, and physically coherent scene transitions. The model's standout feature is its generation speed, producing outputs significantly faster than many competing models while maintaining competitive visual quality, making it suitable for iterative creative workflows. Dream Machine supports both text-to-video mode, where users describe scenes through detailed prompts, and image-to-video mode, where a still image serves as the starting frame and the model generates plausible forward motion. The model demonstrates strong capabilities in generating human motion, environmental dynamics like water flow and wind effects, camera movements, and lighting transitions. It handles various visual styles from photorealistic content to stylized and artistic interpretations. Dream Machine's architecture enables it to understand spatial relationships and maintain 3D consistency throughout generated sequences, producing videos where objects maintain relative positions across frames. Available as a proprietary service through Luma AI's platform and accessible via API through fal.ai and Replicate, Dream Machine operates on a credit-based pricing model with free tier access. The model has become popular among content creators, filmmakers, and designers who value the combination of generation speed and output quality for rapid visual prototyping and content production.

詳細版は英語で提供されています。

プロプライエタリ
Luma Image-to-Video icon

Luma Image-to-Video

Luma AI|N/A

Luma Image-to-Video is the image animation capability of Luma AI's Dream Machine, designed to create compelling video content from still images by generating natural motion dynamics with the model's transformer-based architecture. Released in June 2024, this feature enables users to transform photographs, illustrations, and digital artwork into animated sequences where subjects move naturally, environments come alive, and camera perspectives shift with cinematic fluidity. The model analyzes the input image to understand spatial composition, depth layers, and semantic content, then generates contextually appropriate motion maintaining the source's visual identity throughout. Dream Machine's image-to-video mode benefits from the same fast generation speed as the text-to-video capability, producing results significantly faster than many competitors and enabling rapid iteration. The model demonstrates competence in generating human movement and expressions, environmental dynamics like flowing water and swaying vegetation, camera movements, and atmospheric effects. Users can optionally provide text prompts alongside the reference image to guide generated motion direction. The model supports various output resolutions and durations adapting to different platform requirements. Available through Luma AI's platform and via API through fal.ai and Replicate, it operates on the Dream Machine credit system with free tier access. The feature has become popular among social media creators, digital artists, and marketing professionals who need to quickly produce animated content from existing visual assets without specialized animation skills.

詳細版は英語で提供されています。

プロプライエタリ
Meshy icon

Meshy

Meshy AI|N/A

Meshy is a proprietary AI-powered 3D generation platform developed by Meshy AI that creates detailed, production-ready 3D models from text descriptions and images. The platform combines text-to-3D and image-to-3D capabilities with advanced AI texturing features, positioning itself as a comprehensive solution for rapid 3D content creation. Meshy uses a transformer-based architecture that generates textured 3D meshes with PBR-compatible materials, making outputs directly usable in game engines like Unity and Unreal Engine without additional processing. The platform offers multiple generation modes including text-to-3D for creating objects from written descriptions, image-to-3D for converting photographs into 3D models, and AI texturing for applying realistic materials to existing untextured meshes. Generated models include proper UV mapping, normal maps, and physically based rendering materials suitable for professional workflows. Meshy provides both a web-based interface and an API for programmatic access, making it accessible to individual artists and scalable for enterprise pipelines. The platform is particularly popular among game developers, animation studios, and AR/VR content creators who need to produce large volumes of 3D assets efficiently. As a proprietary commercial service launched in 2023, Meshy operates on a subscription model with free tier access for limited generations. The platform continuously updates its models to improve output quality, topology optimization, and texture fidelity, competing directly with other AI 3D generation services in the rapidly evolving market.

詳細版は英語で提供されています。

プロプライエタリ
Midjourney v6 icon

Midjourney v6

Midjourney|N/A

Midjourney v6 is the sixth major version, released on December 20, 2023, from Midjourney Inc., widely regarded as the industry leader in AI-generated art for its distinctive aesthetic quality and photorealistic capabilities. Accessible exclusively through Discord and the Midjourney web interface, v6 introduced significant improvements in prompt understanding, coherence, and image quality over its predecessors. The model excels at producing visually stunning images with remarkable attention to lighting, texture, composition, and mood that many users describe as having a distinctive cinematic quality. Midjourney v6 demonstrates strong performance in photorealistic rendering, achieving results that are frequently indistinguishable from professional photography in controlled comparisons. It handles complex artistic directions well, understanding nuanced descriptions of style, atmosphere, and emotional tone. The model supports various output modes including standard and raw styles, upscaling options, and aspect ratio customization. While it is a closed-source proprietary model with no publicly available weights, its consistent quality and ease of use have made it the most popular commercial AI image generator. Creative professionals, illustrators, concept artists, marketing teams, and hobbyists rely on Midjourney v6 for everything from professional portfolio work to social media content and creative exploration. The subscription-based pricing model offers different tiers to accommodate casual users and high-volume professionals. Version-specific feature compatibility should be checked before reusing an older prompt.

詳細版は英語で提供されています。

プロプライエタリ
MiniMax H3 icon

MiniMax H3

MiniMax

MiniMax H3 generates video with audio from text and media references. This source-based evaluation separates hosted access, restricted open weights and the checks a product scene needs.

詳細版は英語で提供されています。

公開ウェイト
Minimax Video-01 icon

Minimax Video-01

MiniMax|undisclosed

Minimax Video-01 is MiniMax's flagship video generation model that powers the Hailuo AI platform, capable of generating high-quality video clips from text descriptions and images. Released in September 2024, the model quickly gained attention for producing remarkably natural motion, cinematic camera movements, and consistent character depiction across video frames. Video-01 generates clips up to 6 seconds at 720p resolution with smooth 25fps playback. The model demonstrates particular strength in realistic human movement, facial expressions, and environmental dynamics like water flow, fire, and wind effects. Unlike many competitors that produce visually impressive but physically implausible motion, Video-01 maintains strong physical consistency throughout generated clips. The model supports both text-to-video and image-to-video generation modes, allowing users to animate still images with natural motion while preserving the original image's style and composition. MiniMax's approach combines a large-scale transformer architecture with temporal attention mechanisms to ensure frame-to-frame coherence. The model is accessible through the Hailuo AI web platform with a freemium model offering limited free generations and paid plans for higher volume usage. Video-01 competes with Runway Gen-3, Kling 1.5, and Luma Dream Machine in the consumer video generation space, with particular advantages in natural motion quality and free-tier accessibility.

詳細版は英語で提供されています。

プロプライエタリ
Mochi 1 icon

Mochi 1

Genmo|10B

Mochi 1 is an open-source video generation model developed by Genmo that delivers high motion fidelity and temporal consistency, establishing itself as one of the most capable freely available video generation models. Released in October 2024 with 10 billion parameters, Mochi 1 produces clips with remarkably smooth motion, consistent character appearances, and natural scene dynamics that rival some proprietary alternatives. Built on a transformer architecture that processes text prompts through a language encoder and generates video through iterative denoising, it features architectural innovations focused on maintaining temporal coherence across extended frame sequences. Mochi 1 demonstrates strong capabilities in generating realistic human motion, facial expressions, camera movements, and physical interactions between objects, areas where many competing open-source models produce noticeable artifacts. The model supports text-to-video generation with detailed prompt interpretation, producing clips that accurately reflect specified scenes, actions, and styles. At 10 billion parameters, it is one of the largest open-source video generation models, and this scale contributes to superior ability to capture complex visual details and maintain consistency throughout sequences. The model handles diverse visual styles including photorealistic content, stylized animation, and artistic interpretations. Available under the Apache 2.0 license, Mochi 1 is accessible on Hugging Face and through fal.ai and Replicate, enabling both research and commercial applications. The model has received particular praise for its motion quality, setting a new standard for open-source video generation and providing a compelling alternative for developers who need capable video generation without the constraints and costs of proprietary API services.

詳細版は英語で提供されています。

公開ウェイト
Mochi 1 Preview icon

Mochi 1 Preview

Genmo|10B

Mochi 1 Preview is an open-source text-to-video AI model developed by Genmo that sets a new standard for motion quality and physical realism in generated video content. With 10 billion parameters built on an Asymmetric Diffusion Transformer architecture, Mochi 1 Preview produces videos with remarkably natural and physically plausible motion that distinguishes it from competing models. The asymmetric architecture processes spatial and temporal information through dedicated pathways optimized for their respective characteristics, resulting in videos where objects move with realistic momentum, gravity, and interaction dynamics. Mochi 1 Preview generates 480p resolution videos at 30 frames per second with smooth, continuous motion free from the temporal flickering and object morphing artifacts common in earlier video generation models. The model demonstrates strong understanding of real-world physics including fluid dynamics, rigid body interactions, and natural phenomena like fire, smoke, and water, producing content that feels grounded in physical reality. Mochi 1 Preview responds well to detailed text prompts describing camera movements, scene transitions, and specific motion choreography, giving creators meaningful control over the generated output. Released under the Apache 2.0 license, the model is fully open source and represents one of the strongest open alternatives to proprietary video generation services. It is available through Hugging Face and supported by cloud inference providers for accessible deployment. Key applications include creating concept videos for film and advertising pre-production, generating social media video content, producing animated product demonstrations, creating visual references for motion design projects, and prototyping video ideas before committing to expensive live-action production.

詳細版は英語で提供されています。

公開ウェイト
ModelScope T2V icon

ModelScope T2V

Alibaba DAMO|1.7B

ModelScope T2V is an early open-source text-to-video generation model developed by Alibaba DAMO Academy that pioneered accessible video generation research by making a functional text-to-video pipeline freely available. Released in March 2023, ModelScope T2V was among the first open-source models to demonstrate practical text-to-video capabilities, establishing an important baseline for subsequent developments. Built on a 1.7 billion parameter diffusion architecture, it extends latent diffusion to the temporal domain, incorporating temporal convolution and attention layers for generating short video clips from text descriptions. The architecture processes text prompts through a CLIP encoder and generates video through a modified U-Net with temporal dimensions, producing clips with basic motion coherence and prompt alignment. While output quality is modest compared to recent models like Sora or Runway Gen-3, ModelScope T2V played a crucial historical role in democratizing video generation technology by providing the first truly accessible open-source implementation that researchers could experiment with, modify, and build upon. The model supports generation of short clips at moderate resolutions, handling simple scene descriptions with recognizable subjects and basic motion patterns. Common use cases include research experimentation, educational demonstrations of video generation concepts, rapid prototyping, and serving as a baseline for training more advanced models. Available under the Apache 2.0 license on Hugging Face and Replicate, ModelScope T2V remains relevant as a lightweight, resource-efficient option for scenarios where state-of-the-art quality is not required but functional video generation capability is needed with minimal computational overhead.

詳細版は英語で提供されています。

公開ウェイト
MODNet icon

MODNet

ZHKKKe|N/A

MODNet (Matting Objective Decomposition Network) is an open-source portrait matting model developed by ZHKKKe, designed for real-time human portrait background removal without requiring a pre-defined trimap or additional user input. Unlike traditional matting approaches needing manually drawn trimaps, MODNet achieves fully automatic portrait matting by decomposing the complex matting objective into three sub-tasks: semantic estimation for identifying the person region, detail prediction for refining edge quality around hair and clothing boundaries, and semantic-detail fusion for combining both signals into a high-quality alpha matte. This decomposition enables efficient single-pass inference at real-time speeds, making it practical for video conferencing, live streaming, and mobile photography where latency is critical. The model produces smooth and accurate alpha mattes with particular strength in handling hair strands, fabric edges, and other fine boundary details challenging for segmentation-based approaches. MODNet supports both image and video input with temporal consistency optimizations for stable video matting without flickering. The model is lightweight enough for mobile devices and edge hardware, with ONNX export supporting deployment across iOS, Android, and web browsers through WebAssembly. Common applications include video call background replacement, portrait mode photography, social media content creation, virtual try-on systems, and film post-production green screen alternatives. Released under Apache 2.0, MODNet provides a free and efficient solution widely adopted in both research and production portrait matting applications.

詳細版は英語で提供されています。

公開ウェイト
MotionDiffuse icon

MotionDiffuse

Mingyuan Zhang et al.|200M

MotionDiffuse is a pioneering diffusion model developed by Mingyuan Zhang and collaborators that generates realistic 3D human motion sequences from natural language text descriptions. The model takes text prompts such as 'a person walks forward and waves' or 'someone performs a backflip' and produces corresponding 3D skeleton-based animation data with natural body dynamics and physical plausibility. Built on a diffusion architecture with approximately 200 million parameters, MotionDiffuse introduces probabilistic motion generation that captures the inherent diversity of human movement, generating multiple plausible motion variations for the same text input. The model supports both single-action and sequential multi-action generation, enabling the creation of complex motion sequences that smoothly transition between different activities. MotionDiffuse was trained on large-scale motion capture datasets including HumanML3D and KIT-ML, learning to map semantic descriptions to physically realistic joint rotations and translations across the full body skeleton. The generated motion data can be exported in standard formats compatible with 3D animation software including Blender, Maya, and Unity, making it practical for professional production workflows. Released under the MIT license, the model is fully open source and available for both research and commercial applications. Key use cases include generating character animations for games and films, creating training data for pose estimation models, prototyping choreography, producing VR and AR avatar movements, and automating repetitive animation tasks that traditionally require skilled motion capture artists and extensive studio equipment.

詳細版は英語で提供されています。

公開ウェイト
MusicGen icon

MusicGen

Meta|3.3B

MusicGen is a single-stage transformer-based music generation model developed by Meta AI Research as part of the AudioCraft framework. Released in June 2023 under the MIT license, MusicGen uses a single autoregressive language model operating over compressed discrete audio representations from EnCodec, unlike cascading approaches that require multiple models. The model comes in multiple sizes ranging from 300M to 3.3B parameters, allowing users to balance quality against computational requirements. MusicGen generates high-quality mono and stereo music at 32 kHz from text descriptions, supporting a wide range of genres, instruments, moods, and musical styles. Users can describe desired music using natural language prompts specifying genre, tempo, instrumentation, and atmosphere, and the model produces coherent musical compositions that follow the specified characteristics. Beyond text-to-music generation, MusicGen supports melody conditioning where an existing audio clip guides the melodic structure of the generated output, enabling more controlled music creation. The model achieves strong results across both objective metrics and subjective listening evaluations, producing music that sounds natural and musically coherent for durations up to 30 seconds. As a fully open-source model with code and weights available on GitHub and Hugging Face, MusicGen has become one of the most widely adopted AI music generation tools in both research and creative communities. It integrates easily into existing audio production workflows through the Audiocraft Python library and various community-built interfaces. MusicGen is particularly popular among content creators, game developers, and musicians who need royalty-free background music generated on demand.

詳細版は英語で提供されています。

公開ウェイト
MusicLM icon

MusicLM

Google|N/A

MusicLM is a text-to-music generation model developed by Google Research that generates high-fidelity music from text descriptions at 24 kHz. Published in January 2023 alongside a research paper, MusicLM was one of the first models to demonstrate that AI could generate coherent, high-quality music spanning multiple minutes from natural language descriptions alone. The model employs a hierarchical sequence-to-sequence architecture combining SoundStream for audio tokenization and w2v-BERT for audio representation learning, generating music tokens at multiple temporal resolutions that are then decoded into waveforms. MusicLM can produce music in diverse genres and styles based on text prompts describing instruments, tempo, mood, and musical characteristics, maintaining musical coherence and structural consistency across extended durations. The model also supports melody conditioning where users can hum or whistle a melody that guides the generated output, enabling more intuitive music creation workflows. MusicLM generates audio with rich timbral quality and natural-sounding dynamics that represent a significant improvement over earlier text-to-music approaches. As a proprietary Google model, MusicLM is not open source and was initially accessible only through the AI Test Kitchen experimental platform before being integrated into broader Google services. While newer models like MusicGen and Suno have since achieved wider adoption, MusicLM remains historically significant as a pioneering demonstration of high-quality text-to-music generation. The model influenced subsequent research and commercial developments in the AI music generation space and helped establish text-to-music as a viable and rapidly advancing field of AI research.

詳細版は英語で提供されています。

プロプライエタリ
Nano Banana icon

Nano Banana

Google DeepMind|undisclosed

Nano Banana (technically Gemini 2.5 Flash Image) is Google DeepMind's breakthrough text-to-image model that went viral upon release in August 2025. Built on diffusion-based multimodal technology integrated within the Gemini ecosystem, it generates photorealistic images from text prompts and enables conversational image editing through chat. The model gained massive popularity through its distinctive 3D figurine-style outputs that became a social media phenomenon. Unlike standalone image generators, Nano Banana operates within the Gemini chat interface, allowing users to describe desired images, receive results, and iteratively refine them through natural conversation. It supports a wide range of styles from photorealism to illustration, handles complex multi-element compositions, and produces readable text within images. Available for free through the Gemini app with SynthID watermarking, it democratized AI image generation for millions of users who had never used dedicated image AI tools before.

詳細版は英語で提供されています。

プロプライエタリ
Nano Banana 2 icon

Nano Banana 2

Google DeepMind|Undisclosed

Nano Banana 2, also called Gemini 3.1 Flash Image, creates and edits images. Evaluate reference fidelity, bilingual lettering, search grounding and API costs for your delivery.

詳細版は英語で提供されています。

プロプライエタリ
Nano Banana Pro icon

Nano Banana Pro

Google DeepMind|undisclosed

Nano Banana Pro (technically Gemini 3 Pro Image) is Google DeepMind's premium image generation model released in November 2025, offering studio-grade visual quality with up to 4K resolution output. It builds upon the original Nano Banana with dramatically improved capabilities including multi-language readable text rendering for infographics and marketing materials, face consistency for up to 5 different people across multiple images, advanced creative controls for camera angles, lighting, depth of field, and color grading, and web search grounding for real-time data integration. The model features a 'thinking mode' for complex prompts that require multi-step reasoning, and supports localized image editing to modify specific regions without affecting the rest. Scoring 9.0+ out of 10 for photorealism and ranking as the second-best model for text generation after GPT Image, it delivers professional-grade results suitable for commercial design, marketing, and brand content creation. Available through the Gemini app (free with watermark), Google AI Studio, Gemini API, and Google Ads, with pricing starting at $0.134 per 1K image through the API.

詳細版は英語で提供されています。

プロプライエタリ
Neural Style Transfer icon

Neural Style Transfer

Leon Gatys|N/A

Neural Style Transfer is the pioneering algorithm introduced by Leon Gatys, Alexander Ecker, and Matthias Bethge in their landmark 2015 paper that demonstrated how convolutional neural networks can separate and recombine the content and style of images. The algorithm takes two input images, a content image and a style reference, then iteratively optimizes a generated output to simultaneously match the content structure of one and the artistic style of the other using feature representations extracted from a pre-trained VGG-19 network. Deep layers capture high-level content information like object shapes and spatial arrangements, while shallow layers encode style characteristics including textures, colors, and brush stroke patterns. By defining separate content and style loss functions based on these feature representations and minimizing their weighted combination through gradient descent, the algorithm produces images that preserve the recognizable content of photographs while adopting the visual aesthetic of paintings or other artistic works. This foundational work sparked an entire field of AI-powered artistic image transformation and inspired numerous real-time variants, mobile applications, and commercial products. While the original optimization-based approach requires several minutes per image on a GPU, subsequent feed-forward network approaches by Johnson et al. and others achieved real-time performance. The algorithm is fully open source with implementations available in PyTorch, TensorFlow, and other frameworks. Neural Style Transfer remains a cornerstone reference in computer vision education and continues to influence modern style transfer research and generative AI development.

詳細版は英語で提供されています。

公開ウェイト
One-2-3-45 icon

One-2-3-45

UC San Diego|N/A

One-2-3-45 is a single-image 3D reconstruction system developed by researchers at UC San Diego that generates textured 3D meshes from a single input image through a two-stage pipeline combining multi-view generation with sparse-view 3D reconstruction. The name reflects the core process: from one image, generate two to three to four to five views, then reconstruct a complete 3D object. In the first stage, a fine-tuned Zero123 model generates multiple novel views of the object from different angles based on the single input photograph. In the second stage, these generated multi-view images are fed into a cost-volume-based sparse-view reconstruction network that produces a textured 3D mesh with consistent geometry. Released in June 2023 under the MIT license, One-2-3-45 was among the first systems to demonstrate that combining 2D diffusion models with 3D reconstruction could produce reasonable 3D assets in under a minute. The model handles a variety of object types including everyday items, animals, vehicles, and artistic objects. Unlike optimization-based approaches like DreamFusion that require per-object optimization taking tens of minutes, One-2-3-45 runs in a feed-forward manner making it significantly faster. The output meshes include color and texture information and can be exported for use in standard 3D applications. As a fully open-source project with code available on GitHub, it has served as an influential reference for subsequent research in single-image 3D generation. The system is particularly useful for researchers and developers exploring rapid 3D content creation from limited input data.

詳細版は英語で提供されています。

公開ウェイト
OpenJourney icon

OpenJourney

PromptHero|1B

Openjourney is an open-source Stable Diffusion fine-tuned model created by PromptHero, trained specifically to replicate the distinctive artistic style of Midjourney outputs. The model was fine-tuned on a curated dataset of Midjourney-generated images, learning to produce the characteristic vibrant colors, dramatic lighting, cinematic compositions, and painterly aesthetic that made Midjourney famous. By using the trigger keyword in prompts, users can generate images with Midjourney-like quality without requiring a Midjourney subscription. Openjourney is built on Stable Diffusion 1.5, making it lightweight and accessible to run on consumer GPUs with as little as 4GB VRAM. The model became hugely popular in the early days of the open-source AI art movement as it democratized access to a Midjourney-inspired aesthetic for users who could not afford or access the subscription service. It supports all standard Stable Diffusion features including img2img, inpainting, and ControlNet conditioning. Available on Hugging Face and CivitAI, Openjourney integrates with ComfyUI, Automatic1111, and other popular Stable Diffusion interfaces. Digital artists, hobbyists, content creators, and developers building creative applications form its primary user base. While newer models like SDXL and FLUX.1 have surpassed its output quality and the Midjourney style has evolved significantly beyond what Openjourney captures, the model remains relevant as a lightweight option for artistic image generation and as a historically significant example of style transfer through fine-tuning in the open-source AI community.

詳細版は英語で提供されています。

公開ウェイト
OpenLRM icon

OpenLRM

Zexiang Xu|N/A

OpenLRM is an open-source implementation of the Large Reconstruction Model architecture for single-image 3D reconstruction, developed by Zexiang Xu and collaborators. The project provides a fully open and reproducible implementation of the LRM approach, which uses a transformer-based architecture to predict 3D representations from single input images in a feed-forward manner. OpenLRM processes an input image through a pre-trained vision encoder like DINOv2, then feeds the resulting features into a transformer decoder that generates a triplane-based neural radiance field representation, which can be rendered from novel viewpoints or converted to a textured 3D mesh. The entire reconstruction takes only a few seconds on a modern GPU, making it practical for interactive applications and batch processing workflows. Released under the Apache 2.0 license in December 2023, OpenLRM fills a critical gap in the 3D AI research community by providing an accessible reference implementation that researchers can study, modify, and build upon. The model supports various output formats and can be integrated into existing 3D pipelines for applications ranging from game development to e-commerce product visualization. OpenLRM handles diverse object categories including furniture, vehicles, characters, and everyday items with reasonable geometric fidelity. Pre-trained model weights are available on Hugging Face for immediate use. As one of the foundational open-source projects in feed-forward 3D reconstruction, OpenLRM has directly influenced and enabled numerous downstream projects and research efforts in the rapidly evolving single-image 3D generation space.

詳細版は英語で提供されています。

公開ウェイト
OpenPose icon

OpenPose

CMU|25M

OpenPose is the pioneering real-time multi-person pose estimation system developed at Carnegie Mellon University that simultaneously detects body, face, hand, and foot keypoints of multiple people in images and videos. As the first open-source system to achieve real-time multi-person pose detection, OpenPose has become a foundational tool in computer vision research and creative AI applications. Built on a CNN (Convolutional Neural Network) architecture with approximately 25 million parameters, the model uses Part Affinity Fields (PAFs) to associate detected body parts with the correct individuals in crowded scenes, enabling accurate pose estimation even when people overlap or partially occlude each other. OpenPose detects up to 135 keypoints per person covering the full body skeleton with 25 points, each hand with 21 points, and the face with 70 points, providing comprehensive pose information for detailed motion analysis. The system processes both images and video streams, delivering real-time performance on modern GPUs that makes it suitable for interactive applications. OpenPose has been extensively integrated into AI image generation workflows, particularly as the standard pose extraction method for ControlNet conditioning in Stable Diffusion and FLUX-based generation pipelines. Released under a custom non-commercial license, the source code is available on GitHub and has accumulated one of the highest star counts among computer vision repositories. Key applications include motion capture for animation and gaming, fitness and rehabilitation tracking, sports biomechanics analysis, sign language recognition, dance analysis, human-computer interaction research, and providing pose conditioning for AI image generation tools.

詳細版は英語で提供されています。

公開ウェイト
PaddleOCR icon

PaddleOCR

Baidu|15M

PaddleOCR is a comprehensive optical character recognition system developed by Baidu on the PaddlePaddle deep learning framework, supporting over 80 languages with industry-grade accuracy and speed. The latest PP-OCRv4 architecture employs a three-stage pipeline consisting of text detection, direction classification, and text recognition, each optimized independently for maximum performance. With approximately 15 million parameters in its lightweight configuration, PaddleOCR achieves an exceptional balance between accuracy and inference speed, running efficiently on both server GPUs and edge devices including mobile phones and embedded systems. The system excels at recognizing text in complex real-world scenarios including curved text, rotated text, dense multi-line layouts, and text overlaid on textured backgrounds. PaddleOCR supports Latin, Chinese, Japanese, Korean, Arabic, Cyrillic, and dozens of other scripts with dedicated recognition models for each language family. Beyond basic OCR, the toolkit includes document structure analysis for extracting tables, headers, and paragraphs from scanned documents, as well as key information extraction capabilities for invoices, receipts, and forms. Released under the Apache 2.0 license, PaddleOCR is fully open source and has become one of the most starred OCR repositories on GitHub. It provides pre-trained models, training scripts, and deployment tools for ONNX, TensorRT, and OpenVINO formats. Common applications include document digitization, license plate recognition, receipt processing, handwriting recognition, and industrial text inspection in manufacturing quality control.

詳細版は英語で提供されています。

公開ウェイト

160 件のモデル · Sayfa 4 / 7