Real-Time
Commercial Offerings
Runway Gen 3
Luma Dream Machine?
- Luma Dream Machine is a browser-based AI video generator developed by Luma Labs, a San Francisco-based startup. It allows users to generate short videos (around 5 seconds) by simply entering a text or image prompt.
- Free to Use: Luma Dream Machine is free to try, with no waiting list or subscription required. Users get 30 free video generations per month.
- High-Quality Output: The AI produces impressively clean and detailed videos, adhering to prompts accurately and generating relatively coherent motion.
- Fast Generation: Videos are generated in around 2 minutes after entering the prompt.
- Consistent Subjects: Characters and subjects appear consistent throughout the video, capable of expressing emotion better than many previous AI video models.
Limitations
While groundbreaking, and crucially, “available”, Luma Dream Machine still has some limitations, as acknowledged by the company:
- Morphing, warping, and unnatural movements
- Difficulty with complex scenes or full-body shots
- Text in videos may appear garbled
- Anatomical issues like extra limbs or heads
- (1) Professor John Keeting on X: “this was created with Luma AI I am very impressed. Made by Kevin Van Witt and the talented team at The Monster Library https://t.co/IXLWO1Be91” / X (twitter.com)
https://twitter.com/ProfKeeting/status/1801632319536607623
- proprietary proprietary OpenAI’s Sora model represents a notable advancement in AI video generation. It demonstrates the ability to generate videos up to one minute in 1080p resolution and produce high-resolution images. Sora’s flexibility in handling various aspect ratios and resolutions indicates its adaptability in content creation. Its development leverages insights from prior research, including Vision Transformers and advanced training methodologies. Introduction to Sora A groundbreaking AI video generation model by OpenAI, Sora is designed to transform text instructions into realistic and imaginative video scenes, marking a significant advancement in creative AI technologies. Technical Overview Advanced Diffusion Model Employs a sophisticated diffusion process that starts from static noise and incrementally refines to generate high-resolution videos, showcasing an unparalleled leap in video realism and complexity. Transformer Architecture Leverages the Transformer model’s capabilities for deep understanding and generation of content, adapted here to interpret and create complex visual narratives, ensuring dynamic and coherent video storytelling. twitter link to the render loading below https://twitter.com/sainingxie/status/1758433676105310543 twitter link to the render loading below https://twitter.com/thatguybg/status/1759935959792312461 Patch-Based Data Representation Innovatively represents videos and images as collections of smaller data units, akin to language model tokens, enabling precise and granular control over video generation and editing. https://twitter.com/drjimfan/status/1758355737066299692?s=46 Creative and Professional Applications Opens up endless possibilities for filmmakers, advertisers, educators, and content creators to produce cinema-quality visuals, educational materials, and immersive experiences effortlessly. Democratization of Video Production Simplifies the video creation process, enabling individuals and small teams to produce content that rivals big studio outputs. Enhancement of Creative Expression Allows creators to bring intricate visions and stories to life through simple text prompts, expanding visual storytelling horizons. Technical Insights Designed to scale language model capabilities to visual data, converting videos into patches for efficient processing and diverse video/image handling. Features a video compression network for temporal and spatial video compression, operating within a latent space. Uses a diffusion transformer architecture, effectively scaling video generation and improving sample quality with increased compute. Innovative Features Works with videos at native sizes to offer sampling flexibility and improve composition and framing. Leverages descriptive captioning technique, enhancing video fidelity and quality from text prompts. Can animate still images and extend videos, including seamless interpolation between two videos, showcasing versatility. Emerging Capabilities Exhibits capabilities like 3D consistency, long-range coherence, object permanence, and world interaction simulation. Suggests potential as a tool for simulating physical and digital environments, aiding in the development of capable simulators. Videos can serve as a basis for constructing detailed 3D scenes using techniques like Neural Radiance Fields (NeRFs), potentially revolutionizing 3D content creation and interaction. Rapid prototyping and realization of 3D environments and narratives enhance VR and AR immersion and interactivity. Enables generation of characters, objects, and worlds through text and voice prompts, making 3D content creation more intuitive and accessible. Already being used to create 360 spherical video. … — working/pages/Proprietary AI Video.md
- Google DeepMind on X: “Introducing Veo: our most capable generative video model. 🎥 It can create high-quality, 1080p clips that can go beyond 60 seconds. From photorealism to surrealism and animation, it can tackle a range of cinematic styles. 🧵 GoogleIO https://t.co/6zEuYRAHpH” / X (twitter.com) Google DeepMind on X: “Introducing Veo: our most capable generative video model. 🎥 It can create high-quality, 1080p clips that can go beyond 60 seconds. From photorealism to surrealism and animation, it can tackle a range of cinematic styles. 🧵 GoogleIO https://t.co/6zEuYRAHpH” / X (twitter.com) https://twitter.com/GoogleDeepMind/status/1790435824598716704 — working/pages/Proprietary AI Video.md
The Rest
- Lumiere: Google’s Contribution Google’s Lumiere project also signifies progress in video generation capabilities, though full details remain undisclosed. This suggests ongoing competition and development in the field.
- Meta’s Approach: Foundational World Modeling Meta (formerly Facebook) is taking a distinct approach, focusing on the underlying world modeling needed for video encoding and generation. This emphasis on understanding the principles of physics and object interactions could contribute to more realistic AI-generated videos.
- Technical Capabilities and Limitations
- Capabilities Current AI video generators demonstrate proficiency in producing high-resolution images and videos. They are capable of style adaptation, simulating complex scenes with multiple elements, and handling variations in aspect ratio and resolution.
- Limitations Despite their strengths, these models still struggle to accurately simulate physics and lack a complete understanding of cause and effect. Occasional errors regarding object permanence highlight the existing gap between pattern recognition and a comprehensive understanding of the world.
- Ethical and Creative Considerations
- Potential Impacts Advancements in AI video generation raise questions about the future of creative professions and the ethical implications of AI-generated content. Balancing technological innovation with safeguarding the integrity of human creativity is an important consideration.
- Challenges Distinguishing between pattern recognition and genuine understanding is pivotal in the ethical use of AI. The potential for misuse or the creation of harmful content underscores the need for clear guidelines and responsible practices.
Open systems
OpenSora
Stable Video
Pika
Misc links being integrated.
- MotionDirector, with a dual-path LoRAs architecture to decouple the learning of appearance and motion. Further, we design a novel appearance-debiased temporal loss to mitigate the influence of appearance on the temporal training objective. Experimental results show the proposed method can generate videos of diverse appearances for the customized motions. Our method also supports various downstream applications, such as the mixing of different videos with their appearance and motion respectively, and animating a single image with customized motions.
- RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models (rave-video.github.io)
-
https://discord.com/channels/1076117621407223829/1192162917395730635/1192162917395730635
-
Here’s one way to use the brand new RAVE node from here: https://github.com/spacepxl/ComfyUI-RAVE
- First pass often has flickering (depending a lot on the input), so I made a workflow to smooth even harsh flickering with AD. This allows for utilizing the transformative and often more detailed vid2vid from RAVE and still get smooth results in ComfyUI
- Updated LCM version: https://discord.com/channels/1076117621407223829/1192162917395730635/1192212692354748427 using the “video/controlgif/animatediff” contolnet from here: https://huggingface.co/crishhh/animatediff_controlnet/blob/main/controlnet_checkpoint.ckpt
- First pass often has flickering (depending a lot on the input), so I made a workflow to smooth even harsh flickering with AD. This allows for utilizing the transformative and often more detailed vid2vid from RAVE and still get smooth results in ComfyUI
-
Style transfer for humans
- Multiple techniques tested with the same LoRA DoRA etc for comparison
- ActAnywhere
- AI-Enhanced Creator (beehiiv.com)
- AnimateAnyone for ComfyUI MrForExample/ComfyUI-AnimateAnyone-Evolved: Improved AnimateAnyone implementation that allows you to use the opse image sequence and reference image to generate stylized video (github.com)
- [CG Renders to AI ANIMATION
- NIKE video — MOONWALKERS PICTURE](https://www.moonwalkerspicture.com/newslounge/cg-renders-to-ai-workflow-vol-02-anim)
- Motion Control
- [2401.12945] Lumiere: A Space-Time Diffusion Model for Video Generation (arxiv.org)
- [I2VGen-XL
- a Hugging Face Space by damo-vilab](https://huggingface.co/spaces/damo-vilab/I2VGen-XL)
- ali-vilab/i2vgen-xl: Official repo for VGen: a holistic video generation ecosystem for video generation building on diffusion models (github.com)
- Interpolation and interframe consistency
- controlnet and ebsynth temporal consistency
- Motion-Conditioned Diffusion Model for Controllable Video Synthesis
- Interframe consistency is now here
- Interpolation between two frames
- FILM frame interpolator
- ProPainter for Video Inpainting (shangchenzhou.com)
- zengyh1900/Awesome-Image-Inpainting: A curated list of image inpainting and video inpainting papers and resources (github.com)
- Runway AI video editing
- Gen2 examples
- Multishot VideoDrafter: Content-Consistent Multi-Scene Video Generation with LLM
- vienna with prompts
- Video slowmo and enhance
- deforum stable diffusion video
- Phenaki
- Collaborative video pipeline
- Magicvideo (faster)
- Production ready re aging
- distilled models for 25fps
- Stable warpfusion
- Video talking heads from text service
- Tune a video
- Vidyo: Generates videos for social networks from longer videos.
- Stylegan-T video transformer from google
- Houdini
- Dream Mix video to video remix
- RIFE frame interpolation
- example github for sd
- Synthesia corporate video generation
- pix2pixHD nextframe google colab
- minecraft demo codebase
- animation from mixamo
- Intel enhance photorealism in realtime
- custom SD video to video script
- Testing a custom video2video script I’m working on. (These used RealisticVision1.4 & ControlNet) : r/StableDiffusion
- consistency tools for character tooning
- Alibaba system
- 9 new tools
- Automatic1111 plugin
- Next frame prediction with controlnet
- Will smith eating spaghetti
- Transform Video to Animation in Stable Diffusion | How to Install + BEST Consistency Settings: Learn how to use AI to create animations from real videos. We’ll use Stable Diffusion and other tools for maximum consistencyProject Files:https://bit.ly/3…
- How to Use ModelScope text2video with Automatic1111’s Stable Diffusion Web UI | kombitz: Enable the Extension Click on the Extension tab and then click on Install from URL. Enter https://github.com/deforum-art/sd-webui-modelscope-text2video in the URL box and click on Install. Click on Installed and click on Apply and restart UI. Go to your stable-diffusion-webui/models folder and create a folder called ModelScope and then create a folder called t2v under ModelScope. This is your models folder for text2video.
- This article provides instructions on how to use ModelScope’s text2video feature with Automatic1111’s Stable Diffusion Web UI.
- latent consistency pipeline
- [GitHub
- Picsart-AI-Research/Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators: Text-to-Image Diffusion Models are Zero-Shot Video Generators
- GitHub
- Picsart-AI-Research/Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators](https://github.com/Picsart-AI-Research/Text2Video-Zero)
- The Picsart-AI-Research/Text2Video-Zero repository contains code for a text-to-image diffusion model that can be used to generate videos from text input. The model is a zero-shot video generator, meaning that it does not require any training data in order to generate videos.
- LVDM for long video creation
- The Text2Room algorithm generates textured 3D meshes from a given text prompt by leveraging pre-trained 2D text-to-image models. The core idea is to select camera poses that will result in a seamless, textured 3D mesh. The algorithm iteratively fuses scene frames with the existing geometry to create the final mesh. Evaluation shows that the algorithm is able to generate room-scale 3D geometry with compelling textures from only text as input.
- The VMesh system models a scene with a triangular mesh and a sparse volume for efficient view synthesis. It is trained on multi-view images of an object to create a contiguous representation of the object’s surface and volume. This representation is then used to generate a simplified triangular mesh and a sparse volume, which can be stored and rendered efficiently. The system is designed for real-time applications and can render at 2K 60FPS on common consumer devices.
- LLM guided video generation paper
- LVM video gen using LLM paper
- Temporal stable automatic plugin
- We present a method for high-resolution video synthesis using latent diffusion models (LDMs). Our approach first pre-trains an LDM on images, then introduces a temporal dimension to the latent space diffusion model and fine-tunes it on encoded image sequences (i.e. videos). We focus on two real-world applications: simulation of in-the-wild driving data and creative content creation with text-to-video modeling. Our method achieves state-of-the-art performance on real driving videos of 512 x 1024 resolution. Additionally, our approach can leverage off-the-shelf pre-trained image LDMs, turning the publicly available, state-of-the-art text-to-image LDM Stable Diffusion into an efficient and expressive text-to-video model.
- This script allows for the automation of video stylization using StableDiffusion and ControlNet.
- Really easy videos in A1111
- Dancer 4 keyframes, low noise, controlnet approach
- Flicker free video workflow paper (good!)
- Pika labs
- Realtime lip-sync API
- ms image to video on huggingface
- model to video blender modules
- videocomposer in python 3.9
- motionagent image to video
- Animatediff comfy workflows on discord
- fluid animation youtube
- Controlnet tutorial
- LCM loras for fast inferencing
- Animatediff is a new animation software that provides a range of tools and features for creating high-quality animations. It offers a user-friendly interface and supports various animation techniques, such as 2D, 3D, stop motion, and more. With Animatediff, users can easily bring their ideas to life and express their creativity through unique and captivating animations. Whether you’re a professional animator or a beginner, Animatediff offers a comprehensive set of features to help you create stunning animations in a fast and efficient manner. title:: Animatediff and Stablevideo
- Youtube tutorials
- IF_Animator ComfyUI workflow LCM+Animatediff+IPA+CN (youtube.com)
- [[Part 2] Tips and Tricks
- AnimateDiff ControlNet Animation in ComfyUI
- YouTube](https://www.youtube.com/watch?v=aysg2vFFO9g)
- TianxingWu/FreeInit: FreeInit: Bridging Initialization Gap in Video Diffusion Models (github.com)
- CiaraStrawberry/svd-temporal-controlnet (github.com)
- ProjectNUWA/DragNUWA (github.com)
AnimateDiff
- (1461) Discord | ad_resources | banodoco animatediff resources