Image Generation is the synthesis of realistic or stylised images using generative AI models including Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and Diffusion Models. Modern systems such as DALL-E, Stable Diffusion, Midjourney, and Flux produce high-fidelity images from text descriptions, sketches, or latent representations, enabling creative applications, data augmentation, virtual production, and content creation at scale.

Semantic Classification

Content

  • Image Generation is the synthesis of realistic or stylised images using generative AI models including Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and Diffusion Models. Modern image generation systems (DALL-E, Stable Diffusion, Midjourney) produce high-fidelity images from text descriptions, sketches, or latent representations, enabling creative applications, data augmentation, and content creation.

Collaborative Control for Geometry-Conditioned PBR Image Generation - * Holo-Gen is a research project exploring methods for generating 3D holographic content, especially for mixed reality applications.

  • The project aims to develop tools and techniques that simplify the process of creating holograms, making it more accessible to a wider range of users.
  • One focus is on using neural networks and machine learning to automatically generate holographic representations from 2D images or videos.
  • Another aspect involves creating interactive holographic experiences that allow users to manipulate and interact with virtual objects in a mixed reality environment.
  • The research investigates optimization of holographic displays for improved image quality, brightness, and field of view.
  • Holo-Gen seeks to address challenges such as computational complexity and data management requirements associated with holographic rendering.
  • The project explores different holographic display technologies, including spatial light modulators (SLMs) and computational holography.
  • Colour reproduction in holograms is also a key area of investigation, with research aiming to improve the accuracy and vibrancy of colours.
  • The research encompasses the development of algorithms for efficiently calculating and rendering holograms in real-time.
  • Ultimately, Holo-Gen strives to enable more intuitive and immersive mixed reality experiences through advanced research.

BlenderGPT

  • BlenderGPT GitHub - - BlenderGPT is a project that allows users to control Blender through natural language processing instructions using artificial intelligence models.
  • The tool aims to streamline the 3D modelling process by automation of repetitive tasks and enabling users to create and manipulate objects with simple text commands.
  • The project provides a framework for connecting Blender’s Python API with machine learning, enabling users to translate natural language processing into executable Blender code.
  • Functionality includes object creation, scene organisation, material application (including colour changes), and animation control, all via text prompts.
  • Users can install BlenderGPT as a Blender add-on and configure it with their API key to access the language model’s capabilities.
  • The project is designed to be extensible, allowing developers to add custom functions and improve the integration between natural language processing and Blender actions.
  • The repository offers examples, documentation and a troubleshooting guide to help users get started and resolve common issues.
  • The system allows for iterative design changes, where users can refine their creations by giving further instructions to the machine learning model based on previous results.
  • It streamlines the workflow by eliminating the need to switch between Blender and separate Stable Diffusion interfaces.
  • The add-on provides features for controlling image generation parameters, such as prompts, sampling methods, and image size.
  • Users can use generated images as textures, backgrounds, or reference images within their Blender projects.
  • It enables the creation of new and imaginative assets and content without extensive traditional modelling or texturing.
  • The system requires a local installation of Stable Diffusion and necessary dependencies, configured to work with the Blender add-on.
  • The add-on is designed to be customisable, enabling users to fine-tune image generation based on specific project needs.
  • The installation and use process is organised through a user-friendly interface within Blender.
  • It offers a way to enhance the creative possibilities within Blender by leveraging the power of artificial intelligence image generation.

GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting - //huggingface.co/papers/2402.07207) in UK English spelling, presented as bullet points:

  • The paper introduces a new method for improving the colourisation of greyscale images using diffusion models.
  • It addresses the problem of colour ambiguity in greyscale images by incorporating semantic information and user guidance.
  • The approach uses a diffusion model conditioned on both the greyscale image and semantic segmentation maps, allowing for more accurate and consistent colour assignments.
  • A user interface is provided, enabling users to interactively influence the colourisation process through colour hints or strokes.
  • The model can be used to organise and visualise large collections of greyscale images, by applying consistent colourisation styles across the dataset.
  • The framework achieves state-of-the-art performance compared to existing greyscale image colourisation techniques.
  • The user controlled component allows for finer control over colour choices compared to automatic systems.
  • The colourisation process is designed to be flexible and adaptable to different types of images and user preferences.

Point·E System

  • Point·E GitHub - - Point-E is a system developed by OpenAI for efficiently creating 3D point clouds from text prompts.
  • It offers a fast and direct method for 3D object generation, bypassing the slower and more complex process of first creating a mesh and then rendering.
  • The system utilises a series of models: a text-to-image model followed by an image-to-3D point cloud model.
  • It provides code for training and sampling these models, allowing users to experiment with custom datasets and text prompts.
  • The code includes utilities for visualising and manipulating the generated point clouds, including features for altering their colour and density.
  • A significant advantage of Point-E is its speed; it can produce 3D models significantly faster than previous approaches.
  • The repository provides pre-trained models, enabling immediate use without the need for extensive training on the user’s part.
  • The technology allows for easy integration into existing 3D pipelines and applications.
  • The project encourages further research into improving the quality and complexity of generated 3D assets.
  • The documentation helps users to organise the code and understand the underlying techniques for text-to-3D generation.
  • A search engine that uses AI to search for images and videos.

A website for my company (free hosting auto push to github pages)

image.png

Most Adopted Enterprise AI Use Cases

  • Code generation: 51%
  • Customer support chatbots: 31%
  • Enterprise search: 28%
  • Retrieval and data extraction: 27-28%
  • Meeting summarisation: 24%
  • Copywriting: 21%
  • Image generation: 20%
  • Use cases reflect a shift from consumer-focused tasks to enterprise-specific applications.

Novel VP Render Pipeline:

  • Putting the ML image generation on the end of a real-time tracked camera render pipeline might remove the need for detail in set building. The set designer, DP, director, etc., will be able to ideate in a headset-based metaverse of the set design, dropping very basic elements. If the interframe consistency (img2img) can deliver, the output on the VP screen can simply inherit the artistic style from the text prompts and render production quality from the basic building blocks. This “next level pre-vis” is being trailed in the Vircadia collaborative environment described in this book.

Retrieval Augmented Generation (RAG)

GitHub Repositories

Specialised Models

image.png

3D models for AR and VR

image.png

image.png

Section 4: The Multimodality War

  • Midjourney launched v6 and a web UI.
  • Assembly AI raised $50m for “Stripe for AI models”.
  • Replicate raised $40m to serve AI engineers.
  • Suno AI launched for audio generation.
  • OpenAI and Google continue work on “God Models”.

Image Generation & Editing

  • Task: Create unique images, enhance product photos, generate backgrounds, or design visual assets for marketing and branding.
  • Midjourney
    • Description: High-quality AI image generator accessible via web and Discord. Known for artistic and detailed outputs. As of June 2026, V8.1 is the default model, offering 2K HD image generation and image-to-video generation up to 21 seconds. Earlier notable versions include V6 (text generation within images, Vary Region inpainting) and V7 (April 2025, improved prompt adherence and texture quality).
    • Cost: Subscription-based, starting around $10 USD/month (Basic plan with limited generations).
    • Website: Midjourney
  • Dall-E 3 (via ChatGPT/Copilot)
    • Description: AI image generator accessible through ChatGPT Plus or Microsoft Copilot. Good for generating diverse images from text prompts, including logo concepts and backgrounds.
    • Cost: Included with ChatGPT Plus ($20 USD/month) or free/paid tiers of Microsoft Copilot.
    • Website: (Accessed via ChatGPT or Copilot)
  • Adobe Firefly
    • Description: Adobe’s suite of generative AI tools, integrated into Photoshop/Express. Features include text-to-image generation, generative fill, and text effects. Trained on Adobe Stock (commercially safe).
    • Cost: Included in many Adobe Creative Cloud subscriptions. Free plan with monthly credits available.
    • Website: Adobe Firefly
  • Flair.ai
    • Description: AI design tool for branded content, particularly product photoshoots. Removes backgrounds, places products in new scenes using templates or drag-and-drop interface.
    • Cost: Free plan available. Paid plans based on usage/features.
    • Website: Flair.ai
  • Magnific AI
    • Description: AI tool specialising in upscaling and enhancing image details, making them look higher resolution and more professional. Often used with Midjourney outputs.
    • Cost: Subscription-based, plans start around $39 USD/month.
    • Website: Magnific AI
  • Patterned AI
    • Description: AI tool specifically for generating unique, seamless patterns and backgrounds, useful for product backdrops or branding elements. Often used with Canva.
    • Cost: Check website for pricing (likely subscription or credits).
    • Website: Patterned AI
  • Photoroom
    • Description: Photo editing app/tool with strong AI features for background removal and creation, particularly for product photography. Can generate studio-quality backgrounds.
    • Cost: Free plan available. Pro plans unlock more features, often around £10-£15 GBP/month.
    • Website: Photoroom
  • ML Blocks
    • Description: No-code platform for building automated image processing workflows using AI. Useful for repetitive tasks without needing to write code.
    • Cost: Check website for pricing structure.
    • Website: ML Blocks

Google sub second inferencing on a phone.

Consumer Hardware

Avatar Generation: Creating Digital Beings from Scratch

  • This section focuses on platforms and research enabling the generation of complete avatars, encompassing both visual representation and underlying technologies.
    • REPLIKANT: An AI-assisted 3D avatar and animation platform designed for creators.
    • Meta Research Paper: A research paper from Meta exploring an unspecified aspect of avatar generation.
    • Heygen: A platform for generating and animating realistic avatars from text prompts and images.
    • Synthesia: A leading platform for creating AI-powered videos featuring realistic avatars.

Oobabooga

  • Strengths:

    • Broad feature set including image generation and voice capabilities.
    • Stable for solo usage.
  • Limitations: Slower performance compared to newer backends like TabbyAPI or vLLM.

  • Link: Oobabooga GitHub


Multimodal Capabilities

  • Open WebUI: Can integrate vision models, image generation, TTS, and more with third-party tools like Azure or Together.ai.
  • Koboldcpp: Supports some multimodal backends but lacks native syntax highlighting.

Other Notable Research

Image, Video and 3D

Text-to-Image Generation

  • Stable Diffusion generates realistic and imaginative images from descriptive text prompts. This core functionality allows users to translate their creative visions into visual form with remarkable accuracy and detail. Whether it’s a photorealistic portrait, a surreal landscape, or an abstract concept, Stable Diffusion can bring your ideas to life with just a few words.

  • A lot of the products you see on the market are either wrappers for the big AI companies, or else leveraging Stability models on rented cloud compute.

    ComfyUI_temp_exgja_00013_.png|800

Community Support

Images

Vision Mamba

  • Majority of Mamba papers (over 60%) address vision/image processing, especially biomedical image segmentation
  • Key themes:
    • Representing data as sequences is crucial
    • Images are not inherently sequential like language, music, or DNA
    • Multi-scan approaches enable handling non-sequential data
    • Hybrid architectures leverage strengths of different models
  • Open questions and challenges:
    • Scaling to larger models and datasets
    • Developing state regularization methods
    • Integrating Mamba with other architectural advances (e.g., memory tokens)
  • Potential for transformative impact, especially in biology and vision applications
  • Approaches:
    • U-Mamba (U-Mamba): Hybrid CNN-SSM architecture outperforming CNN and Transformers in biomedical image segmentation
    • Swin-UMamba (Swin-UMamba): Combines Mamba with ImageNet pre-training, outperforms U-Mamba
    • Vision Mamba: Bidirectional (forward and backward) scanning for learning visual representations
    • VMamba: Four-way “cross-scan” starting at each corner, combining representations
    • VM-UNet (VM-UNet): Applies VMamba’s four-way scan to medical image segmentation
    • Mamba-ND: Multi-dimensional sequencing for video and weather data, using sequential SSMs
    • SegMamba: 3D image segmentation with three-way scan, handling long sequences
    • Vivim: Three-way scan for video
    • MambaMorph: Aligns two input images by generating deformation field
  • Key insights:
    • Turning images into sequences is crucial, can be done through multi-scan approaches
    • Combining scans sequentially may be more effective than parallel
    • Mamba enables memory-efficient processing of high-resolution images, promising for edge applications (e.g., robotics)

Artificial Intelligence

Walkthrough Animations or Flythrough Videos

  • Twinmotion (guide)
    • Real-time link via Datasmith; create MP4 or interactive 360 panoramas.
    • Supports AI-powered denoising and crowd/traffic generation.
  • Enscape for Vectorworks (blog)
    • Live rendering inside Vectorworks, with keyframe-based video export and VR standalone packages.
  • Experimental AI Video Tools
    • Runway ML Gen-2 or PromeAI for short motion clips from still frames (low resolution but quick concept demos).

Need or Challenge:

Identity

Need or Challenge:

  • The project addresses the reluctance in the film industry to adopt AI and ML technologies due to tight margins and complexity.
  • VisionFlow introduces “parallax plates as a service”, integrating robotics with ML-based video generation.
  • Key benefits include increased productivity in pre-visualization and improved collaboration.
  • Assessor Feedback: Positive recognition of the project’s potential to improve productivity in video content production. However, a closer association with a video production company could enhance the application’s relevance and impact.

GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting - //huggingface.co/papers/2402.07207) in UK English spelling, presented as bullet points:

  • The paper introduces a new method for improving the colourisation of greyscale images using diffusion models.
  • It addresses the problem of colour ambiguity in greyscale images by incorporating semantic information and user guidance.
  • The approach uses a diffusion model conditioned on both the greyscale image and semantic segmentation maps, allowing for more accurate and consistent colour assignments.
  • A user interface is provided, enabling users to interactively influence the colourisation process through colour hints or strokes.
  • The model can be used to organise and visualise large collections of greyscale images, by applying consistent colourisation styles across the dataset.
  • The framework achieves state-of-the-art performance compared to existing greyscale image colourisation techniques.
  • The user controlled component allows for finer control over colour choices compared to automatic systems.
  • The colourisation process is designed to be flexible and adaptable to different types of images and user preferences.

OnePose++ for Object Pose Estimation

  • OnePose++ Page - - OnePose++, an extension of the OnePose framework, is a streamlined solution for robust and scalable 6D object pose estimation from a single RGB image.
  • Imagine 3D - Luma Labs Imagine allows users to create realistic 3D models from text descriptions, streamlining the design workflow.
  • It offers an intuitive interface to easily generate, edit and visualise 3D assets.
  • Users can control the colour, texture, and shape of the generated 3D models using natural language processing.
  • The tool enables users to iterate quickly on design ideas by making adjustments to the text prompt and regenerating the model.
  • Imagine facilitates the creation of customised 3D models for various applications, including gaming, product visualisation and animation.
  • The platform encourages experimentation with different prompts to explore the creative potential of artificial intelligence-powered 3D generation.
  • This technology could be used for rapid prototyping, game development, and creation of virtual environments.
  • GET3D aims to democratise 3D content creation by simplifying the process and reducing reliance on expert 3D modellers.

Novel VP Render Pipeline:

  • Putting the ML image generation on the end of a real-time tracked camera render pipeline might remove the need for detail in set building. The set designer, DP, director, etc., will be able to ideate in a headset-based metaverse of the set design, dropping very basic elements. If the interframe consistency (img2img) can deliver, the output on the VP screen can simply inherit the artistic style from the text prompts and render production quality from the basic building blocks. This “next level pre-vis” is being trailed in the Vircadia collaborative environment described in this book.

3.2 Llama.cpp

  • Strengths:
    • Minimalist server UI with OpenAI-compatible API.
    • Excellent for developers due to fast updates.
  • Limitations: Limited feature set compared to Oobabooga and Open WebUI.
    • Open WebUI: Can integrate vision models, image generation, TTS, and more with third-party tools like Azure or Together.ai.
    • Koboldcpp: Supports some multimodal backends but lacks native syntax highlighting.

2025 State of Consumer AI Report by Menlo Ventures

  • 91% of AI users default to their preferred general tool for most tasks.
  • AI spending increased from 13.8 billion in 2024, a sixfold increase.
  • Code generation: 51%
    • While adoption has been swift, researchers caution that the long-term economic impact will depend on how deeply generative AI becomes integrated into daily work processes over time.
  • 61% of business leaders confirmed their organizations accelerated AI usage in 2024, with more than half planning to increase their AI budgets in 2025.
  • 72% of enterprises plan to increase their spending on generative AI in the next year, with nearly 40% of those indicating an investment exceeding $250,000 in the current calendar year.

Controlnet

Stable diffusion

  • is a company that specializes in developing advanced artificial intelligence models. They are known for their expertise in creating generative models, which are capable of producing high-quality and realistic outputs in various domains such as image synthesis, language generation, and music composition. Stable Diffusion’s cutting-edge research and innovative approaches have made significant contributions to the field of generative AI.
  • Vlads next SD
  • InvokeAI simple interface

Features

Multi-Modal Large Language Models (LLMs)

  • Introduction:
    • Large Language Models are adept at generating coherent text sequences, predicting word probabilities and co-occurrences.
    • Multimodal models extend LLMs capabilities to not just output text, but images and understand multimodal inputs.
  • Core Concepts:
    • LLMs for Text:
      • LLMs process prompts and generate replies one token at a time, acting as a multiclass classifier.
    • Image Generation:
      • Traditional pixel-by-pixel image generation is intractable; hence, a different approach is needed.
      • The solution is treating image generation as a language generation problem, akin to ancient hieroglyphics.
  • Techniques in Multi-Modal LLMs:
    • Autoencoders:
      • Compress images into a lower-dimensional latent space and then regenerate them, learning crucial properties.
    • Variational Autoencoders (VAE) & VQ-VAE:
      • VAEs add a generative aspect by allowing for new image generation from random latent embeddings.
      • VQ-VAE further discretizes this process, creating a vocabulary of image “words” or tokens.
  • Implementation:
    • Vector Quantization:
      • Creates a discrete set of embedding vectors forming the vocabulary for our image-based language.
    • Encoding and Decoding:
      • Images are encoded to these discrete codes and decoded back to form new or reconstructed images.
  • Training and Inference:
    • A mixed sequence of embeddings (words and image tokens) is created for training.
    • The model learns to generate image tokens, forming a coherent sequence with the text, allowing for the generation of images corresponding to text descriptions.
  • Challenges and Developments:
    • The importance of quality data over quantity, especially for large, complex models.
    • Ongoing efforts focus on refining data quality, applying safety measures, and improving model transparency.

flowchart LR A[Text Input] ⇒|Processed by LLM| B[Text Tokens] B ⇒|Alongside Image Tokens| D[Mixed Embeddings] C[Image Input] ⇒|Encoded via VQ-VAE| E[Image Tokens] E ⇒ D D ⇒|Next Token Prediction| F[Generated Sequence] F ⇒|Decoded| G[Output Image & Text]

Closed Source Image Generation: id:: 659a9229-ed15-4932-a207-eb2daa96786e

Constrained Multi-Modal Retrieval Augmented Generation

Additional Tools and Resources

  • Horde Image and LLM
  • A project integrating images with LLMs for enhanced content generation.
  • LobeHub
  • A technology-driven forum for AIGC, offering modern design components and tools.
  • Microsoft WizardLM 2

Abstract

Recent years have witnessed the strong power of large text-to-image diffusion models for the impressive generative capability to create high-fidelity images. But, it is very tricky to generate desired images using only text prompt as it often involves complex prompt engineering. An alternative to text prompt is image prompt, as the saying goes: “an image is worth a thousand words”. Although existing methods of direct fine-tuning from pretrained models are effective, they require large computing resources and are not compatible with other base models, text prompt, and structural controls. In this paper, we present IP-Adapter, an effective and lightweight adapter to achieve image prompt capability for the pretrained text-to-image diffusion models. The key design of our IP-Adapter is decoupled cross-attention mechanism that separates cross-attention layers for text features and image features. Despite the simplicity of our method, an IP-Adapter with only 22M parameters can achieve comparable or even better performance to a fine-tuned image prompt model. As we freeze the pretrained duffusion model, the proposed IP-Adapter can be generalized not only to other custom models fine-tuned from the same base model, but also to controllable generation using existing controllable tools. With the benefit of the decoupled cross-attention strategy, the image prompt can also work well with the text prompt to accomplish multimodal image generation.

Various image synthesis with our proposed IP-Adapter applied on the pretrained text-to-image diffusion model and additional structure controller.

[Paper]      [Code]      [BibTeX]

Future Directions and Reflections

  • 🔮 Anticipating Future Developments:
    • AI’s capabilities in automating content creation and administrative tasks suggest an imminent shift towards more personalised and efficient educational models.
    • Ongoing advancement of AI tools like GPT-4 and image generation technologies like Midjourney indicates a rapidly evolving educational technology landscape.
    • 🤖 AI as a Collaborative Partner: Emphasising AI’s role as a tool to augment, rather than replace, human educators is key to harnessing its benefits while maintaining the essential human elements of teaching.
    • 💭 Creative Considerations: AI can be an ally in overcoming creative blocks and fostering a culture of innovation and expression in educational settings.

Bots Proliferate

Custom Gen AI models in business

image.png|500

Runway Gen 3

  • Introducing Gen-3 Alpha: A New Frontier for Video Generation (runwayml.com)

    Core Characteristics

  • High-Fidelity Synthesis: Photo-realistic image creation

  • Conditional Generation: Text-to-image, image-to-image, sketch-to-image

  • Controllable Generation: Manipulation of style, content, attributes

  • Diverse Outputs: Stochastic sampling for variety

  • Large-Scale Models: Trained on billion-image datasets

    Relationships

  • Subclass: Computer Vision

  • Related: Generative Adversarial Network, Diffusion Model, Text-to-Image

  • Models: DALL-E, Stable Diffusion, Midjourney, Imagen

  • Applications: Art Generation, Data Augmentation, Design

    Key Literature

    1. Goodfellow, I., et al. (2014). “Generative adversarial nets.” NeurIPS, 2672-2680.

    2. Rombach, R., et al. (2022). “High-resolution image synthesis with latent diffusion models.” CVPR, 10684-10695.

    3. Ramesh, A., et al. (2022). “Hierarchical text-conditional image generation with CLIP latents.” arXiv:2204.06125.

    See Also

  • Generative Adversarial Network

  • Diffusion Model

  • Text-to-Image

Provenance