OpenAI ChatGPT-4o (omni)

https://twitter.com/minchoi/status/1790396782200987662

Google DeepMind Gemini

  • Gemini is a multimodal LLM capable of inputting and outputting text, understanding images, and generating images.
  • While specific architecture details are scarce, it represents a leap in LLMs interacting with multiple data types.

Multi-Modal Large Language Models (LLMs)

  • Introduction:
    • Large language models are adept at generating coherent text sequences, predicting word probabilities and co-occurrences.
    • Multimodal models extend LLMs capabilities to not just output text, but images and understand multimodal inputs.
  • Core Concepts:
    • LLMs for Text:
      • LLMs process prompts and generate replies one token at a time, acting as a multiclass classifier.
    • Image Generation:
      • Traditional pixel-by-pixel image generation is intractable; hence, a different approach is needed.
      • The solution is treating image generation as a language generation problem, akin to ancient hieroglyphics.
  • Techniques in Multi-Modal LLMs:
    • Autoencoders:
      • Compress images into a lower-dimensional latent space and then regenerate them, learning crucial properties.
    • Variational Autoencoders (VAE) & VQ-VAE:
      • VAEs add a generative aspect by allowing for new image generation from random latent embeddings.
      • VQ-VAE further discretizes this process, creating a vocabulary of image “words” or tokens.
  • Implementation:
    • Vector Quantization:
      • Creates a discrete set of embedding vectors forming the vocabulary for our image-based language.
    • Encoding and Decoding:
      • Images are encoded to these discrete codes and decoded back to form new or reconstructed images.
  • Training and Inference:
    • A mixed sequence of embeddings (words and image tokens) is created for training.
    • The model learns to generate image tokens, forming a coherent sequence with the text, allowing for the generation of images corresponding to text descriptions.
  • Challenges and Developments:
    • The importance of quality data over quantity, especially for large, complex models.
    • Ongoing efforts focus on refining data quality, applying safety measures, and improving model transparency.
flowchart LR
A[Text Input] -->|Processed by LLM| B[Text Tokens]
B -->|Alongside Image Tokens| D[Mixed Embeddings]
C[Image Input] -->|Encoded via VQ-VAE| E[Image Tokens]
E --> D
D -->|Next Token Prediction| F[Generated Sequence]
F -->|Decoded| G[Output Image & Text]