Text Generation is the NLP task of producing coherent, contextually appropriate natural language text using neural language models, including applications such as story generation, article writing, code generation, and creative content production. Modern text generation employs transformer-based language models with autoregressive or sequence-to-sequence architectures, controllable generation techniques, and prompt engineering to produce human-quality text across diverse domains and styles.

Semantic Classification

Content

  • Text Generation is the NLP task of producing coherent, contextually appropriate natural language text using neural language models, including applications such as story generation, article writing, code generation, and creative content production. Modern text generation employs transformer-based language models (GPT, T5, BLOOM) with autoregressive or sequence-to-sequence architectures, controllable generation techniques, and prompt engineering to produce human-quality text across diverse domains and styles.

Luma AI Genie - Luma Genie is a tool that allows users to create realistic 3D models from text or image prompts using neural networks technology.

  • Users can describe the desired 3D model with a text prompt, specifying details like shape, colour, and texture.

  • Alternatively, users can upload an image to guide the 3D model generation.

  • The generated 3D models can be downloaded in various formats for use in different applications.

  • Genie aims to simplify the process of 3D content creation, making it more accessible to users of all skill levels.

  • It appears to be still under development, with the website showcasing examples and possibilities.

  • The technology uses neural radiance fields to generate high-quality, realistic 3D models.

  • Users can organise and share their created models through the Luma platform.

ComfyTextures

  • ComfyTextures GitHub - - ComfyTextures is a collection of free, high-quality textures designed for use in 3D rendering and other creative projects.
  • The textures are organised into logical categories such as wood, metal, fabric, and stone, making it easier to find the desired material.
  • Each texture comes with various maps (diffuse, normal, roughness, specular, height) to facilitate realistic material creation in different rendering engines.
  • The textures are generally provided in a tileable format allowing for seamless repetition across surfaces.
  • The repository is actively maintained, with additions and updates being made regularly, enhancing the available resource base.
  • The textures can be downloaded and used for both commercial and non-commercial purposes under a specified licence.
  • The repository aims to provide a valuable resource for artists and developers seeking readily accessible and customisable textures.
  • Many textures include variations in colour and detail allowing for greater control over the final appearance.

Imagine 3D Software

  • Imagine 3D - Luma Labs Imagine allows users to create realistic 3D models from text descriptions, streamlining the design workflow.
  • It offers an intuitive interface to easily generate, edit and visualise 3D assets.
  • Users can control the colour, texture, and shape of the generated 3D models using natural language processing.
  • The tool enables users to iterate quickly on design ideas by making adjustments to the text prompt and regenerating the model.
  • Imagine facilitates the creation of customised 3D models for various applications, including gaming, product visualisation and animation.
  • It allows users to organise and manage generated models within a centralised workspace.
  • The platform encourages experimentation with different prompts to explore the creative potential of artificial intelligence-powered 3D generation.
  • Imagine simplifies 3D content creation, democratising access to 3D modelling for non-experts.

GET3D by Toronto AI Lab

  • GET3D GitHub - GET3D is a generative model that creates high-quality 3D shapes with explicit textures, removing the need for time-consuming 3D modelling.
  • It uses a novel texture generation method directly on 3D volumes, resulting in detailed and realistic surface colour.
  • The model is trained on unlabelled 2D images using a differentiable rendering pipeline, meaning no 3D supervision is required.
  • The generated 3D shapes are mesh-free, represented as neural radiance fields, making them easy to manipulate and edit.
  • GET3D allows users to efficiently generate diverse and high-quality 3D assets, accelerating content creation workflows.
  • The system offers control over object categories during generation, allowing users to specify the type of 3D shape produced.
  • The generated assets are compatible with standard rendering pipelines and can be readily integrated into existing 3D scenes.
  • The research demonstrates significant improvements in 3D shape quality and texture detail compared to previous generative models.
  • This technology could be used for rapid prototyping, game development, and creation of virtual environments.
  • GET3D aims to democratise 3D content creation by simplifying the process and reducing reliance on expert 3D modellers.

Dream Fields for Text-Guided 3D Object Generation

  • Dream Fields - * Dreamfields is a technique for visualising and organising thoughts and ideas, similar to mind mapping but with a focus on colours and spatial arrangement.
  • The system utilises a canvas (physical or digital) where ideas are represented as colourful nodes, allowing for visual categorisation and intuitive connections.
  • Unlike traditional mind maps, Dreamfields encourages free-form arrangement and avoids rigid hierarchical structures, promoting flexible thinking.
  • The colour coding allows for the association of different themes, priorities, or categories to individual ideas, aiding in overall organisation.
  • The spatial arrangement of nodes on the canvas can reflect relationships, importance, or chronological order of the ideas, providing another layer of meaning.
  • Dreamfields is presented as a tool for brainstorming, problem-solving, and project planning, enabling users to explore and connect concepts in a non-linear way.
  • The method is suggested to be particularly useful for visual learners and individuals who benefit from a more fluid and intuitive approach to organisation.
  • Dreamfields encourages ongoing reflection and iteration, allowing the canvas to evolve and adapt as thinking develops.

Research

  • Phi-3 and Phi-4: Powerful language models compact enough to run on a smartphone.
  • Amazon
  • Amazon is making a significant push in AI with its new “Nova” models, a revamped “Alexa Plus,” and next-generation AI chips.

Physically Based Textures from BIM (Revit)

Screenshot 2025-07-24 173949.png

Evolution from Chat to Complex Systems

  • Context engineering emerged as AI systems evolved beyond simple chat interfaces to incorporate:
    • Function calling and tool use
    • Retrieval augmented generation (RAG) systems
    • Multi-agent workflows
    • External API integrations
    • The principle of “garbage in, garbage out” becomes critical when managing complex information flows. Pre-processing and cleaning data before it enters the context window significantly improves output quality.

Text-to-Speech

  • Text-to-speech (TTS) technology can be used to convert written text into spoken audio. This can be used to create podcasts from blog posts, articles, or other written content.

Project Details

  • Technology and Process
    • Utilizing highly-efficient energy generation equipment, the project transforms methane, a natural landfill byproduct, into electricity.
    • This electricity is used for several on-site applications, notably for powering data centers.
  • Environmental and Economic Impacts
    • The initiative aims to reduce greenhouse gas emissions.
    • It also generates revenue, which supports Viridi’s investment in constructing a state-of-the-art Renewable Natural Gas (RNG) facility at the landfill.
    • The RNG facility, expected to be fully operational by the second half of 2024, will produce the equivalent of three million gallons of gasoline annually.

Stable Diffusion in Blender

  • A Blender addon for using Stable Diffusion to render texture bakes for objects.

Dream Textures

  • A Blender addon for applying textures with text prompts.

Prerequisites

  • Before beginning, ensure you have:
    • A modern web browser for testing your wallet.
    • A text editor or IDE (e.g., Visual Studio Code, Sublime Text) for writing and editing code.
    • Access to the internet to fetch resources and documentation from the Cashew GitHub repository.

New AI Model Releases

  • GPT4-x-Alpaca-13B-Native-4bit-128g: Technical discussions on the new model and its capabilities (GitHub Discussion).

Unsorted Links

Advice on AI coding

  • Choose Tools Strategically: Not all AI coding tools are created equal. Select the right tool for the job, considering the project’s scope and complexity:
  • Complex Applications: Cursor, Windsurf, or more established IDE integrations (see below) are often better suited for larger, more intricate projects.
  • Micro-SaaS: Bolt/Lovable are optimised for smaller, Software-as-a-Service applications.
  • Mobile Applications: Replit remains a good choice, alongside framework-specific tools.
  • UI Design: Consider using ‘vo’ or similar specialised tools for user interface design.
  • General Coding Assistance & IDE Integration:
    • GitHub Copilot: A widely used and powerful AI pair programmer that integrates directly into your IDE (VS Code, JetBrains IDEs, etc.).
    • GitHub Copilot Agents: Extend Copilot’s capabilities with specialised agents for tasks like code review, debugging, and test generation.
    • Aider: A command-line tool that helps you write and edit code using GPT models. Good for making changes to existing codebases, particularly for refactoring and adding features.
    • Roo: Provides code generation and chat capabilities within your IDE.
    • Cline: Good for command line interfacing, and code assistance.
  • Context is Paramount: Always provide comprehensive context about your project. AI tools cannot “guess” your intentions. Use Markdown (.md) documents to detail:
  • Product Requirements Document (PRD): Clearly outlines the purpose, features, and functionality of the application.
  • Technical Stack Document: Specifies the programming languages, frameworks, libraries, and databases to be used.
  • File Structure: Defines the organisation of directories and files within the project.
  • Frontend Guidelines: Describes coding standards, styling conventions, and component structure for the user interface.
  • Backend Structure: Outlines the architecture, API endpoints, data models, and business logic for the server-side code.
  • Use CodeGuide (or Similar): Consider using CodeGuide or a similar tool to help generate and manage these AI-specific coding documents. This ensures compatibility across various AI tools and helps maintain a single source of truth.
  • Incremental Development: Avoid overly broad prompts like “build me an AirBNB clone.” Instead, break down the project into manageable steps:
  • Page by Page: Develop the application one page at a time.
  • Component by Component: Within each page, build individual components sequentially.
  • Limited Task Execution: AI models typically perform best with a maximum of 3 concurrent tasks per request. Be mindful of this limitation, and break down larger tasks accordingly. Tools like Aider and Copilot Agents can help manage this complexity.
  • Select AI-Friendly Technologies: Certain technology stacks are better understood by current AI models:
  • Web Applications:
    • React (with NextJS or ViteJS): Provides excellent performance and is well-supported by AI tools.
    • Python (with frameworks like Django or Flask): Widely used and well-understood by AI models.
  • Mobile Applications:
    • React Native: A good choice for cross-platform development.
    • SwiftUI (especially with Claude): Works well, particularly with Claude models.
  • Avoid Older Technologies: Unless absolutely necessary, as AI model support may be limited.
  • Utilise Starter Kits: Save time and reduce token usage by starting with pre-built templates or boilerplates:
  • Example: The “CodeGuide NextJS Starter Kit” can provide a solid foundation.
  • Benefit: Accelerates workflow and provides a structured starting point. Most frameworks have readily available starter kits.
  • Define Rules Within Your Tools: Many AI coding tools allow project-specific rules:
  • Examples: .cursorrules (often “project rules”), .windsurfrules, or similar configuration files within your IDE or tool. Copilot and other IDE-integrated tools often have settings for coding style and preferences.
  • Purpose: Constrain the AI, preventing deviations from your guidelines and coding standards.
  • Coding Standards: Enforce coding standards using linters (e.g., ESLint for JavaScript, Pylint for Python) and integrate their configuration with your AI tools where possible.
  • Employ a Multi-Tool Approach: No single tool handles the entire workflow seamlessly. Combine tools:
  • Research: Perplexity.
  • Brainstorming: ChatGPT (voice features can be helpful).
  • Documentation: CodeGuide, or tools integrated within your IDE.
  • Data Scraping: Firecrawl, or libraries within your chosen language (e.g., Beautiful Soup in Python).
  • Code Generation/Assembly/Refactoring: Your chosen AI coding tool (Cursor, Windsurf, GitHub Copilot, Aider, Roo, Cline, etc.). Choose based on your workflow and project needs.
  • Patience and Persistence: Working with AI requires a specific mindset.
  • Prompt Engineering: Crafting effective prompts is crucial. Experiment with different phrasing and levels of detail.
  • Expect Errors: AI models are not perfect. Be prepared for errors.
  • Iterative Refinement: Stay focused, learn from mistakes, and iteratively refine your prompts and approach.
  • Debugging: Provide the AI with the full code and error message for assistance. Leverage Copilot Agents for debugging tasks.
  • Version Control
  • Use Git for version control.
  • Commit frequently with clear messages.
  • AI can help generate commit messages (Copilot, Aider, and others offer this).
  • Testing
  • Write unit and integration tests.
  • AI can assist in generating test cases (Copilot Agents are particularly useful here). Tools like Aider can help refactor code to improve testability.
  • Agent Frameworks

Random Links

https://twitter.com/tldraw/status/1782443204710674571

Audio & Voice Generation

  • Task: Create voiceovers, audio content like podcasts, or clone voices for various applications.
  • ElevenLabs
    • Description: High-quality text-to-speech and voice cloning AI. Offers a library of voices (Community Voices), can create a synthetic version of your own voice, and provide audio narration for websites (Audio Native). Used for audiobooks, podcasts, voiceovers.
    • Cost: Free tier available. Paid plans based on character usage, starting around $5 USD/month.
    • Website: ElevenLabs
  • PlayHT
    • Description: AI text-to-voice generator with a large library of voices and languages. Suitable for creating audiobooks, podcasts, and voiceovers.
    • Cost: Free plan available. Paid plans based on word count/features, starting around $30 USD/month (billed annually).
    • Website: PlayHT
VRChat
  • none of this makes money yet
  • This text is from wikipedia and will be updated when we have a chance totry VRChat properly. It’s much loved already by the Bitcoin community.
  • “VRChat’s gameplay is similar to that of games such as Second Life andHabbo Hotel. Players can create their own instanced worlds in which theycan interact with each other through virtual avatars. A softwaredevelopment kit for Unity released alongside the game gives players theability to create or import character models to be used in the platform,as well as build their own worlds.
  • Player models are capable of supporting “audio lip sync, eye trackingand blinking, and complete range of motion.
  • VRChat is also capable of running in “desktop mode” without a VRheadset, which is controlled using either a mouse and keyboard, or agamepad. Some content has limitations in desktop mode, such as theinability to freely move an avatar’s limbs, or perform interactions thatrequire more than one hand.
  • In 2020, a new visual programming language was introduced known as”Udon”, which uses a node graph system. While still considered alphasoftware, it became usable on publicly-accessible worlds beginning inApril 2020. A third-party compiler known as “UdonSharp” was developed toallow world scripts to be written in C sharp.”
Vircadia
  • The applications and platforms detailed above have their benefits, butfor the application stack in the next section of the book Vircadia hasbeen chosen. The following text is from their website, and is aplaceholder which gives some idea. This section will be written outcompletely to reflect our use of the product to support emerging market users.
  • Vircadia is open-source software which enables you to create and sharevirtual worlds as virtual reality (VR) and desktop experiences. You cancreate and host your own virtual world, explore other worlds, meet andconnect with other users, attend or host live VR events, and much more.
  • The Vircadia metaverse provides built-in social features, includingavatar interactions, spatialized audio, and interactive physics.Additionally, you have the ability to import any 3D object into yourvirtual environment. No matter where you go in Vircadia, you will alwaysbe able to interact with your environment, engage with your friends, andlisten to conversations just like you would in real life.

Foundation Models

  • Foundation models are large-scale, pre-trained models that can be adapted to a wide range of downstream tasks. They are trained on massive datasets of text and code and can be used for a variety of natural language processing (NLP) tasks, such as text generation, summarization, and question answering.

Animation: Breathing Life into Digital Characters

Logseq

  • Logseq: is very similar to Obsidian, but self hosted and open source. It works on top of plain text files stored in a local system. It supports markdown and Org-mode formatting and allows for hierarchical and networked note-taking. It can be connected to it’s mobile app via github.
  • Integration to Large Language Models can be OpenAI or local.
    • Compare notion, obsidian, and logseq, using a simply markdown table with coloured dots
  • ChatGPT Logseq Summarizer (openai.com)

Screenshot 2024-01-06 120253.png

Screenshot 2024-01-18 103043.png

Screenshot 2024-01-18 102807.png

Key LLM Papers

GPT-1 (2018): This paper introduces the first version of Generative Pre-trained Transformer (GPT), a generative model trained on a massive dataset of text. It demonstrates the ability of LLMs to generate coherent and grammatically correct text, paving the way for future advancements.

GPT-2 (2019): This paper presents a significantly larger GPT model with improved capabilities. It showcases the ability of LLMs to perform various language tasks, including text summarization, question answering, and even code generation.

GPT-3 (2020): This paper introduces GPT-3, a truly massive LLM with billions of parameters. It demonstrates impressive capabilities in diverse tasks, showcasing the emergence of general-purpose language abilities.

GPT-4 (2023): This paper introduces the latest iteration of GPT, featuring multi-modal capabilities and advanced reasoning abilities. It further pushes the boundaries of what LLMs can achieve, demonstrating impressive performance in a wide range of tasks.

Llama-2 (2023): This paper introduces Llama-2, a large language model designed with a focus on efficiency and accessibility. It offers a more resource-friendly alternative to other LLMs, making it more accessible for research and development.

Tools (2023): This paper introduces the “Tools” paradigm for LLMs, allowing them to interact with external tools and resources. It enables LLMs to perform more complex tasks by leveraging the power of external tools, expanding their capabilities significantly.

Gemini-Pro-1.5 (2023): This paper introduces Gemini-Pro-1.5, a large language model developed by Google. It showcases impressive capabilities in various tasks, including code generation, creative writing, and reasoning. It’s a strong contender in the race for developing advanced LLMs.

Generative Models for Molecule Design

  • Generative models based on diffusion and flow-matching approaches enable fine-grained control over the generation of molecules with specific properties. ProteinDT and MoleculeSTM are examples of text-conditioned generative models that allow users to provide natural language prompts to generate molecules with desired properties.
  • RF Diffusion, a diffusion model built on the RoseTTAFold backbone, offers powerful functionalities for protein engineering. It enables unconditional generation of novel proteins, binder design for high affinity and specificity, partial diffusion for refining existing structures, motif scaffolding for combining functional motifs, symmetric generation of protein complexes, and fold conditioning for generating proteins with specific tertiary structures.
  • Complementary models like Ligand and PNN (Protein MPNN) are essential for designing amino acid sequences that fold into the desired 3D structures generated by RF Diffusion.

Some software choices

  • It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.
  • proprietary
  • OpenAI’s Sora model represents a notable advancement in AI video generation. It demonstrates the ability to generate videos up to one minute in 1080p resolution and produce high-resolution images. Sora’s flexibility in handling various aspect ratios and resolutions indicates its adaptability in content creation. Its development leverages insights from prior research, including Vision Transformers and advanced training methodologies.

  • Introduction to Sora
    • A groundbreaking AI video generation model by OpenAI, Sora is designed to transform text instructions into realistic and imaginative video scenes, marking a significant advancement in creative AI technologies.
  • Technical Overview

https://twitter.com/drjimfan/status/1758355737066299692?s=46

  • Creative and Professional Applications
    • Opens up endless possibilities for filmmakers, advertisers, educators, and content creators to produce cinema-quality visuals, educational materials, and immersive experiences effortlessly.
  • Democratization of Video Production
    • Simplifies the video creation process, enabling individuals and small teams to produce content that rivals big studio outputs.
  • Enhancement of Creative Expression
    • Allows creators to bring intricate visions and stories to life through simple text prompts, expanding visual storytelling horizons.
  • Technical Insights
    • Designed to scale language model capabilities to visual data, converting videos into patches for efficient processing and diverse video/image handling.
    • Features a video compression network for temporal and spatial video compression, operating within a Neural Network Latent Space.
    • Uses a diffusion transformer architecture, effectively scaling video generation and improving sample quality with increased compute.
  • Innovative Features
    • Works with videos at native sizes to offer sampling flexibility and improve composition and framing.
    • Leverages descriptive captioning technique, enhancing video fidelity and quality from text prompts.
    • Can animate still images and extend videos, including seamless interpolation between two videos, showcasing versatility.
  • Emerging Capabilities
    • Exhibits capabilities like 3D consistency, long-range coherence, object permanence, and world interaction simulation.
    • Suggests potential as a tool for simulating physical and digital environments, aiding in the development of capable simulators.
    • Videos can serve as a basis for constructing detailed 3D scenes using techniques like Neural Radiance Fields (NeRFs), potentially revolutionizing 3D content creation and interaction.
    • Rapid prototyping and realization of 3D environments and narratives enhance VR and AR immersion and interactivity.
    • Enables generation of characters, objects, and worlds through text and voice prompts, making 3D content creation more intuitive and accessible.
    • Already being used to create 360 spherical video.
  • Research and Discussion
    • Video generation models as world simulators (openai.com) research paper highlights Sora’s technical foundation and its role in simulating the physical world.
    • Discussions emphasize Sora’s potential in democratizing video creation and the need for granular output control for artistic purposes.
  • Google DeepMind on X: “Introducing Veo: our most capable generative video model. 🎥 It can create high-quality, 1080p clips that can go beyond 60 seconds. From photorealism to surrealism and animation, it can tackle a range of cinematic styles. 🧵 GoogleIO https://t.co/6zEuYRAHpH” / X (twitter.com)

https://twitter.com/GoogleDeepMind/status/1790435824598716704

VideoPoet – Google Research

  • Overview: Google’s text to video, linked to Bard, but not yet available.

H200-NeMo-performance

Birme image resizer

Renderings from Plan Drawings

  • Vectorworks AI Visualizer (FAQ)
    • Works inside Vectorworks 2024+, using your active file or view plus a text prompt.
    • Ideal for quick concept iterations (materials, lighting variations).
    • Note: not CAD-accurate rendering but excellent for inspirational visuals.
  • Veras AI for Vectorworks (EvolveLAB announcement)
    • Plugin that uses your 3D model or 2D viewport as a base.
    • Photorealistic or stylised renders in seconds with prompt-driven material and ambience overrides.
  • Mainstream Text-to-Image Generators
    • Export plan or massing views as PNG/JPG and feed into Midjourney, Stable Diffusion (with ControlNet) or DALL·E 3 for high-res concept images.
    • Best for early-stage mood boards rather than precise layouts.

Luma AI Genie - Luma Genie is a tool that allows users to create realistic 3D models from text or image prompts using neural networks technology.

  • Users can describe the desired 3D model with a text prompt, specifying details like shape, colour, and texture.

  • Alternatively, users can upload an image to guide the 3D model generation.

  • The generated 3D models can be downloaded in various formats for use in different applications.

  • Genie aims to simplify the process of 3D content creation, making it more accessible to users of all skill levels.

  • It appears to be still under development, with the website showcasing examples and possibilities.

  • The technology uses neural radiance fields to generate high-quality, realistic 3D models.

  • Users can organise and share their created models through the Luma platform.

ComfyTextures

  • ComfyTextures GitHub - - ComfyTextures is a collection of free, high-quality textures designed for use in 3D rendering and other creative projects.
  • The textures are organised into logical categories such as wood, metal, fabric, and stone, making it easier to find the desired material.
  • Each texture comes with various maps (diffuse, normal, roughness, specular, height) to facilitate realistic material creation in different rendering engines.
  • The textures are generally provided in a tileable format allowing for seamless repetition across surfaces.
  • The repository is actively maintained, with additions and updates being made regularly, enhancing the available resource base.
  • The textures can be downloaded and used for both commercial and non-commercial purposes under a specified licence.
  • The repository aims to provide a valuable resource for artists and developers seeking readily accessible and customisable textures.
  • Many textures include variations in colour and detail allowing for greater control over the final appearance.

Imagine 3D Software

  • Imagine 3D - Luma Labs Imagine allows users to create realistic 3D models from text descriptions, streamlining the design workflow.
  • It offers an intuitive interface to easily generate, edit and visualise 3D assets.
  • Users can control the colour, texture, and shape of the generated 3D models using natural language processing.
  • The tool enables users to iterate quickly on design ideas by making adjustments to the text prompt and regenerating the model.
  • Imagine facilitates the creation of customised 3D models for various applications, including gaming, product visualisation and animation.
  • It allows users to organise and manage generated models within a centralised workspace.
  • The platform encourages experimentation with different prompts to explore the creative potential of artificial intelligence-powered 3D generation.
  • Imagine simplifies 3D content creation, democratising access to 3D modelling for non-experts.

GET3D by Toronto AI Lab

  • GET3D GitHub - GET3D is a generative model that creates high-quality 3D shapes with explicit textures, removing the need for time-consuming 3D modelling.
  • It uses a novel texture generation method directly on 3D volumes, resulting in detailed and realistic surface colour.
  • The model is trained on unlabelled 2D images using a differentiable rendering pipeline, meaning no 3D supervision is required.
  • The generated 3D shapes are mesh-free, represented as neural radiance fields, making them easy to manipulate and edit.
  • GET3D allows users to efficiently generate diverse and high-quality 3D assets, accelerating content creation workflows.
  • The system offers control over object categories during generation, allowing users to specify the type of 3D shape produced.
  • The generated assets are compatible with standard rendering pipelines and can be readily integrated into existing 3D scenes.
  • The research demonstrates significant improvements in 3D shape quality and texture detail compared to previous generative models.
  • This technology could be used for rapid prototyping, game development, and creation of virtual environments.
  • GET3D aims to democratise 3D content creation by simplifying the process and reducing reliance on expert 3D modellers.

Dream Fields for Text-Guided 3D Object Generation

  • Dream Fields - * Dreamfields is a technique for visualising and organising thoughts and ideas, similar to mind mapping but with a focus on colours and spatial arrangement.
  • The system utilises a canvas (physical or digital) where ideas are represented as colourful nodes, allowing for visual categorisation and intuitive connections.
  • Unlike traditional mind maps, Dreamfields encourages free-form arrangement and avoids rigid hierarchical structures, promoting flexible thinking.
  • The colour coding allows for the association of different themes, priorities, or categories to individual ideas, aiding in overall organisation.
  • The spatial arrangement of nodes on the canvas can reflect relationships, importance, or chronological order of the ideas, providing another layer of meaning.
  • Dreamfields is presented as a tool for brainstorming, problem-solving, and project planning, enabling users to explore and connect concepts in a non-linear way.
  • The method is suggested to be particularly useful for visual learners and individuals who benefit from a more fluid and intuitive approach to organisation.
  • Dreamfields encourages ongoing reflection and iteration, allowing the canvas to evolve and adapt as thinking develops.

Research

  • Phi-3 and Phi-4: Powerful language models compact enough to run on a smartphone.
  • Amazon
  • Amazon is making a significant push in AI with its new “Nova” models, a revamped “Alexa Plus,” and next-generation AI chips.

Physically Based Textures from BIM (Revit)

Screenshot 2025-07-24 173949.png

Evolution from Chat to Complex Systems

  • Context engineering emerged as AI systems evolved beyond simple chat interfaces to incorporate:
    • Function calling and tool use
    • Retrieval augmented generation (RAG) systems
    • Multi-agent workflows
    • External API integrations
    • The principle of “garbage in, garbage out” becomes critical when managing complex information flows. Pre-processing and cleaning data before it enters the context window significantly improves output quality.

Text-to-Speech

  • Text-to-speech (TTS) technology can be used to convert written text into spoken audio. This can be used to create podcasts from blog posts, articles, or other written content.

Project Details

  • Technology and Process
    • Utilizing highly-efficient energy generation equipment, the project transforms methane, a natural landfill byproduct, into electricity.
    • This electricity is used for several on-site applications, notably for powering data centers.
  • Environmental and Economic Impacts
    • The initiative aims to reduce greenhouse gas emissions.
    • It also generates revenue, which supports Viridi’s investment in constructing a state-of-the-art Renewable Natural Gas (RNG) facility at the landfill.
    • The RNG facility, expected to be fully operational by the second half of 2024, will produce the equivalent of three million gallons of gasoline annually.

Stable Diffusion in Blender

  • A Blender addon for using Stable Diffusion to render texture bakes for objects.

Dream Textures

  • A Blender addon for applying textures with text prompts.

Prerequisites

  • Before beginning, ensure you have:
    • A modern web browser for testing your wallet.
    • A text editor or IDE (e.g., Visual Studio Code, Sublime Text) for writing and editing code.
    • Access to the internet to fetch resources and documentation from the Cashew GitHub repository.

New AI Model Releases

  • GPT4-x-Alpaca-13B-Native-4bit-128g: Technical discussions on the new model and its capabilities (GitHub Discussion).

Unsorted Links

Advice on AI coding

  • Choose Tools Strategically: Not all AI coding tools are created equal. Select the right tool for the job, considering the project’s scope and complexity:
  • Complex Applications: Cursor, Windsurf, or more established IDE integrations (see below) are often better suited for larger, more intricate projects.
  • Micro-SaaS: Bolt/Lovable are optimised for smaller, Software-as-a-Service applications.
  • Mobile Applications: Replit remains a good choice, alongside framework-specific tools.
  • UI Design: Consider using ‘vo’ or similar specialised tools for user interface design.
  • General Coding Assistance & IDE Integration:
    • GitHub Copilot: A widely used and powerful AI pair programmer that integrates directly into your IDE (VS Code, JetBrains IDEs, etc.).
    • GitHub Copilot Agents: Extend Copilot’s capabilities with specialised agents for tasks like code review, debugging, and test generation.
    • Aider: A command-line tool that helps you write and edit code using GPT models. Good for making changes to existing codebases, particularly for refactoring and adding features.
    • Roo: Provides code generation and chat capabilities within your IDE.
    • Cline: Good for command line interfacing, and code assistance.
  • Context is Paramount: Always provide comprehensive context about your project. AI tools cannot “guess” your intentions. Use Markdown (.md) documents to detail:
  • Product Requirements Document (PRD): Clearly outlines the purpose, features, and functionality of the application.
  • Technical Stack Document: Specifies the programming languages, frameworks, libraries, and databases to be used.
  • File Structure: Defines the organisation of directories and files within the project.
  • Frontend Guidelines: Describes coding standards, styling conventions, and component structure for the user interface.
  • Backend Structure: Outlines the architecture, API endpoints, data models, and business logic for the server-side code.
  • Use CodeGuide (or Similar): Consider using CodeGuide or a similar tool to help generate and manage these AI-specific coding documents. This ensures compatibility across various AI tools and helps maintain a single source of truth.
  • Incremental Development: Avoid overly broad prompts like “build me an AirBNB clone.” Instead, break down the project into manageable steps:
  • Page by Page: Develop the application one page at a time.
  • Component by Component: Within each page, build individual components sequentially.
  • Limited Task Execution: AI models typically perform best with a maximum of 3 concurrent tasks per request. Be mindful of this limitation, and break down larger tasks accordingly. Tools like Aider and Copilot Agents can help manage this complexity.
  • Select AI-Friendly Technologies: Certain technology stacks are better understood by current AI models:
  • Web Applications:
    • React (with NextJS or ViteJS): Provides excellent performance and is well-supported by AI tools.
    • Python (with frameworks like Django or Flask): Widely used and well-understood by AI models.
  • Mobile Applications:
    • React Native: A good choice for cross-platform development.
    • SwiftUI (especially with Claude): Works well, particularly with Claude models.
  • Avoid Older Technologies: Unless absolutely necessary, as AI model support may be limited.
  • Utilise Starter Kits: Save time and reduce token usage by starting with pre-built templates or boilerplates:
  • Example: The “CodeGuide NextJS Starter Kit” can provide a solid foundation.
  • Benefit: Accelerates workflow and provides a structured starting point. Most frameworks have readily available starter kits.
  • Define Rules Within Your Tools: Many AI coding tools allow project-specific rules:
  • Examples: .cursorrules (often “project rules”), .windsurfrules, or similar configuration files within your IDE or tool. Copilot and other IDE-integrated tools often have settings for coding style and preferences.
  • Purpose: Constrain the AI, preventing deviations from your guidelines and coding standards.
  • Coding Standards: Enforce coding standards using linters (e.g., ESLint for JavaScript, Pylint for Python) and integrate their configuration with your AI tools where possible.
  • Employ a Multi-Tool Approach: No single tool handles the entire workflow seamlessly. Combine tools:
  • Research: Perplexity.
  • Brainstorming: ChatGPT (voice features can be helpful).
  • Documentation: CodeGuide, or tools integrated within your IDE.
  • Data Scraping: Firecrawl, or libraries within your chosen language (e.g., Beautiful Soup in Python).
  • Code Generation/Assembly/Refactoring: Your chosen AI coding tool (Cursor, Windsurf, GitHub Copilot, Aider, Roo, Cline, etc.). Choose based on your workflow and project needs.
  • Patience and Persistence: Working with AI requires a specific mindset.
  • Prompt Engineering: Crafting effective prompts is crucial. Experiment with different phrasing and levels of detail.
  • Expect Errors: AI models are not perfect. Be prepared for errors.
  • Iterative Refinement: Stay focused, learn from mistakes, and iteratively refine your prompts and approach.
  • Debugging: Provide the AI with the full code and error message for assistance. Leverage Copilot Agents for debugging tasks.
  • Version Control
  • Use Git for version control.
  • Commit frequently with clear messages.
  • AI can help generate commit messages (Copilot, Aider, and others offer this).
  • Testing
  • Write unit and integration tests.
  • AI can assist in generating test cases (Copilot Agents are particularly useful here). Tools like Aider can help refactor code to improve testability.
  • Agent Frameworks

Random Links

https://twitter.com/tldraw/status/1782443204710674571

Audio & Voice Generation

  • Task: Create voiceovers, audio content like podcasts, or clone voices for various applications.
  • ElevenLabs
    • Description: High-quality text-to-speech and voice cloning AI. Offers a library of voices (Community Voices), can create a synthetic version of your own voice, and provide audio narration for websites (Audio Native). Used for audiobooks, podcasts, voiceovers.
    • Cost: Free tier available. Paid plans based on character usage, starting around $5 USD/month.
    • Website: ElevenLabs
  • PlayHT
    • Description: AI text-to-voice generator with a large library of voices and languages. Suitable for creating audiobooks, podcasts, and voiceovers.
    • Cost: Free plan available. Paid plans based on word count/features, starting around $30 USD/month (billed annually).
    • Website: PlayHT
VRChat
  • none of this makes money yet
  • This text is from wikipedia and will be updated when we have a chance totry VRChat properly. It’s much loved already by the Bitcoin community.
  • “VRChat’s gameplay is similar to that of games such as Second Life andHabbo Hotel. Players can create their own instanced worlds in which theycan interact with each other through virtual avatars. A softwaredevelopment kit for Unity released alongside the game gives players theability to create or import character models to be used in the platform,as well as build their own worlds.
  • Player models are capable of supporting “audio lip sync, eye trackingand blinking, and complete range of motion.
  • VRChat is also capable of running in “desktop mode” without a VRheadset, which is controlled using either a mouse and keyboard, or agamepad. Some content has limitations in desktop mode, such as theinability to freely move an avatar’s limbs, or perform interactions thatrequire more than one hand.
  • In 2020, a new visual programming language was introduced known as”Udon”, which uses a node graph system. While still considered alphasoftware, it became usable on publicly-accessible worlds beginning inApril 2020. A third-party compiler known as “UdonSharp” was developed toallow world scripts to be written in C sharp.”
Vircadia
  • The applications and platforms detailed above have their benefits, butfor the application stack in the next section of the book Vircadia hasbeen chosen. The following text is from their website, and is aplaceholder which gives some idea. This section will be written outcompletely to reflect our use of the product to support emerging market users.
  • Vircadia is open-source software which enables you to create and sharevirtual worlds as virtual reality (VR) and desktop experiences. You cancreate and host your own virtual world, explore other worlds, meet andconnect with other users, attend or host live VR events, and much more.
  • The Vircadia metaverse provides built-in social features, includingavatar interactions, spatialized audio, and interactive physics.Additionally, you have the ability to import any 3D object into yourvirtual environment. No matter where you go in Vircadia, you will alwaysbe able to interact with your environment, engage with your friends, andlisten to conversations just like you would in real life.

Foundation Models

  • Foundation models are large-scale, pre-trained models that can be adapted to a wide range of downstream tasks. They are trained on massive datasets of text and code and can be used for a variety of natural language processing (NLP) tasks, such as text generation, summarization, and question answering.

Animation: Breathing Life into Digital Characters

Logseq

  • Logseq: is very similar to Obsidian, but self hosted and open source. It works on top of plain text files stored in a local system. It supports markdown and Org-mode formatting and allows for hierarchical and networked note-taking. It can be connected to it’s mobile app via github.
  • Integration to Large Language Models can be OpenAI or local.
    • Compare notion, obsidian, and logseq, using a simply markdown table with coloured dots
  • ChatGPT Logseq Summarizer (openai.com)

Screenshot 2024-01-06 120253.png

Screenshot 2024-01-18 103043.png

Screenshot 2024-01-18 102807.png

Key LLM Papers

GPT-1 (2018): This paper introduces the first version of Generative Pre-trained Transformer (GPT), a generative model trained on a massive dataset of text. It demonstrates the ability of LLMs to generate coherent and grammatically correct text, paving the way for future advancements.

GPT-2 (2019): This paper presents a significantly larger GPT model with improved capabilities. It showcases the ability of LLMs to perform various language tasks, including text summarization, question answering, and even code generation.

GPT-3 (2020): This paper introduces GPT-3, a truly massive LLM with billions of parameters. It demonstrates impressive capabilities in diverse tasks, showcasing the emergence of general-purpose language abilities.

GPT-4 (2023): This paper introduces the latest iteration of GPT, featuring multi-modal capabilities and advanced reasoning abilities. It further pushes the boundaries of what LLMs can achieve, demonstrating impressive performance in a wide range of tasks.

Llama-2 (2023): This paper introduces Llama-2, a large language model designed with a focus on efficiency and accessibility. It offers a more resource-friendly alternative to other LLMs, making it more accessible for research and development.

Tools (2023): This paper introduces the “Tools” paradigm for LLMs, allowing them to interact with external tools and resources. It enables LLMs to perform more complex tasks by leveraging the power of external tools, expanding their capabilities significantly.

Gemini-Pro-1.5 (2023): This paper introduces Gemini-Pro-1.5, a large language model developed by Google. It showcases impressive capabilities in various tasks, including code generation, creative writing, and reasoning. It’s a strong contender in the race for developing advanced LLMs.

Generative Models for Molecule Design

  • Generative models based on diffusion and flow-matching approaches enable fine-grained control over the generation of molecules with specific properties. ProteinDT and MoleculeSTM are examples of text-conditioned generative models that allow users to provide natural language prompts to generate molecules with desired properties.
  • RF Diffusion, a diffusion model built on the RoseTTAFold backbone, offers powerful functionalities for protein engineering. It enables unconditional generation of novel proteins, binder design for high affinity and specificity, partial diffusion for refining existing structures, motif scaffolding for combining functional motifs, symmetric generation of protein complexes, and fold conditioning for generating proteins with specific tertiary structures.
  • Complementary models like Ligand and PNN (Protein MPNN) are essential for designing amino acid sequences that fold into the desired 3D structures generated by RF Diffusion.

Some software choices

  • It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.
  • proprietary
  • OpenAI’s Sora model represents a notable advancement in AI video generation. It demonstrates the ability to generate videos up to one minute in 1080p resolution and produce high-resolution images. Sora’s flexibility in handling various aspect ratios and resolutions indicates its adaptability in content creation. Its development leverages insights from prior research, including Vision Transformers and advanced training methodologies.

  • Introduction to Sora
    • A groundbreaking AI video generation model by OpenAI, Sora is designed to transform text instructions into realistic and imaginative video scenes, marking a significant advancement in creative AI technologies.
  • Technical Overview

https://twitter.com/drjimfan/status/1758355737066299692?s=46

  • Creative and Professional Applications
    • Opens up endless possibilities for filmmakers, advertisers, educators, and content creators to produce cinema-quality visuals, educational materials, and immersive experiences effortlessly.
  • Democratization of Video Production
    • Simplifies the video creation process, enabling individuals and small teams to produce content that rivals big studio outputs.
  • Enhancement of Creative Expression
    • Allows creators to bring intricate visions and stories to life through simple text prompts, expanding visual storytelling horizons.
  • Technical Insights
    • Designed to scale language model capabilities to visual data, converting videos into patches for efficient processing and diverse video/image handling.
    • Features a video compression network for temporal and spatial video compression, operating within a Neural Network Latent Space.
    • Uses a diffusion transformer architecture, effectively scaling video generation and improving sample quality with increased compute.
  • Innovative Features
    • Works with videos at native sizes to offer sampling flexibility and improve composition and framing.
    • Leverages descriptive captioning technique, enhancing video fidelity and quality from text prompts.
    • Can animate still images and extend videos, including seamless interpolation between two videos, showcasing versatility.
  • Emerging Capabilities
    • Exhibits capabilities like 3D consistency, long-range coherence, object permanence, and world interaction simulation.
    • Suggests potential as a tool for simulating physical and digital environments, aiding in the development of capable simulators.
    • Videos can serve as a basis for constructing detailed 3D scenes using techniques like Neural Radiance Fields (NeRFs), potentially revolutionizing 3D content creation and interaction.
    • Rapid prototyping and realization of 3D environments and narratives enhance VR and AR immersion and interactivity.
    • Enables generation of characters, objects, and worlds through text and voice prompts, making 3D content creation more intuitive and accessible.
    • Already being used to create 360 spherical video.
  • Research and Discussion
    • Video generation models as world simulators (openai.com) research paper highlights Sora’s technical foundation and its role in simulating the physical world.
    • Discussions emphasize Sora’s potential in democratizing video creation and the need for granular output control for artistic purposes.
  • Google DeepMind on X: “Introducing Veo: our most capable generative video model. 🎥 It can create high-quality, 1080p clips that can go beyond 60 seconds. From photorealism to surrealism and animation, it can tackle a range of cinematic styles. 🧵 GoogleIO https://t.co/6zEuYRAHpH” / X (twitter.com)

https://twitter.com/GoogleDeepMind/status/1790435824598716704

VideoPoet – Google Research

  • Overview: Google’s text to video, linked to Bard, but not yet available.

H200-NeMo-performance

Birme image resizer

Renderings from Plan Drawings

  • Vectorworks AI Visualizer (FAQ)
    • Works inside Vectorworks 2024+, using your active file or view plus a text prompt.
    • Ideal for quick concept iterations (materials, lighting variations).
    • Note: not CAD-accurate rendering but excellent for inspirational visuals.
  • Veras AI for Vectorworks (EvolveLAB announcement)
    • Plugin that uses your 3D model or 2D viewport as a base.
    • Photorealistic or stylised renders in seconds with prompt-driven material and ambience overrides.
  • Mainstream Text-to-Image Generators
    • Export plan or massing views as PNG/JPG and feed into Midjourney, Stable Diffusion (with ControlNet) or DALL·E 3 for high-res concept images.
    • Best for early-stage mood boards rather than precise layouts.

ComfyTextures

  • ComfyTextures GitHub - - ComfyTextures is a collection of free, high-quality textures designed for use in 3D rendering and other creative projects.
  • The textures are organised into logical categories such as wood, metal, fabric, and stone, making it easier to find the desired material.
  • Each texture comes with various maps (diffuse, normal, roughness, specular, height) to facilitate realistic material creation in different rendering engines.
  • The textures are generally provided in a tileable format allowing for seamless repetition across surfaces.
  • The repository is actively maintained, with additions and updates being made regularly, enhancing the available resource base.
  • The textures can be downloaded and used for both commercial and non-commercial purposes under a specified licence.
  • The repository aims to provide a valuable resource for artists and developers seeking readily accessible and customisable textures.
  • Many textures include variations in colour and detail allowing for greater control over the final appearance.

Imagine 3D Software

  • Imagine 3D - Luma Labs Imagine allows users to create realistic 3D models from text descriptions, streamlining the design workflow.
  • It offers an intuitive interface to easily generate, edit and visualise 3D assets.
  • Users can control the colour, texture, and shape of the generated 3D models using natural language processing.
  • The tool enables users to iterate quickly on design ideas by making adjustments to the text prompt and regenerating the model.
  • Imagine facilitates the creation of customised 3D models for various applications, including gaming, product visualisation and animation.
  • The platform encourages experimentation with different prompts to explore the creative potential of artificial intelligence-powered 3D generation.
  • This technology could be used for rapid prototyping, game development, and creation of virtual environments.
  • GET3D aims to democratise 3D content creation by simplifying the process and reducing reliance on expert 3D modellers.

Physically Based Textures from BIM (Revit)

Screenshot 2025-07-24 173949.png

Evolution from Chat to Complex Systems

  • Context engineering emerged as AI systems evolved beyond simple chat interfaces to incorporate:
    • Function calling and tool use
    • Retrieval augmented generation (RAG) systems
    • Multi-agent workflows
    • External API integrations
    • The principle of “garbage in, garbage out” becomes critical when managing complex information flows. Pre-processing and cleaning data before it enters the context window significantly improves output quality.

Text-to-Speech

  • Text-to-speech (TTS) technology can be used to convert written text into spoken audio. This can be used to create podcasts from blog posts, articles, or other written content.

Project Details

  • Technology and Process
    • Utilizing highly-efficient energy generation equipment, the project transforms methane, a natural landfill byproduct, into electricity.
    • This electricity is used for several on-site applications, notably for powering data centers.
  • Environmental and Economic Impacts
    • Many U.S. landfills lack proper methane management systems.
    • Recent studies suggest that landfill methane emissions might be significantly higher than previously estimated.
  • Challenges in Traditional Energy Projects
    • Traditional grid-connected landfill energy projects face high costs and long lead times.
    • Over 70% of the U.S.’s approximately 2,600 municipal landfills lack a viable use for the methane they produce.

Stable Diffusion in Blender

  • A Blender addon for using Stable Diffusion to render texture bakes for objects.

Dream Textures

  • A Blender addon for applying textures with text prompts.

New AI Model Releases

  • GPT4-x-Alpaca-13B-Native-4bit-128g: Technical discussions on the new model and its capabilities (GitHub Discussion).
Vircadia
  • The applications and platforms detailed above have their benefits, butfor the application stack in the next section of the book Vircadia hasbeen chosen. The following text is from their website, and is aplaceholder which gives some idea. This section will be written outcompletely to reflect our use of the product to support emerging market users.
  • Vircadia is open-source software which enables you to create and sharevirtual worlds as virtual reality (VR) and desktop experiences. You cancreate and host your own virtual world, explore other worlds, meet andconnect with other users, attend or host live VR events, and much more.
  • The Vircadia metaverse provides built-in social features, includingavatar interactions, spatialized audio, and interactive physics.Additionally, you have the ability to import any 3D object into yourvirtual environment. No matter where you go in Vircadia, you will alwaysbe able to interact with your environment, engage with your friends, andlisten to conversations just like you would in real life.

LM Studio

  • Integrates advanced tools like text-to-speech (TTS).
  • Highly optimised for macOS environments.
  • Link: Msty App

Logseq

Screenshot 2024-01-06 120253.png

Screenshot 2024-01-18 103043.png

Screenshot 2024-01-18 102807.png

F o u n d a t i o n a l   C o n c e p t s

GPT-1 (2018): This paper introduces the first version of Generative Pre-trained Transformer (GPT), a generative model trained on a massive dataset of text. It demonstrates the ability of LLMs to generate coherent and grammatically correct text, paving the way for future advancements.

GPT-2 (2019): This paper presents a significantly larger GPT model with improved capabilities. It showcases the ability of LLMs to perform various language tasks, including text summarization, question answering, and even code generation.

GPT-3 (2020): This paper introduces GPT-3, a truly massive LLM with billions of parameters. It demonstrates impressive capabilities in diverse tasks, showcasing the emergence of general-purpose language abilities.

GPT-4 (2023): This paper introduces the latest iteration of GPT, featuring multi-modal capabilities and advanced reasoning abilities. It further pushes the boundaries of what LLMs can achieve, demonstrating impressive performance in a wide range of tasks.

Llama-2 (2023): This paper introduces Llama-2, a large language model designed with a focus on efficiency and accessibility. It offers a more resource-friendly alternative to other LLMs, making it more accessible for research and development.

Tools (2023): This paper introduces the “Tools” paradigm for LLMs, allowing them to interact with external tools and resources. It enables LLMs to perform more complex tasks by leveraging the power of external tools, expanding their capabilities significantly.

Gemini-Pro-1.5 (2023): This paper introduces Gemini-Pro-1.5, a large language model developed by Google. It showcases impressive capabilities in various tasks, including code generation, creative writing, and reasoning. It’s a strong contender in the race for developing advanced LLMs.

Agents in Biological Research

  • AI agents have the potential to transform biological research by automating tasks such as literature review, hypothesis generation, experimental design, and data analysis. Companies like Future House are developing AI agents that can identify potential drug targets and design experiments, significantly accelerating the process of discovery. These agents, powered by large language models (LLMs) and other AI technologies, can review thousands of research papers, develop targets or hypotheses to test, and even drive autonomous labs.
  • As these AI agents become more capable, they may play a crucial role in guiding research and helping humans navigate the complex landscape of biological data and interactions. The convergence of AI agents with specific tools for designing molecules, proteins, and nucleic acids could lead to rapid progress in solving challenging problems in biology and medicine.

Some software choices

  • It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.

VideoPoet – Google Research

Text2Mesh

  • The code is organised in a modular fashion, allowing for easy customisation and extension of the system.
  • The repository contains detailed instructions on how to set up the environment, download necessary models, and run the text-to-mesh pipeline.
  • Users can adjust parameters to control the style, complexity, and colour of the generated 3D meshes.
  • The project highlights the potential of automation to simplify 3D content creation and make it more accessible to a wider audience.

Text-to-Speech

  • Text-to-speech (TTS) technology can be used to convert written text into spoken audio. This can be used to create podcasts from blog posts, articles, or other written content.

Dream Textures

Some software choices

  • It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.

Overview

  • Imagine being able to verbally command a virtual design software to create specific CAD primitives or modify existing models. Additionally, the ability to add text annotations or descriptions directly within the virtual space can facilitate collaboration and communication among users.
  • Furthermore, as corporate metaverse like NVIDIA Omniverse Platform expands, the shared virtual spaces will become increasingly complex and vast, accommodating a multitude of digital twin models. This means that users will be able to explore and interact with realistic replicas of real-world objects and environments, such as buildings, vehicles, or even entire cities.
  • By incorporating voice and text input functionalities, developers can empower users to manipulate and navigate these digital twin models more intuitively. Whether it’s adjusting the dimensions of a virtual prototype or performing intricate measurements, the metaverse’s ability to recognize and respond to voice and text commands will revolutionize the way we design, simulate, and experience virtual environments.
  • Table Of Contents — bd_warehouse “0.1.0” # Uncomment this for the next release? documentation (bd-warehouse.readthedocs.io)
  • [Latest General topics

Some software choices

  • It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.

Text to Multiview and Texturing

Multi-Modal Large Language Models (LLMs)

  • Introduction:
    • Large Language Models are adept at generating coherent text sequences, predicting word probabilities and co-occurrences.

April 2024

Text to Multiview and Texturing

November 2024

April 2024

Emerging use cases (AI holding text)

Luma Dream Machine?

  • Luma Dream Machine is a browser-based AI video generator developed by Luma Labs, a San Francisco-based startup. It allows users to generate short videos (around 5 seconds) by simply entering a text or image prompt.
    • Free to Use: Luma Dream Machine is free to try, with no waiting list or subscription required. Users get 30 free video generations per month.
    • High-Quality Output: The AI produces impressively clean and detailed videos, adhering to prompts accurately and generating relatively coherent motion.
    • Fast Generation: Videos are generated in around 2 minutes after entering the prompt.
    • Consistent Subjects: Characters and subjects appear consistent throughout the video, capable of expressing emotion better than many previous AI video models.
    • Difficulty with complex scenes or full-body shots
    • Text in videos may appear garbled
    • Anatomical issues like extra limbs or heads
  • Meta’s Approach: Foundational World Modeling Meta (formerly Facebook) is taking a distinct approach, focusing on the underlying world modeling needed for video encoding and generation. This emphasis on understanding the principles of physics and object interactions could contribute to more realistic AI-generated videos.
  • Technical Capabilities and Limitations
    • Capabilities Current AI video generators demonstrate proficiency in producing high-resolution images and videos. They are capable of style adaptation, simulating complex scenes with multiple elements, and handling variations in aspect ratio and resolution.
    • Limitations Despite their strengths, these models still struggle to accurately simulate physics and lack a complete understanding of cause and effect. Occasional errors regarding object permanence highlight the existing gap between pattern recognition and a comprehensive understanding of the world.
  • Ethical and Creative Considerations
    • Potential Impacts Advancements in AI video generation raise questions about the future of creative professions and the ethical implications of AI-generated content. Balancing technological innovation with safeguarding the integrity of human creativity is an important consideration.

Texturing

Text to Multiview and Texturing

November 2024

April 2024

Evaluation

  • Comparison and Detection
  • LLM QA Evaluation on Wikipedia: An insightful comparison of different LLMs’ performance on QA tasks using Wikipedia as a benchmark. LLM QA Evaluation Wikipedia
  • This study offers a comparative analysis highlighting the strengths and weaknesses of open-source vs closed-source LLMs in handling QA tasks, providing valuable insights for both developers and users.
  • LLM Zoo: A collection of various LLMs to explore and compare their capabilities. LLMZoo GitHub
  • A unique repository that provides access to a wide range of LLMs, facilitating exploration, comparison, and understanding of different models’ functionalities and performance.
  • Can AI-Generated Text be Reliably Detected?: Addresses the critical question of distinguishing between human and AI-generated text. AI-Generated Text Detection Study
  • This paper delves into the challenges and methodologies involved in detecting AI-generated text, offering insights into the reliability of current detection techniques.

Introduction to Large Language Models

  • Large Language Models (LLMs) like OpenAI’s GPT series have revolutionized the field of artificial intelligence, offering unprecedented capabilities in natural language understanding and generation. These models are trained on vast amounts of text data, enabling them to perform a wide range of language-based tasks, from writing and translation to answering questions and generating code.
  • This is a jargon free primer

Mental health Employment Social Contract Under Automation

  • Fraudulent studies are undermining the reliability of systematic reviews – a study of the prevalence of problematic images in preclinical studies of depression | bioRxiv Death of the Internet Deepfakes and fraudulent content

  • Jonathan Haidt Wants You to Take Away Your Kid’s Phone | The New Yorker

    • Jonathan Haidt, a social psychologist and the author of the book “The Anxious Generation: How the Great Rewiring of Childhood is Causing an Epidemic of Mental Illness”. The main points covered in the interview are:
      • Haidt argues that a whole generation has been damaged by growing up with unrestricted access to social media and an overprotected childhood, leading to a sharp increase in anxiety, depression, and self-harm among teenagers, especially girls, starting around 2012.
      • He attributes this to the rapid adoption of smartphones and social media platforms between 2010 and 2015, which radically changed childhood by replacing real-life interactions and play with excessive screen time and exposure to harmful online content.
      • Haidt presents evidence from correlational and experimental studies to support his claim that social media use causes mental health issues, while acknowledging the need for more research.
      • He argues that the benefits of social media are outweighed by its negative impact on child development, as it deprives children of essential real-life experiences, such as play, adventure, and healthy risk-taking.
      • Haidt advocates for changing social norms and implementing restrictions on social media use, such as banning phones in schools and limiting access to social media platforms for children under 16, rather than an outright ban on technology.

Emerging use cases (AI holding text)

Style transfer for humans

Texturing

Text to Multiview and Texturing

Text to 3D

November 2024

April 2024

Evaluation

  • Comparison and Detection
  • LLM QA Evaluation on Wikipedia: An insightful comparison of different LLMs’ performance on QA tasks using Wikipedia as a benchmark. LLM QA Evaluation Wikipedia
  • This study offers a comparative analysis highlighting the strengths and weaknesses of open-source vs closed-source LLMs in handling QA tasks, providing valuable insights for both developers and users.
  • LLM Zoo: A collection of various LLMs to explore and compare their capabilities. LLMZoo GitHub
  • A unique repository that provides access to a wide range of LLMs, facilitating exploration, comparison, and understanding of different models’ functionalities and performance.
  • Can AI-Generated Text be Reliably Detected?: Addresses the critical question of distinguishing between human and AI-generated text. AI-Generated Text Detection Study
  • This paper delves into the challenges and methodologies involved in detecting AI-generated text, offering insights into the reliability of current detection techniques.

Introduction to Large Language Models

  • Large Language Models (LLMs) like OpenAI’s GPT series have revolutionized the field of artificial intelligence, offering unprecedented capabilities in natural language understanding and generation. These models are trained on vast amounts of text data, enabling them to perform a wide range of language-based tasks, from writing and translation to answering questions and generating code.
  • This is a jargon free primer

Mental health Employment Social Contract Under Automation

  • Fraudulent studies are undermining the reliability of systematic reviews – a study of the prevalence of problematic images in preclinical studies of depression | bioRxiv Death of the Internet Deepfakes and fraudulent content

  • Jonathan Haidt Wants You to Take Away Your Kid’s Phone | The New Yorker

    • Jonathan Haidt, a social psychologist and the author of the book “The Anxious Generation: How the Great Rewiring of Childhood is Causing an Epidemic of Mental Illness”. The main points covered in the interview are:
      • Haidt argues that a whole generation has been damaged by growing up with unrestricted access to social media and an overprotected childhood, leading to a sharp increase in anxiety, depression, and self-harm among teenagers, especially girls, starting around 2012.
      • He attributes this to the rapid adoption of smartphones and social media platforms between 2010 and 2015, which radically changed childhood by replacing real-life interactions and play with excessive screen time and exposure to harmful online content.
      • Haidt presents evidence from correlational and experimental studies to support his claim that social media use causes mental health issues, while acknowledging the need for more research.
      • He argues that the benefits of social media are outweighed by its negative impact on child development, as it deprives children of essential real-life experiences, such as play, adventure, and healthy risk-taking.
      • Haidt advocates for changing social norms and implementing restrictions on social media use, such as banning phones in schools and limiting access to social media platforms for children under 16, rather than an outright ban on technology.

Emerging use cases (AI holding text)

Style transfer for humans

Texturing

Text to Multiview and Texturing

Text to 3D

Evaluation

  • Comparison and Detection
  • LLM QA Evaluation on Wikipedia: An insightful comparison of different LLMs’ performance on QA tasks using Wikipedia as a benchmark. LLM QA Evaluation Wikipedia
  • This study offers a comparative analysis highlighting the strengths and weaknesses of open-source vs closed-source LLMs in handling QA tasks, providing valuable insights for both developers and users.
  • LLM Zoo: A collection of various LLMs to explore and compare their capabilities. LLMZoo GitHub
  • A unique repository that provides access to a wide range of LLMs, facilitating exploration, comparison, and understanding of different models’ functionalities and performance.
  • Can AI-Generated Text be Reliably Detected?: Addresses the critical question of distinguishing between human and AI-generated text. AI-Generated Text Detection Study
  • This paper delves into the challenges and methodologies involved in detecting AI-generated text, offering insights into the reliability of current detection techniques.

Introduction to Large Language Models

  • Large Language Models (LLMs) like OpenAI’s GPT series have revolutionized the field of artificial intelligence, offering unprecedented capabilities in natural language understanding and generation. These models are trained on vast amounts of text data, enabling them to perform a wide range of language-based tasks, from writing and translation to answering questions and generating code.

  • This is a jargon free primer

    Core Characteristics

  • Autoregressive Generation: Sequential token-by-token text production

  • Conditional Generation: Text production conditioned on prompts or contexts

  • Controllable Attributes: Style, tone, length, and topic control

  • Few-Shot and Zero-Shot: Generation from minimal examples or instructions

  • Factual Consistency: Grounding in knowledge and reducing hallucination

  • Multi-Domain: News, creative writing, technical documentation, code

    Relationships

  • Subclass: Natural Language Processing

  • Related: Language Modeling, Large Language Model, GPT, Text-to-Text Generation

  • Models: GPT-3/4, T5, BLOOM, LLaMA, PaLM

  • Applications: Content Creation, Code Generation, Creative Writing, Summarisation

    Key Literature

    1. Brown, T., et al. (2020). “Language models are few-shot learners.” NeurIPS, 1877-1901.

    2. Raffel, C., et al. (2020). “Exploring the limits of transfer learning with a unified text-to-text transformer.” JMLR, 21(140), 1-67.

    3. Radford, A., et al. (2019). “Language models are unsupervised multitask learners.” OpenAI Blog.

    4. Holtzman, A., et al. (2020). “The curious case of neural text degeneration.” ICLR.

    2024-2025: Reasoning Models and Multimodal Text Generation Breakthrough

    The period from 2024 through 2025 witnessed transformative developments in text generation, with the emergence of reasoning-optimised models, widespread multimodal integration, and intense competition driving rapid performance improvements across all major frontier language models.

    Reasoning-First Architecture: OpenAI o1 and o3

    In September 2024, OpenAI unveiled o1, experimental models specifically fine-tuned to generate chains of thought before providing answers, scoring particularly high in mathematics, coding, and science benchmarks. Released in full on 5th December 2024, o1 marked a significant shift toward reasoning-first architecture, representing the first model explicitly optimised for chain-of-thought reasoning.

    In December 2024, OpenAI offered a glimpse of o3—o1’s successor with impressive capabilities—whilst Google and DeepSeek unveiled their own reasoning models, establishing reasoning as a core paradigm for text generation going forward.

    Multimodal Text Generation Revolution

    GPT-4o was released on 13th May 2024 as a flagship multimodal model designed to process and generate text, audio, and visual inputs and outputs in real time. The “o” stands for “omni,” signalling the model’s ability to handle longer conversations with better memory whilst understanding both text and images.

    It was truly in 2024 that multimodal LLMs became mainstream. Claude 3.5 Sonnet (released June 2024) excelled in reading, coding, mathematics, and vision tasks. In May 2025, Anthropic introduced the Claude 4 family, including Claude 4 Opus and Claude 4 Sonnet, with Opus 4 optimised for complex reasoning and coding.

    Performance Convergence and Competition

    Some new iterations of fast models (GPT-4o, Gemini 2.0 Flash, and Claude 3.5 Sonnet) became more performant than the flagship models of the previous generation (GPT-4, Gemini 1.5 Pro, Claude 3 Opus), demonstrating accelerating performance improvements. By the end of 2024, OpenAI’s leadership faced stiff competition, with GPT-4o tied with o1 and two versions of Google’s Gemini for first place on the LMSYS Chatbot Arena leaderboard.

    Gemini 2.0 by Google DeepMind launched in December 2024, expanding AI’s multimodal potential and integrating seamlessly with autonomous agents. Gemini 2.0 Flash emerged as one of the fastest options for text generation tasks.

    Competitive Landscape Maturation

    The text generation landscape matured significantly in 2024-2025, transitioning from OpenAI dominance to a highly competitive multi-player market with Google, Anthropic, Meta, and DeepSeek all fielding competitive frontier models. This competition drove rapid capability improvements, pricing reductions, and broader accessibility to state-of-the-art text generation capabilities.

    See Also

  • Natural Language Processing

  • Language Modeling

  • Large Language Model

  • GPT

    Core Characteristics

  • Autoregressive Generation: Sequential token-by-token text production

  • Conditional Generation: Text production conditioned on prompts or contexts

  • Controllable Attributes: Style, tone, length, and topic control

  • Few-Shot and Zero-Shot: Generation from minimal examples or instructions

  • Factual Consistency: Grounding in knowledge and reducing hallucination

  • Multi-Domain: News, creative writing, technical documentation, code

    Relationships

  • Subclass: Natural Language Processing

  • Related: Language Modeling, Large Language Model, GPT, Text-to-Text Generation

  • Models: GPT-3/4, T5, BLOOM, LLaMA, PaLM

  • Applications: Content Creation, Code Generation, Creative Writing, Summarisation

    Key Literature

    1. Brown, T., et al. (2020). “Language models are few-shot learners.” NeurIPS, 1877-1901.

    2. Raffel, C., et al. (2020). “Exploring the limits of transfer learning with a unified text-to-text transformer.” JMLR, 21(140), 1-67.

    3. Radford, A., et al. (2019). “Language models are unsupervised multitask learners.” OpenAI Blog.

    4. Holtzman, A., et al. (2020). “The curious case of neural text degeneration.” ICLR.

    See Also

  • Natural Language Processing

  • Language Modeling

  • Large Language Model

  • GPT

Provenance