Text Generation is the NLP task of producing coherent, contextually appropriate natural language text using neural language models, including applications such as story generation, article writing, code generation, and creative content production. Modern text generation employs transformer-based language models with autoregressive or sequence-to-sequence architectures, controllable generation techniques, and prompt engineering to produce human-quality text across diverse domains and styles.
Semantic Classification
Content
- Text Generation is the NLP task of producing coherent, contextually appropriate natural language text using neural language models, including applications such as story generation, article writing, code generation, and creative content production. Modern text generation employs transformer-based language models (GPT, T5, BLOOM) with autoregressive or sequence-to-sequence architectures, controllable generation techniques, and prompt engineering to produce human-quality text across diverse domains and styles.
Luma AI Genie - Luma Genie is a tool that allows users to create realistic 3D models from text or image prompts using neural networks technology.
-
Users can describe the desired 3D model with a text prompt, specifying details like shape, colour, and texture.
-
Alternatively, users can upload an image to guide the 3D model generation.
-
The generated 3D models can be downloaded in various formats for use in different applications.
-
Genie aims to simplify the process of 3D content creation, making it more accessible to users of all skill levels.
-
It appears to be still under development, with the website showcasing examples and possibilities.
-
The technology uses neural radiance fields to generate high-quality, realistic 3D models.
-
Users can organise and share their created models through the Luma platform.
ComfyTextures
- ComfyTextures GitHub - - ComfyTextures is a collection of free, high-quality textures designed for use in 3D rendering and other creative projects.
- The textures are organised into logical categories such as wood, metal, fabric, and stone, making it easier to find the desired material.
- Each texture comes with various maps (diffuse, normal, roughness, specular, height) to facilitate realistic material creation in different rendering engines.
- The textures are generally provided in a tileable format allowing for seamless repetition across surfaces.
- The repository is actively maintained, with additions and updates being made regularly, enhancing the available resource base.
- The textures can be downloaded and used for both commercial and non-commercial purposes under a specified licence.
- The repository aims to provide a valuable resource for artists and developers seeking readily accessible and customisable textures.
- Many textures include variations in colour and detail allowing for greater control over the final appearance.
Imagine 3D Software
- Imagine 3D - Luma Labs Imagine allows users to create realistic 3D models from text descriptions, streamlining the design workflow.
- It offers an intuitive interface to easily generate, edit and visualise 3D assets.
- Users can control the colour, texture, and shape of the generated 3D models using natural language processing.
- The tool enables users to iterate quickly on design ideas by making adjustments to the text prompt and regenerating the model.
- Imagine facilitates the creation of customised 3D models for various applications, including gaming, product visualisation and animation.
- It allows users to organise and manage generated models within a centralised workspace.
- The platform encourages experimentation with different prompts to explore the creative potential of artificial intelligence-powered 3D generation.
- Imagine simplifies 3D content creation, democratising access to 3D modelling for non-experts.
GET3D by Toronto AI Lab
- GET3D GitHub - GET3D is a generative model that creates high-quality 3D shapes with explicit textures, removing the need for time-consuming 3D modelling.
- It uses a novel texture generation method directly on 3D volumes, resulting in detailed and realistic surface colour.
- The model is trained on unlabelled 2D images using a differentiable rendering pipeline, meaning no 3D supervision is required.
- The generated 3D shapes are mesh-free, represented as neural radiance fields, making them easy to manipulate and edit.
- GET3D allows users to efficiently generate diverse and high-quality 3D assets, accelerating content creation workflows.
- The system offers control over object categories during generation, allowing users to specify the type of 3D shape produced.
- The generated assets are compatible with standard rendering pipelines and can be readily integrated into existing 3D scenes.
- The research demonstrates significant improvements in 3D shape quality and texture detail compared to previous generative models.
- This technology could be used for rapid prototyping, game development, and creation of virtual environments.
- GET3D aims to democratise 3D content creation by simplifying the process and reducing reliance on expert 3D modellers.
Dream Fields for Text-Guided 3D Object Generation
- Dream Fields - * Dreamfields is a technique for visualising and organising thoughts and ideas, similar to mind mapping but with a focus on colours and spatial arrangement.
- The system utilises a canvas (physical or digital) where ideas are represented as colourful nodes, allowing for visual categorisation and intuitive connections.
- Unlike traditional mind maps, Dreamfields encourages free-form arrangement and avoids rigid hierarchical structures, promoting flexible thinking.
- The colour coding allows for the association of different themes, priorities, or categories to individual ideas, aiding in overall organisation.
- The spatial arrangement of nodes on the canvas can reflect relationships, importance, or chronological order of the ideas, providing another layer of meaning.
- Dreamfields is presented as a tool for brainstorming, problem-solving, and project planning, enabling users to explore and connect concepts in a non-linear way.
- The method is suggested to be particularly useful for visual learners and individuals who benefit from a more fluid and intuitive approach to organisation.
- Dreamfields encourages ongoing reflection and iteration, allowing the canvas to evolve and adapt as thinking develops.
Research
- Phi-3 and Phi-4: Powerful language models compact enough to run on a smartphone.
- Amazon
- Amazon is making a significant push in AI with its new “Nova” models, a revamped “Alexa Plus,” and next-generation AI chips.
Physically Based Textures from BIM (Revit)

Evolution from Chat to Complex Systems
- Context engineering emerged as AI systems evolved beyond simple chat interfaces to incorporate:
- Function calling and tool use
- Retrieval augmented generation (RAG) systems
- Multi-agent workflows
- External API integrations
- The principle of “garbage in, garbage out” becomes critical when managing complex information flows. Pre-processing and cleaning data before it enters the context window significantly improves output quality.
Text-to-Speech
- Text-to-speech (TTS) technology can be used to convert written text into spoken audio. This can be used to create podcasts from blog posts, articles, or other written content.
Project Details
- Technology and Process
- Utilizing highly-efficient energy generation equipment, the project transforms methane, a natural landfill byproduct, into electricity.
- This electricity is used for several on-site applications, notably for powering data centers.
- Environmental and Economic Impacts
- The initiative aims to reduce greenhouse gas emissions.
- It also generates revenue, which supports Viridi’s investment in constructing a state-of-the-art Renewable Natural Gas (RNG) facility at the landfill.
- The RNG facility, expected to be fully operational by the second half of 2024, will produce the equivalent of three million gallons of gasoline annually.
Stable Diffusion in Blender
- A Blender addon for using Stable Diffusion to render texture bakes for objects.
Dream Textures
- A Blender addon for applying textures with text prompts.
Prerequisites
- Before beginning, ensure you have:
- A modern web browser for testing your wallet.
- A text editor or IDE (e.g., Visual Studio Code, Sublime Text) for writing and editing code.
- Access to the internet to fetch resources and documentation from the Cashew GitHub repository.
New AI Model Releases
- GPT4-x-Alpaca-13B-Native-4bit-128g: Technical discussions on the new model and its capabilities (GitHub Discussion).
Unsorted Links
- A bunch more GPTs nerority/Advanced-GPTs: Showcase of my custom GPTs, featuring advanced workflows and operational logic. (github.com)
- AIPRM for ChatGPT
- Code Interpreter == GPT 4.5 (w/ Simon Willison, Alex Volkov, Aravind Srinivas, Alex Graveley, et al.)
- Create Your Own ChatGPT ChatBOT With ALL Your Business Content
- GPTZero Case Study (Exploring False Positives)
- Has ChatGPT or me been hacked? Ive never had these conversations..
- How to Create Your Own GPT Voice Assistant with Infinite Chat Memory in Python
- Introducing ChatGPT
- March | 2023 | Ars Technica
- Mintplex-Labs/anything-llm: Open-source ChatGPT equivalent experience for both open and close source LLMs, embedders, and vector databases. Supports unlimited documents, threads, and concurrent users and management all in a very clean UI.
- New model: gpt4-x-alpaca-13b-native-4bit-128g !! · oobabooga text-generation-webui · Discussion #727
- RadioGPT: ‘World’s first’ AI-driven radio station is here
- Scientists begin building AI for scientific discovery using tech behind ChatGPT
- State of GPT | BRK216HFS
- THE DECODER
- TheBloke/starchat-beta-GPTQ · Hugging Face
- Using openai chat gpt to write stable diffusion prompts%22
- Using openai chat gpt to write stable diffusion prompts%7 d%7 btrain
- What is ChatGPT? | OpenAI Help Center
- Yhyu13/30B-Lazarus-gptq-4bit at main
- Zain Kahn on LinkedIn: 1,000+ AI tools were released in March. ChatGPT is just the tip of the… | 345 comments
- mckaywrigley/chatbot-ui: An open source ChatGPT UI.
- sahil280114/chatGPT-multimodal-bot
- ztjhz/BetterChatGPT
- #215 – Wojciech Zaremba: OpenAI Codex, GPT-3, Robotics, and the Future of AI
- #367 – Sam Altman: OpenAI CEO on GPT-4, ChatGPT, and the Future of AI
- 3D-GPT generates 3D worlds in Blender
- 6 Ways ChatGPT Code Interpreter Is Already Being Used
- Can GPT-3 AI write comedy?
- Carlos E. Perez on Twitter / X
- ChatGPT political compass
- Code Interpreter == GPT 4.5 (w/ Simon Willison, Alex Volkov, Aravind Srinivas, Alex Graveley, et al.)
- Developer combines Stable Diffusion, Whisper and GPT-3 for a futuristic design assistant
- George Hotz: Tiny Corp, Twitter, AI Safety, Self-Driving, GPT, AGI & God | Lex Fridman Podcast #387
- Introducing GPT-Furr, Cat-GPT’s meowst sassy, all-knowing system.
- Microsoft just announced a SURPRISE media event taking place tomorrow
- Narrative Manipulation: Convincing Chat GPT to Write a Python Program to Eradicate Humanity
- New model: gpt4-x-alpaca-13b-native-4bit-128g !! · oobabooga text-generation-webui · Discussion #727
- New model: gpt4-x-alpaca-13b-native-4bit-128g !! · oobabooga text-generation-webui · Discussion #727
- Prompt injection attacks against GPT-3
- Sam Altman: OpenAI CEO on GPT-4, ChatGPT, and the Future of AI | Lex Fridman Podcast #367
- Stack Overflow bans ChatGPT as ‘substantially harmful’
- Stephen Wolfram Answers Live Questions About ChatGPT
- Using openai chat gpt to write stable diffusion prompts%22
- Using openai chat gpt to write stable diffusion prompts%7 d%7 btrain
- What is Auto-GPT? | Blog
- You can now run a GPT-3-level AI model on your laptop, phone, and Raspberry Pi
- sahil280114/chatGPT-multimodal-bot
Advice on AI coding
- Choose Tools Strategically: Not all AI coding tools are created equal. Select the right tool for the job, considering the project’s scope and complexity:
- Complex Applications: Cursor, Windsurf, or more established IDE integrations (see below) are often better suited for larger, more intricate projects.
- Micro-SaaS: Bolt/Lovable are optimised for smaller, Software-as-a-Service applications.
- Mobile Applications: Replit remains a good choice, alongside framework-specific tools.
- UI Design: Consider using ‘vo’ or similar specialised tools for user interface design.
- General Coding Assistance & IDE Integration:
- GitHub Copilot: A widely used and powerful AI pair programmer that integrates directly into your IDE (VS Code, JetBrains IDEs, etc.).
- GitHub Copilot Agents: Extend Copilot’s capabilities with specialised agents for tasks like code review, debugging, and test generation.
- Aider: A command-line tool that helps you write and edit code using GPT models. Good for making changes to existing codebases, particularly for refactoring and adding features.
- Roo: Provides code generation and chat capabilities within your IDE.
- Cline: Good for command line interfacing, and code assistance.
- Context is Paramount: Always provide comprehensive context about your project. AI tools cannot “guess” your intentions. Use Markdown (.md) documents to detail:
- Product Requirements Document (PRD): Clearly outlines the purpose, features, and functionality of the application.
- Technical Stack Document: Specifies the programming languages, frameworks, libraries, and databases to be used.
- File Structure: Defines the organisation of directories and files within the project.
- Frontend Guidelines: Describes coding standards, styling conventions, and component structure for the user interface.
- Backend Structure: Outlines the architecture, API endpoints, data models, and business logic for the server-side code.
- Use CodeGuide (or Similar): Consider using CodeGuide or a similar tool to help generate and manage these AI-specific coding documents. This ensures compatibility across various AI tools and helps maintain a single source of truth.
- Incremental Development: Avoid overly broad prompts like “build me an AirBNB clone.” Instead, break down the project into manageable steps:
- Page by Page: Develop the application one page at a time.
- Component by Component: Within each page, build individual components sequentially.
- Limited Task Execution: AI models typically perform best with a maximum of 3 concurrent tasks per request. Be mindful of this limitation, and break down larger tasks accordingly. Tools like Aider and Copilot Agents can help manage this complexity.
- Select AI-Friendly Technologies: Certain technology stacks are better understood by current AI models:
- Web Applications:
- React (with NextJS or ViteJS): Provides excellent performance and is well-supported by AI tools.
- Python (with frameworks like Django or Flask): Widely used and well-understood by AI models.
- Mobile Applications:
- React Native: A good choice for cross-platform development.
- SwiftUI (especially with Claude): Works well, particularly with Claude models.
- Avoid Older Technologies: Unless absolutely necessary, as AI model support may be limited.
- Utilise Starter Kits: Save time and reduce token usage by starting with pre-built templates or boilerplates:
- Example: The “CodeGuide NextJS Starter Kit” can provide a solid foundation.
- Benefit: Accelerates workflow and provides a structured starting point. Most frameworks have readily available starter kits.
- Define Rules Within Your Tools: Many AI coding tools allow project-specific rules:
- Examples: .cursorrules (often “project rules”), .windsurfrules, or similar configuration files within your IDE or tool. Copilot and other IDE-integrated tools often have settings for coding style and preferences.
- Purpose: Constrain the AI, preventing deviations from your guidelines and coding standards.
- Coding Standards: Enforce coding standards using linters (e.g., ESLint for JavaScript, Pylint for Python) and integrate their configuration with your AI tools where possible.
- Employ a Multi-Tool Approach: No single tool handles the entire workflow seamlessly. Combine tools:
- Research: Perplexity.
- Brainstorming: ChatGPT (voice features can be helpful).
- Documentation: CodeGuide, or tools integrated within your IDE.
- Data Scraping: Firecrawl, or libraries within your chosen language (e.g., Beautiful Soup in Python).
- Code Generation/Assembly/Refactoring: Your chosen AI coding tool (Cursor, Windsurf, GitHub Copilot, Aider, Roo, Cline, etc.). Choose based on your workflow and project needs.
- Patience and Persistence: Working with AI requires a specific mindset.
- Prompt Engineering: Crafting effective prompts is crucial. Experiment with different phrasing and levels of detail.
- Expect Errors: AI models are not perfect. Be prepared for errors.
- Iterative Refinement: Stay focused, learn from mistakes, and iteratively refine your prompts and approach.
- Debugging: Provide the AI with the full code and error message for assistance. Leverage Copilot Agents for debugging tasks.
- Version Control
- Use Git for version control.
- Commit frequently with clear messages.
- AI can help generate commit messages (Copilot, Aider, and others offer this).
- Testing
- Write unit and integration tests.
- AI can assist in generating test cases (Copilot Agents are particularly useful here). Tools like Aider can help refactor code to improve testability.
- Agent Frameworks
Random Links
https://twitter.com/tldraw/status/1782443204710674571
- Paper page Design2Code: How Far Are We From Automating Front-End Engineering? (huggingface.co)
- Generative AI Powered Assistant - Amazon Q - AWS Amazons!
- antworks.ai
- OpenBMB/ChatDev: Create Customized Software using Natural Language Idea (through LLM-powered Multi-Agent Collaboration) (github.com)
- Programming AIs worry me • Buttondown:
- Home | Tabby (tabbyml.com)
- The text discusses the concerns around using AI to generate code, specifically around the idea of proofreading the code. The author describes an experience with using voice-to-text where they found it difficult to proofread the text for errors. The text argues that using AI to generate code changes the work from writing code to proofreading code, and that this is a problem.
- Stop whining blog post
- blog post on LLMs for code
- Engshell shell LLM extension
- Github assist
- Locally run 13B coding optimised model
- Programming AIs worry me • Buttondown (other) The article discusses the ethical implications of using machine learning algorithms to generate art. While some see this as a powerful way to create new and interesting works of art, others worry about the potential for misuse and abuse of these technologies.
- GPT synthesizer
- Colab to get codey
- Build prompts using coding keywords, paper
- Continue for VSCode
- Phind technical answers and pair programmer with vscode plugin
- Starchat beta 4bit
- Sweep github pull requests to code system
- Cursor.so coding with gpt interface
- Code llama 2
- Long llama
- Open interpreter
- Open interpreter and autogen local tutorial
- open interpreter github
- codingbuddy
- deepseek 34b q4 AWQ
- Vercel provides front-end Infrastructure to allow developers to build fast, dynamic websites and applications efficiently at global scale. Its open source Next.js framework powers many leading AI products’ user interfaces.
- Vercel’s new vZero product allows developers to visually iterate on UIs with AI assistance.
- Demo/Tutorial: v0 by Vercel AI Code Generation (youtube.com)
- AI code auto-completion tools like Microsoft Copilot have shown the potential for AI to enhance software development. The latest Microsoft Copilot leverages Instruction-Following Conversational AI System 4 and is extremely good.
- AI will likely be incorporated into most software products going forward to enhance capabilities and engagement. Some experiences are better suited to standalone interfaces rather than cramming functionality into chatbots.
- Effective use of AI tools requires developing specialized skills around prompting, understanding system capabilities and limitations, and framing problems appropriately. Different AI systems have strengths in different domains.
- Software development will transition towards more hybrid human-AI teams, with less focus on writing code line-by-line. AI can provide significant productivity gains by automating rote tasks.
- There are open questions around whether to expose functionality through general chatbot interfaces vs company-specific products. There are strategic and technical considerations favouring bespoke solutions.
- Open source software tends to improve quickly over time and should not be underestimated. However, regulations could potentially suppress open source AI progress.
- gptengineer.app is a commercial offering built on GPT Engineer
- Understand a codebase in github with GPT
- Sourcegraph | Code AI platform
- [Bito AI
- Become a 10X Dev with Bito
- Bito](https://bito.ai/)
- Phind
Audio & Voice Generation
- Task: Create voiceovers, audio content like podcasts, or clone voices for various applications.
- ElevenLabs
- Description: High-quality text-to-speech and voice cloning AI. Offers a library of voices (Community Voices), can create a synthetic version of your own voice, and provide audio narration for websites (Audio Native). Used for audiobooks, podcasts, voiceovers.
- Cost: Free tier available. Paid plans based on character usage, starting around $5 USD/month.
- Website: ElevenLabs
- PlayHT
- Description: AI text-to-voice generator with a large library of voices and languages. Suitable for creating audiobooks, podcasts, and voiceovers.
- Cost: Free plan available. Paid plans based on word count/features, starting around $30 USD/month (billed annually).
- Website: PlayHT
VRChat
- none of this makes money yet
- This text is from wikipedia and will be updated when we have a chance totry VRChat properly. It’s much loved already by the Bitcoin community.
- “VRChat’s gameplay is similar to that of games such as Second Life andHabbo Hotel. Players can create their own instanced worlds in which theycan interact with each other through virtual avatars. A softwaredevelopment kit for Unity released alongside the game gives players theability to create or import character models to be used in the platform,as well as build their own worlds.
- Player models are capable of supporting “audio lip sync, eye trackingand blinking, and complete range of motion.
- VRChat is also capable of running in “desktop mode” without a VRheadset, which is controlled using either a mouse and keyboard, or agamepad. Some content has limitations in desktop mode, such as theinability to freely move an avatar’s limbs, or perform interactions thatrequire more than one hand.
- In 2020, a new visual programming language was introduced known as”Udon”, which uses a node graph system. While still considered alphasoftware, it became usable on publicly-accessible worlds beginning inApril 2020. A third-party compiler known as “UdonSharp” was developed toallow world scripts to be written in C sharp.”
Vircadia
- The applications and platforms detailed above have their benefits, butfor the application stack in the next section of the book Vircadia hasbeen chosen. The following text is from their website, and is aplaceholder which gives some idea. This section will be written outcompletely to reflect our use of the product to support emerging market users.
- Vircadia is open-source software which enables you to create and sharevirtual worlds as virtual reality (VR) and desktop experiences. You cancreate and host your own virtual world, explore other worlds, meet andconnect with other users, attend or host live VR events, and much more.
- The Vircadia metaverse provides built-in social features, includingavatar interactions, spatialized audio, and interactive physics.Additionally, you have the ability to import any 3D object into yourvirtual environment. No matter where you go in Vircadia, you will alwaysbe able to interact with your environment, engage with your friends, andlisten to conversations just like you would in real life.
Foundation Models
- Foundation models are large-scale, pre-trained models that can be adapted to a wide range of downstream tasks. They are trained on massive datasets of text and code and can be used for a variety of natural language processing (NLP) tasks, such as text generation, summarization, and question answering.
Animation: Breathing Life into Digital Characters
- Bringing digital characters to life requires compelling animation. This section explores projects and resources focused on achieving realistic and expressive character movement.
-
- Animatable Gaussians (GitHub Repository): Code for “Animatable Gaussians: Learning Pose-dependent Gaussian Maps for High-fidelity Human Avatar Modeling,” presented at CVPR 2024.
- 3D Gaussian Blendshapes: Exploration of the use of 3D Gaussian Blendshapes for head avatar animation.
- Eggnog: A platform for creating infinite AI videos.
- Efficient Portrait Animation (LivePortrait): A project focused on efficient portrait animation with stitching and retargeting control.
- Note by Daniele: A note on LivePortrait by Daniele.
- ComfyUI Nodes for LivePortrait (GitHub Repository): ComfyUI nodes designed for LivePortrait.
- CLARA2 (GitHub Repository): A 3D-rendered AI agent designed to present PowerPoint presentations.
- Consistent Character API (Replicate): An API for running fofr’s consistent character model.
- Joystick-Controlled Character Manipulation (Twitter): A concept for manipulating character features using a joystick.
- Volucap Authentic Digital Avatars: A company specialising in creating authentic digital avatars.
- Openpose Controlnets (Civitai Article): An article explaining how to use and generate poses with Openpose Controlnets V1.1.
- AI Modelling Agency (Creative Bloq Article): An article discussing the emergence of AI modelling agencies.
- Free VRChat Avatars & 3D Assets (VRCMods): A collection of free VRChat avatars and 3D assets.
- AniTalker: A project focused on animating talking heads.
- VLOGGER: A project related to creating virtual vloggers.
- StoryDiffusion: A project exploring consistent self-attention for long-range image and video generation.
- SoccerNet Game State Reconstruction (GitHub Repository): A project focusing on athlete tracking and identification on a minimap.
- Ukraine’s AI Avatar for Consular Affairs (Reddit Post): A discussion on Ukraine’s use of an AI avatar for consular updates.
- Animating Images with Viggle AI (YouTube Tutorial): A tutorial on animating images using Viggle AI.
- PhysAvatar (Hugging Face Paper): Research on learning the physics of dressed 3D avatars from visual observations.
- Vid2Avatar: A project focused on reconstructing 3D avatars from videos.
- Human Tracking and SLAM Capture (YouTube Video): A demonstration of human tracking and SLAM capture technology.
- ComfyUI Character Turntable with SV3D (Reddit Post): A discussion on creating character turntables using ComfyUI and SV3D.
- Animating Characters for Free (LinkedIn Post): A post highlighting methods for animating characters for free.
- Midjourney Character Reference Feature (Medium Article): An article exploring Midjourney’s Character Reference feature.
- Full-Character Consistency with SDXL (Reddit Post): A discussion on achieving full-character consistency using SDXL.
- Create Bot Emotions (Miku.gg Documentation): Documentation on creating bot emotions within the Miku.gg platform.
- Consistent Emotions on a Character with ComfyUI (Reddit User): A Reddit user’s plans to publish a method for achieving consistent emotions on a character using ComfyUI.
- BakedAvatar: A project focused on avatar creation.
- Dreamtalk (GitHub Repository): The official implementation of Dreamtalk, focusing on expressive talking head generation.
- CharTurnerBeta LoRA (Civitai): A LoRA for multi-direction consistency in Stable Diffusion character generation.
- VividTalk: A project focused on one-shot audio-driven talking head generation.
- Gaussian-Based Avatars (Hugging Face Papers):
- Relightable Gaussian Codec Avatars: Research on using Gaussian codecs for avatar representation.
- Gaussian Head Avatar: Research on creating high-fidelity head avatars using dynamic Gaussians.
- NLW Education Discord Projects (Discord Channel): A Discord channel discussing projects related to ElevenLabs and character/avatar creation.
- D-ID AI Video Mobile App: A mobile app for creating AI videos.
- GAIA (Microsoft): A project from Microsoft exploring advanced avatar technologies.
This meticulously curated collection offers a comprehensive overview of the dynamic field of digital human and avatar creation. Explore, learn, and contribute to the ongoing evolution of this exciting frontier!
Note: Some links may lead to projects or resources that are still under development or experimental. Remember to review any licensing information before using code or assets from these projects.
-
Logseq
- Logseq: is very similar to Obsidian, but self hosted and open source. It works on top of plain text files stored in a local system. It supports markdown and Org-mode formatting and allows for hierarchical and networked note-taking. It can be connected to it’s mobile app via github.
- Integration to Large Language Models can be OpenAI or local.
- Compare notion, obsidian, and logseq, using a simply markdown table with coloured dots
- ChatGPT Logseq Summarizer (openai.com)



Key LLM Papers
GPT-1 (2018): This paper introduces the first version of Generative Pre-trained Transformer (GPT), a generative model trained on a massive dataset of text. It demonstrates the ability of LLMs to generate coherent and grammatically correct text, paving the way for future advancements.
GPT-2 (2019): This paper presents a significantly larger GPT model with improved capabilities. It showcases the ability of LLMs to perform various language tasks, including text summarization, question answering, and even code generation.
GPT-3 (2020): This paper introduces GPT-3, a truly massive LLM with billions of parameters. It demonstrates impressive capabilities in diverse tasks, showcasing the emergence of general-purpose language abilities.
GPT-4 (2023): This paper introduces the latest iteration of GPT, featuring multi-modal capabilities and advanced reasoning abilities. It further pushes the boundaries of what LLMs can achieve, demonstrating impressive performance in a wide range of tasks.
Llama-2 (2023): This paper introduces Llama-2, a large language model designed with a focus on efficiency and accessibility. It offers a more resource-friendly alternative to other LLMs, making it more accessible for research and development.
Tools (2023): This paper introduces the “Tools” paradigm for LLMs, allowing them to interact with external tools and resources. It enables LLMs to perform more complex tasks by leveraging the power of external tools, expanding their capabilities significantly.
Gemini-Pro-1.5 (2023): This paper introduces Gemini-Pro-1.5, a large language model developed by Google. It showcases impressive capabilities in various tasks, including code generation, creative writing, and reasoning. It’s a strong contender in the race for developing advanced LLMs.
Generative Models for Molecule Design
- Generative models based on diffusion and flow-matching approaches enable fine-grained control over the generation of molecules with specific properties. ProteinDT and MoleculeSTM are examples of text-conditioned generative models that allow users to provide natural language prompts to generate molecules with desired properties.
- RF Diffusion, a diffusion model built on the RoseTTAFold backbone, offers powerful functionalities for protein engineering. It enables unconditional generation of novel proteins, binder design for high affinity and specificity, partial diffusion for refining existing structures, motif scaffolding for combining functional motifs, symmetric generation of protein complexes, and fold conditioning for generating proteins with specific tertiary structures.
- Complementary models like Ligand and PNN (Protein MPNN) are essential for designing amino acid sequences that fold into the desired 3D structures generated by RF Diffusion.
Some software choices
- It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.
- proprietary
- OpenAI’s Sora model represents a notable advancement in AI video generation. It demonstrates the ability to generate videos up to one minute in 1080p resolution and produce high-resolution images. Sora’s flexibility in handling various aspect ratios and resolutions indicates its adaptability in content creation. Its development leverages insights from prior research, including Vision Transformers and advanced training methodologies.
- Introduction to Sora
- A groundbreaking AI video generation model by OpenAI, Sora is designed to transform text instructions into realistic and imaginative video scenes, marking a significant advancement in creative AI technologies.
- Technical Overview
- Advanced Diffusion Model
- Employs a sophisticated diffusion process that starts from static noise and incrementally refines to generate high-resolution videos, showcasing an unparalleled leap in video realism and complexity.
- Transformer Architecture
- Leverages the Transformer model’s capabilities for deep understanding and generation of content, adapted here to interpret and create complex visual narratives, ensuring dynamic and coherent video storytelling.
- twitter link to the render loading below https://twitter.com/sainingxie/status/1758433676105310543
- twitter link to the render loading below https://twitter.com/thatguybg/status/1759935959792312461
- Patch-Based Data Representation
- Innovatively represents videos and images as collections of smaller data units, akin to language model tokens, enabling precise and granular control over video generation and editing.
- Advanced Diffusion Model
https://twitter.com/drjimfan/status/1758355737066299692?s=46
- Creative and Professional Applications
- Opens up endless possibilities for filmmakers, advertisers, educators, and content creators to produce cinema-quality visuals, educational materials, and immersive experiences effortlessly.
- Democratization of Video Production
- Simplifies the video creation process, enabling individuals and small teams to produce content that rivals big studio outputs.
- Enhancement of Creative Expression
- Allows creators to bring intricate visions and stories to life through simple text prompts, expanding visual storytelling horizons.
- Technical Insights
- Designed to scale language model capabilities to visual data, converting videos into patches for efficient processing and diverse video/image handling.
- Features a video compression network for temporal and spatial video compression, operating within a Neural Network Latent Space.
- Uses a diffusion transformer architecture, effectively scaling video generation and improving sample quality with increased compute.
- Innovative Features
- Works with videos at native sizes to offer sampling flexibility and improve composition and framing.
- Leverages descriptive captioning technique, enhancing video fidelity and quality from text prompts.
- Can animate still images and extend videos, including seamless interpolation between two videos, showcasing versatility.
- Emerging Capabilities
- Exhibits capabilities like 3D consistency, long-range coherence, object permanence, and world interaction simulation.
- Suggests potential as a tool for simulating physical and digital environments, aiding in the development of capable simulators.
- Videos can serve as a basis for constructing detailed 3D scenes using techniques like Neural Radiance Fields (NeRFs), potentially revolutionizing 3D content creation and interaction.
- Rapid prototyping and realization of 3D environments and narratives enhance VR and AR immersion and interactivity.
- Enables generation of characters, objects, and worlds through text and voice prompts, making 3D content creation more intuitive and accessible.
- Already being used to create 360 spherical video.
- Research and Discussion
- Video generation models as world simulators (openai.com) research paper highlights Sora’s technical foundation and its role in simulating the physical world.
- Discussions emphasize Sora’s potential in democratizing video creation and the need for granular output control for artistic purposes.
- Google DeepMind on X: “Introducing Veo: our most capable generative video model. 🎥 It can create high-quality, 1080p clips that can go beyond 60 seconds. From photorealism to surrealism and animation, it can tackle a range of cinematic styles. 🧵 GoogleIO https://t.co/6zEuYRAHpH” / X (twitter.com)
https://twitter.com/GoogleDeepMind/status/1790435824598716704
VideoPoet – Google Research
- Overview: Google’s text to video, linked to Bard, but not yet available.
- NVIDIA/NeMo: NeMo: a toolkit for conversational AI (github.com)
- [Canary
- NVIDIA NeMo](https://nvidia.github.io/NeMo/blogs/2024/2024-02-canary/)

- NeMo/tutorials/tts/FastPitch_Adapter_Finetuning.ipynb at main · NVIDIA/NeMo (github.com)
- ElevenLabs Audio Native
- OpenAI whisper local deploy
- realtime transciber
- high performance CPP
- 30% quantised optimisation
- Brillbits OpenAI whisper demo with mic
- Cleanvoice audio denoise
- Cloud voice change app
- downloadable voice generation systems
- Language AI open libraries
- Language practice
- MUGEN multi modal from facebook
- Oneshot speach to text
- Record and cleanup pro audio with commodity hardware
- Respeecher
- Voice AI voices
- Voice controlled assisted creation
- Voice to text, Lopp
- whisper transcriber
- Wolfram alpha voice chatbot integration
- Microsoft Vall-E voice synthesis
- Uberduck text to speech (plus own voice)
- Eleven labs language and text to speech
- Uberduck open source text to speech
- numen voice control system in linux
- Inworld (steam game plugin AI system) for voice chat and answer
- Bark text to speech from google labs
- https://github.com/TensorSpeech/TensorFlowTTS very configurable from what I see
- VoiceVox engine
- [coqui-ai TTS
- very good samples](https://github.com/coqui-ai/TTS)
- https://github.com/neonbjb/tortoise-tts
- https://github.com/CorentinJ/Real-Time-Voice-Cloning
- custom voices? looks neat
- https://github.com/rhasspy/larynx - very low-spec compatible, acceptable quality
- Voice cloning local
- Meta voicebox
- The Reddit post discusses the different open source voice cloning projects available, including Coqui, Tortoise, and Bark. The advantages and disadvantages of each project are briefly outlined, with ElevenLabs being noted as the best but not open source, while Tortoise is suggested as the closest open source alternative. Other tools for speech to speech and singing conversion, such as so-vits/diff-svc/rvc, are also mentioned. The post suggests that the quality of open source voice cloning projects is improving, and that there may be more options available in the future. https://www.reddit.com/r/MachineLearning/comments/133hanr/d_what_are_the_differences_between_the_major_open/
- The Retrieval-based Voice Conversion WebUI is a simple and useful voice conversion (voice changer) framework based on the VITS algorithm. It can use a small amount of voice data and still achieve good results. It incorporates a top-1 retrieval method to replace the source feature with the training set feature to avoid voice leakage, and it is easy to use with a simple web interface. It also features model fusion to change voice characteristics and the ability to integrate with the UVR5 model to quickly separate vocals and accompaniment. The project requires the installation of PyTorch and its core dependencies, and other pre-models are also needed for inference and training. The repository provides a guide to environment setup and usage, as well as links to relevant resources and contributors. https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI
- The article discusses different open-source voice cloning projects and their advantages and disadvantages. The projects mentioned include Coqui, Tortoise, and Bark, with the author highlighting Coqui’s unlocked platform, while Tortoise and Bark are newer transformer-based projects that can clone much more effectively with much less training and are restricted to prevent custom voice cloning. The author suggests that the ElevenLabs is currently the best voice cloning solution available, but it is not open source and can be expensive. The article also includes comments from other Reddit users, who suggest other open source options and provide additional insights into each option’s strengths and weaknesses. https://www.reddit.com/r/MachineLearning/comments/133hanr/d_what_are_the_differences_between_the_major_open/
- The article provides instructions on how to use OpenAI’s ChatGPT chatbot on an Android device using the Tasker app. The process involves importing a ChatGPT profile into Tasker, obtaining an API key from OpenAI, and setting up home screen shortcuts. The article also notes that ChatGPT can be run through Google Assistant with voice commands. The author suggests that while ChatGPT may not necessarily be better than Google Assistant, it can perform tasks that Google Assistant may not be capable of. https://www.howtogeek.com/882019/how-to-use-chatgpt-like-google-assistant-on-android/
- The Voice Assistant is an AI-powered chatbot that uses several APIs to understand natural language commands and provide helpful responses. It features a wide range of capabilities, including answering general knowledge questions, providing recommendations, performing productivity tasks, and entertaining users. The Voice Assistant was built using ChatGPT, Whisper API, Gradio, and Microsoft’s SpVoice TTS API, and it can be accessed through a web-based interface. The installation process involves cloning the repository and installing the required Python packages. Contributions to the project are welcome. https://github.com/DonGuillotine/chatGPT_whisper_AI_voice_assistant
- The Retrieval-based Voice Conversion WebUI is a voice conversion framework that uses a top-1 retrieval algorithm to eliminate voice leakage. It is capable of quickly training even on relatively poor GPUs and can achieve good results even with just 10 minutes of low noise voice data. It has a user-friendly web interface and the ability to use a model fusion system to change voice timbre. The setup recommends using Poetry and downloading the necessary pre-trained models from their Hugging Face space. It also includes additional files such as ffmpeg and ffprobe that may need to be downloaded. The WebUI can be initiated using the command “python infer-web.py” and Windows users can run the “go-web.bat” file. The project also acknowledges the contributions of related tools and libraries such as Gradio, HIFIGAN, and ContentVec. https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI
- VoicePen is a tool that uses AI to convert audio or video files into blog posts and transcriptions in minutes. The service includes a transcription and SRT file generated by a top speech-to-text model, an English blog post that pulls out key topics from the audio, and the ability to convert audio in 96 different languages. Use cases include repurposing podcasts, webinars, and tutorial videos. Monthly plans are available, with options for one-time conversions. Testimonials praise the accuracy and speed of VoicePen’s service. https://voicepen.ai
- Krisp is a software application designed to improve the productivity of online meetings by using AI-powered voice clarity and a meeting assistant to cancel background noise, echo, and accent localization. It works on both Mac and Windows platforms and processes only the user’s voice on their device, unlike other solutions that transmit voice over the internet. Krisp offers a free forever plan with no credit card required and is trusted by global brands. The insights gathered from calls can be viewed by the user to improve their communication skills over time. Krisp has received recognition from various prestigious awards such as America’s Most Promising AI Companies and has been awarded for its quality of support and ease of use. Krisp also offers SDK for developers, pricing and plans, and use cases such as contact centers and enterprise. The company prioritizes customers’ privacy, security and offers accessible support, including video tutorials and a help center. By accepting all cookies, users consent to the storing of cookies on their device to enhance site navigation, analyze site usage and assist in the company’s marketing efforts. https://krisp.ai/
- Cleanvoice AI is an artificial intelligence platform that assists users in editing their podcasts or audio recordings. The platform offers various features such as filler sound removal, mouth sound removal, stutter removal, and Deadair remover to make the audio recording more professional. Cleanvoice AI is multilingual and can detect filler sounds in multiple languages, including accents from various countries. The platform also allows for manual editing with assistance and offers tools like podcast mixing and background noise remover. Users can try Cleanvoice AI for free for 30 minutes without providing credit card details. However, users must accept the platform’s cookie policy to use the service. https://cleanvoice.ai/
- The article discusses the potential of Central Intelligent Agents (CIAs) and the role of large language models (LLMs) and other next-generation AI technologies in enabling them. It highlights the need for businesses to have a cross-functional team, ethical guidelines, and clear objectives in deploying their own CIA. The article also suggests steps to build a solid foundation for deploying a CIA, assess organizational readiness, assemble a cross-functional team, define objectives, develop the CIA components and evaluate its performance while continuing to learn and adapt. The author discusses the potential of AI tools and voice assistants in transforming the way businesses interact with their customers and suggests that the advent of advanced AI technologies has revolutionized the shift of businesses towards a more personalized and ethically responsible approach to engaging with their customers. Finally, the article ends by highlighting the importance of experimenting through crisis and providing expert guidance tailored to specific business needs. https://www.linkedin.com/pulse/central-intelligent-agent-enabling-next-generation-james-poulter?
- TensorSpeech/TensorFlowTTS: :stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German and Easy to adapt for other languages) Translation Accessibility Speech and Voice Speech and Voice
- Variety Speech and Voice Employment Social Contract Under Automation
- transcriptionstream/transcriptionstream: turnkey self-hosted offline transcription and diarization service with llm summary (github.com) Speech and Voice transcription locally RFC 2119 SHOULD Normative Keyword
- Tincans - Gazelle v0.2 Speech and Voice fast speech engine RFC 2119 SHOULD Normative Keyword
- Speech and Voice Open Voice (myshell.ai) cloning MIT license
- EndlessDreams: Voice directed real-time videos at 1280x1024 : r/StableDiffusion (reddit.com) Speech and Voice Speech and Voice Product Design Real Time
- https://demo.hume.ai/? Speech and Voice Large Language Models empathetic voice to voice
- Speech and Voice metavoiceio/metavoice-src: AI for human-level speech intelligence (github.com) check for PlayerTwo
- NeMo/tutorials/tts/NeMo_TTS_Primer.ipynb at main · NVIDIA/NeMo (github.com) NVIDIA Omniverse Platform Speech and Voice primer and demo.
Birme image resizer
- 2 hour tutorial
- inject your face into any model (dreambooth)
- Guide for dreambooth
- Shivram
- Progen photorealism Miro guide
- rare dreambooth tokens
- Multi subject tokens
- tag editor
- SDXL dreambooth
- Lora guide
- stable swarm distributed comfyui
- Textual inversion
- Img2Img guide from reddit for face mapping
- textual inversion cheaper training
- CIO blog post
- google stable diffusion
- Cross attention replace named items
- 256 x faster speedup
- VoltaML acceleration
- Depth map into blender from SD2
- midjourney tweaks
- and another
- Updates Pastebin
- Game development using SD
- Wildcard manager using ChatGPT
- Depth2Img for text
- train chat GPT to write prompts
- non destructive image manipulation using seeds
- Instruct pix2pix
- reddit post
- Attention heatmap for prompts (youtube)
- enormous link roundup
- Prompt master variations management
- panoramic world builder
- GitHub AbdullahAlfaraj/Auto-Photoshop-StableDiffusion-Plugin: A user-friendly plug-in that makes it easy to generate stable diffusion images inside Photoshop using Automatic1111-sd-webui as a backend.
- GitHub ashawkey/stable-dreamfusion: A pytorch implementation of text-to-3D dreamfusion, powered by stable diffusion.
- Fine tune stable diffusion
- GitHub Sanster/lama-cleaner: Image inpainting tool powered by SOTA AI Model. Remove any unwanted object, defect, people from your pictures or erase and replace(powered by stable diffusion) any thing on your pictures.
- holovolo immersive volumetric VR180 videos and photos, and 3D stable diffusion, for Quest and WebVR
- The Illustrated Stable Diffusion Jay Alammar Visualizing machine learning one concept at a time.
- reddit educational links
- Negative prompt hack tip
- Modify images with text
- Photorealism
- sdtools image v 1.6
- Character plugin
- Checkpoints
- Stability specific tools
- Arible Prompt Database https://www.arible.co/prompts
- [Guide] Make your own Loras, easy and free | Stable Diffusion Other | Civitai: You don’t need to download anything, this is a guide with online tools. Click “Show more” below.
- sdxl lora training
- dylora scripts
- kohya fork with scripts
- lora of loras (compressed sets)
- chart of print size aspect ratios
- SDXL native text lora
- SDXL lcm motion lora
- SDXL universal negative prompt
- text, watermark, low-quality, signature, moiré pattern, downsampling, aliasing, distorted, blurry, glossy, blur, jpeg artifacts, compression artifacts, poorly drawn, low-resolution, bad, distortion, twisted, excessive, exaggerated pose, exaggerated limbs, grainy, symmetrical, duplicate, error, pattern, beginner, pixelated, fake, hyper, glitch, overexposed, high-contrast, bad-contrast
- SDXL prodigy training guide
- Lora training interface for windows
- Refined model
- Fine tuning with captioning and other fine tuning tricks, followfox
- Negative embedding textual inversion for hands etc
- GitHub kpthedev/ez-text2video: Easily run text-to-video diffusion with customized video length, fps, and dimensions on 4GB video cards, as well as on CPU.
- Gligen grounding capability for sd1.5
- This repository contains a ComfyUI Extension for Automated Text Generation. The extension provides nodes which can be used to automate the text generation process. The goal is to build a node-based Automated Text Generation AGI. This extension should ultimately combine all of the features of the existing text generation tools into one tool.
- [R] Text-to-image Diffusion Models in Generative AI: A Survey: r/MachineLearning
- Tutorial: Creating a Consistent Character as a Textual Inversion Embedding
- Segment anything webui
- segment anything training
- Nvidia stable diffusion segment through clip
- Overriding iphone footage with SD characters using controlnet
- Interactive photo manipulation GAN
- 3d plugin for Automatic1111
- Face replace plugin for automatic
Renderings from Plan Drawings
- Vectorworks AI Visualizer (FAQ)
- Works inside Vectorworks 2024+, using your active file or view plus a text prompt.
- Ideal for quick concept iterations (materials, lighting variations).
- Note: not CAD-accurate rendering but excellent for inspirational visuals.
- Veras AI for Vectorworks (EvolveLAB announcement)
- Plugin that uses your 3D model or 2D viewport as a base.
- Photorealistic or stylised renders in seconds with prompt-driven material and ambience overrides.
- Mainstream Text-to-Image Generators
- Export plan or massing views as PNG/JPG and feed into Midjourney, Stable Diffusion (with ControlNet) or DALL·E 3 for high-res concept images.
- Best for early-stage mood boards rather than precise layouts.
Luma AI Genie - Luma Genie is a tool that allows users to create realistic 3D models from text or image prompts using neural networks technology.
-
Users can describe the desired 3D model with a text prompt, specifying details like shape, colour, and texture.
-
Alternatively, users can upload an image to guide the 3D model generation.
-
The generated 3D models can be downloaded in various formats for use in different applications.
-
Genie aims to simplify the process of 3D content creation, making it more accessible to users of all skill levels.
-
It appears to be still under development, with the website showcasing examples and possibilities.
-
The technology uses neural radiance fields to generate high-quality, realistic 3D models.
-
Users can organise and share their created models through the Luma platform.
ComfyTextures
- ComfyTextures GitHub - - ComfyTextures is a collection of free, high-quality textures designed for use in 3D rendering and other creative projects.
- The textures are organised into logical categories such as wood, metal, fabric, and stone, making it easier to find the desired material.
- Each texture comes with various maps (diffuse, normal, roughness, specular, height) to facilitate realistic material creation in different rendering engines.
- The textures are generally provided in a tileable format allowing for seamless repetition across surfaces.
- The repository is actively maintained, with additions and updates being made regularly, enhancing the available resource base.
- The textures can be downloaded and used for both commercial and non-commercial purposes under a specified licence.
- The repository aims to provide a valuable resource for artists and developers seeking readily accessible and customisable textures.
- Many textures include variations in colour and detail allowing for greater control over the final appearance.
Imagine 3D Software
- Imagine 3D - Luma Labs Imagine allows users to create realistic 3D models from text descriptions, streamlining the design workflow.
- It offers an intuitive interface to easily generate, edit and visualise 3D assets.
- Users can control the colour, texture, and shape of the generated 3D models using natural language processing.
- The tool enables users to iterate quickly on design ideas by making adjustments to the text prompt and regenerating the model.
- Imagine facilitates the creation of customised 3D models for various applications, including gaming, product visualisation and animation.
- It allows users to organise and manage generated models within a centralised workspace.
- The platform encourages experimentation with different prompts to explore the creative potential of artificial intelligence-powered 3D generation.
- Imagine simplifies 3D content creation, democratising access to 3D modelling for non-experts.
GET3D by Toronto AI Lab
- GET3D GitHub - GET3D is a generative model that creates high-quality 3D shapes with explicit textures, removing the need for time-consuming 3D modelling.
- It uses a novel texture generation method directly on 3D volumes, resulting in detailed and realistic surface colour.
- The model is trained on unlabelled 2D images using a differentiable rendering pipeline, meaning no 3D supervision is required.
- The generated 3D shapes are mesh-free, represented as neural radiance fields, making them easy to manipulate and edit.
- GET3D allows users to efficiently generate diverse and high-quality 3D assets, accelerating content creation workflows.
- The system offers control over object categories during generation, allowing users to specify the type of 3D shape produced.
- The generated assets are compatible with standard rendering pipelines and can be readily integrated into existing 3D scenes.
- The research demonstrates significant improvements in 3D shape quality and texture detail compared to previous generative models.
- This technology could be used for rapid prototyping, game development, and creation of virtual environments.
- GET3D aims to democratise 3D content creation by simplifying the process and reducing reliance on expert 3D modellers.
Dream Fields for Text-Guided 3D Object Generation
- Dream Fields - * Dreamfields is a technique for visualising and organising thoughts and ideas, similar to mind mapping but with a focus on colours and spatial arrangement.
- The system utilises a canvas (physical or digital) where ideas are represented as colourful nodes, allowing for visual categorisation and intuitive connections.
- Unlike traditional mind maps, Dreamfields encourages free-form arrangement and avoids rigid hierarchical structures, promoting flexible thinking.
- The colour coding allows for the association of different themes, priorities, or categories to individual ideas, aiding in overall organisation.
- The spatial arrangement of nodes on the canvas can reflect relationships, importance, or chronological order of the ideas, providing another layer of meaning.
- Dreamfields is presented as a tool for brainstorming, problem-solving, and project planning, enabling users to explore and connect concepts in a non-linear way.
- The method is suggested to be particularly useful for visual learners and individuals who benefit from a more fluid and intuitive approach to organisation.
- Dreamfields encourages ongoing reflection and iteration, allowing the canvas to evolve and adapt as thinking develops.
Research
- Phi-3 and Phi-4: Powerful language models compact enough to run on a smartphone.
- Amazon
- Amazon is making a significant push in AI with its new “Nova” models, a revamped “Alexa Plus,” and next-generation AI chips.
Physically Based Textures from BIM (Revit)

Evolution from Chat to Complex Systems
- Context engineering emerged as AI systems evolved beyond simple chat interfaces to incorporate:
- Function calling and tool use
- Retrieval augmented generation (RAG) systems
- Multi-agent workflows
- External API integrations
- The principle of “garbage in, garbage out” becomes critical when managing complex information flows. Pre-processing and cleaning data before it enters the context window significantly improves output quality.
Text-to-Speech
- Text-to-speech (TTS) technology can be used to convert written text into spoken audio. This can be used to create podcasts from blog posts, articles, or other written content.
Project Details
- Technology and Process
- Utilizing highly-efficient energy generation equipment, the project transforms methane, a natural landfill byproduct, into electricity.
- This electricity is used for several on-site applications, notably for powering data centers.
- Environmental and Economic Impacts
- The initiative aims to reduce greenhouse gas emissions.
- It also generates revenue, which supports Viridi’s investment in constructing a state-of-the-art Renewable Natural Gas (RNG) facility at the landfill.
- The RNG facility, expected to be fully operational by the second half of 2024, will produce the equivalent of three million gallons of gasoline annually.
Stable Diffusion in Blender
- A Blender addon for using Stable Diffusion to render texture bakes for objects.
Dream Textures
- A Blender addon for applying textures with text prompts.
Prerequisites
- Before beginning, ensure you have:
- A modern web browser for testing your wallet.
- A text editor or IDE (e.g., Visual Studio Code, Sublime Text) for writing and editing code.
- Access to the internet to fetch resources and documentation from the Cashew GitHub repository.
New AI Model Releases
- GPT4-x-Alpaca-13B-Native-4bit-128g: Technical discussions on the new model and its capabilities (GitHub Discussion).
Unsorted Links
- A bunch more GPTs nerority/Advanced-GPTs: Showcase of my custom GPTs, featuring advanced workflows and operational logic. (github.com)
- AIPRM for ChatGPT
- Code Interpreter == GPT 4.5 (w/ Simon Willison, Alex Volkov, Aravind Srinivas, Alex Graveley, et al.)
- Create Your Own ChatGPT ChatBOT With ALL Your Business Content
- GPTZero Case Study (Exploring False Positives)
- Has ChatGPT or me been hacked? Ive never had these conversations..
- How to Create Your Own GPT Voice Assistant with Infinite Chat Memory in Python
- Introducing ChatGPT
- March | 2023 | Ars Technica
- Mintplex-Labs/anything-llm: Open-source ChatGPT equivalent experience for both open and close source LLMs, embedders, and vector databases. Supports unlimited documents, threads, and concurrent users and management all in a very clean UI.
- New model: gpt4-x-alpaca-13b-native-4bit-128g !! · oobabooga text-generation-webui · Discussion #727
- RadioGPT: ‘World’s first’ AI-driven radio station is here
- Scientists begin building AI for scientific discovery using tech behind ChatGPT
- State of GPT | BRK216HFS
- THE DECODER
- TheBloke/starchat-beta-GPTQ · Hugging Face
- Using openai chat gpt to write stable diffusion prompts%22
- Using openai chat gpt to write stable diffusion prompts%7 d%7 btrain
- What is ChatGPT? | OpenAI Help Center
- Yhyu13/30B-Lazarus-gptq-4bit at main
- Zain Kahn on LinkedIn: 1,000+ AI tools were released in March. ChatGPT is just the tip of the… | 345 comments
- mckaywrigley/chatbot-ui: An open source ChatGPT UI.
- sahil280114/chatGPT-multimodal-bot
- ztjhz/BetterChatGPT
- #215 – Wojciech Zaremba: OpenAI Codex, GPT-3, Robotics, and the Future of AI
- #367 – Sam Altman: OpenAI CEO on GPT-4, ChatGPT, and the Future of AI
- 3D-GPT generates 3D worlds in Blender
- 6 Ways ChatGPT Code Interpreter Is Already Being Used
- Can GPT-3 AI write comedy?
- Carlos E. Perez on Twitter / X
- ChatGPT political compass
- Code Interpreter == GPT 4.5 (w/ Simon Willison, Alex Volkov, Aravind Srinivas, Alex Graveley, et al.)
- Developer combines Stable Diffusion, Whisper and GPT-3 for a futuristic design assistant
- George Hotz: Tiny Corp, Twitter, AI Safety, Self-Driving, GPT, AGI & God | Lex Fridman Podcast #387
- Introducing GPT-Furr, Cat-GPT’s meowst sassy, all-knowing system.
- Microsoft just announced a SURPRISE media event taking place tomorrow
- Narrative Manipulation: Convincing Chat GPT to Write a Python Program to Eradicate Humanity
- New model: gpt4-x-alpaca-13b-native-4bit-128g !! · oobabooga text-generation-webui · Discussion #727
- New model: gpt4-x-alpaca-13b-native-4bit-128g !! · oobabooga text-generation-webui · Discussion #727
- Prompt injection attacks against GPT-3
- Sam Altman: OpenAI CEO on GPT-4, ChatGPT, and the Future of AI | Lex Fridman Podcast #367
- Stack Overflow bans ChatGPT as ‘substantially harmful’
- Stephen Wolfram Answers Live Questions About ChatGPT
- Using openai chat gpt to write stable diffusion prompts%22
- Using openai chat gpt to write stable diffusion prompts%7 d%7 btrain
- What is Auto-GPT? | Blog
- You can now run a GPT-3-level AI model on your laptop, phone, and Raspberry Pi
- sahil280114/chatGPT-multimodal-bot
Advice on AI coding
- Choose Tools Strategically: Not all AI coding tools are created equal. Select the right tool for the job, considering the project’s scope and complexity:
- Complex Applications: Cursor, Windsurf, or more established IDE integrations (see below) are often better suited for larger, more intricate projects.
- Micro-SaaS: Bolt/Lovable are optimised for smaller, Software-as-a-Service applications.
- Mobile Applications: Replit remains a good choice, alongside framework-specific tools.
- UI Design: Consider using ‘vo’ or similar specialised tools for user interface design.
- General Coding Assistance & IDE Integration:
- GitHub Copilot: A widely used and powerful AI pair programmer that integrates directly into your IDE (VS Code, JetBrains IDEs, etc.).
- GitHub Copilot Agents: Extend Copilot’s capabilities with specialised agents for tasks like code review, debugging, and test generation.
- Aider: A command-line tool that helps you write and edit code using GPT models. Good for making changes to existing codebases, particularly for refactoring and adding features.
- Roo: Provides code generation and chat capabilities within your IDE.
- Cline: Good for command line interfacing, and code assistance.
- Context is Paramount: Always provide comprehensive context about your project. AI tools cannot “guess” your intentions. Use Markdown (.md) documents to detail:
- Product Requirements Document (PRD): Clearly outlines the purpose, features, and functionality of the application.
- Technical Stack Document: Specifies the programming languages, frameworks, libraries, and databases to be used.
- File Structure: Defines the organisation of directories and files within the project.
- Frontend Guidelines: Describes coding standards, styling conventions, and component structure for the user interface.
- Backend Structure: Outlines the architecture, API endpoints, data models, and business logic for the server-side code.
- Use CodeGuide (or Similar): Consider using CodeGuide or a similar tool to help generate and manage these AI-specific coding documents. This ensures compatibility across various AI tools and helps maintain a single source of truth.
- Incremental Development: Avoid overly broad prompts like “build me an AirBNB clone.” Instead, break down the project into manageable steps:
- Page by Page: Develop the application one page at a time.
- Component by Component: Within each page, build individual components sequentially.
- Limited Task Execution: AI models typically perform best with a maximum of 3 concurrent tasks per request. Be mindful of this limitation, and break down larger tasks accordingly. Tools like Aider and Copilot Agents can help manage this complexity.
- Select AI-Friendly Technologies: Certain technology stacks are better understood by current AI models:
- Web Applications:
- React (with NextJS or ViteJS): Provides excellent performance and is well-supported by AI tools.
- Python (with frameworks like Django or Flask): Widely used and well-understood by AI models.
- Mobile Applications:
- React Native: A good choice for cross-platform development.
- SwiftUI (especially with Claude): Works well, particularly with Claude models.
- Avoid Older Technologies: Unless absolutely necessary, as AI model support may be limited.
- Utilise Starter Kits: Save time and reduce token usage by starting with pre-built templates or boilerplates:
- Example: The “CodeGuide NextJS Starter Kit” can provide a solid foundation.
- Benefit: Accelerates workflow and provides a structured starting point. Most frameworks have readily available starter kits.
- Define Rules Within Your Tools: Many AI coding tools allow project-specific rules:
- Examples: .cursorrules (often “project rules”), .windsurfrules, or similar configuration files within your IDE or tool. Copilot and other IDE-integrated tools often have settings for coding style and preferences.
- Purpose: Constrain the AI, preventing deviations from your guidelines and coding standards.
- Coding Standards: Enforce coding standards using linters (e.g., ESLint for JavaScript, Pylint for Python) and integrate their configuration with your AI tools where possible.
- Employ a Multi-Tool Approach: No single tool handles the entire workflow seamlessly. Combine tools:
- Research: Perplexity.
- Brainstorming: ChatGPT (voice features can be helpful).
- Documentation: CodeGuide, or tools integrated within your IDE.
- Data Scraping: Firecrawl, or libraries within your chosen language (e.g., Beautiful Soup in Python).
- Code Generation/Assembly/Refactoring: Your chosen AI coding tool (Cursor, Windsurf, GitHub Copilot, Aider, Roo, Cline, etc.). Choose based on your workflow and project needs.
- Patience and Persistence: Working with AI requires a specific mindset.
- Prompt Engineering: Crafting effective prompts is crucial. Experiment with different phrasing and levels of detail.
- Expect Errors: AI models are not perfect. Be prepared for errors.
- Iterative Refinement: Stay focused, learn from mistakes, and iteratively refine your prompts and approach.
- Debugging: Provide the AI with the full code and error message for assistance. Leverage Copilot Agents for debugging tasks.
- Version Control
- Use Git for version control.
- Commit frequently with clear messages.
- AI can help generate commit messages (Copilot, Aider, and others offer this).
- Testing
- Write unit and integration tests.
- AI can assist in generating test cases (Copilot Agents are particularly useful here). Tools like Aider can help refactor code to improve testability.
- Agent Frameworks
Random Links
https://twitter.com/tldraw/status/1782443204710674571
- Paper page Design2Code: How Far Are We From Automating Front-End Engineering? (huggingface.co)
- Generative AI Powered Assistant - Amazon Q - AWS Amazons!
- antworks.ai
- OpenBMB/ChatDev: Create Customized Software using Natural Language Idea (through LLM-powered Multi-Agent Collaboration) (github.com)
- Programming AIs worry me • Buttondown:
- Home | Tabby (tabbyml.com)
- The text discusses the concerns around using AI to generate code, specifically around the idea of proofreading the code. The author describes an experience with using voice-to-text where they found it difficult to proofread the text for errors. The text argues that using AI to generate code changes the work from writing code to proofreading code, and that this is a problem.
- Stop whining blog post
- blog post on LLMs for code
- Engshell shell LLM extension
- Github assist
- Locally run 13B coding optimised model
- Programming AIs worry me • Buttondown (other) The article discusses the ethical implications of using machine learning algorithms to generate art. While some see this as a powerful way to create new and interesting works of art, others worry about the potential for misuse and abuse of these technologies.
- GPT synthesizer
- Colab to get codey
- Build prompts using coding keywords, paper
- Continue for VSCode
- Phind technical answers and pair programmer with vscode plugin
- Starchat beta 4bit
- Sweep github pull requests to code system
- Cursor.so coding with gpt interface
- Code llama 2
- Long llama
- Open interpreter
- Open interpreter and autogen local tutorial
- open interpreter github
- codingbuddy
- deepseek 34b q4 AWQ
- Vercel provides front-end Infrastructure to allow developers to build fast, dynamic websites and applications efficiently at global scale. Its open source Next.js framework powers many leading AI products’ user interfaces.
- Vercel’s new vZero product allows developers to visually iterate on UIs with AI assistance.
- Demo/Tutorial: v0 by Vercel AI Code Generation (youtube.com)
- AI code auto-completion tools like Microsoft Copilot have shown the potential for AI to enhance software development. The latest Microsoft Copilot leverages Instruction-Following Conversational AI System 4 and is extremely good.
- AI will likely be incorporated into most software products going forward to enhance capabilities and engagement. Some experiences are better suited to standalone interfaces rather than cramming functionality into chatbots.
- Effective use of AI tools requires developing specialized skills around prompting, understanding system capabilities and limitations, and framing problems appropriately. Different AI systems have strengths in different domains.
- Software development will transition towards more hybrid human-AI teams, with less focus on writing code line-by-line. AI can provide significant productivity gains by automating rote tasks.
- There are open questions around whether to expose functionality through general chatbot interfaces vs company-specific products. There are strategic and technical considerations favouring bespoke solutions.
- Open source software tends to improve quickly over time and should not be underestimated. However, regulations could potentially suppress open source AI progress.
- gptengineer.app is a commercial offering built on GPT Engineer
- Understand a codebase in github with GPT
- Sourcegraph | Code AI platform
- [Bito AI
- Become a 10X Dev with Bito
- Bito](https://bito.ai/)
- Phind
Audio & Voice Generation
- Task: Create voiceovers, audio content like podcasts, or clone voices for various applications.
- ElevenLabs
- Description: High-quality text-to-speech and voice cloning AI. Offers a library of voices (Community Voices), can create a synthetic version of your own voice, and provide audio narration for websites (Audio Native). Used for audiobooks, podcasts, voiceovers.
- Cost: Free tier available. Paid plans based on character usage, starting around $5 USD/month.
- Website: ElevenLabs
- PlayHT
- Description: AI text-to-voice generator with a large library of voices and languages. Suitable for creating audiobooks, podcasts, and voiceovers.
- Cost: Free plan available. Paid plans based on word count/features, starting around $30 USD/month (billed annually).
- Website: PlayHT
VRChat
- none of this makes money yet
- This text is from wikipedia and will be updated when we have a chance totry VRChat properly. It’s much loved already by the Bitcoin community.
- “VRChat’s gameplay is similar to that of games such as Second Life andHabbo Hotel. Players can create their own instanced worlds in which theycan interact with each other through virtual avatars. A softwaredevelopment kit for Unity released alongside the game gives players theability to create or import character models to be used in the platform,as well as build their own worlds.
- Player models are capable of supporting “audio lip sync, eye trackingand blinking, and complete range of motion.
- VRChat is also capable of running in “desktop mode” without a VRheadset, which is controlled using either a mouse and keyboard, or agamepad. Some content has limitations in desktop mode, such as theinability to freely move an avatar’s limbs, or perform interactions thatrequire more than one hand.
- In 2020, a new visual programming language was introduced known as”Udon”, which uses a node graph system. While still considered alphasoftware, it became usable on publicly-accessible worlds beginning inApril 2020. A third-party compiler known as “UdonSharp” was developed toallow world scripts to be written in C sharp.”
Vircadia
- The applications and platforms detailed above have their benefits, butfor the application stack in the next section of the book Vircadia hasbeen chosen. The following text is from their website, and is aplaceholder which gives some idea. This section will be written outcompletely to reflect our use of the product to support emerging market users.
- Vircadia is open-source software which enables you to create and sharevirtual worlds as virtual reality (VR) and desktop experiences. You cancreate and host your own virtual world, explore other worlds, meet andconnect with other users, attend or host live VR events, and much more.
- The Vircadia metaverse provides built-in social features, includingavatar interactions, spatialized audio, and interactive physics.Additionally, you have the ability to import any 3D object into yourvirtual environment. No matter where you go in Vircadia, you will alwaysbe able to interact with your environment, engage with your friends, andlisten to conversations just like you would in real life.
Foundation Models
- Foundation models are large-scale, pre-trained models that can be adapted to a wide range of downstream tasks. They are trained on massive datasets of text and code and can be used for a variety of natural language processing (NLP) tasks, such as text generation, summarization, and question answering.
Animation: Breathing Life into Digital Characters
- Bringing digital characters to life requires compelling animation. This section explores projects and resources focused on achieving realistic and expressive character movement.
-
- Animatable Gaussians (GitHub Repository): Code for “Animatable Gaussians: Learning Pose-dependent Gaussian Maps for High-fidelity Human Avatar Modeling,” presented at CVPR 2024.
- 3D Gaussian Blendshapes: Exploration of the use of 3D Gaussian Blendshapes for head avatar animation.
- Eggnog: A platform for creating infinite AI videos.
- Efficient Portrait Animation (LivePortrait): A project focused on efficient portrait animation with stitching and retargeting control.
- Note by Daniele: A note on LivePortrait by Daniele.
- ComfyUI Nodes for LivePortrait (GitHub Repository): ComfyUI nodes designed for LivePortrait.
- CLARA2 (GitHub Repository): A 3D-rendered AI agent designed to present PowerPoint presentations.
- Consistent Character API (Replicate): An API for running fofr’s consistent character model.
- Joystick-Controlled Character Manipulation (Twitter): A concept for manipulating character features using a joystick.
- Volucap Authentic Digital Avatars: A company specialising in creating authentic digital avatars.
- Openpose Controlnets (Civitai Article): An article explaining how to use and generate poses with Openpose Controlnets V1.1.
- AI Modelling Agency (Creative Bloq Article): An article discussing the emergence of AI modelling agencies.
- Free VRChat Avatars & 3D Assets (VRCMods): A collection of free VRChat avatars and 3D assets.
- AniTalker: A project focused on animating talking heads.
- VLOGGER: A project related to creating virtual vloggers.
- StoryDiffusion: A project exploring consistent self-attention for long-range image and video generation.
- SoccerNet Game State Reconstruction (GitHub Repository): A project focusing on athlete tracking and identification on a minimap.
- Ukraine’s AI Avatar for Consular Affairs (Reddit Post): A discussion on Ukraine’s use of an AI avatar for consular updates.
- Animating Images with Viggle AI (YouTube Tutorial): A tutorial on animating images using Viggle AI.
- PhysAvatar (Hugging Face Paper): Research on learning the physics of dressed 3D avatars from visual observations.
- Vid2Avatar: A project focused on reconstructing 3D avatars from videos.
- Human Tracking and SLAM Capture (YouTube Video): A demonstration of human tracking and SLAM capture technology.
- ComfyUI Character Turntable with SV3D (Reddit Post): A discussion on creating character turntables using ComfyUI and SV3D.
- Animating Characters for Free (LinkedIn Post): A post highlighting methods for animating characters for free.
- Midjourney Character Reference Feature (Medium Article): An article exploring Midjourney’s Character Reference feature.
- Full-Character Consistency with SDXL (Reddit Post): A discussion on achieving full-character consistency using SDXL.
- Create Bot Emotions (Miku.gg Documentation): Documentation on creating bot emotions within the Miku.gg platform.
- Consistent Emotions on a Character with ComfyUI (Reddit User): A Reddit user’s plans to publish a method for achieving consistent emotions on a character using ComfyUI.
- BakedAvatar: A project focused on avatar creation.
- Dreamtalk (GitHub Repository): The official implementation of Dreamtalk, focusing on expressive talking head generation.
- CharTurnerBeta LoRA (Civitai): A LoRA for multi-direction consistency in Stable Diffusion character generation.
- VividTalk: A project focused on one-shot audio-driven talking head generation.
- Gaussian-Based Avatars (Hugging Face Papers):
- Relightable Gaussian Codec Avatars: Research on using Gaussian codecs for avatar representation.
- Gaussian Head Avatar: Research on creating high-fidelity head avatars using dynamic Gaussians.
- NLW Education Discord Projects (Discord Channel): A Discord channel discussing projects related to ElevenLabs and character/avatar creation.
- D-ID AI Video Mobile App: A mobile app for creating AI videos.
- GAIA (Microsoft): A project from Microsoft exploring advanced avatar technologies.
This meticulously curated collection offers a comprehensive overview of the dynamic field of digital human and avatar creation. Explore, learn, and contribute to the ongoing evolution of this exciting frontier!
Note: Some links may lead to projects or resources that are still under development or experimental. Remember to review any licensing information before using code or assets from these projects.
-
Logseq
- Logseq: is very similar to Obsidian, but self hosted and open source. It works on top of plain text files stored in a local system. It supports markdown and Org-mode formatting and allows for hierarchical and networked note-taking. It can be connected to it’s mobile app via github.
- Integration to Large Language Models can be OpenAI or local.
- Compare notion, obsidian, and logseq, using a simply markdown table with coloured dots
- ChatGPT Logseq Summarizer (openai.com)



Key LLM Papers
GPT-1 (2018): This paper introduces the first version of Generative Pre-trained Transformer (GPT), a generative model trained on a massive dataset of text. It demonstrates the ability of LLMs to generate coherent and grammatically correct text, paving the way for future advancements.
GPT-2 (2019): This paper presents a significantly larger GPT model with improved capabilities. It showcases the ability of LLMs to perform various language tasks, including text summarization, question answering, and even code generation.
GPT-3 (2020): This paper introduces GPT-3, a truly massive LLM with billions of parameters. It demonstrates impressive capabilities in diverse tasks, showcasing the emergence of general-purpose language abilities.
GPT-4 (2023): This paper introduces the latest iteration of GPT, featuring multi-modal capabilities and advanced reasoning abilities. It further pushes the boundaries of what LLMs can achieve, demonstrating impressive performance in a wide range of tasks.
Llama-2 (2023): This paper introduces Llama-2, a large language model designed with a focus on efficiency and accessibility. It offers a more resource-friendly alternative to other LLMs, making it more accessible for research and development.
Tools (2023): This paper introduces the “Tools” paradigm for LLMs, allowing them to interact with external tools and resources. It enables LLMs to perform more complex tasks by leveraging the power of external tools, expanding their capabilities significantly.
Gemini-Pro-1.5 (2023): This paper introduces Gemini-Pro-1.5, a large language model developed by Google. It showcases impressive capabilities in various tasks, including code generation, creative writing, and reasoning. It’s a strong contender in the race for developing advanced LLMs.
Generative Models for Molecule Design
- Generative models based on diffusion and flow-matching approaches enable fine-grained control over the generation of molecules with specific properties. ProteinDT and MoleculeSTM are examples of text-conditioned generative models that allow users to provide natural language prompts to generate molecules with desired properties.
- RF Diffusion, a diffusion model built on the RoseTTAFold backbone, offers powerful functionalities for protein engineering. It enables unconditional generation of novel proteins, binder design for high affinity and specificity, partial diffusion for refining existing structures, motif scaffolding for combining functional motifs, symmetric generation of protein complexes, and fold conditioning for generating proteins with specific tertiary structures.
- Complementary models like Ligand and PNN (Protein MPNN) are essential for designing amino acid sequences that fold into the desired 3D structures generated by RF Diffusion.
Some software choices
- It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.
- proprietary
- OpenAI’s Sora model represents a notable advancement in AI video generation. It demonstrates the ability to generate videos up to one minute in 1080p resolution and produce high-resolution images. Sora’s flexibility in handling various aspect ratios and resolutions indicates its adaptability in content creation. Its development leverages insights from prior research, including Vision Transformers and advanced training methodologies.
- Introduction to Sora
- A groundbreaking AI video generation model by OpenAI, Sora is designed to transform text instructions into realistic and imaginative video scenes, marking a significant advancement in creative AI technologies.
- Technical Overview
- Advanced Diffusion Model
- Employs a sophisticated diffusion process that starts from static noise and incrementally refines to generate high-resolution videos, showcasing an unparalleled leap in video realism and complexity.
- Transformer Architecture
- Leverages the Transformer model’s capabilities for deep understanding and generation of content, adapted here to interpret and create complex visual narratives, ensuring dynamic and coherent video storytelling.
- twitter link to the render loading below https://twitter.com/sainingxie/status/1758433676105310543
- twitter link to the render loading below https://twitter.com/thatguybg/status/1759935959792312461
- Patch-Based Data Representation
- Innovatively represents videos and images as collections of smaller data units, akin to language model tokens, enabling precise and granular control over video generation and editing.
- Advanced Diffusion Model
https://twitter.com/drjimfan/status/1758355737066299692?s=46
- Creative and Professional Applications
- Opens up endless possibilities for filmmakers, advertisers, educators, and content creators to produce cinema-quality visuals, educational materials, and immersive experiences effortlessly.
- Democratization of Video Production
- Simplifies the video creation process, enabling individuals and small teams to produce content that rivals big studio outputs.
- Enhancement of Creative Expression
- Allows creators to bring intricate visions and stories to life through simple text prompts, expanding visual storytelling horizons.
- Technical Insights
- Designed to scale language model capabilities to visual data, converting videos into patches for efficient processing and diverse video/image handling.
- Features a video compression network for temporal and spatial video compression, operating within a Neural Network Latent Space.
- Uses a diffusion transformer architecture, effectively scaling video generation and improving sample quality with increased compute.
- Innovative Features
- Works with videos at native sizes to offer sampling flexibility and improve composition and framing.
- Leverages descriptive captioning technique, enhancing video fidelity and quality from text prompts.
- Can animate still images and extend videos, including seamless interpolation between two videos, showcasing versatility.
- Emerging Capabilities
- Exhibits capabilities like 3D consistency, long-range coherence, object permanence, and world interaction simulation.
- Suggests potential as a tool for simulating physical and digital environments, aiding in the development of capable simulators.
- Videos can serve as a basis for constructing detailed 3D scenes using techniques like Neural Radiance Fields (NeRFs), potentially revolutionizing 3D content creation and interaction.
- Rapid prototyping and realization of 3D environments and narratives enhance VR and AR immersion and interactivity.
- Enables generation of characters, objects, and worlds through text and voice prompts, making 3D content creation more intuitive and accessible.
- Already being used to create 360 spherical video.
- Research and Discussion
- Video generation models as world simulators (openai.com) research paper highlights Sora’s technical foundation and its role in simulating the physical world.
- Discussions emphasize Sora’s potential in democratizing video creation and the need for granular output control for artistic purposes.
- Google DeepMind on X: “Introducing Veo: our most capable generative video model. 🎥 It can create high-quality, 1080p clips that can go beyond 60 seconds. From photorealism to surrealism and animation, it can tackle a range of cinematic styles. 🧵 GoogleIO https://t.co/6zEuYRAHpH” / X (twitter.com)
https://twitter.com/GoogleDeepMind/status/1790435824598716704
VideoPoet – Google Research
- Overview: Google’s text to video, linked to Bard, but not yet available.
- NVIDIA/NeMo: NeMo: a toolkit for conversational AI (github.com)
- [Canary
- NVIDIA NeMo](https://nvidia.github.io/NeMo/blogs/2024/2024-02-canary/)

- NeMo/tutorials/tts/FastPitch_Adapter_Finetuning.ipynb at main · NVIDIA/NeMo (github.com)
- ElevenLabs Audio Native
- OpenAI whisper local deploy
- realtime transciber
- high performance CPP
- 30% quantised optimisation
- Brillbits OpenAI whisper demo with mic
- Cleanvoice audio denoise
- Cloud voice change app
- downloadable voice generation systems
- Language AI open libraries
- Language practice
- MUGEN multi modal from facebook
- Oneshot speach to text
- Record and cleanup pro audio with commodity hardware
- Respeecher
- Voice AI voices
- Voice controlled assisted creation
- Voice to text, Lopp
- whisper transcriber
- Wolfram alpha voice chatbot integration
- Microsoft Vall-E voice synthesis
- Uberduck text to speech (plus own voice)
- Eleven labs language and text to speech
- Uberduck open source text to speech
- numen voice control system in linux
- Inworld (steam game plugin AI system) for voice chat and answer
- Bark text to speech from google labs
- https://github.com/TensorSpeech/TensorFlowTTS very configurable from what I see
- VoiceVox engine
- [coqui-ai TTS
- very good samples](https://github.com/coqui-ai/TTS)
- https://github.com/neonbjb/tortoise-tts
- https://github.com/CorentinJ/Real-Time-Voice-Cloning
- custom voices? looks neat
- https://github.com/rhasspy/larynx - very low-spec compatible, acceptable quality
- Voice cloning local
- Meta voicebox
- The Reddit post discusses the different open source voice cloning projects available, including Coqui, Tortoise, and Bark. The advantages and disadvantages of each project are briefly outlined, with ElevenLabs being noted as the best but not open source, while Tortoise is suggested as the closest open source alternative. Other tools for speech to speech and singing conversion, such as so-vits/diff-svc/rvc, are also mentioned. The post suggests that the quality of open source voice cloning projects is improving, and that there may be more options available in the future. https://www.reddit.com/r/MachineLearning/comments/133hanr/d_what_are_the_differences_between_the_major_open/
- The Retrieval-based Voice Conversion WebUI is a simple and useful voice conversion (voice changer) framework based on the VITS algorithm. It can use a small amount of voice data and still achieve good results. It incorporates a top-1 retrieval method to replace the source feature with the training set feature to avoid voice leakage, and it is easy to use with a simple web interface. It also features model fusion to change voice characteristics and the ability to integrate with the UVR5 model to quickly separate vocals and accompaniment. The project requires the installation of PyTorch and its core dependencies, and other pre-models are also needed for inference and training. The repository provides a guide to environment setup and usage, as well as links to relevant resources and contributors. https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI
- The article discusses different open-source voice cloning projects and their advantages and disadvantages. The projects mentioned include Coqui, Tortoise, and Bark, with the author highlighting Coqui’s unlocked platform, while Tortoise and Bark are newer transformer-based projects that can clone much more effectively with much less training and are restricted to prevent custom voice cloning. The author suggests that the ElevenLabs is currently the best voice cloning solution available, but it is not open source and can be expensive. The article also includes comments from other Reddit users, who suggest other open source options and provide additional insights into each option’s strengths and weaknesses. https://www.reddit.com/r/MachineLearning/comments/133hanr/d_what_are_the_differences_between_the_major_open/
- The article provides instructions on how to use OpenAI’s ChatGPT chatbot on an Android device using the Tasker app. The process involves importing a ChatGPT profile into Tasker, obtaining an API key from OpenAI, and setting up home screen shortcuts. The article also notes that ChatGPT can be run through Google Assistant with voice commands. The author suggests that while ChatGPT may not necessarily be better than Google Assistant, it can perform tasks that Google Assistant may not be capable of. https://www.howtogeek.com/882019/how-to-use-chatgpt-like-google-assistant-on-android/
- The Voice Assistant is an AI-powered chatbot that uses several APIs to understand natural language commands and provide helpful responses. It features a wide range of capabilities, including answering general knowledge questions, providing recommendations, performing productivity tasks, and entertaining users. The Voice Assistant was built using ChatGPT, Whisper API, Gradio, and Microsoft’s SpVoice TTS API, and it can be accessed through a web-based interface. The installation process involves cloning the repository and installing the required Python packages. Contributions to the project are welcome. https://github.com/DonGuillotine/chatGPT_whisper_AI_voice_assistant
- The Retrieval-based Voice Conversion WebUI is a voice conversion framework that uses a top-1 retrieval algorithm to eliminate voice leakage. It is capable of quickly training even on relatively poor GPUs and can achieve good results even with just 10 minutes of low noise voice data. It has a user-friendly web interface and the ability to use a model fusion system to change voice timbre. The setup recommends using Poetry and downloading the necessary pre-trained models from their Hugging Face space. It also includes additional files such as ffmpeg and ffprobe that may need to be downloaded. The WebUI can be initiated using the command “python infer-web.py” and Windows users can run the “go-web.bat” file. The project also acknowledges the contributions of related tools and libraries such as Gradio, HIFIGAN, and ContentVec. https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI
- VoicePen is a tool that uses AI to convert audio or video files into blog posts and transcriptions in minutes. The service includes a transcription and SRT file generated by a top speech-to-text model, an English blog post that pulls out key topics from the audio, and the ability to convert audio in 96 different languages. Use cases include repurposing podcasts, webinars, and tutorial videos. Monthly plans are available, with options for one-time conversions. Testimonials praise the accuracy and speed of VoicePen’s service. https://voicepen.ai
- Krisp is a software application designed to improve the productivity of online meetings by using AI-powered voice clarity and a meeting assistant to cancel background noise, echo, and accent localization. It works on both Mac and Windows platforms and processes only the user’s voice on their device, unlike other solutions that transmit voice over the internet. Krisp offers a free forever plan with no credit card required and is trusted by global brands. The insights gathered from calls can be viewed by the user to improve their communication skills over time. Krisp has received recognition from various prestigious awards such as America’s Most Promising AI Companies and has been awarded for its quality of support and ease of use. Krisp also offers SDK for developers, pricing and plans, and use cases such as contact centers and enterprise. The company prioritizes customers’ privacy, security and offers accessible support, including video tutorials and a help center. By accepting all cookies, users consent to the storing of cookies on their device to enhance site navigation, analyze site usage and assist in the company’s marketing efforts. https://krisp.ai/
- Cleanvoice AI is an artificial intelligence platform that assists users in editing their podcasts or audio recordings. The platform offers various features such as filler sound removal, mouth sound removal, stutter removal, and Deadair remover to make the audio recording more professional. Cleanvoice AI is multilingual and can detect filler sounds in multiple languages, including accents from various countries. The platform also allows for manual editing with assistance and offers tools like podcast mixing and background noise remover. Users can try Cleanvoice AI for free for 30 minutes without providing credit card details. However, users must accept the platform’s cookie policy to use the service. https://cleanvoice.ai/
- The article discusses the potential of Central Intelligent Agents (CIAs) and the role of large language models (LLMs) and other next-generation AI technologies in enabling them. It highlights the need for businesses to have a cross-functional team, ethical guidelines, and clear objectives in deploying their own CIA. The article also suggests steps to build a solid foundation for deploying a CIA, assess organizational readiness, assemble a cross-functional team, define objectives, develop the CIA components and evaluate its performance while continuing to learn and adapt. The author discusses the potential of AI tools and voice assistants in transforming the way businesses interact with their customers and suggests that the advent of advanced AI technologies has revolutionized the shift of businesses towards a more personalized and ethically responsible approach to engaging with their customers. Finally, the article ends by highlighting the importance of experimenting through crisis and providing expert guidance tailored to specific business needs. https://www.linkedin.com/pulse/central-intelligent-agent-enabling-next-generation-james-poulter?
- TensorSpeech/TensorFlowTTS: :stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German and Easy to adapt for other languages) Translation Accessibility Speech and Voice Speech and Voice
- Variety Speech and Voice Employment Social Contract Under Automation
- transcriptionstream/transcriptionstream: turnkey self-hosted offline transcription and diarization service with llm summary (github.com) Speech and Voice transcription locally RFC 2119 SHOULD Normative Keyword
- Tincans - Gazelle v0.2 Speech and Voice fast speech engine RFC 2119 SHOULD Normative Keyword
- Speech and Voice Open Voice (myshell.ai) cloning MIT license
- EndlessDreams: Voice directed real-time videos at 1280x1024 : r/StableDiffusion (reddit.com) Speech and Voice Speech and Voice Product Design Real Time
- https://demo.hume.ai/? Speech and Voice Large Language Models empathetic voice to voice
- Speech and Voice metavoiceio/metavoice-src: AI for human-level speech intelligence (github.com) check for PlayerTwo
- NeMo/tutorials/tts/NeMo_TTS_Primer.ipynb at main · NVIDIA/NeMo (github.com) NVIDIA Omniverse Platform Speech and Voice primer and demo.
Birme image resizer
- 2 hour tutorial
- inject your face into any model (dreambooth)
- Guide for dreambooth
- Shivram
- Progen photorealism Miro guide
- rare dreambooth tokens
- Multi subject tokens
- tag editor
- SDXL dreambooth
- Lora guide
- stable swarm distributed comfyui
- Textual inversion
- Img2Img guide from reddit for face mapping
- textual inversion cheaper training
- CIO blog post
- google stable diffusion
- Cross attention replace named items
- 256 x faster speedup
- VoltaML acceleration
- Depth map into blender from SD2
- midjourney tweaks
- and another
- Updates Pastebin
- Game development using SD
- Wildcard manager using ChatGPT
- Depth2Img for text
- train chat GPT to write prompts
- non destructive image manipulation using seeds
- Instruct pix2pix
- reddit post
- Attention heatmap for prompts (youtube)
- enormous link roundup
- Prompt master variations management
- panoramic world builder
- GitHub AbdullahAlfaraj/Auto-Photoshop-StableDiffusion-Plugin: A user-friendly plug-in that makes it easy to generate stable diffusion images inside Photoshop using Automatic1111-sd-webui as a backend.
- GitHub ashawkey/stable-dreamfusion: A pytorch implementation of text-to-3D dreamfusion, powered by stable diffusion.
- Fine tune stable diffusion
- GitHub Sanster/lama-cleaner: Image inpainting tool powered by SOTA AI Model. Remove any unwanted object, defect, people from your pictures or erase and replace(powered by stable diffusion) any thing on your pictures.
- holovolo immersive volumetric VR180 videos and photos, and 3D stable diffusion, for Quest and WebVR
- The Illustrated Stable Diffusion Jay Alammar Visualizing machine learning one concept at a time.
- reddit educational links
- Negative prompt hack tip
- Modify images with text
- Photorealism
- sdtools image v 1.6
- Character plugin
- Checkpoints
- Stability specific tools
- Arible Prompt Database https://www.arible.co/prompts
- [Guide] Make your own Loras, easy and free | Stable Diffusion Other | Civitai: You don’t need to download anything, this is a guide with online tools. Click “Show more” below.
- sdxl lora training
- dylora scripts
- kohya fork with scripts
- lora of loras (compressed sets)
- chart of print size aspect ratios
- SDXL native text lora
- SDXL lcm motion lora
- SDXL universal negative prompt
- text, watermark, low-quality, signature, moiré pattern, downsampling, aliasing, distorted, blurry, glossy, blur, jpeg artifacts, compression artifacts, poorly drawn, low-resolution, bad, distortion, twisted, excessive, exaggerated pose, exaggerated limbs, grainy, symmetrical, duplicate, error, pattern, beginner, pixelated, fake, hyper, glitch, overexposed, high-contrast, bad-contrast
- SDXL prodigy training guide
- Lora training interface for windows
- Refined model
- Fine tuning with captioning and other fine tuning tricks, followfox
- Negative embedding textual inversion for hands etc
- GitHub kpthedev/ez-text2video: Easily run text-to-video diffusion with customized video length, fps, and dimensions on 4GB video cards, as well as on CPU.
- Gligen grounding capability for sd1.5
- This repository contains a ComfyUI Extension for Automated Text Generation. The extension provides nodes which can be used to automate the text generation process. The goal is to build a node-based Automated Text Generation AGI. This extension should ultimately combine all of the features of the existing text generation tools into one tool.
- [R] Text-to-image Diffusion Models in Generative AI: A Survey: r/MachineLearning
- Tutorial: Creating a Consistent Character as a Textual Inversion Embedding
- Segment anything webui
- segment anything training
- Nvidia stable diffusion segment through clip
- Overriding iphone footage with SD characters using controlnet
- Interactive photo manipulation GAN
- 3d plugin for Automatic1111
- Face replace plugin for automatic
Renderings from Plan Drawings
- Vectorworks AI Visualizer (FAQ)
- Works inside Vectorworks 2024+, using your active file or view plus a text prompt.
- Ideal for quick concept iterations (materials, lighting variations).
- Note: not CAD-accurate rendering but excellent for inspirational visuals.
- Veras AI for Vectorworks (EvolveLAB announcement)
- Plugin that uses your 3D model or 2D viewport as a base.
- Photorealistic or stylised renders in seconds with prompt-driven material and ambience overrides.
- Mainstream Text-to-Image Generators
- Export plan or massing views as PNG/JPG and feed into Midjourney, Stable Diffusion (with ControlNet) or DALL·E 3 for high-res concept images.
- Best for early-stage mood boards rather than precise layouts.
ComfyTextures
- ComfyTextures GitHub - - ComfyTextures is a collection of free, high-quality textures designed for use in 3D rendering and other creative projects.
- The textures are organised into logical categories such as wood, metal, fabric, and stone, making it easier to find the desired material.
- Each texture comes with various maps (diffuse, normal, roughness, specular, height) to facilitate realistic material creation in different rendering engines.
- The textures are generally provided in a tileable format allowing for seamless repetition across surfaces.
- The repository is actively maintained, with additions and updates being made regularly, enhancing the available resource base.
- The textures can be downloaded and used for both commercial and non-commercial purposes under a specified licence.
- The repository aims to provide a valuable resource for artists and developers seeking readily accessible and customisable textures.
- Many textures include variations in colour and detail allowing for greater control over the final appearance.
Imagine 3D Software
- Imagine 3D - Luma Labs Imagine allows users to create realistic 3D models from text descriptions, streamlining the design workflow.
- It offers an intuitive interface to easily generate, edit and visualise 3D assets.
- Users can control the colour, texture, and shape of the generated 3D models using natural language processing.
- The tool enables users to iterate quickly on design ideas by making adjustments to the text prompt and regenerating the model.
- Imagine facilitates the creation of customised 3D models for various applications, including gaming, product visualisation and animation.
- The platform encourages experimentation with different prompts to explore the creative potential of artificial intelligence-powered 3D generation.
- This technology could be used for rapid prototyping, game development, and creation of virtual environments.
- GET3D aims to democratise 3D content creation by simplifying the process and reducing reliance on expert 3D modellers.
Physically Based Textures from BIM (Revit)

Evolution from Chat to Complex Systems
- Context engineering emerged as AI systems evolved beyond simple chat interfaces to incorporate:
- Function calling and tool use
- Retrieval augmented generation (RAG) systems
- Multi-agent workflows
- External API integrations
- The principle of “garbage in, garbage out” becomes critical when managing complex information flows. Pre-processing and cleaning data before it enters the context window significantly improves output quality.
Text-to-Speech
- Text-to-speech (TTS) technology can be used to convert written text into spoken audio. This can be used to create podcasts from blog posts, articles, or other written content.
Project Details
- Technology and Process
- Utilizing highly-efficient energy generation equipment, the project transforms methane, a natural landfill byproduct, into electricity.
- This electricity is used for several on-site applications, notably for powering data centers.
- Environmental and Economic Impacts
- Many U.S. landfills lack proper methane management systems.
- Recent studies suggest that landfill methane emissions might be significantly higher than previously estimated.
- Challenges in Traditional Energy Projects
- Traditional grid-connected landfill energy projects face high costs and long lead times.
- Over 70% of the U.S.’s approximately 2,600 municipal landfills lack a viable use for the methane they produce.
Stable Diffusion in Blender
- A Blender addon for using Stable Diffusion to render texture bakes for objects.
Dream Textures
- A Blender addon for applying textures with text prompts.
New AI Model Releases
- GPT4-x-Alpaca-13B-Native-4bit-128g: Technical discussions on the new model and its capabilities (GitHub Discussion).
Vircadia
- The applications and platforms detailed above have their benefits, butfor the application stack in the next section of the book Vircadia hasbeen chosen. The following text is from their website, and is aplaceholder which gives some idea. This section will be written outcompletely to reflect our use of the product to support emerging market users.
- Vircadia is open-source software which enables you to create and sharevirtual worlds as virtual reality (VR) and desktop experiences. You cancreate and host your own virtual world, explore other worlds, meet andconnect with other users, attend or host live VR events, and much more.
- The Vircadia metaverse provides built-in social features, includingavatar interactions, spatialized audio, and interactive physics.Additionally, you have the ability to import any 3D object into yourvirtual environment. No matter where you go in Vircadia, you will alwaysbe able to interact with your environment, engage with your friends, andlisten to conversations just like you would in real life.
LM Studio
- Integrates advanced tools like text-to-speech (TTS).
- Highly optimised for macOS environments.
- Link: Msty App
Logseq



- AI-Powered Search: Embedding-Based Retrieval and Retrieval-Augmented Generation (RAG) | by Daniel Tunkelang | Apr, 2024 | Medium
- AutoRAG documentation (marker-inc-korea.github.io)
- llmware-ai/llmware: Providing enterprise-grade LLM-based development framework, tools, and fine-tuned models. (github.com) Large Language Models Infrastructure Knowledge Graphing
- turbopuffer Knowledge Graphing serverless vector database
- Using agents over Knowledge Graphing Forget RAG: Embrace agent design for a more intelligent grounded ChatGPT! | by James Nguyen | Nov, 2023 | Medium
- Instruction-Following Conversational AI System threatens the Knowledge Graphing model with better capabilities Chat GPT 4 Turbo for Tech Leaders | Medium
- CLI tool to deploy a GPT model from a directory of data Knowledge Graphing
- VECTORDB open source Knowledge Graphing database
- https://nux.ai/guides/chaining-rag-systems Knowledge Graphing
- win4r/GraphRAG4OpenWebUI: GraphRAG4OpenWebUI integrates Microsoft’s GraphRAG technology into Open WebUI, providing a versatile information retrieval API. It combines local, global, and web searches for advanced Q&A systems and search engines. This tool simplifies graph-based retrieval integration in open web environments. (github.com) Open Webui and Pipelines Knowledge Graphing Knowledge Graphing
- Elicit search around Knowledge Graphing
- https://elicit.com/notebook/c4b29508-b134-429d-bda3-88a3b947375f
- For instance, this old and simple system
- https://elicit.com/notebook/c4b29508-b134-429d-bda3-88a3b947375f#17e74118b78497a92f941b07a460dd99
F o u n d a t i o n a l C o n c e p t s
GPT-1 (2018): This paper introduces the first version of Generative Pre-trained Transformer (GPT), a generative model trained on a massive dataset of text. It demonstrates the ability of LLMs to generate coherent and grammatically correct text, paving the way for future advancements.
GPT-2 (2019): This paper presents a significantly larger GPT model with improved capabilities. It showcases the ability of LLMs to perform various language tasks, including text summarization, question answering, and even code generation.
GPT-3 (2020): This paper introduces GPT-3, a truly massive LLM with billions of parameters. It demonstrates impressive capabilities in diverse tasks, showcasing the emergence of general-purpose language abilities.
GPT-4 (2023): This paper introduces the latest iteration of GPT, featuring multi-modal capabilities and advanced reasoning abilities. It further pushes the boundaries of what LLMs can achieve, demonstrating impressive performance in a wide range of tasks.
Llama-2 (2023): This paper introduces Llama-2, a large language model designed with a focus on efficiency and accessibility. It offers a more resource-friendly alternative to other LLMs, making it more accessible for research and development.
Tools (2023): This paper introduces the “Tools” paradigm for LLMs, allowing them to interact with external tools and resources. It enables LLMs to perform more complex tasks by leveraging the power of external tools, expanding their capabilities significantly.
Gemini-Pro-1.5 (2023): This paper introduces Gemini-Pro-1.5, a large language model developed by Google. It showcases impressive capabilities in various tasks, including code generation, creative writing, and reasoning. It’s a strong contender in the race for developing advanced LLMs.
Agents in Biological Research
- AI agents have the potential to transform biological research by automating tasks such as literature review, hypothesis generation, experimental design, and data analysis. Companies like Future House are developing AI agents that can identify potential drug targets and design experiments, significantly accelerating the process of discovery. These agents, powered by large language models (LLMs) and other AI technologies, can review thousands of research papers, develop targets or hypotheses to test, and even drive autonomous labs.
- As these AI agents become more capable, they may play a crucial role in guiding research and helping humans navigate the complex landscape of biological data and interactions. The convergence of AI agents with specific tools for designing molecules, proteins, and nucleic acids could lead to rapid progress in solving challenging problems in biology and medicine.
Some software choices
- It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.
VideoPoet – Google Research
- Overview: Google’s text to video, linked to Bard, but not yet available.
Text2Mesh
- The code is organised in a modular fashion, allowing for easy customisation and extension of the system.
- The repository contains detailed instructions on how to set up the environment, download necessary models, and run the text-to-mesh pipeline.
- Users can adjust parameters to control the style, complexity, and colour of the generated 3D meshes.
- The project highlights the potential of automation to simplify 3D content creation and make it more accessible to a wider audience.
Text-to-Speech
- Text-to-speech (TTS) technology can be used to convert written text into spoken audio. This can be used to create podcasts from blog posts, articles, or other written content.
Dream Textures
- A Blender addon for applying textures with text prompts.
- Stable Diffusion Image Model
Some software choices
- It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.
Overview
- Imagine being able to verbally command a virtual design software to create specific CAD primitives or modify existing models. Additionally, the ability to add text annotations or descriptions directly within the virtual space can facilitate collaboration and communication among users.
- Furthermore, as corporate metaverse like NVIDIA Omniverse Platform expands, the shared virtual spaces will become increasingly complex and vast, accommodating a multitude of digital twin models. This means that users will be able to explore and interact with realistic replicas of real-world objects and environments, such as buildings, vehicles, or even entire cities.
- By incorporating voice and text input functionalities, developers can empower users to manipulate and navigate these digital twin models more intuitively. Whether it’s adjusting the dimensions of a virtual prototype or performing intricate measurements, the metaverse’s ability to recognize and respond to voice and text commands will revolutionize the way we design, simulate, and experience virtual environments.
- Table Of Contents — bd_warehouse “0.1.0” # Uncomment this for the next release? documentation (bd-warehouse.readthedocs.io)
- [Latest General topics
Some software choices
- It is possible at this stage to put more flesh on the bones through example software stack choices. Such specificity likely introduces overlaps, technical challenges, and contradictions, but has been generated in the main by GenAI based on the wider corpus of text and demonstrates the direction of travel well.
Text to Multiview and Texturing
Multi-Modal Large Language Models (LLMs)
- Introduction:
- Large Language Models are adept at generating coherent text sequences, predicting word probabilities and co-occurrences.
April 2024
- 1 Apr, Do Language Models Plan Ahead for Future Tokens?, https://arxiv.org/abs/2404.00859
- 1 Apr, Bigger is not Always Better: Scaling Properties of Latent Diffusion Models, https://arxiv.org/abs/2404.01367
- 1 Apr, The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis, https://arxiv.org/abs/2404.01204
- 1 Apr, Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models, https://arxiv.org/abs/2404.04478
- 2 Apr, Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models, https://arxiv.org/abs/2404.02258
- 2 Apr, Long-context LLMs Struggle with Long In-context Learning, https://arxiv.org/abs/2404.02060
- 2 Apr, Emergent Abilities in Reduced-Scale Generative Language Models, https://arxiv.org/abs/2404.02204
- 2 Apr, Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, https://arxiv.org/abs/2404.02151
- 3 Apr, On the Scalability of Diffusion-based Text-to-Image Generation, https://arxiv.org/abs/2404.02883
- 3 Apr, BAdam: A Memory Efficient Full Parameter Training Method for Large Language Models, https://arxiv.org/abs/2404.02827
- 3 Apr, Cross-Attention Makes Inference Cumbersome in Text-to-Image Diffusion Models, https://arxiv.org/abs/2404.02747
- 4 Apr, Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences, https://arxiv.org/abs/2404.02151
- 4 Apr, Training LLMs over Neurally Compressed Text, https://arxiv.org/abs/2404.03626
- 4 Apr, CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues, https://arxiv.org/abs/2404.03820
- 5 Apr, ReFT: Representation Finetuning for Language Models, https://arxiv.org/abs/2404.03592
- 5 Apr, Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data, https://arxiv.org/abs/2404.03862
- 5 Apr, Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation, https://arxiv.org/abs/2404.04256
- 8 Apr, AutoCodeRover: Autonomous Program Improvement, https://arxiv.org/abs/2404.05427
- 8 Apr, Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence, https://arxiv.org/abs/2404.05892
- 8 Apr, CodecLM: Aligning Language Models with Tailored Synthetic Data, https://arxiv.org/abs/2404.05875
- 9 Apr, MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies, https://arxiv.org/abs/2404.06395
- 9 Apr, Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models, https://arxiv.org/abs/2404.06209
- 9 Apr, LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders, https://arxiv.org/abs/2404.05961
- 10 Apr, Adapting LLaMA Decoder to Vision Transformer, https://arxiv.org/abs/2404.06773
- 10 Apr, Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention, https://arxiv.org/abs/2404.07143
- 11 Apr, LLoCO: Learning Long Contexts Offline, https://arxiv.org/abs/2404.07979
- 11 Apr, JetMoE: Reaching Llama2 Performance with 0.1M Dollars, https://arxiv.org/abs/2404.07413
- 11 Apr, Best Practices and Lessons Learned on Synthetic Data for Language Models, https://arxiv.org/abs/2404.07503
- 11 Apr, Rho-1: Not All Tokens Are What You Need, https://arxiv.org/abs/2404.07965
- 12 Apr, Pre-training Small Base LMs with Fewer Tokens, https://arxiv.org/abs/2404.08634
- 12 Apr, Dataset Reset Policy Optimization for RLHF, https://arxiv.org/abs/2404.08495
- 13 Apr, LLM In-Context Recall is Prompt Dependent, https://arxiv.org/abs/2404.08865
- 15 Apr, State Space Model for New-Generation Network Alternative to Transformers: A Survey, https://arxiv.org/abs/2404.09516
- 15 Apr, Chinchilla Scaling: A Replication Attempt, https://arxiv.org/abs/2404.10102
- 15 Apr, Learn Your Reference Model for Real Good Alignment, https://arxiv.org/abs/2404.09656
- 16 Apr, Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study, https://arxiv.org/abs/2404.10719
- 16 Apr, Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies, https://arxiv.org/abs/2404.08197
- 16 Apr, How Faithful Are RAG Models? Quantifying the Tug-of-War Between RAG and LLMs’ Internal Prior, https://arxiv.org/abs/2404.10198
- 17 Apr, A Survey on Retrieval-Augmented Text Generation for Large Language Models, https://arxiv.org/abs/2404.10981
- 18 Apr, When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes, https://arxiv.org/abs/2404.12365
- 2 Jun, Show, Don’t Tell: Aligning Language Models with Demonstrated Feedback, https://arxiv.org/abs/2406.00888
- 3 Jun, Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models, https://arxiv.org/abs/2406.06563
- 3 Jun, OLoRA: Orthonormal Low-Rank Adaptation of Large Language Models, https://arxiv.org/abs/2406.01775
- 3 Jun, The Geometry of Categorical and Hierarchical Concepts in Large Language Models, https://arxiv.org/abs/2406.01506
- 3 Jun, Towards Scalable Automated Alignment of LLMs: A Survey, https://arxiv.org/abs/2406.01252
- 4 Jun, Scalable MatMul-free Language Modeling, https://arxiv.org/abs/2406.02528
- 4 Jun, Block Transformer: Global-to-Local Language Modeling for Fast Inference, https://arxiv.org/abs/2406.02657
- 6 Jun, Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models, https://arxiv.org/abs/2406.04271
- 6 Jun, The Prompt Report: A Systematic Survey of Prompting Techniques, https://arxiv.org/abs/2406.06608
- 6 Jun, Transformers Need Glasses! Information Over-Squashing in Language Tasks, https://arxiv.org/abs/2406.04267
- 6 Jun, Are We Done with MMLU?, https://arxiv.org/abs/2406.04127
- 6 Jun, Step-aware Preference Optimization: Aligning Preference with Denoising Performance at Each Step, https://arxiv.org/abs/2406.04314
- 7 Jun, Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven Approach, https://arxiv.org/abs/2406.04594
- 4 Nov, “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM Quantization, https://arxiv.org/abs/2411.02355
- 4 Nov, Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study, https://arxiv.org/abs/2411.02462
- 5 Nov, HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems, https://arxiv.org/abs/2411.02959
- 6 Nov, Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination, https://arxiv.org/abs/2411.03823
- 6 Nov, Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding, https://arxiv.org/abs/2411.04282
- 6 Nov, Number Cookbook: Number Understanding of Language Models and How to Improve It, https://arxiv.org/abs/2411.03766
- 7 Nov, Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models, https://arxiv.org/abs/2411.04996
- 7 Nov, BitNet a4.8: 4-bit Activations for 1-bit LLMs, https://arxiv.org/abs/2411.04965
- 7 Nov, Scaling Laws for Precision, https://arxiv.org/abs/2411.04330
- 8 Nov, Energy Efficient Protein Language Models: Leveraging Small Language Models with LoRA for Controllable Protein Generation, https://arxiv.org/abs/2411.05966
- 8 Nov, Balancing Pipeline Parallelism with Vocabulary Parallelism, https://arxiv.org/abs/2411.05288
- 11 Nov, Toward Optimal Search and Retrieval for RAG, https://arxiv.org/abs/2411.07396
- 12 Nov, Large Language Models Can Self-Improve in Long-context Reasoning, https://arxiv.org/abs/2411.08147
- 12 Nov, Stronger Models are NOT Stronger Teachers for Instruction Tuning, https://arxiv.org/abs/2411.07133
- 12 Nov, Direct Preference Optimization Using Sparse Feature-Level Constraints, https://arxiv.org/abs/2411.07618
- 13 Nov, Cut Your Losses in Large-Vocabulary Language Models, https://arxiv.org/abs/2411.09009
- 15 Nov, Does Prompt Formatting Have Any Impact on LLM Performance?, https://arxiv.org/abs/2411.10541
- 17 Nov, SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization, https://arxiv.org/abs/2411.11909
- 17 Nov, SageAttention2 Technical Report: Accurate 4 Bit Attention for Plug-and-play Inference Acceleration, https://arxiv.org/abs/2411.10958
- 18 Nov, Bi-Mamba: Towards Accurate 1-Bit State Space Models, https://arxiv.org/abs/2411.11843
- 19 Nov, RedPajama: an Open Dataset for Training Large Language Models, https://arxiv.org/abs/2411.12372
- 20 Nov, Hymba: A Hybrid-head Architecture for Small Language Models, https://arxiv.org/abs/2411.13676
- 20 Nov, Loss-to-Loss Prediction: Scaling Laws for All Datasets, https://arxiv.org/abs/2411.12925
- 21 Nov, When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training, https://arxiv.org/abs/2411.13476
Text to Multiview and Texturing
November 2024
- 1 Nov, Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations, https://arxiv.org/abs/2411.00640
- 1 Nov 2024, Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation, https://arxiv.org/abs/2411.00412
- 1 Nov 2024, Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models, https://arxiv.org/abs/2411.00492
- 3 Nov, Sample-Efficient Alignment for LLMs, https://arxiv.org/abs/2411.01493
- 4 Nov 2024, A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness, https://arxiv.org/abs/2411.03350
- 4 Nov, “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM Quantization, https://arxiv.org/abs/2411.02355
- 4 Nov, Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study, https://arxiv.org/abs/2411.02462
- 5 Nov, HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems, https://arxiv.org/abs/2411.02959
- 6 Nov, Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination, https://arxiv.org/abs/2411.03823
- 6 Nov, Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding, https://arxiv.org/abs/2411.04282
- 6 Nov, Number Cookbook: Number Understanding of Language Models and How to Improve It, https://arxiv.org/abs/2411.03766
- 7 Nov, Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models, https://arxiv.org/abs/2411.04996
- 7 Nov, BitNet a4.8: 4-bit Activations for 1-bit LLMs, https://arxiv.org/abs/2411.04965
- 7 Nov, Scaling Laws for Precision, https://arxiv.org/abs/2411.04330
- 8 Nov, Energy Efficient Protein Language Models: Leveraging Small Language Models with LoRA for Controllable Protein Generation, https://arxiv.org/abs/2411.05966
- 8 Nov, Balancing Pipeline Parallelism with Vocabulary Parallelism, https://arxiv.org/abs/2411.05288
- 11 Nov, Toward Optimal Search and Retrieval for RAG, https://arxiv.org/abs/2411.07396
- 12 Nov, Large Language Models Can Self-Improve in Long-context Reasoning, https://arxiv.org/abs/2411.08147
- 12 Nov, Stronger Models are NOT Stronger Teachers for Instruction Tuning, https://arxiv.org/abs/2411.07133
- 12 Nov, Direct Preference Optimization Using Sparse Feature-Level Constraints, https://arxiv.org/abs/2411.07618
- 13 Nov, Cut Your Losses in Large-Vocabulary Language Models, https://arxiv.org/abs/2411.09009
- 15 Nov, Does Prompt Formatting Have Any Impact on LLM Performance?, https://arxiv.org/abs/2411.10541
- 17 Nov, SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization, https://arxiv.org/abs/2411.11909
- 17 Nov, SageAttention2 Technical Report: Accurate 4 Bit Attention for Plug-and-play Inference Acceleration, https://arxiv.org/abs/2411.10958
- 18 Nov, Bi-Mamba: Towards Accurate 1-Bit State Space Models, https://arxiv.org/abs/2411.11843
- 19 Nov, RedPajama: an Open Dataset for Training Large Language Models, https://arxiv.org/abs/2411.12372
- 20 Nov, Hymba: A Hybrid-head Architecture for Small Language Models, https://arxiv.org/abs/2411.13676
- 20 Nov, Loss-to-Loss Prediction: Scaling Laws for All Datasets, https://arxiv.org/abs/2411.12925
- 21 Nov, When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training, https://arxiv.org/abs/2411.13476
April 2024
- 1 Apr, Do Language Models Plan Ahead for Future Tokens?, https://arxiv.org/abs/2404.00859
- 1 Apr, Bigger is not Always Better: Scaling Properties of Latent Diffusion Models, https://arxiv.org/abs/2404.01367
- 1 Apr, The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis, https://arxiv.org/abs/2404.01204
- 1 Apr, Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models, https://arxiv.org/abs/2404.04478
- 2 Apr, Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models, https://arxiv.org/abs/2404.02258
- 2 Apr, Long-context LLMs Struggle with Long In-context Learning, https://arxiv.org/abs/2404.02060
- 2 Apr, Emergent Abilities in Reduced-Scale Generative Language Models, https://arxiv.org/abs/2404.02204
- 2 Apr, Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, https://arxiv.org/abs/2404.02151
- 3 Apr, On the Scalability of Diffusion-based Text-to-Image Generation, https://arxiv.org/abs/2404.02883
- 3 Apr, BAdam: A Memory Efficient Full Parameter Training Method for Large Language Models, https://arxiv.org/abs/2404.02827
- 3 Apr, Cross-Attention Makes Inference Cumbersome in Text-to-Image Diffusion Models, https://arxiv.org/abs/2404.02747
- 4 Apr, Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences, https://arxiv.org/abs/2404.02151
- 4 Apr, Training LLMs over Neurally Compressed Text, https://arxiv.org/abs/2404.03626
- 4 Apr, CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues, https://arxiv.org/abs/2404.03820
- 5 Apr, ReFT: Representation Finetuning for Language Models, https://arxiv.org/abs/2404.03592
- 5 Apr, Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data, https://arxiv.org/abs/2404.03862
- 5 Apr, Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation, https://arxiv.org/abs/2404.04256
- 8 Apr, AutoCodeRover: Autonomous Program Improvement, https://arxiv.org/abs/2404.05427
- 8 Apr, Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence, https://arxiv.org/abs/2404.05892
- 8 Apr, CodecLM: Aligning Language Models with Tailored Synthetic Data, https://arxiv.org/abs/2404.05875
- 9 Apr, MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies, https://arxiv.org/abs/2404.06395
- 9 Apr, Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models, https://arxiv.org/abs/2404.06209
- 9 Apr, LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders, https://arxiv.org/abs/2404.05961
- 10 Apr, Adapting LLaMA Decoder to Vision Transformer, https://arxiv.org/abs/2404.06773
- 10 Apr, Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention, https://arxiv.org/abs/2404.07143
- 11 Apr, LLoCO: Learning Long Contexts Offline, https://arxiv.org/abs/2404.07979
- 11 Apr, JetMoE: Reaching Llama2 Performance with 0.1M Dollars, https://arxiv.org/abs/2404.07413
- 11 Apr, Best Practices and Lessons Learned on Synthetic Data for Language Models, https://arxiv.org/abs/2404.07503
- 11 Apr, Rho-1: Not All Tokens Are What You Need, https://arxiv.org/abs/2404.07965
- 12 Apr, Pre-training Small Base LMs with Fewer Tokens, https://arxiv.org/abs/2404.08634
- 12 Apr, Dataset Reset Policy Optimization for RLHF, https://arxiv.org/abs/2404.08495
- 13 Apr, LLM In-Context Recall is Prompt Dependent, https://arxiv.org/abs/2404.08865
- 15 Apr, State Space Model for New-Generation Network Alternative to Transformers: A Survey, https://arxiv.org/abs/2404.09516
- 15 Apr, Chinchilla Scaling: A Replication Attempt, https://arxiv.org/abs/2404.10102
- 15 Apr, Learn Your Reference Model for Real Good Alignment, https://arxiv.org/abs/2404.09656
- 16 Apr, Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study, https://arxiv.org/abs/2404.10719
- 16 Apr, Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies, https://arxiv.org/abs/2404.08197
- 16 Apr, How Faithful Are RAG Models? Quantifying the Tug-of-War Between RAG and LLMs’ Internal Prior, https://arxiv.org/abs/2404.10198
- 17 Apr, A Survey on Retrieval-Augmented Text Generation for Large Language Models, https://arxiv.org/abs/2404.10981
- 18 Apr, When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes, https://arxiv.org/abs/2404.12365
- 2 Jun, Show, Don’t Tell: Aligning Language Models with Demonstrated Feedback, https://arxiv.org/abs/2406.00888
- 3 Jun, Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models, https://arxiv.org/abs/2406.06563
- 3 Jun, OLoRA: Orthonormal Low-Rank Adaptation of Large Language Models, https://arxiv.org/abs/2406.01775
- 3 Jun, The Geometry of Categorical and Hierarchical Concepts in Large Language Models, https://arxiv.org/abs/2406.01506
- 3 Jun, Towards Scalable Automated Alignment of LLMs: A Survey, https://arxiv.org/abs/2406.01252
- 4 Jun, Scalable MatMul-free Language Modeling, https://arxiv.org/abs/2406.02528
- 4 Jun, Block Transformer: Global-to-Local Language Modeling for Fast Inference, https://arxiv.org/abs/2406.02657
- 6 Jun, Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models, https://arxiv.org/abs/2406.04271
- 6 Jun, The Prompt Report: A Systematic Survey of Prompting Techniques, https://arxiv.org/abs/2406.06608
- 6 Jun, Transformers Need Glasses! Information Over-Squashing in Language Tasks, https://arxiv.org/abs/2406.04267
- 6 Jun, Are We Done with MMLU?, https://arxiv.org/abs/2406.04127
- 6 Jun, Step-aware Preference Optimization: Aligning Preference with Denoising Performance at Each Step, https://arxiv.org/abs/2406.04314
- 7 Jun, Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven Approach, https://arxiv.org/abs/2406.04594
- 7 Jun, CRAG — Comprehensive RAG Benchmark, https://arxiv.org/abs/2406.04744
- 7 Jun, WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild, https://arxiv.org/abs/2406.04770
- 7 Jun, Mixture-of-Agents Enhances Large Language Model Capabilities, https://arxiv.org/abs/2406.04692
- 7 Jun, BERTs are Generative In-Context Learners, https://arxiv.org/abs/2406.04823
- 7 Jun, 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination, https://arxiv.org/abs/2406.05132
- 8 Jun, Creativity Has Left the Chat: The Price of Debiasing Language Models, https://arxiv.org/abs/2406.05587
- 10 Jun, Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation, https://arxiv.org/abs/2406.06525
- 10 Jun, Margin-aware Preference Optimization for Aligning Diffusion Models Without Reference, https://arxiv.org/abs/2406.06424
- 10 Jun, Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning, https://arxiv.org/abs/2406.06469
- 10 Jun, Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters, https://arxiv.org/abs/2406.05955
- 10 Jun, Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching, https://arxiv.org/abs/2406.06326
- 11 Jun, An Image is Worth 32 Tokens for Reconstruction and Generation, https://arxiv.org/abs/2406.07550
- 11 Jun, TextGrad: Automatic “Differentiation” via Text, https://arxiv.org/abs/2406.07496
- 11 Jun, Simple and Effective Masked Diffusion Language Models, https://arxiv.org/abs/2406.07524
- 11 Jun, Never Miss A Beat: An Efficient Recipe for Context Window Extension of Large Language Models with Consistent “Middle” Enhancement, https://arxiv.org/abs/2406.07138
- 11 Jun, Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling, https://arxiv.org/abs/2406.07522
- 12 Jun, Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing, https://arxiv.org/abs/2406.08464
- 12 Jun, What If We Recaption Billions of Web Images with LLaMA-3?, https://arxiv.org/abs/2406.08478
- 12 Jun, Large Language Model Unlearning via Embedding-Corrupted Prompts, https://arxiv.org/abs/2406.07933
- 12 Jun, Large Language Models Must Be Taught to Know What They Don’t Know, https://arxiv.org/abs/2406.08391
- 12 Jun, An Empirical Study of Mamba-based Language Models, https://arxiv.org/abs/2406.07887
- 12 Jun, Discovering Preference Optimization Algorithms with and for Large Language Models, https://arxiv.org/abs/2406.08414
- 13 Jun, Transformers Meet Neural Algorithmic Reasoners, https://arxiv.org/abs/2406.09308
- 13 Jun, MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding, https://arxiv.org/abs/2406.09297
- 13 Jun, An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels, https://arxiv.org/abs/2406.09415
- 13 Jun, FouRA: Fourier Low Rank Adaptation, https://arxiv.org/abs/2406.08798
- 14 Jun, Bootstrapping Language Models with DPO Implicit Rewards, https://arxiv.org/abs/2406.09760
- 14 Jun, Be like a Goldfish, Don’t Memorize! Mitigating Memorization in Generative LLMs, https://arxiv.org/abs/2406.10209
- 14 Jun, Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs, https://arxiv.org/abs/2406.10216
- 16 Jun, THEANINE: Revisiting Memory Management in Long-term Conversations with Timeline-augmented Response Generation, https://arxiv.org/abs/2406.10996
- 17 Jun, Task Me Anything, https://arxiv.org/abs/2406.11775
- 17 Jun, How Do Large Language Models Acquire Factual Knowledge During Pretraining?, https://arxiv.org/abs/2406.11813
- 17 Jun, mDPO: Conditional Preference Optimization for Multimodal Large Language Models, https://arxiv.org/abs/2406.11839
- 17 Jun, Nemotron-4 340B Technical Report, https://arxiv.org/abs/2406.11704
- 17 Jun, DataComp-LM: In Search of the Next Generation of Training Sets for Language Models, https://arxiv.org/abs/2406.11794
- 17 Jun, Tokenization Falling Short: The Curse of Tokenization, https://arxiv.org/abs/2406.11687
- 17 Jun, DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence, https://arxiv.org/abs/2406.11931
- 17 Jun, Unveiling Encoder-Free Vision-Language Models, https://arxiv.org/abs/2406.11832
- 17 Jun, Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level, https://arxiv.org/abs/2406.11817
- 17 Jun, HARE: HumAn pRiors, a key to small language model Efficiency, https://arxiv.org/abs/2406.11410
- 17 Jun, Measuring memorization in RLHF for code completion, https://arxiv.org/abs/2406.11715
- 18 Jul, Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies, https://arxiv.org/abs/2407.13623
- 19 Jul, BOND: Aligning LLMs with Best-of-N Distillation, https://arxiv.org/abs/2407.14622
- 19 Jul, Compact Language Models via Pruning and Knowledge Distillation, https://arxiv.org/abs/2407.14679
- 19 Jul, LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference, https://arxiv.org/abs/2407.14057
- 22 Jul, Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training, https://arxiv.org/abs/2407.15892
- 22 Jul, DDK: Distilling Domain Knowledge for Efficient Large Language Models, https://arxiv.org/abs/2407.16154
- 23 Jul, Generation Constraint Scaling Can Mitigate Hallucination, https://arxiv.org/abs/2407.16908
- 23 Jul, Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach, https://arxiv.org/abs/2407.16833
- 23 Jul, Course-Correction: Safety Alignment Using Synthetic Preferences, https://arxiv.org/abs/2407.16637
- 26 Jul, Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?, https://arxiv.org/abs/2407.16607
- 28 Jul, Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge, https://arxiv.org/abs/2407.19594
- 29 Jul, Improving Retrieval Augmented Language Model with Self-Reasoning, https://arxiv.org/abs/2407.19813
- 29 Jul, Apple Intelligence Foundation Language Models, https://arxiv.org/abs/2407.21075
- 30 Jul, ThinK: Thinner Key Cache by Query-Driven Pruning, https://arxiv.org/abs/2407.21018
- 31 Jul, The Llama 3 Herd of Models, https://arxiv.org/abs/2407.21783
- 31 Jul, Gemma 2: Improving Open Language Models at a Practical Size, https://arxiv.org/abs/2408.00118
Emerging use cases (AI holding text)
Luma Dream Machine?
- Luma Dream Machine is a browser-based AI video generator developed by Luma Labs, a San Francisco-based startup. It allows users to generate short videos (around 5 seconds) by simply entering a text or image prompt.
- Free to Use: Luma Dream Machine is free to try, with no waiting list or subscription required. Users get 30 free video generations per month.
- High-Quality Output: The AI produces impressively clean and detailed videos, adhering to prompts accurately and generating relatively coherent motion.
- Fast Generation: Videos are generated in around 2 minutes after entering the prompt.
- Consistent Subjects: Characters and subjects appear consistent throughout the video, capable of expressing emotion better than many previous AI video models.
- Difficulty with complex scenes or full-body shots
- Text in videos may appear garbled
- Anatomical issues like extra limbs or heads
- Meta’s Approach: Foundational World Modeling Meta (formerly Facebook) is taking a distinct approach, focusing on the underlying world modeling needed for video encoding and generation. This emphasis on understanding the principles of physics and object interactions could contribute to more realistic AI-generated videos.
- Technical Capabilities and Limitations
- Capabilities Current AI video generators demonstrate proficiency in producing high-resolution images and videos. They are capable of style adaptation, simulating complex scenes with multiple elements, and handling variations in aspect ratio and resolution.
- Limitations Despite their strengths, these models still struggle to accurately simulate physics and lack a complete understanding of cause and effect. Occasional errors regarding object permanence highlight the existing gap between pattern recognition and a comprehensive understanding of the world.
- Ethical and Creative Considerations
- Potential Impacts Advancements in AI video generation raise questions about the future of creative professions and the ethical implications of AI-generated content. Balancing technological innovation with safeguarding the integrity of human creativity is an important consideration.
Texturing
Text to Multiview and Texturing
November 2024
- 1 Nov, Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations, https://arxiv.org/abs/2411.00640
- 1 Nov 2024, Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation, https://arxiv.org/abs/2411.00412
- 1 Nov 2024, Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models, https://arxiv.org/abs/2411.00492
- 3 Nov, Sample-Efficient Alignment for LLMs, https://arxiv.org/abs/2411.01493
- 4 Nov 2024, A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness, https://arxiv.org/abs/2411.03350
- 4 Nov, “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM Quantization, https://arxiv.org/abs/2411.02355
- 4 Nov, Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study, https://arxiv.org/abs/2411.02462
- 5 Nov, HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems, https://arxiv.org/abs/2411.02959
- 6 Nov, Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination, https://arxiv.org/abs/2411.03823
- 6 Nov, Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding, https://arxiv.org/abs/2411.04282
- 6 Nov, Number Cookbook: Number Understanding of Language Models and How to Improve It, https://arxiv.org/abs/2411.03766
- 7 Nov, Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models, https://arxiv.org/abs/2411.04996
- 7 Nov, BitNet a4.8: 4-bit Activations for 1-bit LLMs, https://arxiv.org/abs/2411.04965
- 7 Nov, Scaling Laws for Precision, https://arxiv.org/abs/2411.04330
- 8 Nov, Energy Efficient Protein Language Models: Leveraging Small Language Models with LoRA for Controllable Protein Generation, https://arxiv.org/abs/2411.05966
- 8 Nov, Balancing Pipeline Parallelism with Vocabulary Parallelism, https://arxiv.org/abs/2411.05288
- 11 Nov, Toward Optimal Search and Retrieval for RAG, https://arxiv.org/abs/2411.07396
- 12 Nov, Large Language Models Can Self-Improve in Long-context Reasoning, https://arxiv.org/abs/2411.08147
- 12 Nov, Stronger Models are NOT Stronger Teachers for Instruction Tuning, https://arxiv.org/abs/2411.07133
- 12 Nov, Direct Preference Optimization Using Sparse Feature-Level Constraints, https://arxiv.org/abs/2411.07618
- 13 Nov, Cut Your Losses in Large-Vocabulary Language Models, https://arxiv.org/abs/2411.09009
- 15 Nov, Does Prompt Formatting Have Any Impact on LLM Performance?, https://arxiv.org/abs/2411.10541
- 17 Nov, SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization, https://arxiv.org/abs/2411.11909
- 17 Nov, SageAttention2 Technical Report: Accurate 4 Bit Attention for Plug-and-play Inference Acceleration, https://arxiv.org/abs/2411.10958
- 18 Nov, Bi-Mamba: Towards Accurate 1-Bit State Space Models, https://arxiv.org/abs/2411.11843
- 19 Nov, RedPajama: an Open Dataset for Training Large Language Models, https://arxiv.org/abs/2411.12372
- 20 Nov, Hymba: A Hybrid-head Architecture for Small Language Models, https://arxiv.org/abs/2411.13676
- 20 Nov, Loss-to-Loss Prediction: Scaling Laws for All Datasets, https://arxiv.org/abs/2411.12925
- 21 Nov, When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training, https://arxiv.org/abs/2411.13476
- 21 Nov, Multimodal Autoregressive Pre-training of Large Vision Encoders, https://arxiv.org/abs/2411.14402
- 21 Nov, Natural Language Reinforcement Learning, https://arxiv.org/abs/2411.14251
- 22 Nov, Large Multi-modal Models Can Interpret Features in Large Multi-modal Models, https://arxiv.org/abs/2411.14982
- 23 Nov, MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs, https://arxiv.org/abs/2411.15296
- 23 Nov, TÜLU 3: Pushing Frontiers in Open Language Model Post-Training, https://arxiv.org/abs/2411.15124
- 24 Nov, LLMs Do Not Think Step-by-step In Implicit Reasoning, https://arxiv.org/abs/2411.15862
April 2024
- 1 Apr, Do Language Models Plan Ahead for Future Tokens?, https://arxiv.org/abs/2404.00859
- 1 Apr, Bigger is not Always Better: Scaling Properties of Latent Diffusion Models, https://arxiv.org/abs/2404.01367
- 1 Apr, The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis, https://arxiv.org/abs/2404.01204
- 1 Apr, Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models, https://arxiv.org/abs/2404.04478
- 2 Apr, Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models, https://arxiv.org/abs/2404.02258
- 2 Apr, Long-context LLMs Struggle with Long In-context Learning, https://arxiv.org/abs/2404.02060
- 2 Apr, Emergent Abilities in Reduced-Scale Generative Language Models, https://arxiv.org/abs/2404.02204
- 2 Apr, Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, https://arxiv.org/abs/2404.02151
- 3 Apr, On the Scalability of Diffusion-based Text-to-Image Generation, https://arxiv.org/abs/2404.02883
- 3 Apr, BAdam: A Memory Efficient Full Parameter Training Method for Large Language Models, https://arxiv.org/abs/2404.02827
- 3 Apr, Cross-Attention Makes Inference Cumbersome in Text-to-Image Diffusion Models, https://arxiv.org/abs/2404.02747
- 4 Apr, Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences, https://arxiv.org/abs/2404.02151
- 4 Apr, Training LLMs over Neurally Compressed Text, https://arxiv.org/abs/2404.03626
- 4 Apr, CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues, https://arxiv.org/abs/2404.03820
- 5 Apr, ReFT: Representation Finetuning for Language Models, https://arxiv.org/abs/2404.03592
- 5 Apr, Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data, https://arxiv.org/abs/2404.03862
- 5 Apr, Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation, https://arxiv.org/abs/2404.04256
- 8 Apr, AutoCodeRover: Autonomous Program Improvement, https://arxiv.org/abs/2404.05427
- 8 Apr, Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence, https://arxiv.org/abs/2404.05892
- 8 Apr, CodecLM: Aligning Language Models with Tailored Synthetic Data, https://arxiv.org/abs/2404.05875
- 9 Apr, MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies, https://arxiv.org/abs/2404.06395
- 9 Apr, Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models, https://arxiv.org/abs/2404.06209
- 9 Apr, LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders, https://arxiv.org/abs/2404.05961
- 10 Apr, Adapting LLaMA Decoder to Vision Transformer, https://arxiv.org/abs/2404.06773
- 10 Apr, Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention, https://arxiv.org/abs/2404.07143
- 11 Apr, LLoCO: Learning Long Contexts Offline, https://arxiv.org/abs/2404.07979
- 11 Apr, JetMoE: Reaching Llama2 Performance with 0.1M Dollars, https://arxiv.org/abs/2404.07413
- 11 Apr, Best Practices and Lessons Learned on Synthetic Data for Language Models, https://arxiv.org/abs/2404.07503
- 11 Apr, Rho-1: Not All Tokens Are What You Need, https://arxiv.org/abs/2404.07965
- 12 Apr, Pre-training Small Base LMs with Fewer Tokens, https://arxiv.org/abs/2404.08634
- 12 Apr, Dataset Reset Policy Optimization for RLHF, https://arxiv.org/abs/2404.08495
- 13 Apr, LLM In-Context Recall is Prompt Dependent, https://arxiv.org/abs/2404.08865
- 15 Apr, State Space Model for New-Generation Network Alternative to Transformers: A Survey, https://arxiv.org/abs/2404.09516
- 15 Apr, Chinchilla Scaling: A Replication Attempt, https://arxiv.org/abs/2404.10102
- 15 Apr, Learn Your Reference Model for Real Good Alignment, https://arxiv.org/abs/2404.09656
- 16 Apr, Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study, https://arxiv.org/abs/2404.10719
- 16 Apr, Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies, https://arxiv.org/abs/2404.08197
- 16 Apr, How Faithful Are RAG Models? Quantifying the Tug-of-War Between RAG and LLMs’ Internal Prior, https://arxiv.org/abs/2404.10198
- 17 Apr, A Survey on Retrieval-Augmented Text Generation for Large Language Models, https://arxiv.org/abs/2404.10981
- 18 Apr, When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes, https://arxiv.org/abs/2404.12365
- 18 Apr, Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing, https://arxiv.org/abs/2404.12253
- 18 Apr, OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data, https://arxiv.org/abs/2404.12195
- 19 Apr, The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions, https://arxiv.org/abs/2404.13208
- 22 Apr, How Good Are Low-bit Quantized LLaMA3 Models? An Empirical Study, https://arxiv.org/abs/2404.14047
- 22 Apr, Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone, https://arxiv.org/abs/2404.14219
- 22 Apr, OpenELM: An Efficient Language Model Family with Open-source Training and Inference Framework, https://arxiv.org/abs/2404.14619
- 22 Apr, A Survey on Self-Evolution of Large Language Models, https://arxiv.org/abs/2404.14662
- 23 Apr, Multi-Head Mixture-of-Experts, https://arxiv.org/abs/2404.15045
- 23 Apr, NExT: Teaching Large Language Models to Reason about Code Execution, https://arxiv.org/abs/2404.14662
- 23 Apr, Graph Machine Learning in the Era of Large Language Models (LLMs), https://arxiv.org/abs/2404.14928
- 24 Apr, Retrieval Head Mechanistically Explains Long-Context Factuality, https://arxiv.org/abs/2404.15574
- 25 Apr, Layer Skip: Enabling Early Exit Inference and Self-Speculative Decoding, https://arxiv.org/abs/2404.16710
- 25 Apr, Make Your LLM Fully Utilize the Context, https://arxiv.org/abs/2404.16811
- 28 Apr, LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report, https://arxiv.org/abs/2405.00732
- 30 Apr, Better & Faster Large Language Models via Multi-token Prediction, https://arxiv.org/abs/2404.19737
- 30 Apr, RAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language Processing, https://arxiv.org/abs/2404.19543
- 30 Apr, A Primer on the Inner Workings of Transformer-based Language Models, https://arxiv.org/abs/2405.00208
- 30 Apr, When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively, https://arxiv.org/abs/2404.19705
- 30 Apr, KAN: Kolmogorov–Arnold Networks, https://arxiv.org/abs/2404.19756
Evaluation
- Comparison and Detection
- LLM QA Evaluation on Wikipedia: An insightful comparison of different LLMs’ performance on QA tasks using Wikipedia as a benchmark. LLM QA Evaluation Wikipedia
- This study offers a comparative analysis highlighting the strengths and weaknesses of open-source vs closed-source LLMs in handling QA tasks, providing valuable insights for both developers and users.
- LLM Zoo: A collection of various LLMs to explore and compare their capabilities. LLMZoo GitHub
- A unique repository that provides access to a wide range of LLMs, facilitating exploration, comparison, and understanding of different models’ functionalities and performance.
- Can AI-Generated Text be Reliably Detected?: Addresses the critical question of distinguishing between human and AI-generated text. AI-Generated Text Detection Study
- This paper delves into the challenges and methodologies involved in detecting AI-generated text, offering insights into the reliability of current detection techniques.
Introduction to Large Language Models
- Large Language Models (LLMs) like OpenAI’s GPT series have revolutionized the field of artificial intelligence, offering unprecedented capabilities in natural language understanding and generation. These models are trained on vast amounts of text data, enabling them to perform a wide range of language-based tasks, from writing and translation to answering questions and generating code.
- This is a jargon free primer
Mental health Employment Social Contract Under Automation
-
Fraudulent studies are undermining the reliability of systematic reviews – a study of the prevalence of problematic images in preclinical studies of depression | bioRxiv Death of the Internet Deepfakes and fraudulent content
-
Jonathan Haidt Wants You to Take Away Your Kid’s Phone | The New Yorker
- Jonathan Haidt, a social psychologist and the author of the book “The Anxious Generation: How the Great Rewiring of Childhood is Causing an Epidemic of Mental Illness”. The main points covered in the interview are:
- Haidt argues that a whole generation has been damaged by growing up with unrestricted access to social media and an overprotected childhood, leading to a sharp increase in anxiety, depression, and self-harm among teenagers, especially girls, starting around 2012.
- He attributes this to the rapid adoption of smartphones and social media platforms between 2010 and 2015, which radically changed childhood by replacing real-life interactions and play with excessive screen time and exposure to harmful online content.
- Haidt presents evidence from correlational and experimental studies to support his claim that social media use causes mental health issues, while acknowledging the need for more research.
- He argues that the benefits of social media are outweighed by its negative impact on child development, as it deprives children of essential real-life experiences, such as play, adventure, and healthy risk-taking.
- Haidt advocates for changing social norms and implementing restrictions on social media use, such as banning phones in schools and limiting access to social media platforms for children under 16, rather than an outright ban on technology.
- Jonathan Haidt, a social psychologist and the author of the book “The Anxious Generation: How the Great Rewiring of Childhood is Causing an Epidemic of Mental Illness”. The main points covered in the interview are:
Emerging use cases (AI holding text)
Style transfer for humans
- Multiple techniques tested with the same LoRA DoRA etc for comparison
- ActAnywhere
- AI-Enhanced Creator (beehiiv.com)
- AnimateAnyone for Node-Based Diffusion Pipeline Interface MrForExample/ComfyUI-AnimateAnyone-Evolved: Improved AnimateAnyone implementation that allows you to use the opse image sequence and reference image to generate stylized video (github.com)
- [CG Renders to AI ANIMATION
- NIKE video — MOONWALKERS PICTURE](https://www.moonwalkerspicture.com/newslounge/cg-renders-to-ai-workflow-vol-02-anim)
- Motion Control
- MotionCtrl (wzhouxiff.github.io)
- [2401.12945] Lumiere: A Space-Time Diffusion Model for Video Generation (arxiv.org)
- [I2VGen-XL
- a Hugging Face Space by damo-vilab](https://huggingface.co/spaces/damo-vilab/I2VGen-XL)
- ali-vilab/i2vgen-xl: Official repo for VGen: a holistic video generation ecosystem for video generation building on diffusion models (github.com)
- MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation (magicvideov2.github.io)
- Interpolation and interframe consistency
- controlnet and ebsynth temporal consistency
- Motion-Conditioned Diffusion Model for Controllable Video Synthesis
- Interframe consistency is now here
- Interpolation between two frames
- FILM frame interpolator
- ProPainter for Video Inpainting (shangchenzhou.com)
- zengyh1900/Awesome-Image-Inpainting: A curated list of image inpainting and video inpainting papers and resources (github.com)
- Runway AI video editing
- Gen2 examples
- Multishot VideoDrafter: Content-Consistent Multi-Scene Video Generation with LLM
- vienna with prompts
- Video slowmo and enhance
- deforum stable diffusion video
- Phenaki
- Collaborative video pipeline
- Magicvideo (faster)
- Production ready re aging
- distilled models for 25fps
- Stable warpfusion
- Video talking heads from text service
- Tune a video
- Vidyo: Generates videos for social networks from longer videos.
- Stylegan-T video transformer from google
- Houdini
- Dream Mix video to video remix
- RIFE frame interpolation
- example github for sd
- Synthesia corporate video generation
- pix2pixHD nextframe google colab
- minecraft demo codebase
- animation from mixamo
- Intel enhance photorealism in realtime
- custom SD video to video script
- Testing a custom video2video script I’m working on. (These used RealisticVision1.4 & ControlNet) : r/StableDiffusion
- consistency tools for character tooning
- Alibaba system
- website
- github
- model cards
- 9 new tools
- Automatic1111 plugin
- Next frame prediction with controlnet
- Will smith eating spaghetti
- Transform Video to Animation in Stable Diffusion | How to Install + BEST Consistency Settings: Learn how to use AI to create animations from real videos. We’ll use Stable Diffusion and other tools for maximum consistencyProject Files:https://bit.ly/3…
- How to Use ModelScope text2video with Automatic1111’s Stable Diffusion Web UI | kombitz: Enable the Extension Click on the Extension tab and then click on Install from URL. Enter https://github.com/deforum-art/sd-webui-modelscope-text2video in the URL box and click on Install. Click on Installed and click on Apply and restart UI. Go to your stable-diffusion-webui/models folder and create a folder called ModelScope and then create a folder called t2v under ModelScope. This is your models folder for text2video.
- This article provides instructions on how to use ModelScope’s text2video feature with Automatic1111’s Stable Diffusion Web UI.
- latent consistency pipeline
- [GitHub
- Picsart-AI-Research/Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators: Text-to-Image Diffusion Models are Zero-Shot Video Generators
- GitHub
- Picsart-AI-Research/Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators](https://github.com/Picsart-AI-Research/Text2Video-Zero)
- The Picsart-AI-Research/Text2Video-Zero repository contains code for a text-to-image diffusion model that can be used to generate videos from text input. The model is a zero-shot video generator, meaning that it does not require any training data in order to generate videos.
- LVDM for long video creation
- The Text2Room algorithm generates textured 3D meshes from a given text prompt by leveraging pre-trained 2D text-to-image models. The core idea is to select camera poses that will result in a seamless, textured 3D mesh. The algorithm iteratively fuses scene frames with the existing geometry to create the final mesh. Evaluation shows that the algorithm is able to generate room-scale 3D geometry with compelling textures from only text as input.
- The VMesh system models a scene with a triangular mesh and a sparse volume for efficient view synthesis. It is trained on multi-view images of an object to create a contiguous representation of the object’s surface and volume. This representation is then used to generate a simplified triangular mesh and a sparse volume, which can be stored and rendered efficiently. The system is designed for real-time applications and can render at 2K 60FPS on common consumer devices.
- LLM guided video generation paper
- LVM video gen using LLM paper
- Temporal stable automatic plugin
- We present a method for high-resolution video synthesis using latent diffusion models (LDMs). Our approach first pre-trains an LDM on images, then introduces a temporal dimension to the latent space diffusion model and fine-tunes it on encoded image sequences (i.e. videos). We focus on two real-world applications: simulation of in-the-wild driving data and creative content creation with text-to-video modeling. Our method achieves state-of-the-art performance on real driving videos of 512 x 1024 resolution. Additionally, our approach can leverage off-the-shelf pre-trained image LDMs, turning the publicly available, state-of-the-art text-to-image LDM Stable Diffusion into an efficient and expressive text-to-video model.
- This script allows for the automation of video stylization using StableDiffusion and ControlNet.
- Really easy videos in A1111
- Dancer 4 keyframes, low noise, controlnet approach
- Flicker free video workflow paper (good!)
- Pika labs
- Realtime lip-sync API
- ms image to video on huggingface
- model to video blender modules
- videocomposer in python 3.9
- motionagent image to video
- Animatediff comfy workflows on discord
- fluid animation youtube
- Controlnet tutorial
- LCM loras for fast inferencing
- Animatediff is a new animation software that provides a range of tools and features for creating high-quality animations. It offers a user-friendly interface and supports various animation techniques, such as 2D, 3D, stop motion, and more. With Animatediff, users can easily bring their ideas to life and express their creativity through unique and captivating animations. Whether you’re a professional animator or a beginner, Animatediff offers a comprehensive set of features to help you create stunning animations in a fast and efficient manner. title:: Animatediff and Stablevideo
- Youtube tutorials
- IF_Animator ComfyUI workflow LCM+Animatediff+IPA+CN (youtube.com)
- [[Part 2] Tips and Tricks
- AnimateDiff ControlNet Animation in ComfyUI
- YouTube](https://www.youtube.com/watch?v=aysg2vFFO9g)
- TianxingWu/FreeInit: FreeInit: Bridging Initialization Gap in Video Diffusion Models (github.com)
- CiaraStrawberry/svd-temporal-controlnet (github.com)
- ProjectNUWA/DragNUWA (github.com)
Texturing
Text to Multiview and Texturing
Text to 3D
November 2024
- 1 Nov, Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations, https://arxiv.org/abs/2411.00640
- 1 Nov 2024, Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation, https://arxiv.org/abs/2411.00412
- 1 Nov 2024, Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models, https://arxiv.org/abs/2411.00492
- 3 Nov, Sample-Efficient Alignment for LLMs, https://arxiv.org/abs/2411.01493
- 4 Nov 2024, A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness, https://arxiv.org/abs/2411.03350
- 4 Nov, “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM Quantization, https://arxiv.org/abs/2411.02355
- 4 Nov, Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study, https://arxiv.org/abs/2411.02462
- 5 Nov, HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems, https://arxiv.org/abs/2411.02959
- 6 Nov, Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination, https://arxiv.org/abs/2411.03823
- 6 Nov, Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding, https://arxiv.org/abs/2411.04282
- 6 Nov, Number Cookbook: Number Understanding of Language Models and How to Improve It, https://arxiv.org/abs/2411.03766
- 7 Nov, Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models, https://arxiv.org/abs/2411.04996
- 7 Nov, BitNet a4.8: 4-bit Activations for 1-bit LLMs, https://arxiv.org/abs/2411.04965
- 7 Nov, Scaling Laws for Precision, https://arxiv.org/abs/2411.04330
- 8 Nov, Energy Efficient Protein Language Models: Leveraging Small Language Models with LoRA for Controllable Protein Generation, https://arxiv.org/abs/2411.05966
- 8 Nov, Balancing Pipeline Parallelism with Vocabulary Parallelism, https://arxiv.org/abs/2411.05288
- 11 Nov, Toward Optimal Search and Retrieval for RAG, https://arxiv.org/abs/2411.07396
- 12 Nov, Large Language Models Can Self-Improve in Long-context Reasoning, https://arxiv.org/abs/2411.08147
- 12 Nov, Stronger Models are NOT Stronger Teachers for Instruction Tuning, https://arxiv.org/abs/2411.07133
- 12 Nov, Direct Preference Optimization Using Sparse Feature-Level Constraints, https://arxiv.org/abs/2411.07618
- 13 Nov, Cut Your Losses in Large-Vocabulary Language Models, https://arxiv.org/abs/2411.09009
- 15 Nov, Does Prompt Formatting Have Any Impact on LLM Performance?, https://arxiv.org/abs/2411.10541
- 17 Nov, SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization, https://arxiv.org/abs/2411.11909
- 17 Nov, SageAttention2 Technical Report: Accurate 4 Bit Attention for Plug-and-play Inference Acceleration, https://arxiv.org/abs/2411.10958
- 18 Nov, Bi-Mamba: Towards Accurate 1-Bit State Space Models, https://arxiv.org/abs/2411.11843
- 19 Nov, RedPajama: an Open Dataset for Training Large Language Models, https://arxiv.org/abs/2411.12372
- 20 Nov, Hymba: A Hybrid-head Architecture for Small Language Models, https://arxiv.org/abs/2411.13676
- 20 Nov, Loss-to-Loss Prediction: Scaling Laws for All Datasets, https://arxiv.org/abs/2411.12925
- 21 Nov, When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training, https://arxiv.org/abs/2411.13476
- 21 Nov, Multimodal Autoregressive Pre-training of Large Vision Encoders, https://arxiv.org/abs/2411.14402
- 21 Nov, Natural Language Reinforcement Learning, https://arxiv.org/abs/2411.14251
- 22 Nov, Large Multi-modal Models Can Interpret Features in Large Multi-modal Models, https://arxiv.org/abs/2411.14982
- 23 Nov, MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs, https://arxiv.org/abs/2411.15296
- 23 Nov, TÜLU 3: Pushing Frontiers in Open Language Model Post-Training, https://arxiv.org/abs/2411.15124
- 24 Nov, LLMs Do Not Think Step-by-step In Implicit Reasoning, https://arxiv.org/abs/2411.15862
April 2024
- 1 Apr, Do Language Models Plan Ahead for Future Tokens?, https://arxiv.org/abs/2404.00859
- 1 Apr, Bigger is not Always Better: Scaling Properties of Latent Diffusion Models, https://arxiv.org/abs/2404.01367
- 1 Apr, The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis, https://arxiv.org/abs/2404.01204
- 1 Apr, Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models, https://arxiv.org/abs/2404.04478
- 2 Apr, Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models, https://arxiv.org/abs/2404.02258
- 2 Apr, Long-context LLMs Struggle with Long In-context Learning, https://arxiv.org/abs/2404.02060
- 2 Apr, Emergent Abilities in Reduced-Scale Generative Language Models, https://arxiv.org/abs/2404.02204
- 2 Apr, Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, https://arxiv.org/abs/2404.02151
- 3 Apr, On the Scalability of Diffusion-based Text-to-Image Generation, https://arxiv.org/abs/2404.02883
- 3 Apr, BAdam: A Memory Efficient Full Parameter Training Method for Large Language Models, https://arxiv.org/abs/2404.02827
- 3 Apr, Cross-Attention Makes Inference Cumbersome in Text-to-Image Diffusion Models, https://arxiv.org/abs/2404.02747
- 4 Apr, Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences, https://arxiv.org/abs/2404.02151
- 4 Apr, Training LLMs over Neurally Compressed Text, https://arxiv.org/abs/2404.03626
- 4 Apr, CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues, https://arxiv.org/abs/2404.03820
- 5 Apr, ReFT: Representation Finetuning for Language Models, https://arxiv.org/abs/2404.03592
- 5 Apr, Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data, https://arxiv.org/abs/2404.03862
- 5 Apr, Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation, https://arxiv.org/abs/2404.04256
- 8 Apr, AutoCodeRover: Autonomous Program Improvement, https://arxiv.org/abs/2404.05427
- 8 Apr, Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence, https://arxiv.org/abs/2404.05892
- 8 Apr, CodecLM: Aligning Language Models with Tailored Synthetic Data, https://arxiv.org/abs/2404.05875
- 9 Apr, MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies, https://arxiv.org/abs/2404.06395
- 9 Apr, Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models, https://arxiv.org/abs/2404.06209
- 9 Apr, LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders, https://arxiv.org/abs/2404.05961
- 10 Apr, Adapting LLaMA Decoder to Vision Transformer, https://arxiv.org/abs/2404.06773
- 10 Apr, Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention, https://arxiv.org/abs/2404.07143
- 11 Apr, LLoCO: Learning Long Contexts Offline, https://arxiv.org/abs/2404.07979
- 11 Apr, JetMoE: Reaching Llama2 Performance with 0.1M Dollars, https://arxiv.org/abs/2404.07413
- 11 Apr, Best Practices and Lessons Learned on Synthetic Data for Language Models, https://arxiv.org/abs/2404.07503
- 11 Apr, Rho-1: Not All Tokens Are What You Need, https://arxiv.org/abs/2404.07965
- 12 Apr, Pre-training Small Base LMs with Fewer Tokens, https://arxiv.org/abs/2404.08634
- 12 Apr, Dataset Reset Policy Optimization for RLHF, https://arxiv.org/abs/2404.08495
- 13 Apr, LLM In-Context Recall is Prompt Dependent, https://arxiv.org/abs/2404.08865
- 15 Apr, State Space Model for New-Generation Network Alternative to Transformers: A Survey, https://arxiv.org/abs/2404.09516
- 15 Apr, Chinchilla Scaling: A Replication Attempt, https://arxiv.org/abs/2404.10102
- 15 Apr, Learn Your Reference Model for Real Good Alignment, https://arxiv.org/abs/2404.09656
- 16 Apr, Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study, https://arxiv.org/abs/2404.10719
- 16 Apr, Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies, https://arxiv.org/abs/2404.08197
- 16 Apr, How Faithful Are RAG Models? Quantifying the Tug-of-War Between RAG and LLMs’ Internal Prior, https://arxiv.org/abs/2404.10198
- 17 Apr, A Survey on Retrieval-Augmented Text Generation for Large Language Models, https://arxiv.org/abs/2404.10981
- 18 Apr, When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes, https://arxiv.org/abs/2404.12365
- 18 Apr, Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing, https://arxiv.org/abs/2404.12253
- 18 Apr, OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data, https://arxiv.org/abs/2404.12195
- 19 Apr, The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions, https://arxiv.org/abs/2404.13208
- 22 Apr, How Good Are Low-bit Quantized LLaMA3 Models? An Empirical Study, https://arxiv.org/abs/2404.14047
- 22 Apr, Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone, https://arxiv.org/abs/2404.14219
- 22 Apr, OpenELM: An Efficient Language Model Family with Open-source Training and Inference Framework, https://arxiv.org/abs/2404.14619
- 22 Apr, A Survey on Self-Evolution of Large Language Models, https://arxiv.org/abs/2404.14662
- 23 Apr, Multi-Head Mixture-of-Experts, https://arxiv.org/abs/2404.15045
- 23 Apr, NExT: Teaching Large Language Models to Reason about Code Execution, https://arxiv.org/abs/2404.14662
- 23 Apr, Graph Machine Learning in the Era of Large Language Models (LLMs), https://arxiv.org/abs/2404.14928
- 24 Apr, Retrieval Head Mechanistically Explains Long-Context Factuality, https://arxiv.org/abs/2404.15574
- 25 Apr, Layer Skip: Enabling Early Exit Inference and Self-Speculative Decoding, https://arxiv.org/abs/2404.16710
- 25 Apr, Make Your LLM Fully Utilize the Context, https://arxiv.org/abs/2404.16811
- 28 Apr, LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report, https://arxiv.org/abs/2405.00732
- 30 Apr, Better & Faster Large Language Models via Multi-token Prediction, https://arxiv.org/abs/2404.19737
- 30 Apr, RAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language Processing, https://arxiv.org/abs/2404.19543
- 30 Apr, A Primer on the Inner Workings of Transformer-based Language Models, https://arxiv.org/abs/2405.00208
- 30 Apr, When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively, https://arxiv.org/abs/2404.19705
- 30 Apr, KAN: Kolmogorov–Arnold Networks, https://arxiv.org/abs/2404.19756
Evaluation
- Comparison and Detection
- LLM QA Evaluation on Wikipedia: An insightful comparison of different LLMs’ performance on QA tasks using Wikipedia as a benchmark. LLM QA Evaluation Wikipedia
- This study offers a comparative analysis highlighting the strengths and weaknesses of open-source vs closed-source LLMs in handling QA tasks, providing valuable insights for both developers and users.
- LLM Zoo: A collection of various LLMs to explore and compare their capabilities. LLMZoo GitHub
- A unique repository that provides access to a wide range of LLMs, facilitating exploration, comparison, and understanding of different models’ functionalities and performance.
- Can AI-Generated Text be Reliably Detected?: Addresses the critical question of distinguishing between human and AI-generated text. AI-Generated Text Detection Study
- This paper delves into the challenges and methodologies involved in detecting AI-generated text, offering insights into the reliability of current detection techniques.
Introduction to Large Language Models
- Large Language Models (LLMs) like OpenAI’s GPT series have revolutionized the field of artificial intelligence, offering unprecedented capabilities in natural language understanding and generation. These models are trained on vast amounts of text data, enabling them to perform a wide range of language-based tasks, from writing and translation to answering questions and generating code.
- This is a jargon free primer
Mental health Employment Social Contract Under Automation
-
Fraudulent studies are undermining the reliability of systematic reviews – a study of the prevalence of problematic images in preclinical studies of depression | bioRxiv Death of the Internet Deepfakes and fraudulent content
-
Jonathan Haidt Wants You to Take Away Your Kid’s Phone | The New Yorker
- Jonathan Haidt, a social psychologist and the author of the book “The Anxious Generation: How the Great Rewiring of Childhood is Causing an Epidemic of Mental Illness”. The main points covered in the interview are:
- Haidt argues that a whole generation has been damaged by growing up with unrestricted access to social media and an overprotected childhood, leading to a sharp increase in anxiety, depression, and self-harm among teenagers, especially girls, starting around 2012.
- He attributes this to the rapid adoption of smartphones and social media platforms between 2010 and 2015, which radically changed childhood by replacing real-life interactions and play with excessive screen time and exposure to harmful online content.
- Haidt presents evidence from correlational and experimental studies to support his claim that social media use causes mental health issues, while acknowledging the need for more research.
- He argues that the benefits of social media are outweighed by its negative impact on child development, as it deprives children of essential real-life experiences, such as play, adventure, and healthy risk-taking.
- Haidt advocates for changing social norms and implementing restrictions on social media use, such as banning phones in schools and limiting access to social media platforms for children under 16, rather than an outright ban on technology.
- Jonathan Haidt, a social psychologist and the author of the book “The Anxious Generation: How the Great Rewiring of Childhood is Causing an Epidemic of Mental Illness”. The main points covered in the interview are:
Emerging use cases (AI holding text)
Style transfer for humans
- Multiple techniques tested with the same LoRA DoRA etc for comparison
- ActAnywhere
- AI-Enhanced Creator (beehiiv.com)
- AnimateAnyone for Node-Based Diffusion Pipeline Interface MrForExample/ComfyUI-AnimateAnyone-Evolved: Improved AnimateAnyone implementation that allows you to use the opse image sequence and reference image to generate stylized video (github.com)
- [CG Renders to AI ANIMATION
- NIKE video — MOONWALKERS PICTURE](https://www.moonwalkerspicture.com/newslounge/cg-renders-to-ai-workflow-vol-02-anim)
- Motion Control
- MotionCtrl (wzhouxiff.github.io)
- [2401.12945] Lumiere: A Space-Time Diffusion Model for Video Generation (arxiv.org)
- [I2VGen-XL
- a Hugging Face Space by damo-vilab](https://huggingface.co/spaces/damo-vilab/I2VGen-XL)
- ali-vilab/i2vgen-xl: Official repo for VGen: a holistic video generation ecosystem for video generation building on diffusion models (github.com)
- MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation (magicvideov2.github.io)
- Interpolation and interframe consistency
- controlnet and ebsynth temporal consistency
- Motion-Conditioned Diffusion Model for Controllable Video Synthesis
- Interframe consistency is now here
- Interpolation between two frames
- FILM frame interpolator
- ProPainter for Video Inpainting (shangchenzhou.com)
- zengyh1900/Awesome-Image-Inpainting: A curated list of image inpainting and video inpainting papers and resources (github.com)
- Runway AI video editing
- Gen2 examples
- Multishot VideoDrafter: Content-Consistent Multi-Scene Video Generation with LLM
- vienna with prompts
- Video slowmo and enhance
- deforum stable diffusion video
- Phenaki
- Collaborative video pipeline
- Magicvideo (faster)
- Production ready re aging
- distilled models for 25fps
- Stable warpfusion
- Video talking heads from text service
- Tune a video
- Vidyo: Generates videos for social networks from longer videos.
- Stylegan-T video transformer from google
- Houdini
- Dream Mix video to video remix
- RIFE frame interpolation
- example github for sd
- Synthesia corporate video generation
- pix2pixHD nextframe google colab
- minecraft demo codebase
- animation from mixamo
- Intel enhance photorealism in realtime
- custom SD video to video script
- Testing a custom video2video script I’m working on. (These used RealisticVision1.4 & ControlNet) : r/StableDiffusion
- consistency tools for character tooning
- Alibaba system
- website
- github
- model cards
- 9 new tools
- Automatic1111 plugin
- Next frame prediction with controlnet
- Will smith eating spaghetti
- Transform Video to Animation in Stable Diffusion | How to Install + BEST Consistency Settings: Learn how to use AI to create animations from real videos. We’ll use Stable Diffusion and other tools for maximum consistencyProject Files:https://bit.ly/3…
- How to Use ModelScope text2video with Automatic1111’s Stable Diffusion Web UI | kombitz: Enable the Extension Click on the Extension tab and then click on Install from URL. Enter https://github.com/deforum-art/sd-webui-modelscope-text2video in the URL box and click on Install. Click on Installed and click on Apply and restart UI. Go to your stable-diffusion-webui/models folder and create a folder called ModelScope and then create a folder called t2v under ModelScope. This is your models folder for text2video.
- This article provides instructions on how to use ModelScope’s text2video feature with Automatic1111’s Stable Diffusion Web UI.
- latent consistency pipeline
- [GitHub
- Picsart-AI-Research/Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators: Text-to-Image Diffusion Models are Zero-Shot Video Generators
- GitHub
- Picsart-AI-Research/Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators](https://github.com/Picsart-AI-Research/Text2Video-Zero)
- The Picsart-AI-Research/Text2Video-Zero repository contains code for a text-to-image diffusion model that can be used to generate videos from text input. The model is a zero-shot video generator, meaning that it does not require any training data in order to generate videos.
- LVDM for long video creation
- The Text2Room algorithm generates textured 3D meshes from a given text prompt by leveraging pre-trained 2D text-to-image models. The core idea is to select camera poses that will result in a seamless, textured 3D mesh. The algorithm iteratively fuses scene frames with the existing geometry to create the final mesh. Evaluation shows that the algorithm is able to generate room-scale 3D geometry with compelling textures from only text as input.
- The VMesh system models a scene with a triangular mesh and a sparse volume for efficient view synthesis. It is trained on multi-view images of an object to create a contiguous representation of the object’s surface and volume. This representation is then used to generate a simplified triangular mesh and a sparse volume, which can be stored and rendered efficiently. The system is designed for real-time applications and can render at 2K 60FPS on common consumer devices.
- LLM guided video generation paper
- LVM video gen using LLM paper
- Temporal stable automatic plugin
- We present a method for high-resolution video synthesis using latent diffusion models (LDMs). Our approach first pre-trains an LDM on images, then introduces a temporal dimension to the latent space diffusion model and fine-tunes it on encoded image sequences (i.e. videos). We focus on two real-world applications: simulation of in-the-wild driving data and creative content creation with text-to-video modeling. Our method achieves state-of-the-art performance on real driving videos of 512 x 1024 resolution. Additionally, our approach can leverage off-the-shelf pre-trained image LDMs, turning the publicly available, state-of-the-art text-to-image LDM Stable Diffusion into an efficient and expressive text-to-video model.
- This script allows for the automation of video stylization using StableDiffusion and ControlNet.
- Really easy videos in A1111
- Dancer 4 keyframes, low noise, controlnet approach
- Flicker free video workflow paper (good!)
- Pika labs
- Realtime lip-sync API
- ms image to video on huggingface
- model to video blender modules
- videocomposer in python 3.9
- motionagent image to video
- Animatediff comfy workflows on discord
- fluid animation youtube
- Controlnet tutorial
- LCM loras for fast inferencing
- Animatediff is a new animation software that provides a range of tools and features for creating high-quality animations. It offers a user-friendly interface and supports various animation techniques, such as 2D, 3D, stop motion, and more. With Animatediff, users can easily bring their ideas to life and express their creativity through unique and captivating animations. Whether you’re a professional animator or a beginner, Animatediff offers a comprehensive set of features to help you create stunning animations in a fast and efficient manner. title:: Animatediff and Stablevideo
- Youtube tutorials
- IF_Animator ComfyUI workflow LCM+Animatediff+IPA+CN (youtube.com)
- [[Part 2] Tips and Tricks
- AnimateDiff ControlNet Animation in ComfyUI
- YouTube](https://www.youtube.com/watch?v=aysg2vFFO9g)
- TianxingWu/FreeInit: FreeInit: Bridging Initialization Gap in Video Diffusion Models (github.com)
- CiaraStrawberry/svd-temporal-controlnet (github.com)
- ProjectNUWA/DragNUWA (github.com)
Texturing
Text to Multiview and Texturing
Text to 3D
Evaluation
- Comparison and Detection
- LLM QA Evaluation on Wikipedia: An insightful comparison of different LLMs’ performance on QA tasks using Wikipedia as a benchmark. LLM QA Evaluation Wikipedia
- This study offers a comparative analysis highlighting the strengths and weaknesses of open-source vs closed-source LLMs in handling QA tasks, providing valuable insights for both developers and users.
- LLM Zoo: A collection of various LLMs to explore and compare their capabilities. LLMZoo GitHub
- A unique repository that provides access to a wide range of LLMs, facilitating exploration, comparison, and understanding of different models’ functionalities and performance.
- Can AI-Generated Text be Reliably Detected?: Addresses the critical question of distinguishing between human and AI-generated text. AI-Generated Text Detection Study
- This paper delves into the challenges and methodologies involved in detecting AI-generated text, offering insights into the reliability of current detection techniques.
Introduction to Large Language Models
-
Large Language Models (LLMs) like OpenAI’s GPT series have revolutionized the field of artificial intelligence, offering unprecedented capabilities in natural language understanding and generation. These models are trained on vast amounts of text data, enabling them to perform a wide range of language-based tasks, from writing and translation to answering questions and generating code.
-
Core Characteristics
-
Autoregressive Generation: Sequential token-by-token text production
-
Conditional Generation: Text production conditioned on prompts or contexts
-
Controllable Attributes: Style, tone, length, and topic control
-
Few-Shot and Zero-Shot: Generation from minimal examples or instructions
-
Factual Consistency: Grounding in knowledge and reducing hallucination
-
Multi-Domain: News, creative writing, technical documentation, code
Relationships
-
Subclass: Natural Language Processing
-
Related: Language Modeling, Large Language Model, GPT, Text-to-Text Generation
-
Models: GPT-3/4, T5, BLOOM, LLaMA, PaLM
-
Applications: Content Creation, Code Generation, Creative Writing, Summarisation
Key Literature
-
Brown, T., et al. (2020). “Language models are few-shot learners.” NeurIPS, 1877-1901.
-
Raffel, C., et al. (2020). “Exploring the limits of transfer learning with a unified text-to-text transformer.” JMLR, 21(140), 1-67.
-
Radford, A., et al. (2019). “Language models are unsupervised multitask learners.” OpenAI Blog.
-
Holtzman, A., et al. (2020). “The curious case of neural text degeneration.” ICLR.
2024-2025: Reasoning Models and Multimodal Text Generation Breakthrough
The period from 2024 through 2025 witnessed transformative developments in text generation, with the emergence of reasoning-optimised models, widespread multimodal integration, and intense competition driving rapid performance improvements across all major frontier language models.
Reasoning-First Architecture: OpenAI o1 and o3
In September 2024, OpenAI unveiled o1, experimental models specifically fine-tuned to generate chains of thought before providing answers, scoring particularly high in mathematics, coding, and science benchmarks. Released in full on 5th December 2024, o1 marked a significant shift toward reasoning-first architecture, representing the first model explicitly optimised for chain-of-thought reasoning.
In December 2024, OpenAI offered a glimpse of o3—o1’s successor with impressive capabilities—whilst Google and DeepSeek unveiled their own reasoning models, establishing reasoning as a core paradigm for text generation going forward.
Multimodal Text Generation Revolution
GPT-4o was released on 13th May 2024 as a flagship multimodal model designed to process and generate text, audio, and visual inputs and outputs in real time. The “o” stands for “omni,” signalling the model’s ability to handle longer conversations with better memory whilst understanding both text and images.
It was truly in 2024 that multimodal LLMs became mainstream. Claude 3.5 Sonnet (released June 2024) excelled in reading, coding, mathematics, and vision tasks. In May 2025, Anthropic introduced the Claude 4 family, including Claude 4 Opus and Claude 4 Sonnet, with Opus 4 optimised for complex reasoning and coding.
Performance Convergence and Competition
Some new iterations of fast models (GPT-4o, Gemini 2.0 Flash, and Claude 3.5 Sonnet) became more performant than the flagship models of the previous generation (GPT-4, Gemini 1.5 Pro, Claude 3 Opus), demonstrating accelerating performance improvements. By the end of 2024, OpenAI’s leadership faced stiff competition, with GPT-4o tied with o1 and two versions of Google’s Gemini for first place on the LMSYS Chatbot Arena leaderboard.
Gemini 2.0 by Google DeepMind launched in December 2024, expanding AI’s multimodal potential and integrating seamlessly with autonomous agents. Gemini 2.0 Flash emerged as one of the fastest options for text generation tasks.
Competitive Landscape Maturation
The text generation landscape matured significantly in 2024-2025, transitioning from OpenAI dominance to a highly competitive multi-player market with Google, Anthropic, Meta, and DeepSeek all fielding competitive frontier models. This competition drove rapid capability improvements, pricing reductions, and broader accessibility to state-of-the-art text generation capabilities.
See Also
-
-
Core Characteristics
-
Autoregressive Generation: Sequential token-by-token text production
-
Conditional Generation: Text production conditioned on prompts or contexts
-
Controllable Attributes: Style, tone, length, and topic control
-
Few-Shot and Zero-Shot: Generation from minimal examples or instructions
-
Factual Consistency: Grounding in knowledge and reducing hallucination
-
Multi-Domain: News, creative writing, technical documentation, code
Relationships
-
Subclass: Natural Language Processing
-
Related: Language Modeling, Large Language Model, GPT, Text-to-Text Generation
-
Models: GPT-3/4, T5, BLOOM, LLaMA, PaLM
-
Applications: Content Creation, Code Generation, Creative Writing, Summarisation
Key Literature
-
Brown, T., et al. (2020). “Language models are few-shot learners.” NeurIPS, 1877-1901.
-
Raffel, C., et al. (2020). “Exploring the limits of transfer learning with a unified text-to-text transformer.” JMLR, 21(140), 1-67.
-
Radford, A., et al. (2019). “Language models are unsupervised multitask learners.” OpenAI Blog.
-
Holtzman, A., et al. (2020). “The curious case of neural text degeneration.” ICLR.
See Also
-