Narrative Goldmine

Home

❯

working

❯

Speech and voice

Speech and voice

01 Oct 20268 min read

Properties

Type
  • Note
Status
  • stable
Generated
  • by: process:vault-migrate/1.0 · at: 2026-09-22T12:41:58.55607396Z

  • NVIDIA/NeMo: NeMo: a toolkit for conversational AI (github.com)

    • [Canary
      • NVIDIA NeMo](https://nvidia.github.io/NeMo/blogs/2024/2024-02-canary/)

    H200-NeMo-performance

  • NeMo/tutorials/tts/FastPitch_Adapter_Finetuning.ipynb at main · NVIDIA/NeMo (github.com)

  • ElevenLabs Audio Native

  • OpenAI whisper local deploy

  • realtime transciber

  • high performance CPP

  • 30% quantised optimisation

  • Brillbits OpenAI whisper demo with mic

  • Cleanvoice audio denoise

  • Cloud voice change app

  • downloadable voice generation systems

  • Language AI open libraries

  • Language practice

  • MUGEN multi modal from facebook

  • Oneshot speach to text

  • Record and cleanup pro audio with commodity hardware

  • Respeecher

  • Voice AI voices

  • Voice controlled assisted creation

  • Voice to text, Lopp

  • whisper transcriber

  • Wolfram alpha voice chatbot integration

  • Microsoft Vall-E voice synthesis

  • Uberduck text to speech (plus own voice)

  • Eleven labs language and text to speech

  • Uberduck open source text to speech

  • numen voice control system in linux

  • Inworld (steam game plugin AI system) for voice chat and answer

  • Bark text to speech from google labs

  • https://github.com/TensorSpeech/TensorFlowTTS very configurable from what I see

  • VoiceVox engine

  • [coqui-ai TTS

    • very good samples](https://github.com/coqui-ai/TTS)
  • https://github.com/neonbjb/tortoise-tts

  • https://github.com/CorentinJ/Real-Time-Voice-Cloning

    • custom voices? looks neat
  • https://github.com/rhasspy/larynx - very low-spec compatible, acceptable quality

  • Voice cloning local

  • Meta voicebox

  • The Reddit post discusses the different open source voice cloning projects available, including Coqui, Tortoise, and Bark. The advantages and disadvantages of each project are briefly outlined, with ElevenLabs being noted as the best but not open source, while Tortoise is suggested as the closest open source alternative. Other tools for speech to speech and singing conversion, such as so-vits/diff-svc/rvc, are also mentioned. The post suggests that the quality of open source voice cloning projects is improving, and that there may be more options available in the future. https://www.reddit.com/r/MachineLearning/comments/133hanr/d_what_are_the_differences_between_the_major_open/

  • The Retrieval-based Voice Conversion WebUI is a simple and useful voice conversion (voice changer) framework based on the VITS algorithm. It can use a small amount of voice data and still achieve good results. It incorporates a top-1 retrieval method to replace the source feature with the training set feature to avoid voice leakage, and it is easy to use with a simple web interface. It also features model fusion to change voice characteristics and the ability to integrate with the UVR5 model to quickly separate vocals and accompaniment. The project requires the installation of PyTorch and its core dependencies, and other pre-models are also needed for inference and training. The repository provides a guide to environment setup and usage, as well as links to relevant resources and contributors. https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI

  • The article discusses different open-source voice cloning projects and their advantages and disadvantages. The projects mentioned include Coqui, Tortoise, and Bark, with the author highlighting Coqui’s unlocked platform, while Tortoise and Bark are newer transformer-based projects that can clone much more effectively with much less training and are restricted to prevent custom voice cloning. The author suggests that the ElevenLabs is currently the best voice cloning solution available, but it is not open source and can be expensive. The article also includes comments from other Reddit users, who suggest other open source options and provide additional insights into each option’s strengths and weaknesses. https://www.reddit.com/r/MachineLearning/comments/133hanr/d_what_are_the_differences_between_the_major_open/

  • The article provides instructions on how to use OpenAI’s ChatGPT chatbot on an Android device using the Tasker app. The process involves importing a ChatGPT profile into Tasker, obtaining an API key from OpenAI, and setting up home screen shortcuts. The article also notes that ChatGPT can be run through Google Assistant with voice commands. The author suggests that while ChatGPT may not necessarily be better than Google Assistant, it can perform tasks that Google Assistant may not be capable of. https://www.howtogeek.com/882019/how-to-use-chatgpt-like-google-assistant-on-android/

  • The Voice Assistant is an AI-powered chatbot that uses several APIs to understand natural language commands and provide helpful responses. It features a wide range of capabilities, including answering general knowledge questions, providing recommendations, performing productivity tasks, and entertaining users. The Voice Assistant was built using ChatGPT, Whisper API, Gradio, and Microsoft’s SpVoice TTS API, and it can be accessed through a web-based interface. The installation process involves cloning the repository and installing the required Python packages. Contributions to the project are welcome. https://github.com/DonGuillotine/chatGPT_whisper_AI_voice_assistant

  • The Retrieval-based Voice Conversion WebUI is a voice conversion framework that uses a top-1 retrieval algorithm to eliminate voice leakage. It is capable of quickly training even on relatively poor GPUs and can achieve good results even with just 10 minutes of low noise voice data. It has a user-friendly web interface and the ability to use a model fusion system to change voice timbre. The setup recommends using Poetry and downloading the necessary pre-trained models from their Hugging Face space. It also includes additional files such as ffmpeg and ffprobe that may need to be downloaded. The WebUI can be initiated using the command “python infer-web.py” and Windows users can run the “go-web.bat” file. The project also acknowledges the contributions of related tools and libraries such as Gradio, HIFIGAN, and ContentVec. https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI

  • VoicePen is a tool that uses AI to convert audio or video files into blog posts and transcriptions in minutes. The service includes a transcription and SRT file generated by a top speech-to-text model, an English blog post that pulls out key topics from the audio, and the ability to convert audio in 96 different languages. Use cases include repurposing podcasts, webinars, and tutorial videos. Monthly plans are available, with options for one-time conversions. Testimonials praise the accuracy and speed of VoicePen’s service. https://voicepen.ai

  • Krisp is a software application designed to improve the productivity of online meetings by using AI-powered voice clarity and a meeting assistant to cancel background noise, echo, and accent localization. It works on both Mac and Windows platforms and processes only the user’s voice on their device, unlike other solutions that transmit voice over the internet. Krisp offers a free forever plan with no credit card required and is trusted by global brands. The insights gathered from calls can be viewed by the user to improve their communication skills over time. Krisp has received recognition from various prestigious awards such as America’s Most Promising AI Companies and has been awarded for its quality of support and ease of use. Krisp also offers SDK for developers, pricing and plans, and use cases such as contact centers and enterprise. The company prioritizes customers’ privacy, security and offers accessible support, including video tutorials and a help center. By accepting all cookies, users consent to the storing of cookies on their device to enhance site navigation, analyze site usage and assist in the company’s marketing efforts. https://krisp.ai/

  • Cleanvoice AI is an artificial intelligence platform that assists users in editing their podcasts or audio recordings. The platform offers various features such as filler sound removal, mouth sound removal, stutter removal, and Deadair remover to make the audio recording more professional. Cleanvoice AI is multilingual and can detect filler sounds in multiple languages, including accents from various countries. The platform also allows for manual editing with assistance and offers tools like podcast mixing and background noise remover. Users can try Cleanvoice AI for free for 30 minutes without providing credit card details. However, users must accept the platform’s cookie policy to use the service. https://cleanvoice.ai/

  • The article discusses the potential of Central Intelligent Agents (CIAs) and the role of large language models (LLMs) and other next-generation AI technologies in enabling them. It highlights the need for businesses to have a cross-functional team, ethical guidelines, and clear objectives in deploying their own CIA. The article also suggests steps to build a solid foundation for deploying a CIA, assess organizational readiness, assemble a cross-functional team, define objectives, develop the CIA components and evaluate its performance while continuing to learn and adapt. The author discusses the potential of AI tools and voice assistants in transforming the way businesses interact with their customers and suggests that the advent of advanced AI technologies has revolutionized the shift of businesses towards a more personalized and ethically responsible approach to engaging with their customers. Finally, the article ends by highlighting the importance of experimenting through crisis and providing expert guidance tailored to specific business needs. https://www.linkedin.com/pulse/central-intelligent-agent-enabling-next-generation-james-poulter?

  • TensorSpeech/TensorFlowTTS: :stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German and Easy to adapt for other languages) Translation Accessibility Speech and voice Speech and voice

  • Variety Speech and voice Social contract and jobs

  • transcriptionstream/transcriptionstream: turnkey self-hosted offline transcription and diarization service with llm summary (github.com) Speech and voice transcription locally SHOULD

  • Tincans - Gazelle v0.2 Speech and voice fast speech engine SHOULD

  • Speech and voice Open Voice (myshell.ai) cloning MIT license

  • EndlessDreams: Voice directed real-time videos at 1280x1024 : r/StableDiffusion (reddit.com) Speech and voice Speech and voice Product Design Real Time

  • https://demo.hume.ai/? Speech and voice Large language models empathetic voice to voice

  • Speech and voice metavoiceio/metavoice-src: AI for human-level speech intelligence (github.com) check for PlayerTwo

  • NeMo/tutorials/tts/NeMo_TTS_Primer.ipynb at main · NVIDIA/NeMo (github.com) NVIDIA Omniverse Speech and voice primer and demo.


Graph View

Backlinks

  • Automated Podcasting
  • Speech and voice

Created with Quartz v4.5.2 © 2026

  • Ontology (Turtle)
  • Search index
  • JSON-LD context
  • Explorer