English
Discover amazing AI tools and applications
An upgraded TTS system featuring multilingual support, real-time style switching and efficient inference
a powerful node-based AI workflow platform that enables local image, video, voice, and music generation with highly flexible and customizable workflows
Turn Text into Realistic Podcasts with Multi-Speaker, Multilingual & Emotional Speech
Zero-shot TTS system by Bilibili featuring fast voice cloning, cross-lingual synthesis, precise emotion/speed control
Automatically generates HD short videos with script, footage, voiceover, subtitles and music
Clone voice in 5 seconds — GPT-SoVITS enables multilingual AI speech.
Zero-shot voice cloning and natural-language-driven voice design across multiple languages and dialects
An ultra-fast, lightweight open-source music model, delivering commercial-grade audio on local hardware with less than 4GB VRAM
Supporting 600+ languages, voice design, voice cloning, natural speech, and ultra-fast inference
Supporting Mandarin, English and Cantonese, with natural speech synthesis and zero-shot voice cloning
A fast multimodal audio generation model supporting text and video conditioning for audio/music synthesis
Generate 1-minute videos quickly with only 6GB of low VRAM.
A web-based 3D tool built with Three.js and React that lets users manually arrange models, set character poses, and adjust camera angles for quick storyboarding
Automatically handles video translation, subtitle generation and dubbing
An open-source LLM-DiT speech model by Xiaohongshu featuring zero-shot cloning for 24 languages and 21 dialects
A fast local AI audio platform supporting dozens of AI models, enabling voice generation, recognition, conversion, and music creation through an easy-to-use Web interface
Zero-shot cloning and emotional control across 23 languages
Delivers single-pass transcription, speaker diarization, and timestamping for long audio across 50+ languages
Highly expressive, long-form, multi-speaker conversational audio generation
Generates lightweight interactive 3D Gaussian scenes compatible with mainstream 3D software from one single image rapidly
A bilingual (English/Chinese) text-to-speech model built for ultra-low-latency streaming that enables text-driven voice creation, voice cloning, and fine-grained emotional control
Upgraded model architecture with greatly improved sound quality and song integrity, supports ultra-long duration and multilingual creation
Creates high-quality realistic sound effects from text or video with only English prompts supported
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Animate Digital Humans' Lip Movements
Featuring ultra-realistic 48kHz zero-shot voice cloning with rich emotional and lifelike expressiveness
Voice cloning with text-tag control over 11 emotions, 4 paralinguistic sounds (like laughter/breathing), and 14 Chinese dialects
Transforms simple text prompts into up to 6 minutes of studio-quality music and lifelike sound effects
Turns lyrics into songs in seconds, generates music by style keywords.
A speech synthesis system that generates multi-speaker conversations with voice cloning and multilingual support.
Multilingual support for 52 languages/dialects and exceptional robustness in song and contextual transcription
An open-source fast-thinking multilingual translation model supporting 33 languages
Multi-language emotional expression, and real-time streaming generation to enable human-like natural speech with low-resource cross-scenario deployment.
An AI cover and music remixing model that supports reference audio, lyrics, MIDI, and controllable musical styles
Clones voices from short audio and generates natural speech.
Multilingual speech recognition, emotion & audio event detection—efficient and accurate
multilingual, real-time/offline recognition, easy to use and efficient
Real-time talking-head framework, high-fidelity, long-duration stable audio-visual synchronization
An open-source translation model, 33 languages + 5 dialects, accurate & flexible
A joint zero-shot SVS project supporting multilingual synthesis and dual-mode control
An expressive open-source text-to-speech model supporting 31 languages, featuring stable zero-shot voice cloning and precise inline pause control
Turns static portraits into video/audio-driven 3D models in real time
Generate high-quality classical music, supports generation by period, composer and instrumentation
An all-in-one local AI assistant supporting cross-platform operation and mobile remote control
A desktop graphical tool developed by ValueCell-ai based on OpenClaw, featuring one-click installation
Zero-shot voice cloning, supports multiple languages, allows voice parameter control
A large model runner with a visual interface.
Supporting 17 languages, with accurate dialect and low-volume recognition
Generation and understanding, featuring high-fidelity song synthesis and controllable structural creation
Taming Bad Noise for Effective Video Object Removal
A desktop client for ChatGPT, Claude and other LLMs
PartPacker enables part-level 3D object generation from single-view images
Turns a simple description into a high-fidelity sound effect up to 30 seconds long, perfect for video voiceovers and game asset creation.
Removing hard-coded subtitles from videos and text watermarks from images with lossless resolution
Generates depth-aware 3D panoramas and scene models from single images
Supporting high-quality TTS and zero-shot voice cloning with extremely high timbre similarity
Add perfectly fitting foley sounds to silent videos.
Get up and running with large language models
A workflow automation platform supporting no-code/code dual-mode building
A security-first workflow automation tool with enhanced reliability and efficiency, supporting visual drag-and-drop operation
Clones voice and emotion from a short sample for accent-free dubbing across 14 languages
An easy-to-deploy, extensible open-source AI chatbot
An ultra-lightweight TTS tool, supporting multilingual speech synthesis and zero-shot voice cloning on ordinary CPU
A smart-searchable library of 2000+ ready-to-use n8n automation workflows
A lightweight, efficient open-source document parser that accurately converts PDFs, images, and e-books into Markdown/JSON
Zero-shot voice cloning, emotion expression capabilities
Zero-shot voice cloning and expressive editing of emotion, style, and paralinguistic cues
Making local speech-to-text and translation easy.
Zero-shot voice cloning, emotion expression and streaming inference
Generates pixel-aligned high-fidelity 3D models with PBR textures from a single image
A vision-only, single-step generative image super-resolution model, featuring ultra-fast inference and superior restoration of fine text and structural details
A free video editor featuring drag-and-drop simplicity, rich effects, and 4K video export
Easily displays technical and metadata information of audio and video files
A self-hosted machine translation tool supporting multi-language translation, usable offline with controllable data privacy
A cross-platform video editor offering broad format support, powerful filters, and easy timeline editing
Seamless integration with chat apps like WeChat, DingTalk, and Feishu for multi-agent collaborative workflows
A lightweight and easy-to-use multi-terminal personal AI assistant, supporting multi-channel connection and custom skills
A cross-platform software for high-performance live streaming and screen recording
An open-source enterprise AI assistant integrating RAG pipelines, multi-modal interaction, and workflow orchestration.