KittenTTS logo

KittenTTS

KittenTTS is an open-source, lightweight text-to-speech library built on ONNX that ships state-of-the-art voice synthesis in models ranging from 15M to 80M parameters, only 25 to 80 MB on disk, and runs entirely on CPU without a GPU. It exposes a clean Python API for synthesis, writes audio directly to a file, normalises numbers and currencies per locale, switches voices by name, and adjusts speed with a multiplier, all without external services or accounts. The v0.8 release added 15M, 40M, and 80M parameter variants so you can trade size for fidelity on anything from an edge device to a server batch job. It is built for developers who want a dependency-minimal, offline TTS for embedded apps, agents, accessibility tooling, and speech output where shipping a multi-gigabyte model is not feasible. KittenTTS matters now because high-quality, CPU-only speech under 25 MB makes on-device voice practical at the long tail of constrained hardware.

Reader rating

No ratings yet

Visit website

You might also like

Related tools

View all
Demon favicon
Demon
No ratings yet

Demon (Diffusion Engine for Musical Orchestrated Noise) is an open-source real-time music generation system that runs locally on consumer GPUs at 25Hz. It is built for musicians, sound designers, music producers, and AI researchers who want to generate, iterate, and perform with AI music in real time without relying on cloud APIs. The system uses diffusion-based synthesis to produce musical audio streams with low latency, enabling live experimentation and performance workflows. Demon launched on Hacker News with 15 points and the project page at daydreamlive.github.io/DEMON describes a fully local, GPU-accelerated approach to music generation. What makes it notable is the combination of real-time performance with diffusion models — a technical achievement that opens up live music creation use cases that were previously impossible with slower batch-generation approaches.

View details
Udio favicon
Udio
No ratings yet

Meet Udio, your AI-powered music creation companion. With Udio, you can effortlessly create and share music using cutting-edge AI technology. This free platform offers a range of tools to produce and refine audio content, from generating diverse music genres to creating vocals and instrumentals in seconds. Whether you're a music enthusiast, content creator, or just looking to add a unique touch to your projects, Udio's AI audio tools have you covered. From crafting melodies to experimenting with text-to-speech capabilities, Udio empowers you to explore the endless possibilities of AI-generated music and audio content. Unleash your creativity and dive into the world of AI music with Udio today!

View details
Magenta RealTime 2 favicon
Magenta RealTime 2
No ratings yet

Magenta RealTime 2 is a Google Magenta AI model or feature for real-time music and media generation workflows. It is aimed at creators who want responsive audio experimentation, live composition support, and AI-assisted musical ideation without waiting for slow offline rendering cycles. Musicians, sound designers, creative coders, and media teams can use this kind of tool to sketch melodies, explore variations, prototype interactive audio, or build performance-focused experiences. Its value is the real-time angle: instead of treating generative audio as a batch process, it supports a more immediate creative loop. Magenta RealTime 2 fits Smartoolbox as a niche but relevant AI audio generator.

View details