Skip to content
Aoden Teo@AodenTeoMT · Jun 3, 2026

Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an…

7 tweets1 min read♥ 9.8Koriginal

Summary

Miso One is a new 8-billion-parameter text-to-speech model that generates highly expressive, human-like speech with 110ms latency. The model weights have been open-sourced with API access coming soon, and it can be tried directly online.

Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.

  1. #1

    Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of latency. We’ve open-sourced the model weights, with API access coming soon. Hear how Miso One sounds in the thread below.

  2. #2

    1/3 Miso One generated every voiceover in this thread.

  3. #3

    2/3 Miso One can narrate shorts like a human content creator.

  4. #4

    To download Miso One, check out the repo: github.com/MisoLabsAI/Mis…

  5. #5

    You can also try Miso One directly at misolabs.ai.

  6. #6

    3/3 We even cloned Sal Khan’s voice and had Miso One teach us math!

  7. #7

    Yes, me.