Voice and narration

Voice and narration

Plottery can use a text-to-speech model to read the story aloud with AI voices. Like the story text itself, the speech is generated on your own machine. You design voices by describing how they should sound. The voice tools are under Settings > Voice.

Design a voice by describing it

There's no fixed list of preset voices. You type how a voice should sound in plain words ("Calm female narrator, neutral American accent, warm and clear"; "Gruff old man, low and gravelly, weary"), and a model designs it from that. A rework button can have the AI write or tweak the description for you (older, younger, opposite gender, slower, more theatrical, different accent).

Takes and locking a voice

Generate 5 takes creates different variations on your description. When you name and save the voice, it is locked, so every future line uses that exact voice instead of generating a new one from the description.

The narrator and reading beats aloud

One saved voice has to be set as the Narrator. It reads prose and dialog if that character doesn't have its own voice. The first voice you save becomes the narrator automatically, and you can reassign the role with the mic button on any voice. You can set the Auto-narration to automatically read every new beat as soon as the language model finishes generating it.

Character voices

A character with a voice description can speak in their own voice. You can design it from the character's Voice section in the scenario editor, or reuse a saved voice. A character with a description but no assigned voice gets one generated automatically the first time one of their lines is read. However this can result in badly chosen voices. Published scenarios can include their characters' voices, so imported characters have their voices ready without any setup.

In a chat, every line uses the assigned character's voice, or the narrator voice if none is assigned.

Read any selection aloud

Select text anywhere in the app and a small speaker button appears. Click it to hear the selection in the narrator voice.

The voice models

Narration runs on Qwen3-TTS through a bundled engine with two models. You can choose between the models in the settings. Quality uses the 1.7B model. Speed uses the 0.6B one, which is faster and lighter at lower quality. You can also choose between GPU or CPU. Consider using the CPU, because the performance difference is not big, and it leaves more memory for the text model on the GPU. The models download on first use and can be deleted to reclaim disk space if you don't want to use them.