Text → voice → YouTube: the AI does the narration itself
The AI works with the APIs itself: it turns long text into natural narration, builds one seamless MP3, sends it to Telegram with a player and files it on Google Drive, and then the YouTube autopublishing system turns that audio into a video.
In-house buildNarrates Bible studies for the “Slovo i Dukh” channel
- A chapter in ~35 seconds
- Up to 25 minutes of audio at once
- Feeds YouTube autopublishing
01Diagram: text → AI agent → Gemini TTS → FFmpeg → Telegram and Google Drive → YouTube
01Before
Turning long text into publishable audio meant running text-to-speech piece by piece by hand, stitching the pieces, naming the files, archiving and sending them. Every chapter was the same routine, and one naming mistake would throw the whole archive off.
02What changed
- A chapter in half a minute
15 minutes of audio are generated in about 35 seconds, and up to 25 minutes in a single request.
- A natural voice
long text is narrated in one go, with no chopping or stitching.
- The file lands where it's needed
an MP3 with a player in Telegram and in the Google Drive archive, consistently named.
- Then YouTube
the autopublishing system turns the audio into a video for the “Slovo i Dukh” channel.
03How it works
- 01
The chapter text goes to the AI agent
- 02
The agent strips markup and checks every word
- 03
A neural voice narrates it and the file becomes an MP3
- 04
The audio goes to Telegram, Drive and on to YouTube
The same approach, built for your task
Here it's Bible studies. The same “text → voice → publish” chain works for any written content.
- 01Blogs and media
an audio version of every article for a podcast and a Telegram channel.
- 02Online schools
narrated lessons and notes, including in several languages.
- 03Business
voice instructions for staff, company news, narrated presentations.
- 04Books and courses
an audiobook by chapter, with a version archive and consistent file names.
Voice, language, format and destination are all set up for you.
Under the hoodTechnical details for specialists
- Two versions of the system
an n8n pipeline with parallel narration, and an AI agent that calls the speech API and assembles the file itself.
- Word check
after stripping markup, the system compares the word sequence with the source, so nothing is lost or changed.
- Reliability
in the n8n version, a failed request is retried and missing chunks are re-voiced in a second pass, so the file arrives complete.
- MP3 is a must
25 minutes of WAV is 74 MB against Telegram's 50 MB limit, so FFmpeg always builds an MP3.
Stack
Need a system like this?
Describe your process and I'll tell you what to automate first. Free.