Case study · February 2026
Serverless Text-to-Speech Pipeline
AWS Polly · GitHub Actions · S3 · Python · boto3
The scenario
A training scenario for a fictional client, Pixel Learning Co., that wanted course content stored in GitHub converted to audio for visually impaired users and for people who learn by listening. The brief was a lightweight, cost-effective pipeline that stays fully serverless.
How it works
- A Python script using boto3 sends
speech.txtto Amazon Polly and writesexample.mp3. - Two GitHub Actions workflows run it.
on_pr.ymlruns when a pull request targetsmainand uploads to abeta/folder in S3.on_merge.ymlruns on a push tomainand uploads toprod/. - Each run overwrites the file in its folder, and I verified the result in the S3 console, with the AWS CLI and in the Actions log.
What I learned
- Polly’s synchronous synthesis is limited to 3,000 characters. Longer content needs the asynchronous
start_speech_synthesis_taskmethod. - Most of my errors were in the YAML: spelling, indentation and naming consistency across two near-identical workflow files. One of them had the same command typed twice and a missing file name in the production upload.