Generative media pipeline
Script to synthesised speech to rendered video. The interesting engineering is not the generation, it is staying reliable on top of third-party APIs that are slow, asynchronous and fail more often than their documentation implies.
The problem
- Sector
- Automated content production
- Flow
- Text · speech · video
- Pattern
- Provider adapter
- Role
- Sole engineer
Generative media APIs are long-running and unreliable in ordinary operation. A render takes minutes, the job can fail halfway, and a retry issued carelessly bills you twice for the same video. Cost and correctness both depend on machinery that has nothing to do with the models.
So the pipeline is built around job state, backoff, duplicate prevention and a dead-letter path, with the speech provider isolated behind an interface so it can be changed on price or quality without touching the video stage.
Architecture
Decisions worth defending
The provider is swappable by design
The speech step sits behind a small interface, so changing voice engine means changing configuration rather than rewriting the video path. In a market where pricing and quality move every quarter, that is the difference between switching on a benchmark and switching on a project plan.
Duplicate prevention before retry logic
Retries are mandatory when renders fail, and they are dangerous when each attempt costs money. Keying every job so a repeat is recognised has to come before the retry mechanism, not after the first double bill.
A dead-letter state, not silent failure
After the retry budget is spent the job parks with full context rather than disappearing. Someone can look at it, understand why, and requeue it. Jobs that vanish quietly are how a content pipeline loses a day's output without anyone noticing.
Stack
Generation
- Avatar video API
- Neural speech synthesis
- External voice provider
- Asset upload
Reliability
- Job state machine
- Exponential backoff
- Duplicate prevention
- Dead-letter handling
Runtime
- Python
- Standard library only
- Bounded polling
- Zero-cost dry run
Output
- Rendered MP4
- Committed samples
- Per-path comparison