Synthesize speech in documents
🤖/text/speak synthesizes speech in documents.

You can use the audio that we return in your application, or you can pass the audio down to other Robots to add a voice track to a video for example.
Another common use case is making your product accessible to people with a reading disability.
Usage example
Synthesize speech from uploaded text documents, using a female voice in American English:
{
"steps": {
"synthesized": {
"robot": "/text/speak",
"use": ":original",
"provider": "aws",
"voice": "female-1",
"target_language": "en-US"
}
}
}Parameters
interpolateboolean | Record<string, boolean>Controls whether Assembly Variables are interpolated for individual instruction fields.
By default, most Robot instruction fields interpolate Assembly Variables. Set this to
falseto treat every instruction field as literal text, or set an individual field path tofalseto treat only that field as literal text. For Robot-specific fields that are literal by default, set this totrueor set that field path totrueto opt back into interpolation.Use field names such as
path, or dotted paths such asffmpeg.vffor nested objects.output_metaRecord<string, boolean> | boolean | Array<string>Allows you to specify a set of metadata that is more expensive on CPU power to calculate, and thus is disabled by default to keep your Assemblies processing fast.
For images, you can add
"has_transparency": truein this object to extract if the image contains transparent parts and"dominant_colors": trueto extract an array of hexadecimal color codes from the image.For images, you can also add
"blurhash": trueto extract a BlurHash string — a compact representation of a placeholder for the image, useful for showing a blurred preview while the full image loads.For videos, you can add the
"colorspace": trueparameter to extract the colorspace of the output video.For videos, you can also add
"interlaced": trueto detect whether the video is interlaced. This combines the cheap ffprobefield_orderflag with a boundedidetsampling pass over the first frames of the source, exposinginterlaced,field_order, and a diagnosticinterlace_detectionobject underfile.meta. This is computationally expensive and billed accordingly.For audio, you can add
"mean_volume": trueto get a single value representing the mean average volume of the audio file.You can also set this to
falseto skip metadata extraction and speed up transcoding.user_metaRecord<string, any>(default:{})Adds custom metadata to each file emitted by this Robot without modifying the file’s contents.
The values are merged with any existing
user_metacarried by the input file. If both objects contain the same key, this Robot’s value takes precedence. Assembly Variables are supported, for example{ "internal_file_id": "${file.id}" }.resultboolean(default:false)Whether the results of this Step should be present in the Assembly Status JSON
queuebatchSetting the queue to 'batch', manually downgrades the priority of jobs for this step to avoid consuming Priority job slots for jobs that don't need zero queue waiting times
force_acceptboolean(default:false)Force a Robot to accept a file type it would have ignored.
By default, Robots ignore files they are not familiar with. 🤖/video/encode, for example, will happily ignore input images.
With the
force_acceptparameter set totrue, you can force Robots to accept all files thrown at them. This will typically lead to errors and should only be used for debugging or combatting edge cases.ignore_errorsboolean | Array<meta | execute>(default:[])Ignore errors during specific phases of processing.
Setting this to
["meta"]will cause the Robot to ignore errors during metadata extraction.Setting this to
["execute"]will cause the Robot to ignore errors during the main execution phase.Setting this to
trueis equivalent to["meta", "execute"]and will ignore errors in both phases.usestring | Array<string> | Array<object> | objectSpecifies which Step(s) to use as input.
- You can pick any names for Steps except
":original"(reserved for user uploads handled by Transloadit) - You can provide several Steps as input with arrays:
{ "use": [ ":original", "encoded", "resized" ] } - You can also tag input Steps with
asto pass semantic intent to robots:{ "use": [ { "name": ":original", "as": "image" }, { "name": ":original", "as": "mask" } ] }
TipThat's likely all you need to know about
use, but you can view Advanced use cases.- You can pick any names for Steps except
promptstring | nullWhich text to speak. You can also set this to
nulland supply an input text file.providerauto | aws | gcp(default:"auto")Chooses the best provider based on your request.
Set this to
"aws"or"gcp"to force a specific provider.target_languagestring(default:"en-US")The written language of the document. This will also be the language of the spoken text.
The language should be specified in the BCP-47 format, such as
"en-GB","de-DE"or"fr-FR". Please consult the list of supported languages and voices.voicefemale-1 | female-2 | female-3 | female-child-1 | male-1 | male-child-1(default:"female-1")The gender to be used for voice synthesis. Please consult the list of supported languages and voices.
ssmlboolean(default:false)Supply Speech Synthesis Markup Language instead of raw text, in order to gain more control over how your text is voiced, including rests and pronounciations.