{"id":68828,"date":"2019-08-02T04:19:55","date_gmt":"2019-08-02T11:19:55","guid":{"rendered":"http:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/"},"modified":"2019-08-02T04:19:55","modified_gmt":"2019-08-02T11:19:55","slug":"text-to-speech-with-aws","status":"publish","type":"post","link":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/","title":{"rendered":"Text To Speech With AWS"},"content":{"rendered":"<link rel=\"canonical\" href=\"https:\/\/www.smashingmagazine.com\/2019\/08\/text-to-speech-aws\/\"><title>Text To Speech With AWS<\/title><\/p>\n<article>\n<header>\n<h1>Text To Speech With AWS<\/h1>\n<address>Philip Kiely<\/address>\n<p>                  <time datetime=\"2019-08-01T14:00:00+02:00\">2019-08-01T14:00:00+02:00<\/time><time datetime=\"2019-08-02T11:07:55+00:00\">2019-08-02T11:07:55+00:00<\/time><\/header>\n<p>This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a spoken .mp3 file to give more options to blind and dyslexic users of your site.<\/p>\n<p>In the next article, we will embark on the return journey, from speech to text, and consider the accuracy of these transcriptions by sending various samples through a round-trip translation. To follow these tutorials, you will need an AWS account with billing enabled, though the tutorials will stay well within the constraints of free-tier resources. Examples will focus on using the AWS console, but I will also demonstrate the AWS CLI (Command Line Interface), which requires basic command line knowledge.<\/p>\n<h3>Introduction And Motivation<\/h3>\n<p>Most of the internet is text-based. Text is lightweight (1 byte per letter), widely supported, easy to interpret, and has a precedent as old as the internet as the default medium of online communication. Sending written text predates the internet: telegraphs carried text over wires hundreds of years ago and physical mail has transmitted writing for centuries. Voice transmission over radio and telephone also predates the internet, but did not translate to the same foundational medium that text did online. This is in almost all cases a good thing, again, text is lightweight and easy to interpret compared to audio. However, transforming between voice and text can add powerful functionality to and improve the accessibility of a wide variety of applications.<\/p>\n<p>It has always been possible to transform between audio and text, you can read a written speech or transcribe an oral sermon. Indeed, if we think back to the telegram, trained operators transcoded Morse Code messages to words. In each example, it has always been very labor intensive to move from speech to writing or back, even with specialized training and equipment. With a variety of cloud services, we can automate these processes to allow transitioning between mediums in seconds without any human effort, which expands the possible use cases.<\/p>\n<p>The most obvious benefit of implementing appropriate text to speech and speech to text options is accessibility. A visually impaired or dyslexic user would benefit from a narrated version of an article, while a deaf person could become a member of your podcasting audience by reading a transcript of the show.<\/p>\n<div data-component=\"FeaturePanel\" data-audience=\"non-subscriber\" data-remove=\"true\"><\/div>\n<h3>Text to Speech Project<\/h3>\n<p>Say you wanted to add narrated versions of every post to your blog. You could purchase a microphone and invest hours into recording and editing spoken renditions of each post. This would result in a superior listener experience, but if you want most of the benefit for only a couple of minutes and a few pennies per post, consider using AWS instead. If you are the sort of person who regularly updates and revises older or evergreen content, this method also helps you keep the spoken version up to date with minimal effort.<\/p>\n<p>We will begin with text to speech using <em>Amazon Polly<\/em>. For simple exploration, AWS provides a graphical user interface through its online console. After logging in to your AWS account, use the \u201cServices\u201d menu to find \u201cAmazon Polly\u201d or go to <a href=\"https:\/\/us-east-1.console.aws.amazon.com\/polly\/home\/SynthesizeSpeech\">https:\/\/us-east-1.console.aws.amazon.com\/polly\/home\/SynthesizeSpeech<\/a>.<\/p>\n<h4>Using the Polly Console<\/h4>\n<figure><a href=\"https:\/\/i0.wp.com\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/43ecd49d-9684-4942-81ff-20dd435fc40f\/pollyconsole.png?ssl=1\"><\/p>\n<p>    <img data-recalc-dims=\"1\" decoding=\"async\" srcset=\"http:\/\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/pollyconsole.png 400w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_800\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/43ecd49d-9684-4942-81ff-20dd435fc40f\/pollyconsole.png 800w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1200\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/43ecd49d-9684-4942-81ff-20dd435fc40f\/pollyconsole.png 1200w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1600\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/43ecd49d-9684-4942-81ff-20dd435fc40f\/pollyconsole.png 1600w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_2000\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/43ecd49d-9684-4942-81ff-20dd435fc40f\/pollyconsole.png 2000w\" src=\"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/pollyconsole.png?w=900\" sizes=\"100vw\" alt=\"Amazon Polly Console\"><\/a><figcaption>\n      Amazon Polly provides a console to perform text-to-speech operations. (<a href=\"https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/43ecd49d-9684-4942-81ff-20dd435fc40f\/pollyconsole.png\">Large preview<\/a>)<br \/>\n    <\/figcaption><\/figure>\n<p>You can use the Amazon Polly console to read 3,000 characters (about 500 words) and get an audio stream or immediate download. If you need up to 100,000 characters (about 16,600 words) read, your only option is to have AWS store the result in S3 after it has finished processing, which can take a couple of minutes. At the time of writing, Amazon Polly does not support inputs of over 100,000 billable characters, if you want to convert a longer text like a book you will most likely have to do so in chunks and concatenate the audio files yourself.<\/p>\n<p>A \u201cbillable character\u201d is one that the service actually pronounces. Specifically, that means that SSML tags are not billable characters, which we will cover later. For your first year of using Amazon Polly, you get 5 million billable characters per month for free, which is more than enough to run the examples from this article and do your own experimentation. Beyond that, Amazon Polly costs four dollars per million billable characters at the time of writing, meaning that converting a standard-length novel would cost about two dollars.<\/p>\n<p>The console also allows you to change the language, region, and voice of the reader. Though this article only covers English, at the time of writing AWS supports 21 languages and 29 distinct language-region pairs. While most regions only have one or two voices, popular ones like United States English have several options to chose between.<\/p>\n<blockquote>\n<p>\n    <a aria-label=\"Share on Twitter\" href=\"http:\/\/twitter.com\/share?text=Amazon%20Polly%20narrated%20text%20is%20very%20obviously%20read%20by%20a%20robot,%20but%20the%20resulting%20audio%20is%20quite%20listenable.%0A&#038;url=https:\/\/smashingmagazine.com%2F2019%2F08%2Ftext-to-speech-aws%2F\"><br \/>\n      Amazon Polly narrated text is very obviously read by a robot, but the resulting audio is quite listenable.<\/p>\n<p>    <\/a>\n  <\/p>\n<div>\n<div>\n      <span>\u201c<\/span><\/div>\n<\/p><\/div>\n<\/blockquote>\n<p>I often prefer to use the UK English voice \u201cBrian.\u201d To my American ears, the British accent covers some of the inflections in robotic speech and makes for a smoother listening experience. To be clear, Amazon Polly narrated text is very obviously read by a robot, but the resulting audio is quite listenable.<\/p>\n<p>It is significantly better than the built-in reader that the MacOS <code>say<\/code> terminal command uses, and is comparable to the speech quality of voice assistants like Siri and Alexa.<\/p>\n<h4>Writing SSML<\/h4>\n<p>If you want full control over the resultant speech, you can take the time to tag your input with SSML. <a href=\"https:\/\/docs.aws.amazon.com\/polly\/latest\/dg\/supported-ssml.html#supportedtags\">SSML<\/a> (Speech Synthesis Markup Language) is a standardized language for representing verbal cues in text. Like HTML, XML, and other markup languages, it uses opening and closing tags. Amazon Polly supports SSML input, and tags do not count as \u201cbillable characters.\u201d Alexa skills also use SSML for pre-programmed responses, so it is a worthwhile language to know.<\/p>\n<p>The foundational tag, <code><speak><\/code>, wraps everything that you want read. Like HTML, use <code><\/p>\n<p><\/code> to divide paragraphs, which results in a significant pause in the narration. Smaller pauses come from punctuation, and you always have the option to insert pauses of up to ten seconds with <code><break><\/code>.<\/p>\n<p>SSML provides <code><say-as><\/code>, a very flexible tag that supports everything from pronouncing phone numbers to censoring expletives using the <code>interpret-as<\/code> argument. Consider the options from this tag with the following sample.<\/p>\n<div>\n<pre><code><speak>\nCall 5551230987 by 11'00\" PM to get tips on writing clean JavaScript.<break time=\"1s\"\/>\nCall <say-as interpret-as=\"telephone\">5551230987<\/say-as> by 11'00\" PM to get tips on writing clean <say-as interpret-as=\"expletive\">JavaScript<\/say-as>\n<\/speak><\/code><\/pre>\n<\/div>\n<p>Further flexibility comes from the <code><prosody><\/code> tag, which provides you with control over the rate, pitch, and volume of speech. Unfortunately, at the time of writing Polly does not support the <code><voice><\/code> tag, which Alexa skills can use to speak in multiple standard voices, but does support the <code><lang><\/code> tag that allows voices in one language to correctly pronounce words from other languages. In this example, <code><lang><\/code> corrects the pronunciation of \u201ctag\u201d from American to German.<\/p>\n<div>\n<pre><code><speak>\n    Guten tag, where is the airport?<break time=\"1s\"\/>\n    <lang xml:lang=\"de-DE\">Guten tag<\/lang>, where is the airport>\n<\/speak><\/code><\/pre>\n<\/div>\n<div><\/div>\n<p>Finally, if you want to customize pronunciation within a language, Amazon Polly supports the <code><phoneme><\/code> tag.<\/p>\n<blockquote>\n<p>\n    <a aria-label=\"Share on Twitter\" href=\"http:\/\/twitter.com\/share?text=My%20last%20name,%20Kiely,%20is%20spelled%20differently%20than%20it%20is%20pronounced.%20Using%20the%20x-sampa%20alphabet,%20I%20am%20able%20to%20specify%20the%20correct%20pronunciation.%0A&#038;url=https:\/\/smashingmagazine.com%2F2019%2F08%2Ftext-to-speech-aws%2F\"><br \/>\n      My last name, Kiely, is spelled differently than it is pronounced. Using the x-sampa alphabet, I am able to specify the correct pronunciation.<\/p>\n<p>    <\/a>\n  <\/p>\n<div>\n<div>\n      <span>\u201c<\/span><\/div>\n<\/p><\/div>\n<\/blockquote>\n<div>\n<pre><code><speak>\n    Philip Kiely<break time=\"1s\"\/>\n    Philip <phoneme alphabet=\"x-sampa\" ph=\"?kaI.li\">Kiely<\/phoneme>\n<\/speak><\/code><\/pre>\n<\/div>\n<p>This is not an exhaustive list of the customization options available with SSML. For a complete reference, <a href=\"https:\/\/docs.aws.amazon.com\/polly\/latest\/dg\/supported-ssml.html#supportedtags\">visit the documentation<\/a>.<\/p>\n<h4>Writing Lexicons<\/h4>\n<p>If you want to specify a consistent custom pronunciation or expand an abbreviation without tagging each instance with a phoneme tag, or you are using plain text instead of SSML, Amazon Polly supports lexicons of custom pronunciations. You can apply up to five lexicons of up to 4,000 characters each per language to a narration, though larger lexicons increase the processing time.<\/p>\n<p>As with before, I want to make sure that Amazon Polly says my name correctly, but this time I want to do so without using SSML. I wrote the following lexicon:<\/p>\n<div>\n<pre><code><?xml version=\"1.0\" encoding=\"UTF-8\"?>\n<lexicon version=\"1.0\"\n      xmlns=\"http:\/\/www.w3.org\/2005\/01\/pronunciation-lexicon\"\n      xmlns:xsi=\"http:\/\/www.w3.org\/2001\/XMLSchema-instance\"\n      xsi:schemaLocation=\"http:\/\/www.w3.org\/2005\/01\/pronunciation-lexicon\n        http:\/\/www.w3.org\/TR\/2007\/CR-pronunciation-lexicon-20071212\/pls.xsd\"\n      alphabet=\"x-sampa\"\n      xml:lang=\"en-US\">\n  <lexeme><grapheme>Kiely<\/grapheme><alias>?kaIli<\/alias><\/lexeme>\n<\/lexicon><\/code><\/pre>\n<\/div>\n<p>The <code><?xml?><\/code> header and <code><lexicon><\/code> tag will stay mostly constant between lexicons, though the <code><lexicon><\/code> tag supports two important arguments. The first, <code>alphabet<\/code>, lets you choose between x-sampa and ipa, two standard pronunciation alphabets. I prefer x-sampa because it uses standard ASCII characters, so I am unlikely to encounter encoding issues. The <code>xml:lang<\/code> argument lets you specify language and region. A lexicon is only usable by a voice from that language and region.<\/p>\n<p>The lexicon itself is a sequence of <code><lexeme><\/code> tags. Each one contains a <code><grapheme><\/code> tag, which contains the original text, and the <code><alias><\/code> tag, which describes what you want said instead. Aliases go beyond pronunciation, you can use them for expanding abbreviations (\u201cJr\u201d becomes \u201cJunior\u201d) or replacing words (\u201cBruce Wayne\u201d becomes \u201cBatman\u201d). A lexicon can have as many lexeme tags as it can fit in the 4,000 character limit.<\/p>\n<figure><a href=\"https:\/\/i0.wp.com\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/cc5215fc-ca40-4574-923c-58d5da1f3422\/lexicons.png?ssl=1\"><\/p>\n<p>    <img data-recalc-dims=\"1\" decoding=\"async\" srcset=\"http:\/\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/lexicons.png 400w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_800\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/cc5215fc-ca40-4574-923c-58d5da1f3422\/lexicons.png 800w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1200\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/cc5215fc-ca40-4574-923c-58d5da1f3422\/lexicons.png 1200w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1600\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/cc5215fc-ca40-4574-923c-58d5da1f3422\/lexicons.png 1600w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_2000\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/cc5215fc-ca40-4574-923c-58d5da1f3422\/lexicons.png 2000w\" src=\"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/lexicons.png?w=900\" sizes=\"100vw\" alt=\"Amazon Polly Console with lexicon loaded\"><\/a><figcaption>\n      The included lexicon will modify the pronunciation of the input text. (<a href=\"https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/cc5215fc-ca40-4574-923c-58d5da1f3422\/lexicons.png\">Large preview<\/a>)<br \/>\n    <\/figcaption><\/figure>\n<p>The screenshot shows the plain text that would be mispronounced and the applied lexicon. Use the \u201cCustomize Pronunciation\u201d menu to select up to five uploaded lexicons, uploaded from the left navbar tab \u201cLexicons.\u201d Listening to the speech verifies that my name is said correctly.<\/p>\n<p>Now that we have full control over the resultant speech, let\u2019s consider how to save the output for use in our application.<\/p>\n<h4>Saving and loading from S3<\/h4>\n<p>If you want to re-use spoken text in your application, you\u2019ll want to choose the \u201cSynthesize to S3\u201d option in the Amazon Polly console. In this example, I am using the voice \u201cBrian\u201d to perform a surprisingly capable reading of <a href=\"http:\/\/shakespeare.mit.edu\/Poetry\/sonnet.XXIX.html\">Shakespeare\u2019s sonnet XXIX<\/a>. We begin by copying in the poem as plain text and selecting \u201cSynthesize to S3,\u201d which launches the following modal.<\/p>\n<figure><a href=\"https:\/\/i0.wp.com\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/bbdfdb18-5e85-4162-bf3c-2c8a64e948d3\/s3synthesizemodal.png?ssl=1\"><\/p>\n<p>    <img data-recalc-dims=\"1\" decoding=\"async\" srcset=\"http:\/\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/s3synthesizemodal.png 400w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_800\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/bbdfdb18-5e85-4162-bf3c-2c8a64e948d3\/s3synthesizemodal.png 800w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1200\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/bbdfdb18-5e85-4162-bf3c-2c8a64e948d3\/s3synthesizemodal.png 1200w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1600\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/bbdfdb18-5e85-4162-bf3c-2c8a64e948d3\/s3synthesizemodal.png 1600w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_2000\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/bbdfdb18-5e85-4162-bf3c-2c8a64e948d3\/s3synthesizemodal.png 2000w\" src=\"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/s3synthesizemodal.png?w=900\" sizes=\"100vw\" alt=\"S3 Synthesize Modal\"><\/a><figcaption>\n      The &#8216;Synthesize to S3&#8217; button gives you options for where to save the resultant file. (<a href=\"https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/bbdfdb18-5e85-4162-bf3c-2c8a64e948d3\/s3synthesizemodal.png\">Large preview<\/a>)<br \/>\n    <\/figcaption><\/figure>\n<p>S3 buckets have globally unique names, and you can enter any S3 bucket that you own or have the appropriate permissions to. Make sure the bucket allows for making its contents public, as that will be required in a future step. You should also set a \u201cS3 key prefix,\u201d which is a string that will help you identify the output in the bucket. After clicking Synthesize and giving it a moment to process, we navigate to the S3 bucket that we synthesized the speech into.<\/p>\n<figure><a href=\"https:\/\/i0.wp.com\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/4a633f47-262e-4b10-8880-ec6fa8c6cacf\/s3bucket.png?ssl=1\"><\/p>\n<p>    <img data-recalc-dims=\"1\" decoding=\"async\" srcset=\"http:\/\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/s3bucket.png 400w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_800\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/4a633f47-262e-4b10-8880-ec6fa8c6cacf\/s3bucket.png 800w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1200\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/4a633f47-262e-4b10-8880-ec6fa8c6cacf\/s3bucket.png 1200w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1600\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/4a633f47-262e-4b10-8880-ec6fa8c6cacf\/s3bucket.png 1600w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_2000\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/4a633f47-262e-4b10-8880-ec6fa8c6cacf\/s3bucket.png 2000w\" src=\"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/s3bucket.png?w=900\" sizes=\"100vw\" alt=\"S3 Bucket main page\"><\/a><figcaption>\n      A S3 bucket stores your project&#8217;s files. (<a href=\"https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/4a633f47-262e-4b10-8880-ec6fa8c6cacf\/s3bucket.png\">Large preview<\/a>)<br \/>\n    <\/figcaption><\/figure>\n<p>The arrow points to the entry in the bucket that we just created. Selecting that item will bring us to the following page.<\/p>\n<figure><a href=\"https:\/\/i0.wp.com\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/fd0d3826-3246-43bb-a2dc-f63a08bfffaa\/s3makepublic.png?ssl=1\"><\/p>\n<p>    <img data-recalc-dims=\"1\" decoding=\"async\" srcset=\"http:\/\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/s3makepublic.png 400w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_800\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/fd0d3826-3246-43bb-a2dc-f63a08bfffaa\/s3makepublic.png 800w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1200\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/fd0d3826-3246-43bb-a2dc-f63a08bfffaa\/s3makepublic.png 1200w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_1600\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/fd0d3826-3246-43bb-a2dc-f63a08bfffaa\/s3makepublic.png 1600w,\n\t\t\t        https:\/\/res.cloudinary.com\/indysigner\/image\/fetch\/f_auto,q_auto\/w_2000\/https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/fd0d3826-3246-43bb-a2dc-f63a08bfffaa\/s3makepublic.png 2000w\" src=\"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/s3makepublic.png?w=900\" sizes=\"100vw\" alt=\"S3 Bucket file view\"><\/a><figcaption>\n      For each file, you can make it public using this button. (<a href=\"https:\/\/cloud.netlifyusercontent.com\/assets\/344dbf88-fdf9-42bb-adb4-46f01eedd629\/fd0d3826-3246-43bb-a2dc-f63a08bfffaa\/s3makepublic.png\">Large preview<\/a>)<br \/>\n    <\/figcaption><\/figure>\n<p>Follow the arrow to select the \u201cMake Public\u201d option, which will make the file accessible to anyone with a link. Scroll down and copy the link and use it in your application. For example, you can <a href=\"https:\/\/philipkiely-essays-spoken.s3.amazonaws.com\/sonnetxxix.9175c38c-9635-450f-a0a9-2267dcdca8ec.mp3\">download the poem here<\/a>. For many applications, you may wish to pass the url to an html <code><audio><\/code> tag to allow for web playback.<\/p>\n<div><\/div>\n<p>We have covered every necessary component for transforming text to speech on AWS. Next, we turn our attention to a more advanced interface that can provide automation potential and save time.<\/p>\n<h3>Using the AWS CLI<\/h3>\n<p>Back to our hypothetical blog post. The simplest workflow would be to take the final written version of each article, copy it into the console, click the \u201cSynthesize to S3 button,\u201d and embed a download link to the resultant .mp3 file in the blog. Honestly, this is a pretty decent workflow; it is exactly what I do for my personal website. However, AWS offers another option: the AWS CLI.<\/p>\n<p>Make sure that you have <a href=\"https:\/\/docs.aws.amazon.com\/cli\/latest\/userguide\/cli-chap-install.html\">installed<\/a> and <a href=\"https:\/\/docs.aws.amazon.com\/cli\/latest\/userguide\/cli-chap-configure.html\">configured<\/a> the AWS CLI appropriately. Begin by entering <code>aws polly help<\/code> to make sure that Polly is available and to read a list of supported commands. For troubleshooting, see the <a href=\"https:\/\/docs.aws.amazon.com\/polly\/latest\/dg\/setup-aws-cli.html\">documentation<\/a>.<\/p>\n<p>To perform a conversion from the command line, I first copied the <a href=\"http:\/\/shakespeare.mit.edu\/Poetry\/sonnet.XXIX.html\">poem from earlier<\/a> into a .txt file. I then ran the following command in terminal (MacOS\/Linux):<\/p>\n<pre><code>aws polly synthesize-speech \n    --output-format mp3 \n    --voice-id Joanna \n    --text \"`cat sonnetxxix.txt`\" \n    poem.mp3\n<\/code><\/pre>\n<p>In a few seconds, the resulting .mp3 file was downloaded to my machine, ready for inclusion in my CMS or other application. Note the special characters around the <code>--text<\/code> argument, this passes the contents of the file rather than just the file name.<\/p>\n<p>Finally, for more advanced applications, Amazon Polly has <a href=\"https:\/\/aws.amazon.com\/polly\/resources\/\">an SDK for 9 languages\/platforms<\/a>. The SDK would be overkill for these examples, but is exactly what you want for automating Amazon Polly calls, especially in response to user actions.<\/p>\n<h3>Conclusion<\/h3>\n<p>Text to speech can help you create more versatile, accessible content. Beginning in the Amazon Polly console, we can transform up to 100,000 billable characters in plain text or SSML, make the resulting .mp3 file public, and use that file in an application. We can use the AWS CLI for automation and more convenient access.<\/p>\n<p>Stay tuned for the second installment of the series, we will convert media in the other direction, from speech to text, and consider the benefits and challenges of doing so. Part two will build on the technologies that we have used so far and introduce Amazon Transcribe.<\/p>\n<h4>Further Reference<\/h4>\n<ul>\n<li><a href=\"https:\/\/aws.amazon.com\/polly\/\">AWS Polly<\/a><\/li>\n<li><a href=\"https:\/\/aws.amazon.com\/s3\/\">AWS S3<\/a><\/li>\n<li><a href=\"https:\/\/docs.aws.amazon.com\/polly\/latest\/dg\/supported-ssml.html#supportedtags\">SSML Reference<\/a><\/li>\n<li><a href=\"https:\/\/docs.aws.amazon.com\/polly\/latest\/dg\/managing-lexicons-console.html#managing-lexicons-console-synthesize-speech\">Managing Lexicons<\/a><\/li>\n<\/ul>\n<div>\n  <img data-recalc-dims=\"1\" decoding=\"async\" src=\"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/logo-red-1.png?w=900\" alt=\"Smashing Editorial\"><span>(yk,ra)<\/span>\n<\/div>\n<\/article>\n<p class=\"wpematico_credit\"><small>Powered by <a href=\"http:\/\/www.wpematico.com\" target=\"_blank\">WPeMatico<\/a><\/small><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Text To Speech With AWS Text To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00 This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content&#8230;<a class=\"moretag\" href=\"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/\"> Read the full article&#8230;<\/a><\/p>\n","protected":false},"author":4,"featured_media":68829,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[371],"tags":[],"class_list":["post-68828","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-user-experience"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 5.0.0.1 - aioseo.com -->\n\t<meta name=\"description\" content=\"Text To Speech With AWSText To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<meta name=\"author\" content=\"Guest Contribution\"\/>\n\t<meta name=\"keywords\" content=\"user experience\" \/>\n\t<link rel=\"canonical\" href=\"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 5.0.0.1\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"Design for Immersive Technologies | User Experience for Games, Virtual Reality (VR) and Augmented\/Mixed Reality (AR\/MR)\" \/>\n\t\t<meta property=\"og:type\" content=\"article\" \/>\n\t\t<meta property=\"og:title\" content=\"Text To Speech With AWS | Design for Immersive Technologies\" \/>\n\t\t<meta property=\"og:description\" content=\"Text To Speech With AWSText To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/\" \/>\n\t\t<meta property=\"article:published_time\" content=\"2019-08-02T11:19:55+00:00\" \/>\n\t\t<meta property=\"article:modified_time\" content=\"2019-08-02T11:19:55+00:00\" \/>\n\t\t<meta name=\"twitter:card\" content=\"summary\" \/>\n\t\t<meta name=\"twitter:title\" content=\"Text To Speech With AWS | Design for Immersive Technologies\" \/>\n\t\t<meta name=\"twitter:description\" content=\"Text To Speech With AWSText To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a\" \/>\n\t\t<script type=\"application\/ld+json\" class=\"aioseo-schema\">\n\t\t\t{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#article\",\"name\":\"Text To Speech With AWS | Design for Immersive Technologies\",\"headline\":\"Text To Speech With AWS\",\"author\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/author\\\/guestcontribution\\\/#author\"},\"publisher\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/#organization\"},\"image\":{\"@type\":\"ImageObject\",\"url\":\"https:\\\/\\\/i0.wp.com\\\/www.taterboy.com\\\/blog\\\/wp-content\\\/uploads\\\/2019\\\/08\\\/pollyconsole.png?fit=400%2C240&ssl=1\",\"width\":400,\"height\":240},\"datePublished\":\"2019-08-02T04:19:55-07:00\",\"dateModified\":\"2019-08-02T04:19:55-07:00\",\"inLanguage\":\"en-US\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#webpage\"},\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#webpage\"},\"articleSection\":\"User Experience\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#breadcrumblist\",\"itemListElement\":[{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog#listItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.taterboy.com\\\/blog\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/category\\\/user-experience\\\/#listItem\",\"name\":\"User Experience\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/category\\\/user-experience\\\/#listItem\",\"position\":2,\"name\":\"User Experience\",\"item\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/category\\\/user-experience\\\/\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#listItem\",\"name\":\"Text To Speech With AWS\"},\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog#listItem\",\"name\":\"Home\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#listItem\",\"position\":3,\"name\":\"Text To Speech With AWS\",\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/category\\\/user-experience\\\/#listItem\",\"name\":\"User Experience\"}}]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/#organization\",\"name\":\"Design for Immersive Technologies\",\"description\":\"User Experience for Games, Virtual Reality (VR) and Augmented\\\/Mixed Reality (AR\\\/MR)\",\"url\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/author\\\/guestcontribution\\\/#author\",\"url\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/author\\\/guestcontribution\\\/\",\"name\":\"Guest Contribution\",\"image\":{\"@type\":\"ImageObject\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#authorImage\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d2aad73eb0f48c67b141e9fc978a0e498969791ed623346e361d02d19775b0bb?s=96&d=mm&r=g\",\"width\":96,\"height\":96,\"caption\":\"Guest Contribution\"}},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#webpage\",\"url\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/\",\"name\":\"Text To Speech With AWS | Design for Immersive Technologies\",\"description\":\"Text To Speech With AWSText To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a\",\"inLanguage\":\"en-US\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/#website\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#breadcrumblist\"},\"author\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/author\\\/guestcontribution\\\/#author\"},\"creator\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/author\\\/guestcontribution\\\/#author\"},\"image\":{\"@type\":\"ImageObject\",\"url\":\"https:\\\/\\\/i0.wp.com\\\/www.taterboy.com\\\/blog\\\/wp-content\\\/uploads\\\/2019\\\/08\\\/pollyconsole.png?fit=400%2C240&ssl=1\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#mainImage\",\"width\":400,\"height\":240},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/2019\\\/08\\\/text-to-speech-with-aws\\\/#mainImage\"},\"datePublished\":\"2019-08-02T04:19:55-07:00\",\"dateModified\":\"2019-08-02T04:19:55-07:00\"},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/\",\"name\":\"Design for Immersive Technologies\",\"description\":\"User Experience for Games, Virtual Reality (VR) and Augmented\\\/Mixed Reality (AR\\\/MR)\",\"inLanguage\":\"en-US\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.taterboy.com\\\/blog\\\/#organization\"}}]}\n\t\t<\/script>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"Text To Speech With AWS | Design for Immersive Technologies","description":"Text To Speech With AWSText To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a","canonical_url":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/","robots":"max-image-preview:large","keywords":"user experience","webmasterTools":{"miscellaneous":""},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#article","name":"Text To Speech With AWS | Design for Immersive Technologies","headline":"Text To Speech With AWS","author":{"@id":"https:\/\/www.taterboy.com\/blog\/author\/guestcontribution\/#author"},"publisher":{"@id":"https:\/\/www.taterboy.com\/blog\/#organization"},"image":{"@type":"ImageObject","url":"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/pollyconsole.png?fit=400%2C240&ssl=1","width":400,"height":240},"datePublished":"2019-08-02T04:19:55-07:00","dateModified":"2019-08-02T04:19:55-07:00","inLanguage":"en-US","mainEntityOfPage":{"@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#webpage"},"isPartOf":{"@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#webpage"},"articleSection":"User Experience"},{"@type":"BreadcrumbList","@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#breadcrumblist","itemListElement":[{"@type":"ListItem","@id":"https:\/\/www.taterboy.com\/blog#listItem","position":1,"name":"Home","item":"https:\/\/www.taterboy.com\/blog","nextItem":{"@type":"ListItem","@id":"https:\/\/www.taterboy.com\/blog\/category\/user-experience\/#listItem","name":"User Experience"}},{"@type":"ListItem","@id":"https:\/\/www.taterboy.com\/blog\/category\/user-experience\/#listItem","position":2,"name":"User Experience","item":"https:\/\/www.taterboy.com\/blog\/category\/user-experience\/","nextItem":{"@type":"ListItem","@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#listItem","name":"Text To Speech With AWS"},"previousItem":{"@type":"ListItem","@id":"https:\/\/www.taterboy.com\/blog#listItem","name":"Home"}},{"@type":"ListItem","@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#listItem","position":3,"name":"Text To Speech With AWS","previousItem":{"@type":"ListItem","@id":"https:\/\/www.taterboy.com\/blog\/category\/user-experience\/#listItem","name":"User Experience"}}]},{"@type":"Organization","@id":"https:\/\/www.taterboy.com\/blog\/#organization","name":"Design for Immersive Technologies","description":"User Experience for Games, Virtual Reality (VR) and Augmented\/Mixed Reality (AR\/MR)","url":"https:\/\/www.taterboy.com\/blog\/"},{"@type":"Person","@id":"https:\/\/www.taterboy.com\/blog\/author\/guestcontribution\/#author","url":"https:\/\/www.taterboy.com\/blog\/author\/guestcontribution\/","name":"Guest Contribution","image":{"@type":"ImageObject","@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#authorImage","url":"https:\/\/secure.gravatar.com\/avatar\/d2aad73eb0f48c67b141e9fc978a0e498969791ed623346e361d02d19775b0bb?s=96&d=mm&r=g","width":96,"height":96,"caption":"Guest Contribution"}},{"@type":"WebPage","@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#webpage","url":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/","name":"Text To Speech With AWS | Design for Immersive Technologies","description":"Text To Speech With AWSText To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a","inLanguage":"en-US","isPartOf":{"@id":"https:\/\/www.taterboy.com\/blog\/#website"},"breadcrumb":{"@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#breadcrumblist"},"author":{"@id":"https:\/\/www.taterboy.com\/blog\/author\/guestcontribution\/#author"},"creator":{"@id":"https:\/\/www.taterboy.com\/blog\/author\/guestcontribution\/#author"},"image":{"@type":"ImageObject","url":"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/pollyconsole.png?fit=400%2C240&ssl=1","@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#mainImage","width":400,"height":240},"primaryImageOfPage":{"@id":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/#mainImage"},"datePublished":"2019-08-02T04:19:55-07:00","dateModified":"2019-08-02T04:19:55-07:00"},{"@type":"WebSite","@id":"https:\/\/www.taterboy.com\/blog\/#website","url":"https:\/\/www.taterboy.com\/blog\/","name":"Design for Immersive Technologies","description":"User Experience for Games, Virtual Reality (VR) and Augmented\/Mixed Reality (AR\/MR)","inLanguage":"en-US","publisher":{"@id":"https:\/\/www.taterboy.com\/blog\/#organization"}}]},"og:locale":"en_US","og:site_name":"Design for Immersive Technologies | User Experience for Games, Virtual Reality (VR) and Augmented\/Mixed Reality (AR\/MR)","og:type":"article","og:title":"Text To Speech With AWS | Design for Immersive Technologies","og:description":"Text To Speech With AWSText To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a","og:url":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/","article:published_time":"2019-08-02T11:19:55+00:00","article:modified_time":"2019-08-02T11:19:55+00:00","twitter:card":"summary","twitter:title":"Text To Speech With AWS | Design for Immersive Technologies","twitter:description":"Text To Speech With AWSText To Speech With AWS Philip Kiely 2019-08-01T14:00:00+02:002019-08-02T11:07:55+00:00This two-part series presents three projects that teach you how to use AWS (Amazon Web Services) to transform text between its written and spoken states. The first project will use text to speech to turn a blog post or other written content into a"},"aioseo_meta_data":{"post_id":"68828","title":null,"description":null,"keywords":null,"keyphrases":null,"canonical_url":null,"og_title":null,"og_description":null,"og_object_type":"default","og_image_type":"default","og_image_custom_url":null,"og_image_custom_fields":null,"og_custom_image_width":null,"og_custom_image_height":null,"og_video":null,"og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":false,"twitter_card":"default","twitter_image_type":"default","twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_title":null,"twitter_description":null,"schema_type":null,"schema_type_options":null,"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":null,"robots_max_videopreview":null,"robots_max_imagepreview":"large","priority":null,"frequency":null,"location":null,"local_seo":null,"created":"2021-02-07 18:39:38","updated":"2026-09-01 18:36:44","focus_keyword":null,"additional_keywords":null,"truseo_locale":null,"primary_term":null,"og_image_url":null,"og_image_width":null,"og_image_height":null,"twitter_image_url":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"Article","isEnabled":true},"graphs":[]},"limit_modified_date":false,"ai":null,"breadcrumb_settings":null,"seo_analyzer_scan_date":null},"aioseo_breadcrumb":"<div class=\"aioseo-breadcrumbs\"><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/www.taterboy.com\/blog\" title=\"Home\">Home<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/www.taterboy.com\/blog\/category\/user-experience\/\" title=\"User Experience\">User Experience<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\tText To Speech With AWS\n\t\t<\/span><\/div>","aioseo_breadcrumb_json":[{"label":"Home","link":"https:\/\/www.taterboy.com\/blog"},{"label":"User Experience","link":"https:\/\/www.taterboy.com\/blog\/category\/user-experience\/"},{"label":"Text To Speech With AWS","link":"https:\/\/www.taterboy.com\/blog\/2019\/08\/text-to-speech-with-aws\/"}],"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p8wzr5-hU8","jetpack_featured_media_url":"https:\/\/i0.wp.com\/www.taterboy.com\/blog\/wp-content\/uploads\/2019\/08\/pollyconsole.png?fit=400%2C240&ssl=1","_links":{"self":[{"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/posts\/68828","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/comments?post=68828"}],"version-history":[{"count":0,"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/posts\/68828\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/media\/68829"}],"wp:attachment":[{"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/media?parent=68828"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/categories?post=68828"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.taterboy.com\/blog\/wp-json\/wp\/v2\/tags?post=68828"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}