Transcripts
1. Welcome: Welcome to the ElevenLabs
AI Master Class. AI-generated voices have become incredibly popular and capable, but ElevenLabs can do so much more than
just text-to-speech. You can design and clone voices, create sound effects on music, produce audiobooks and podcasts, dub content in other languages, generate visual
content, and even build AI agents that can
interact with users. In this course, I'll show you how to use these tools through practical examples and show you exactly what ElevenLabs
is capable of. You can understand not
only how ElevenLabs works, but how you could apply it and use it for your own
professional work. My name is Jose Kuchii and
I'm a visual designer with over five years of experience
working with AI tools, creative content, and
digital workflows. I have spent a lot of time
working with these AI tools to figure out the best way to integrate them into
real projects. Throughout this course, I'll share with you the techniques, workflows, and tools that
have helped me along the way. We will start with
the fundamentals. You will get familiar
with the interface, the accounts, the pricing, how you can set it
up the best way, and then we're going to
move into prompting, which is essentially how you
communicate with this tool. Once we have those
under our belt, we're going to move into all the tools that
ElevenLabs has for us. You will learn how to find
and create your own voices, how you can transform
existing ones, how you can dub audio for
podcasts or audiobook, how you can create sound
effects and music, and also how to make visual
content with ElevenLabs. Along the way, we
will also explore studio and other production
tools for larger projects. Once you're comfortable
with the tools, we're going to move into
some practical projects. You will work with
assets and templates, image and video
generation, lip syncing, upscaling and flows to
see how these tools can be integrated into a much
larger creative process. And finally, we're going
to explore 11 agents. Over there, you're
going to learn what they are, how
you can use them, and we're also
going to create one together where we
test and deploy them. By the end of this course, you will have a broader
understanding of 11 labs, but more importantly,
you're going to have had hands on experience
creating those projects. So let's go ahead
and get started.
2. What is ElevenLabs: Before we get into how
to use ElevenLabs, let's first talk about what it is and what are the use cases. So ElevenLabs is actually a AI company and they
specialize in generative audio. Generative audio basically
means when you generate audio, and given that this
is an AI tool, you're making audio with AI. Now, previously,
there has been a lot of technologies regarding
generative audio, such as text-to-speech,
speech-to-text, and other stuff. However, those rarely involved AI recently with
this implementation, the workflow has become a lot easier and a
lot more accurate. When you type something
and you use punctuation, instead of it sounding
super robot, now with AI, it sounds a lot more realistic and the other way
around speech-to-text, it can catch onto your
words a lot better compared to the previous
technology that was out there. This right now is the
homepage of ElevenLabs. We will talk about account
and pricing in a little bit. But just to walk you
through what you can do with this tool, obviously, you can make a lot of
instant stuff such as text-to-speech and
then speech-to-text, but you can generate audios
for your audio book. You can make podcast with it. You can maybe just
talk to an AI for YouTube video or can even generate your
own sound effects. There's a lot you
can do here and this is a very popular
tool if you're into making content or have
some channel where you would want to post these contents or
perhaps even monetize. A lot of room for
experimenting with this tool, you can generate
a ton of things, see which one works
best and then fully commit with a longer audio clip. The way it works is that
you have your account, you get a certain amount of
credits and per generation, you're going to
use a few credits depending on how
long your audio is. If you've used AI tools before, the workflow is
relatively the same. You give it a prompt,
you give it an input, choose your model, and
then you get an output. That we know what
the platform is, briefly, just a
general overview, we're going to talk about the account and pricing
in the next lesson.
3. Account and Pricing: Before we get into how
to use ElevenLabs, let's first talk about what it is and what are the use cases. So ElevenLabs is actually a AI company and they
specialize in generative audio. Generative audio basically
means when you generate audio, and given that this
is an AI tool, you're making audio with AI. Now, previously,
there has been a lot of technologies regarding
generative audio, such as text-to-speech,
speech-to-text, and other stuff. However, those
rarely involved AI, recently with this
implementation, the workflow has become a lot easier and a
lot more accurate. When you type something
and you use punctuation, instead of it sounding
super robotic, now with AI, it sounds a lot more realistic and the other way
around speech-to-text, it can catch onto your
words a lot better compared to the previous
technology that was out there. This right now is the
homepage of ElevenLabs. We will talk about account
and pricing in a little bit. But just to walk you
through what you can do with this tool, obviously, you can make a lot of
instant stuff such as text-to-speech and
then speech-to-text, but you can generate audios
for your audio book. You can make podcast with it. You can maybe just
talk to an AI for YouTube video or can even generate your
own sound effects. There's a lot you
can do here and this is a very popular
tool if you're into making content or have
some channel where you would want to post these contents or
perhaps even monetize. A lot of room for
experimenting with this tool, you can generate
a ton of things, see which one works
best and then fully commit with a longer audio clip. The way it works is that
you have your account, you get a certain
amount of credits, and per generation,
you're going to use a few credits depending on
how long your audio is. If you've used AI tools before, the workflow is
relatively the same. You give it a prompt,
you give it an input, choose your model, and
then you get an output. That we know what
the platform is, briefly, just a
general overview, we're going to talk about the account and pricing
in the next lesson.
4. AI Ethics Overview: With all of these new AI tools coming out almost
every month now, it's really important
for you to brush up on the ethics of using AI if you're not already
familiar with them. If you're a beginner using 11
labs or any other AI tools, you may be wondering why sometimes it's
considered unethical. Even if you know and you're
still not convinced, it's important for
you to realize the risk and why it could
potentially harm other people. For this lesson, we're
going to talk a bit about the ethics
of using AI tools, whether it's Chau BT, runway, M journey,
or even 11 labs. There is a boom in
AI tools nowadays, and it's important for you
as the user to know about the risk and why you need
to be using this tool, these really cool tools in
a more responsible way. A lot of times you come to these AI tools to either
make your work easier, help you with a project,
or just have some fun. Those are all okay
as long as you're not using it to harm anybody. Now, what I mean by harm? A lot of times these AI tools are being used to
impersonate people. The inputs that are used
to train these models are often taken from people
who did not consent to it. For example, with a
lot of art tools, there are many artists
claiming that their works have been used to train these models when they did not consent. What happens when that artist who worked really
hard gives their art it's being used to
train these models and then you make something
with it and then sell it. You right there have
knowingly or unknowingly have contributed
to this bad cycle that is harming these artists. Now with 11 labs, even though it's a great tool, it's important for
you to take some time to read about how
these models are being trained and how you could just use it for something
that will not harm people. Give you some examples with
11 labs in particular, but this applies to any other
AI tool that's out there. You should never be using these audios to impersonate
real human beings. There are a lot of
rules and regulations coming out every day about how to basically rules that prevent the harm
of other people. If you generate a
podcast fully using AI, you do need to label it as AI. If you use the script from the Internet that was also
made with AI, maybe Chat GBD, it is your responsibility
as a user to inform your audience
that you're in no way trying to impersonate
someone else or trick people into thinking that what you just generated is a person. You're going to be using 11 labs for your job or for your school, bear in mind that each
company and institution has their own set of rules and
regulations regarding it. So make sure you do your
research and you don't put yourself at risk
when using these tools. Now, for the purposes
of this course, I'm just showing
you how to use it. Use it to make podcasts, audio books, and
stuff like that. These are just meant for
you to explore and learn. If you do want to
commercialize any of your works that you make
by the end of this lesson, please take the time to read
the rules and regulation of all the platforms
where you would be uploading these products too, and again, label
your work as AI. Because what happens is that
when you submit your work in the same space as
other human creators do, you will be flagged and according depending on which
platform you're doing it on, there will be consequences. So be responsible with
your generations. Don't use it to harm or impersonate anyone,
label all your work. And if you choose to make money, monetize your work with
either 11 labs or any other out there, do it in
the proper fashion. Go to the platform, see
what their guide is. Almost every platform has it, whether it's on
Instagram, Bhand, Adobe, wherever you want
to do it, go and navigate on their website, read about it, and
follow their steps. Now you know a little bit
more about the risk and the potential harm
that you might be doing if you were to use
AI in the wrong way. We can now move on
and continue with our course where we learn more about this particular tool. So be careful out
there and let's go. Oh
5. Prompting Basics: In this lesson,
we're going to go over how to prompt
with ElevenLabs. So a prompt is the directions that you give an
AI tool to follow. And in this case, these are the words that you want
11 laps to execute. Whether it's a voice
over, a text-to-speech, voice changer, or any of the other tools that we're
going to take a look at later, there is a general style
that you need to follow, and it's actually pretty easy compared to the
other AI tools out. Head over to
text-to-speech where we're going to practice with
the different prompts. Now, don't mind the
tools on the right side. I will get into what this
is in a different lesson. But first, let's go ahead
and do a simple practice. We're going to start
typing in something. I'm going to have ChatPT give us a very simple
story and then we're going to convert
it into a decent audio. Give me a four sentence story about a dog. Just very simple. To make this better, let's
do something like this. Make a podcast episode based on this with only four
sentences for dialogue. All right, so let's
copy the podcast part. The reason why I
chose podcast is because this way we can play around with the
different emotions. I'm just going to
paste this here. Feel free to write
your own text. But let's get rid of all of these tags and just put in
the text that they say. Um, I'll put a paragraph
break after each sentence. So let's first start
by choosing a voice. For me, by default, it's George. The voice is the model
that is reading your text. So if I just hit George, I'm not gonna change
anything here. Let's actually reset the
values just to make sure that we're all using
the same settings. Let's hit Generate speech. Every morning, a
scruffy little dog named Max wandered
the quiet streets, searching for scraps
and friendly faces. One rainy afternoon, Max
huddled beneath a cafe awning, shivering until a kind barista spotted him
and stepped outside. Hey, there, buddy. You hungry. From that day on, Max came
back until one morning, the barista opened the door with a smile and said, want to stay. That's just the model
reading our text. It's very monotone, it's not
exactly exciting or dull, it's just very flat. So here are some things
that already are included in this text that partake in that style that I
was talking about earlier. Every time you add a period, that's an indication
for the model to pause. Then when we have a, there's three dots, it's
going to be a longer pause. If you do this, this is
going to be a different cut. We're going to incorporate these within our
prompt going forward. But right now we have dashes, we have the three dots, just a period, we have a comma. Those things are
already included. And when I played the audio, you saw how it paused. When the sentence ended, it again paused here between the two words and it
read it just fine. Now we're going to talk
about how we can add emotions because we have a Brista here and then
we have just a narrator. When you use
different adjectives, ElevenLabs is going
to understand that as the emotion it
needs to execute. We can say, let's see. Right here when
there's dialogue, we can say set the
excited Brista let's see. Set the Brista in a excited way. I'm using the word
excited here and then let's see what
else can we add? Instead of this, I'm going to do an exclamation point and then
maybe two question marks. So just this small change is going to make
a big difference. Let's generate this
speech one more time. Every morning, a scruffy
little dog named Max wandered the quiet streets searching for scraps and friendly faces. One rainy afternoon, Max
huddled beneath a cafe awning, shivering until a kind barista spotted him
and stepped outside. Hey, there buddy. You hungry? Said the barista in excited
way from that day on. So you saw how it says how
the George here said that, Hey, they're buddy, you hungry. So that was the
difference between adding two different signs
versus just having one. Now, if you go to if I want
to take this a step further, you can add pauses
and all of that, too, which we'll get
into later because there's a tool within studio where you
just click a button, but you can add Pause is here, like three dots, and then
maybe another one here. Now let's see how that works. One thing that you
can do is when you combine the two sentences, put them right after each other, it's not going to pause as long. Every morning, a scruffy
little dog named Max wandered the quiet streets searching for scraps and friendly faces. One rainy afternoon, Max
huddled beneath a cafe awning, shivering until a kind barista spotted him
and stepped outside. Hey, there, buddy. You hungry? Said the barista in excited way? From that day on, Max. So that's the changes
that we're seeing. And moving forward, the
only prompting you would be doing is giving 11 labs a
story line like this one. We're going to look at how
we can add sound effects, longer pauses, and
different sorts of audios with the other
advanced tools. The traditional form of AI prompting where you describe or you command for something, for example, make
me a sound effect of heavy rain within the forest. That's an example. That
will only take place when you're making sound
effects with 11 laps. But for the rest of the other
tools and there's so many, this is the prompting
that you're doing. You can either generate
these story lines yourself, type it in here and as you
saw, it's really easy. Or you can have
another tool like ChaGPT make something for
you in a matter of seconds. This is just a lesson for you to know what we mean
when we say prompt. It's not your
traditional prompt, where there's specific
formulas to follow. Just keep in mind
the punctuation. Each of them mean
different things, as we saw with the commas, dashes, three dots,
stuff like that, and be very generous with your adjectives
because those can mean your AI voice model going from monotone
to either excited, angry, SAD, whatever
adjective you want to use. And we're going to
be seeing that a lot in the further lessons. Without further ado, let's
jump into our next chapter where we look at the
mini ElevenLabs tools.
6. Voices Library : Let's take a quick look
around this platform and see how we can right away start making voices with ElevenLabs. This is the page you will see the homepage after you've
signed into your account. So right away, there are some latest from the library voices. Essentially, we have a library where people can
add in new voices. Of course, there's a
procedure that they have to go through and it's not
low quality voices at all. But this is a library
that ElevenLabs manages and they're releasing
new voices every day. Let's give this one a
listen, for example. AubukieraOsk Terra. That's in Hindi,
and then we have, for example, a
young woman Silas. A So this voice right now you can see it's being called female mature voice, and it sure did sound
that way, calm and deep. So I just refresh my page to show you that
every time you do it, you're going to get
a new set of voices. So, for example,
let's go on this. Evil is evil, Lesser, greater. Idling. It's all the same. So this is serious and grim, and you can see it definitely
did hit that mark. Now, down here, you will notice
it says English preview. If you click on this, you get to switch between languages. These are the current
languages that are available. It's going to be the same voice, the same tone, and all, but it'll be in a
different language. So let's go ahead and try
Russian, for example. Chive Katorn D Ciba Nemo
switching Devi Camibe. City er SchimaFantasia,
solitaria, Poggi transformer completa Menchi HmunogunKab Hari Hunaja, JonikOtmaFatra and 40s Helps
Dantwsen The man lipped. You can see how seamlessly
it switches between these languages while maintaining
the voice and the tone. If you speak any of
these languages, you can see if it sounds
realistic or not, but from what the
community says, it does hit the
mark pretty well. Okay, so if you wanted
to explore more, you can go over here or
simply go down to voices. This will take you to
the voice library, and we have a lot of
options to choose from. Right away, you will see that there are checkmarks on some of them and that just
indicates that they are high quality and they've
been used a lot. You can see the name of this person and if you
just hover over it, you're going to see a little
bit more information. The description
of the voices can really help you choose
the one that you. Example, David here
is a news reader, and he has a clear
and crisp voice. He's middle aged,
professional American, and then you can just gauge on whether or not that's going
to fit to your project. Let's go ahead and
preview this voice. Just click on it. Later that same evening, Detective Carlson
received an anonymous tip directing her to a
second crime scene. What had at first
seemed a random event. And then once again, you can switch between the languages. Now, you can see here that the language selection is
different from the first voice, and that's because each voice has its own language selection. So here we're actually
getting Arabic, Polish, Alien, I think those are the only
ones that are different. We're getting Hindi, English and French like we did before. And down here, you
can see the flags. So it comes in these two
languages, and then five more. These two, two more, and it's just going to
change as you go on. In the Explore page, you can search for
either the name of that voice or
the way it sounds. For example, I'm looking
for a calm voice, and I'm getting a lot
of narrative and story, informative, educational
going to give this a listen. Here, we have a calm,
well spoken narrator, full of intrigue and
wonder for nature, science, mystery, and history with a
smooth and velvety tone. You can see that definitely
did sound calm and it's actually one
of the adjectives that's being listed here. Right away without having to click on it or hover over it, I know exactly what
voice I'm going to get. Now, you can take a
look on the right side, which is the categories
that they go into. If you open these in a new
tab, let's just open one. We're going to get a library just for informative
and educational. You can see that's the
category that I went into. Close that guy.
Right next to it, you can see two Y or
zero D. That just means that how long this
voice is going to stay if it were
to be removed. If the owner of this voice removes it from
the voice library, it will become unavailable
for you after two years. But here, for example, it will be removed immediately. If you want to use
this long term, definitely go for
the ones that have a longer retention compared
to something like this. Next, we can see how many
users have used this voice. The first one by Adam Stone, you can see it's pretty popular. 438,000 people have
used this voice, meaning that you've probably
heard it at some point. The people who are crazy
enough to think they can change the world
are the ones who do. There we go. Then
there's a plus button, which means that you can favor
this voice for later use. Click on at once, and then it
basically goes in this tab, which we'll get into later. Now you can see that it went
from a plus to this icon. I can use this voice to either generate an audio
to text-to-speech, speech-to-text, whatever
that I have to do. Next is more actions. You can copy the voice link the voice ID, and
then view similar. If you've used other AI tools, think of the voice
ID as the ST number. When you have the ID and you're trying to
generate your voice, it will keep it a
lot more consistent. It won't functuate as you're putting in different
adjectives, different tones. I will remain similar. But usually, if you use one
of these pre made voices, you don't really have to use it. The link is just a link to
that voice if you want to share it with someone and
then we have similar, which gives you more voices
that sound like Adam Stone. Now let's explore
the things up here. This was our search bar. We were able to look for an adjective, maybe
someone's name. Let's look for someone
called John Doe. There we go. We
have three options. You can also look with
search with age or gender, maybe senior can do male or
maybe teen male or something. Anything that we want. So for example, English teen youth. Hi there. I'm Archie. If you're after a
young, energetic and authentic British
teen voice, then Hey, I'm Ethan, and I would
be a great choice if you're looking for
a teenage voicedor. You can see how those
search bars affect. You can see how searching for a certain thing does give you the correct voice selection. Some of these voices will be charged a bit more
than the others. With the plan that I have, you can see these ones
are okay for me to use. They're the popular ones, and there's no dollar
sign over here. But for these, you can see it's $0.02 and that's only for
a few of these voices. You can see most of
them don't have that. Let's go back up here. This was our search bar. Right next to it, you
will notice this icon. Which is basically
when you get to upload a sample, close that up. You get to upload
a sample and then 11 labs will find one
that's similar to that. Let's say you have this voice that you want
to replicate with AI, of course, be very
mindful of copyright, but you can put it over here. For example, if you want a nice confident
voice, for example, Michel Obama and you want to get something relatively close, can upload a snippet of her speeches over here
to get something close. For example, if you're doing
some animation and you have a character after Michel
Obama, you can do that. Again, you really want to be mindful about how
you're doing this. We will talk about AI
ethics in a later lesson. But you can't just scrap
any voice and try to get the exact same thing
because that could lead to some big problems. Right next to that,
we have a trending. If you click on it,
there are a lot of popular voices as well
as unique voices. What I mean by that is that the trending ones
have been used a lot. These are the ones that
we see right here. I'm going to click
on this triangle just to go into the list again. You can see a large number
of people have used it. If you want to be unique here, you shouldn't probably
go for something like this and that's where
this tab comes in handy. If I switch to latest,
you can see this one, for example, it's not verified, but only 14 people have used it. This way, I'm being
more unique when I have my voices in my videos and there's also
options that are lower, seven people, five people, and the list just
goes on two people. Getting left behind with AI
feel like you can't keep up. You business needs a prompt
engineer who matters. So you can see, even
though they're new, some of them are
not as high quality as the most popular ones. So it's a good idea
to preview them, make sure it's something
that works for you, and then proceed to use it. We have Latest, we
have most users. So you can see this
one is really high. I think this is a
really nice way to just talk naturally together, you know, T plainly.
Let's give it a shot. You can see it's conversational, it sounds really
natural and that's perfect if you want
to do a podcast or if you want to use your own voice and
convert it to this tone. That's something really
cool that 11 labs have, and we'll get into that
in a separate lesson. But that's what most
users will get you. Character usage basically means how many characters you get to convert to audio
within your credit range. This is something we'll
talk about later when I show you how to
do text-to-speech. You will see exactly how many
characters you're using. For some of the voices, it will take more credits while the others won't is
what it's referring to. Don't mind this for now.
We'll get into this later. At the same time, we
have the filters tab, which lets you dig deeper and
find what you really need. We have languages, first of all, a long range of languages. It's like Welsh, for example. You can even choose an
accent if it applies. For example, it
doesn't for Welsh, but let's see what
would have an accent. Let's see if English
has one. There we go. If you're choosing English, there's a lot of accents
that you can explore. Let's say I want someone from
Chicago, let's choose that. Next is the category, which is the tone that the
voices sound and speaking. Think about what
you want this voice for and simply choose
your category. Going to do a very natural
tone, conversational. Click on that. You can also
click more than one category. I will do maybe
something like this. Quality, any or high quality. Remember, high quality are
the ones with the checkmark. Those may ask for more credits, but you can just click on any to get to have a bunch
of options to look at. Next is gender. I think I'll do a female age. Let's do an old person. Notice period. These were
the numbers on the side, how long you get
to keep them once the owner deletes that voice. Let's say two years. Custom rates,
include or exclude. This is for that character
usage that we saw. Then live moderation enabled. This is going to
moderate the type of text that you're
giving to this voice. I'm going to hit
Include for now. Let's apply the filters. Right now, you can see I don't have any and this will happen, but you get to create
your own voices. So that's something
we'll explore later. I'm going to remove a
few of these just so I can get something
to show you guys. Okay, so just a US
Chicago accent. Exercise makes our heart stronger and our
smiles brighter. That was one example.
You can see it's a very recent audio. Only 32 people have used it. It's not verified, but it's still a voice that you can use. Let's try something else. We can do, again, English, try Canadian maybe. So here's a high quality audio. Met Gram, this handsome
cat at our shelter. This boy is lovable. With his tabby stripes
and cute face, he's sure to steal your
heart in an instant. So that was one example. Let's try to find another one. Ladies and gentlemen,
welcome to today's match. We have an exciting
match ahead between EFC United and Team Impact. The atmosphere is electric. These are all Canadian accents
in different categories. Next, we have create
or clone a voice. This is that cool feature that I talked about
where you get to either design an entirely new
voice from a text prompt. You would say maybe an older gentleman
with a thick accent, maybe a deep voice, calm, energetic, whatever
you want to put. And it's going to ElevenLabs. We'll give you a bunch
of options, samples. You get to build up
on that and then make your own original audio. Now, the things that you create, you can also have
other people use it. That's where all those
new voices come from. Less than a minute, and we'll get into this in
a different lesson. We have a voice clone option, so you can clone your own
voice with only 10 seconds. I will speak into
the microphone for 10 seconds and then 11
labs will study my voice. Then I get to use my voice
with different text prompts. For example, if I don't want to record a audio book
that's 3 hours long, I will just speak into
the mic for 10 seconds and then paste the story line or just the story into
the box and then 11 labs will generate that
three hour audio for me. Next is the professional
voice clone. This is not included in
the plan that I have, but this is just for
the most realistic one. But for most of you, since you are using this tool
for the first time, these two should be enough. Up here, we have a feedback box, which you could just
write your feedback. We have some other
stuff up here. This was the Explore tab. My voices will have the voices that you favorited
or made yourself. We have default voices. These are voices by ElevenLabs, so these aren't user generated. Just trust yourself, then
you will know how to live. You can take this, use it
like all the other voices, and just take a look
at the many options. We also have collections, which you get to create when you are trying to
organize your voices. For example, I could create a collection for my
YouTube channel, one for my podcast, one for my audio books, and that way, everything will
be a lot more organized. You're previewing a voice,
we have the name here, the different languages it's
available in Play button, you can skip back and forth, see how long it is over here. You can download the voice, and then you can just hide the player if you don't need it. That's the voices panel. If you click on
this Plus button, it's going to bring
you the same menu that came up when
we clicked here. Now, in the next lesson, we're going to go over the
playground tools that are located right underneath and see what else we can
do with this tool.
7. Text to Speech: With text-to-speech,
you can easily turn written text into genuine
good quality audio. And this is the interface
where you're going to do that. So this right here is the
first tool available. Once you click on this, you're
going to have this area, which is essentially
your playground, and then you have
all the properties, the different voice libraries, any sort of setting that
you're going to need, it's going to be on the right. And down here, we have
some quick templates for you to get started with. But of course, you don't need to use these unless you want to. This right here is where
you get to write the text, and you're going to basically
pretend to be the voice. So right now, I have
Liam from the library, but you can choose
anyone else here. We already went over
the voice library. There's so many to choose from. And you can also use the
voices that you have saved. Again, you can filter
through any of these categories and
find the perfect voice. But even if you have a voice
that you think is perfect, you still have further
ability to adjust that voice. For example, if I
choose Adam right here, I can still come over
change the stability, the panning, the language
that he's going to speak, the output format, and
everything else that's needed. So let's go ahead and see how
this works exactly because there is some cool
things that you could do with ElevenLabs in
text-to-speech. So here you can type in the text that you
want Adam to say, but at the same time, ElevenLabs has something called audio tags. And this is an
example right here. So, for example,
laughs, cries, screams, any sort of audio tag that
you want to be implemented, you can add that in
using the brackets. Can also use punctuation
to work with the pauses, the full stops, and all that. And at the same time, we
can have multiple speakers. So what we're going
to do is make a very simple conversation between Adam and
another speaker, and we're going to type in some sentences and use some of that audio text
that we talked about. This way, you will get
a better idea about how text-to-speech
actually works and how you can use it to
the full potential. So let's start with Adam. I'm going to have
him you say, Hey, by using the exclamation point, it's going to be a
little bit more excited. So Adam will be kind of
energetic when he's saying that. And just to show you
what that sounds like, I'm going to put another
sentence after this. So, hey, how is it going? So, right now, I have used
two punctuation marks, and we're going to see
what this sounds like. I'm not going to
change anything here, and we're simply just gonna
click on Generate speech. Hey, how's it going? So you can see that
he kind of stretched the hay and then it
asked a question. Now, if I remove these guys and just not use
anything at all, that's going to sound
a little different. Hey, how's it going? So he did not pause, and it's less of a question now. So one way that you can control the voice is using those marks. So let's add in
those question mark. And then we're going to
add our second speaker. So let's make this someone else. I'll just go with Ellen. And now Ellen is going
to do something else. So maybe she gets startled
when Adam asks this question. So maybe yelps. Oh, my God. You scared me. Hey, how was it going?
Yeah. Oh, my God. You scared me. So you can see
how well that played out. And to show you the
power of the audio tags, I'm going to remove this now. So now it's just, Oh, my God, you scared me with the
exclamation points. Hey, how's it going? Oh, my God, you scared me. So you can see, even though
I didn't put the audio tag, the voice is still
pretty natural, and that's where I want to point to the models that we're using. So right now we are on 113. If you click on this, you get to see all the different models. But essentially, 11 V three
is the most expressive. So the way that, you know, he comes in, I'll coal and, like, a deep voice, and then she comes in a little startled. That's because of our model. We do have multilingual
models here as well, but right now, I'm just
sticking to English. So if you do want to do this
in a different language, I recommend moving to this one. There's at least 26 other
languages that it supports. So if you see the
language that you like, you can put it right here. And we have even more models. So you can see how
right now it's saying that it recommends
this model for Ellen. Sometimes it does that depending on the voice that you choose. There's also flash turbo
and different turbo models. So it really depends on
how much stability you want with the voices and how much expression
you want with them. So these guys down here, they're usually good
for quick conversions. Like, you don't want a lot of detail going into the
voice generation. So you can see that it mentions
conversational use cases. So maybe if you have
a YouTube video where the person
is just speaking, there isn't much going on. You can use a model like this. And we also have other models that it recommends us
for to use it for, like, developer use cases. For example, if there's, like, AI chatbot on the website, you can use one of these models. But if you're going
to build a story where there's a
lot of audio tags, there's a lot of
pauses, and motions, you want to go for
the first model. So I'm going to
stick to this guy, and I'm going to put
back that audio tag. Just say what it is and
then put it in brackets. It should turn pink. If it didn't either because it doesn't recognize
that audio tag, like, it doesn't support
that expression, or you may have put
in an extra space, a comma, misspelled
or other causes. But once it's pink, it's
for sure and audio tag. You can add Adam back. You can see it picked
up immediately, but you can add a third
speaker if you want. So I'm going to do laughs. You're so easy to scare. Did you finish the homework? Let's add Alan. Again, you can see it picked up that there is a conversation. We'll do this size. No, da da dt. I was so busy yesterday. You can also use the
audio tags between words, so it doesn't always
have to be at the start of a sentence
or at the end. So right here, we could
do Snifs audio tag. They think I am coming
down with something. And now we're going to generate. Hey, how's it going? Oh, my God. You scared me. You're
so easy to scare. Did you finish the homework? No. I was so busy yesterday, but I think I am coming
down with something. So there's our very
simple conversation. But now we could go
into each voice and further define the way the
expressions are being made. So, for example, Adam right now is a little bit too intense. It's not really a
conversational voice. So I'm just gonna click on Adam, but then go back to
the page right here and make it a little
bit more creative. So let's see what
that sounds like. Hey, how's it going? Huh. Oh, my God. You scared me. You're so easy to scare. Did you finish the homework? No. I was so busy yesterday, but I think I am coming
down with something. So now his voice is
a little lighter. That's because I pulled it
towards the creative side. Just be careful with this
because when you're telling ElevenLabs to not be that consistent through
the voice generation, the voice may kind of glitch if you're doing a very
long paragraph or, like, a big story. But for a situation like
this where it's, like, a sentence or two,
works pretty well. Now when you change the models, you do get additional
things that you can adjust. So this right here
is for 11 V three, but I'm going to go to
Ellen right now and switch this to the second model. You can see that
it is recommended right here. Going
to click on it. And now you can see that
there is a lot more options. The only problem is
now that we don't have that conversational
layout with B three. We don't have that
different speaker mode. Right now, everything
that's written here is for Ellen only. So that's the downside. So you can kind of move around each model depending on
what you need it for. But we can use Ellen right here. I'm just going to delete
the lines that Adam had. And I'm going to, you know, play around with some of
these sliders down here. So the first thing is speed. Is she talking really
fast, really slow? You can adjust it right here. So, for me, I'm just
going to lower it a bit. And with these sliders, you do want to be careful again about not going overboard, because if you do go all the way back or
all the way forward, the voice may not be able to
keep its consistency across, you know, different sentences. So whatever adjustment that you do try to keep it very minimal. So I was one here. I just came down to
96, so not like here. Next thing is stability. So do you want her to have different volumes at points or do you want her
to be just monotone? You know, that's
something you get to decide with this slider. So if I go all the way here, she's going to sound
like a, you know, robot. But if I go all the way here, she may just be very hyper. So again, we can make
very slight adjustments. The next thing is voice clarity. I'm going to keep this as it is. I think it's, you
know, pretty clear. Next, we have style
exaggeration. So do you want it
to be, you know, the way she sounds right now, or you want her to
be more excited? Do you want it to exaggerate all the different
punctuation marks, audio tags, and all of that? Again, I'm going to
leave mine as it is. And then, lastly,
we have panning. So it's just about,
like, the stereo field. You can think about
like two speakers. Do you want it to
be more towards the right or more
towards the left? So just to show you
what this sounds like, I'm going to bring this
all the way to the right. And that way, it will only come out from one
side of the speaker. We also have language overwrite. If you want to manually select the language that ElevenLabs
produces for you, you can turn this on. So, you get to choose
a language here. If you don't, it's going to detect what you have up here and then make
that decision itself. Next is audio effects. These are additional
files that you can add in here just to make the voice
maybe more realistic, more suitable for a
certain environment. These are all easy
to add as well. For example, I could
add in some noise, say she's in a cafe. I could filter
something as well. I maybe she's on the phone. You can work with the space. This is the echo, and then
we have the distance. So how far is she
from the microphone? So I'm just going to do
the first two first, and you can do a little preview before you waste any credits. Here's a preview of your
voice with effects applied. Here's a preview of your
voice with effects. So you can see that it only came out of the right side
because of this. So let me just put this back. Here's a preview of your
voice with effects applied. Here's a preview of your voice. And that's what
it's sounding like. Lastly, is the output format. These are different formats
you get to choose from. It depends on what you
need this voice for. So you can see some of
them are completely grade out because it's for a
more professional plan. I'm currently on 11 Creative, so these two should
be enough for me. I'm just going to keep
it on the default, and, of course, we want the
speech to be boosted. If you ever wanted to go back
to what it was initially, you can click on Reset values. Let's generate the speech. Oh, my God, you scared me. No, I was so busy yesterday, but I think I am coming
down with something. So now you can see that the
voice is very different from the first model because we completely got rid
of the audio tags, as well as a lot
of the expression that was available
in that first model. But I'm just gonna
turn this loud, just so you guys can
hear that cafe thing, and I'll remove the phone. I think that's making
it harder to hear. So let's do this one more time. Here's a preview of your
voice with Effects applied. Now we can better hear that
cafe background noise. If you go onto a
different model, you may get additional sliders. But again, only
come down to these if you want something
specific done. So this right here
is the 11 turbo. It's pretty much
the same sliders. So it really comes
down to, like, the languages that you want to use or the speed
of the generation. But V three is, you know, the best for conversations, and honestly, it sounds
the most natural. You have an audio, you
can come down here, you can play it, move
forward, backward, give feedback if you want,
share it with people, download it for yourself, and you can even
hide the player. So if you download
it, it's going to just go straight
to your computer. It's going to be the same format that you chose right here. Now, if you're using V three, you can further enhance
audio with audio tags, and that's something
specific to this model only. That's how the text-to-speech
feature works. It's really helpful, and the more that you
play around with it, the more you're
going to be able to find the perfect voice for your project and be able to keep that consistency
across all of them.
8. Voice Changer: Now, while the voice changer seems like a very
straightforward tool, there's a lot of
things that you have to consider in order to get a decent and accurate
voice conversion. This is the interface
when I first click on it. We're still in the
playground group. I just click on voice changer and this is what you should see. Essentially you input an audio, whether you record it now or
have a pre recorded audio, and then you choose the
voice and the settings, it converts it from your voice to whichever voice you choose. What we're going to do is
record our audio directly into the computer and
then just go from there. So click on record
audio and when you do, you're going to
get the option of which microphone you'd
like to use for this. When you're ready, you
can hit Start and you can see that we have a
maximum of 5 minutes. I'm just going to
read a piece from a short story maybe and then we can see how well it
does over there. This is my original audio. Everyone looks up. A few people start jumping up and down, waving their arms in the air. Jake looks over at the
towers surrounding the compound and notices
the guards on full alert, training their rifles down into the compound where the men and women are
making a commotion. So here I read things
with a slower pace, but later I'll show you what
you would get if you were to speak too fast or if your
voice was just unclear. Now that I have my audio, I can either delete
it to get another one or just proceed
with the conversion. First, you have to
choose your audio. I'm going to try
both a female and a male so that you can
see the difference. Let's go ahead and try
female voice that is not similar to mine.
Search for female. Just trust yourself.
Then you will know how to live. She tore her gaze. Hi. Government of the people. There is no greater harm. Ideas are the beginning
points of all fortunes. So here we have Dorothy
with a British accent, which is an accent I don't have. So let's go ahead
and click on this. This is the model
that it chose for us, and I'm just going to leave
it with the default setting. Now, one thing that I do
want to point out here are these two options that you
get to turn on or off. The first one is remove
background noise. If you're not recording this
directly into a microphone, you will need to turn that on. That way you will
not focus so much on the cars passing by or
someone hitting the table. And just focus on
your voice alone. Now, speaker boos by
default, it turns on, but I found that
when it's turned on, it does distort your
voice a little bit. I'm just going to run this
with the default settings, which is what you see on
the screen right now. Then later I'll turn it off and then show
you the difference. Everyone looks up. A few people start jumping up and down, waving their arms in the air. Jake looks over at
the tower surrounding the compound and noises
the guard on full alert, training their rifles down into the compound where the men and women are
making a commotion. You can see it
struggled with some of the words and use a lot of s, which is something that I believe is caused by
the speaker boost. Because I do have access to
a professional microphone, I don't really need this one. So what I'm going to
do is turn it off. And then reduce the
similarity so that I could have more
of British voice, I guess, something closer to
what Dorothy sounded like. If you increase this
all the way to high, it's going to really try to mimic your exact tone and pitch, and that could lead to
further distortion. Let's reduce this to 40%. I will leave stability and
exaggeration as they are. Now let's see the difference. Everyone looks up. A few people start jumping up and down, waving their arms in the air. Hake looks over at the
towers surrounding the compound and noices
the guards on full alert, triming their rifles down into the compound where the men and women are making
a commercial. So that's still a little better, but we can go even lower. We can try for more stability.
Give that another try. Everyone looks up. A few people start jumping up and down, waving their arms in the air. Jake looks over at the
towers surrounding the compound and noices
the guard on full alert, training their rifles down into the compound where the men and women are making
a commercion. It's a lot better than
what we had initially. It's still struggling
with the accent part. Words like commotion, it
does commercion an RN there. But that's something that we can fix with a
different voice. But instead of going
through all the voices, I'm just going to switch over to a male voice and try
a different accent. Let's try something without an accent, like American male. With the dawn, fresh
solutions light the path. To climb steep hills requires
a slow pace at first. Hey, come here. Hi, you've reached Acme Corp. There comes a moment when the familiar path no
longer feels like enough. A sunset is a day's way
of saying goodbye, Grace. Okay, so I'm going to
go with Liam over here. Let's click on his voice. So we have a new
audio right now. Over and over, he hears
the sound of the doors being slammed and bolted
slammed and bolted. So now we're gonna
try a male voice. God has given you one face, and you make yourself another. So there's George down there. We do have the option
of different languages, but let's start with English. Over and over, he hears
the sound of the doors being slammed and bolted,
slammed and bolted. So it's interesting I said
those two words twice. The first round, I
didn't say well, but the second round it did. So there is some inconsistencies as there always is with AI. But for the most part, you can see that it's followed along the places where I changed my tone and the way
I said the words. So it's pretty identical except for the fact
that it's George. I'm going to do it one
more time without the speaker boost and see if
that's any different. Over and over, he hears
the sound of the doors being slammed and bolted,
slammed and bolted. Let's actually reduce
the similarity to 20%. Over and over, he hears
the sound of the doors being slammed and bolted,
slammed and bolted. Was a lot better.
I said the words right and that's
probably because I reduced the similarity to 20% and removed
these two options. What you can do with this
audio is download it. You can also download it
with the original name, delete it, and share
it if you'd like. That's pretty much how
the voice changer works. You could also try it
with external voices, maybe do something
in text-to-speech, and then you change your mind, you can just come
to Voicehanger, upload what you got
from here into here, switch to a different voice. Just keep in mind that
text-to-speech is a lot more accurate
than voice to voice. So keep that in mind if you're doing this for a more
serious project. In the next lesson,
we're going to go over to sound effects, which is little
sound bites that you get to put in your
video to either emphasize a certain action or to make something
more dramatic. Let's move on to that lesson.
9. Sound Effects: Now, let's move on
to sound effects, which is when you get to generate a sound effect
with just a few words. And right away, you
can see that there are some examples
for you to explore. These are sample prompts. When you click on them,
they'll show up into this box, which is where your
prompt is going to go in and then you
just hit generate. Let's do a car whizzing by. We can also try something
else like a cat maybe. Down here, it's the duration. So how long do we want
this sound effect to be? You can grab a slider
and if you're not sure how many minutes or
seconds you want it to be, you can just hit Auto. Then we have our
prompt influence. How strict do we want the AI
module to follow our prompt? If you wanted to have
some creativity, you can keep it on
the lower side. But if you want it to be
exactly what you asked for, just bring the
slider towards high. The default is 30, and that's what I'm
going to leave it as. Over here, we have the history, so you can get different samples based on a prompt
that you give it. Previously I did this one
semiautomatic weapon, and I got four samples. Here you can press play
to hear the difference. You can see all of
them are weapons, but they start differently and
they even sound different. Based on whichever
what I'm trying to do, I can choose any
of these samples and simply click on Download. What you can do is
click on D plus to reuse the prompt
and maybe you can add a new detail over here. These are organized by date. You can just have a look
at what you've made so far and just keep track
of what you're building. Is the duration
and the influence. Okay, so that's about
it with what you can do right away in the
sound effects panel. We do have a Explore tab, just like the voices
library where you get to view some of these sound
effects that already exist. For example, let's
do, let's see. A Lion. We can even do cinematic effects. This is urgent emergency
broadcast. Attention resonates. A contentment breach has you can see all sorts of
sound effects can be found here if you
wanted to fine tune it. For example, this
basketball thing is great, but maybe you want it to
start with a whistle. All you have to do
is click on use and it's going to come
into your prompt box, and then you can
add that detail. Maybe whistles
blowing in the back. I added one more detail. Let's hit generate
and see what we get. While that's happening, I'm
going to just there we go. Let me replay the original
for you so you can compare. And then we have four samples. All of them sound like a very
intense basketball match. We could also in a similar
fashion, remove a condition. Let's remove the
whistle and crowd cheering and just do a
silence basketball match. Let's hit Generate and I could just toggle that so
that it's organized. If you have a lot of
stuff in your history, there's also a search bar, that's really handy
if you plan on generating a lot of
sound effects here. So we can hear that basketball dribbling and there's the
sneaker squeaking in the back. But as you saw, there is no more crowd,
there's no more whistling. And that's because we removed those two parts from the
prompt. That's it X. You can download any of
these if you'd like. If you click on them, you can
also just copy the prompt. Just like the voices library,
there's a search bar. You can click on More to
also look at it this way, most downloaded, most recent
for something unique, and then trending audios. Up here, there's a bunch of categories to make
search easier. Let's try something
like UI elements. If you want to add
a sound effect to your website, for example. Welcome to VCStom Graphics. Let's go back to Explore. There's one more
thing that's really cool in the sound effects panel, and that is soundboard. If you click on it, it will take you to this whole
different page. If you've worked
with soundboards, this should look pretty
familiar to you, but if not, don't
worry about it. A soundboard is basically
a board where you get to add different voices
anytime you click on it. Think of these boxes
as buttons and say you are having a podcast where
you're reacting to something. Anytime you want
there to be a drum, you would click on
the drum button, and that's going to be
recorded into your podcast. We have a bunch of
sounds here already. Let's go ahead and try a few. I just went in the
drum category. We can do some nature noises. You can see how I can stack
them up as well and create that perfect cinematic
sound effect for my video. There's also movie slash TV. These were your categories. You can also add a new row and not all of them
have something in them. For example, movie slash TV
doesn't have anything here. What you can do is click
on them and generate one. Let's say I want to create
one where people say boo. Let's say a crowd says
boom in the satisfaction. Once I do that, I
could give it a name. Let's call it boo. We get a bunch of samples. Then I could add
any of these that I like to fill up
this empty spot. Let's do number three. Now you can see there is
Bo this empty button. You can save the preset,
share it with someone, can also assign a device
here if you have a keyboard or something that you can
press keys on essentially. You can grant the permission and then have that
connected to your computer. Instead of pressing these with
your mouse or your cursor, you can click on the keys on a keyboard or a
different device. You can also do is
create presets. Right now, these
are the defaults. But you can generate
your own presets if there are certain noises
that you need every time. For example, let's
say I want to work with a lot of singular noises like someone saying something. I could start doing that. Just go to my
preset, new preset, let's call it singular
or just singulars, save the preset,
and then you will get all these empty boxes. Click on any of them and
start filling them in. A man saying, Hey, Hi. Hey. Hey. Hey. Hey. We got four generations. Pick the one that
is most accurate. If it did not give you
something you were looking for, you can just click
on Generate again. I think I will add Hey. Hey. Hey. This one. You can see
there's three versions, and that's fine because I
could just cut it down. Maybe here we will
do a woman gasping. Then here I will do
maybe a car crash. Call it crash. Let's add this one car
crash and traffic, maybe. I'm still on number three, so I get to replace it if
I find something better. Sounds more like an
explosion, but it's okay. And then number four,
we can do sirens. Maybe police car, sirens. So there is the police car. Let's do one last one for
traffic horns or honking. Traffic honking. So maybe the first one. So we have all of these sounds, and let's say I want to
generate a car crash scene. All I have to do is
press the buttons at certain times and make
that perfect audio. So maybe we start with traffic, car crash, woman gasp,
this, and then hey. Hey. Hey. So that was an example, but you can just
play around with it. You can do That's
just how it works. Then once you have something
that you're happy with, you can even loop
it. Turn on loop That's just how it works. This was the volume, if you ever want to play
around with it. We also have the
edit mode if you wanted to edit your sound. I'm just clicking
on the man saying, Hey, and I could change my
prompt and replace this. If you need, you can add
more rows down here and just fill it up with whatever that you need for your project. Then up here we have
the play button. We have the stop button or just clear whatever
that's in here. This currently is not saved, so you can just save
it into your preset. If you scroll down,
there's a lot of explanation on
how this works. We also have a bunch of
soundboards for you to try. Let's try Goofy of, for example, and now we have a bunch
of voices to try. G to hit loop and then lower the audio and start
filling it in. What was that noise? That was a good so that's
just how it works. You're just pressing a bunch of buttons to make
your sound effect. Let's go back over here and close that up and then
go back to generate. That was the sound
effects panel. It's a lot to work
with and you really have a lot of flexibility on how your sound effects sound. Now if you're not familiar on how to use these
effects, basically, if you have a video
and you want to emphasize a certain
action or certain dialog, you can add the sound effects. For example, if there is a
car passing by in your video, you can ask 11 laps to do a wish sound so
that you don't have to actually get super close to a car and record that sound
with your own microphone. It is also really helpful
if you do any games or sketches as it saves you time by just not having you record
sound effects from scratch, which can take a long time. For the most part, it's
pretty accurate if you saw that you're not getting
the sound that you want. Try going in the Explore tab, searching for something similar, and simply change it
up to your liking. Now let's move on
to voice isolator, which is the last tool within
the playground category.
10. Voice Isolator: The voice isolator is another straightforward
tool within 11 labs, and essentially what it
does is that it removes the background noise from your audio and
isolates your voice. So to test this, we're going to do one directly into the mic, and I already recorded an audio where I'm speaking
next to the AC. So there is some noise
in the background, and we're going to see
how well it performs. This is the interface when you first come to this
tool right over here. The only thing you can do
here is upload an audio or record into your
computer live. So let's go ahead and do one where I'm just recording
my audio with you guys. I have a microphone with me, but I'm going to keep a distance from the mic itself,
and then we'll compare. First, actually, let's
do one where I'm speaking directly
at a good distance, and then we'll move
back and compare. That is my audio 17 seconds. I just read something
off a piece of paper, and let's hear what it sounds like without any
sort of isolation. Find a sister still living in her small apartment
above a grocery store. The flat is crowded with people, friends who had fled the city
and are now really tired. No one has appetite for food. Okay. What you can do is
delete this record again, or simply click
on Isolate Voice. They find a sister
still living in her small apartment
above a grocery store. The flat is crowded with people, friends who had fled the city
and are now really tired. So that was with me speaking into a mic at a good distance. Let's just your ideal
recording scenario. But what if I was standing
next to a noisy environment? Maybe there's a fan turned on, and there's just a lot of
white noise around me. So that's what I'm
going to do next, and then we can compare
how well this does. So I just recorded
an audio where I have a fan in the background. They find a sister
still living in her small apartment
above a grocery store. The flat is crowded with people, friends who had fled the city
and are now really tired. No one has appetite
for any food. So it's really subtle
in the background, but if I want to do
this professionally, that subtlety is also what
I want to get rid of. I just want my voice isolated so let's see
how this performs. They find a sister
still living in her small apartment
above a grocery store. The flat is crowded with people, friends who had fled the city, and are now really tired. So notice how it completely removed
that background noise. I'll play the original
just for reference. They find a sister
and then this one. They find a sister still
this was only 15 seconds, but you can also do this
for longer audios and it comes really in handy when you don't have a
professional microphone, maybe you're recording
in your phone. There's some stuff in the back and then you can just
download it right over here. That's all you can do
with voice isolator. You can isolate your voice. In the next lesson, we're
going to move on to the products section
of ElevenLabs. There's a lot of cool
stuff there too, and there are a lot more in depth compared
to what we're here. I hope you guys were following
along and trying to mimic the demonstrations
that I was doing in the lessons because in the
next couple of months, things are going to get a
little bit more advanced and feel free to pause or refer back to the older
lessons because essentially what these
products do is that they combine a bunch of these
playground tools and they give you an option to
make more complex audios. Let's go ahead and get started with our first product studio.
11. Studio Part 1: So studio allows you to create
something from scratch. And right away, you
can see up here we have some presets to follow, such as starting from scratch, which is you starting
an audio that could go in any way, audiobook. So it will have tools
designed for audiobook audio. We have podcasts and then we
can just import from URL. Can access studio from the product category
right over here. These are the options that could get you
started right away. You can search for your
projects right here. Right now, I don't
have anything, so there's no results. But you can see there are
some guides down here if you wanted to learn more about sound effects
in game development, voice changer dubbing
and all that stuff. First, let's start
from the first option. Start from scratch.
I click that, I think you saw that it
made untitled project. Let me just go
back and show you. This is my first project and they're just going to
line up right over here. You can open the project, duplicate it or delete it, and when you click on it,
you go into the project. This is how you can
make audios from scratch and decide if you
want to add sound effects, make it an audio book,
make it a podcast, make it a narration,
anything you want, really. First, let's go ahead and
give our project a name. I will call this my scary story. Can start typing
right over here. Then if we don't want
to type right here, there's some stuff
that we can add in like uploading a PDF. If you have essay
you want to narrate, you have a short story
written in the form of a PDF. This is where you upload them. We can add a video upload a video to add
the voiceovers too. You can add a URL. These are some presets for you to get started with,
and that's about it. Let's go ahead and try the
stuff up here and then we'll come down for the other ways
to add an input for this. Right over here, you can change your text into either
different headings or just your regular body text. If you have a paragraph,
you get to use this and show a history
of what you've done. You got some undo buttons. We have a break. So when you add a break, the speaker is not
going to speak. For example, if you
have two paragraphs, you want to break
between the paragraphs. You can add sound effects. Can even alter the
pronunciation of our voice, which right now is George. We can direct the speech
with our own voice, and then we can lock them, then finally get
some information about what we're doing here. On the right side, we
will choose our voices. If you hit Select voice, this is the same library that we saw in the previous lessons.
You can search for it. Go to your saved audios, use professional voices
only the default ones, what you used recently, and even your own voices. Let's go ahead and choose
someone that whispers maybe. The moonlight dances
across the room. I'm here to help you
drift into deep, gentle sleep. You're safe. So I want to do a scary story. So I want a voice that
can sort of whisper. So ASMR titles are kind
of great for this. I'm just going to
click on ASMR mic. Let's add it to our voices so we can easily access it later, and then we can just hit Apply. So there is this
toolbar next to it, which is what we saw previously in I believe one of
the other tools. So you still have a
lot of control as to how the voices
apply to your text. First let's go ahead and write
something very basic for our scary story and see how we can get Mike to
bring it to life. You can see right away
as I start writing, it's assigning Mike to
what I'm writing here. Then I could switch
between George, maybe seven other people or just have one person
saying this entire thing. So here's a little text
that I just wrote. It's only three sentences long. Right over here, I get to
change the layout for the text. For example, I could
make one of them bigger, and maybe if I had
two paragraphs, so let's just write one more. Following him as a long
slender figure lingering in the shadows coming for his soul which he
cherished once. So another paragraph,
and then you can see how I can change the way
the paragraphs look, and this could show that this is a really important
opening and then maybe I would assign a command to it and do
a bunch of other stuff. That's one of the stuff. If you hit Inter and hit
a forward slash, you can actually
add the commands. Instead of going from here, you can also do forward
slash. Let's do a break. For 1 second, Mike is
going to stop reading. You can also click on
this and make it more. I do 1 second. Can do the same thing in between the words,
forward slash. Let's do another
break and I will make this one 3 seconds
maybe 1 second. 0.1 is really little. Let's try 1.6. Then we can also do other stuff like forward
slash, sound effect. Let's see where we can
add a sound effect here. We can do right before on
a rainy day, sound effect, hit Enter and then click
on preview Raining let's make the duration like 1.4 seconds generate
the preview. It's the same thing that we made in the sound
effect generator. You get your previews, hit
Auto if you're not sure, and then see which
one you like best. So none of these are raining. It's still raining on a stone road with thunder
in the background. Something more for 11 labs
to work with. I regenerate. Let's describe the sound, and I will make the
duration a little longer for, like, one sentence. So I like preview two better, hit a flive and there you go. So for this whole sentence, there's going to be
this rain sound effect. If you change your mind about the sound effect
being too long, just click on it and you
get to mute it after a bit. So I'm going to make it a
little lower and tear it. Yeah. So that's a lot better. I want it to be a
background noise, not so much like the main focus. So that's our sound effect. We can also add the pronunciation
editor if we wanted to add maybe a foreign word or something we want to make
sure it's pronounced right. I'm just going to
find a unique name that has to be pronounced. Unique name with right
pronunciation, something random. I do want to test the
accuracy of ElevenLabs. Let's pick up this name. Maybe N would be nice. Maybe 11 labs will
read this as CN, but that's where we get
to fix it with this guy. Can looks back and
holds his breath. He knows what's coming next. Okay. So this is my
story. Really short. Let's go ahead and have
Mike start reading them. So he needs to read everything. 39 seconds. We can change the voice later. Down the road on a rainy day, there was a lone man
walking back from work. His eyes were red from no sleep, his skin sunken from
a lack of food. He drifts in and out of sleep but keeps on walking
with his heavy body, for he knows what will
happen once he stops. Following him was a
long and slender figure lingering in the shadows
coming for his soul, which he cherished once. Kian looks back and
holds his breath. He knows what's coming next. So for me, it read
key and write, but if for your case,
it didn't work out, you have a specific
word or different name, and ElevenLabs is just
not picking up on it. You can go ahead and use the pronunciations editor.
This is what it looks like. You either get to connect a dictionary which you
uploaded or you made. And then you assign rules
to this dictionary. For example, every time
you see this letter, it needs to be pronounced
in this manner. You can find your
dictionaries here, you can add rules. Add new rule, input the word, the Alias, and then the output. Then if you don't want to use this dictionary ever again,
you just disconnect it. You make a new one right
here and these are the default voices that you
get to get a preview with. If you're really into using this for foreign words or words
that are hard to pronounce, I would highly recommend
checking out ElevenLabs on tutorial on how
to do it properly. You get to manage pronunciation dictionaries
with programming. You have to install the SDK via your terminal
using the Linux language. If you plan on doing
this the right way, definitely check this out. Can find it in the
docs from ElevenLabs. Normally it picks up
on the pronunciations. For example, here it didn't
read Can it read Ken, it should work for
the most part. Let's go ahead and give this another regenerate now that we added more sound
effects and this name. Once again, it's giving us free regeneration so that we're sure with our text before actually using our
credits with Mike. Down the road, on a rainy day, it was a lone man
walking back from work. His eyes were red from no sleep, his skin sunken from
a lack of food. He drifts in and out of sleep but keeps on walking
with his heavy body, for he knows what will
happen once he stops. So that was my regeneration. Up here, we also have direct
speech with your voice. For example, if it's reading a sentence in a
really monotone way, this tool will help
you use the voice to voice capability
within ElevenLabs, something we looked
at previously, and I get to guide 11 labs on how to read this sentence
the way I wanted to. That was how we
can use the tools, most of them for this one
particular scary story. In the next part of this lesson, we're going to continue with
the tools up here and then see how we can apply
this to a document, a video, a URL, and expand our applications. I will see you guys in the
next part of this lesson.
12. Studio Part 2: Now we're going to continue looking over the tools up here. So previously, we made
this little short story. We had ASMR mice read it, added some sound effects
and some pauses. But you can actually take
this a step further. So say you have this sentence
and you just don't like how either monotone or exaggerated the
voice is reading it. What you can do is read
it in the way you want and then have the voice
copy that exact style. You can do that
right over here with direct speech with
your voice, right now, I just highlighted this
sentence for reference, but you can do the entire thing or just one word,
whichever you want. Once you have that selected, we just click on that and
you can see it right here. What you can do is upload an audio or record
your voice directly. Basically, let's first hear
how Mike is reading this, and then we'll do our version. Let's regenerate the selection in case there's
anything we missed. Down the road on a rainy day, it was a lone man
walking back from work. His eyes were red from no sleep, his skin sunken from
a lack of food. He drifts in and out of sleep
but keeps on walking with his heavy body for he knows what will
happen once he stops. Was a long and slender figure lingering in the shadows
coming for his soul, which he cherished once. Can looks back and
holds his breath. He knows what's coming next. That's what we have so
far and interestingly, you read this as SN, but we can just direct
it in a different way. For example, here, I wanted to pause between
saying lingering. What I'm going to do is just
cover this entire section, go over here and start reading this in the
way that I want to. Choose your microphone
and just start hit Start. Following him was a
long and slender figure lingering in the shadows
coming for his soul, which he cherished once. Kian looks back and
holds his breath. He knows what's coming next. Once you're done, it's
gonna show up like this, and you can either generate it or just start over for
a different recording. I'm going to hit Generate, and we can now compare
what we're getting here. So if you go right
over following him was a long and slender figure lingering in the shadows
coming for his soul, which he cherished once. Ken looks back and
holds his breath. He knows what's coming next. So now that it was applied, let's hear what this new
version sounds like. You can see on the side here, we have that conversion, and you can just click on it to see which part you converted. Following him was a
long and slender figure lingering in the shadows
coming for his soul, which he cherished once. Ken looks back and
holds his breath. He knows what's coming next. And there we go. We were
able to first of all, have this sentence read very slowly with the
pause that we wanted. Also we had Mike read this
as Kan and not seeing, which is what he
was reading before. You can just really
expand on this with some pronunciations if you would rather use this and this, especially for
foreign languages, this will come in handy. Next tool allows you to lock paragraphs to
prevent changes. For example, I don't
want anything done to this top paragraph by
accident because doing it, it's going to be a
little bit annoying. All you really have
to do is highlight that text and click on the lock. Now you can see a lock icon on the left side,
that's about it. You can unlock it by grabbing it again and just
clicking on the same icon. This one right here
sends a report. Nothing too much you
can do over there. You have your voices,
change your settings if you need to whenever
you make a change, just don't forget
to hit Generate. You can also generate a section only by grabbing it and hitting. Use the bar down here
to play a selection, change the speed, skip through, hit play your
pause, and just see how many minutes or
seconds your audio is. Can also zoom in and out. There's a timeline
actually going to hit Zoom a little bit on this
button first click Expand. You can see that when it starts, this is where the
range starts playing, but I could just
grab this and have it play at a different time. You can zoom in and out of
the timeline like this. I believe it through
so there's the pause. This is the for sound. I could just grab
it somewhere like this and just make my
changes in that way. This was with your classic
text-to-speech feature. But what if you wanted to
do this from a document? What I'm going to do now
is just copy this and open a text document,
nothing too fancy. Just paste it here, remove the different words and
just leave the script. Once I have this, I'm just
going to save it as a dot TxD. Call this my script. Then I will just delete all of this to get started
with a document. Click on this and upload
your document and see how ElevenLabs handles it. Can see we can upload a PDF, a EPAP, THD, HDL,
or even a document. There's a lot of
options. There we go. Now I have my script
brought in from a document. If you're doing a audiobook, this is going to be really handy because you don't have to
copy and paste anything. Then when we have let's
try different voice here. Have Dorth read it.
Now the name switched, go to hit generate, and we
get to see it down here. Down the road, on a rainy day, there was a lone man
walking back from work. His eyes were red from no sleep, his skin sunken from
a lack of food. He drifts in and out of sleep but keeps on walking
with his heavy body. Following him was a long and
slender figure lingering in the shadows coming for his
soul, which he cherished once. Kian looks back and
holds his breath. Okay. So just like that, I imported text, it
starts reading it. But what if you wanted to
add something to a video? I have this video that
I made with runway, which is another AI tool where you get to
generate videos, and there is a voiceover
currently on there, but I do want to try
it with ElevenLabs. There's a lot of
scenery in there. There's a lot of
scevibes and it's a perfect opportunity for us
to combine sound effects, breaks, voices and all
of that in one project. So what I'm going
to do is upload that video and I'll be
back so we can continue. There is my uploaded video. You can preview it. It's a
minute and 10 seconds long. Once again, we have
a similar layout in terms of the voices
and the tools up here. I I play this, you're not going to hear anything because
there is no audio. But what we can do is
go to the beginning. Let's open this guy right here and bring the video
in the beginning. I think it should be shorter. Yeah, it's actually
53 seconds long. So what I'm going
to do is just have this play and write
a script for it. Let's start with a sentence. So that's our first sentence and based on how long it would take for it to become an audio, you can see the
estimate down here. This is going to take
about 5 seconds. If I had play, it's
going to generate. There she was the lone survivor
walking down the ruins. That's just an example.
Obviously, we have to change the voice and do some
other adjustments. So let's move on
to our next scene, maybe somewhere around here. I wrote my script right here. I will add some pauses between
each one just to make sure that every sentence is
dedicated to one scene. This right here is going
to be for the first scene. So right over here,
we do need a break. Add a pause, 1 second. Then I'll add a pause
after each sentence. Okay. And then we can add
some sound effects. So here I think some crows in the back a little
wind would be nice. Let's click on this right here and then go on the preview. So pros in the background. Let's try to keep it for the whole part where
she's in this location. So both of these clips,
once we're done, we can generate preview and decide which one would
work for this scene. Okay. Now you can see that
it goes over both of them. The next part is still about
her being in this forest, so we can move on
to the next one where she or she is
outside the forest. This is when she
actually goes in. So what we can do
for both of these, actually, highlight
this whole section and add another sound effect. So let's do Erie forest sounds with don't want to
do crows again, maybe wind. I'm going to do another
full ten second, and I think the first
one is the best one. I added the pauses and the
sound effects that I wanted. Just to walk you through, we had the first one with the
crows in the background. There's a 1 second pause, the sentence, another second. I added a slow footstep
sound effect in here because we're seeing
there's something following her and
that's about it. Now in terms of the voice, I do want to go for something
a little different. Let's try searching
for thriller maybe. Be it so dirty. No Atmos Segun Mosm Ignorovch. God has given you
Let's do deep meal, something that would
suit this horror look. Eso Hey, come here.
I see that look. Our distrust Changing
all things is sweet. Changing all things is sweet. Okay, I think I like
Josh right now. So what we're gonna do is just grab all of this, choose Josh. Now you can see the
name showing up. So she took to the woods
looking for any signs of life. She found only memories of it. Okay, so I won't play
you the entire thing. But essentially, this
is how it would work. You have your video here. The timeline down
below really helps you figure out what sort
of sound effects you want, what sort of pauses, how long the pauses should be. And overall, you can create really good effects
with this technology. So one thing that people like to do when they have
a video like this, this video will be available for you guys to use, by the way. You can also put in, like,
a vlog or something. Just focus on getting the
sound effects right and then go ahead and add some music
maybe or a voiceover. That's only because that way you can play around
with the volumes. Listen without any voiceover and just hear the sound effects. Do they appear at
the right times? If not, you can go
back and maybe select a different time or simply move them around like we did before. What you can also
do is do it with the sound effects export
bring it in again for the voiceover or add
another set of sound effects. Because this is an online tool when you add a lot of
overlapping layers, it may crash, which is
a very normal thing. If it happens to
you, don't panic. But one way to mitigate that is to export it and
then bring it back in. That way, the first changes
you make will be secured and permanent because you
exported the video and then you can build it up to
as many layers as you'd like. That's about it for
the studio tool, pretty straightforward. There's a lot that
you can do with it, and when we go back,
as we mentioned, we can view our project, go back into it, make
additional changes. In the next lesson, we're
going to learn how to generate an audiobook
using ElevenLabs. This is another
straightforward tool. Essentially, you
just click upload, choose your voice,
and that's it. The next one is
going to be podcast, which gets a little
bit more in depth, but we're going
to look at all of those in the next
couple of lessons.
13. Studio Updates: Now, we already looked at how the studio feature works in
ElevenLabs, but since then, there have been
some upgrades and some new features
that can really transform the way
you use ElevenLabs. So we're still in studio, and usually it would just be a box where you type in
something for voice. But ElevenLabs now allows
you to make videos, music, and create auto captions, as you can see so these are just a bunch of examples
from what you can do. So a faceless video
captions, dubbing, voiceover, video to music,
and generating audio. Your projects are all
the way down here. If you ever made one,
they're going to be listed down below, and you can even filter
through them using these tags. When you go up here, you
can immediately jump in, say what you want to create. You can think of studio as
the all purpose platform. So instead of going
through each one, maybe combining two things, you can just go
to Studio and get a mix of all of these features. You can attach something as reference from
your computer by clicking on this or upload
an entire folder using this. Instead of creating
a video right here, I'm just going to create a new blank project and
we get two options here. First is a video project and
the second is Audio Project. The audio project is what we looked at in the
previous two lessons. But the video project is
that additional thing. So just to go over
audio project, I'm going to click on that. And the idea is
pretty much the same. So we are in Audio Project, and the main features are
pretty much the same. So you get your little
workspace in the center. We can break it into chapters. You can think of this as tabs. We have our voice library, some sound effects that we get to incorporate here, music. This is another new feature
that ElevenLabs has, and, of course, things that you have uploaded for your project. Could be a reference
for a character. It could be an audio sample,
anything that you want. They're going to be listed
right here for you. Now, you can upload to
the project like this. You can directly record
into the project, or you can import files. So, you can use
all three options. And then, of course,
we have the edit tab where all the familiar
sliders are going to be here. We get to type things describing basically that
final project that we want. So it could be a short
story about XYZ, and you just type it in here, and then all those
elements come together. It could be a voiceover
for YouTube video. You can upload the
YouTube video as reference and get the
help of ElevenLabs to create an AI voice for you that would match
the vibe of that video. We're going to do a
very simple exercise right here where I'm
going to type in, like, instant of a horror video where our character is a
little bit scared. There's going to be
some sound effects. We can pull in some music here, get a good suitable voice, and, of course, add breaks
and audio tags. Now, over here we
have our models. We also have three. So I'm going to use that
because of the audio tags, and then we can work
around with the type. I'll leave mine for as text, and then you can add in
some voice, as well. So let's just
explore the library, something that will
work well with, like, a scared character. In the theater of life, we are all actors sharing the
same Just trust yourself. Then you The thing
always happens that you as we are liberated so let's go for default voices
and look for scared. Life isn't about
finding yourself. The years teach
much, which the day? It is not so important
to know everything as to appreciate I have never been hurt by anything
I didn't say. We make our own fortunes
and we call them fat. God has given you one face,
and you make yourself. Gratitude is rich
as complaint is. So I will go with
Eric for this case. We're just going to
click on that name. There's going to be a checkmark, and we could just
go back to Edit. So now we can see
Eric as our voice. We have version three text, and, you know, everything
else is good so far. I'm now just going to start
typing something very simple. So maybe, like, I think
I heard something. It sounds like it sounds
like something is coming. We can add in some
audio text here. So but heavy breathing. Oh, no, no time. Now, here I'm using
the three dots, but if you're using another
model like the V two, you're able to use the audio or the speech breaks, actually. But if you want to use V three for something
more natural, you could just, you know,
do the same thing here. So that's our first line, and here I'm going to
put some Erie music. So we have some songs here, but we could also
generate our music. So I'm going to explain this further in the later chapters. But for now, just describe
the music that you want. So maybe Erie suspense
background music. And I'll just leave it
with the current settings. We're getting three
versions, which is a lot, but you can maybe
change it to one, so it focuses more
on the first one. Choose your model right here. And then the duration, maybe you can do like 5
seconds or something. If you want a zinger in there, there could be, you know, someone speaking, and then
we can fine tune it further. So now, when I clear my search, we can see the three
generations that were made. I'm just going to add
the first one in. It's importing the music. We just have something
for the background, and we get this little timeline. So I get to maybe have it
start here, extend it, change my music if needed, and build this
audio from scratch. I think I heard something. It sounds like
something is coming. Around here, I want him
to say something else. So let's, like, go
to the next line. We're like, Oh, what is that? And then we'll add a
sound effect here, which is another feature. I'm gonna do, like, rustling in the bush. I'm going to do
something very short. Maybe like 5 seconds. So I'll just add this here. Importing, we have
another element now, and you get to just
build upon this. I think I heard something.
It sounds like something is coming. Oh, what is that? Maybe a little sooner. Oh, what is that? Adding more text. And then let's say it's
like a cat that comes out, so we can do something
like Alright. And then while I'm here, I'm just gonna do cat. We can see if there
is something. Let's do cat meow. Alright. And there's
our final sound effect. Now, I just got
to clean this up, and we have our first
updated studio project. Now, before I generate this, there are some
things you could do with each of these layers. The first thing is to cut something if there's
too much of it. For example, right here, I
don't need all of this audio, so I could just zoom
out right here and pull the end of
this whole audio. You could kind of counter this by generating
something shorter. But this is also another way
that you can cut things. And another faster way is to just grab, you
know, your segment. Another thing you could
do is just click on that Audio segment and
hit S on your keyboard, and it's going to split it. You can simply delete that
extra bit with the backspace. Next thing is to
work with the volume and the fade effects for
each of these segments. So, for example,
we had the meow. It just starts out of nowhere. So maybe I could kind
of have this audio fade out and the audio here
fade out as well. So basically, it's going
to lower the volume to zero over a longer
period of time. I think I heard something. It sounds like
something is coming. Oh, what is that? Oh,
no, I can't handle this. Oh, it was just a cat. Silly me. And for some reason, we have a female voice here, so let me just change that back. You can change the
audio just like that. The cat meow is a
little too loud, so let's just bring that
down, and we're good to go. I think I heard something. It sounds like
something is coming. Oh, what is that? Oh,
no, I can't handle this. Oh, it was just a cat. Silly me. And that's the updated
studio interface. Again, it's not that different
from what we looked at, but there are some
additional things like music and this timeline
that we have right here, and you get to import different parts of ElevenLabs
into your next audio file. Down here, we have a bunch of other AI tools that
you could utilize. So maybe enhancing text. If you want to write
something a little better, you can use the
enhanced text feature. So I'm just going to highlight
that. Click on this. Could do that. If
you're importing audio and there's
background music, you can use this
option right here. You can change the voice
with voice changer. So maybe change the pitch, the speed, and all of that. And you can also direct the
speech with your voice. So basically, if he's
speaking in a weird way, I could record my voice with
the way I want him to speak, and then 11 labs will
just follow through. Again, you can export very
easily with this button. You can share,
name your project, and create different
chapters if needed.
14. Audiobook Tool: The next tool that we're
going to take a look at is how we can
make an audio book. When you click on it, you can straightaway upload
your short story. I have a short PDF for you
guys in a resource pack. It's a children's
book by Monkey Pen. They make free children's
book available for you. You can upload anything with these formats or even
write your own story. So let me just show
you what the PDF looks like beforehand,
and from there, we can understand how ElevenLabs converts
text into audio. So this right here is
a short story that I'm going to upload.
This is monkey pen. You can see that
there's a lot of logos, a little preface text, we get illustrations
and all of that stuff. So how does ElevenLabs turn
this into an audiobook? Instead of copying all of these, all you have to do
is upload it here. Once you've done that, you
can remove the document to upload something else and
immediately choose your voice. Go over there. You can preview them and look at the
different categories. There is no great the
years Our distrust. Love. Life without love is
like a Just trust yourself. So I'm going to go with Alice
for this one, click on it. Then we have this option
for auto assigned voices. Essentially, that's going to
switch between characters. This story doesn't really have dialogue, so I
don't need to do that. But if you wanted
to just turn it on and you can see that it's
going to take a bit longer. Simply hit Create and you're going to be
brought into this page, which is pretty
similar to studio. We get all of our text here, and then you choose your voice and add the sound effects
as we learned before. These are the preface stuff, and I could simply
delete them so that the story the storybook
starts right away. Let's go ahead and hit Generate. Command or Control A, so you select everything, your voice, confirm it, then begin the generation. So I generated the audio. First thing you
can do is, again, you can change the name, and then we can just hit play and see what
it sounds like. Hi, I'm Professor Moisture, and I will be telling
you about water. You can call it rain.
You can call it snow. You can call it sleet. You can call it hail, but
it's water all the same. Did you ever wonder how old water is or where it comes from? The answers may surprise you. The next time you see a pond
or even a glass of water, think about how old
that water might be. Do you really want to
know? I thought you did. Did you brush your
teeth this morning? Well, some of the water that you used could have fallen
from the sky yesterday. So as you can see, does a pretty good job with
the capitalization. I did not do that. It came
straight from the PDF. And even the way the voice pronounces this sentence,
it's pretty good. And her voice in general, works perfectly with
the children's book. But again, just
like last lesson, you get to add pauses or even go to the settings here to add more exaggeration to 37%,
lower the similarity. We slow down her speech a little bit and
at the same time, I'm going to add some pauses
in between these sentences. Let's do a pause here
and that's about it. Let's do another Commander
Control A and hit generate. Let me actually save
this, then generate this. Hi, I am Professor Mois Tui, and I will be telling
you about water. You can call it rain.
You can call it snow. You can call it sleet.
You can call it hail, but it's water all the same. Did you ever wonder
how old water is or where it comes from? The answers may surprise you. The next time you see a pond
or even a glass of water, think about how old
that water might be. Do you really want to
know? I thought you did. So it's pretty straightforward, works just the same as the basic tools that we
learned about already. When you're done, you can hit Export and either
publish it to one of the platforms that you see
right here or simply export it as the single file or
the paragraph separately, the audio format,
and there you go. This is where it tells you the amount of
credits it will use. So all of your credits
are over here and yeah. You again have that
timeline down here, zoom in and out, delete parts, add to it, and you can just bring your stories to
life using this tool. Let's move on to the next
lesson where we see how we can create podcasts
with ElevenLabs.
15. Podcasts: Our next tool is creating
a podcast with AI. When you click on it, again, it's like the audiobook. Straightaway, you can
upload your content, import it from a URL, or even use a existing project. If that's the case,
you get to select it like that and then
put in your voices. But let's start over here, which is where you get
to upload your document. Now, one thing that
I'm going to do for the purposes of this lesson is use HGPT to make
us a sample podcast. In your case, you would
be having your own file. If you've written a
script for your podcast, you just save it as one of the formats down here
and simply upload it. But I'm just going to go
into HGPT and ask for a sample podcast with a host and one guest
where they discuss the weather and
how fun Summer is. Make this 5 minutes long. So we even got a name,
which is pretty cool. Our first the host is named
Alex and the guest is Maya. I'm just going to let this generate and we're
basically going to copy and paste it
into a Word document, and that way we can save
it to one of the formats. I'm just going to open one here, a new document, and just
copy the podcast itself. Maybe all of this stuff because
ideally in your sample, you do have a bunch of
information as well. There is my podcast. Let's save it. Name
it my podcast, save it as a dotxem. It's save and
upload. There it is. You can remove it to
put in another one. The next thing
that you can do is choose the format
of your podcast. Focus podcast with key updates, casual conversation between
a host and a guest. We did ask for two
people in the podcast. This is what I'm
going to go for. Then we get to
choose the duration. I'm just going to
go with default, then hit X when I'm done. These are the voices for
our people, the host. A single rose can be my garden. And then the gift. If you spend your whole life
waiting for the storm, If you go on any of these, so let me just search
for Jessica real quick, you can see that one of her
tags is conversational, and I did ask for a podcast where two people are
having a conversation. You can also look for other conversational people and then make your
choice from there. Hi there and welcome.
Nicht to the Dao. The default ones
are by ElevenLabs. They're pretty good usually. Nature as a mutable cloud. We make our own fortune. But I'm just going to keep the ones that ElevenLabs
chose for me. Down here, you do
have some settings. So what is the podcast focus? I think we said it
was hot weather. Let's add another highlight
summer. It's safe. So what it's going
to do is that it will emphasize those words
within your podcast, and I'll just show you
what that looks like. Lastly, we have the language. I'm going to stick with English, but we will come back
to try a different one. Let's generate and see what
our podcast looks like. Okay, so there is my audio. And we have 6 minutes
and 26 seconds. That's because it
has the pauses in. But if you saw that it's way
too long or way too short, you can always go
back to Cha hBT and ask for it to either
make it 3 minutes, 20 minutes, however
many long you want it. Let's go ahead and listen. Did you know that
heat waves now kill more Americans each
year than hurricanes, floods and tornadoes combined? The way we experience summer
is fundamentally changing, and today we're diving
into how communities are adapting to these
rising temperatures. Those statistics are
truly eye opening. I've been reading about
how different cities are developing innovative solutions to deal with extreme heat. Well, what's fascinating
is how this has become a major public
as you can see, it does switch between the
people, and down here, we're actually seeing the colors assigned to Chris and Jessica. One thing that you may have
noticed is that by default, ElevenLabs removed
the first part of our script where it just talks about the description
about the podcast. So it cuts immediately to
what these people are saying. I'm not seeing guess parentheses Maya or the podcast title
or anything like that. That's pretty impressive.
And what you can do is alter the voices
at the same time and have it regenerate
everything. For Chris, I think
that he could talk a little slower and be a
little bit more stable, actually a little less stable. He sounded a little bit too robotic and a little
exaggeration. Let's do something
similar with Jessica. I think she could definitely
slow down with her voice. I think we didn't
save. Chris actually lower this and make this a
little less stable. It's safe. One more thing that
I do want to do is add in a few words of, like, pauses, human errors to
make this poadcast a little bit more realistic and that's something ChatBT
could help you with. I'm just going to paste my text. Take this text and add and
pauses to make it sound. More human. Also, let's
reduce this to 3 minutes. Quotation marked,
paste and quotes, and then hits Enter. Just a disclaimer,
this is by AI, the information here
may not be true. So just take all of this
with a grain of salt. If you want to make
statistically correct podcast, be sure to choose
reliable sources, but this is just for
demonstration purposes. These numbers, these places, programs, they're
probably not real. Remove the titles
for host and guest. I just want the dialogue
separated by paragraph breaks. All right so this
will work just fine. You could have it do that to remove the titles
for host and guest, but then you would
have to assign Chris and Jessica to each
of these sentences. What I'm going to
do is just copy the previous response and then re upload it
into our project. Command or Control S, let's go ahead and first of all, delete everything and
choose a document. Let me actually go back here, start a new one so I don't
have to edit anything. The same settings as before. I will just add the highlights
here, maybe summer. Here, what you want to
do is actually look at your scripts and find something that you'd
like to emphasize. I see kids here maybe,
maybe temperature. If it turns green,
that works for this focus because
it's in the script. I think I'll just
go with one since there's a lot of different information being
passed around here, but temperature is definitely something we're talking about. Save the changes. Let's
put this on English. The voices are fine,
and start generating. Okay, let's play and see
what we got this time. Here's something that
stopped me in my tracks. Heat waves are now killing more Americans than hurricanes, floods and tornadoes combined, and it's completely changing
how we live in our cities. That's such a
striking statistic. And what really gets me is how this isn't some
distant threat. It's affecting
communities right now. Okay, so I have my
awards right here. It kind of brought
in the last script, but I could just switch
it out like I just did. Hey, um, did you know
that heat waves now kill more Americans each year than hurricanes floods and
tornadoes combined? Yeah, it's wild. The way we experience Summers, So if that happens to you, just go back to your script
and copy what you want. And here, I'll add it to what
was there before, paste it. And what you need to do is
highlight this entire section and only apply Jessica or
whichever voice you want. And this way you
can switch it out. But now that there's
ms and pauses, it sounds a lot more human. That's exactly why we
went to ChachipiT. Once again, I will lower
down his speed a little bit, add more exaggeration, save it. For Jessica, I think she's fine. Maybe make it a
little less stable. Her speed is perfect
for a podcast. But once we're done, we can
regenerate just a bit of it so we don't use a lot
of credits generate. Hey, did you know that
heat waves now kill more Americans each year than hurricanes floods and
tornadoes combined? Yeah, it's wild. That stat is honestly
kind of scary. I've been reading about how some cities are coming up with these really creative ways
to deal with extreme heat. The way we experience summer is, like, fundamentally changing. And today we're just
diving into how communities are
actually adapting to these rising temperatures. That's such a
striking statistic. And what really gets me is how this isn't some
distant threat. It's affecting
communities right now. Alright, so there's our
simple podcast example. Let's go back to the first menu and take a
look at what these are. So input from URL could
be like a blog post, something on the website, on your social
media or something, and it's going to use this technology to make a
podcast that you can listen to. I will try this with a blog post maybe and then see how it
uses that technology for it. I will find a URL and
then I'll be right back. Let's go to Wikipedia and
get a very simple article. I will search for maybe
apples and simply paste this URL into the bar here. Then again, you can choose
which type of podcast. This time I will go with a bulletin because it's
facts about an apple. Let's make it really short. A single rose can be my choose a different voice this time. God has given you one fame. Just a British voice. Language. We can do English. And then we can see
what that sounds like. Okay, here is our Apple
podcast. Let's hear it. Apple's remarkable
journey began with its wild ancestor malls
ciersi in Central Asia, evolving into today's
malls domestica through millennia
of cultivation. Commercial orchards rely on clonal grafting rather
than seed propagation, as seeds can produce
unpredictable results. This careful
cultivation has led to over 7,500 distinct cultivars, each developed for
specific uses from. So as you can see, I'm not
going to play the whole thing. It basically summarizes
entire Wikipedia page and gave it a more
conversational speech. So the way this started
is that it talked about, you know, the
remarkable journey, which is from, I believe, the cultivation section of the the history
section, actually. Started from here
because it knew that you're going to
need some context. Then it went into the ancestry, when it was sequenced,
even the genome, which is a whole
different section. It's actually pretty
impressive how it takes little bits of information
from an entire web page, such as Wikipedia, pretty large, and it makes it into a podcast. That's pretty good.
There's always you can change the voice, add break, sound effects, or anything else that you want. But the last thing that
we want to try with the podcast is existing project. Right now I have four
in the background. All you have to do is choose
one. Choose the format. I will make it into
a conversation, short, choose my host and guest. She tore her gaze away from her ruins to Arabella and Adam. People who are crazy. Let's make it Japanese and see how the different
language works. So as you recall, my scary story was just me writing a
bunch of sentences. We added some sound effects,
but there wasn't much. Let me just show you
real quick what we had. Not a lot of sentences going on. I'm going to have a podcast
be made out of this. So let's choose our
project conversation, short, and I'll choose my
voices again in Japanese. So here it is our
story in Japanese. I'll give you a preview, and then we'll turn it
back to English so we can see what sort of
changes ElevenLabs did. You show sakaka dee tomo Koa
Nova Mirena, kommen Kiowa, Tata tore Umacho So done. Hi, ShotunaaKo jo too. So great conversation
right there, I'm assuming. Let's go back over here and do the same
thing, but in English. That way, we'll know
how it witches. I'm going to use the
Japanese version as my project. Same settings. Well let's actually go
back to our chosen ones, make it English this time, and then we can compare. Here is my first of
all, converted text. You can see from our
short story that I wrote, it turned it into
a conversation. Let's go ahead and hit
Play and see what we got. Ever wonder why the
scariest moments in horror aren't the monsters, but those quiet seconds when
you're completely alone. There's something deeply
unsettling about true isolation. That's exactly what makes
this story so effective, the abandoned factory setting where literally no
one is breathing. So we took our
very simple story. I'm gonna open it
for a second to show you so this couple of sentences, I'm just describing this video, it's basically it made a
podcast based up of this story. So these two people are talking about our character
lost in the Forest, but you can see what
a great conversation was brought up all using AI, and we can switch between
the different sentences. Podcast feature is really powerful and there's a lot
that you can do with it. In the further chapters, we're going to do a
podcast from scratch. Don't worry if you didn't get a lot of practice
with this right now. We're just going over the
basic capabilities of each of these tools so that we are ready for those
interactive projects. Now, that concludes
the studio section. Import from URL just lets you put in a Wikipedia
page and for example, it could be a blog post,
Wikipedia, anything, and it's just going to
simplify it for you. Whether it's a podcast, a audiobook, a narration,
it doesn't matter. ElevenLabs will
decide what works best and will give
you an audio instead. This is pretty straightforward. You just paste your URL, choose your voice,
and that's it. If there were a conversation involved, you can
just turn this on. But it's a combination of the tools that we
already looked at. I will not go over
this right now. That is all for studio. In the next lesson, we're
going to go over dubbing, which is basically
where you get to put a audio on top of an
existing audio to, as it says, localize
the content. Currently, there's 29 languages, and I'm pretty sure in the following months,
there's going to be more. Let's move on to dubbing.
16. Dubbing: With dubbing, you can translate your original clip using one of the 29 languages
that are available and basically decide
where the voiceovers go. So this is the
interface right now. There are some tutorials for you to look at to get started. But the first thing
you have to do is simply click on
Create a new dub. When you do that,
you're going to be making a dub project. So I will call this
my scary movie dub. The source language is the language that your
original video has, for me, that is in English, and then you choose
your target language. I think I will try
something, let's see. Let's do Spanish, for example, click away, and then you can put in your audio
or video source. Apart from your
classic upload button, we have YouTube links, Tik Tok Links, other
URL, and manual. So if you want to do manual, just to walk you through this, this is the video file. A CSV file is going to be
your transcription file. Foreground audio file is
going to be the speech, whatever that's in there, and then background could be
a little music or something. So you would only
come here if you have the separate files
for each of these. If you don't, it's
the same thing as doing all of that
in one file here. So I will upload this video, which I'll show
you guys in a bit. They say she woke up
with blood in her eyes, wrapped in metal and silence. The plane was gone.
Everyone else gone. So this is a video that
I made with Runway AI, which is another tool, as I mentioned in the
previous lesson. But this time it
comes with audio. And as you heard,
it's in English. So I want to turn
this English audio and have a Spanish dub on top. And this video is available for you guys to download
in the resource pack. So let's go ahead and
upload it right here. You can preview it once again. They say she. Then when you scroll down, you can delete it to put it in another one or simply continue. If you create the
dubbing project, you're going to be able to get a timeline where you adjust
the different audios. Definitely turn this on unless you have a very
straightforward project, and it's just like a plug and play thing where you
don't need to edit any. Number of speakers, it
can detect it for you, but you can make sure it does it correctly by
choosing one of the numbers. I only have one speaker,
so I'm choosing one, and then the time
range to dub for me, it's going to be the full
thing. 02 53 seconds. Disabled voice cloning,
turn that on so it doesn't make an identical
voice. That's it. You just hit Create Dup, it's going to upload it, and now it's going to
convert it to Spanish. It's using the playground tools and some of the more
audio into Spanish, maintaining the natural tone, the mood, and all that stuff. I will let this
run and then I'll be back so we can
see what we have. As you can see, it's done, I can simply click on it
and move into my project. As you recall, we enabled the dubbing project and that's
why we have this option. If you do not enable that, you will not see this interface. So here is the
transcribed audio. You can see the
text, and it does match to what we hear
in the video itself. So I don't have a
lot of information, a lot of audio in this video. It was like 53
seconds, I believe. But right away, we can see that it separated the
background music, the foreground music,
and then the speaker. So right now, I'm
not going to do anything and just simply
play it for you guys. They say she woke
up with blood in her eyes wrapped in metal. So you can see it's still
the original video, and what you can do here is play around with the different
audios it extracted, add something, remove it, and do a bunch of stuff. But before I get into
that, let's first look at this interface and
see what's happening. On the top left, we have the transcription
of the audio, as you can see,
can also click on this so that it
redoes it for you. That's about it for this
section right here. On the right is the preview
of the actual video. If I move my playhead, you can see that I'm
going back and forth, and that is synced with
the transcription. Speaking of the timeline,
we have layers. These are three
layers right now, but if there's nothing on there, that means that there's no
foreground music right here. This is the background
music alone, and this is the
speaker speaking. You can see here the speaker
doesn't say anything, but the music is still playing. Let's hear what that sounds. Not empty. Every steps. You see it says
something, pauses, says it again, but there was still music
playing in the back. Over here is the
names of the layers, and you can change the name, not for background and foreground,
but for your speakers. If you have multiple speakers, you can give them a
name and this will keep things a lot
more organized. I can call this Jake. It's labeled original because we're still in our
original video. Everything you do is going
to be saved automatically, so don't worry about there
not being any safe changes. On the right, you can
only enable that layer. If I click on background, you can see that it grade
out the foreground and Jake. When I hit Play, even
though I'm on this part, I'm not getting any speech. Let's do the same thing
with the Jake audio. And the trees. She
walked for hours. It's just our speaker. You can see how clean
it separated the two. I only gave it one video. On the side, you
are able to delete the speaker and add a new one. You can use this to go
across your timeline. On the top right, we have how many seconds you've elapsed. We have the select, which
lets you select things. We have the trim if you
wanted to trim something. I'm going to select
the background, hit the scissor icon, and then you can see that I now have two sections
of the same audio. Go to hit Command or
Control Z to undo that. That is only one on this end, we have add clip so you can add up different
video if you want. I'm not going to do that,
but if you did want to, you can decide where it goes, maybe here and then maybe turn this into
an Saffx or something. We'll get into that later. To delete, just select it and hit Backspace
on your keyboard. Now, with every
one of the audios, you have the option
to change the volume. This is what it sounds
like right now. Let's only highlight
the background music. Now let's go back and
reduce the volume. We can barely hear it. If you see something
is too loud, you have the option to adjust. All right. That is
the basic interface. Over here, you're able to change the name of your
dubbing project. Down below, you're able to see
how many credits you have, and this is where all
the dubbing happens. You have the original,
which is the original clip. If you go to the
Spanish version, we're going to get
Jake in Spanish. Here's how it's translating the original audio into
the Spanish version. Let's see what that sounds like. Inches Esperto consangre
L o Sojos and well t. Let's actually
increase the audio for the background music. DenchesEsperto, consangren Los Sojos and
Welt metal is IlensioEabon, Able Zaparsdo, Todos
Los mas, desaparzidos, solo Aga, Ilosarbos, caminodor
teoras, Talves dilas, lb so I don't speak Spanish, but if you do, you can see how it translated
these word by word. If for some reason
you saw that it's missing something or it
didn't translate right, you have the option to
change things up here. I could just delete that word, for example, and
quote something else. When I do that, I can hit
Generate audio again. If you go to Jake Spanish, there's this gear icon, and this is where you get
to choose the voice and the familiar settings that we've been working with so far. First, if you go to voice, you can choose someone here. Let's see. Not a clone. I have crossed hope. Let's actually look for Spanish. Be ambitious in your goals and
diligent in amateur voice. Let's, um, try to
do, like, a male. Let my voice be all right,
this is pretty good. Let's add it to our voices
so we can access it easily. Choose our model. I'm going to go with the
recommended and then if we want, we can change these sliders. I'm just going to keep it at
default and then close this. That's the control you
have over the voice, and then you are able to
add even more languages. If you want to do a Spanish dub, a Japanese dub, you are
able to do that easily. Let's do a Japanese dub. It's going to
process everything. Immediately, we can see
the Japanese version. That is looking pretty good. Metals lensoabon Able
So that's in Spanish. Let's sit still and get the, you know, generated
audio for Japanese. Gabe, KinsoctoKs in its marite
Mesa metato, Yareto. He co So I use that new voice. I believe we chose Christopher. That's why he's sounding
a bit different. Now, you are able to create a voice from the selection
or generate audio. These are the familiar settings. Going to see if I can
extend this a little bit. That's a volume from
your original clip. If you grab something like this, the language that
you're currently on, you can create a voice
from this selection only. So if I do that, let's
give it like Paul. As a name, create. Now, whatever adjustments I did to this particular clip becomes a voice model that I could use with other tools
within ElevenLabs. Just like how we
have these voices, I just made one called Paul, going to see if I can access
it from here. There we go. There's my Paul that I just made and you can create
as many as you want. Now, for voice settings, you can inherit the
track settings, which is whatever you do in the track is going
to be applied. If you disable this,
you can do what we were doing in this skeer
icon but over here. Choosing your voice, your model, and then configuring the
voice with these sliders. Now, you're also able to
dictate it yourself and it will follow the
way you're reading it to adjust that voice. But you can also
have your own voice playing over this if you want, and that's all done
with this button. This clip is only 20 seconds. That's why we see it right here. Whereas the clip
the entire movie, let's call it, is
a little longer. Okay. Then down here, we have a bunch of other stuff. First one is add
dubbed speaker Track. I could add a different speaker. You can see we have another one. Add voiceover track, which you can do with
the text-to-speech, for example, like that,
let's increase it. I don't have any audio
right now, the source, so I could put something, something like that,
let's generate. Now because I'm in
Japanese is in Japanese. Oops. Let's only highlight this. H, I'm lost. This is the original, but we
can do the Japanese version. I stretched it out, but it's really short. Going to hit Generate
audio again. Myota. There is our short, Oh, I'm lost in Japanese. Can have both of them
play at the same time. If you wanted to, let's
move them around. Ho, I'm lost. Myota. I could just switch between languages
and go from there. You can extend like this to
make it a little slower. It sounds like a
monster right now, but the more tight
these bars are, the more natural
it's going to sound. L. So I make it really tiny, it's going to be super fast. Lyota. Then we have
add SFX Track, which is your sound effects, which we already
know how to use. When you hit Plus, you're
going to get a new layer, and all you really have
to do is first of all it's enable sound
effect, go down. Let's first actually
make a space for it and describe our sound. Let's do a parrot. Generate the audio. It's showing the SFX in English and Japanese,
but it's the same thing. That's our parrot.
We can alter it. We don't get different
samples like we usually do with the
sound effects panel, but if you don't like this, you can hit regenerate right on top or delete it with
backspace and start over. You can also do a longer one, a longer sound effect. Let's do raining or
just heavy rain maybe. Generate audio. The
reason why you see both English and Japanese is because that's going to be
copied into all of these guys. I don't know how that
sounds like heavy rain, but you get the point. You just have to
be very specific. Let's try to do heavy
rain pouring down on Stone Road with thunder
in the background. Something more specific.
And there we go. Sometimes, you got to make
the prompt a little longer. Then I could adjust the volume like we did
with the other ones. And when I have everything play, this becomes really important. So let me just go
over, you know, this section where we're
getting a bunch of things. Mechintakega, NokoTta
Kano JawananjiKamo, Aukiso, Ariwa Nanihim. So you see how this
needs to be a lot lower because we have
our speaker negative 27. Amo, Arakio. We actually have to
alter the Japanese one. Amo Aukisok Ariwa
NanihimKamo Shini now we can see that it's
a lot more natural. You can go to the
Spanish version or stick to the original. The last thing you can
do is upload audio. You can upload your
very own sound effects, your own voice overs or
anything else that you want. If you made an audio
with a different tool, you can export that and
import it right over here. Existing music, sound effects, backgrounds, and
all of that stuff. Now, keep in mind that the
speakers won't be detected. The only way to
bring in speakers is through the original
way that we did, and that way,
ElevenLabs is using the correct technology to do
this whole dubbing process. Okay, once you're done, let's say this is my
finalized project, you're able to hit
Export down here, export it in whichever
language you want, the ones that you made down here as the output
that you want. You can do the video and
the audio, just the audio, only the caption, transcribe, a zip file, anything
that works best for you. So let's hit X, and that's about it with the dubbing tool. It may look complicated when
you first come in here, but it's pretty intuitive
when you look at it. You have different sounds
on different layers, different languages
well organized down here. Got a preview. You can alter the text, and it's a pretty handy tool. Now that we learned
about this new tool, we can move on to the
next one and see what other features
ElevenLabs has to offer.
17. Speech to Text: Speech-to-text is
another great tool that you can utilize when
it comes to ElevenLabs. Now, it's very accurate and it's getting better as we go on, but this could be used for meetings that you want
to transcribe or a separate audio file that you want to just convert
into a different format. Now, the way it works
is that you have different transcriptions
right over here and different speakers. Let's say you are always using ElevenLabs to transcribe
meetings at your company. There's always a certain
number of people, and you know everyone by name. You can add those speakers, those participating
members as speakers here, and ElevenLabs will
mark that for you. So when you put in a
recorded meeting file, it's going to identify
the speakers and then mark that for you in
the transcription file. So that's a really cool
thing that it does, and you can get started by
clicking on transcribe file. This could be different things. It could be something you
upload from your computer. You could record something
yourself right now. You can put in a YouTube
URL or any other URL. You can have the language
said right here, so many options, scroll down and choose the
language of your choice. And you can also have 11 labs write down
the audio events. And the way it's
describing that is that if someone's
laughing in that audio, it's going to put laughing. If someone's crying, it's
going to put crying, and that's going to be
included in the final file. If you wanted to
ignore those things, you can just turn this off. But it's a pretty cool
feature that it has. Now, the file that
you're going to get can also include a subtitle. Let's say that you want to transcribe an audio and
then use it for your video. Instead of putting that back in for a different
subtitle file, you can have it all
be done in one go. That's a separate thing
that you can turn on. And this feature is going to remove everything
that's a filler sound. So if someone's coughing, if there's a pause, any ums, you know, these voices are
going to be removed, and 11 labs can ignore those and not include
them in the final file. Now, we also have the assigned
speakers from library. These are these speakers
that are going to show up from the tab that
we looked at earlier. So if I turn this on, I
currently have no one saved, but we're going to
see what that looks like once we add some speakers. We have some key terms. Finally, we have key terms. Say you are working on a very important project and
you want 11 labs to make sure that it doesn't
miss anything with the word project
or something specific. You can add those key
terms right here, and 11 labs will
take extra caution to bold those items
and not miss them. So if you don't put these here, it may miss a couple of things when they talk about
that boosted term. But that's only if you
have a boosted term. If you don't don't need
to put anything in here. Let's get out of this and take a look at what the
speaker looks like. We're going to go
there at speaker, give them a name, speaker ID, and an audio sample. So this sample is how ElevenLabs
is going to detect who is speaking and will tag it with their name in
the transcripted file. What we're going to do
is first get an audio. For this, I'm going
to use this video. This way we can all
grab the same clip, but feel free to put in your
own audio files as well. Because we are able to
put in YouTube links, we can easily just grab
the URL right here, copy that, and paste
it right here. So just switch to YouTube. Past it, and we're just going
to change a few things. So this is in English, but if you leave it detect, it will do that, you
know, on its own. I could turn this on
or off as we said. I'm just going to turn it off. Say, I just want to
get a quick summary of what happened
during this meeting, and I just want to remove
anything that's unnecessary. Now, we can add the
key terms right here. I'm just going to take a look. I think we're going to
talk about meetings here. I've heard some of my
mentes talking about how it's really hard to
get out of bed on Fridays. So we're just going to
focus on the key term bid. Say, that's, like, something
really important for me. I'm just going to
put this right here. And if you hit Inter, you
can add more keyterms. I'm going to skip all
the way to the end. Let's get another key term.
Are you meet with him. It's a lot to deal with for middle schooler so
middle schooler. I'm going to put
that in as well. Alright, so we're going to
start transcribing this. It's only a minute long
without any assigned speakers. And then I'm going
to show you how to do the speaker end as well. So let's click on this and
have it work on this file. Now, if you want and you have certain people in mind
that you want to save, you can use the speaker
feature right here. Add those people in,
and those will be kind of bolded throughout
the transcribed file. And that way, you know, when you have a meeting with 20 people, you can easily filter
through and find out what John Smith said
during that meeting. Only thing that you
are going to need is a sample of
their voice alone. Now, because this
is a YouTube URL, I can't really, you know,
put the whole thing. But if you have that
person in mind, you can just put in
their audio alone. So that's that. Let's
take a look at what the transcribed file looks
like. So there we go. We have our speakers
listed as speaker, but we can now switch
their names out and then later save
them to the library. So you can either start
with the speakers saved or save it from
a video like this. The first one we found
out their name is Mara. I just scrolled down on this video and someone
has it all transcribed. So I'm just going to
copy that over and, you know, change the
default things right here. So we have our first speaker, and I think this
is the same one, so I'm just going to,
you know, click on it, and immediately it
did update it for me. The first speaker is
the next speaker, I'm assuming is this person. Oh, we're actually
getting some cues, so glasses is f. I'm
just gonna switch out these names so we can
save them to the library. So speaker two, who's this guy? Just skip a little
bit here to find out. So that's her. I'm
assuming it's this person. We're just gonna put
that in as Speaker two. Alright, click away
Speaker three. So this person must be Thomas, and then we have Matthew. And we can see that 11 labs
did detect five people. It's pretty accurate right now, and I have these
speakers now saved. Now, I could add a speaker, say it forgot to add someone, but usually when
the audio is clear, it doesn't do that, but just
in case you can come in. And if for some reason the
speakers were mixed up, you can grab the edges
right here and decide or maybe fix the duration
of each person's speech. So that's that. I'm
just going to undo what I did by clicking
on that button. Now, on each of these guys, we can change their order, but these orders are based on the normal flow of the video. So unless it's wrong, you don't really
need to change them. We can also change the
color of the icon. If you want, just
to set them apart. And here is where we get to add them to the speaker library. So I could give them a name. And now that it has the
sample from the video, I just need to give
it an ID and a name. So I'm just going to
copy this person's name and then put it right here. So the next time I have another meeting file with
this person speaking, it's going to recognize
that it is Maha, and it will hold that
in a interface like. Alright, so we're going
to check all the text. There's also a spell check feature here if
you don't want to, you know, read
through everything. But let's take a look
at what's on this side. So first of all, is
the name of this file, this audio file, which
is just the UT BRL name. I'm going to leave it as that. Language is English. I did not add any subtitles, but if I wanted
to, I could just, you know, that and I'm going to get a separate subtitle file. So that's something you can
explore if you need to. We have a spell checker. I click on it, it's going
to go through every word in English and let you know
if there's any errors. Thankfully, we don't
have any of those, so I could just collapse this. Next, we have the
segment properties. So this right here
is the segment, the duration, the start time
and time, and that's about. So if everything looks good, all you have to do is go
over to either Export it or copy to clipboard if you're just
pasting it in a Google Doc. But clicking on Export is
usually the way to go. I'm going to click on
PDF just so that it's, you know, accurate
and easy to share. Over here, you can
add timestamps. So at this second,
this person said this, and we can now also include the speaker names because
we did change each one. Click Download, and now we have ourselves a transcribed
file. There we go. First speaker, second
speaker, third one, and it's telling
me the duration, the end of that segment, the start of that segment, and, of course, what
they are saying. If there is a pause, it's going to break it down
into multiple segments. So technically, the first
few seconds of the video, only Mo is speaking. But because she paused
during her speech, there is two segments now. Back here, you can just share it directly on ElevenLabs
and not anload a file. So you can search for people
on your ElevenLabs team and, you know, create a workspace. Say you're working with different editors
and you want to, you know, make sure
everything is good to go. This is a cool feature for that. Going back to the page, I should have one speaker saved in my library,
and there we go. So I can have the
speaker for later use, and if I wanted to delete it, I could just click on this here. Now, apart from the classic
speech-to-text feature, we also have a real time model that is still on the demo mode, but we can still give it a go. Basically, you can have 11 labs detect the language and just speak to
it in real time. So this could be in between a bunch of people
that are speaking. It's going to be all
real time transcription. So when I click on transcribe, let's give it some permission, and we're going to have 11
laps take a look at what I'm saying and write it
down right here in front you can see that
this is pretty accurate. And just to check out the
whole multilingual claim, we can switch between languages. I could now switch
to, let's say, Turkish and say, Mava, it's going to write
it like that as well. But once you're
done transcribing, you can click on
this button and I could copy this over
to wherever I want. Maybe I don't want to spend
a lot of time typing. I'll just say it into
my mic, copy it, put it into Slack or
any other platform. That's the scribe feature. So go ahead and give this a try. It's a really fun feature
that ElevenLabs has. This one is also really cool, especially if you can switch
to different languages. Although it doesn't currently
support every language. If the one that you speak
is within this list, feel free to try it
and see how it works.
18. Production Tools: Now, let's go over
productions with ElevenLabs. Right away, on the left side, you can see that it's
labeled with Alpha, which means you're going
to need to upgrade or get that specific plan to be
able to use productions. So like I said before, I'm not going to go too
deep into how to use this because these are
human edited tools. And in this course,
we're trying to focus on the AI aspect of ElevenLabs. So right away, you can see the services that they provide. We have transcripts, captions,
subtitles, and dubbing. These guys are coming soon, but basically what
we do and did with AI is now being done
by professionals. So think about if you were to go to Fiber to get a
transcripted file, you can now do that
within 11 labs. As we said before, if you see that the AI
tools are just not working well for you and this is a very high quality project, you can switch over to the production services to get something a lot
more professional. As you can see, you do
get charged additionally, $2 per minute and
the prices change. But let's just go ahead in the transcript and I will
show you around a little bit. When you click on it, it's the same interface
as the AI one. You upload your file. It gets reviewed by
the native speakers. Choose your language right here, give it a name, and then you
can even add a custom style. Any instructions, for example, really emphasize the
word temperature within this transcription or add emotions in there,
something like that. We can also add the
sounds which aren't words such as coughing
or dog barking. If you want to get rid
of that instruction, you just click on remove. It's pretty straightforward. The rest of the three
things up here are not available right now at the time that I'm
making this course. But what they have is
that you can tell them what services you want for captions and they will
notify you when it's ready. You just fill this out and
then talk about your use case. Now, down here,
you're able to make a new production and create folders to keep things organized if you plan on
using this tool very often. When you click on
New Production, it only lets you do
the human transcript because that's the only
thing that's available. But later on, hopefully soon you will be able to
do the other stuff. Captions, a human will be
writing your captions, subtitles, another human will be making you subtitles
for different languages. They would be the ideal speakers within that target language. Then dubbing, they
will translate your work and then record
the dubbed version for you. That's the production
tabs right over here. If you go to Audio Tools, there are some more
things that you can do. For example, Audio native. This one is again with a
different subscription. If you, for example,
have a blog or a news website where you
publish a lot of texts, you would like to make your
website more accessible by having an audio
version of that text. This is where you would come and if you click on Play Video, they will show you an example. For example, this
is a news story, and this is how the
audio native would work. Not only are you letting people read your work
in a podcast way, but for people
with disabilities, they have the option to come on your website and be one
of your frequent users. It's pretty cool,
and this is what the bar would look like if
you were to implement it. That's Audio native. We also have Voiceover Studio, which is going to basically
combine a bunch of the tool. We already looked at how to do sound effects and make
voiceovers with dubbing. This studio basically
combines the two. If you scroll down, we
have a demo right here. So we have a voice track and
then we can put in our text. It's just like the dubbing
studio but for voiceovers. The next thing is AI
speech classifier. Now, this one is really cool because you get to
see whether or not this audio that you're about to upload was made
using ElevenLabs. So if you are suspicious and
you just want to make sure that the person who sent you that audio is actually
a real person, you can let ElevenLabs
find that out for you. You just upload it here. There's not much you can do. The next thing is
the soundboard, which we already
looked at previously, but I'll go in it once more, takes you to this
whole different URL, and it's basically
what we did where we combine different
sounds over here, we made folders and presets, looked at some of the
available default values. Soundboard is free. You don't need to
upgrade your plan. Going to the
subscriptions real quick, you can see the differences. This is probably the
one where you get to use a lot of these
production tools. If you were a bigger company, you would go up to scale. This is what I'm currently on, which is the most
popular package. But if you felt like you
want to use the other ones, simply click on Upgrade
and it's going to charge you more to the card that's
attached to your account. Okay. So those were some of the production tools that I
just wanted to show you that they are available
and you are able to upgrade your plan to use
them whenever you want it. This concludes our
second chapter, which is looking at
different ElevenLabs tools. In the next chapter,
we're actually going to start building things
with ElevenLabs, starting from basic projects and then moving on to
more advanced projects. This is a great time
for you guys to follow along and build with me and that way you know exactly how to use this tool
for specific cases, troubleshoot using the
knowledge that you've earned in this chapter and
after this course, hopefully you're able to create your own audios for your
videos, websites, et cetera. Now let's move on to
the next chapter.
19. ElevenLabs Music: A recent feature that ElevenLabs now provides is
the music feature. You can now describe
the song that you want, and it will generate it for you using different references, and you have so much control than just getting
a random music. So this is where you can
access it, and immediately, we have this chatbot that we
could you describe the song. You can see the templates
just coming up. But it's very similar to any other AI generation platform in the sense that we could
get different versions, choose our model,
set the duration, and then of course, we
are using some credits. Down here, if you scroll down, you can see some samples of the music that was
made in ElevenLabs. We can take a look at
their beats per minute, download them if you want, add them to your favorite list, and even adapt it
to your project. So when it says adapt to
your project, basically, you're kind of remixing that music for your
either voiceovers, videos, images, and all of
that within ElevenLabs. Let's hear a few
of these samples. So we have a few
instrumental examples, but you could also have
someone sing in these songs. So that's something I
will show you as well. But if you want to
quickly browse through these songs and find
something for your podcast, you can click on this, and it will list the ones that
are most suitable. So what I could
do with this one, let's say my podcast
is 30 minutes, not 2 minutes, is to
adapt it to my project. I would have to use credits for this because someone
else make this song, and then I could extend
it to 30 minutes. But I could also
make my own songs, so I don't have to use, a lot more credits, basically. Can also browse
for similar songs. It's going to list
them for you based on the texts that the song
has and the duration. We get a bunch of different
things to look at, and this is all within
the marketplace. Now, if you don't want
to download anything, you can go to Generations and just simply
describe the song. Now you may find it a little
intimidating because, you know, you have to put
in the right keywords. But the good thing about
this is that it gives you a little helper down here where it will generate
the prompt for you. So you can either start with a prompt where
you describe the song. You can also put in that
lyrics that I told you about, add some reference if you want. And if you're not sure how
to even write what you need, you can use this
section down here. We made these in the previous
lessons with sound effects, but they still use
the music model. I could further edit
them if I wanted to, but first let's do
something here. So I'm going to just
put it in here. I want a song that is suitable for a log where I
am exploring the forest. I want someone to sing
in the background. So let's send this
and there we go. I now have a prompt based on
what I you know, described. So it's a song for Travel
Flogs which is true, features finger picked
acoustic guitar. We got some birds in the back. And here's the lyrics. So we get some, you
know, openings, and then these are the verses
some chorus and an outro. If I'm happy with
this, I could let it go and press on
the generate button, but I could also come in
here and make changes. You can either type directly,
add something in here, or have it regenerate
the lyrics for you. And as I mentioned, you
could add a reference here. So this could be another
song that you really liked, but I'm just going to
let it be like this. And down here, we get to choose
the duration of the song. Because I have some
lyrics over here, I can't really do that
because it's predetermined. So this is a few
seconds for each part. I'm just going to
set it to Auto, which is however long
11 laps determines. Down here, you can choose
the number of variants. When you add more variants, it's going to take more credits. I'm just going to do one
for now and you can see that it just becomes 900/minute. We have some fine tuning as
well that you could add, so you can go for a different style
of singing if you want. We have male, female. You can preview them as well. Have male or female, and you can add those as well. I'm going to leave
mine as it is, nothing added because this
is a very low stake song, and I'm just going to
click on Generate. So once you generate it, it comes to the side, and we
can see them being built. As we speak, it's a
minute and 48 seconds. Let's hear what it sounds like. Steps on the mars. Softly, softly, green canopy
high above M I rivers. So that's our song.
If it's, you know, too upbeat, I could go back and maybe make
some adjustments. But overall, you can
see that it does fit the travel lock
style very well. Now, if you like
the song and you want to make further
adjustments, you can just go to Edit Song, and it's going to load
the project for you. So the name of my song
is lost in the wind. You can see the
history of the chat. On the right side, we have this little timeline where I get to play around
with the verses. We have a play button. We can work with the different lyrics
and the different tags, add more stuff if we want, maybe we want to add some maybe two more
minutes of this song. So you just put it in this chat right here, make it longer, and maybe add a second singer in there, whatever
you want to do. Since there is a chatbot here, you don't need to do
a perfect prompt. As we saw, initially, ElevenLabs just took my idea and turned it into
a decent prompt. Once you have your song
and you're happy with it, you can download it
on the right side, or you could publish it
on the platform itself. So this is where people can use your song and then
you earn something. I think the rates differ depending on how
many you've made, how many songs you've made, and how many people
download your song. But it's a pretty cool option. You can just play
around with these, make some decent music, and have it available for
other people to download. We can also share this, share it as a video,
give feedback. And these are, like, different segments of the song, so we could cut it short or make it longer,
whatever works. Since I just made an adjustment, we would have to basically
update this song. So it's pretty straightforward. I'm just going to do a
slight change to the song. Say, I don't want it to
be, you know, this chill, I want to do something
with fine tuning, maybe like a rock kind of style. I could do that right here. So let me just undo
this extension, and we're going to
add a fine tune. Let's go for the country. Since rock wouldn't
really fit this style, and I'm going to say introduce a female singer
after the course. So that's our version
one. We're now going with a version two, and I can see that there
is some suggestions here, so I'm just going to hit tab, and that will be added, as well. So let's put this in again, one variation so it
doesn't take too long, and I'm going to
send that in for 11 laps to update my song. What it needs to do is that
it needs to add more lyrics. So here's my female vocal solo, right after the chorus, as I asked, and here
is the section. So let's hear what
it sounds like. I'm going to skip to this part. So right now, we just
added these things, but I'll just, like, copy this part just
to show you that you can change the
lyrics yourself. So back to the start, let's make that adjustment.
And let's see what we have. To the s back to the star. And there we go. We
were able to make a song out of one prompt
or should I say an idea? We added another singer. We added our own lyrics, and we had this country
style music at the end. So go ahead and try this, make a bunch of songs, and it would be a good idea to publish them on
the marketplace. You never know when someone might just download
your song and you'll earn a few
credits on the way too. So if you do want
to explore this, just click on Get Started, get yourself a nice profile, and then your songs will
go onto the marketplace. This is a great feature since ElevenLabs also introduced
video and images. So you can now do
your voice overs, background music, the content, the visual content, captions, and basically anything else you need with
content generation.
20. Voice Design: The voice library offers
so many different voices, all that are very diverse. There's different languages,
different tones, styles. But sometimes it's just not the one that
you're looking for. But, luckily, ElevenLabs
allows you to create your own voice
using voice design. So right now we are
in D voices tab, but we can already see all the different voices in
the different languages. So right now, I
can easily go in, choose a tag that is related to what I need and
download that song, use it, change the speed, tone, and, you know, go about
making my content. We're going to explore
this feature right here, which is when you create
your own custom voice. And the cool thing about it is that once you have
your own voice, you can actually upload it to the marketplace where others can use your voice and you
make some earnings on the way. So let's go ahead and click on this and create our first voice. So with voice design, you can either design
something completely from scratch or you could provide
your own voice for a clone. This is the fastest way, but you do have the option
to clone your own voice, remix with an existing
voice or try to, you know, make a more
professional clone. So here, it actually needs 30 minutes of your
audio in order to make a very clean and
professional audio clone. But you don't really need
to do this if you want something quick
for your own use. So let's explore the
first one. Voice design. And we get this box where we
get to describe the voice. So basically, you can let me
extend this a little bit. Talk about the gender, the age, the pitch of the voice, what you need the voice for. Sometimes that helps. Like if you say you
want ASMR voice, it's going to make
it kind of whispery. So you can explore that as well. And down here, we can get some random prompts
to get started. So if I click on
this, we're going to get a high pitched female voice. Click again, an older woman
with a southern accent, and you can see how the prompts are differing in the length. It doesn't necessarily have
to follow one template. You can just put
in what you want and get a nice voice
to get started. If you click on settings, we can work around with the loudness of the voice
and the guidance scale. So the guidance
scale is how closely this new voice is going
to follow the prompt. You want something
to be basically very controlled based
on your prompt, you will increase that to high. But if you want 11
labs to be a little creative with how it
interprets your prompt, you can keep it on the low side. I would recommend keeping it somewhere in the middle
just to get started. And once you have that voice, you can edit it to either side. So instead of using one
of these templates, we're just going to
describe our prompt, so that way you guys
can follow along, and we can all make one
voice to get started. So what I'm going to do is try to make a voice
for a news anchor. And I'm going for a
female, maybe middle aged, confident voice, serious
voice suitable for the news. And I'm just going to
describe that right here. So a confident middle aged female voice for news
anchor purposes, we can say speaks very clearly. So here's my prompt.
Very straightforward. I just basically wrote
down what I need. Now, on the right side, we
can work with the loudness. We don't want her to be
too quiet or too loud, so we can go for maybe 60%. And as for the guidance scale, I do want it to bear
these key terms in mind. So I'm simply going to increase
the guidance scale to 60, as well, because I
don't want, like, a very excited news anchor. I want someone who's
confident senior and someone who can relate
the news very well. So a confident for a senior? Let's put that in. And we can also mention the
accent, if you want. So American accent,
English. This is optional. You can leave that out and then see what ElevenLabs
makes for you. So once we're done,
we can see that it will cost us 350 credit. We can generate the
voice and then make adjustments after it
has been generated. Good evening, everyone.
Tonight, we're bringing you the latest developments
on the economic outlook, following the Federal
Reserve's recent announcement. Our team has been tracking the
market's reaction closely. That's the first variant. As you can see,
it's really good, suitable for relaying news. Let's look at the
other two variants. Good evening, everyone.
Tonight, we're bringing you the
latest developments on the economic outlook. Good evening, everyone.
Tonight, we're bringing you the latest developments
on the economic outlook. So I will go with the first one. Good evening,
everyone. Just because it's a little louder and
a little more clear. Now, the clearing
throat part is actually an interesting addition because news anchors are often live, and that may happen, so I could even see
what that sounds like. And once I'm happy with it, I could select the voice. But if you felt like
it's, you know, too fast, too high pitched, you can make an adjustment here and then generate
the voice again. So I'm just going to
select voice one. Now we're going to
give it a name. So let's call her Maria, use Anchor, and then
we could label it. So language is English. This way, it could be used
for organizing your voices, but it could also be used when people try to
download your voice. So the more labels you have, the easier it is for them to find your voice and download it. So accent we can
put what we have gender female. And there we go. I have four labels that will help people find the
voice and for my own use. I could change this if I want, but I think the prompt
is a good indicator. So let's save this voice. And now it will be
added to my library. So with this new voice, we can either
generate this speech. So I'm going to
type something in. This will be the voice
that will be used. I could create an
agent with this voice, or I could have it,
narrate a story. So audiobooks is something that we're going to look at
in the later lessons, but this is how you
get to use your voice. These are just a few examples. Of course, there's studio
and other stuff down here, even dubbing now
when I go to voices, my voices, Maria is right here. So I could use this
for text-to-speech, have it narrate a book for me, change the voice if I want to and use it for
different purposes. So what we can do
is just click on use voice and do a little
script right here. Let me just choose Maria first. Going to do V three for
a very natural style. I'm going to grab a
sample news segment just so she could use the right wording and the right style, and we're just going to
paste this right here. So I'm just going
to paste that in. We have something very simple. I will not use audio tags
just because the voice that we generated should kind of follow that news
anchor style anyway. But if you want, you can put in different audio
tags like laughter, coughing, you know,
whatever you want. But now we're just going
to generate that speech. If you want, you
can further enhance the paragraph with audio tags, but I'm just going to see what this looks like on its own. Last Friday, students from Sunshine Elementary
School planted Oh, so we do have the
cafe audio effect from the previous lesson. Let's cut that out
and regenerate that. Last Friday, students from Sunshine Elementary
School planted 20 young trees in
the school garden to keep the environment
clean, green and cool. Last Friday, students from Sunshine Elementary
School planted 20 young trees in
the school garden. So as you can see,
that's really good for my application.
She's confident. She's very clear
with what she says, and you can easily
use the voices that you generate using this section. Other things that you could
do is clone your own voice, and that's going to be something that we'll look at
in the next lesson. But now let's take a
look at voice remixing. But let's go into this guy, and we could first add
a voice reference. So I'm just going to
add Maria right here, and we're going to transform
her into a different style. So let's say I want
to make her British, and let's actually change
the gender as well, so two things at the same time. So turn this voice into
middle aged British male. Here we can look on
the prompt strength. So how much the prompt
alters the original voice. I will keep it medium because that's exactly
what I'm doing, the accent and the gender. Could get some inspiration. There's also a guide if
you're not sure what to do, and you could even
give it a script so you can see the
full application. So I'm going to paste the same news segment that
we saw earlier. I go to save that right here, and we're going to send this in. Last Friday, students from Sunshine Elementary
School planted 20 young trees in
the school garden to keep the environment
clean, green, and cool. Last Friday, students from Sunshine Elementary
School planted 20 young trees Last Friday, students from Sunshine
Elementary School planted 20 young trees in the school garden to keep
the environment clean. So, these were three samples. I do feel like he seemed
a little younger. I was asking for a
middle aged British man. So let's say generate more, and I'm gonna increase
the prompt strength. So let's actually do it here. Generate you can also
use these as reference, combine it with Maria, but I'm first gonna work with
the age of the new voice. Sunshine Elementary
School planted 20 young trees in
the school garden. Last Friday, students from
Sunshine Elementary School. Last Friday, students from Sunshine Elementary
School planted 20 young trees in the school garden to so
this is pretty good. I can now save this as a
voice, the same thing. You just put in a name, give it some labels, and a description. Then this will be added
into your voice library. And this way, you could have different variations
of your first voice, and it's really good for
different sorts of content. So say I have one news piece, I want to try a
female anchor and a male anchor and then
finally make my decision. So this is a pretty easy
thing that you could do. You could make another
remix right here, give it a different script, maybe choose one of the main voices and
see what comes out. Go ahead and give
this a try and see what sort of voices
you can generate.
21. Voice Clone: The next thing we're going
to do is clone our voice. So just as I'm
speaking into this mic right now and all of these
voices are available, we can actually have our voice, the same voice you're hearing
be one of these options. So when I do text-to-speech, I can have it read the text. As myself. Right away in voices, there's two ways to access this. Number one is hitting
the plus right here and then
instant voice clone, or you could go over
to this blue thing. If you click on it, it's
the exact same thing. It doesn't matter
which way you go. And we already
looked at this one, now we're going to
go down to this. This is a really simple way
to create your own voice. If piece of audio that you don't want to record
the whole thing for, you may consider
cloning your voice. Now, I will say that
this is not exactly the perfect replacement for your own voice recording
because as we know with AI, it's inconsistent and if you plan on doing
this professionally, you may want to find
a better alternative. But nonetheless, it's
a cool thing to do. And all it asks from you
is a ten second audio. You need to speak for ten
second in a microphone, make sure there's no noise around and your voice
is crystal clear. Say everything in a slow and clear voice,
don't rush anything. And that way the 11 laps can study your voice and
then replicate it. All you have to do
is hit record audio. Choose the mic and it will do a three second countdown and
you get to start reading. I will say a really
random sentence. It doesn't have to be
the full 10 seconds, that's just the ideal amount. I will start talking
about the weather. Lately, it's been really hot and I have planned not to
go outside as often. I'd rather sit home
and relax all day. So 10 seconds of audio. I did 11, but it's fine.
Here is your recording. If you'd like to do another one, you can just hit
delete and replace it. Here you can listen to
what you sound like. Lately, it's been really hot and I have planned not to
go outside as often. I'd rather sit home
and relax all day. So there we have it.
There is an option for ElevenLabs to remove
the background noise. I'm going to disable that
because I didn't have any, but if you feel like your
audio does, just turn that on. Now it's ready and I just
have to click on next. Now I get to put up some
information regarding my voice. You can preview the
voice right here. Airplanes soar through the sky as autumn leaves drift down. That's how 11 labs interpreted my voice.
Let's give it a name. Call it My voice.
Then once again, we have to add labels. I'm speaking English. And then we can give
it a description. My casual voice. Next, you just have to confirm
that this is your voice. We talked about this
in the ethics lesson. Do not replicate
other people's voice, especially not without
their written consent. But this is my voice,
so I have all rights over it and I'm just
going to click it, save the voice. And there we go. Now we have the option
to try out our voice. We have generate speech
with text-to-speech, speak with your own clone, where we can create a
conversational AI agent, and then we have that dubbing
feature using studio. I'm just going to
hit Skip for now because we already know
where the tools are. There is my voice. You can
go to my voices as well, and it should be the
first thing, zero days. I could begin using
the voice by clicking that icon and start
typing something. The model it recommends, you can just use that and then I'll leave
everything as default. Let's turn this off and
generate the voice. What a wonderful day. We can make it slower,
less unstable. And add some exaggeration. What a wonderful day. That's my voice right there. It doesn't I said, sound exactly like my voice as
I'm speaking right now. It's just a fun
thing you can do. I wouldn't see you guys using this professionally
in any way, but I believe the more they
work with their models, the better this feature
is going to be. That's how you can
clone your voice. It's really straightforward. If you wanted to change or delete your voice, you
just go to voices, my voices, and then
delete. And there we go. Now, we don't have
any more clones. Now that we know how to clone
and design our own audios, let's go ahead and start
a project from start to finish where we take a web page and turn
it into a podcast. I will see you guys there.
22. Scary Podcast Part 1: Now we're going to start
our very first project. We know how to use the tools. It's now time to put it
into one massive audio. With this first project,
we're going to do a scary story podcast
where we take a web page, turn it into a podcast and then have three people
in that podcast. We have a host and two guests. What we're first
going to do is find that web page that has that
story we're looking for. I'm just going to head over to the Internet and just
go to Wikipedia. Usually they have really
long and detailed pages regarding a certain incident, a certain object or whatever. Going to look for a natural disaster one because those can be really well converted
into different emotions, fear and all that stuff. Let's go to Wikipedia. You can use any other website, just make sure that you
have the rights to use it. I will look for a tornado. So right away, we
have some examples. This one in 2007. We can look for other
ones, even this one. I think I will go with
this extremely powerful. I'm just going to
make sure there is a lot of
information about it. We have comparison to other
tornadoes that's good, the aftermath, significance,
the damage, so it's perfect. I'm just going to copy this URL, and we're going to go over to Go Studio and import from URL. Let's paste our scary
podcast and choose a natural conversational voice since we are going
for a podcast. Lovell, trust a few,
do wrong to none. We make our own fortunes
and we call them fate. Nicht Due to daraufin as. Hi there, and welcome.
Grab a coffee, settle in, and relax. And
let's chat about this. Gratitude is riches complaint. A single rose can be my garden, a single friend my world. We make our own fortunes
and we level, trust a few. I think I'm liking
Charlie for now, so I'm just going to
click on his name. Right now, we don't even have
multiple voices in the URL, so this will not
be useful for us. I'm just going to have Charlie be the voice for
this podcast first. Once we are happy with
the way he sounds and the way the podcast
is extracted, you can see right
away, it's 17 minutes. If you have a longer text,
it's going to be longer. Over here, it's
trying to convert this table into something
conversational. I'm going to have
Chad GBT just take out the numbers and leave
out the details for this. As you can see, it removed
all of these numbers and only kept something
that's more natural. It kept the dates, the number of the miles
and all that stuff. You can do this and refine
further if you don't want any numbers in this section or you want to take
out certain words. I'm just going to copy
this and put it in 11 labs right around here. Here it talks about this tornado and then talks
about the other four. Okay. You can see
that it also has the titles from the
Wikipedia page, but we do want to make it sound
a little bit more normal. Let's add in a few words. The reason why
we're changing from normal text to a heading
is because it's going to emphasize it and read it in a way where we're moving
to a new section. There's going to be
pauses, and normally, it's going to be a little slower than how it reads the
other normal text. Now, just like as
we learned before, you could add the pauses
and the sound effects, but for now just focus
on finalizing your text. If you want, you
could go back to Chat JBT and ask
for more inputs, maybe different reactions
to what Charlie is saying. For now, I'm just
going to finalize what Charlie says and it generate. While that's generating,
let's give this a name. I command A and then
generate one more time. Let's talk about the 2007 El
tornado and its aftermath. The 2007 El tornado was a small but extremely
powerful and erratic tornado that occurred in Canada during the evening
hours of Friday, June 22, 2007, powerful Five
tornado that struck the town of Eli in the
Canadian province of Manitoba, 40 kilometres west of Winnipeg, was known for its unusual path, how it was during its path, its rope to cone structure, as opposed to a wedge structure,
and how it is unique. But what were the
meteorological synopsis? Okay. Notice how it did pause between the normal
text and the headings. I did heading two, but you
can do heading one as well. Any of these headings do signify that the voice has to pause and emphasize moving to a different segment
of the podcast. Let's start from here
since this is where we took it from Taji BT and
see how normal it sounds. Were there more?
Yes, there were. In addition to the
Eli F five tornado, four more tornadoes
also affected Canada on June 22 to 23rd. After the historic F five
tornado hit Eli, Manitoba, another powerful
tornado touched down just 10 miles to the west
near Oakville on June 22. Rated F three, this tornado
damaged trees, outbuildings, and a couple of grain
storage bins along now that the base information
is up to my liking, I can now take this back to hat TPT to make it a
little bit more human. It's going to add in ms, pauses, maybe some stuttering, anything that would make this
like an actual podcast. Because right now it only
took the information from this webpage and
converted it into voice. But since we are
going for a podcast, we want to start
with a strong base, and once we have Charlie's
parts put in properly, we can add the other two guests. Let's take all of this, then go back to our chat and
first describe the scenario, then paste a text. Okay, I'm going to let this
run and as you can see, it's assuming that you want
this to be the script, but we don't really need
that capability right now. When it's done building this, I'm going to have remove these segments so that it's
only the text on its own. But it doesn't look
like it's going to generate more. Let's see. But right away, you can
see that it's adding the words, the three dots, dashes, and it does
emphasize on certain words, just like a regular podcast. It's done with the conversion,
and as you can see, didn't add anymore
of these segments, so we don't need to
regenerate anything. Let's copy this body text. And replace it with
what we have here. Now we're going to use the different headings
on the places where we want Charlie to pause. This one's a beginning
definitely like that. Here's another segment. Okay, so we have two segments here not counting the beginning. Let's go ahead and
hit Command or Control A, and hit regenerate. Okay, so let's talk about the 2007 Eli tornado and
what happened after? This wasn't just any tornado. I mean, this was Canada's first ever
official F five tornado. And yeah, it might
have been small, but it was
ridiculously powerful. And, honestly, just
straight up weird. It hit Ely Manitoba, a quiet little town about
40 kilometres west of Winnipeg on the evening
of June 22, 2007. Now, here's the creepy part. This tornado didn't
follow the usual rules. It had this strange
roping shape, not the massive wedge you'd expect from
something that strong, and it moved slowly, like weirdly slow. It twisted. Let's move on to
this last sentence. The total damage estimated at about $39 million back then, which would be over
$56 million today. So what made that
day so volatile? Well, meteorologically speaking, the setup was like
textbook chaos. A low pressure system moved
in from Saskatchewan, and a warm front parked
itself north of Eli. And yeah, things exploded. Now, here's what
happened on the ground. The tornado touched
down just north of the really important to
use those heading words, and that's the perfect
opportunity for you to add some music in there with a
pause or some sound effects. But other than that,
you can see how using HGBT was able to add
that human touch to it. It does pause at certain times, but feel free to move these words to other places or just remove them altogether. But given that this is
a podcast and we're trying to make this sound like Charlie is just
your average guy, this works perfectly
for our case. So now that we have
Charlie's part all done, we're going to move to
adding our two other guests. So that's something
we would have to do with a different tool. And that will be for
the next lesson.
23. Scary Podcast Part 2: So now we have something in our studio which we call
depict our Native podcast. Previously we took a
Wikipedia page and we converted it into a
podcast using JAGPT. Now, you could fully
rely on just 11 labs for this for turning that
Wikipedia page into a podcast. However, for my case, I did want them
to have our host, which is Charlie Charlie
to have more pauses, some ms, some s in there just to make it
sound more human. But you're totally
free to not even use JAGPT and only rely on 11 laps. So now we're going to move
from here to create a podcast. Immediately, you can see that we can use existing project. Now, you did also
have the option to put your Wikipedia page here. However, I wanted to start
from scratch so you can see the progression and how you can put in your own
adjustments as we go. Let's go to use Existing Project and choose our tornado podcast. We want to keep it
conversational. Let's keep it a little longer
because it was 17 minutes. Our host is still Chris, but we do have a
guest right now. So let's decide
who our guest is. She tore her gaze away
from her ruined footwear, still very much
grieving the loss. Let's search for conversational
There's a Nadal on. If you spend your whole life. So I will go with Jessica. You can see she has
that conversational tag on it, so it's perfect. And then the language
is in English. If you want to have
some sort of focus, you can put in the words. So let's put the word
tornado. Let me see. Let's to highlight the dangers of a tornado pause
on the numbers, I guess, damage numbers. I just gave ElevenLabs
two directions to follow. You can see start
with some sort order, highlight, pause, and then
tell it what to do exactly. You can add as many of
these as you want as we saw earlier or delete the
ones that you no longer need. Save the changes and
let it generate. I'm going to let this run
since it's a little longer. You can see that it generated
another project here. So once it's done, I'll be back so we can hear
what it sounds. Okay, our podcast is ready. It's 10 minutes, 24 seconds. It did cut down from
the 17 minutes, and that's mostly because it
shortened the paragraphs. As you can see,
there aren't any of those bulky paragraphs
that we had before. That just shows how amazing this podcast conversion is within 11 laps. And
look over here. I even gave us a cool title, the unexpected F five, a tornado surprising power, even though we had it named the tornado podcast
or something. When I hover over them, I can see who's
saying what green is Chris and blue is Jessica. So without changing anything, let's hear what these
two have to say. Tornados are often measured
in miles of destruction. But in 2007, Canada's
most powerful twister ever recorded was barely
wider than a one lane road, and it rewrote everything we thought we knew
about these storms. Well, that's
fascinating. How does something so narrow
pack that much power? You know what's really
mind bending about this? This F five tornado
that hit Eli, Manitoba was just 35 yards
across at its strongest point, but it literally picked up entire houses and spun them
through the air like toys. Hold on. F five. That's the highest
rating possible, right? So that was a little sneak peek. I'm not going to have you
listen to the entire thing. But if you wanted
to skip through, remember that there's
a timeline down here that you could just
put your plate head over and then skip to maybe Chris or Jessica
whenever you wanted to. Here, everything is
in the regular text. You can see the
attribute right here, but we could change
the way they read it by changing some things to
headings. So let's see. So here's our transition
that we had initially. You can see it kind
of changed it. Initially, it was like, but
what about the damages? Now they turned it into, so what kind of
damage are we talking about here with that kind
of focused destruction? So let's change this to heading one and then look
for another one. This is about the damage. All right, I'm just going
to keep these two as our key transitions and then
simply regenerate this part. Well, you're not
wrong. At one point, it just parked itself over the town's flour mill for
four straight minutes. Just imagine watching a
tornado hover in place, methodically destroying
everything beneath it. So, what kind of damage
are we talking about here? With that kind of
focused destruction? Let me paint you a picture, and this is where the numbers
get truly staggering. The total damage was estimated
at $39 million in $2,007, which would be over
$56 million today. But here's the most
incredible part. Not a single person died. Wait, what? How is that even
possible with an F five? So we did pause with
the heading just as it did with the
first tool we used. And I just forgot
to mention that we did ask ElevenLabs to
highlight the damage. So you can see how it
has multiple sentences. We had one here. There's
another one here. Damage. Let's see. More damage. It's reiterating the damage part of the podcast numerous times. You could go back and
change that into something else and really make
this podcast your own. I do have all of these words. What I do want to do is maybe add even more pauses between these two headings just so that it's really a
full transition. Let's add a pause
right here, 1 second, and then another one
or one more second. And then another thing that
I want to do is do an intro for Chris and Jessica and then
bring in our third person. We mentioned having a
host and two guest. So let's hit Inter right
here. Hello, everyone. I'm your host. Chris. And this is I'll
just use the title here. Okay. So let's mention Jessica and then
find another voice. Waves crashed. I think Alexis is a good one, so let's just say Jessica and Alexis and then have each of
them introduce themselves. All right. So I have
four new sentences or lines, should I say, and this just makes
it a lot more natural considering we do have three people
in this podcast. Now all you have to
do is assign each of these sentences to their lines. Alexis is this line actually. When I go over it, I could
see that it's not Alexis, so I'm just going to
click once on this. Our voices are right over here. We accidentally brought in
Archer, but that's fine. Let's turn this into Chris. Apply this one to Jessica. That one's already Alexis, and this is back to
Chris. The rest is Chris. Here's Jessica reacting. Let's turn this to Alexis. Just go back and forth until we have a
conversation going on. Seeing how Chris is the host, all of these lines where there's a lot of information
about the tornadoes, I'm assigning that to Chris. Jessica and Alexis are just reacting to what he says
because they are the guest. Let's two of them,
Alexis this time and not have a routine transition. Okay, so we have a good back and forth
between our subjects, and now we are ready to
add some sound effect. When he is introducing
the podcast, we do want an audio
to play before it, and maybe at the headings, put some sort of sound effect, maybe a commercial
break in the middle. Something that will make
this work like a podcast. We're going to do that
in the next lesson. Just make sure that by
the end of this lesson, you are happy with the words that each of the guests have. You could even add more
people to this podcast. I'm just going to
keep it at three. But go over each of the words, make sure that they do belong
to that person you want. And you don't have a random
person in your podcast. When I removed Archer, you can see that he's no longer here and I only
have three people. If you see another voice here, it means that somewhere
in your podcast, you did assign that
voice to a certain text. Down here, again, you can
see a colorful transition. If you saw that there's way
too many greens back to back, then that's assigned to switch it out with
a different voice. Let's move on to the
next lesson where we add our audio and just finish
up this scary podcast.
24. Scary Podcast Part 3: I let's finalize
our scary podcast and add some sound effects. Previously, we took
a Wikipedia page, converted it into a
conversational podcast, then brought it in here where
we added two other voices. All I did right now was put the pauses after the headings, but you can add more
pauses wherever you want. Even between the sentences,
you're free to do that. Now, what I'm going to do is make an intro for this podcast. So right before hello everyone, I'm going to put in
the sound effect. Let's do a tornado Okay. I like the first one more
duration wise, let's just do a shorter one. Hello everyone. Hit apply. Then after everyone, I'm
going to add a pause, it's the intro essentially. Then over here, we can add, let's see, either a drum roll or a next page sound effect. Let's just do that.
To page flip. Okay. And in terms
of the duration, once again, I'm going
to do a little bit of it and not the
entire sentence. I like number three
more and just be sure to keep this on
the low side since we don't want it to go over our Cris basically. All right. And then what we're
going to do is just adjust the voices a little bit. Jessica's is a little
bit too exaggerated, so I'm going to
lower the stability. Sorry, increase
the stability and just make her tone
a little slower. It's safe. Alexis was fine, so there's no need to
do anything with her. Chris, I do want his voice to
be a little slower as well. S save and I'm going to find
other places to add a pause. Okay, so we added
our sound effect. What I'm going to
do now is just have a small section of this generated so you guys don't have to listen
to the entire thing. Let's do this one part where we had the sound effect, the pause, and of course, we can see how the voices sound like after
we made our adjustments. Hello, everyone. I'm
your host Chris, and this is the
unexpected F five, the podcast that reports on the world's most dangerous
F five tornadoes. I'm here with our two
guests, Jessica and Alexis. Hey, guys, I'm Jessica, and I'm so happy to be
here, and I'm Alexis. So excited to talk to you guys
about these spooky events. Alright, let's get
into it, then, guys. Starting with what we know about the tornadoes.
Tornadoes are often. Alright, then that
sounds a lot better. I do feel like I want
to change Alexis. It seems a little
bit too robotic. So what I will do
is go to this part or part and find another
voice, basically. Let's look for
conversational and find a female voice that sounds, you know, a lot more normal. In the ocean of competition,
be a swift fish. The Ocean heights treasure. He up yours. Just
trust yourself. There is no greater. I think Area sounds good here, so let's switch Alexis out
for Area. Then let's see. All we have to do is just grab this part and click
on Apply with Area. And then repeat for
the rest of the thing. I'm going to try
another generation for you guys since we
added a page flip here. It starts with Chris. Let's
apply Area over here. Once again, I do want a
slow and steady voice. Let's do that, it safe, and then do it here. I think I'll change this to Jessica so there's
some variety. All right, let's grab
this whole part, generate, and listen
to another segment. What kind of damage are
we talking about here? With that kind of
focused destruction, let me paint you a picture. And this is where the numbers
get truly staggering. The total damage was estimated
at $39 million in $2,007, which would be over
$56 million today. But here's the most
incredible part. Not a single person died. Wait. What? How is that even
possible with an F five? You know, that's the
real story here. It's about preparation meeting. But before we get into
that, let me break down exactly what
made this day so perfect for tornado
formation because it's fascinating from a
meteorological standpoint. Hmm. I'm guessing
multiple factors had to come together just right. Oh, you better believe it. We had this low pressure
system moving in from Saskatchewan temperatures
in the high 20s Celsius. That's around 80 Fahrenheit. And humidity levels that would make a rainforest feel dry. Okay, and that is
sounding a lot better. The page flip I do feel like has to be
moved a little bit, so I could just grab
it down here and move it till the end, maybe here. After she says the word
destruction, page flip. Then I would just continue altering the rest of my podcast, being sure to add the
pauses where it needs it, and then adjusting my voices further if they're not
reading something right. But for my case, they
read everything fine. I didn't need to do
much adjustments over my scary podcast. And now we're basically done. We took a Wikipedia page and
turned it into a podcast. What I could do is just
export it as a single file, p three, and then just
listen to it when I want to. This is also a nice way to take articles that are way
too long and then just turn them into a podcast
so you can listen to it when you're walking around
or you're in transport. That concludes our
first project. The next project, we're
going to try to do a different type of voice generation
using the same tools, but application wise,
it'll be for something other than a podcast. I
will see you guys there.
25. Audiobook Part 1: Now we're going to move on
to a different project, which will be to make an
audio book from scratch. What we're first
going to do is get that book that we're trying
to turn into an audiobook. I'm starting here in ChachPT, but this would be
assuming if you already have a
book that you have the right to turn into an audiobook or that you
have your own stories. Just for practice sake and for you guys to have some
hands on experience, we're going to use this tool to make a book out of certain
words and prompts. Open up Chat GPT and first decide what your book
is going to be about. You have to know the subject, what type of book it is, how long it is, and who
is like the characters. One thing that you could do
if you don't know where to start is have Chat JBT give you a few ideas as to what you should turn into
a book in the beginning. For example, if you have a YouTube channel where you
talk a lot about real estate, you can come to Chat GPT and
give a prompt like this. I enter, and then we're going
to get some suggestions. From there, we get
to continue on. What's interesting
is that Chat Tu But will give you the
title and then it tells us why it works and some
details inside that book. We got what they don't teach you in school
from rent to Riches, civilized, real estate Diary, and all of that stuff. What you could also do is
combine these ideas into one big book or maybe just take out a few of these bullet points and replace them
with something else. Let's see. I think I like the House
hunter survival guide. That name is pretty good to me. I will say I like number one. However, remove the quiz part. And replace it with let's
see what we can replace it with the quote real
estate jargon. Here I just did a little switch and now it's giving
me more details. Who are the audience? For some buyers, renters,
real estate browsers. If you already
have your channel, just double check that this
is indeed your audience. If it's not, you
can switch it out. For example, I will
say my audience include middle aged
to young adults. Age range, I can put 50, maybe 25 to 40, let's see. Then this is just a
random age range. Those who live in California, Los Angeles, California, those
who have no kids and dogs. I'm getting more specific
regarding this book. This is our audience
and let's see. Maybe we can add one more thing. And work in the
fashion industry. Here's a cool snippet that
it does, the tone and vibe. Think a stylish friend
who already bought a place and is walking you
through it over Macha. Sharp but casual language, real talk about money, references to fashion design
and LA specific quirks. Then this is the chapter. There are ten chapters,
which is good. And now that it looks like this is going
to be a good book, I'm going to specify
from what point of view this is written from or if there's no
characters at all. Then lastly, how
long is this book? Let's say, in terms of who whose point of view
this book is from, I think this person, sharp a casual language is good. But we can make the
narrator use this one. But make them a well
informed real estate agent who knows a lot about
this particular audience. Then give me only
three chapters, totaling to 50 pages. Make this a quick and fun
read for my audience. So this will be a
whole different page and here Armour chapters. As you can see, it's giving me bullet points and you
could keep that idea. You can see it's giving us a snippet, summary or something. What I want to do is convert these bullet points
into paragraphs. Let's close this real quick, come back here and ask for a
full version. This is good. Give me the full
version in paragraphs. Then I'll cut it down even more. So I'm going to let this
run and then I'll be back to show you what our
script looks like. Then if we need to
do any changes, we get to do that before
we move to ElevenLabs. Okay. So here is
my book right now. It's not that long,
which is perfect. But as you can see, it went from bullet points to
actual paragraphs. Now, keep in mind
that you do want this to be a audiobook, so make sure yours
has the indications for chapters and anything
else that you want to put in. You could have bullet points, but I would recommend
putting it in between paragraphs so that it's a
more natural audio book. And that is the
name of our book. In the next lesson,
we're going to put this into ElevenLabs and start matching that audio to fit perfectly with this tone and for our audience who are
people that are looking for a nobs guy to
buy their first. Let's move on to
the next lesson.
26. Audiobook Part 2: Alright. Now that
we have our book, I'm simply going to copy
everything from here. I think this entire
thing is the book, so you can just hit
copy right here. Go to studio and
create an audiobook. So uploading a
document is optional. We can also copy paste. So let's not do anything here, but go over which
voice we want to use. So let's look for narratives
and see what we have. She tore her gaze away from her ruined God has
given you one fact. Let my voice become your voice. The people who I have
crossed. Let my voice be. So, like, I guess, for me,
the main thing is, like, I don't understand why
people have to get upset. It's like Theresa
She tore her gait. Net. I'm Sophie Lang. Hi there and welcome.
Grab a coffee, settle in. I will go with this guy
that we didn't use in the last project just
to give him a chance, and then we can hit Create. So we're brought onto this page that you could also go
in with the studio, but we're simply
going to paste what we got from Chat
GPT in this box. Then notice here we
have plus chapter. We have three chapters. So we're going to hit
Plus chapter here and copy the Chapter
two and beyond. Delete it from this one, go
to Chapter two, paste it, and then copy Chapter three over to another
chapter. There we go. We have Chapter one, two, three, and I'm going to
match the names over here with what we
have over here. Go to the three
little dots, rename, paste it with Command
or Control V, click away, and
repeat this process. We got our chapters all done. I'm going to remove some of the things that
I don't need this guy, instead I'll call it
preface. All right. This is the title of our book. I'm going to make that
heading one so that it pauses and registers this
as a separate thing. In terms of our voice, let's lower the speed and keep it more stable
and more professional. I'll lower the
audio a little bit. It's safe, and let's generate. Let's start by generating a small bits and then
make our changes. I'll grab all of these
guys, then it generate. The House Hunters survival
guide, LA edition. A OBS guide to buying
your first home in Los Angeles without losing
your cool, dog or sanity. Preface. Hey, I'm not your mum, your financial advisor
or that guy on YouTube who thinks every
home should have a panic. Okay, so right away, I noticed that I want this
line to come right after this. So let's just delete
it and bring it up. Colin, put it here. And I think I'll keep A and change this to
maybe disclaimer. Then since it's a
male, let's do Tad. Then what I'm going to do
is do one more change, increase the style exaggeration, and the audio is
a little too low. Let's bring the playhead
back and redo this part. The House Hunters survival
guide, LA edition, ANBS guide to buying
your first home in Los Angeles without losing
your cool dog or sanity. Disclaimer. Hey,
I'm not your dad, your financial advisor
or that guy on YouTube, who thinks every home
should have a panic room. I'm a real estate agent who's
helped people just like you stylish, dog
owning creatives, juggling a dream and a deadline, buy their first place in LA without having
a total meltdown. I'm here to make
this whole process feel way less overwhelming, a little more empowering. And honestly, kind of fun. Right, so that's a lot better. And now we can move on to how it reads the
different chapters. So I'm gonna continue
pressing play, so we can see how that sounds. Chapter one, welcome to the
LA circus AKA the Market. So for the title, there should be a break here. That's what I'm going to do. A five second break,
half a second. And then here, I'll
do the same thing. Okay, let's hear
it one more time. Chapter one, welcome to the LA circus AKA the market. Let's get
something straight. Buying a home in Los Angeles is not for
the faint of heart. This is the city where
a two bedroom craftsman can cost more than your
entire college education, and you'll still need to beg
someone for street parking. Competition is wild. Cash buyers are everywhere, and the market changes faster
than your favorite pop up coffee shop closing
for renovations. Right and that's pretty good. So what you're going to do is continue building your
chapters if you'd want to. But honestly, because ElevenLabs has a built in audio tool, there isn't much
for you to do here. So you just declare
your chapters, add a pause whenever you want
to and adjust your voice. Now for the title in
the chapter titles, I'm going to do the same
thing where we added a pause. Let me just go over here. This is heading
three. Keep things consistent so that your
voice can read it the same. Let's add another pause here. I will do maybe 0.4 seconds,
and then the last one. One more time, heading
three, another pause. Now that I have all of my
pauses and all my text, I can go ahead and
export this and perhaps in a different
software pair it with a music if you want to or add sound effects as we
already learned how to do. However, I'm going to
keep it plain because that's just how I
prefer my audio books, but you're free to do
whatever you want. Because you know how to use each and every one
of these tools. That's how we can
make an audiobook. You started from HAGPT
finding book ideas. We tailored those ideas into a very specific
niche of an audience. Once we had that, we
got some bullet points. We turned those points
into paragraphs, limited that to 25 pages, and then we brought it
into 11 labs where we declare the chapters and
started the translation. Now Archer is reading our
book for the audience. I hope you guys enjoy
this short project. Let's move on to
the next lesson.
27. ElevenReader: Once you have a story file, you can easily turn it into an audiobook using ElevenLabs. We saw how to do that, but there has been an update, which is the ElevenReader. So basically, you can
upload your eBook, and it will start narrating it for you and you can earn
from that narration. So previously this
is where we went. We would upload our PDF file or EPAP to the narration
style chooser model. But now when you go down here, you can upload the same story and just have it go
to this platform. So I have a sample story in dot TxDFmat with
two characters. Going to upload that right here. I just uploaded that
from my computer, and we're just going to click on create and preview book.
So here's my story. It's, you know,
written line by line, and we're getting our
different characters. The first one is, I could, give it a name since we
have two characters, we can either have it
be multiple characters, narrators, or just one. So that's up to you. But if
you click on either one, you can assign who is speaking. Here I have some tags
about the characters. I don't really need this
to be set anywhere. So we're just going
to delete that using the backspace key. So you can easily make
adjustments in this page. And the same thing here. That's the name of our story. Scroll down. Everything
else looks good. Now, what we could do after that is play around
with the volume. This is pretty similar
to the studio layout, but we're just getting
that eBook capability. So if you want, you can add
in some music from here, make your own or drag one in. You can make
different characters. So this is the current
narrator for a story. The lantern at the
edge of the sea. Mara found the lantern
on a rainy Tuesday. So this right here is
not something I want. I want something
a little deeper. So I'm just gonna go here
and change the voice. So let's go for Explore. We will go for English, see if there's an
audiobook option. I want to get a preview first. Ne IguodaTokEviacos,
OEtoato Meg ir. Let's just make
sure it's English. The winds of the lands between
whisper of your deeds. Do not be afraid to ask largely. So what I'm gonna do
is click on Marcus. That's the voice that I want,
and it's going to basically reassign the narration
from Eric to Marcus. So there we go.
Everything is now blue, and we could do a
little preview here. Mara found the lantern
on a rainy Tuesday, half buried beneath
the wooden steps of the Old Harbor Lighthouse. She had come there
looking for shelter from the storm, not treasure. But the small brass lamp
caught her attention. So once I'm happy
with the story, the speech, the type of voice, I can now click on Export. Can export it as we saw either as an audio or a
different file structure, but I want to point
to ElevenReader now. So the cool thing about this is that you can earn
from this feature. You get 0.25/hour, that is streamed and 60 per direct sale. Now, this is a sample story, but we're going to
try it out anyway. You can do dynamic narration, so they could kind of
people that download your story can choose
their own voices, and this way, there's more
flexibility on their end, but you can also just
publish what you have and not give
them that option. So I'm just going to do
original audio for this. We can click on
Publish. All right. So once we come
here, we can give some more information
about the storybook. We can create our author
profile that's us. Upload your picture,
give yourself a name, the language, of your profile, and if you have a portfolio, you can put the URL
as well as a bio. I'm just going to do some
random stuff right now. Say my name is Megan. You don't have to
fill out the rest. Going to the next one, we can talk about the
book itself, the story. That's going to be the metadata. About the category, the
genre, and all of that. And on the right side, we can see the story being uploaded. So while this is loading, I'll just go here
and you can share this QR code with people or
download the app yourself. But going back, we can now
put in a name, a cover image, subtitles, our own author profile description of the book, and it's pretty much like
any other eBook platform. The only difference
is that you made the voice completely
with ElevenLabs. This is another output that you could do from
your creations here. I'm not going to go
through the full thing because it's going to
be like payout details, distribution details, and this, like I said, is a sample story. But if you are
interested, you can take your audio books a step further and upload
them on ElevenReader. There's also other options if you want to publish to Spotify. There's in Audio, the
video and Audio Native. So these are pretty cool things, and you can explore
them further.
28. Advertisement Part 1: For our last project, we're going to include some
things that are outside of ElevenLabs and see how we can combine them into one video. What we're going to do is get some footage
from the Internet. This could be made
with another AI tool such as runway or Mid Journey, or you could use copyright
free videos from Pexels, combine them, and then
make a dubbed footage, adubb audio for your footage. First, let's go ahead and decide on what the
video is about. I'm going to head over to
Pexels video. Right over here. And while we're here, let's
get an audio tab as well. I'm holding down Command to open a new tab just so we can have some music
in there as well. So I meant Pixabe for audio
and then pixels for video. So I think what I will do is an advertisement where first
we make it in English. We have some nature
shots maybe and then have an upbeat
music in the background. Once we have our English audio, we will translate that into
a different language so that it's like a
multilingual commercial. Let's get some nature
shots or maybe tourism. This will be a touristy ad, and there's a ton of videos
here for free for you to use. Just be sure that if you plan
on publishing it anywhere, you give credit to the creator. All the people who
made these videos are at the bottom left. Let's download a few of them. I'm going to go in and
get a less heavy file. There's our first one.
This is our second one. Let's do an island
shot and try to get different angles because that's what commercials usually do. Some birds over here, Okay. I got a bunch of videos. I'm going to head over to Chat GPT to get a script real quick. Feel free to do your own script. I'm just doing this for the
purposes of this course. Let's say I want to
make a hoops 60, let's say 30/32 ad
about traveling. Help me write a script. And then let's say who
is this for and what should the script I want to emphasize be who you are
anywhere and being free. Use inspirational and
professional tone that motivates the audience
to start traveling. Give a shout out at the
end to example channel. Let's assume example channel
is my YouTube channel. Very short and we got
a couple sentences. I'm just going to
have ChachBT remove these snippets and just give
me the voiceover script. Very short, very
straightforward, and this works well
with our tourism shots. That's exactly what
I'm going to do. Let's copy this and
then go into dubbing, create a dub example channel, source language in English, and let's say the
other one is French. French or Spanish, any
other language you want. Now, the video source, this is where we're
going to have to combine our videos and then put
them into one file. You can use any software
you want for that. You're just basically taking the videos that you downloaded, sticking them together, and then exporting them
for this part. Whichever source you choose, you still need that file. What I'm going to do is go to Premiere Pro and put these
four clips together, and then I'll be right back
so we can upload them here. Put the clips together. I came up to 35 seconds, but I did that just in
case I wanted to stretch out one of the dubs or
add some dramatic music. As you can see, it's just a very simple
compilation of these videos, no transitions or anything. But one thing to keep in mind is keeping the source
clips similar. They're all slow pans of touristy destinations
and because I'm talking about individual freedom
and being who you are, I do have people walking
in different places. Then last part where I'm
subscribe to our channel, it's a long shot of this deep. Later, I could put
text and stuff. Now that I have my base clip, before I go into dubbing, I'm actually going to get
some background music. I forgot to do that.
You can use Pixaba. It has free, no copyright music. Again, if you plan on
posting this anywhere, be sure to attribute the artists that you
can see right here. Let's look for commercial
or let's look here. Going to go with piano. This audio works for me. You can also look for other
things like travel music, log music, anything
else that you want. And now that I have
my base stuff, I'm going to go to dubbing
and start filling these in. You could put the
background music on the video file itself, so you're only
uploading one thing, but I'm going to go
to manual so I could perhaps have option
over the layers. Now that I have my source files, the video file and
the audio file, we can go ahead and
start making the dub. I will see you guys in the
next lesson where we do just that and finalize our
example channel advertisement.
29. Advertisement Part 2: Welcome back. In
the last lesson, we came up with an ad, and now it's time to
actually create the text, the text to voice, and then do a Spanish tub or another language
stub on top of it. Now, before I get into that, I do want to try making
a music for this ad. We did download that
soft piano music from Pixabe in the
previous lesson, but you also have the option to generate music right over here. So let's open that in a new tab. So you have variance, we just like the
different versions you get to create. This
is the duration. I will set it to 30 seconds
because that's my ad, but you can put it on
auto or any other values. These are your recent projects, but all you have to do
is describe your song. Let's do a inspirational let's try something
like this and we can also then change it up
if it's not to our liking. I set it to one variant and you can see that I'm only
getting one right now and essentially
we're getting the different parts made for us. 30 seconds like we asked for. And this is the styles
that excludes styles, you know, things that
we did not ask for. We asked for inspirational,
piano cinematic, but we obviously did not ask for a rock guitar or
anything like that. Then we got the
different sections. Let's hear our first variant and see what we're dealing with. That was actually pretty good. Notice here, it's interesting. I added a sample text for us to better imagine that voice
over over this intro. For me, this is perfect, but if you wanted
to change it up, you can either continue
this conversation, which means adding
a new section, maybe you want something else, we hit plus and you
add a new section, upbeat remix section
or something, then we can once again do the
include and exclude parts. So that's how you get to change
up your song composition. Once you're done,
I'm just going to delete this extra section. You can hit Download or
share it with other people. So when I hit Share,
it's kind of like a ElevenLabs audio library. You can put the song title, give it a little color, and then we can generate
the link to this song, preview it and Export. So when you hit Export, it actually exports this
animation. I just downloaded it. Let's see what we got. So that's pretty cool. I just
made all of that with AI. Okay, that's my
music for the ad. Feel free to still use the music you downloaded
from Pixabay, combine them. But I just wanted
to show you guys that this product exists. Now we can go to text-to-speech and paste
our advertisement. Okay, so I went back to Chachi
BT and I got our script. Let's go ahead and
choose a voice that's somewhat deep,
very slow tone, confident, and just like those classic commercial
travel commercials you see on the TV. So let's look for
deep mail, I guess. Life is like water. This is your deep. It was one of those rare smiles with a quality of
eternal reassurance in it. Welcome to my demo. A voice that doesn't just speak. Hey there. This is Dante. With a deep, easy going tone. The world of artificial. Okay, I think I really
like. Hey there. This is with a deep Angelo. And then we had Hey, guys. Life is like So it was
one of those Seko, I think, that's how
you say his name. Let's try both of them and
see which one works better. Let's click on add
to My voices and just generate first to
see how it deals with it. I will make this a little slower because he
did speak a little fast when I did the preview,
and then we can decide. There's a version of you that's waiting to be
unlocked out there. When you travel, you
don't escape who you are. You become more of it. Okay, so that is not the
sound I'm going for. Okay, so this is an example of when you could
actually design your own voice
instead of searching through the thousands
of available voices. So what I'm going to do
is head over to home, which I have it
opened right here and just go to voice Design. So let's just see if we have something to start
will be trailer voice. Let's say dramatic male voice. And then here what we have. In a world on the
brink of chaos, one hero will rise. Prepare yourself for a story of epic In a world on
the brink of chaos. In a world on the brink. Okay, so this is giving
a little bit more, you know, thriller and action. So I'm actually going to
copy our script and put it into texture preview and
give this a different thing. So dramatic male voice
say deep male voice. Inspiration and Travel
advertisements. Let's try that. There's a version of you that's waiting to be
unlocked out there. When you travel, you don't escape who There's
a version of you that's There's a version of you that let's try something else. Let's say, slow and
confident male voice. So let's go to Chat GBT
and have it basically describe that classic
travel ad voice. The ones that are,
like, you know, slow, they're inspirational,
deep, confident. But I just got to put
that sound into words. So let's say I'm trying to
design a voice with 11 laps, ElevenLabs for this
travel commercial. This is what I have right now. Help me get it closer to the
classic travel commercials. So this is perfect,
national geographic. That's exactly what
I'm trying to get to. Now you can see it took this one sentence and
built it up into more. Male voice, slow
pace and confident, warm and rich tone, and I'm
just going to copy this. If we saw that it's not getting
us exactly what we need, we can just take away from
the prompt little by little. But for now, let's see what the direct result
of using hachPT is. And recall we do have the settings for you to
either put in a seat, work with the loudness
and all that stuff. I'm just going to keep
everything at default. There's a version of you that's waiting to be
unlocked out there. When you travel, there's a version of you that's
waiting to be unlocked. There's a version of you. Okay, so it's definitely
a lot better, but it is a little bit too slow. So let's try to fix
that up over here, and I'm actually going to
add a British voice here. British male voice, slow
paced, arm enrich tone. Remove this line. So let's
remove the articulation part. And see what we got. There's a version of
you that's waiting to be unlocked out there.
When you travel. There's a version of
you that's waiting to be un There's a version of you that's waiting to
be unlocked out there. Okay, so this one
is really good, but I just want to make
this a little deeper. So British deep male voice. Just in terms of the pitch, I'm gonna let's try actually
generating this one first, and then we'll make another one. So let's say travel
at voice one. Language is English
accent is British. Save that voice. I will go to a different window
while that's happening. So let's just replace
this prompt with our deep male voice and
then put in our script. There's a version of
you that's waiting to be unlocked out there. There's a version of you that's waiting to there's a
version of you that's waiting to There's a
version of you Okay, these three are really good. I'm just gonna go back
and listen to our music so I can kind of
compare and decide. Let's open it right over
here. This is the video. There's a version of you
that's waiting to be unlocked. There's a version of
you that's waiting. There's a version of
you that's waiting. There's a version of you
that's waiting to be unlocked. Okay, I think I'll
go with voice. There's a version of you let's select it and call it
Travel ad voice two, British. And there we have it. So now we have our music
and our very own voice. Previously, we made our video, and now it's time to put it
all together and then do our translation to Spanish or any other
language you choose. So I will see if it's in the next lesson where
we do just that.
30. Advertisement Part 3: Okay, let's go ahead and
translate our video. Before this lesson,
I just combined my base clip and the audio
itself, not the music. And this is essentially
what I have. There's a version of you
that's waiting to be unlocked. Out there. You unlearn, you breathe these too. So the reason why I did
that first is because ElevenLabs will have an
easier time translating, and this will save you the
hassle of uploading the video and trying to match the word that you want at
that exact minute. You could go ahead and
skip this and just upload your audio
first in dubbing. So when you go here, when
you want to create a dub, you're able to upload the
audio or video source. So I'm doing video source because it has the
audio attached to it, but you could just do audio and then add the video later on. I suggest always starting
with a video source. Let's call this example channel from English to whichever
language you prefer, and I'm going to upload that very same
video. There it is. It has the idea once again. Let's create a dubbing project. And if you want to save usage, you can have the
ElevenLabs watermark, but I'm just going to
leave it be if you wanted to experiment and you're not sure
about your video, just check this and it will
save you a lot of credits. Number of speakers, we only have one and then we want the
full range to be dubbed, and that's it. Let's hit Create. It's going to upload
this video and begin the translation
from English to French. All right. So I'm going to
wait for this to happen. It shouldn't take too long, and then I'll be right
back so we can continue. It's finished processing. I'm just going to open it
and see our translations. Right now, we're familiar
with this interface. You have the original,
which is English, and then you have the French. So when you go to French,
you can see how it translated the English
version into here. Again, you're able to edit this if you know French and it's wrong or if it transcribed the English words
in the wrong way. I'm just going to
double check this. Okay, so not a lot of
texts to deal with, but it's all correct, at least. So let's go to original
and try to add our music. So all you have to do is put it right over
here for foreground. You can upload an
audio right over here, so bring the music. There it is. You can see it's even telling us to import something without a speaker because it won't transcribe that.
There is my music. I could rename this, click on the name,
call this Music. This is an imported audio. If this original
video had some music, it would have showed
up here in foreground. But since ours didn't
have any sort of music, that's why it's empty. Let's go ahead and play this and see if we want to
move anything around. There's a version of
you that's waiting to be unlocked out there when you travel. Escape. Right away, I could tell
that the music is way too loud and just
visually looking at it, the background video
ends at this time, but the speaker and
audio end a lot sooner. Now, with the speaker part, it's not a big deal
because you want some natural pauses
in there and when they tell you to subscribe
to example channel, you do want there
to be a little bit of just nothing until
the end of the video, so it ends in a smooth way. But for the music, we
could slow this down by grabbing the edge here and
matching it to our video. Doing this, it's going
to lower the pitch of whatever audio you do this
too, so keep that in mind. But because I have some
piano and some violin, that's not going to
be a problem for me. Of course, I don't
want to go overboard. Let me just show you
what that sounds like here without the speaker. You can see the music is
still pretty much the same. It's just that the
pitch is different. But if I overdo this or
if I go the other way, I squeeze it, it's still music. It's just very different. So keep that in mind. When you do that,
I'm going to hit Command or Control Z to undo. Now, the music is also
a little bit too loud. Travel, you don't escape. We want the speaker's
voice to be the main part. So just click on that audio
and reduce the volume. You you play around with
it until it's perfect. Become more of it. You learn, you unlearn, you breathe easier. B who you are anywhere. That's a lot better. Now I have my English speaker and my music. It's time to add
some sound effect. With commercials,
you want to have the viewers imagine
themselves in that setting. Since we are talking
about travel, we have some scenery
of putter balloons, that could sea, all that stuff, we could add and generate sound effects to put in
between these clips. So let's scroll down right here. You can see at Sex track. When you click it once, it's
going to just make the track and this screen bar lets you generate
those sound effects. Let me first see where I want
the sound effects to be. For the first part,
I can't think of a sound effect I want here. We could do people
chatting in the back, but because they're
not in the frame, that's going to be
a little weird. But for the second
clip, we have the C, and you can put
in some SGLs fine by or the waves crashing,
something like that. Here again, we don't
have much to do, but this we could do something
like a crowded area, so this person is taking that freedom and walking
through a bizarre or something, and here we can do
some wave crash. Maybe we do Segal for this
one and waves for this one. Let's go ahead and bring our playhead right where we
want this audio to start. We could move the SFX
after we've built it too. Click once, and I'm going to decide on the
duration pretty long, this long, stretch it over. No audio right now. We didn't put anything here.
Don't worry about that. Let's skip over to
the other part that we wanted an audio,
start one here. I'm trying to overlap the clips because that makes it
a lot more natural. Now here, when there's
two SFX back to back, so pack light, dream. So we have one of the person walking and then we have
one with the beach. So instead of putting them
right after each other, you could also layer them. Essentially, we scroll
down and we add another Saffx track and we
make our second one down here. The reason why
we're doing this is because when this
guy is fading out, this guy could fade
in and that creates a very natural
seamless transition. Let's make sure they overlap for a couple of seconds like that, and I have three
Aafcs to work with. Let's go to the first
one, and I will do seagulls flying by near a beach. You'll be flying by distance. Fade in and then fade out. Let's try something
really simple. Avon, you don't
escape who you are. You become more
of it. You learn. There we go. Pretty good. Siegel's flying in the back. I'm just going to mute
just highlight this track so click on this headphone icon so we can just hear the SFX. Notice how good that
fade in Fade Out was. If I remove that, I'm just going to copy this. Just get rid of this part,
generate it one more time. You can see that's not as good. It just suddenly starts, but by adding that
fade in fade out, it's going to be more natural. Fade in, audio and
then fade out. Just leave that out. Seagulls
flying by in a distance. I don't think we need to put the sound of the ocean
because it's really far away. We'll reserve that for
this audio. All right. This one, I'll leave it as
the volume it has right now. Later, we can listen to
the entire thing and then decide or how loud we
want the audio to be. Let's move on to the second one where we have this
person walking through, let's assume a bizarre,
a crowded one. Click on this can do a
crowded bizarre with people. Talking in the distance, sellers yelling their
prices, I guess. We can experiment with this. I will do one without
fade and fade out. Now, the reason why I keep doing in the distance
is because I want a faded audio and sometimes the volume alone is not going to be
that helpful for you. If you do one without the
in the distance prompt, it will be way too
clear and way too loud, but that's not the point here. We don't want the SFX to be the main character in
this advertisement, so that's why I'm adding in the distance for
all of these SFC. If you had the sound really
close up, for example, super close up shot of a bird, then you may want to
not use in distance. Just gauge what
distance there is between us as the viewer and that sound
that's in the video. This gave us a shorter audio. If you change prompt, you can regenerate that audio or click on this refresh icon
to get a different one. We can just click
on this right now, right click and do a
fixed duration one. I meant to make this like
this first. Fixed duration. Now it meets that
requirement that we wanted. The words that the people are saying don't make any sense, but that's okay because
it's a background noise. Now, by default, when
you generate here, it's doing the dynamic duration. If you did what we did here and you saw that it keeps giving
you a shorter version, don't try to stretch
it out because that just slows down the clip. But instead just right click
and do fixed duration. Now, our last scene is, let me see if this
fades in and fades out. Okay, doesn't. So let's
do Let's go here. Do fade in and
fade out all here. Generated one more time.
Meant to do it the other way. So this is an example,
fixed duration. It's gonna use the same prompt, but keep it the same
length that you want it. A Okay, so that's good. Didn't really fade in and
fade out like the Segel one, but I think it'll be fine. For the last one, let's do
the sound of the waves. Go in this one down here
and do waves crashing. We just do waves at the shore. Got slow at the beginning. Okay, right click. Fixed duration Audio. Hi. Turn it on like that. Okay. I think this will be
fine. I don't want two Segel sound
effects in one click. So when I play
everything together, I just uncheck this so everything is now
being considered. Where That's the alluding. You can see that there is no balance between
the different audios. So let's go to the
first sound effect. When you travel, you
don't escape who you are. Click on the first sound effect and just lower that audio. You travel, you don't
escape who you are. You become more
of it. You learn, you are a little
lower than that. Let's try the one here. The bizarre one, I'll
go a little bit lower just because it's the labs
one and then the beach, we can go 16 decibels,
negative 16. B who you are anywhere.
That's the real freedom. So pack light, dream big, and start your journey today. Subscribe to example channel. So notice how that
transition was really smooth from that bazaar
to, um, the beach. And that's because of this
overlapping thing that we did. Now, to make this one
a little smoother, because you can clearly tell when it starts and when it ends. You travel, you don't escape who you are. You become more of it. I think I will just
extend it to maybe here. So it's less noticeable. Right click, fixed duration. And because we're so focused
on this audio ending, we could pay less
attention to the SFX. You don't escape who you are. You become more of it.
You learn, you unlearn. It's a little too low. So let's increase that. Become more of it. You learn, you unlearn. You breathe easier. There. When you travel, it. You learn, you unlearn,
you breathe easier. Okay, that's pretty good. I think I'll do another
overlapping thing. So let's move this to its
own layer as a vex track, and I will just put my playhead to get
some direction here, put it like that, and
now give it a listen. Make it a little louder. You become more
of it. You learn, you unlearn, you breathe easier. Be who you are anywhere. Okay. So that's pretty good. Let's just extend the
SQLs to the beginning, maybe a little bit further back, something like that,
generated again. So it's like this
ends, this starts, then this starts kind
of like a go situation. When you travel,
you don't escape who you are. You
become more of it. You learn, you unlearn,
you breathe easier. Be who you are anywhere.
That's the real. Okay, so I'm pretty happy
with this original clip, and I'll just play it for
you guys as it is so you can know what we're dealing with before and then what the French version is going to be. For the French version, there's not much you can do. You can see that
it just duplicates the same audio track and
simply translates it. So you're not doing
much editing here. Let's start over and
see what we got. There's a version of
you that's waiting to be unlocked out there. When you travel,
you don't escape who you are. You
become more of it. You learn, you unlearn. You breathe easier. Be who you are anywhere. That's the real freedom.
So pack light, dream big, and start your journey today, subscribe to example channel and see where the
world takes you. So that's the original
version. It's all good. What does the French
version sound like? Now, when I play
this, you may hear, like a duplicate voice. Iliana. So if this happens
to you, don't panic. Right now, I'm in
the French version, but we can hear something like the English version
in the background. Now, right now, I have
my French speaker, my original, and the background. So there's three speakers
in this clip right now, and when we listen to it, it sounds a little weird. We have the French speaker and someone's muffled
in the background. Don't panic if this
is happening to you. This only happened
because we had the speaker as part
of the original clip. Even though the speaker one
original is grayed out, it's still reading it weird because we have the
background music. If I just highlight this
you don't who you are. You can see that's where the
muffled sound comes from. All you really have
to do is delete it, and then it should be back to normal. That's a simple fix. Now to export this, you can just go down
here, click on Export. These are the stuff
that you've done before, but down here, you can decide in which of the languages you want to
export and then as what. Pfore video will have the video, the sound effects, the speaker,
and all of that stuff. You can even export
this as a zip file every layer is its own audiora you can do the timeline data. If you have any captions, we don't audio in this format, in this format or this one. Then when you're done,
you can hit Export. It's going to load it up here. You get time stamps
and in what language, then you have the option to do a preview and then download
it when you're ready. Tvantage oye Kivu said, part two, Seliberte. Hello, Voyage grand. He come and save a voyage shows. I believe example, channel, I voye will amount for salmon. That's how you can make a very simple advertisement
with ElevenLabs. We started everything from
scratch, from the script, the voice, the audio, speakers, sound effects,
and all that stuff. The only thing that
we didn't make was the videos that we
got from pexels.com. But if you want this to
be a full AI experience, there are some tools
out there that can make AI videos for you. This concludes our third
and final project. I hope you guys were able
to get the same result, if not similar to what
we have right here. Sure to try even more projects and practice the tools
that we learned about. Now let's move on to
the next chapter.
31. Creator Tools Introduction: You can now generate image
and video within 11 labs. Apart from generating
those assets, you can also upscale them, create your avatars,
and really have a full production
on one platform. This is a new edition, image and video,
and essentially, this is what the
page looks like. We're first greeted
with avatars, which is a way that you can
keep consistent characters. So say you have one face, you want it to be the same thing across different
video generations. Instead of trying to perfect
that with your prompt, you can just generate
an avatar and that way, everything will look
very consistent. Here, we're seeing a bunch of examples of what
other people made, as you can see, they're
pretty high quality. I could go into each
video, either remix it, meaning that combine
my own take on this video or just download
it or share it as it is. We have the option for
both video and image. And with the image, you can
actually see the full prompt. So basically the instruction
that this person provided ElevenLabs in
order to get this result. And the same thing goes
for video as well. So here we have an animation. We have a long prompt, as well. Now the cool thing about
ElevenLabs when it comes to image and video
generation is that it combines all of your
favorite AI models. So just to show you a preview, when you go down here, you may be familiar with
most of these models. So we have GPT, we have craft, we have SDream, we
have nano Banana, and so many others if
you just scroll down. And that's what makes this
platform a little bit more unique because apart from their own powerful
models for music, voice and sound effect, it also combines all of
the other models that you may have relied on and
brought it onto one place. Apart from the models
that generate content, we also have editing models. So upscaling, this is
another famous tool. It's been brought
into ElevenLabs. We also have background
removal here as well, and there's some tabs for you
where you could explore it. So if you want to
upscale to edit, we can use these models, and they take credit as well, so you could go based
on the cost here. Same thing applies to
video that was images. We have video, and
then we have lip sync. So lip sync is where you made the audio or you
just uploaded it, and you wanted to
basically generate a video where you
match the audio. So if you made a voice over
here or did a text-to-speech, you can now have that speech set by an avatar
in the lip sync. So that's another cool feature. So let's dive into this new feature in the
next couple of lessons, explore how they work, how you could get your first asset made, and how you could organize
them moving forward.
32. Assets & Templates: So anything that you
make is going to be saved in your assets folder. So if you click on that, we
have different things already because we have been using
ElevenLabs for a little bit. Any avatar that you make
is going to be saved here. You flows for the AI agents
are going to be saved here. Anything you upload to ElevenLabs will be
saved here as well. So I have two images here. Can upload assets so
that you could easily drag them into image
or video generation. You can make a new folder where you put certain things inside, and this way, you can
have a storage area online for your content. So I just made a folder. We could give it
a different name. So I will say my video
and images just so we could save the
things that we're going to make in the next
couple of lessons. This with other people,
delete it if we need to. When we go in the folder, we get to upload things
from our computer. And once we make something
in image and video, we get to put it onto
that area as well. So I could just go over
here, download this, and then upload
it to that folder to use as a reference
for another project. If you're not sure how to create a full content using
this new feature, ElevenLabs also has
the templates tab, so that's right
above the assets. And when you click on them, we have some easy
drag and drop options for you to create different
sorts of production. So a couple of examples here, we have let's look
at the categories. Say we want to do mockups. I have this packing
mockup generator. I could click on it, upload my blank product and the
two D packaging design, and it's going to
be turned into this really nicely polished up. So the templates essentially do something for you
from start to finish. This is an example for
a mockup generator, but we have a bunch
of other things. So let's close this. Here we have a UGC
video template. So if you want to create
that sort of content, you can upload the inputs
for this template. So the influencer portrait, which this is an example. But in the next
couple of lessons, I'll show you how you could do the AI image generator
in ElevenLabs, and that way you can use
your own generations, too. Next, we have the product. So whatever product you want to show, you will
put it right here. Choose your voice.
This could be from your own voice library or the many voices
available to you. And then basically the benefits that you want this
model to now highlight. So here's an example. Then we have the shoot location. So right now, it's like
a pink vanity corner, but you could put in the forest, an office or anything else. You're done, you can
just click Generate, and you got yourself
a nice UGC video. Now, of course, whatever you do, you still have the chance
to edit it further, make refinements, and
make it really your own. Down here, we can see how much credits this
is going to take. But just to look at the flow, basically how this is going to generate that
final product you, it's taking the following inputs that we looked at on the side. So if you're curious,
you can always come into the flow section. So influence a portrait. This is an input, the
product, the voice, perfect match, product
benefits, and shoot location. So it's going to take these guys and write a script first, decide on the length and
then have the models run. Then it creates a
couple of versions. Down here, it's working
with the voiceover and then it puts the two together
for the final product. And then this is like
the final final product. You could use the flow to figure out basically how it
approaches this issue, it's kind of fun to look at too. So you can edit
this flow as well, say you wanted to at some point, bring in another input
for a second character. You can add that right here. We're going to learn
about flows later on. But for now, this is something
that you can explore. Now, these were the
examples that we looked at and also things that
other people have made. So you can see the
image product. These are all really
well done, really. You can just see the
quality of the work. If you go to Generations, these are the things
that you have made by filling these in and
then hitting Generate. So right now there's
nothing there. Think of templates
as little tools that you could use on top
of all the other ones, and they're really
cool to work with. The only thing about
them is that they do take a lot of
credits to make sure that you're allocating
enough credits for that and you are aware of how much
is going to be used. So here's video
effects for dubbing, upload your input video and
then you take the dub out. Also make your own template if these tools are not
really for your need, you can click on
plus New Template. And here you can kind of talk with ElevenLabs to
build your own flow. Now, again, I am going to dive into flows a little bit later. But basically, if you have something that you
need to do repeatedly, like tasks that you have to
do every day where you upload a video and it
needs to be broken down into a vertical reel, horizontal reel, and
maybe a square one, you can build a flow based on that and automate
that process. So basically the input would be the the music and
maybe some captions. You'd have 11 labs, break it down into three
versions of the video, and then refine it with
maybe more AI content, AI music, voice over, and deliver you
that final product. That's an example of a flow. If you want, you could just leave it up to 11
labs completely. It's very user friendly. You don't need to use
specific wording. You just say what you want, and then it will help
you out along the way. So going back, we
can now move on to how to start generating
things with ElevenLabs. We're going to first get started with images. So right over here. And we're going to see how
this whole interface works. So let's jump right in.
33. Image Generator: So the first option here is
using this as a reference. So again, this is a
template that I could give 11 labs to generating new. Maybe this time, I could have a penguin in this
exact same pose. But in order to best
describe the pose, I could just upload
this image as reference for the
tool to follow. That's a reference, and then
we have upscale image that's going to increase the size and the resolution of
the current image, make it high quality. For example, if it's
a 500 by 500 image, I could upscale
it to the size or even larger and use it
for different purposes. You can recreate something, so basically put in this prompt, but see what comes out of it. So it's going to look a
little bit different. We can use this for lips. So if I have an audio already, I could have this subject
basically speak that audio. So it's going to this option is going to work on the lips, the movement, facial
expressions, and all of that. And lastly, we
have create video, so this could just be emotion. It could be the subject doing something or
just standing there, things that we can define
using this option. So a lot of cool things
are being made here. We have mockups. We have art. We have some like photography
techniques here. We have some designs, posters, illustrations,
and so on forth. So we're going to do a very simple exercise
right now where we try our first prompt and then see how we
could build upon that. So what I'm going to
do is type my subject. I want a woman, and I'm just going
to send this in. So I'm trying to show you guys the power of a
prompt and how well the output is going to be if you describe things better and
you follow a certain format. I now need to see
what this woman is doing. Is she sitting? Is she standing, lying
down, what is happening? I need to mention
that. So a woman sitting down on a chair. So this right here
basically sets a scene. Not enough details,
but we're just going to send this in with
one of our models. So again, we can go to image, and there's a lot of models
that you guys can explore. To get started, you can try
one of the trending models. GPT right now is very
famous for image. Nano Banana is also really
good for characters, realism, and all of that. I'm going to stick
with the newer model right here and click
on that alone. So we have our model set. We can now describe basically set the aspect
ratio for the final product. I'll go with the landscape shot. Next, we have the resolution, just to get a fast generation. I'm going to keep
it to the minimum. Quality, I'll just go with
high somewhere in the middle, and then we have generations. So how many versions do you want 11 labs to make based
off of this prompt? I'll do four just so we
can see how they differ. And lastly, we have this editing tool which
helps improve your prompt. But since I'm trying to show you the difference between prompts, I'm going to turn that off. So just make sure you're
image, put in your prompt, this exact prompt or
something else that you prefer these settings,
and let's send this in. Here are my different versions. We have a similar room. We have a similar subject. They're just in different
shirts and different poses. But it looks like
it's the same couch. There's a slight
difference between them. So right now, it did give me
a woman sitting on a chair, but it doesn't really have
anything unique to it. It looks like a stock
image, basically. But now, whatever
I have in my mind, whatever I'm envisioning
for this woman, I can now describe it
building upon this prompt. So either click on this
or type the prompt again. Now we're going to add
a few more details. Since I can see the
different versions, I could just cut
this down to two, so it generates faster, and we can go through
this a little quicker. So a woman sitting
down on a chair, before we even
mention the woman, we can see what she looks like. So in terms of age, in terms of maybe ethnicity, and then after that, what
does her hair look like? What's the hair color, things that we can better
describe the person. So we can see a young woman
this case with, let's say, blonde hair this time, we'll
do the opposite blonde hair and blue eyes sitting
down on a chair. So right now, just
the woman alone, we're adding more
details that can make her basically pop from
the following results, and we could even add some more. So she's sitting
down on a chair, maybe, what is the chair? Exactly, what does it look like? Is there a certain color?
We can mention that here. Sitting down on a velvety chair, and then we can mention
more about her pose. You can say with a
leaned forward pose and a confident face. Something like that. So more
details, we can put that in. And we could also mention
things about the background. So right now, I'm
just going to send this in with the female only, not going to mention
the background, and we can see how
the new result is going to be a little bit
different from the first one. This ended up being
locked because probably because of the leaned forward pose, sometimes
with wording, it may send some flares, but I just removed the lean forward thing and
just did confident pose. So if you run into this problem, just alter your
prompt a little bit. So this is the pose that I got. It's not exactly what
I was looking for, but you can see the difference between this image
and this image. The chair is indeed velvety. She is blonde, and she
does have blue eyes. Now if I want, I could set the same character
in this position, say, I really like
this background. I could now use this
pose as reference, say, this is what
I was going for. Instead of trying to alter this, I could reference this image. So I'm just going to reference that first image and
have my prompt as it is. And basically, the tool
needs to replace her hair and the chair and
give me a new photo. So there's my new model. She's exactly in this pose, but she's blond and the
chair looks different. So this is what we
have right now. It's looking pretty good. And since I have my
model finalized, I can now use this as reference to have
her do other things. For example, she has glasses
on and she's on her phone. That could be a different
pose in this exact setting. But by providing the reference, I could put less pressure on the prompt and more
pressure on the visual. If you're not sure how to
best write a prompt or you're struggling by struggling to refine it to what you need, you can turn on this feature. So basically, this is going to work through the current prompt, improve it, and give you
something that is more suitable. So I could now click on Enhance, and it basically transformed my very simple prompt into
something more detail. So flowing blond hair, using words that would
better describe texture, light, contrast and all of that. Instead of just blue eyes, it's piercing blue eyes. So more detail.
And then the color of the chair, what
type of chair? And then the lighting, we didn't even
mention the lighting, but it's doing that for us now. Something about the
texture of the chair and, of course, the
posture that she has. Different things you
could do here as well, instead of just enhancing, you could further
refine the pose. So a subtle, knowing
smile gracing her lips. So this is more about her,
the way she's sitting. And then, you know,
you could go on and on by adding
further details. So let's generate this with the same model and everything
and see what we get. I did remove the
reference image. Sometimes you may run into
issues with certain models. For example, this is my prompt, but it's saying that I may have violated their
terms of service. So certain models prevent you from using certain
combination of words, and that's just to keep
their community safe. But if I just switch
to a different model, it may interpret my prompt in a different way that is not, you know, deemed harmful. So we could try
maybe nano banana instead and try that
out. So there we go. You can see that we didn't
run into that problem. But here we can just see the big difference
between the models. So this is the nano banana one. You know, it looks like
she's in a studio, whereas the other one
looked very natural. So changing models really
helps when it comes to the visual delivery
of that prompt. So that's one example. We could go down to Runway AI. That's another model and
see what that looks like. We have Gen four Image. That's one of the
runway AI models. And so far, we're just
building upon our prompt. We first started with a woman
sitting down on a chair. This is where we have ended up. And, you know, with
different models switching between references, we are fine tuning that
image that we want. So there is my new image. It's very different
from the previous ones, but that's just how this model
is interpreting my prompt. So a lot of models out there, each AI model is
built differently, trained differently,
so expect to see some big differences when you
switch between the models. That's pretty much how the image generator
works in ElevenLabs. The purpose of making
one is so that you could maybe switch it
into a video later on, combined with the audios
that you have made already. It could be to create something with the
templates that we looked at and you could either keep
it within this platform, use it for different
output or export it into something
else externally. So maybe using it on another AI tool or just
posting it somewhere. You could explore this tool. It's new, it's improving. But since there's a lot of
external tools combined here, it really makes it
convenient for you to make all of your
content in one place. So that's the image generator. Let's move on to the
video generator, which is the next option.
34. Video Generator: Similar to how you made images, you can now make videos. So that's the next
tab right here. Once you switch over, there are a couple of
different things, but the core concept
is still the same in the sense that you
still choose your model, the aspect ratio, the
quality, and your references. The only difference between image and video is
that you now have starting frames and frames and video and audio references. And of course, the duration because we're now
working with motion. Way this works is that you can either just type
what you want and have 11 labs generate a video for you purely based
on the prompt, or you could have an image generated and then use
that as your reference. So in this lesson,
we're going to do a very simple exercise where we generate the image and then use that
as our reference. We're just going to
switch over here, and we're going to write
something about our character. So I want a video where the subject is running
away from the camera, and I want him to wear, like, a winter coat and
maybe run in the snow. So that's my idea. I just have to bring it to life. And it all starts with what
my character looks like. Switch over to image. I'm just going to
type in a man wearing a winter coat looking
very serious. And then I want this to be
full body and hyper realistic. So there is my prompt. I'm going to switch
to nano Banana, and then I'll keep this at 41k, and we'll do a
vertical aspect ratio. So no references
just typing this in, and we're going
to generate this. So here are the characters. We have our first one, second, third and fourth. I could now choose the
one I like and then use that as a reference to
further define my character. So I'm just going to
go with this guy. I think, you know, he
has the full setup, and now I want him
to hold a bag. So we're going to use this
guy as our image reference, and then we're going to
type in a new prompt. So a man holding a
backpack in his hand, you can say which hand, what color the backpack is, but I'm going to keep
it just like this. Same model, same settings. We're now going to send this in. Here are the models. They're now holding a backpack. Again, I didn't really say
what color or which hand, but I think this one
looks the best because it matches with his outfit
and he has the gloves on. Now, this can be my image
reference for the video, and I could also add in other references as we
looked at it earlier. We can now switch over to video. And have this image as
our starting frame. So I'm not going to use any
reference for this example, but only the starting frame because that's technically
a reference of its own. But actually, if
we try to do both, you can see that we're
not allowed to do that. So you cannot have
a start end frame alongside image or
video references. So now we just have to describe
what this man is doing. So we could say man running
away from camera chasing. I want him to be running
really fast. Name settings. I'm just going to switch
this to 10 seconds. And I'll keep this to 1720, and I'll use the S dance 2.5
model. Let's send this in. So that's going to take
a while to generate. So at the same time, I'm
going to have another example where we remove the start frame
and use three references. So same prompt, we're just going to use our image reference, which is our character.
Drop that in. Video reference, I could
find a video online where I want my character
to run that way. If we just go to the main page, explore video and search
for Running Away, I could use something
like this as my video reference. Like that. And for audio, I
actually made one earlier where I wrote
that man is running away, so running sound, heavy
breathing, that sort of stuff. If you don't want to
make your own audio, you could also look for
something in the sound effects tab downloaded and
simply upload it here. There's my audio.
Again, same settings. I'm now just gonna hit Sent. So my videos are done. This is our first one,
where we had him run away. So we can see that it provided the footstep noise, you know, suspenseful music,
and, of course, him breathing a little
bit at the start. Now, I could guide the audio, maybe have maybe describe
the sound of the footsteps, which is going to be
running in the snow. This here sounded like a rubber boots sound,
which wasn't suitable. But you could see
how it implemented that chase scene audio
using just two words. So that's our first video. Let's take a look
at our second video where we provided three
types of reference. So that was our subject. You can see that he's wearing the same outfit, pretty much. Then we had our video reference, which is pretty identical. Lastly, we had our
reference audio, which was just the
heavy breathing. So all of that combined
gave us this video. And you can see, at some
point he turns back to the camera and he looks
exactly like my subject. So that's the video generator that ElevenLabs
recently introduced. It's really powerful.
It's only getting better. You can try this with
different models, but bear in mind the amount
of credits that it's going to use because this is on the
heavier side of production. You can see that this
five second clip used 15,000 credits, and I did combined three
references to make this video. Down here, you could
see the size as well as the dimension,
the resolution. So you could use this
information to better plan your production without running out of credits every time. Now, let's go ahead and look at the lip sync feature
inside ElevenLabs.
35. Lip Sync: We saw how to make video images, and now we're going to have our characters say
something with lip sync. So that's the third tab, and all you do really
is upload an avatar and a speech that you either made on this platform or had
it from another place. You can guide the
generation here, maybe talk about the way the facial expressions are made or the speed that
the person is speaking. So we're going to
just click on Avatar. You can either
upload an image for your avatar or select
one that already exists. If you want to make
your own avatar, you can do that with the
image generator tab. The one thing that you
need to bear in mind is that you want the
image to be close enough to the subject's
face because a lot of the focus and generation is going to be around the
face and the mouth. So this image right here, the character is a little
further away from the camera, so we may end up with
some glitching the face when ElevenLabs is trying to
make the characters speak. But when you go to the homepage, you can see that we
have some avatars. So when we click on View All, notice how close the
subject is to the camera, and this is going to
really help with the most basically making
natural gestures, natural lip movement,
and all of that. Basically, you're going
to click on Avatar, select Avatar, and then
choose whichever you want. Once you click on
them, you actually get more options about the
positioning of the subject, the clothing they're
wearing, the location. And you can see that I
have a lot of options. So if I go with this one, I could click on you style
and then go for the speech. It could be text-to-speech or something that
I have uploaded. So if you do text-to-speech, this is our saved paragraph
from a couple of lessons ago, and I could just either
use the speech that I made in that lesson or generate a new one with
a different voice. I'm just going to click
on the latest one, and you could just click on text-to-speech or
upload your audio. If you want to make
your own Avatar, you just have to click
on New and either create the face from a prompt
or upload reference images. So basically, if I go
into any of these, you can see that there is a lot of different angles from
that one character. And that's basically what
you want with your avatar. So if I head over
to upload images, I'm going to need different
angles of my subject. And that's something that again, I could make in ElevenLabs or something I could do
with my own camera. So upload the images. Ten is recommended. You can use these as reference if you want to take
your own images, or you could let 11 labs create each one for you
and just upload them here. You can give the Avatar a name and a voice,
and that's it. You just click on Create Avatar. If you want to start
from a prompt, you don't need to upload any images or any
references at all. You just describe
what they look like, the way they move, and that is. Only issue with this is that
you may have to do a lot of editing to have 11 labs
understand your ideal avatar. So if you have the images or
if you've made them here, it's better to go
for this option. So I just chose my avatar
and I'm going to do a text-to-speech,
something very simple. You just have to change
the voice to, you know, the avatar's voice because
this is a template. So I'm just going to cut the majority of this
paragraph and have him say this very informative
sentence, generate speech. Last Friday, students from Sunshine Elementary
School planted 20 young trees in
the school garden to keep the environment
clean, green, and cool. Once I'm happy with this, I could use the speech
and then send that in. Down here, we do
have a few options. We already know
what these two are, but this guy is basically
the model for lip syncing. So it's the same idea as the
image and the video models. Only these are made for
a different purpose. You can go with any of these, but I just went with
the most trending one. Now, while this is loading, we could spend some
time generating our avatar to see
how well that works. Our avatar is now
saying our speech. Last Friday, students from Sunshine Elementary
School planted 20 young trees in
the school garden to keep the environment
clean, green, and cool. You could see the lip movement. According to the words
that are being said, he does have some
motion in his body. He looks at the camera, and it has that podcast
like wipe to it. Last Friday, students from
Sunshine Elementary School so, of course, the motion
is not that natural. This is AI lip syncing, but it could be helpful for
certain types of content. And you can explore this
with other avatars, even make your own avatar. You can see how you can combine your audio with either an avatar like this or a video content that we made in the
previous lesson. So these are two options to make your clips more engaging
while showcasing the quality of your
AI-generated voices. Mm hm
36. Upscaling with ElevenLabs: If you ever ended up with
an image that is not the best quality or you just want to change
the size of it, you can now use
upscaling in ElevenLabs. So I just went to my
assets folder and I uploaded an image
that I made with AI. So currently, it
looks like this. It's a closeup of a man's eye, but I want to add more detail. I want to make it larger. And instead of going to
another tool paying for it, I could just, you know,
come in here with the same subscription
and upscale. Click on Upscale, and it's going to bring you to the
image and video tab. Right now, this is the image
that we're going to upscale, and it just automatically
chose this model. So this is the only model right now that does the upscaling. And then on the right side, we can decide the size of it. So I'm just going to go all
the way to the max just so we can see the difference,
and that's pretty much it. So you don't need to type anything or give any
other references. You just send it in, and it's going to upscale your image. It's done upscaling. It's up by four times. This is the original
image that I made. It looks like this,
pretty detailed already. But pay attention
to the size here. This is the size and the current dimension in
pixels of this image. I then upscaled it
using the model, and now we can look at the size as well as the dimension
again in pixels. So that is an improvement
by four times, and I can now use this for bigger projects without
losing any quality. That's pretty much how
you could upscale. You could further
upscale upscaled image. So I'm just going to go to
the one that I just upscale. I still have the option
to make it bigger. So I could go to this size
now and send that in. But that's only going to build upon the
previously upscaled image. It's a pretty straightforward
tool if you felt like, you know, the image that you
cut is just lacking detail. It could be more crisp. You can use that feature to, you know, make it
highly detailed, much larger, but
bear in mind that the size is also going
to be increased. So you can go ahead and try that out with the
images that we've been making and compare
the different sizes.
37. Flows: All of the new tools
here really give you the possibility to expand
your creative outputs. So this could be music, sound effects, video, image, and the classic speech-to-text
or text-to-speech tools. But flows takes all of that to the next level and
allows you to basically have a canvas for generating the content
with all of these tools. So it's a very different sort of layout to what we've been
used to this whole time. We're going to be exploring
it in this lesson. We know that studio can be
used for larger productions, but flows gives you a whole new layout where you get to drag and
drop different tools, connect them, and kind of work with them as
if they're nodes. So if you're
familiar with nodes, it's pretty much that same idea. You create one, you connect it, and you build a flow. So here are some examples you can think of a final product, but then go backwards and see what the different
components were. So this guy right here,
if we click on it, we can see the flows that
is behind this production. It looks like a lot. And these are things
that we may not pay attention to because we're
only seeing the final product. So basically, if we zoom in, you can see how everything came about and where the
starting points were. So we have the initial designs. We have the text. You can use the space bar
to move the canvas. We then have different text, again, some more text because
this is an explainer video. Um, and then I'm going
to find the visuals. So these are the visuals
which were not generated, and then I could move
each of these notes. So right now if you look
at this, it's a lot, but I'm going to
show you how you could make your own flow, and then this whole
thing will make sense. So this is just an example
of what it looks like. If we go back and
hover over this, this is what that
flow chart two. So let's click on New flow, and we're going to be
met with a blank Canvas. So just to go over the
interface real quick, we have the menu
on the left side. So you can make a new flow, quick actions, give
it a new name, duplicate share,
create a template based on this flow that
we're going to make. Even see your version history. So you could go back
to an older version. First, let's give
our flow a name. We're going to call this
my flower video, maybe. And the concept that I
have in my mind is like a flower pot where a new
plant is growing within. So that is my concept, and I'm going to make
a flow based on that. So before you click
or write anything, you're going to need some
different components that are going to lead
to that final product. I know that my final product
is going to have a visual, a video, and a video is going to need some
image references. So that's already two things. It's going to need an audio. Maybe that could be a voiceover, as well as some
background music, and I want to add in some
sound effects as well. To help write the
voiceover content, I'm also going to
use AI for that. So if you count all of those, they're going to be
multiple components. But from now on, I'll call
those components nodes. So simply right click
on this Canvas, you're going to have an
option to add a new node. You can also hit N on your
keyboard for a quick action. So notes can differ. It could be an image, video, text-to-speech, lip
sync, and other things. So these are all things that we have looked at so
far. It's nothing new. The only difference
is that there are now smaller parts that play a
role in a bigger production. So these are all like actors in a bigger movie, basically. So like I said, I'm going to need an image to get started. So let's click on this, and
we're going to get this box. Now, the interface here is what we've been
doing this whole time. You choose your
model, the ratio, resolution, quality, and then you write your
prompt, you hit Run. Only difference here is that we have these
things on the side. So these are the connectors. So what is going to be done with this image
after it's done? Is it going to link to
a video production or is it going to be the
result of a text box? So this is where the node idea is going to make more sense. So instead of writing
something here, I'm going to make a text no that way I could
have different models, make that same prompt, and then I could choose the
one that is best for me. That's my image right here. Let's right click
again, add node, and we're going to go down to or just go here for
text and choose text. That's my text box. I could start writing or I could use the help of a
large language model. So say I don't know how to write my prompt for that flower pot, I could create another node. I need the help of a LLM. Let's put that here. Let's go down. And now
I have three things. And just to tell you what
the flow is like so far, we're going to have
a prompt generated by Gemini or any other
model that we want. That's going to lead to this text box where I could
see my finalized prompt. This text prompt is going to lead to the image generation. So the image is going to
be made based on this guy, and this is going to be
the result of this guy. So there is a connector,
0.10 0.2 0.3. So let's start right
here with our LLM. You can see we still have
some reference points, but I'm just going to have
this be the starting point. So let's say give me
an image prompt for a nice plant pot that has no plant in
it and is only soil. Make this indoors
with good lighting. So that's my prompt. I just basically have an idea so far. I'm not going to have
the final prompt here, and I'm just going to
choose my model now. So this could be anything. Let's say you want to
do Gemini is fine here. Then you can choose the
mode for that model. So if you're each of these models are going
to have their own mode. So, for example,
if I chose Claude, I'm going to have
different options for the reasoning and thinking, and these will all result in different things because we're dealing with different models. And then we have
more here as well. So let's just go
with the Gemini 3.5. I'll do minimal. There's not
a lot of thinking needed. Here we can decide
on the length. This could be automated. I'm going to leave it
entirely up to Gemini to decide the length
of my final prompt, and then we can
start running this. But before I do that, I'm
going to point the output, which is on the right
side to my text box. So these are inputs.
These are outputs. Use the space bar to
move the canvas around. So I could also use this
actually as my input idea. So let's break this
down even more. Let's copy what I
wrote and put it here instead and use
this as an input. So just connect it to. That's our first connection. You can see that
the outline became darker and we're
seeing an actual line. I could always cut this connection so they're
not related anymore, or I could just make
a new one like that. So now it's doing at text. That's the name of this box. If I change the name, let's call this my idea. That's going to be
changed here as well. So it's telling
this box to look at this guy to get started
with that prompt. So at my idea, we're going to hit Run now. Using our Gemini 3.5. So it's going to generate a
nice prompt for us that now I could connect to my image
node. So I'll call this. So this name is fine.
I'm now going to output this as a text input
for my image node. If you want, you could
also make your changes, run it again and then do this, but I'm going to be using
this prompt as it is. So now notice it says
at LLM right here, and I'm going to have this image node generate
with my favorite model. So I'll do nano Banana, choose the Aspec ratio. I'll do a square for now, the quality and everything, things that we have used so far, and I will hit Run. So that's going to do its thing. But say I want to
do multiple models. So this is not enough for me. I also want to try JAGBT. I want to try basically
a bunch of other models to determine which one
is best for my project. We are going to right click
again or just click and drag. And once you let go,
it's going to ask you, well, what is this
going to lead to? So I'll just do image
generation again. And this time we can
choose something else. So let's go for Cream
and then hit Run. Now my text is leading
to two things. We had our first
image, second image. I could do a third
one, a fourth one. As long as they are
connected to this, this is going to
be their prompt. You could, of course, write
the prompt directly here, but creating a flow
like this gives you a lot more flexibility when
it comes to editing and, you know, refining your project. So we're going to have
this guy do its thing. You can see that it's
about 30% done right now, and I could continue
building based off of this. Now, you could also add a image reference by right clicking,
adding another image. So let's go to image right here. Click Image generation.
Now, over here, I could just write something. We can say basically this idea, but not counting the pot. I want to talk about the room. So I'm gonna say a
cozy living room with open windows
and blue furniture, something that could be a little different
from what we have. So you could again use a large language model to
help you write the prompt, but I'm going to do
something very simple, and I will use nano
Banana for this run it. So here's our second image. You can see how
different they are. And I could now connect
this to something else. But first, I want to use this room and have it be an
image reference for this guy. Say, This is more ideal
for me rather than this. I could now connect
this image output as an image input for this pot. So let's move this here. Let's give it a name, actually. So I'll say this is
image reference. This is Image one. Let's say. That's image two, just so you know what I'm
referring to. All right. Once we have our image labeled, we can now take this as
an input for a video. So we learned about videos. We learned that we have starting frames and image
references involved. So this could now lead
to my video footage. So click and drag, let go, and we're going to choose
Video generation. There it is. Here you can see that we
have multiple inputs, just like we saw in
the previous lessons, we have starting
frame, end frame, image reference, video
reference, and audio reference. We also have the prompt input. This could be whatever
you type here, or it could be a
text box like this. I'm just going to have
this be a image reference. So let's hit X and then
connect it again using this. So that's what it's going to be. This way, I could reconnect
it to whichever input I want. But I actually realized that we wanted to redo the background. So now that I have
two things connected. Original prompt and
this new background, we're going to hit run again. So that's going to redo the whole work and give
me the same flower pot. So there's my flower pot. We could see that it
integrated the blue, not so much the couch, but that's okay for me. Since I just wanted
to show you guys how an input reference is going
to change the result. So now that I have
a blue flower pot, I'm going to use it as an image reference
or a starting frame. Ing on what you need.
And I'm going to have one of these models. I'll do seed dance. Have a basically seed sprout into a flour from this very pot. So seed sprouts in the pot
turning into a nice red. Flower. Choose the settings. I'll do squares, since that's
what we're doing so far, the resolution, and then you
can choose the duration. I'm going to mute it here
because I want to make another node for the
audio. So let's run this. If you want, you could have
another large language model. Node, be here to create
a better prompt for you. But we're just going
to do what we have right now and see
what comes as we saw, we could also make
multiple notes for video where we
try different models, and then we can choose
the one that we prefer. Now, while this is running, we are going to
work on the audio. So somewhere along the canvas,
I'm going to write click. Add node, but this time, we're going to go for audio. This could be text-to-speech, sound effect, music, or anything else that
you want to get started. We're going to right
click here and make a simple audio note. I'll do text-to-speech, and we'll do a very simple sentence. So we can say new beginnings and new flowers
all here to stay. Something like that,
choose your voice, and you can also work
with the stability, choose your model, and
then click on Run. So it's 3 seconds. New beginnings and new flowers all here to stay. Very good. And I could then
combine this into, you know, my video right here
and get my final result. So just drag this audio output and put it in as an audio
input. Move this to the side. And so far, we have
this, you know, mixed up map all leading to
our final exported footage. Bear in mind that each of these do take their own credits. So make sure that you are allocated enough
credits for that. I'm going to leave
mine like this, but if you hit rerun, it's going to apply this audio
into the footage as well. Once you're done
with your video, although this is not
really finished yet, you have the option to
copy it, duplicate it, maybe connect it
to another where you have someone grab the
pot or a different scene. You can make a template out of this whole thing
with this button. You can download
it when it's done. And since this is a canvas, you can also invite other team members
to come in here and, you know, help you along
with this project. There are some tools for that. We have our Move tool,
our select tool. We have our Move tool, which is what we've been using
with a Space Bar. Have the comment tool. Maybe a friend comes in here and they think that the
couch should be read. They could leave a comment, and it will go here. So make red, save the comment, and now it goes here. I could resolve the comment, disagree or agree, or
leave another comment. You can also delete people's
comment if you want. Now if you're not sure how to get started with
all these notes, you can also just start
chatting in this box. So say you can write about the
final scene that you want, a flower pot where a
new plant is growing. You can write it here, and then ElevenLabs will guide
that production for you. If you have references
saved on your local device, you can attach them here, and you could send
that request in. Now, say this is a
workflow that I do very often for other projects
for my social media maybe. Instead of creating
this whole node layout, I could create a template. I click on this, I just have
to identify the inputs, outputs, and the publish format. So things that only
have something coming out the right
side are the inputs. So this guy had no inputs, therefore, it is its own input. So it's just a lone like
this. This is my input. I can select it, and it's
going to be marked as orange. And then I could do the
same thing with output, give it a name, and then I
could publish my template. So that's something
you could do as well. Here's my video. We have, you know, a seed sprouting. It looks kind of magical, but it is a red flower. We could then extend this
to a different node. So a video extender note, if we want to turn it from
4 seconds to 20 seconds, we could do dubbing, mix it with an audio,
use it as a reference. And so much more can be done. So go ahead and
try out the flows. It is a new tool, and it's really cool, actually. And you could really work
on larger production, combine different
tools as notes, and build a final
product where you use different things in
ElevenLabs to tell a story, to promote a product, do a podcast, or anything
else that you want to do.
38. What is ElevenAgents: So we looked at the creative
side of ElevenLabs. We looked at voice. We looked at image,
we looked at video, and so many other tools
on the left side. But they also introduce
something called ElevenAgents, and that's something
you can switch to by going down here. So just click on that
and we're going to go over to the next
interface, basically. So ElevenAgents are
basically AI agents that you can use to
automate a certain task, either integrating the tools
that ElevenLabs has or, you know, just keeping
it as a simple agent. So one of the
templates that you can look at is customer support. That's just an example
of an 11 agent. And on the side,
it actually tells you so much about the
workflow behind it, how it's working, and you
can also preview it here. So right now, I'm
going to pretend that, you know, I need support, and this is currently
taking in two integrations. So Saleforce and Zendas we have many different
options here. But the main feature
is basically the tools that you
already know how to use being integrated
in these workflows. So once we see an example, it's going to make more sense. Hey, this is Jamie from support. What can I help you with today? I need help with a
technical issue. I'd be happy to help you
with that technical issue. Could you tell me a bit more about what you're experiencing and maybe share your name so I can look into your account? My name is Alice and
my computer broke. I'm sorry to hear your
computer broke, Alice, but since I support our AI platforms APIs and integrations, I can only assist with
issues like API errors, integration problems or
performance concerns. So that's just an example. You can see, even
though I didn't really ask what the support is for, it still understood my concern, and it took in all those inputs. So the start was me pressing
on the call button. The first thing was to
identify the issue. It opened up warmly,
which it did. And then I just said that
I had a technical issue. Then we went here to propose
one concrete first step, and then you could either
resolve or escalate. Say, it says, Okay, you need to provide us with
a serial number, and then I agree to them. They're going to
resolve this now, and then it ends. If I said that I have an account problem and not so much a technical issue, it's going to take
a different route. So it's going to
look at billing with my name and then see if
it reached a solution, either go with the solution
it has or look further, and then end the call. So this whole map
is basically how this AI agent is going about
being a customer support. We have different things, too, something a little bit less technical is a language
practice tutor. So you can start again, and the first thing
it's going to do is access your level
for that language. If I'm beginner,
it goes this way. If I'm advanced,
it goes this way. Beginners, it's going
to have a certain mode. So maybe using certain dialects, steering away from
large vocabularies, speaking, slower, those things. And then the session ends. It gives us a summary. So lots of templates you could
look at, but, of course, we could always make our
own by creating an agent. So you can browse the templates, make adjustment or
create one from scratch. So we're going to
learn all about that. But just to give you a
little bit of a tour, we have the different categories on the left side to
help us get started. We have creating agents, which is where we
were configuration, the knowledge bases that it has, I could create and upload some documents here that
would be used for the agents. For example, if I
have a agent that is supposed to help customers
with broken laptops, that agent needs to know
a lot about laptops. So instead of typing
everything in, I could just have a
document uploaded here and then connect this
base to the agent. So it can go back to this, learn a little bit more about laptops and then get
back to that collar. Could be a URL, could you trust a website. You
can put it here. You can upload a file,
type something in, create a folder
with multiple files or sync documents
from other platforms. We have 21 megabytes
of storage here, which is perfectly it's
plenty for documents, and you could sort
and search as well. Of our different tools here
that you could add in. So any sort of integration
that you want to do. So I would have to add
a connection first. We have a lot of things that currently work with ElevenLabs, and all you have to do is sign
in, make that connection, and it will be
available for you to use in your agent
building process. So these are the
different categories, but we have all of
them listed here. You can search, and
once you again sign in, they're going to be listed
here and you just select them. So that's an integration tool. We have client tools. Client tools are things
that require an input. So, for example, the template
that we looked at where the agent was asking about the type of issue
that I was having, that is a client tool. So it's like technical issue
or just type of issue, wait for user response, and then do certain things. We can add parameters,
different variables. Even edit this as a JSOd file. We have webhook as well and then a list of other
integrations that you could add. Of course, we have our voices. This is 11 labs after all. These can be integrated as
the voices of your agents. To monitor, you can look at the conversations that are happening as you deploy agents, the users, and then the
different tests that you do. So when you make an agent, you do want to test it out, see if it works fine, and
those will all be logged here. Deploy, you are going
to have the following. So it could be deployed
with phone numbers to Whatsapps or even batch calling. So we have all the
levels that we need to deploy a successful AI agent. Now, we're first going
to start with our agent. We're going to build
one in the next lesson, integrate a few tools, test it out, and then deploy it. So you guys can
follow along with this simple agent that
we're going to build and that we are going to have
a better understanding of how this works and
how you can apply it to your daily work. O
39. Create First Agent: So let's build our first agent. This is going to be
a very simple agent, just to show you how it works, and then based off of that, you could build
more complex agents and integrate different things. So to create an agent, you click on this
button right here. And instead of
choosing templates, we're going to
create a blank one. Now, the templates, we
already looked at them, you can choose one that
is similar to what you need and simply switch out the things that
are irrelevant. We're going to start
with a blank agent. So first, let's give
our agent a name. Going to say
information provider, and we'll call him Ben. So you can choose
here for this to be a chat only agent
or an audio agent. So by default, it is audio
because this is ElevenLabs, but if you want to exclude that, you just keep it at chat only. So basically, it
would be similar to the chatbots that appear on
the side of the website. But the audio is going to be something like a call in a way, where it will wait for a user input speech
in order to operate. I will leave this off, and we're going to create Ben. So here's the new agent. We can see that
there's a bunch of things around that
we can now change. So on the right is the
preview of our agent. And because we said
we want audio, you can see that it has, like, a little phone call symbol, and I could talk with Ben right now and have him help
me with something. Now we do need to set Ben up because right now he's
just an avoid, basically. The first thing we need to do
is tell Ben what they are. So this could be you are
a customer support agent. You are a personal assistant
or something like that. Whatever you want to do,
you can put it right here. We also have our generated with AI option if you're not sure how to go about writing
the system prompt. But basically, if you
click on the little book, it's going to show you
the official guide. So you do need to set
up the environment, the personality, and different
aspects of this agent. So you can see tone
is one of them. The goal is one of them, and the way that this agent should interact and
process the inputs. So you could read the guide
if you want to learn more. But to get started, I'm just
going to have Ben act like an automated voice agent whenever someone
calls the business, and our imaginary
business is a hair salon. So I did make a knowledge
based document, and that is available for
you guys to download. And I'll show you where
you can upload that. But for now, we are
going to tell Ben Act. So let's say you're
a helpful assistant for a hair salon and people will call to ask about
business operations and hours. So very simple, his
job isn't that hard. I could now use the tool here
to enhance this further. So generate the AI. Let's actually copy
this and put him in and see what prompt ElevenLabs will give us that
is more suitable. So you can see that it followed the same guide have
personality, the environment. So who are they
interacting with, the tone of voice here, the goal, the whole purpose of this agent, and guard rails. So the only thing that I need to change is, you
know, the name. So you are Ben. And I forget what the name of
our hair salon is called. Here's the document that I
was telling you guys about. This is what I ended up
calling the hair salon, so I'm just going
to copy that in. The Locke hair salon or
Hair Studio, cut that out. And let's check the rest. He knows everything, and then you're interacting with colors, and I'm just going to replace
the name as much as I can. So that's the name
of our hair salon. I'll do Control F,
look for the old name, and just switch it out. This is Ben, not Stella. So just do a little cleaning up. Alright, so no more glamour. And I'm just gonna check for Stella in case there's
more, and there is not. Alright, so we have
everything set up for us. I didn't have to do
much, as you saw. Then we're going to decide
on the first message. So when the person calls Ben, what should be the response? So, hello, how can
I help you today? I could say hello, I'm Ben. How can I help you today? Now, this option allows the user to kind of
cut him off and, you know, say their
urgent message. But if you want him to go
on and say what he wants, you can just turn this off. So this is all good.
Next is the voice. So we have a lot of
voices to choose from. I'm going to go with Eric because it's one
of the defaults, and it's pretty smooth
for this purpose. We do have something
called expressive mode, which is going to be more, as they call it,
emotionally intelligent. So when it comes to
pauses to, like, making different sounds, it's
going to be less robotic. So if the user said, I fell down the stairs, they're not going
to be like, Oh, that's great. Let's do this. They're going to be more
expressive and sound more human. So that's something you
could add if you want, or you could completely
dismiss it. Up to you. So I will enable
it for my purpose, and I'll just close
this right now. You can see it's turned
on, and it's Eric. Next, we have the language. We can only have people
speak in English with Ben or add other languages.
That's up to you. And then we have
our language model. So there are a couple of
things to get started, I'm going to keep the default. But if you want to
build your own model, oh, that's something that
you can do right here. It is a little bit
more intensive, so I'll leave this
out for now just so we can get our
simple agent going. So that's the default
and then agent behavior. You can basically
set up behavior where it works
with another tool, and there is a routine that
this agent needs to follow. For example, we
all know WhatsApp. That's like a texting platform. So if someone texts them, they have to first
read the text and then reply to the text and also call. So these are things that
WhatsApp allows people to do. And therefore, that agent
needs to follow that behavior. This is going to be
different for fresh desk, so that's going to be
responding to tickets only. So each of these guys
are a little different. And if you choose
them, there's going to be additional options for them. But again, I'm going
to keep everything as default just so
we can get started. So turn that off, and we
have now made our agent. Let's go ahead and
add a knowledge base, which is where we create a space full of information
about our brand. So in my case, that's
the hair studio, and I need to provide
Ben with business hours, the type of services
so that when someone calls our
number and he picks up, he knows exactly how to
respond to the questions. Mmm.
40. Integrate: So, we made our agent
in the previous lesson. Let's go ahead and give
Ben some knowledge. We're just going to head
down to knowledge base, and we're going to
upload the document that I mentioned that is
available for you guys. This could also be a URL
or a different file. Mine is going to be a file, so let's just upload it now. So there's the file. We can keep it to
our current folder. We're just going
to add the file. So right now, whatever
you add here is going to the knowledge base for Ben because that's the
workspace that we are in. So if you want to
remove something, you could just click
detach from agent, and that's going to be
out of reach for Ben. But I do want to keep that, so we're just going to now use this integration to help Ben
with his job, basically. Apart from document as a base, you can integrate different
tools for your agents. So there's a bunch of things
that you could connect. I'm just going to go over
to tools to add a tool, you are going to need
something called API keys, which is the unique
key that you get, and you bring into this
platform to help the connection between ElevenLabs and whatever tool that you're
trying to connect. So, for example, if I go to Ad Tool Integrations
add a new connection, let's find something fairly
easy. I'll go with Exa. I could give this a
name, web search. I need an API key. So let's head over to EA
and get that for ourselves. You just search a API Key, you're going to be
directed to this link. Create APIKey and we're going
to just create an account. This will be the case for
any tool that you choose. So that's what I'm
going to do real quick and then we're going
to get a key for Bed. So I just signed
in using an email, and we're going to
create our account. I'll just skip this for now. And I'm just going to go down
to APIKeys and create one. So I'll call this
ElevenLabs. Create the key. Make sure you copy it, go back here, and we're just going to paste it in this area. Once you're done, you
can click on Connect, and now it's all good to go. So once we have this connection, we can decide on the tool and the service that
we want EA to give us. And for my case,
that is a search. So click on search
and click Create. So once we're done in tools, we could just click on EA and basically say how Ben
should use this tool. So the condition prompt is the following.
This is the default. So we can say choose content type, use
highlights, recommended. I think this is fine
for now because the user could ask
anything from Ben, and I don't want to
restrict it here. So I'm just going to
let this be as it is. And we have now made our
agent an integrated a tool. In the next lesson,
we're going to try Ben out and see what
sort of changes we can make and then
how we can publish it onto somewhere else where
other people can use it.
41. Testing: M So Ben is all set up and we now have to test to see how
well he performs. So in the same page for agents, you can basically
start talking to Ben. You can see on the right side, we have the little basically caller
interface where I get to just speak into the microphone and see how Ben responds to me. So let's start a call. Hello, I'm Ben. How can I help you today? Hello. When is the
business open? We are closed on Mondays, but open from 9:00 A.M. To 6:00 P.M. On Tuesdays and Wednesdays, 9:00 A.M. To 8:00 P.M. On
Thursdays and Fridays. 8:00 A.M. To 5:00
P.M. On Saturdays and 10:00 A.M. To
3:00 P.M. On Sundays. Does that help you
plan your visit, or is there anything else
I can assist you with? Yes. Can you tell me
a little bit about the latest hair trends that I could tell my stylist
at the business? I don't have specific
information on the latest hair trends
in my current resources, but our stylists are very
knowledgeable and can definitely discuss the newest styles with you when you arrive. Would you like to go ahead and
book an appointment so you can chat with him about
what's popular right now? No. Can you do a web
search about it? I'm unable to perform
web searches, but I'd be happy to
connect you with one of our stylists who can
share the latest trends. Would you like me to transfer you to someone who
can help with that? So I just ended the
call, but so far, when it came to talking
about the business, Ben was able to help us out. But when it came
to the web search, he still didn't have the
permission to do so. So what we're going
to do is go back to the system prompt and include
the web search aspect. So I'm just gonna go here
and find out about the goal. So welcome. It's fine. Okay, so we have appointments,
location, directions, general inquiries, and then
we could add another one. I'll just copy the same line just so that the
format is the same. We're going to say for
web search inquiries, answer questions using A for websearch and provide
user with the highlights. So let's try this again
and see what happens. I'm gonna close this right now and then start
another conversation. Let's actually also change his first message just so
we can see the update. So how can I help you today? Do you want to style your hair? Alright? Let's try this out. Hello, I'm Ben. How
can I help you today? Do you want to style your hair? Yes, tell me more about the business services and when I could come in
to get my hair done. We offer a wonderful range of services, including women's, men's and kids haircuts
blow I don't want any kids, just for adult female. Oh, absolutely. We have fantastic services
specifically for adult women, such as our signature
women's haircut, blowouts and various
coloring options like Baliage or root touch ups. You can visit us Tuesday through Wednesday from 9:00
A.M. To 6:00 P.M. Thursday and Friday until
8:00 P.M. Saturday from 8:00 A.M. To 5:00 P.M. Or Sunday
from 10:00 A.M. To 3:00 P.M. As we are closed on Mondays. Does that sound like a good
time for you to come in? Yes, but do a web search about the latest hair trends so I could come in and tell
my stylists about it. That is a fantastic idea. The latest trends for late
2026 are all about polished, healthy looking hair
with a natural feel. Popular styles include
sleek blowouts, modern chignons and
the boyfriend Bob, while colors are
shifting toward rich, glossy brunettes, warm cognac
brond and soft coppers. Would you like to book an
appointment to discuss trying one of these looks
with one of our stylists? Yes, book me an appointment. I would be happy to help
you get that scheduled. To find the perfect time, could you please
tell me your name, which service you are interested in and what day or time
works best for you. Also, do you have
a preference for a specific stylist or would you like to see
whoever is available? Actually, I changed my mind and I don't want an appointment. No problem at all. I completely understand that
plans can change. If you ever decide you'd
like to book an appointment later or have any other
questions about our services, please feel free
to call us back. Is there anything else I
can help you with today? No, that is all. You are very welcome. Have a wonderful day, and we look forward to seeing you at Lux Hair Studio whenever
you are ready. Goodbye. So as you can see, that was a very natural way
of speaking for Ben, and now it was able to
look onto the web via the tool that we connected and provide me some information. And I did interrupt
Ben at the start, and that is because we
have this turned on. Again, you can turn
it off if you want your agent to just say
their sentences fully. And the back and
forth between me and the agent also shows
the power of 11 agents. You can see that I went
to make an appointment, but then I changed my mind, and been was able to use the correct tags to
respond to me as a user. So understanding, friendly, and here for goodbye,
went for warm. And then when I finally said yes to the appointment,
it was cheerful. So these are all signs that this conversation
went very well, and we set up Ben correctly. So we tested Ben, and we made some adjustments. Also do more intensive tests. So that is where you
set up a template, and you have that go on for repeat until a problem is found. For example, I could set
up a test where Ben is called five times in a minute, and I could see
how well he jumps back and forth with
the five callers. But that is something a
little bit more intensive, and for now I want to shift the focus towards
publishing this new agent.
42. Publishing : M So Ben is all set up and we now have to test to see how
well he performs. So in the same page for agents, you can basically
start talking to Ben. You can see on the right side, we have the little basically caller
interface where I get to just speak into the microphone and see how Ben responds to me. So let's start a call. Hello, I'm Ben. How can I help you today? Hello. When is the
business open? We are closed on Mondays, but open from 9:00 A.M. To 6:00 P.M. On Tuesdays and Wednesdays, 9:00 A.M. To 8:00 P.M. On
Thursdays and Fridays. 8:00 A.M. To 5:00
P.M. On Saturdays and 10:00 A.M. To
3:00 P.M. On Sundays. Does that help you
plan your visit, or is there anything else
I can assist you with? Yes. Can you tell me
a little bit about the latest hair trends that I could tell my stylist
at the business? I don't have specific
information on the latest hair trends
in my current resources, but our stylists are very
knowledgeable and can definitely discuss the newest styles with you when you arrive. Would you like to go ahead and
book an appointment so you can chat with him about
what's popular right now? No. Can you do a web
search about it? I'm unable to perform
web searches, but I'd be happy to
connect you with one of our stylists who can
share the latest trends. Would you like me to transfer you to someone who
can help with that? So I just ended the
call, but so far, when it came to talking
about the business, Ben was able to help us out. But when it came
to the web search, he still didn't have the
permission to do so. So what we're going
to do is go back to the system prompt and include
the web search aspect. So I'm just going to go here
and find out about the goal. So welcome. It's fine. Okay, so we have appointments, location, directions,
general inquiries. And then we could
add another one. I'll just copy the same line just so that the
format is the same. We're going to say for
web search inquiries, answer questions using A for web search and provide
user with the highlights. So let's try this again
and see what happens. I'm gonna close this right now and then start
another conversation. Let's actually also change his first message just so
we can see the update. So how can I help you today? Do you want to style your hair? Alright? Let's try this out. Hello, I'm Ben. How
can I help you today? Do you want to style your hair? Yes, tell me more about the business services and when I could come in
to get my hair done. We offer a wonderful range of services, including women's, men's and kids haircuts
blow I don't want any kids, just for adult female. Oh, absolutely. We have fantastic services
specifically for adult women, such as our signature
women's haircut, blowouts and various
coloring options like Baliage or root touch ups. You can visit us Tuesday through Wednesday from 9:00
A.M. To 6:00 P.M. Thursday and Friday until
8:00 P.M. Saturday from 8:00 A.M. To 5:00 P.M. Or Sunday
from 10:00 A.M. To 3:00 P.M. As we are closed on Mondays. Does that sound like a good
time for you to come in? Yes, but do a web search about the latest hair trends so I could come in and tell
my stylists about it. That is a fantastic idea. The latest trends for late
2026 are all about polished, healthy looking hair
with a natural feel. Popular styles include
sleek blowouts, modern chignons and
the boyfriend Bob, while colors are
shifting toward rich, glossy brunettes, warm cognac
brond and soft coppers. Would you like to book an
appointment to discuss trying one of these looks
with one of our stylists? Yes, book me an appointment. I would be happy to help
you get that scheduled. To find the perfect time, could you please
tell me your name, which service you are interested in and what day or time
works best for you. Also, do you have
a preference for a specific stylist or would you like to see
whoever is available? Actually, I changed my mind and I don't want an appointment. No problem at all. I completely understand that
plans can change. If you ever decide you'd
like to book an appointment later or have any other
questions about our services, please feel free
to call us back. Is there anything else I
can help you with today? No, that is all. You are very welcome. Have a wonderful day, and we look forward to seeing you at Lux Hair Studio whenever
you are ready. Goodbye. So as you can see, that was a very natural way
of speaking for Ben, and now he was able to
look onto the web via the tool that we connected and provide me some information. And I did interrupt
Ben at the start, and that is because we
have this turned on. Again, you can turn
it off if you want your agent to just say
their sentences fully. And the back and
forth between me and the agent also shows
the power of 11 agents. You can see that I went
to make an appointment, but then I changed my mind, and been was able to use the correct tags to
respond to me as a user. So understanding, friendly, and here for goodbye,
went for warm. And then when I finally said yes to the appointment,
it was cheerful. So these are all signs that this conversation
went very well, and we set up Ben correctly. So we tested Ben, and we made some adjustments. Can also do more
intensive tests. So that is where you set
up a template and you have that go on for repeat
until a problem is found. For example, I could set
up a test where Ben is called five times in a minute, and I could see
how well he jumps back and forth with
the five callers. But that is something
a little bit more intensive, and for now, I want to shift the focus towards publishing
this new agent.
43. Other AI Tools to Use with ElevenLabs
: Now that we know
how to fully use ElevenLabs to make
different sorts of audios, ranging from a voiceover
to an audio book, we're now ready to create audios in the best way possible. But you can also take
your skills further by combining different AI
tools with ElevenLabs. So ElevenLabs is mainly
known for audio. As we saw, it's really good and it's constantly improving. But how can you combine the other very good AI tools to make some amazing projects? So I'm going to recommend some tools that you
may or may not know already and tell
you how you could combine it with ElevenLabs. The first tool that we already saw how to use was Chat GPT. I'm sure you're mostly
familiar with this. But what you can do with
it is first of all, turn A URL into a summarized version so that you could use it for a
podcast or an audio book. As we also saw,
you're able to change a language and to add in few elements that can enhance the speech that 11 Lab creates. Which at GPT, you can also
start things from scratch, such as writing a
book about apples or write me a script for a podcast talking
about tornadoes. In terms of scripting
and brainstorming, this platform is a
great place to come to *** not only free, but it's pretty good
at what it does. Now, you have your script, you have your audio.
What about footage? One place that you could get
images from is Mid journey. This is another AI platform. Nowadays, they also do
short snippets of videos. I will talk about
videos in a bit, but if you wanted to do
voiceover over images, then this is the place to come. This is not a free tool. I believe even if you do want
to use the free version, there's a lot of limitations. But as you can see, these are some samples that you can
make with this y tool. Once you're done
making them here, you can export them
into ElevenLabs, do your voiceovers and make
some content with them. Now, what if you
want to make videos? Runway AI is a great
platform for videos. There are others out
there, but this one is probably the most
easiest to use. The way this works is that just like any other AI platform, you give it a prompt,
as you can see here. You can play around with
angles, with the mood, the characters, keep
characters stable, and just do a bunch
of different things. This, once again, if you're
using the free version, there are some limits. You could upgrade to
a better version, but if you don't plan on
doing this that often, you don't really need
to buy a full account. ChaGPT also does
generate images for you. If you didn't want to pay
for Mid Journey or DaVinci, you can just come back
to Chat GIP Team. This way, you're combining
scripting, imaging, and video capabilities with
your already made audios. This way, you have the option
to generate different of content and not fully rely
on audio only content. Now that we have all of
these tools at our disposal, let's talk about how you can monetize these content
if you choose to. I will see you guys
in the next lesson.
44. Monetization and Next Steps: There are a lot of ways
that you can monetize with ElevenLabs just as you
can with other AI tools. So apart from being able to
make your YouTube channel and monetize through that platform by making content
with ElevenLabs, you can also join in on
ElevenLabs own Creator program. So when we go to
the Voices Library, there are some that
charge you per minute. And you can see
there are not that many, but they do exist. Essentially, what you
can do when you make your own voice is
that you would be getting paid passively
after you have uploaded a high quality
audio on 11 labs. For example, this
audio is I go over it, $0.02 for me to use. Dude, if riding waves
is my religion, then tacos are straight
up communion, bro. So that's an example. You can
see it's very high quality, very clear, and it has the perfect voice
for this category. So that's one way that you can monetize
ElevenLabs creations. On that note, they do have
a carriers page as well, where you get to be one
of those specialists that you can use your
services from in production. Recall when we went over
here, we were getting, if we were to use it, charge $2 per minute and
all these pricing. These would then be made by
real humans as we discussed. You could get qualified to
be one of those people. Simply by going into
the slash careers page, then open positions, you can see how much stuff
they have available. Now, those are full time jobs, but if you go in the
transcription slash subtitling, you can see that this
is a freelance thing. So whenever there's a
demand for your service, you're going to get
reached out to and then record an audio with
your voice for someone else. You can see it's even remote, and this is all
within this platform. You should definitely
check it out if you plan on making money
with ElevenLabs. Other way that you could do
it is make content online, such as YouTube videos,
Instagram videos, Tik Tok, or whatever,
using ElevenLabs as audio. This will save the
time for you to record everything with your
own voice and you get to use these really quick and
accurate audio generations to build your content. For example, you could cover scary natural disasters like we did with one of our projects. You can do a biography overview where you tell the
story of famous people. There are so many things
out there that you can do. But the role of ElevenLabs in that is that it's
just the audio. You still have to
work really hard to advertise your content, get it out there, have
people watching it, be engaged, and
all of that stuff. Another thing that you could do is freelance with ElevenLabs. Basically, you could go on
platforms like Fiber, Upwork, and there's so many out
there nowadays and offer the service of making
AI audios for a client. If someone doesn't want
to use their own audio, they can come to you
so that you could accurately make them and audio. The same way we altered the
voices for a scary podcast, a advertisement, a
regular podcast, you would be doing
the same thing, but based on the client's need. Those are just some
really quick ways that you can monetize
with 11 labs. Just like any other AI tools, there's going to
be a lot of demand for this type of
content and work. Keep an eye out and make
sure that you're in the right communities
so that you're aware of all of the updates and all of the cool tricks that you
can do with this platform. Redit has a lot of them, and then you can go
to maybe Facebook, Discord and just
keep on practicing making audios and learn new
ways to make them better. I hope you guys enjoy this
course and that you continue using this really cool
program for your own content. Hope to see you guys soon.
45. Class Project: Create Your Own Audio/Visual Content: Now it's your turn to take
what you have learned and create your own project
using ElevenLabs. For your class project, I would like you guys to
create either an audio or a visual content using all the tools
that we have learned. You can also combine
it to where you have the visuals made and then combined with the audio
that you made. You plenty of freedom here. You can create a short
video, an audiobook, a podcast, lip sync video, and anything else
that you want to do. Start by deciding what you
want to make and who it's for. Then think about which of the many tools can you use
to create this project. You can choose or design
an appropriate voice for your project and see how that could fit
with your idea. Could just be an audio project. But if you want to take
it a step further, you can add some visual content made from the creative tools. We learned how to
generate images, videos, and how we can make a full
production using flows. So choose the tools that
genuinely help with your idea. When you're finished,
you can upload the project to the
class project library. Depending on what you have made, you can include
the full content, screenshots of it,
and alongside that, you can provide an explanation talking about your process, your idea, some
challenges you face, and what you were curious about. And remember, when using
AI to generate voices, you want to do this responsibly. Especially if you're cloning other voices or you're creating content
based on real people. Looking forward to seeing
what you guys create.
46. Congratulations! What’s Next?: Mm hm. Congratulations on finishing the ElevenLabs
AI master class. We have covered a lot in
this course, but by now, you know that 11 labs is so much more than
just text-to-speech. The important thing now is to keep experimenting
with the tools. You don't have to use every
single tool in your project, but you can see which
tools would help improve your workflow or help you
create better content. Or even larger workflows, you can combine 11 labs with
other AI tools out there. This could be for whether you're working for a business with a single client or for
your own personal project. If you haven't already,
be sure to upload your project to the
class Project library. I would love to see how
you guys use the tools and how you combine them
to generate your ideas. And if you enjoy the course, feel free to leave us
a review as it really helps us improve our content.
Thank you for joining me. I hope this course have left
you with a lot of ideas on how you could use ElevenLabs
for your own projects. Keep creating, and I'll see
you guys in our next class.