Previous Post

We need your help choosing the voices for the characters.

Next Post
no preview
DESCRIPTION

Hello everybody!

We need your help choosing the voices for the characters. 

At first, we were going to release it without sound at all, or with text like in the comics... but at the last moment, we realized that it would be a huge mistake, and to do it later with sound, we would have to redo 95% of the work anew. After consulting, we decided to make a version with sound. Because this allows you to more accurately calculate the speed of events in a particular scene, the speed of camera movement, and the like. Not to mention the movements of the characters' lips to match the words. As you probably know, our team mostly does not speak English well enough. And the seemingly simple lipsinc operation has become a problem, more than we could have imagined. We had to ask an announcer with good English to record the movements of the lips and mouth so that we could start filling out the lipsinc library. Try 4 different technologies and 6 different plug-ins for this. Neural networks are already able to do this with the lips of live human actors and it is used in many films... But it's expensive and works well mostly only with people. Furry characters... There are a lot of questions and a lot of defects. Therefore, for now we will have to do everything using old-fashioned technologies in a Blender.

The second part of the problem is that the easiest way would be to hire live actors, but it's expensive. It is more expensive than our monthly budget now for a small part of their work. Considering the huge amount of work we have done as a team... this did not seem to us to be a fair and good decision in proportion to our work. Nevertheless, we tried a lot and settled on a variant in which we will have 1-2 actors (male and female voices), and with the help of neural networks we will be able to change these voices to a large extent so that each character's voice is unique. It took us almost two months of all this fuss, but we figured out how it works and trained the neural network for hundreds of different voices. There were 600 of them in total, and we pre-selected several dozen of them. They chose the sound quality, since not all voices can be copied well without a lengthy retraining of the neural network. Not everyone has clean enough material for learning - without noise, music, and the like. Not everyone has enough different expressions and intonations. In general, we have reduced the list of 600 votes to less than a hundred of the most successful options.

1.5 GB of audio

https://dropmefiles.com/QpXPO

copy this list to notepad

Rose

Tilda

Jay

Fat Rabbit

Glass Rabbit

Houligan Rabbi1

Random Rabbit 1

and write 3-5 names of your favorite files from our sample that you consider most suitable for this character. If everything is "not suitable", write it down, and we will consider these estimates.

as a result, you should get something like:

test.txt

Rose 237w, 11R, 239w_pch85neploh

Tilda 15T

Jay 27JJ-Jay -6, 10J

Fat Rabbit 77z-5oct,

Send this file to our email. [email protected]

if there are no suitable votes, leave the place empty. If there are many suitable votes, specify all of them, as many as three as possible. Specify them in the order in which you consider these voices to be suitable. That is, move the most suitable voice to the first place, and the slightly less suitable one to the second. After the third place, the order in which the votes go is no longer important.

Your choice now will determine which voices the characters will have. They will not change, even if the speaker's voice changes, because it is specially trained neural networks that make all these voices. From the same voice of our announcer, a girl.

If you find that the voices are not suitable, we will look for more.

So, we are looking forward to hearing your verdicts on the sound samples.

I think about 3-4 days is quite enough to gather quite a lot of opinions on this, right? After that, we will finalize the video based on your choice.

You can also add any advice, suggestions, or criticism to the end of the letter. The more opinions, the better. We really want to do everything well enough so that we don't have to redo the "early pages" anymore. So that this whole project has a single style and is good enough for any segment. That's why there's so much fuss and that's why the approach is so thorough.

Oh, yeah. another important question. there is a "narrator's text" in this sample. I'm not sure if everyone will like this idea, maybe it's better without him. In this case, the video will be a little shorter (and perhaps noticeably shorter), more concentrated. Because it will be more difficult to show more subtle nuances and the video will become (in my subjective opinion) a little more boring in a number of places. It's a bit difficult to do a lot of translations and text, sound, etc. under the "narrator's voice", but when we tried to do it with only simple dialogues... It seemed depressingly boring and strange. We will show you both versions of the video when it is fully ready so that you can make the final choice. But we would also like to hear the preliminary pros and cons

P.S.

2. If you have an expressive voice, like to "play with intonation", and are ready to spend several hours a month (maybe even half a day) participating in this project, we would be happy to talk. So far, we will not be able to offer payment worthy of good actors and interesting to residents of rich countries - you can see that our funding is now literally microscopic. Nevertheless, we can agree on a certain %, which may become quite a significant figure in the future when we start moving forward and making ads. We will also be able to offer a free subscription and a number of other bonuses. It's okay if you've never worked with sound, applications in this topic, etc. We will easily and quickly teach you all this and show you everything. All you need is a decent quality microphone (it usually costs from 50-60 dollars on different Internet sites, that's enough) and a little patience. And of course, native English with an expressive voice. We need someone who understands this fetish and all the nuances in it. We need a male and a female voice. We will be able to make many different-sounding versions of them (you can see them in the samples now). All these voices are artificial, but one of them is the real voice of a living person. And that's what all the other voices are made of. Nevertheless, our speaker, as I said above, is not a native speaker of English, and perhaps you will notice this. If no one notices and everything is fine, then this removes the problem of finding speakers for us. But if the accent is noticeable, the problem is relevant.

And yet, it's easier to make variants of male voices out of a male voice. And the female ones are female. Perhaps we are just being overly careful about this issue, and if you find the male voice options in our samples acceptable, then this will remove a very significant part of the problem for us.

In any case, we will try to thank you well for your time later, but at the start our resources are extremely scarce and we have little to offer now.

If you are interested in participating in this, please send us a sample of your acting. An audio file in any format that is convenient for you. Demonstrate expressive reading in it for a short period (well, you need at least about 3-5 minutes, preferably a little more). First of all, the expressiveness of speech and acting are important. Everything else is not critical. We will show you everything and teach you everything if you want to work professionally with sound later. It's not difficult. Only expressiveness is important. And native English.

By sending us an excerpt, you should understand that we will listen to it within our team and show it to a couple of English-speaking people so that they appreciate your acting talents. But it won't be used anywhere else. Basically, in the future, we plan to use such audio exclusively for neural network processing (so that one of your voices can sound with different timbres, like the voices of completely different characters). The voice is needed solely to make the neural network sound like a living person. Your voice will be like a puppeteer's hand inserted into a doll. The neural network will sound, and your voice will control its intonations. But it will be an additional advantage if you can make any artistic sounds implied in this fetish. This is especially true of the female voice ;) Well, you know. We have to pay for soda and cake expenses, which are the props that contribute to this;)

And the most important thing is, of course, the preference for young voices. It is more difficult to make an 18-year-old voice out of an adult's voice of 30-40 years old, although we can try, but it would be easier if the voice initially corresponded to the age of the characters at least minimally. However, the "narrator's voice" can probably be of any age.

vorecomicsclub PATREON 61 favs
VIEWS1
FILES1 file
POSTEDNov 16, 2025
ARCHIVEDNov 16, 2025