Speech To Text

PollDo you want Speech to Text Wrapper/Source Code?
7 votes
Posts 16–30 of 30 · Page 2 of 2
Quote Originally Posted by Hell_Demon View Post
How extremely godlike would it be to shout "AIMBOT ON" ingame and instantly dominating everyone cause of the aimbot?! :O
You never fail to make me laugh. It would be so funny if you turned on your mic to talk to the others and start being like: "Chams on", "Aimbot Head", "Aimangle 90", "Aimspeed Legit". I would buy that VIP hack and rape EVERYTHING. I'd use the command "Rape mode activate"
Quote Originally Posted by fukojr2 View Post

I'd use the command "Rape mode activate"
And then an extremely sexy computer voice comes on and says: "Rape Mode Activated"


... then I guess people would get raped. That be sweeEEEeeeeeet! =)
Quote Originally Posted by why06 View Post


And then an extremely sexy computer voice comes on and says: "Rape Mode Activated"


... then I guess people would get raped. That be sweeEEEeeeeeet! =)
I like the way u think Shaun.
I really don't see the big deal in this?

SAPI is a suit by microsoft with all this built in.
Technically with the "Robot" stuff as well.

Example: Using Microsoft Dictation Component, You can have the application "Listen" on command & stop accepting commands other then listen, meaning You can have it respond or ignore voices (which comes in handy) , You can also using dictation implement (perhaps) 5 lines of code and build a function the will listen for a specific list of commands & then perform the actions, I have been doing this since VB5 (as some of you know from when I made the post about the proximity sensors) .

Example>

List of commands
You Speak: Listen
Bot uses TTS and responds: "what would you like me to do"
You Speak: Search Google: MPGH
Bot speaks: Results are in.
You Speak: Read Results
Bot lists results
You speak : Select 4th

or whatever. But the best part, is it's already developed for you with the SAPI, which the OP is using, so the OP will code a wrapper using SAPI, Sounds redundant and pointless to me.

No offense. (btw)

Proof : Simple Dictation (SAPI 5.3)

Any of that look hard? So why not just use SAPI , why create a SAPI on a SAPI that does the exact same thing as the aforementioned SAPI
What would be cooler is a web based voice recognition software like google has, you link to google and they provide you with a text interpretation of the sound byte. Next the text question is inputed to another AI, like Watson, which will find an answer to your question. Or if it's not a question, but a command such as something you want to do on your computer, another external Watson can translate this into commands for your computer up to a point.

Say you said: "I want to move all my pictures to a new folder on my desktop"

The Watson can extract information and turn it into a command: "Move [user created ][TAG:PICTURES][possible multiple file location] to [new][desktop folder]"

Lastly a Personaly AI, which monitors your file activity will have tagged each file with a tag based on its usage, so if it sees other files being used in the same way pictures are it will add that file to a tag humans would define as {PICTURES} though the computer only thinks of it as {TAG:0x000053BA} or something. because certain file types such as [jpeg, png, gif, etc. are considered images or pictures, the personal AI can find all tags similar to pictures without understanding what pictures are. If it could find no tags for pictures, an online thesaurus might be used to find other tags to search.

Finally the Personal AI will come up with a list of resulting actions that best exemplifies the previous command.

So we have multiple levels of AI, working at different levels to assist each other. Alone no single AI could give clear and accurate results, but with multi-tiered approach it creates a transparently unitarily smart system to the user.
Quote Originally Posted by why06 View Post
What would be cooler is a web based voice recognition software like google has, you link to google and they provide you with a text interpretation of the sound byte. Next the text question is inputed to another AI, like Watson, which will find an answer to your question. Or if it's not a question, but a command such as something you want to do on your computer, another external Watson can translate this into commands for your computer up to a point.

Say you said: "I want to move all my pictures to a new folder on my desktop"

The Watson can extract information and turn it into a command: "Move [user created ][TAG:PICTURES][possible multiple file location] to [new][desktop folder]"

Lastly a Personaly AI, which monitors your file activity will have tagged each file with a tag based on its usage, so if it sees other files being used in the same way pictures are it will add that file to a tag humans would define as {PICTURES} though the computer only thinks of it as {TAG:0x000053BA} or something. because certain file types such as [jpeg, png, gif, etc. are considered images or pictures, the personal AI can find all tags similar to pictures without understanding what pictures are. If it could find no tags for pictures, an online thesaurus might be used to find other tags to search.

Finally the Personal AI will come up with a list of resulting actions that best exemplifies the previous command.

So we have multiple levels of AI, working at different levels to assist each other. Alone no single AI could give clear and accurate results, but with multi-tiered approach it creates a transparently unitarily smart system to the user.

You work on that and tell us how it goes ok?
Quote Originally Posted by NextGen1 View Post
I really don't see the big deal in this?

SAPI is a suit by microsoft with all this built in.
Technically with the "Robot" stuff as well.

Example: Using Microsoft Dictation Component, You can have the application "Listen" on command & stop accepting commands other then listen, meaning You can have it respond or ignore voices (which comes in handy) , You can also using dictation implement (perhaps) 5 lines of code and build a function the will listen for a specific list of commands & then perform the actions, I have been doing this since VB5 (as some of you know from when I made the post about the proximity sensors) .

Example>

List of commands
You Speak: Listen
Bot uses TTS and responds: "what would you like me to do"
You Speak: Search Google: MPGH
Bot speaks: Results are in.
You Speak: Read Results
Bot lists results
You speak : Select 4th

or whatever. But the best part, is it's already developed for you with the SAPI, which the OP is using, so the OP will code a wrapper using SAPI, Sounds redundant and pointless to me.

No offense. (btw)

Proof : Simple Dictation (SAPI 5.3)

Any of that look hard? So why not just use SAPI , why create a SAPI on a SAPI that does the exact same thing as the aforementioned SAPI
MS SAPI is extremely crap. The accuracy is way off... (Compared to the Nuance line of products, that is)

Quote Originally Posted by fukojr2 View Post



You work on that and tell us how it goes ok?
Totally. Sure. Of course...
Quote Originally Posted by freedompeace View Post

MS SAPI is extremely crap. The accuracy is way off... (Compared to the Nuance line of products, that is)



Totally. Sure. Of course...
I'm making a rather ignorant comment when I say this (As I don't do audio related programming)

I don't beleive MS SAPI is supposed to be perfect out of the box; it provides a framework, you'll need to do some fine-tuing to get it working they way you need it. A free base which has the potential to be powerful with a bit of work is, IMHO, probably a lot better than getting an expensive software license.

Again, never worked with the API; just my experience working with SDKs.
Quote Originally Posted by radnomguywfq3 View Post
I'm making a rather ignorant comment when I say this (As I don't do audio related programming)

I don't beleive MS SAPI is supposed to be perfect out of the box; it provides a framework, you'll need to do some fine-tuing to get it working they way you need it. A free base which has the potential to be powerful with a bit of work is, IMHO, probably a lot better than getting an expensive software license.

Again, never worked with the API; just my experience working with SDKs.
Just to make sure, does everyone know that im using the SAPI and upgrading it? Just like what Jetamay said
Meow...?

This is turning into a somewhat heated debate.
I'm making a rather ignorant comment when I say this (As I don't do audio related programming)

I don't beleive MS SAPI is supposed to be perfect out of the box; it provides a framework, you'll need to do some fine-tuing to get it working they way you need it. A free base which has the potential to be powerful with a bit of work is, IMHO, probably a lot better than getting an expensive software license.

Again, never worked with the API; just my experience working with SDKs.
Yeah, it's not perfect, which is why the accuracy is off for freedom.
Really, all Dictation software / API / etc is going to require training (although I found the google search on my tab and android devices are surprisingly accurate)

@aanthonyz , well that's my point, your not building a custom wrapper at all, and I am curious what will be upgraded that isn't already a supported feature.
Quote Originally Posted by NextGen1 View Post


Yeah, it's not perfect, which is why the accuracy is off for freedom.
Really, all Dictation software / API / etc is going to require training (although I found the google search on my tab and android devices are surprisingly accurate)

@aanthonyz , well that's my point, your not building a custom wrapper at all, and I am curious what will be upgraded that isn't already a supported feature.
Basically after I read everything in this thread, I have become so confused.
Im just going to post what I did for my ROBOT and you can decide.

Good luck mate, sounds interesting.
Quote Originally Posted by Hassan View Post


Dragon Naturally Speaking

^^ I used it for quite a long time. It's accuracy is amazing ;]
You have to pay for that dont you ? :S
Lol...epic bump...
Posts 16–30 of 30 · Page 2 of 2
This thread is closed for replies.

Tags for this Thread

None

Need help?