You never fail to make me laugh. It would be so funny if you turned on your mic to talk to the others and start being like: "Chams on", "Aimbot Head", "Aimangle 90", "Aimspeed Legit". I would buy that VIP hack and rape EVERYTHING. I'd use the command "Rape mode activate"
Speech To Text
PollDo you want Speech to Text Wrapper/Source Code?
7 votes
I really don't see the big deal in this?
SAPI is a suit by microsoft with all this built in.
Technically with the "Robot" stuff as well.
Example: Using Microsoft Dictation Component, You can have the application "Listen" on command & stop accepting commands other then listen, meaning You can have it respond or ignore voices (which comes in handy) , You can also using dictation implement (perhaps) 5 lines of code and build a function the will listen for a specific list of commands & then perform the actions, I have been doing this since VB5 (as some of you know from when I made the post about the proximity sensors) .
Example>
List of commands
You Speak: Listen
Bot uses TTS and responds: "what would you like me to do"
You Speak: Search Google: MPGH
Bot speaks: Results are in.
You Speak: Read Results
Bot lists results
You speak : Select 4th
or whatever. But the best part, is it's already developed for you with the SAPI, which the OP is using, so the OP will code a wrapper using SAPI, Sounds redundant and pointless to me.
No offense. (btw)
Proof : Simple Dictation (SAPI 5.3)
Any of that look hard? So why not just use SAPI , why create a SAPI on a SAPI that does the exact same thing as the aforementioned SAPI
SAPI is a suit by microsoft with all this built in.
Technically with the "Robot" stuff as well.
Example: Using Microsoft Dictation Component, You can have the application "Listen" on command & stop accepting commands other then listen, meaning You can have it respond or ignore voices (which comes in handy) , You can also using dictation implement (perhaps) 5 lines of code and build a function the will listen for a specific list of commands & then perform the actions, I have been doing this since VB5 (as some of you know from when I made the post about the proximity sensors) .
Example>
List of commands
You Speak: Listen
Bot uses TTS and responds: "what would you like me to do"
You Speak: Search Google: MPGH
Bot speaks: Results are in.
You Speak: Read Results
Bot lists results
You speak : Select 4th
or whatever. But the best part, is it's already developed for you with the SAPI, which the OP is using, so the OP will code a wrapper using SAPI, Sounds redundant and pointless to me.
No offense. (btw)
Proof : Simple Dictation (SAPI 5.3)
Any of that look hard? So why not just use SAPI , why create a SAPI on a SAPI that does the exact same thing as the aforementioned SAPI
What would be cooler is a web based voice recognition software like google has, you link to google and they provide you with a text interpretation of the sound byte. Next the text question is inputed to another AI, like Watson, which will find an answer to your question. Or if it's not a question, but a command such as something you want to do on your computer, another external Watson can translate this into commands for your computer up to a point.
Say you said: "I want to move all my pictures to a new folder on my desktop"
The Watson can extract information and turn it into a command: "Move [user created ][TAG:PICTURES][possible multiple file location] to [new][desktop folder]"
Lastly a Personaly AI, which monitors your file activity will have tagged each file with a tag based on its usage, so if it sees other files being used in the same way pictures are it will add that file to a tag humans would define as {PICTURES} though the computer only thinks of it as {TAG:0x000053BA} or something. because certain file types such as [jpeg, png, gif, etc. are considered images or pictures, the personal AI can find all tags similar to pictures without understanding what pictures are. If it could find no tags for pictures, an online thesaurus might be used to find other tags to search.
Finally the Personal AI will come up with a list of resulting actions that best exemplifies the previous command.
So we have multiple levels of AI, working at different levels to assist each other. Alone no single AI could give clear and accurate results, but with multi-tiered approach it creates a transparently unitarily smart system to the user.
Say you said: "I want to move all my pictures to a new folder on my desktop"
The Watson can extract information and turn it into a command: "Move [user created ][TAG:PICTURES][possible multiple file location] to [new][desktop folder]"
Lastly a Personaly AI, which monitors your file activity will have tagged each file with a tag based on its usage, so if it sees other files being used in the same way pictures are it will add that file to a tag humans would define as {PICTURES} though the computer only thinks of it as {TAG:0x000053BA} or something. because certain file types such as [jpeg, png, gif, etc. are considered images or pictures, the personal AI can find all tags similar to pictures without understanding what pictures are. If it could find no tags for pictures, an online thesaurus might be used to find other tags to search.
Finally the Personal AI will come up with a list of resulting actions that best exemplifies the previous command.
So we have multiple levels of AI, working at different levels to assist each other. Alone no single AI could give clear and accurate results, but with multi-tiered approach it creates a transparently unitarily smart system to the user.
I'm making a rather ignorant comment when I say this (As I don't do audio related programming)
I don't beleive MS SAPI is supposed to be perfect out of the box; it provides a framework, you'll need to do some fine-tuing to get it working they way you need it. A free base which has the potential to be powerful with a bit of work is, IMHO, probably a lot better than getting an expensive software license.
Again, never worked with the API; just my experience working with SDKs.
I don't beleive MS SAPI is supposed to be perfect out of the box; it provides a framework, you'll need to do some fine-tuing to get it working they way you need it. A free base which has the potential to be powerful with a bit of work is, IMHO, probably a lot better than getting an expensive software license.
Again, never worked with the API; just my experience working with SDKs.
Meow...?
This is turning into a somewhat heated debate.
This is turning into a somewhat heated debate.
I'm making a rather ignorant comment when I say this (As I don't do audio related programming)
I don't beleive MS SAPI is supposed to be perfect out of the box; it provides a framework, you'll need to do some fine-tuing to get it working they way you need it. A free base which has the potential to be powerful with a bit of work is, IMHO, probably a lot better than getting an expensive software license.
Again, never worked with the API; just my experience working with SDKs.
I don't beleive MS SAPI is supposed to be perfect out of the box; it provides a framework, you'll need to do some fine-tuing to get it working they way you need it. A free base which has the potential to be powerful with a bit of work is, IMHO, probably a lot better than getting an expensive software license.
Again, never worked with the API; just my experience working with SDKs.
Really, all Dictation software / API / etc is going to require training (although I found the google search on my tab and android devices are surprisingly accurate)
@aanthonyz , well that's my point, your not building a custom wrapper at all, and I am curious what will be upgraded that isn't already a supported feature.
Good luck mate, sounds interesting.
Lol...epic bump...
This thread is closed for replies.

