Python Voice Lab - Text-to-Speech + 10 Voice Effects (Audio Processing Project)

Posts 1–5 of 5 · Page 1 of 1
Python Voice Lab - Text-to-Speech + 10 Voice Effects (Audio Processing Project)
I created this Python project to experiment with text-to-speech generation and audio processing.

The script works as a small voice effects studio. It converts user text into speech using installed system voices, then allows the user to apply different audio effects and export the final result as a WAV file.

The project was created to learn more about:

- Text-to-speech engines
- Audio waveform manipulation
- Voice processing techniques
- Working with WAV files
- Python audio libraries


Features:

- Detects available system voices
- Allows selecting different TTS voices
- Converts text input into speech
- Applies multiple voice effects
- Exports processed audio into WAV format
- Automatically plays the final audio file


Available voice effects:

1. Robot voice
2. Cartoon voice
3. Low narrator voice
4. Deep cinematic voice
5. Electronic voice
6. Metallic robot effect
7. Room echo effect
8. High cartoon voice
9. AM/FM radio effect
10. Computer assistant style


Libraries used:

pyttsx3
- Offline text-to-speech engine using installed system voices

pydub
- Audio processing and effects

winsound
- WAV playback on Windows

tempfile / os
- Temporary file creation and cleanup


Installation:

pip install pyttsx3 pydub


Note:
FFmpeg may be required for additional pydub audio features.


This is a learning project focused on Python audio processing and experimenting with different voice styles.

Future improvements:
- More advanced audio filters
- Real-time voice processing
- GUI interface
- Saving custom effect presets


Feedback and suggestions are welcome.

Code:
import pyttsx3
from pydub import AudioSegment
import tempfile
import os
import winsound


# Initialize text-to-speech engine
engine = pyttsx3.init()

voices = engine.getProperty('voices')

print('Available system voices:')

for i, voice in enumerate(voices):
    print(f'{i}: {voice.name} - {voice.languages}')


voice_index = int(input('Select voice number: '))

voice = voices[voice_index]

engine.setProperty('voice', voice.id)


text = input('Enter text for speech generation: ')


# Create temporary WAV file
with tempfile.NamedTemporaryFile(delete=False, suffix='.wav') as file:
    temp_file = file.name


engine.save_to_file(text, temp_file)
engine.runAndWait()


# Load generated audio
audio = AudioSegment.from_wav(temp_file)


print('\nSelect voice effect:')
print('1 - Robot')
print('2 - Cartoon')
print('3 - Low narrator')
print('4 - Deep cinematic')
print('5 - Electronic')
print('6 - Metallic robot')
print('7 - Room echo')
print('8 - High cartoon')
print('9 - AM/FM radio')
print('10 - Computer assistant')


effect = input('Choose effect (1-10): ')


# Apply selected effect

if effect == '1':
    audio = audio._spawn(
        audio.raw_data,
        overrides={'frame_rate': int(audio.frame_rate * 1.1)}
    )

elif effect == '2':
    audio = audio._spawn(
        audio.raw_data,
        overrides={'frame_rate': int(audio.frame_rate * 1.6)}
    )

elif effect == '3':
    audio = audio._spawn(
        audio.raw_data,
        overrides={'frame_rate': int(audio.frame_rate * 0.85)}
    )

elif effect == '4':
    audio = audio._spawn(
        audio.raw_data,
        overrides={'frame_rate': int(audio.frame_rate * 0.65)}
    )

elif effect == '5':
    audio = audio.high_pass_filter(800).low_pass_filter(3000)

elif effect == '6':
    shifted = audio._spawn(
        audio.raw_data,
        overrides={'frame_rate': int(audio.frame_rate * 1.2)}
    )

    audio = audio.overlay(shifted, position=10)

elif effect == '7':
    echo = audio - 10
    audio = audio.overlay(echo, position=200)

elif effect == '8':
    audio = audio._spawn(
        audio.raw_data,
        overrides={'frame_rate': int(audio.frame_rate * 2.0)}
    )

elif effect == '9':
    audio = audio.high_pass_filter(200)
    audio = audio.low_pass_filter(3000)
    audio = audio.apply_gain(-5)

elif effect == '10':
    audio = audio._spawn(
        audio.raw_data,
        overrides={'frame_rate': int(audio.frame_rate * 1.05)}
    )

    echo = audio - 12
    audio = audio.overlay(echo, position=120)


# Export final audio

output_file = 'voice_output.wav'

audio.export(output_file, format='wav')


print(f'\nAudio saved: {output_file}')


# Play generated audio

winsound.PlaySound(
    output_file,
    winsound.SND_FILENAME
)


# Remove temporary file

os.remove(temp_file)
Quote Originally Posted by arunforce View Post
not sure if trolling or not
Not trolling Just a learning project to experiment with TTS and audio effects in Python.
Nothing too advanced, but maybe useful for someone learning audio processing.

- - - Updated - - -

It's mainly an audio processing experiment, but it also has practical uses.

For example, I often use TTS notifications in my own scripts and bots, so I don't have to constantly watch the monitor. Instead of checking logs every few seconds, the program can speak events like "account found", "task completed", "price increased", or "price dropped".

It's a simple project, but it demonstrates how TTS and basic audio processing can be integrated into Python automation tools.
Ah my bad. Saw that it mostly imported libraries instead of actually synthesizing text but I see that was the intent. Nice demo!
No worries, I understand what you mean.

The goal was not to build a TTS engine from scratch, but to combine existing TTS engines with audio processing and voice effects. More like a small voice effects playground that can also be used in scripts and automation for voice notifications.

Thanks for checking it out!
Posts 1–5 of 5 · Page 1 of 1

Post a Reply

Similar Threads

Tags for this Thread

Need help?