Rondelek TWST-1

GAME · KIDS · SPEECH

A little therapeutic platformer for kids that you play with your voice. Shout a vowel into the microphone and the hero jumps, ducks or leaps high - so saying the sounds out loud is the whole game.

ROLEDesign, code, AI development
STACKRust, Go, raylib
SHIPPEDIn progress

A work in progress. I’m currently looking for a speech therapist to help me make this game more useful for kids. If you are one, or know one, please get in touch. Also looking for a designer to help me make it look better.

My e-mail: [email protected]

What it is

A simple therapeutic platformer game for kids to help them train their speech in a fun way. The player controls a hero with their voice only, saying vowels out loud to make the hero jump, duck, leap high or shoot at obstacles.

Voice is the controller

Before the game begins, the therapist or parent must assist the kid in training the engine to recognize their voice by saying each vowel a few times. The game setup screen allows you to bind specific vowels to player actions. The player then shouts those vowels into the microphone to control the hero’s actions.

Why vowels

The vowel shouting game started as a side effect of what Rondelek TWST-1 was initially built for: a simple-to-use sampler for kids to record their own voice and play it back.

At some point I brought Claude into the process and thought that we could build a vowel-recognition-based visualisation system into it. When I finally saw it working, I couldn’t resist making a little game out of it.

Initially, the Rondelek sampler was written by hand in Golang and Raylib, and it focused on the sound recording/playing part only. It looked like this, and it even worked:

My aim was to make it look a bit like the EP-133 K.O. II sampler. What you see here is truly my best attempt at designing it. I never finished it, and I wasn’t proud of my Go skills either. Being jobless and having some time on my hands, I decided to let an LLM rewrite the whole thing in Rust, out of pure curiosity.

Combobulating the load-bearing noodles

The first model I asked for the rewrite was DeepSeek V4 Pro. It churned and churned, but ended up with a broken audio layer, messed-up looks and utterly wild controls.

Claude Sonnet, the next contestant, actually didn’t make the looks much better, but at least the UI worked and the core parts behaved very well on eframe / egui. Eventually we revisited the looks and introduced simple skinning.

After fixing high CPU usage, then latency and UI issues, we were about to vibe-code a simple dot-matrix visualiser, when I suddenly remembered a class from my English studies, where Dr. Wiktor Gonet showed us a spectral voice analyser. I thought that maybe with the generous help of Claude we could build a working vowel recognition system into the sampler, thus turning it into a simple speech therapy tool for kids.

To recognize or not to recognize, that is the question

This part probably ate more tokens than it should have, all because of my lack of voice processing knowledge. I made a bad assumption and tried to coax Claude into building a usable vowel recognition system based on formants.

I dropped the idea for a while, but one night I came back and simply asked the right model (Claude Opus) the right question: what would be the most bulletproof way to recognize vowels in a voice signal? The answer was shockingly quick: use MFCCs (Mel-Frequency Cepstral Coefficients - gotcha!).

Of course, I still have only a vague idea of what they are, but at least I know that the right approach was to compare the fingerprint of the incoming signal with the fingerprints recorded during the calibration phase.

My only additions to Claude’s idea were a smoothing control and making the ‘nearest match’ work across several recordings of the same vowel, which made the recognition more robust and less prone to false positives.

I was amazed at where it took us: six samples of each vowel were enough to make the recognition practically flawless.

Where it is now

It is not yet publicly available, and I’m still looking for a speech therapist to challenge or support the idea of the game. The game is playable, but the graphics are placeholders and the game / sampler shell still leave much to be desired, though it wouldn’t take much to release it.