I play guitar and I record myself playing and singing, and a while back I hit a plateau. What kicked this off was Peter Gabriel's "In Your Eyes." I'd been listening to it and loving it, and then I tried it at karaoke and I just couldn't get it, which bugged me, because I can usually land other singers pretty close, both their notes and their style. I've sung songs higher than that one and songs lower than it, so I couldn't work out why the middle of my range was the part tripping me up, or where my voice was actually failing me. I never studied music, so I didn't really have a way to answer any of that, or even the vocabulary to describe what I was hearing.
That's the reason Timbr exists. It's browser-based (gettimbr.com), and the idea is simple: you record a short clip, it measures what your voice is actually doing (pitch, range, steadiness, vibrato, where you spend your time), and it shows that back to you as a voice map with a plain-language read next to it. The options I had before I built it didn't really help. An app that just hands you a label like "you're a tenor" is a black box I can't check, and a coach costs money and means having someone in the room. I wanted to be able to see it on my own.
One note on scope before you read too much into any of this. It's a private preview with a handful of users, so there aren't growth charts to show you, and the paid tier is built but switched off for now. I did the whole thing solo: the research, the product design, the data visualization, and the part I care most about, which was tuning how the model reads the measurements and judging whether the measurements were any good to begin with.
Why I didn't just bolt on a chatbot
The default move right now is to drop a general LLM into your product and call it a feature, and I didn't want to do that, for a few reasons.
Most of the trouble with LLMs in products comes from the input. If you give users an open text box, they'll take it wherever they want, which opens the door to prompt injection and a hundred conversations that have nothing to do with the point of the tool. So instead I gave people a set of preloaded questions, the ones I'd found genuinely useful when I was poking at my own voice. That keeps people on task, it keeps my costs under control (this is a free tool, and I wasn't about to hand out an unlimited chatbot), and it keeps the whole experience narrow and within my control.
The deeper reason is sycophancy. Chatbots get tuned so that people like them, and a lot of that likability comes from flattery, which is fine for a companion app but works against you in a diagnostic one. If the whole point is to tell you something true about your own voice, flattering you doesn't move you forward, it just feels good for a second. So the model never hears your audio at all. It receives a structured, plain-text summary of the measurements and reasons from those numbers, and the system prompt tells it as much: you're reasoning over numbers, not audio, and you did not hear the recording. Working only from numbers means it can't make up a performance it never actually heard.
The product kept telling people they'd failed
There were two problems I only found by using the thing myself and watching other people use it.
The first one is a little embarrassing. I build on a MacBook, and most of my users sing on their phones. Someone was using the live tuner on their phone and the whole screen kept bouncing, because every time the readout changed width the row would re-wrap and shove everything underneath it down a bit. I'd never once seen it happen, because on my laptop the layout had plenty of room. The fix was simple enough (reserve the worst-case width and drop the hint text on small screens), but it was a reminder that I'd been designing from the wrong device, and the only reason I caught it was that I put the tool in front of someone who actually lives where my users live.
The second problem was bigger, and it was about what the tool was willing to judge. Early on it treated every single note as though you'd meant to hold it, so I'd sing a phrase and it would come back and tell me I hadn't hit any of them and that my vibrato was all over the place, when all I'd really done was slide through those notes on my way somewhere, the way you do when you're singing expressively. If every note gets scored as a sustain, most of a song comes back as a failure, which isn't how a song works, it's how practicing a single note works. So I had to teach the tool the difference between a note you're parking on and a note you're traveling through. The signal that ended up working was slope: a wide, clean sweep gets read as a glide and is never scored for steadiness, and any note whose pitch is still moving quickly is treated as moving rather than held. I also smooth the vibrato out first so that expression doesn't get mistaken for a wobble.
That shaped the visualization as much as the scoring. The tool is measuring flatness, not whether you matched a melody, because it has no melody to compare you against, so it's deliberately blind to that and won't raise an alarm just because your pitch moved around. On the voice map itself, the vertical bar through each note shows how far that note drifted, the color runs from green to red depending on steadiness, and the size of the dot shows how often you've sung it.

Saying what I don't know
The thing I keep coming back to is that I don't want to flatter people about their voice, because I wouldn't want anyone doing that to me about mine. The best feedback I've ever gotten was clear about its own limits and about where it was coming from, and that kind of integrity does more than just feel right, it gets the user a better outcome and it sets honest expectations for the business too.
So the honesty is built into the data the model works from rather than added on as a polite tone at the end. The thresholds that decide whether a note counts as steady or drifting are hand-set and flagged as provisional in the code, because that's exactly what they are. They haven't been validated against a real coach, and the product says so on every screen: "This is an informed starting hypothesis from measurement, not a verdict. Register and voice-type reads are inferred from provisional thresholds. A real voice coach is the ground truth." Register only ever shows up as a rough zone estimated from your glides, and when the glides disagree with each other, the tool tells you that directly instead of guessing.

Leaving register out entirely was a deliberate call, and it's a decent example of how I work with AI. I'm not a vocal expert. The tool harvests your timbre data, the spectral shape of your sound, and the model could easily have handed you a register or a voice type based on it. Early on I wanted to understand how that's actually determined, so rather than taking the model's default answer I interrogated it and dug into what that design decision would really mean for a user. What I came away with was that reading someone's true singing register takes a human in the room, watching multimodal cues like breath and how the body is working while the person sings, and that's not something I can rebuild from a spectral number. So I left it out, and the tool harvests that data without pretending to know what it means.
Choosing the model was part of the same effort. What I needed was structural honesty: state the measured things plainly, hedge the inferred things, and never hand someone a confident single number for something that's genuinely fuzzy, so instead of saying "your break is at F4" it says something more like "somewhere around E4 to G4, and the signal there is messy." One model I tried had great formatting but was far too sure of itself, making claims about your physiology that I had no business making, so I ended up on a cost-efficient model (DeepSeek-V4.1-Flash via DeepInfra, per the project's docs) and did the real work in the prompt, with evidence tiers that don't blur together and a rule to report the confidence band rather than its midpoint. The provider sits behind a clean abstraction, so switching to any model with an API key is a single environment variable, which matters a lot more for how fast I can iterate than it does as a feature on a list.
What this is, if you're hiring
I came up through design and research, and I took the initiative to work directly with engineering to build this. I can have the conversation with AI the way I'd have it with an engineer, because I've spent years on cross-functional teams, and I'm not shy about pushing back on the AI I'm building with. Since I'm not an expert in vocal physiology, I made a point of understanding the consequences of a design decision before shipping it rather than accepting whatever the model gave me by default. I also think like a product person, in that I have limited bandwidth and rely on research, real usage, feedback, and strategy to decide where it goes. The question sitting underneath the whole project is how to give singers a fuller picture of the data around their voice without being prescriptive about what it means, and the best work I did on Timbr came down to answering that honestly, which mostly meant being disciplined about what not to claim.