A vet finishes an examination with both hands busy and a dog on the table. The notes get typed later, from memory, or not at all. Groomers have the same problem with session notes.
So voice to notes is on our plan for the Happy Pet Tech SaaS: a staff member speaks, and the app turns it into a written note in the right place. We have not built it. This post is my research before we start: how the technology works, and where I expect it to go wrong.
The pipeline has three steps
microphone → speech to text → LLM → draft note → human confirms → saved
People often imagine one AI that listens and writes. In practice these are separate jobs, and keeping them separate makes each one easier to test.
Step 1: recording in the browser
Our app is a PWA, so recording happens in the browser. getUserMedia asks for the microphone, and MediaRecorder turns the stream into audio chunks.
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
const recorder = new MediaRecorder(stream);
recorder.ondataavailable = (e) => chunks.push(e.data);
recorder.start(1000); // a chunk every second
Three things to know before writing this:
- It only works on HTTPS, and the user must give permission. If they deny it once, the browser remembers.
- The audio format is not the same in every browser. Chrome records WebM with Opus. Safari records MP4 with AAC. The server or the speech service has to accept both.
- On a phone, recording can stop when the screen locks or the app goes to the background.
Step 2: speech to text
The audio goes to a speech-to-text model, which returns a transcript. There are two ways to do it.
Batch. Record the whole note, upload it, get the text back. It is simple and cheap and the accuracy is good, because the model hears the full sentence before deciding. The user waits a few seconds at the end.
Streaming. Send audio over a WebSocket while the person is speaking and get words back live. It feels better, and it is much more work: a long-lived connection, partial results that change as more audio arrives, and reconnecting when the network drops.
For notes, I think batch is the right start. A vet dictating for 40 seconds does not need to watch the words appear. They need the note to be right.
The hard part of this step is the audio itself, not the model. A clinic or a salon is loud. Dogs bark, dryers run, other people talk. And our users are in India, the UAE, Australia, the Philippines and Thailand, so accents vary a lot, and people mix languages in one sentence.
There is also vocabulary. Drug names, breed names and procedure names are not common words, and a general speech model gets them wrong. Most speech services accept a list of expected terms or a prompt with hints. We already have lists of breeds and services in the product, and I expect feeding those in to matter more than which model we pick.
Step 3: from transcript to a note
A transcript is not a note. It has “um”, repeats, corrections (“two ml, no, three ml”), and no structure.
This is the LLM’s job. It gets the transcript and returns the note in a fixed structure. For a vet consultation, a common format is SOAP: subjective, objective, assessment, plan. The model is asked to return JSON that matches a schema, so the app can put each part in its own field.
{
"subjective": "Owner reports reduced appetite for three days.",
"objective": "Temperature 39.4 C. Mild dehydration.",
"assessment": "",
"plan": "Blood panel. Recheck in 48 hours."
}
The instruction that matters most is about what the model must not do. It must not add anything the speaker did not say. If the vet did not give an assessment, that field stays empty. A model that fills gaps with reasonable-sounding medical text is dangerous here, because the result looks exactly like a real note.
I would also keep the original transcript, and the audio for a limited time, next to the note. If someone later questions the note, there is something to compare it with.
Step 4: a person confirms
Nothing should be saved to a pet’s record without a person reading it first.
The app shows the draft in the normal note form, already filled in. The vet fixes what is wrong and saves. Numbers need the most care. “Fifteen” and “fifty” sound close, and a wrong dose in a record is a real problem. I want the draft to mark numbers and drug names so the eye goes to them.
This step also solves a trust problem. Staff will not use a feature that writes into medical records by itself. They will use one that types for them and lets them check.
What I still have to decide
- Where the audio goes. Audio of a consultation can include the owner’s name, phone number and address. Which provider processes it, where, and how long they keep it has to be answered before any code.
- Cost. Speech to text is priced by the minute of audio, and the LLM by tokens. Short notes are cheap. Someone who leaves the recorder on for a whole appointment is not. There needs to be a time limit.
- Languages. We have to pick which languages to support first and test with real recordings from real clinics, not with clean audio from a quiet room.
- Failure. If the transcript is poor, the app should say so and keep the audio, and not produce a confident note from bad input.
I’ll write the second half of this when it is built and I can say which of these guesses were wrong.