The practice · · 5 min
Voice journaling: what changes when you talk instead of type
Typing gives your inner editor time to work. Speaking does not. What comes out when you talk your journal, and three prompts to try tonight.
In short
Typed journal entries often come out tidy and true and somehow beside the point, because typing runs slow enough for the editor to keep up. Speaking does not. This post names four things that change when the journal gets spoken instead of written: wants that the composed voice will not say, patterns that only show up when you review weeks of transcripts at once, introspection that stays with the felt thing instead of turning into an argument, and a next step that arrives rather than gets figured out. It closes on the actual practice, which is speak first, transcribe later, distill on a schedule, and then go back and read the specific lines the distillation points at. The argument is simply that speaking makes different material than typing does.
You open the notes app. The cursor blinks. And what comes out is neat, and true, and somehow beside the point.
You read it back a week later and it sounds like a report about someone else. Correct in every line. Nothing in it surprises you.
The problem is not that you are bad at journaling. It is that typing runs slow enough for your editor to keep up.
Voice journaling is the practice of speaking the entry out loud and keeping the transcript. What changes is not the topic. It is who is holding the pen. In 2016 Sherry Ruan and a team from Stanford, Baidu and the University of Washington tested speech input against the iPhone keyboard, and speaking came out roughly three times faster in English. Your editor is calibrated to the slower one. Speak, and you get out ahead of it.
You compose when you type. You confess when you talk.
Why does speaking get past the editor?
The editor is not your enemy. It is the part that checks a sentence before it goes out, and most days it is doing you a favor.
But it needs time. Between the thought and the finished sentence, typing leaves a gap wide enough for a small decision: is this the version I want on the page. Usually the answer is no, and a tidier version goes down instead. That happens hundreds of times an entry and you don’t even notice it happening.
Speech closes the gap. The mouth is already moving before the check completes. So the first version gets out, and the first version is the one with the useful stuff in it.
There is a second, quieter thing. When you write, you are performing for a reader, even if the reader is you in five years. When you speak alone, out on a walk or in the car, the audience drops away. Sentences trail off. You contradict yourself in the same breath. That mess is not a flaw in the record. It is the record being honest.
What shows up that typing never surfaces?
Four things, in my own practice.
Wants the composed voice will not say. The typed version of what I want is always slightly noble and slightly vague. The spoken version says the specific thing: the number, the room, the summer I actually want. I have said things into a recording that I would not have written down, not because they were shameful, but because writing them would have made them a claim. Speaking them made them true.
Patterns across weeks. No single entry shows you the whole picture. Fifteen of them, looked at together, do. The same phrase keeps arriving in different months, attached to different situations. I never caught that when I was typing, because a typed entry is already summarized, and a summary hides recurrence by design.
Introspection that stays with the feeling. When I type about something hard, it turns into an argument within a few sentences: here is the case, here is the counter-case, here is my verdict. When I speak about it, I stay in the room with it longer. I describe where it sits in the body before I explain what it means. Sometimes I never get to the meaning at all, and the entry is better for it.
A next step that arrives instead of getting figured out. This is the one that convinced me. After a long swim last month I recorded for a while about nothing in particular, and partway through I heard myself say that I move now for my own sake and not to impress anyone. I did not work that out. It showed up in my mouth. A few sentences later the next thing I was supposed to be working on was sitting there in plain language, and I had not been looking for it.
Is this just journaling with extra steps?
Journaling itself has a long research trail. James Pennebaker, at the University of Texas at Austin, has been running expressive-writing studies since the mid-1980s, and the finding held up repeatedly: writing about what is bothering you does measurable good.
None of that is at issue. The claim here is narrower. Speaking and typing produce different material out of the same person on the same day. If the typed version keeps coming out flat, the fix may not be more discipline about journaling. It may be a different input.
What do you actually do with the transcripts?
Speaking is capture, and capture without a second pass is just a bigger pile. Every screenshot you saved and never went back to is the same failure in a different medium.
Personally, I find it hard to muster up the motivation to read my transcripts. I have months of transcripts and I have read almost none of them straight through. What I do is hand the batch to an AI, to extract and distill insights. What keeps coming up? What did I say more than once? What am I circling without naming?
It comes back with a short list and the actual lines it pulled that from. Then I read those lines. Whether you analyze with AI or read yourself, review is not optional.
The distillation is an index, not a verdict. It will flatten things. It can miss the one entry that mattered most. It will hand you a neat summary of something you were being messy about on purpose. The transcript is the record. The summary is only a way back into it.
The recurring theme is the signal. One entry is a mood. Four entries carrying the same phrase across six weeks is a pattern, and patterns are the part you can actually do something with.
What does the practice look like?
Speak first. Transcribe later. Distill on a schedule.
Speak into whatever is already in your pocket. Voice Memos is fine. Do it while walking, or driving, or right after something physical, when the body is settled and the mouth is loose. Ten minutes is plenty. Say the boring parts.
Transcribe it later, not while you are speaking, so nothing on a screen is watching you talk. Whisper, or the transcription already built into your phone, will do it well enough. Perfect accuracy is not needed. My transcripts are full of mangled words and that’s ok. They are still invaluable to my growth.
Then the review day. Same day each time so it is not a decision. Feed it everything since the last one and ask what repeats.
One caution. Speaking will bring things up faster than typing does, and it does not always give you a place to put them. If something arrives and stays, take it to a practice built for that, or to a person. The four stages is one place to start, and it is five minutes.
Voice journaling stuck where typing never did. The review cadence is still loose: some months I run it weekly, some months I let the pile get long. What I no longer do is compose the entry. The composed version was always the one I already knew.
Written by Martin Ronquillo, founder of Genuai. He reads every reply.
Send me Sunday's noteA short note · every Sunday
The Sunday Alignment
For the Sunday night when the week already feels heavy. One short note before Monday: three minutes, one thing to do. By Thursday you have done it once.
You're in. The next note lands Sunday.
Something went wrong. Please try again.
One email a week. Free. Unsubscribe in one click.