When a live translation reads badly the instinct is to blame the translation. In our experience it is almost always the audio. Speech recognition can only work with what reaches it, and a room full of echo, a lapel microphone at belt height, or a wandering handheld quietly ruins the whole chain before translation is even attempted.
This is the most useful article on this site for anyone whose results are disappointing, because the fix is nearly always upstream of the software and nearly always cheap.
How much difference it makes
More than people expect. On our own recordings, a microphone close to the speaker delivered several times the signal level of the same person picked up by a device across the room — and the consequence is not simply that quiet audio is a bit worse.
Below a certain level, speech recognition stops failing honestly. It does not go quiet or leave gaps; it produces confident, fluent, grammatical sentences that nobody said. We have watched a system generate an entire paragraph out of room noise during a pause.
That is why this matters more than any other setting. A system that fails visibly is safe. A system that fails fluently puts words in your preacher's mouth, in a language nobody in the room can check.
Take the feed from the desk, not the room
The single biggest improvement available is to stop using the laptop's built-in microphone.
A device sitting on a table hears the room: the air conditioning, the chairs, the coughing, the reverberation off the walls, and — somewhere in there, distantly — the speaker. Take a line out of your mixing desk instead and the recognition hears exactly what the sound engineer hears, which is the speaker's microphone and nothing else.
Most churches with a sound desk can do this for the price of a cable and a simple USB audio interface. If there is a spare output — a monitor send, a record output, even a headphone socket — you already have what you need.
Get the microphone close and keep it there
- Headset or lapel beats handheld. A fixed distance from the mouth means consistent volume. A handheld drifts every time the speaker gestures, and preachers gesture.
- A lapel microphone belongs high on the chest, near the collarbone — not clipped to a belt or a lower button. Every extra centimetre adds room and loses clarity, and the difference between collar and waist is much larger than it looks.
- Watch for clothing. A scarf, a lanyard or a jacket lapel brushing the capsule produces bursts of noise that recognition reads as syllables — which is how invented words appear in the middle of a good sentence.
- Beware the pulpit microphone on a gooseneck. It is fine for a speaker who stays put and useless for one who steps away to make a point.
- Fresh batteries, every time. A radio microphone fading through the second half of a talk produces exactly the gradual degradation described above, and nobody in the room will hear it happening.
Levels: consistent beats loud
Recognition does not need a loud signal. It needs a steady one.
Aim for healthy, unclipped levels and resist heavy compression. Clipping destroys the consonants that distinguish similar words — and consonants carry most of the information in speech, which is why a clipped signal produces text that is confidently wrong rather than obviously broken. Aggressive compression does the opposite damage: it pulls up the room noise between sentences until the engine starts trying to transcribe it.
If your desk has a gate on the speaking channel, a gentle one helps. A harsh one clips the start of words and is worse than none.
Mind the music
Instruments and singing are not speech, and feeding them to a speech engine produces confident nonsense.
If your desk allows it, send the translation system a mix that carries the speaking microphone and leaves the band out — a dedicated auxiliary send is ideal. Where that is not practical, pause during worship. It keeps the archive clean, it stops a congregation reading garbled lyrics, and on a metered arrangement it stops you paying for minutes nobody reads.
Rooms fight you
A hard, tall, beautiful church is an acoustically difficult room. Sound arrives at a distant microphone several times over, at slightly different moments, and the smearing that results is precisely what makes speech hard to recognise.
You cannot fix the building, and you do not need to. Every point above is really a way of not letting the room into the signal in the first place. A close microphone in a cathedral beats a distant microphone in a carpeted hall.
The five-minute test that settles it
- Speak a normal sentence into the microphone the speaker will actually use, from where they will actually stand.
- Read the transcript. Properly — not a skim.
- Do the same again after walking two steps away, and compare.
Problems that would otherwise have run for forty minutes show up in thirty seconds. And judge by the text, never by listening to the room: your ears know the preacher, the topic and the vocabulary, and they fill gaps without telling you. A speech engine has none of that.
If it is still poor after all that
Work through this order, because it is the order of how much difference each thing makes:
- Is the audio coming from the desk rather than a room microphone?
- Is the microphone on the speaker, high, and fixed in place?
- Are levels healthy and unclipped, without heavy compression?
- Is the band out of the feed?
- Have you supplied a list of the names and terms that recur?
Only the last of those is about software, and it only helps once the first four are right.
Where the platform helps
InMyTongue takes audio straight from a browser, so it works with whatever feed you can route to the machine — a desk output included — and shows you the live text as it arrives, so a bad feed is visible in seconds rather than after the service. New accounts include free credit. Try it with your own setup.