Do You Need a Bot on the Call? Device-Audio vs. Bot-Based Translation Tools
A bot that joins your call as a participant needs to be admitted by the host and shows up for everyone; a tool that listens through one person's device audio needs neither, at the cost of only hearing what that one device can hear.
How a bot-based tool works
A bot-based translation tool integrates with a video platform's API and joins the meeting as its own participant — often with a name like "Translation Bot" and its own video tile. This requires the meeting host to admit it (or configure auto-admit in advance), and it's visible in the participant list to everyone on the call, which is sometimes a deliberate feature (transparency about who's listening) and sometimes an unwanted extra step or a source of questions from other participants.
It also only works with platforms it has been built to integrate with — a bot built for one platform's API doesn't automatically work on a platform it wasn't built for.
How a device-audio tool works
A device-audio tool runs as software on one participant's computer and captures whatever that device's microphone or speaker output is producing — the same audio that person is already hearing or speaking into. It never joins the call as a separate entity, so there's nothing for a host to admit and nothing extra in the participant list.
Because it works at the operating-system or browser level rather than through a specific platform's API, it works identically regardless of which video platform the call happens to be on — the tool doesn't need to know or care what's generating the audio. This is the approach A6 Relay's Meeting Mode is built around.
The real tradeoff
Bot-based tools can be a better fit when everyone on the call should visibly know translation is active, or when the translation needs to be shared live to every participant automatically rather than just the one running the tool. Device-audio tools fit better when a single participant wants translation for themselves without changing anything about the meeting for everyone else, or when the call is on a platform no bot has been built to integrate with. Neither approach is strictly better — they answer different questions about who the translation is for.
Be first to try real-time interpretation yourself.
In active testing — join the waitlist and we'll email you when it's ready.
Read next
Simultaneous Interpretation for Video Calls: How It Actually Works
On a video call, simultaneous interpretation is a short pipeline running continuously in the background: audio in, transcribed speech, translated text, and captions or synthesized speech out — fast enough that it never asks the conversation to pause and wait for it.
Browser-Based Translation Tools vs. Downloadable Apps: Tradeoffs
The real tradeoff isn't quality — it's reach: a browser tool works instantly for anyone with a link, while an installed app can do things a browser can't, like listening to audio from other applications on the same device.
How Fast Does Real-Time Translation Need to Be? A Latency Guide
Latency in real-time translation isn't just a speed metric — it's the difference between a tool that feels like part of the conversation and one that feels like reading minutes from a meeting that already ended.