On July 8, 2026, OpenAI replaced the entire voice stack inside ChatGPT with GPT-Live—full-duplex (think of it as simultaneous listening and speaking, where it can be interrupted mid-sentence), with the added bonus of real-time translation while you listen [Source: https://openai.com/index/introducing-gpt-live/]. Free accounts get GPT-Live-1 mini by default; paid accounts can use the larger GPT-Live-1 [Source: https://openai.com/index/introducing-gpt-live/].
These past few days, the community has been flooded with screenshots showing how smooth its translation is. But there’s a question almost nobody’s answering: if you’re on a Zoom or Google Meet call with an overseas client, can you actually pipe it in? Can the person on the other end hear the translation too?
Let’s get this out of the way first: this “cable” isn’t just for client meetings. You’re watching a foreign YouTube livestream and want to understand it in real time, you’re taking an online class that’s only in English, you’re doing a daily with a remote teammate in a totally different time zone, you’re interviewing at a foreign company, you’re attending a webinar in another language—anytime the sound is coming out of your computer, the setup is the same. That’s why it’s worth spending five minutes to learn this: you’re learning one cable you can plug in anywhere, not some narrow trick that only serves international business.
The answer is yes, you can hook it in—and you can do it today—as long as you know where that invisible “cable” lives. This article walks you through the full setup, and lays out the three things that’ll trip you up, so you’re not just playing with something new—you actually have a working tool for tomorrow’s meeting or tomorrow’s livestream.
Why “Just Open a Second Phone and Speak Into It” Is a Dead End
Video meetings are two-way audio: you need to hear them, and they need to hear you. You put a phone next to you running GPT-Live for translation, and the translated audio only goes into your ears—Zoom on the computer doesn’t even know that audio exists, and the other side receives absolutely nothing. The reverse is just as broken: the phone can’t clearly hear the other party’s voice coming out of your Zoom speakers, with air between them plus echo on top.
For one-way scenarios like watching a YouTube livestream or taking an online class, pointing a phone at the screen barely gets you something usable—but it still loses accuracy to air, noise, and echo, plus you can’t take notes. So “just use another device” isn’t hard to use—it’s architecturally disadvantaged from the start. To get the translation cleanly into your side, the audio needs to be connected with a “cable” inside the same machine, not shouting across a room. Once you see this, you’re already ahead of most people still pointing their phone at the screen and yelling.
The Real Setup: An Invisible “Virtual Audio Cable”
The keyword is virtual audio cable—you can think of it as an invisible cable that takes Program A’s speaker output and feeds it directly into Program B’s microphone input. Common options are BlackHole or Loopback on Mac, VB-CABLE on Windows, all with free versions [Source: https://github.com/ExistentialAudio/BlackHole].
Let me draw out the signal flow first, so every step below makes sense:
| |
In one sentence: you’re not speaking to Zoom directly—you’re letting GPT-Live’s translated audio “pretend to be your microphone,” Zoom thinks that’s you talking, and the other side hears it. Follow the steps below; on Mac it takes about five minutes.
Mac version (BlackHole, free)
- Install BlackHole. Download BlackHole 2ch from the official GitHub, or use Homebrew with one line:
brew install blackhole-2ch. Quit any apps currently playing audio before installing [Source: https://github.com/ExistentialAudio/BlackHole]. - Open “Audio MIDI Setup” (in Applications → Utilities). Click the + in the bottom-left, create a “Multi-Output Device,” and check both BlackHole and your own headphones/speakers—this step is so you can hear the audio yourself too.
- In the same window, right-click that Multi-Output and select “Use This Device for Sound Output.” The system audio will split a copy to BlackHole and a copy to your ears simultaneously.
- Go to Zoom → Settings → Audio, and in the “Microphone” dropdown select BlackHole 2ch; keep your headphones as the speaker. Same for Google Meet and Teams—just pick BlackHole in their audio settings.
- Test it: in ChatGPT, turn on GPT-Live and say something to have it translated. If the microphone level bar in Zoom’s settings moves, the other side is receiving it.
Two gotchas upfront so you don’t waste time: First, when building the Multi-Output, don’t use AirPods as the “primary” device—their low sample rate will make the whole thing sound garbled; use your built-in speakers or BlackHole 2ch as the primary device [Source: https://github.com/ExistentialAudio/BlackHole]. Second, with this free setup, all system audio gets piped into the microphone, so during meetings remember to turn off notification pings and background music. If you want only ChatGPT’s audio to go in and nothing else, you’ll need the paid Loopback for per-app routing.
Windows version (VB-CABLE, free)
- Download VB-CABLE, unzip it, and run VCCABLE_Setup_x64.exe as Administrator. Reboot after installing so the system recognizes the cable [Source: https://help.livevoice.io/article/146-using-vb-audio-virtual-cable-to-send-audio-from-zoom-or-ms-teams-to-livevoice].
- Route the audio you want translated into this cable: go to Settings → System → Sound → Volume Mixer, and set ChatGPT (or your browser)’s output device to CABLE Input only; other apps won’t be affected.
- In Zoom → Settings → Audio, select “CABLE Output” as the microphone [Source: https://help.livevoice.io/article/146-using-vb-audio-virtual-cable-to-send-audio-from-zoom-or-ms-teams-to-livevoice].
- To hear it yourself: go to “Sound Control Panel → Recording → CABLE Output → Properties → Listen,” check “Listen to this device” and pick your headphones; or install VoiceMeeter (free, same vendor) to split audio between your headphones and Zoom.
- Test the same way—say a sentence and check if Zoom’s microphone level bar moves.
Just want to one-way listen to a foreign livestream or an English class? Even easier: you don’t need the “send back to Zoom” half at all. Just feed the source audio through the same virtual cable into GPT-Live and let it translate for you—skip the outgoing side entirely. Same cable, plugged into both ends for meetings, one end for pure listening. Learn it once, use it everywhere.
Honestly though: the truly tricky part is “two-way”—translating your words for them and their words back for you in the same meeting means running a second cable to feed Zoom’s incoming audio back into GPT-Live, while dodging echo loops. That’s a hard nut to crack, so when I cover the three bottlenecks below I’ll recommend: for two-way, multi-person, high-accuracy scenarios, don’t try to make GPT-Live carry the whole load—pair it with a dedicated meeting translation tool. Start with the stable path above—“you speak, the other side hears the translation”—and you’re already half a step ahead of most people.
This part is the one you can actually do today, right now. It relies on OS-level audio routing, not waiting for some official feature to ship. In other words, even though GPT-Live hasn’t opened up its developer API yet (OpenAI only says “coming soon,” with no date and no pricing [Source: https://openai.com/index/introducing-gpt-live/]), you don’t have to sit and wait—the virtual audio cable itself is the “start using it now” solution.
No API Yet—What Can Developers Do Right Now?
People who want to build products with GPT-Live see “no API” and instinctively think it’s game over—but this needs to be looked at on two layers.
Layer one: the official API isn’t open yet, so you genuinely can’t programmatically call GPT-Live as a backend service. Layer two—and this is what most people miss—the end result you’re delivering, “real-time voice translation,” doesn’t necessarily have to wait for that API. The virtual audio cable approach itself can be packaged into a repeatable internal workflow: lock in the routing, write it as a setup script, pair it with existing meeting or livestream transcription tools, and you’ve got something you can deliver to a team, to a client, or even package for general users today—no need to wait on an API with no release date.
And the demand has already been validated, far beyond just business meetings. Tools like Felo, which are specifically bolted onto audio, already do real-time bilingual captions and transcription for Zoom, Google Meet, and Teams using ChatGPT/OpenAI, and even cover YouTube livestreams [Source: https://chromewebstore.google.com/detail/felo-subtitles-chatgpt-li/ponokiofkijoolhebggofhhibnafebna?hl=zh-CN]. In other words, “AI riding the audio stream for real-time translation” covers the needs of people in meetings, classes, and livestreams alike—not some new concept you have to educate the market on. It’s a mature demand already being served. Your job isn’t to prove from zero whether anyone wants it, but to make the experience fit your scenario better than existing tools.
When the API eventually opens, the audio routing, meeting and livestream integration, and latency trade-offs you familiarize yourself with today won’t go to waste—they’re the foundation that lets you run faster than everyone else the moment it lands.
The Three Things That’ll Trip You Up: Latency, Accent, Terminology
Getting it hooked up doesn’t mean it runs smoothly. What actually decides whether this setup is usable comes down to three things. Know them upfront, and you can pick the right scenario before the meeting or the stream starts, instead of embarrassing yourself in the moment.
Latency is the biggest one to watch. OpenAI still hasn’t published any first-token or trailing latency numbers for GPT-Live [Source: https://openai.com/index/introducing-gpt-live/]. Real-time voice translation inherently carries a time gap—it has to hear a full sentence, send it to the model for translation, then speak it out; that pipeline is naturally slower by half a beat. In one-on-one slow chats, or one-way livestream watching, it’s tolerable. But in a multi-person meeting where everyone’s going back and forth, every sentence getting pulled apart a bit longer and the whole rhythm falls apart. Exactly how many seconds? OpenAI hasn’t said, and we won’t make it up—we’ll update once they publish or credible benchmarks come out. The practical fix is simple: start with one-on-one, one-way viewing, or slow-paced scenarios, and lean into its strengths.
Accent and language coverage is the second hurdle. OpenAI itself has disclosed that beyond “most common languages,” other languages will have non-native accents and insufficient fluency, with full multilingual parity “explicitly listed as not achieved” at launch [Source: https://openai.com/index/introducing-gpt-live/]. So if your meeting or that livestream is in a major language like English, Japanese, or Korean, the experience will feel much better; the more niche the language, the easier it is to feel out of place. Pick the right language, and you’re most of the way past this one.
Specialized terminology is the third hurdle, and there’s no official data on it—only reasoning. GPT-Live feeds the hard problems back to the text model behind it to compute while it speaks, which is fine for everyday conversation and everyday livestreams. But for financial figures, legal clauses, medical terms, or technical specs where you can’t get a single word wrong, the judgment call is: it can be an assistant, but you can’t treat it as your only transcript. That’s why, in real ongoing conversations, dedicated real-time translation tools tend to be more reliable: the built-in real-time translation captions in platforms like Zoom and Google Meet [Source: https://support.google.com/meet/answer/10964115?hl=zh-Hant&co=GENIE.Platform%3DDesktop][Source: https://www.timekettle.co/blogs/tips-and-tricks/how-to-translate-zoom-or-google-meet-conversations-in-real-time], or tools like Felo that are specifically bolted onto meetings and livestreams for transcription and translation [Source: https://chromewebstore.google.com/detail/felo-subtitles-chatgpt-li/ponokiofkijoolhebggofhhibnafebna?hl=zh-CN], hold down the “just translate” role and also leave you with a bilingual transcript you can review—hugely useful for people who need to organize meeting notes afterward or review class highlights. The smart move is to run both: GPT-Live covers the flow of real-time conversation, and the dedicated tool covers the accuracy you can look back on.
One more expectation trap to watch out for: GPT-Live (and really, any OpenAI interface) can’t make regular phone calls for translation calls—that’s an entirely different product category [Source: https://openai.com/index/introducing-gpt-live/]. If you need to take phone calls, don’t wait for it—go find the right tool directly.
So, Is It Actually Usable Right Now?
Yes, and the barrier to entry is lower than you’d think. To get a feel for it first, just turn on GPT-Live in ChatGPT’s voice mode and try some translation. To actually pipe it into Zoom or that livestream you’re watching, look up the virtual audio cable for your system (BlackHole for Mac, VB-CABLE for Windows) and set up that “pretend to be a microphone” path above—plug into both ends for meetings, just one end for pure viewing, and start with a major language and a slower pace.
“Hooking GPT-Live into your audio stream”—whether that’s a client, a colleague, or a foreign livestream—is technically doable today; the only thing left is whether you’ve picked the right scenario. The virtual audio cable part is real and actionable, and the three bottlenecks—latency, niche languages, specialized terminology—aren’t mysticism; they’re known boundaries you can route around ahead of time. Once you understand how to set up the cable and where the edges are, you’re half a step ahead of everyone else in that same meeting or livestream: while they’re still debating whether to hold a phone up and talk into it, you’ve already got the translation cleanly in your ears. Tools will keep running forward, but the person who gets their hands dirty first and wires it up is always in the best position to catch the next wave.