📰 Key Takeaways
Synthesia is a startup focused on AI digital avatar video generation, valued at $4 billion earlier this year, with annual recurring revenue (ARR) surpassing $100 million last year; competitors in the space include D-ID, HeyGen, and Colossyan. The company’s core business is helping enterprises create interactive AI avatar training videos, and it recently launched a new product called Roleplay Sessions that lets employees practice scenarios like sales pitches with interactive AI avatars, with the system responding and scoring in real time. The article’s author was invited to visit Synthesia’s new New York office in September, where she was offered the chance to have an interactive digital avatar made of herself — the first time Synthesia has created one for a journalist rather than someone internal to the company. During the process, she stepped into a mini studio inside the office, where the team shot a large number of photos and recorded a two-minute voice sample, and with her consent, generated the avatar. The team made two static avatars for her (reading a fixed script only, with and without glasses) and two interactive avatars (able to converse and listen to questions, also with and without glasses). The tech stack behind the interactive avatar combines speech-to-text, image generation, a language model, and text-to-speech, trained on one of her articles about elevated fraud rates at VC-backed startups — meaning it can only answer questions related to that specific article. Synthesia provides its own video and voice models, but also lets customers swap in alternatives from vendors like Cartesia, ElevenLabs, and Google. See the original article for full details.
💬 JudyAI Lab Take
Synthesia, the AI digital avatar video startup now valued at $4 billion with ARR that crossed $100 million last year, just let a journalist go through the process of building an interactive avatar firsthand — and that “opening the black box to outsiders” move is itself a signal worth paying attention to.
What stands out from the reporting is that the interactive avatar runs on a combination of speech-to-text, image generation, a language model, and text-to-speech — and critically, the training material was scoped to a single article, with the answer space locked down accordingly. That reflects a tradeoff a lot of AI builders are making right now: a sprawling general-purpose capability sounds appealing, but products that actually ship tend to narrow the interaction scope, letting the avatar only converse within a defined set of material and context, in exchange for controllability and trustworthiness. Synthesia also lets customers plug in outside models from Cartesia, ElevenLabs, and Google alongside its own — another sign that “in-house model plus third-party models” is becoming the standard architecture for this category, rather than single-vendor lock-in.
If you’re building an AI product: next time you’re designing a conversational feature, ask yourself — does this AI really need to “know everything,” or would it actually be more reliable if you scoped it down to one clear domain?
📅 Source Info
- Published: 2026-09-26T14:00
- Original source: https://techcrunch.com/2026/09/26/i-created-an-interactive-digital-avatar-of-myself-and-you-can-talk-to-it/