今年夏天,当视频生成初创公司Synthesia的企业事务负责人Alexandru Voica给我发来链接,介绍其公关团队的新成员时,我感到十分惊讶。那是他的一个交互式虚拟化身,经过训练可以回答关于Synthesia的常见媒体问题,比如它是做什么的、如何运作等。就在前一天,我还在一个小组讨论中,有公关人员问我是否介意使用AI生成文本的推介材料。但Alexandru的化身超越了这一点——在我看来,这似乎是AI在公关领域应用的“最终Boss”。
9月,Synthesia邀请我前往其在纽约的新办公空间。Synthesia最初总部位于英国,与D-ID、HeyGen和Colossyan等其他公司一起,成为炙手可热的数字虚拟化身初创企业之一。该公司今年早些时候估值达到40亿美元,并曾宣布去年其年度经常性收入(ARR)已突破1亿美元。
Synthesia让企业能够利用AI虚拟化身构建交互式培训视频,最近还推出了一款名为“Roleplay Sessions”的产品,允许员工与交互式AI虚拟化身进行练习,例如模拟销售推介,该化身会对员工的回答做出回应并评分。
当我参加新办公室开业活动时,他们问我是否想要一个属于自己的AI虚拟化身,我甚至没有犹豫就答应了。当然,我想要一个自己的数字分身。那天我的穿搭很可爱,发型也很完美。
在见到我的数字分身之前,我对虚拟化身一直持 indifferent(无感)态度,但我认为它们不可避免地会成为日常网络生活的一部分。我听说有人在Instagram上创建以他们为原型的化身,以帮助制作社交媒体内容。我觉得这一切都非常有趣,这也可能是我现在毫无顾虑地向大家展示我的数字分身的原因。这是Synthesia首次为记者(或除Voica以外的任何人)制作数字虚拟化身。它基于我关于“为何获得风险投资支持的初创公司比未获风投支持的初创公司更容易实施欺诈”的报道进行训练,且仅能回答与该报道相关的问题。
下方只需点击“在新窗口中开始”即可启动。
你可以向它提问,例如:
我为何决定撰写这篇报道?
那篇研究论文是关于什么的?
研究人员发现了什么?
为了制作这个化身,我进入了Synthesia办公室内一个迷你摄影棚,他们拍摄了我的大量照片,并录制了我两分钟的声音。我必须同意制作这些虚拟化身,于是,“数字Dom”就此诞生。他们为我创建了一个个人虚拟化身(仅朗读我提供的脚本)——有戴眼镜和不戴眼镜两个版本;还为我制作了两个交互式虚拟化身(可以对话并聆听我的声音),同样也提供了戴眼镜和不戴眼镜的版本。
我们选择了一篇文章作为训练交互式化身的基础,随后其中一个团队构建了我的交互式化身,其技术核心结合了语音转文本、视频、语言以及文本转语音模型。我的化身的技术栈包括Synthesia自有的视频和语音模型,尽管该公司也允许客户从Cartesia、ElevenLabs、Google或OpenAI等其他实验室选择替代方案。企业还可以选择将他们的虚拟化身托管在任意云平台上,或者付费让Synthesia代为托管。
语音转文本模型将人们说的话转化为文字,代理式语言模型理解文字内容并可根据其采取行动,文本转语音模型将回应转化为音频,最后由Synthesia构建的视频模型在化身说话时驱动其动画表现。
总体而言,Synthesia构建了三种类型的产品:一是带有经典虚拟化身的视频创建和分发平台,用户输入脚本后,虚拟化身会复述;二是名为Sessions的代理式平台,用户可以在调查或角色扮演中与虚拟化身互动;三是API平台,用户可以将Synthesia的视频和语音模型与其他技术服务结合,构建交互式虚拟化身或其他类型的产品。
Synthesia团队花了我几天时间才制作好我的虚拟形象。我先试用了个人版,输入了一段相当通用的脚本,以测试我的AI声音效果。我让它讲述纽约秋天到来的故事,那是我最喜欢的季节。声音相当逼真,我很高兴它没有录进我在录制音频样本时带有的沙哑声。
我把视频展示给一些非技术背景的朋友看,他们觉得既有趣又令人毛骨悚然。
接着,我向他们展示了我的互动版,它是确定性的,意味着它只会说出经过训练以回应的内容。在这种情况下,那就是我讲述的风投欺诈故事。我问了它诸如我在加入TechCrunch之前在哪里、住在纽约的哪个区域等问题,但每次它都将话题引回那个故事。
我的朋友觉得声音不太像我的,认为相似度不如个人版,但尽管如此,其逼真程度仍足以让人感到些许不安。我妈称之为“太神奇了”,这在她看来是极高的评价。她和爸爸一直试图问它一些“只有他们才知道的”关于我的问题——但模型并未作答,每次都把话题转回风投欺诈故事。
在测试完这个模型后,她开玩笑说:“我不记得给你们俩接生过。”
这次经历让我思考新闻业的未来会是什么样子。人们能接受打开电视看新闻时,由虚拟形象来播报吗?一位投资者立刻表示不行。当然,如今充斥着大量侵入社交媒体和其他新闻分享平台的AI生成垃圾内容,引发了强烈的抵触情绪。但我问过的其他人则不那么确定。虚拟形象能否增强——甚至取代——记者?首席执行官们是否更愿意与AI虚拟形象的记者交谈,而不是真人记者?
我喜欢我职业生涯中与人联系、撰写故事以及研究新话题的部分。但新闻业最核心的部分是信任。这似乎永远无法外包给AI。
除了新闻业之外,我确信“克隆自己”这一概念可能颇具吸引力。再也不用在假期或度假后补做工作,因为你的一个分身可以随时随地待命,回答问题。
我们还得看看虚拟形象在美国企业界的使用将如何演变。但现在我自己拥有了一个,心情颇为复杂。在度过最初的猎奇阶段后,我在它停止说话后仔细端详着它,并等待着——等什么呢,我也不确定。也许我在等它眨眼。或者等它说些新话,或者微笑,或者只是让我知道它“存在”。
因为我的数字分身是确定性的,它们永远不会那样做。但我能想象,如果有人面对一个非确定性(即由可以自由发表高论的聊天机器人驱动)的分身,很容易陷入轻微的AI精神错乱状态。
我曾对一位投资者说过,无论数字分身的未来如何,我预测我们这一代Z世代可能很难适应它们。它们感觉太像科幻了:电影里看到的一切如今都成了现实。但我得说,我发现虚拟形象比人形机器人没那么让人不适。至少对于虚拟形象来说,如果情况变得诡异,我总是可以下线。
不过在那之前,这是我的个人虚拟形象,为您梳理本周我们网站上的所有头条新闻!
When Alexandru Voica, head of corporate affairs at the video-generation startup Synthesia, sent me a link this summer to the newest addition to their PR team, I was surprised. It was an interactive virtual avatar of him, trained to answer common press questions about Synthesia, like what it does and how it works. The day before, I was on a panel where PR people asked me if I minded pitches that used AI-generated text. But Alexandru’s avatar was beyond that — this seemed to me the final boss of using AI in PR.
In September, Synthesia invited me to its new office space in New York. Originally based in the U.K., Synthesia is a hot digital avatar startup alongside others like D-ID, HeyGen, and Colossyan. It hit a $4 billion valuation earlier this year and said last year it had crossed $100 million in ARR .
Synthesia lets enterprises build interactive training videos with AI avatars and recently launched a product called Roleplay Sessions that lets employees practice, for example, sales pitches with an interactive AI avatar that responds and scores their responses.
When I went to the new office opening and they asked me if I would like my own AI avatar, I didn’t even hesitate to say yes. Of course I would like my own digital twin. My outfit was cute that day, and my hair was in place.
Until I met my digital twin, I had been indifferent toward avatars, but I felt they would inevitably become part of everyday online life. I heard of people on Instagram creating them in their likenesses to help them make social content. I find it all to be very interesting, and it’s perhaps why I have no qualms now presenting to you my digital twin. This is the first time Synethsia has made a digital avatar for a journalist (or for anyone, period, outside of Voica). It is trained on my story about why venture-backed startups commit more fraud than non-VC-backed startups and will only answer questions about that story.
Below, simply press “start in a new window” to begin.
You can ask it questions like:
Why did I decide to write this story?
What is the research paper about?
What did the researchers find?
To make this, I entered a mini film studio nestled inside Synthesia’s office where they took numerous photos of me and captured a two-minute recording of my voice. I had to consent to these avatars being made and, well, digital Dom was born. They created a personal avatar for me (one that just reads whatever script I give it) — with and without glasses — and they made two interactive avatars for me (ones that can talk back and listen to me), also with and without glasses.
We picked an article to train the interactive avatar on, and then one of the teams built my interactive avatar , which is powered by a combination of voice-to-text, video, language, and text-to-voice models. My avatar’s tech stack includes Synthesia’s own video and voice models, although the company also allows customers to choose alternatives from other labs like Cartesia, ElevenLabs, Google, or OpenAI. Enterprises can also choose to host their avatars on whatever cloud they want or pay Synthesia to host them.
The voice-to-text model turns what people say into text, the agentic language model makes sense of text and can take actions based on it, the text-to-voice model turns a response into audio, and finally a video model (built by Synthesia) animates the avatar as it talks.
Overall, Synthesia builds three types of products — a video-creation and distribution platform with classic avatars, where someone types in a script and the avatar repeats it; an agentic platform called Sessions where people can interact with the avatars in surveys or roleplay; and an API platform where people can take Synthesia video and voice models and combine them with other tech services to build interactive avatars or other types of products.
It took the Synthesia team a couple of days to make my avatars. I played first with the personal ones and typed a fairly generic script to see how my AI voice sounded. I had it talk about how fall has arrived in New York, my favorite time of year. The voice was fairly accurate, and I was glad it didn’t pick up any of the hoarseness I had when I recorded my audio sample.
I showed it to some non-tech friends who found it both interesting and creepy.
Then I showed them my interactive one, which is deterministic, meaning it will only say what it was trained to respond to. In this case, that was my venture fraud story. I asked it questions like where I was before TechCrunch and what part of New York I lived in, but each time it directed me back to the story.
My friends didn’t think the voice sounded much like mine and thought the likeness was not as good as the personal avatar, but nonetheless it was close enough to be somewhat creepy. My mom called it “amazing,” which is pretty high praise from her. She and my father kept trying to ask it questions “only they would know” about me — but the model didn’t answer, redirecting them each time to the venture fraud story.
“I don’t remember giving birth to two of you,” she joked after testing the model.
This experience has made me think about what the future of journalism could be. Would people be OK with turning on the news and it being presented by an avatar? One investor told me no immediately. Certainly there is a lot of pushback today on AI slop that has infiltrated social media and other news-sharing platforms. But others I asked weren’t so sure. Could avatars augment — or even replace — journalists? Would CEOs want to talk to an AI avatar of a journalist rather than a human one?
What I like about my career is connecting with people, writing stories, and researching new topics. But the biggest part of journalism is trust. That doesn’t seem like it could ever be outsourced to an AI.
Outside of journalism, I’m confident the idea of cloning yourself could be appealing. No more catching up on work after a holiday or vacation, because a version of you can always be around, answering questions.
We’ll have to see how avatar usage evolves in corporate America. But now that I have my own, I have mixed feelings. After I got over the initial novelty of it, I examined my avatar closely after it stopped talking and waited — for what, I’m not sure. Maybe I’m waiting for it to blink. Or to say something new, or to smile, or to just let me know that it knows.
Because my digital twins are deterministic, they will never do that. But I can see how easy it would be for someone to fall into a tinge of AI psychosis with one that was nondeterministic, meaning powered by a chatbot that was free to pontificate.
I told one investor that, whatever the future of digital twins is, I predict that my generation, Gen Z, likely won’t get used to them. They feel sci-fi: Everything we’ve ever seen in a movie is here now. But I will say, I find digital avatars less jarring than humanoids . At least with an avatar, if stuff gets weird, I can always log off.
Until then, however, here is my personal avatar, giving you a rundown of all the top stories on our site this week!
首次收录 · 2026-09-27 · 12.55 分